model : re-enable -sm tensor for qwen4exp (#28569)
#27941 disabled -sm tensor for qwen4exp because test-llama-archs asserted on the Meta device once the fixture carried a PLE layer: GGML_ASSERT(ggml_backend_buffer_is_meta(tensor->buffer)) at ggml-backend-meta.cpp:476. With host-resident embeddings the PLE gather is a CPU node and hc_init (the REPEAT that fans the embedding out to the hc streams) was first reached through layer 0's PLE path, after that gather. ggml_backend_sched_split_graph pass 2 expands a device assignment upwards only until it meets a CPU node, so the REPEAT stayed on the CPU and the later reshape of hc_init inside the meta split viewed a host-resident node. Expanding hc_init right after it is built puts the REPEAT directly before the first device node, where pass 2 assigns it; the embedding reshape stays in the CPU split and is copied in as a split input, as in deepseek4.
K
Kevin Hopper committed
10f340d1a2b07d2b222b76df9eab70764e63b01b
Parent: 0c1e570
Committed by GitHub <noreply@github.com>
on 10/1/2026, 5:16:49 AM