llama : clamp kpool re-pool bound to existing pools (#29805)
* tests : simplify function signature * llama : clamp kpool re-pool bound to existing pools The n_tokens/kpool + n_seqs_unq bound on n_new_g overshoots when a batch fills the whole cache: n_ctx tokens complete exactly n_ctx/kpool pools, so the +1 pads new_pool_idxs/new_pool_rep one entry past n_pool_real. Graph reserve only covers n_pool_real entries, so the first full-context decode builds bigger tensors than reserved and ggml-alloc demands a graph reallocation (abort under GGML_SCHED_DEBUG_REALLOC=1). Clamp the bound to n_pool_real: a ubatch can never mark more pools than the cache holds, and reserve's n_pool_max already covers that. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * cont : cap to n_pool_max
G
Georgi Gerganov committed
81e39ad34368329b5db77620721510fd91107e80
Parent: dcd387a
Committed by GitHub <noreply@github.com>
on 10/1/2026, 4:55:34 PM