COMMITS
September 2, 2026
J
J
fix: preserve Qwen4 PLE mmap admission fallback
jundot committed
P
fix(engine_pool): keep the Qwen4 PLE table resident across a model swap (#3299)
pedroberaldo87 committed
J
refactor: simplify Qwen4 prefill memory accounting (#3350)
jundot committed
J
fix(qwen4): reject corrupt QSA prefix states (#3294)
jundot committed
J
fix(qwen4): isolate gathered prefill memory history (#3350)
jundot committed
J
fix(qwen4_exp): price gathered QSA prefill memory (#3351)
jundot committed
J
fix(qwen4): preserve gathered prefill across mRoPE rebinds (#3355)
Jonathan Spangler committed
J
fix(qwen4): harden QSA batch cache joins (#3294)
jundot committed
A
fix(qwen4): handle mixed-rank QSA cache joins (#3369)
astro-15-alive committed
September 1, 2026
August 31, 2026
J
fix: validate boundary snapshots before materialization
jundot committed
J
build: upgrade MLX to 0.32.2 with runtime compatibility (#3332)
Jun Kim committed
August 29, 2026
G
formula: bump to 0.6.4
github-actions[bot] committed
J
chore: bump version to 0.6.4
jundot committed
J
fix(qwen4): gate wide prefill on sparse native kernel (#3244)
jundot committed
J
perf(qwen4): flatten exact QSA prefill and accelerate Lightning MTP (#3244)
Jonathan Spangler committed
J
fix(cache): initialize boundary diagnostics in snapshot test
jundot committed
J
fix(admin): validate Lightning MTP draft depth (#3280)
jundot committed
J
fix(cache): keep boundary diagnostics accurate
jundot committed
C
fix(admin): persist mtp_num_draft_tokens via settings PUT (#3280)
Cold Cloud committed
A
fix: allow TurboQuant with Lightning MTP in macOS app (#3254)
Aka.Fido committed
M
feat(cache): expose hybrid boundary diagnostics (#3249)
MDX Tom committed
R
fix(cache): normalize mixed Qwen4 QSA positions (#3219)
rsnow committed
J
fix(qwen4_exp): bound QSA long-prefill memory (#3283)
jundot committed