Add disk-based caching and continuous batching for VLMs (#326)
* Add batched vision backend * Add VLM prompt spill cache * Split VLM safetensor spool * Split VLM prompt cache payload * Simplify batched vision scheduler * Split VLM prompt cache planning * Add span-aware VLM prompt spill cache * Clarify VLM prefix cache chunk naming * Add VLM hot prompt cache saves * Refine VLM spill cache disk budget * Refine VLM spill cache eviction * Align VLM spill cache with physical records * Refine VLM cache save points * Clarify VLM cache save points * Refactor batched vision cache modules * Simplify batched vision cache store * Simplify batched vision generator * Add batched VLM logits processors * Improve batched vision generation * Update batched vision tests * Refine batched vision prompt cache restore * Add batched vision prompt cache tests * Add batched vision cache tests * Add batched vision prompt input tests * Simplify generate model kit typing * Stop tracking batched vision smoke script * Tighten batched vision cache lifecycle * Simplify batched vision cache chunking * Improve Qwen prompt cache checkpointing * Guard rotating prompt cache saves * Preserve Gemma batched vision masks * Avoid default VLM top logprobs work * Fix cached VLM prompt kwargs restore * Validate KV prompt cache record coverage * Group VLM generation row state * Tighten batched vision response metadata * Tighten batched VLM request handling * Optimize single-row VLM batching * Sync scalar Qwen RoPE state * Fix wrapped rotating cache records * Preserve budget-fitting prompt cache restores * Evict failed prompt cache records * Avoid scalar cache response aliasing * Tighten batched vision failure boundaries * Make VLM cache snapshots best effort * Update VLM prefill alignment test * Restore semantic cache test assertions * Document Qwen MRoPE workaround * Clarify VLM prompt cache restore names * Preserve restore chains after partial cache saves * Keep hot cache fallback after disk restore miss * Tighten VLM prompt cache restore handoffs * Fix VLM test compatibility * Fix VLM cache budget failure and trim tests * Disable disk cache on budget estimate failure * Stabilize Gemma3n cache prompt test * Stabilize Gemma3n cache reuse test * Restore Gemma3n cache semantic assertions * Stabilize Gemma3n long prompt cache test * Stabilize Gemma3n long prompt author test * Stabilize Gemma long prompt cache tests * Fix Gemma3 batched vision mask handling * Gate Gemma3 attention mask drop * Route VLM attention masks in batched prefill * Adjust VLM prompt cache disk budget * Fix Granite4 batched vision DeepStack slicing * Speed up batched vision repeat penalty * Fix worker-safe sampling * Fix thread-local prompt cache tokens * Fix Qwen3.5 text fast path in batched vision * Fix batched vision structured processors * Avoid rebuilding VLM detokenizer per request * Run sequential model kit on owner thread * Log lifetime prompt cache evictions * Fix batched VLM prompt progress reporting * Fix VLM model kit setup edge cases * Normalize VLM concurrency limit * Add Qwen3.5 VLM restore parity test * Add Qwen3.5 VLM generation trace parity test * Add Qwen3.5 continuous batching parity test * Add batched VLM parity tests * Relax VLM concurrency logprob parity
N
Neil Mehta committed
e2f0e8933047eb07618e5af59dd8bb8873e2d887
Parent: aea0911
Committed by GitHub <noreply@github.com>
on 5/28/2026, 7:00:29 PM