SIGN IN SIGN UP

Add disk-based caching and continuous batching for VLMs (#326)

* Add batched vision backend

* Add VLM prompt spill cache

* Split VLM safetensor spool

* Split VLM prompt cache payload

* Simplify batched vision scheduler

* Split VLM prompt cache planning

* Add span-aware VLM prompt spill cache

* Clarify VLM prefix cache chunk naming

* Add VLM hot prompt cache saves

* Refine VLM spill cache disk budget

* Refine VLM spill cache eviction

* Align VLM spill cache with physical records

* Refine VLM cache save points

* Clarify VLM cache save points

* Refactor batched vision cache modules

* Simplify batched vision cache store

* Simplify batched vision generator

* Add batched VLM logits processors

* Improve batched vision generation

* Update batched vision tests

* Refine batched vision prompt cache restore

* Add batched vision prompt cache tests

* Add batched vision cache tests

* Add batched vision prompt input tests

* Simplify generate model kit typing

* Stop tracking batched vision smoke script

* Tighten batched vision cache lifecycle

* Simplify batched vision cache chunking

* Improve Qwen prompt cache checkpointing

* Guard rotating prompt cache saves

* Preserve Gemma batched vision masks

* Avoid default VLM top logprobs work

* Fix cached VLM prompt kwargs restore

* Validate KV prompt cache record coverage

* Group VLM generation row state

* Tighten batched vision response metadata

* Tighten batched VLM request handling

* Optimize single-row VLM batching

* Sync scalar Qwen RoPE state

* Fix wrapped rotating cache records

* Preserve budget-fitting prompt cache restores

* Evict failed prompt cache records

* Avoid scalar cache response aliasing

* Tighten batched vision failure boundaries

* Make VLM cache snapshots best effort

* Update VLM prefill alignment test

* Restore semantic cache test assertions

* Document Qwen MRoPE workaround

* Clarify VLM prompt cache restore names

* Preserve restore chains after partial cache saves

* Keep hot cache fallback after disk restore miss

* Tighten VLM prompt cache restore handoffs

* Fix VLM test compatibility

* Fix VLM cache budget failure and trim tests

* Disable disk cache on budget estimate failure

* Stabilize Gemma3n cache prompt test

* Stabilize Gemma3n cache reuse test

* Restore Gemma3n cache semantic assertions

* Stabilize Gemma3n long prompt cache test

* Stabilize Gemma3n long prompt author test

* Stabilize Gemma long prompt cache tests

* Fix Gemma3 batched vision mask handling

* Gate Gemma3 attention mask drop

* Route VLM attention masks in batched prefill

* Adjust VLM prompt cache disk budget

* Fix Granite4 batched vision DeepStack slicing

* Speed up batched vision repeat penalty

* Fix worker-safe sampling

* Fix thread-local prompt cache tokens

* Fix Qwen3.5 text fast path in batched vision

* Fix batched vision structured processors

* Avoid rebuilding VLM detokenizer per request

* Run sequential model kit on owner thread

* Log lifetime prompt cache evictions

* Fix batched VLM prompt progress reporting

* Fix VLM model kit setup edge cases

* Normalize VLM concurrency limit

* Add Qwen3.5 VLM restore parity test

* Add Qwen3.5 VLM generation trace parity test

* Add Qwen3.5 continuous batching parity test

* Add batched VLM parity tests

* Relax VLM concurrency logprob parity
N
Neil Mehta committed
e2f0e8933047eb07618e5af59dd8bb8873e2d887
Parent: aea0911
Committed by GitHub <noreply@github.com> on 5/28/2026, 7:00:29 PM