SIGN IN SIGN UP

Add Gemma 12b Unified Support (#334)

* Keep model-prefixed Mistral3 vision weights

* Sync Qwen3.5 quantization config keys

* Accept deserialized Qwen nested configs

* Remap Qwen3.5 vision weight keys before filtering

* Route Qwen3.5 target verify attention to VLM

* Materialize batch state before cache mutation

* Handle Qwen3.5 attention position embeddings

* Slice Gemma4 token type prompt kwargs

* Remove legacy vision kit

* Remove legacy vision model kit

* Remove legacy vision add-on tests

* Port vision feature memoizer to batched vision

* Align vision feature cache with upstream

* Load VLM image processor before processor

* Handle Gemma4 unified visual APC

* Simplify Gemma4 unified visual prefill

* Reuse Gemma4 prefix before new images

* Force safe mlx-vlm model loading

* Restore Qwen decode fast path

* Remove legacy Qwen image parity test

* Use loaded processors in VLM parity tests

* Work around MLX threaded compile cache cleanup

* Fix VLM cache restore CI failures

* Fix Gemma4 cache follow-up test prompt

* Hardcode Gemma4 cache test prompt

* Restore prompt cache save priority

* Update generated requirements

* Handle Qwen left-padded decode mask

* Handle Qwen left-padded text decode

* Limit Qwen left-padded positions to decode

* Add Gemma4 12B VLM parity coverage

* Preserve VLM backend errors for active requests
N
Neil Mehta committed
9445b319d62ebc1f377fd5db8f156d423369abe2
Parent: e47768b
Committed by GitHub <noreply@github.com> on 6/11/2026, 7:02:42 PM