Add Gemma 12b Unified Support (#334)
* Keep model-prefixed Mistral3 vision weights * Sync Qwen3.5 quantization config keys * Accept deserialized Qwen nested configs * Remap Qwen3.5 vision weight keys before filtering * Route Qwen3.5 target verify attention to VLM * Materialize batch state before cache mutation * Handle Qwen3.5 attention position embeddings * Slice Gemma4 token type prompt kwargs * Remove legacy vision kit * Remove legacy vision model kit * Remove legacy vision add-on tests * Port vision feature memoizer to batched vision * Align vision feature cache with upstream * Load VLM image processor before processor * Handle Gemma4 unified visual APC * Simplify Gemma4 unified visual prefill * Reuse Gemma4 prefix before new images * Force safe mlx-vlm model loading * Restore Qwen decode fast path * Remove legacy Qwen image parity test * Use loaded processors in VLM parity tests * Work around MLX threaded compile cache cleanup * Fix VLM cache restore CI failures * Fix Gemma4 cache follow-up test prompt * Hardcode Gemma4 cache test prompt * Restore prompt cache save priority * Update generated requirements * Handle Qwen left-padded decode mask * Handle Qwen left-padded text decode * Limit Qwen left-padded positions to decode * Add Gemma4 12B VLM parity coverage * Preserve VLM backend errors for active requests
N
Neil Mehta committed
9445b319d62ebc1f377fd5db8f156d423369abe2
Parent: e47768b
Committed by GitHub <noreply@github.com>
on 6/11/2026, 7:02:42 PM