Enable disk-backed KV cache for text-only models (#352)
* Route batched text models through mlx-vlm * Correct GPT-OSS context auto-fit
N
Neil Mehta committed
3ff8b81ca8159c4b2bdb87916649ab20b8c6f25f
Parent: ec2f585
Committed by GitHub <noreply@github.com>
on 7/24/2026, 4:47:16 PM