SIGN IN SIGN UP

Enable disk-backed KV cache for text-only models (#352)

* Route batched text models through mlx-vlm

* Correct GPT-OSS context auto-fit
N
Neil Mehta committed
3ff8b81ca8159c4b2bdb87916649ab20b8c6f25f
Parent: ec2f585
Committed by GitHub <noreply@github.com> on 7/24/2026, 4:47:16 PM