Context length auto fit (#345)
* Add MLX batched VLM context fitter * Document MLX context fit calculation * Clarify context memory budget calculation * Leave unsupported VLM contexts unchanged * Best-effort fit compatible VLM caches * Report fitted VLM context after load * Fit batched VLM context automatically * Leave context unchanged when fitting fails * Skip context fitting for unchunked prefill * Account for per-layer prompt inputs
N
Neil Mehta committed
464fd25daf1a52058e07623485dc0b62dbe8c1f4
Parent: cd68a5f
Committed by GitHub <noreply@github.com>
on 7/14/2026, 6:32:03 PM