SIGN IN SIGN UP

Add experimental server with chat completions endpoint (#353)

* Add MLX server

* Simplify MLX server request handling

* Clarify inline image schema names

* Keep chat template rendering model-owned

* Stream HTTP generation on request threads

* Admit batched requests between generation steps

* Handle structured output and assistant prefills

* Validate server request boundaries

* Validate generation requests and prompt usage

* Harden server request and usage accounting

* Format server startup test

* Normalize text messages and close error responses

* Bound stream writes and correct prefill progress

* Harden chat request validation

* Harden server request handling

* Reject decompression bomb images

* Validate complete image requests

* Prevent batched request starvation
N
Neil Mehta committed
35a325f1fa1b3814737d7dbf7e3d7bbcf19f38e3
Parent: 3abfcbf
Committed by GitHub <noreply@github.com> on 8/10/2026, 2:55:13 PM