SIGN IN SIGN UP

fix(onnx-export): don't freeze pylate's per-call truncation into tokenizer.json

`tokenizer.backend_tokenizer.save()` serializes whatever truncation/padding
state pylate happened to have configured. pylate sets these per call
(query_length for queries, document_length for documents), so the exported
tokenizer.json froze a query-shaped limit.

Exporting lightonai/Reason-ModernColBERT produced truncation.max_length=127
and padding=Fixed(127), which silently clipped every document to 127 tokens
instead of document_length=8192 — a 375-token document came back as 113
embeddings. The Rust reader cannot compensate: it applies query_length /
document_length itself and builds its own padding, but tokenizer.encode_batch
truncates first, so a baked-in limit wins.

Clear both before saving. Verified against pylate afterwards: token counts
match exactly and embeddings agree to 1.0e-06 (worst per-token cosine
0.99999991, identical MaxSim rankings).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
R
raphaelsty committed
0f82c780862446922f96a41dbfc3196604264ddd
Parent: 4ff801e