fix(onnx-export): don't freeze pylate's per-call truncation into tokenizer.json
`tokenizer.backend_tokenizer.save()` serializes whatever truncation/padding state pylate happened to have configured. pylate sets these per call (query_length for queries, document_length for documents), so the exported tokenizer.json froze a query-shaped limit. Exporting lightonai/Reason-ModernColBERT produced truncation.max_length=127 and padding=Fixed(127), which silently clipped every document to 127 tokens instead of document_length=8192 — a 375-token document came back as 113 embeddings. The Rust reader cannot compensate: it applies query_length / document_length itself and builds its own padding, but tokenizer.encode_batch truncates first, so a baked-in limit wins. Clear both before saving. Verified against pylate afterwards: token counts match exactly and embeddings agree to 1.0e-06 (worst per-token cosine 0.99999991, identical MaxSim rankings). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
R
raphaelsty committed
0f82c780862446922f96a41dbfc3196604264ddd
Parent: 4ff801e