perf: cut bounded-string token-mask fills by classifying the repetition body with its parent (#903)
**Incremental review:** [Files
changed](https://github.com/mlc-ai/xgrammar/pull/903/files) against
`main`.
# perf: cut bounded-string token-mask fills by classifying the
repetition body with its parent
Related to #852.
A `maxLength`-only JSON string compiles into a repetition of the string
body, and the compiled token
mask is shared by every repeat count. Two things followed from
classifying a body state on its own:
* every token that leaves the repetition (the closing quote, an escape)
came out **uncertain** and was
replayed through the Earley parser on every fill;
* the unrolled `{0,127}` head plus the `{128,N}` tail with its lookahead
kept about **130 leaf
states** alive while matching inside the string.
This branch changes only the body of a counted repetition of a
**negative** character class with
`lower == 0` - the shape a `maxLength`-only string produces:
* the repetition is compiled into one counted repeat edge instead of the
unrolled head and the repeat
tail with the lookahead;
* the body state is classified with the repetition parent **seeding the
parse stack** while the
first-character mask, the speculative path and the per-token simulation
still come from the body
rule, so tokens that stay inside the class are accepted from a
precomputed per-codepoint-count
bitset and tokens that leave are decided as before.
## Measurements
151669-id tokenizer, the reported schema, 50-token in-string replay,
`fill_next_token_bitmask` per step:
| build | fill/step median | compiled masks | leaf states | per-state
uncertain |
|---|---|---|---|---|
| base | 0.3904 ms | 7.40 MB | 130 | 896 |
| this branch | **0.0614 ms** | 1.79 MB | **2** | 742 |
| unbounded control (`C*`) | 0.0441 ms | 27 KB | 1 | 264 |
## Verification
| check | command | result |
|---|---|---|
| exhaustive mask <-> accept oracle | `oracle_852.py`, bounds
8/129/512/100000, every token id replayed against a fresh matcher | **0
mismatches**, allowed-token set identical to base |
| Python suite | `pytest tests/python -q` | **3918 passed, 815 skipped,
0 failed** |
| bounded-string tests | `pytest
tests/python/test_grammar_matcher_bounded_string.py -q` | 25 passed (on
base and on this branch) |
| C++ suite | `ctest` (73 registered tests) | **73 / 73 passed** |
| reproduction/benchmark harness | `repro_852.py`, `diag_852.py`,
`probe_852.py`, `route_852.py`, `split_852.py`, `budget_852.py` |
retained in the artifact tree |
## Limits
* Measured on CPU mask filling only: no vLLM / MLC-LLM end-to-end run,
no production tokenizer, no
WASM build. The C++ suite ran with `XGRAMMAR_BUILD_CXX_TESTS=ON` (the
repository's
`cmake/config.cmake` forces it OFF by default and shadows the `-D` cache
variable).
* `minLength` with `maxLength` and every positive-character-class
repetition keep the previous
expansion (`lower == 0` and a negative class are required), so their
cost is unchanged.
* Alternatives that were measured and rejected are documented in the
artifact tree: compiling the
bound as one counted repeat without the parent seeding (117x slower), a
run-time budget fast path
without it (2.5x slower), and the character-budget representation (17x
slower).
---------
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: yuchuan <yuchuan.7streams@gmail.com>
Co-authored-by: yuchuan <yuchuan.7streams@gmail.com> 0
0z5a committed
a2faca36fea00963b745f156adc0ddbcbf2a413e
Parent: eee144a
Committed by GitHub <noreply@github.com>
on 9/28/2026, 7:01:11 PM