SIGN IN SIGN UP

regression: source_model for 24 entries — the rebake refused every one without it

The chr1s4 rebake run (crispasr-regression-rebake v1) did not fail obscurely; it
produced a verdict. 28 of the 45 manifest entries came back

    ✗ <name>  0.0s  manifest entry has no `source_model`; add it before re-baking

so the bake was blocked on missing metadata, not on infrastructure. Two entries
(moonshine-tiny, moonshine-base) baked fine, which is what makes the diagnosis
trustworthy: the harness works, the data was incomplete.

Each id is sourced, not guessed:
 - 19 from this model's own published GGUF card in hf_readmes/, matched by exact
   filename against the entry's gguf repo;
 - 5 with no card, each HEADed against the HF API and only written on a 200.

A FUZZY MATCH PRODUCED A WRONG ID AND IT WAS CAUGHT BY MAKING THE TOOL SAY WHERE
IT LOOKED. My first pass matched cards by prefix and mapped granite-speech-4.1-2b
to ibm-granite/granite-speech-3.2-8b — a different model, and one that would have
baked a reference dump silently attributed to the wrong upstream. Printing which
card each id came from made it obvious in one line. The corrected id
(ibm-granite/granite-speech-4.1-2b) was then confirmed 200 by the API.

Four remain unresolved and are deliberately left blank: firered-lid,
silero-lid-lang95, moss-audio-4b-instruct, kugelaudio-0-open. My candidate ids
for those returned 401/307, so writing them would have put unverified metadata
in the manifest that looks exactly like verified metadata. They stay empty and
the rebake will keep naming them, which is the correct behaviour.

Also outstanding, separately: 4 entries failed `dump_reference exit=1`
(parakeet-tdt-0.6b-ja, canary-1b-v2, firered-asr2-aed, glm-asr-nano) — they HAVE
a source_model, so that is a different fault and needs the reference-module logs.
C
crispasr integration committed
f8711215efdde447650a337fd05fc21feabb2417
Parent: 915c553