SIGN IN SIGN UP

ci(pr-proof): decide proof requirement from the diff, not the PR title (#1645)

* ci(pr-proof): decide proof requirement from the diff, not the PR title

The proof gate read only the PR title, which was wrong in both
directions: a workflow- or docs-only change titled `fix(` was forced to
fabricate a RelayFlow case, and a change that genuinely touched the
broker or CLI skipped the gate entirely by being titled `chore(`.

classifyPullRequest now takes the changed-file list and decides from the
diff. The exemption is an allowlist (workflows, docs, .github,
.agentworkforce, scripts, tests, top-level markdown), so an unrecognised
path counts as runtime and still demands a proof — it fails closed. When
the diff cannot be read the classifier falls back to the previous
title-based behaviour rather than silently exempting anything.

No changelog entry: this is repo CI tooling, not a user-facing Relay
surface.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014N3p9VEngj9kLDFFhrzHNd

Session-Id: 246fba6e-6436-46ae-bc50-6bb3cca0d95e

* ci(pr-proof): close four bypasses in the diff-based classifier

Review found four ways the new gate could still be evaded or could block
honest work. All four are fixed and each is covered by a test that fails
when the fix is reverted.

- `scripts/` was exempt wholesale, but publish.yml runs
  scripts/inject-posthog-key.mjs to rewrite the compiled CLI before it
  ships. Narrowed to scripts/pr-proof/ and scripts/evals/, the subtrees
  that only serve CI. Any other script now counts as runtime.

- A rename reports only its destination path, so moving a runtime file
  into docs/ read as a docs-only change. pullRequestFiles now collects
  previous_filename alongside filename.

- classifyPullRequest documented a title-only fallback for an unreadable
  diff, but the fetch threw before classification, so the fallback was
  unreachable and an outage of the files endpoint blocked every PR. The
  fetch is now guarded. An over-large PR is tagged ambiguousScope and
  still refuses to fall back, because scope that big must not be
  resolved by wording.

- The non-functional guard required a feat/fix title, so a `chore(`
  titled runtime change declaring non-functional with a valid case id
  had that contradiction silently coerced to `bugfix`. The diff now
  overrules the title in both directions.

Tests: 83 passing, 6 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014N3p9VEngj9kLDFFhrzHNd

Session-Id: 246fba6e-6436-46ae-bc50-6bb3cca0d95e

* style: auto-format with Prettier

* ci(pr-proof): stop the title-only fallback crashing the required path

Two more from review, both in the fallback added by the previous commit.

The fallback left `changedFiles` as null and then walked straight into
`changedFiles.some(...)`, so a bugfix PR with valid metadata hit a bare
TypeError whenever the files endpoint was down — a crash where there had
previously been a clear API error. A required proof genuinely cannot be
validated without the file list (it is what confirms the declared case is
touched), so that combination now fails closed with a message saying so.
The fallback still stands for the not-required path, where the title is
sufficient and the old behaviour is preserved.

Page 30 returning a full 100 entries means "3,000 so far", not "more
than 3,000", so an exactly-3,000-file PR was rejected as ambiguous
scope. Probing one page past the cap separates the two; the extra
request only happens at exactly 3,000.

Tests: 85 passing, 6 skipped. The null-diff case is driven through
main() with a stubbed fetch rather than an extracted helper, so it
proves the real path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014N3p9VEngj9kLDFFhrzHNd

Session-Id: 246fba6e-6436-46ae-bc50-6bb3cca0d95e

Session-Id: 246fba6e-6436-46ae-bc50-6bb3cca0d95e

* ci(pr-proof): revert the page-31 probe; it accepted a truncated diff

The previous commit added a page-31 probe to tell "exactly 3,000 files"
from "more than 3,000". That reasoning was wrong and I should have
checked it before implementing it.

GitHub documents this endpoint as "Responses include a maximum of 3000
files." The cap means page 31 is empty for a 3,000-file PR AND for a
30,000-file one, so the probe cannot distinguish a complete diff from a
truncated one — it accepted the truncation as complete. A PR large
enough to be capped would then have been classified on a partial file
list, and any runtime file beyond the cap would have gone unseen and
read as non-runtime. That is precisely the bypass this PR exists to
close, reintroduced by the fix for a cosmetic edge case.

Restored: a full 30th page refuses as ambiguous. An exactly-3,000-file
PR is refused too. That is a false positive, but it fails closed and no
such PR has ever existed here.

The pagination test stub now parses the `page` query parameter and
returns distinct filenames per page, so it fails if pagination stalls
rather than passing on identical content.

Tests: 84 passing, 6 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014N3p9VEngj9kLDFFhrzHNd

Session-Id: 246fba6e-6436-46ae-bc50-6bb3cca0d95e

* test(pr-proof): cover successful multi-page file accumulation

Deleting the exactly-3,000-file test removed the only coverage of a
successful multi-page fetch. The replacement only exercised the refusal
path, so nothing verified that entries from page 2 onward are appended
rather than replacing page 1. Had accumulation regressed, runtime files
from earlier pages would have dropped out of the list and the PR could
have read as non-runtime — the same class of bypass the reverted probe
would have opened.

Covers a three-page sub-cap diff and asserts a file from every page.
Mutation-checked by clearing the accumulator each page: the test fails.

Tests: 85 passing, 6 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014N3p9VEngj9kLDFFhrzHNd

Session-Id: 246fba6e-6436-46ae-bc50-6bb3cca0d95e

---------

Co-authored-by: Proactive Runtime Bot <agent@agent-relay.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
K
Khaliq committed
3e363af70feb54daf277fa5cba3d9c548604ba2e
Parent: d3b3f68
Committed by GitHub <noreply@github.com> on 9/2/2026, 7:03:23 PM