fix(cli): fetch a page's assets as the same agent that loaded the page (#3726)
A capture is one session with two halves: Chrome navigates the page with a browser User-Agent, then Node fetches the assets that page referenced. Those halves sent three different identities — "HyperFrames/1.0" from the asset and media downloaders, a bare "Mozilla/5.0" from the stylesheet inliner, and the real Chrome UA from the navigation itself. An origin is free to answer those differently, and anti-bot edges do. Capturing one large site, GET /favicon.svg answers 403 text/html to "HyperFrames/1.0" and 200 image/svg+xml to the UA the very same capture had just navigated with. The favicon ranker had already picked that SVG as the best declared icon; the 403 discarded it and the downloader fell through to the next candidate, so the icon written to assets/ was chosen by the CDN's bot rules rather than by the ranker. The capture reported it as one "unavailable" drop and carried on. Hoist the navigation UA into CAPTURE_USER_AGENT and use it for every out-of-band fetch the capture makes: favicons, images, og:image, fonts, stylesheets, Lottie JSON and videos. One constant is what stops the two halves drifting apart again. Verified end to end against that site: before, assets/favicon.png (the apple-touch icon) plus one unavailable drop; after, assets/favicon.svg, byte identical to the file the site itself serves.
M
Miguel Ángel committed
7a07ea9ac34804e77cd054ca10ea84f3e0a0c774
Parent: be86a1e
Committed by GitHub <noreply@github.com>
on 9/6/2026, 2:20:15 AM