SIGN IN SIGN UP

Pin WebGPU KV cache for ImageAudioTextToText models (Gemma3n, Gemma4) (#1757)

ImageAudioTextToText was the only generating model type in
MODEL_SESSION_CONFIG without a cache_sessions declaration, so
constructSessions passed cache_config=false for decoder_model_merged and
getSession never set preferredOutputLocation='gpu-buffer' for the
present.* outputs on WebGPU. Every decode step of
Gemma3nForConditionalGeneration and Gemma4ForConditionalGeneration
therefore round-tripped the full, growing KV cache through the CPU,
while the same weights loaded text-only kept it on-GPU.

Sibling precedent: same file, [MODEL_TYPES.AudioTextToText]
(session_config.js:74) and [MODEL_TYPES.ImageTextToText]
(session_config.js:64) both declare
cache_sessions: { decoder_model_merged: true }; text-only Gemma3n and
Gemma4 already get pinning via MODEL_TYPES.DecoderOnly
(session_config.js:24).

Adds a regression test asserting the ImageAudioTextToText decoder is
cache-pinned, plus an invariant that every model type declaring a
generation_config also declares at least one cache session.
J
Jeremy Schoemaker committed
6fe1ccc2432c4d975a62dc62d8e37cf54a396930
Parent: f09d000
Committed by GitHub <noreply@github.com> on 9/4/2026, 4:10:27 AM