COMMITS
October 1, 2026
S
[ROCm][CI] Drop two no-GPU AMD mirrors from the CPU test areas (#59593)
stefankoncarevic committed
D
H
[MyPy] Fix mypy errors in `vllm/model_executor/models/[kK]*` (#59402)
Harry Mellor committed
J
[ROCm][Triton] Migrate Kimi-K3 kernels from make_block_ptr to tensor … (#58769)
Jaden Mathias committed
C
[KV Connector][NIXL] Coalesce host-buffer KV copies across cache groups (#54483)
Chanbin Lim committed
R
M
[CI/Build] Add agents auto-label rule (#59641)
Misha Goin committed
M
[Bugfix][HiSparse] Fix a chunked-prefill preemption livelock (#59494)
Matthew Bonanni committed
F
M
[Agents] Expose PR checklist skill to Claude (#59638)
Misha Goin committed
M
[Bugfix][ROCm] Use a zero default for masked scales in the MXFP8 GEMM (#59454)
Mehmet Cagri committed
S
A
[Bugfix] Tie lm_head.weight for Nemotron Parse when checkpoint omits it (#53020)
aniskumar-nv committed
S
[ROCm][CI] Drop four no-GPU CPU groups from the legacy AMD pipeline (#59595)
stefankoncarevic committed
A
M
I
[Rust][Benchmark] Warn when temperature is left to the server default (#59247)
Idder Ghanbaja committed
N
[CI/Build] Add mooncake auto-label rule and assign topic owners (#59192)
Nicolò Lucchesi committed
J
[Core] Remove deprecated mamba_cache_mode "all" (#58997)
Jiangyun Zhu committed
Y
S
[Docs] Remove references to removed env vars (#59530)
Sundri Lai committed
B
[Rust Frontend] Add `hf` parser for response templates (#59005)
Bugen Zhao committed
M
[Core][Logging] Add built-in JSON formatter (#58739)
Mark McLoughlin committed
J
[Bugfix][Spec Decode] Per-module LM heads for multi-layer MTP on Model Runner V2 (#58921)
Jiangyun Zhu committed
R
[Model Runner V2] Support randomized dummy inputs (#58411)
Robert Shaw committed
G
[CPU][Zen] Pass f32 weight scales to the zentorch INT8 MoE (#59434)
Ganesh R committed
J
[Test] Make Anthropic messages test compatible with SDK 1.x via extra_body (#57780)
jiangyunfan1 committed
G
[CPU][Zen] Add DA8W4 (W4A8) int4 support for dense and MoE layers (#54024)
Ganesh R committed