COMMITS
October 3, 2026
T
model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (#29862)
Tarek Dakhran committed
P
qwen4exp : halve the indexer score memory (#29825)
Pascal committed
X
model: add support for clef decision model (text-only) (#29831)
Xuan-Son Nguyen committed
October 2, 2026
A
CUDA: fuse shared experts into MMVQ (#29184)
Aman Gupta committed
S
ci : use t4-medium for cuda jobs (#29842)
Sigbjørn Skjæret committed
P
spec : add probabilistic sampling for simple draft and MTP (#27694)
Pranesh Gonegandla committed
Y
ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (#27663)
Yash Raj Pandey committed
Y
ggml-quants : avoid invalid rounding in qkx3 scale search (#29817)
Yash Raj Pandey committed
Y
ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096)
Yash Raj Pandey committed
X
model: support nimble decision model (#29844)
Xuan-Son Nguyen committed
G
readme : add cmd install commands (#29850)
Georgi Gerganov committed
E
metal : add tensor API flash attention kernel for F16 KV (#29570)
Ethan Guo committed
X
llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (#29818)
Xuan-Son Nguyen committed
A
common : remove fs_open_ifstream() by using u8path() (#29841)
Adrien Gallouët committed
A
llama : silence unused-result warnings (#29839)
Adrien Gallouët committed
A
llama : use GGML_ABORT instead of throw (#29840)
Adrien Gallouët committed
L
opencl: use sigmoid f16 for bf16 (#29787)
lhez committed
C
SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (#29186)
cwriter committed
J
vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (#28531)
Jiwoong Song committed
T
sycl: large register file for D=512 FA vec kernels (#29062)
Titaniumtown committed
Ł
sycl : do not use slow oneDNN reference matmul and fattn (#28985)
Łukasz Ślusarczyk committed
G
qwen4exp : optimize mask constructions (#29824)
Georgi Gerganov committed
G
ggml : add `alloc_buffer_n` to buffer type interface (#23671)
Georgi Gerganov committed
A
ci : fix missing zdnn backend check (#29837)
Aaron Teo committed
R
vulkan: add logging to pipeline compile issues (#29794)
Ruben Ortlam committed
K
pyproject : add linux platform marker to uv torch source (#29177)
Kasimir Tanner committed
K
hexagon: install rebuilt HTP skels (#29828)
kurquhar committed
A
qwen4exp: fix tests (#29819)
Aman Gupta committed
October 1, 2026
J
hexagon: add q2_k and q3_k quant type support (#29717)
Jhen-Jie Hong committed
J
CUDA: fix 2 broken Volta FA cases (#29803)
Johannes Gäßler committed