COMMITS
October 1, 2026
J
llama: refer to segment documentation [no ci] (#29074)
Johannes Gäßler committed
A
common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (#29816)
Adrien Gallouët committed
X
skill: note about model-specific CLI arguments + testings (#29808)
Xuan-Son Nguyen committed
P
ci: fix Fusion / metal by updating the qwen4exp baseline (#29812)
Pascal committed
G
llama : clamp kpool re-pool bound to existing pools (#29805)
Georgi Gerganov committed
Y
hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)
Yiwei Shao committed
S
server: return HTTP 400 for invalid embedding requests (#29060)
Sam Malayek committed
Y
convert : write Gemma embedding scale for DFlash drafts (#29802)
Yu Chengye committed
T
cuda : route sm70 to the Turing MMVQ nwarps table (#29753)
Theera K. committed
M
metal : release temporary private transfer buffers (#29777)
Mike van Lammeren committed
M
webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (#29358)
Masashi Yoshimura committed
G
llama : fix invalid assert in recurrent memory (#29799)
Georgi Gerganov committed
Y
CUDA: Handle compute type for NVFP4 on cublass path (#29173)
ynankani committed
A
Qwen4Exp: add MTP (#29761)
Aman Gupta committed
A
llama: fix qwen4exp (#29751)
Aman Gupta committed
O
CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
Oliver Simons committed
X
mtmd: cap max_image to ubatch for non_causal models (#29773)
Xuan-Son Nguyen committed
G
meta: clear inactive AllReduce shards with FILL, not SCALE (#29793)
Georgi Gerganov committed
A
jinja : skip copying loop scope unless a loop filter needs it (#29776)
a u s t i n committed
P
llama-mmap : avoid a second full-size copy of each tensor with direct-io (#29749)
Pranesh Gonegandla committed
U
M
hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (#29785)
Max Krasnyansky committed
P
BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640)
Pradeep Rao committed
E
common : add LLM-jp-4.1 Harmony dialect handler (#29681)
e-mon committed
T
vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734)
Toki Nasin committed
K
llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
Kushal Garg committed
G
metal : use bf16 math for mxfp4 mul-mat (#29770)
Georgi Gerganov committed
M
llama-bench : fix docs (#29464)
Marlon Paz committed
C
docs : refresh CPU ops support matrix (#29666)
CaramelizedCUDA committed