llama: llama_prefetch_rows (#29599)
* llama: llama_prefetch_rows * llama: support row prefetch on Windows Apply the Windows port contributed by @praneshgo unchanged. Source: https://github.com/ggml-org/llama.cpp/pull/29599#issuecomment-5887721014 * avoid exposing llama-mmap in model code, route via llama-impl * add windows check, only prefetch in lazy mode * cont : clean-up * cont : fix build * cont : clarify padding token for gemma4 --------- Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
A
Aman Gupta committed
185103dcf53222165ecd15ce8e406606e25bd091
Parent: 2090f60
Committed by GitHub <noreply@github.com>
on 9/30/2026, 12:27:18 PM