SIGN IN SIGN UP

llama: llama_prefetch_rows (#29599)

* llama: llama_prefetch_rows

* llama: support row prefetch on Windows

Apply the Windows port contributed by @praneshgo unchanged.

Source: https://github.com/ggml-org/llama.cpp/pull/29599#issuecomment-5887721014

* avoid exposing llama-mmap in model code, route via llama-impl

* add windows check, only prefetch in lazy mode

* cont : clean-up

* cont : fix build

* cont : clarify padding token for gemma4

---------

Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
A
Aman Gupta committed
185103dcf53222165ecd15ce8e406606e25bd091
Parent: 2090f60
Committed by GitHub <noreply@github.com> on 9/30/2026, 12:27:18 PM