SIGN IN SIGN UP

perf(index): throttle index saves and report each leg's save state (#1415)

* perf(index): throttle index saves and report each leg's save state

Every observation, compression and memory save scheduled a full index
save 5 s later, and each save re-serializes the whole BM25 and vector
index and rewrites every shard. On a busy store that is a full index
rewrite almost continuously.

Saves now run at most once per AGENTMEMORY_INDEX_SAVE_INTERVAL_MS
(default 10 minutes), measured from the start of the previous save.
Explicit saves (deletes, shutdown) still run immediately. Only one save
runs at a time; saves requested while one is running collapse into a
single follow-up save.

The BM25 and vector saves now fail independently. Before, a BM25
failure skipped the vector save in the same run and both were reported
as one warning. Each leg records when it last saved, its last error,
since when it has unsaved changes, and its serialized size.

/agentmemory/status reports a failing leg as an error, unsaved changes
older than twice the interval as a warning, and a vector index above
80% of the Node.js string limit (536,870,888 characters) as a warning,
since past that limit the index can no longer be saved or loaded. The
status page shows each leg's save state and the vector index size.

* chore: credit co-authors

The per-leg save failure handling and the single save queue build on
the fix in #1256. The problem was measured and reported in discussion
#1358.

Co-authored-by: Omar Gerardo <omargrard@gmail.com>
Co-authored-by: Vladimir <7998636+MarvinFS@users.noreply.github.com>

* fix(index-persistence): snapshot dirtyEpoch per leg before serializing

runSave() captured one dirtyEpoch before the BM25 leg's write and reused
it for the vector leg's dirty check. When scheduleSave() ran while the
BM25 write was in flight, the vector leg would serialize the current
(already up to date) state but still compare against the stale epoch,
leaving vector.dirtySince set even though the just-written snapshot
already reflected the change. status() would report the vector leg as
dirty until an unrelated follow-up save cleared it.

saveLeg() now snapshots dirtyEpoch itself, immediately before invoking
the leg's own serialize+write callback, so each leg judges its dirty
state against the epoch as of its own write rather than one captured
before an earlier leg ran.

Added a regression test that gates the BM25 manifest write, triggers a
scheduleSave() while it is held, and asserts the vector leg's
dirtySince clears once its own write completes.

---------

Co-authored-by: Omar Gerardo <omargrard@gmail.com>
Co-authored-by: Vladimir <7998636+MarvinFS@users.noreply.github.com>
R
Rohit Ghumare committed
cfc193408c09cd7087c7c03cd682c3d92a187840
Parent: 39658a9
Committed by GitHub <noreply@github.com> on 9/26/2026, 8:27:20 AM