[AMD][DSV4] Add a DP-attention arm and refresh the MI355X SGLang AgentX key (#2800)
* [AMD][DSV4] Add a DP-attention arm and refresh the MI355X SGLang AgentX key
Image: lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260829 -> ...-20260831.
Search space:
- add tp 8, ep 1, dp-attn true on the hicache host KV tier at concurrency
[64, 96, 128]
- drop concurrency 2 and 10 from the TP4 arm, leaving [1, 4, 8]
- drop concurrency 16 from the TP8 hicache arm, leaving [32, 48]; conc 16
stays on the TP8 no-offload arm
Serving, split per-arm so the tensor-parallel arms keep their current settings:
- shared-experts fusion stays enforced without DP and is disabled under DP
attention, matching the DP baselines
- swa-full-tokens-ratio stays 0.10 without DP and is 0.15 under DP attention
- DP attention adds --enable-two-batch-overlap,
--enable-dp-attention-local-control-broadcast, --tokenizer-worker-num equal
to TP, --stream-interval 20 and --prefill-decode-interval 10 alongside the
existing --enable-prefill-delayer
- chunked prefill under DP scales the per-TP base (16384 at TP8, 8192 at TP4)
by the DP degree rather than a fixed 8192
- mem-fraction-static 0.89 -> 0.90
Co-authored-by: Cursor <cursoragent@cursor.com>
* Point the changelog entry at PR #2800
Co-authored-by: Cursor <cursoragent@cursor.com>
* Set GPU_MAX_HW_QUEUES per-arm: 2 without DP, 5 under DP attention
The DP branch previously exported 2, and the tensor-parallel arms did not set
it at all, taking the container default. Both are now explicit: 2 on TP4/TP8
and 5 under DP attention, the documented companion to two-batch overlap.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Re-append the changelog entry so the diff is additions-only
The merge resolution reordered main's trailing entries, which check-changelog
rejects ("Deletions are not allowed in perf-changelog.yaml"). Restore main's
file verbatim and append this entry at the end instead.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Set mem-fraction-static per-TP, add the FP4 indexer flag, and extend the DP arm
- mem-fraction-static 0.86 on both TP8 arms, with and without DP attention;
TP4 stays at 0.89
- serve with --enable-deepseek-v4-fp4-indexer
- export HSA_NO_SCRATCH_RECLAIM=0
- DP arm concurrency [64, 96, 128] -> [64, 96, 128, 160]
Co-authored-by: Cursor <cursoragent@cursor.com>
* Lower mem-fraction-static to 0.86 on TP4 as well
Applies the same 0.86 to every arm rather than keeping TP4 at 0.89.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Update image to lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260902
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com> K
karverma-amd committed
650ee9e90a72329c4da0d5c386731a5216e55c79
Parent: 9bf93d0
Committed by GitHub <noreply@github.com>
on 9/5/2026, 12:05:13 AM