SIGN IN SIGN UP

[AMD][DSV4] Add a DP-attention arm and refresh the MI355X SGLang AgentX key (#2800)

* [AMD][DSV4] Add a DP-attention arm and refresh the MI355X SGLang AgentX key

Image: lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260829 -> ...-20260831.

Search space:
- add tp 8, ep 1, dp-attn true on the hicache host KV tier at concurrency
  [64, 96, 128]
- drop concurrency 2 and 10 from the TP4 arm, leaving [1, 4, 8]
- drop concurrency 16 from the TP8 hicache arm, leaving [32, 48]; conc 16
  stays on the TP8 no-offload arm

Serving, split per-arm so the tensor-parallel arms keep their current settings:
- shared-experts fusion stays enforced without DP and is disabled under DP
  attention, matching the DP baselines
- swa-full-tokens-ratio stays 0.10 without DP and is 0.15 under DP attention
- DP attention adds --enable-two-batch-overlap,
  --enable-dp-attention-local-control-broadcast, --tokenizer-worker-num equal
  to TP, --stream-interval 20 and --prefill-decode-interval 10 alongside the
  existing --enable-prefill-delayer
- chunked prefill under DP scales the per-TP base (16384 at TP8, 8192 at TP4)
  by the DP degree rather than a fixed 8192
- mem-fraction-static 0.89 -> 0.90

Co-authored-by: Cursor <cursoragent@cursor.com>

* Point the changelog entry at PR #2800

Co-authored-by: Cursor <cursoragent@cursor.com>

* Set GPU_MAX_HW_QUEUES per-arm: 2 without DP, 5 under DP attention

The DP branch previously exported 2, and the tensor-parallel arms did not set
it at all, taking the container default. Both are now explicit: 2 on TP4/TP8
and 5 under DP attention, the documented companion to two-batch overlap.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Re-append the changelog entry so the diff is additions-only

The merge resolution reordered main's trailing entries, which check-changelog
rejects ("Deletions are not allowed in perf-changelog.yaml"). Restore main's
file verbatim and append this entry at the end instead.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Set mem-fraction-static per-TP, add the FP4 indexer flag, and extend the DP arm

- mem-fraction-static 0.86 on both TP8 arms, with and without DP attention;
  TP4 stays at 0.89
- serve with --enable-deepseek-v4-fp4-indexer
- export HSA_NO_SCRATCH_RECLAIM=0
- DP arm concurrency [64, 96, 128] -> [64, 96, 128, 160]

Co-authored-by: Cursor <cursoragent@cursor.com>

* Lower mem-fraction-static to 0.86 on TP4 as well

Applies the same 0.86 to every arm rather than keeping TP4 at 0.89.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Update image to lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260902

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com>
K
karverma-amd committed
650ee9e90a72329c4da0d5c386731a5216e55c79
Parent: 9bf93d0
Committed by GitHub <noreply@github.com> on 9/5/2026, 12:05:13 AM