SIGN IN SIGN UP

[Klaud Cold] dsv4-fp4-b200-vllm: refresh nightly and use CuTeDSL EP backend / 更新 DSV4 B200 vLLM nightly 并使用 CuTeDSL 专家并行后端 (#2534)

* dsv4 fp4 b200 vllm: prefill-schedule-interval 16 at CONC>=256, bump to nightly-700d39b

* fill pr-link

* DEP arm: prefill-schedule-interval 16, FULL_DECODE_ONLY, 2x seqs/capture, gmu 0.94; fix perf-changelog

* DEP max-num-seqs and cudagraph-capture-size = CONC (not 2x)

* DEP: MAX_NUM_SEQS=2*CONC/TP, cudagraph sizes 1..MAX_NUM_SEQS, FULL_DECODE_ONLY

* cudagraph capture sizes linear 1..DEP_MAX_NUM_SEQS

* add --attention-backend FLASHMLA_SPARSE_DSV4

* revert attention-backend, set image to vllm/vllm-openai:v0.26.0

* max-num-seqs only for DP_ATTENTION=true

* disable eplb

* dsv4 b200: revert script, megamoe when ep>1, new search space, nightly image

* gmu 0.95, cap max-model-len to 12288

* fix(config): route DSV4 B200 vLLM to Nscale

将 DSV4 B200 vLLM 配置路由至 Nscale,并固定本地 NVFP4 模型路径。

* fix(config): refresh dsv4 B200 vLLM nightly

将 dsv4 B200 vLLM 配置更新至最新发布的不可变 nightly 镜像。

* fix(config): update DSV4 B200 vLLM nightly image

* fix(vllm): use NVFP4-compatible EP backend

Use the FlashInfer NVFP4 MegaMoE backend for expert-parallel scenarios and keep EPLB disabled because that backend does not support it.\n\n中文:为专家并行场景使用兼容 NVFP4 的 FlashInfer MegaMoE 后端,并禁用该后端不支持的 EPLB。

* fix(vllm): use standard CuTeDSL EP backend

Use the standard FlashInfer CuTeDSL NVFP4 backend for expert-parallel scenarios and keep EPLB disabled.

中文:在专家并行场景中使用标准 FlashInfer CuTeDSL NVFP4 后端,并保持禁用 EPLB。

---------

Co-authored-by: Ankur Singh <ankusingh@nvidia.com>
Co-authored-by: Rohit Nagraj <rohitnagraj.99@gmail.com>
Co-authored-by: Rohit Pujar Nagraj <rpujarnagraj@nvidia.com>
Co-authored-by: Kedar Potdar <kepotdar@nvidia.com>
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com>
X
Xin Li committed
d13df3789276e7b34ab69abd506e082cd9a687de
Parent: 62c140a
Committed by GitHub <noreply@github.com> on 9/3/2026, 10:38:08 PM