[Klaud Cold] dsv4-fp4-b200-vllm: refresh nightly and use CuTeDSL EP backend / 更新 DSV4 B200 vLLM nightly 并使用 CuTeDSL 专家并行后端 (#2534)
* dsv4 fp4 b200 vllm: prefill-schedule-interval 16 at CONC>=256, bump to nightly-700d39b * fill pr-link * DEP arm: prefill-schedule-interval 16, FULL_DECODE_ONLY, 2x seqs/capture, gmu 0.94; fix perf-changelog * DEP max-num-seqs and cudagraph-capture-size = CONC (not 2x) * DEP: MAX_NUM_SEQS=2*CONC/TP, cudagraph sizes 1..MAX_NUM_SEQS, FULL_DECODE_ONLY * cudagraph capture sizes linear 1..DEP_MAX_NUM_SEQS * add --attention-backend FLASHMLA_SPARSE_DSV4 * revert attention-backend, set image to vllm/vllm-openai:v0.26.0 * max-num-seqs only for DP_ATTENTION=true * disable eplb * dsv4 b200: revert script, megamoe when ep>1, new search space, nightly image * gmu 0.95, cap max-model-len to 12288 * fix(config): route DSV4 B200 vLLM to Nscale 将 DSV4 B200 vLLM 配置路由至 Nscale,并固定本地 NVFP4 模型路径。 * fix(config): refresh dsv4 B200 vLLM nightly 将 dsv4 B200 vLLM 配置更新至最新发布的不可变 nightly 镜像。 * fix(config): update DSV4 B200 vLLM nightly image * fix(vllm): use NVFP4-compatible EP backend Use the FlashInfer NVFP4 MegaMoE backend for expert-parallel scenarios and keep EPLB disabled because that backend does not support it.\n\n中文:为专家并行场景使用兼容 NVFP4 的 FlashInfer MegaMoE 后端,并禁用该后端不支持的 EPLB。 * fix(vllm): use standard CuTeDSL EP backend Use the standard FlashInfer CuTeDSL NVFP4 backend for expert-parallel scenarios and keep EPLB disabled. 中文:在专家并行场景中使用标准 FlashInfer CuTeDSL NVFP4 后端,并保持禁用 EPLB。 --------- Co-authored-by: Ankur Singh <ankusingh@nvidia.com> Co-authored-by: Rohit Nagraj <rohitnagraj.99@gmail.com> Co-authored-by: Rohit Pujar Nagraj <rpujarnagraj@nvidia.com> Co-authored-by: Kedar Potdar <kepotdar@nvidia.com> Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com>
X
Xin Li committed
d13df3789276e7b34ab69abd506e082cd9a687de
Parent: 62c140a
Committed by GitHub <noreply@github.com>
on 9/3/2026, 10:38:08 PM