SIGN IN SIGN UP

feat(config): add GLM-5.2 GB300 AgentX concurrency-1 disaggregated point / 添加 GLM-5.2 GB300 AgentX 并发度 1 分离式配置 (#2720)

* feat(config): add GLM-5.2 GB300 AgentX disaggregated point

新增 GLM-5.2 GB300 AgentX 并发度 1 的预填充与解码分离配置,并移除已替代的聚合配置。

* chore(changelog): link PR #2720

在性能变更日志中补充 PR #2720 链接。

* fix(config): colocate GLM-5.2 AgentX clients

将 GLM-5.2 AgentX 基准客户端与 Dynamo 前端统一部署在头节点,避免本地端口连接失败。

* fix(config): require TRT-LLM metrics for GLM-5.2 c1

为 GLM-5.2 GB300 并发度 1 配置启用迭代性能统计,并要求采集 TensorRT-LLM KV 缓存指标。

* chore: refresh PR #2720 for sweep reuse [skip-sweep]

---------

Co-authored-by: Cam Quilici <cjquilici@gmail.com>
Co-authored-by: Ankur-singh <ankusingh@nvidia.com>
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com>
R
Rohit Nagraj committed
f1e2e19d727c736d359fef9ebe87a696af461165
Parent: 131d052
Committed by GitHub <noreply@github.com> on 9/3/2026, 10:27:04 PM