NVDA, AMD, GOOGL, MSFT, AMZN, MU · 美股
AgentX InferenceX v3:Agentic 推理里 CUDA 护城河还稳吗?|SemiAnalysis 笔记
Apache 2.0 AgentX:$3M 数据集、100 万+ 上下文、393 Claude Code trace、95%+ KV 复用;~2MW/1000+ 芯片测全栈。DCP/PCP composability 仍是 NVDA 优势;8/21 后 B200 perf/$ 反超 MI355X。对照 TileRT、开源模型追赶、AMD。
章节深链示例:/zh/r/sa-agentx-inferencex-v3-2026#thesis
数据时点:2026-09-25(周度刷新;个股按 §A 标准档对齐,缺数用 N/A/null,非投资建议)
AgentX InferenceX v3:Agentic 推理里 CUDA 护城河还稳吗?
来源:SemiAnalysis — AgentX InferenceX v3: Does CUDA Moat Hold up in Agentic Inferencing?(2026-08-24)。本页为 Investor Research 摘要笔记,非原文转载。
核心判断
SemiAnalysis 发布 AgentX InferenceX v3——Apache 2.0 开源 agentic coding 基准,配套 $3M 数据集、100 万+ token 上下文、393 条 Claude Code trace(95%+ KV cache 复用模式),在 ~2MW、1000+ 芯片(GB300/GB200 NVL72、MI355X、B200/B300、H200 等)上测 routing → KV offload → vLLM/SGLang/TRT-LLM/ATOM 全栈。
核心结论:agentic 长上下文 workload 把竞争从单 kernel 拉到 composability;DCP/PCP(context parallelism) 等特性在 vLLM 中对 AMD backend 全 unsupported,构成 CUDA 护城河 在 agentic 场景的新证据。2026-08-21 后 NVIDIA vLLM 优化使 B200 perf/$ 反超 MI355X,但 race 仍近。
前序基准 TileRT InferenceX 聚焦 tile/kernel;v3 扩展为 多轮、子 agent、KV 生命周期 的生产代理 workload,并与 开源模型追赶 中 Kimi K3、GLM-5.2 等 frontier open weights 实测直接挂钩。
AgentX 设计与硬件矩阵
数据集与 trace 特征
- 393 条 Claude Code session trace;P90 input sequence length 317k;95%+ KV cache hit rate(多轮 prefix 复用)。
- $3M 量级开源数据集;测试 sub-agent traffic、cold prefill spike、session stickiness 等单轮 benchmark 无法覆盖的模式。
- 与 Meta infra 文化、MSFT Copilot / GOOGL Gemini agent 产品所隐含 workload 形状一致。
硬件与 perf/$ 快照
| 维度 | SA 要点 |
|---|---|
| MI355X | ATOM 单点 e2e 可 beat B200 vLLM;但 Western/中国 lab 生产多仍用 vLLM/SGLang,ATOM 功能缺口大(阿里广告 BU 外少规模采用) |
| B200/B300 | 8/21 前 MI355X SGLang 与 B200 vLLM perf/$ 接近;8/21 后 Inferact+NVIDIA vLLM 优化使 B200 perf/$ 超 MI355X |
| GB300/GB200 NVL72 | Dynamo TRT-LLM / vLLM + PD disagg + wide-EP;高并发下 TTFT 对 subagent cold prefill 更敏感 |
| H200 | 低并发 DeepSeek v4 仍有 perf/$ 竞争力;高吞吐受 HBM 限制 |
| Rubin / TPU / MI455X | 文章时点:Rubin 当月到货;TPU 与 UALoE72 年内——对照 Vera/Rubin NVL72 TCO、Google–AMD TPU |
Frontier open-weight 模型
- Kimi K3(~2.8T 总参数):宽 EP/TP/PP 才能上单 B200 机架;B200 早期 PP 与 speculative decoding 不 compose,被 MI355X「mog」;长上下文 multiturn 首周 MI355X vLLM 不可用。
- MiniMax M3、DeepSeek v4、GLM-5.2:expose wide EP、DCP、MSA indexer、hybrid KV layout 等组合优化需求——与 开源模型追赶 agentic era 叙事一致。
CUDA 护城河与软件栈
DCP/PCP 与 composability
- PCP(prefill context parallelism):分 query chunk prefill,缓解 giant-prompt spike。
- DCP(decode context parallelism):分 KV shard 并行 decode,缓解 memory-BW bound。
- SA 明确:DCP/PCP 为 NVIDIA Research 贡献,构成 CUDA moat 一部分;vLLM support matrix 上 所有 AMD backend 对 DCP/PCP unsupported。
- Agentic 场景要求 DCP 与 prefix caching、chunked prefill、FP8 KV、MTP、spec decode 同时 compose——AMD 栈在 组合特性 上明显落后,而非单一 kernel 分数。
Routing 与 Dynamo
- Dynamo router 工作量随 live prefix 数量与长度 缩放,agentic trace 下 router 可成瓶颈(非 GPU kernel)。
- 一系列 PR:batched KV matching(+22.2% median throughput @ c512)、arena ownership / request leases(-23.7% replay time)、agentic router preset(p95 TTFT -43.1% 等)。
- Rack-scale GB200/300 在 AgentX 上 perf/TCO 优势不如单节点明显:router bottleneck + 缺少 tuned wide EP/DCP kernel + 更高 TCO。
KV offload 生态
- vLLM connector:LMCache、Mooncake;silent wrong KV placement 案例(127,500 token test:2/128 → 128/128 needles)——吞吐 benchmark 可能 又快又错。
- LMCache:chunked load 防 deadlock;DCP-aware CPU offload;AMD 路线用 Triton 替代 flashinfer、hipFile 替代 cuFile、预编译 gfx942/gfx950 wheel。
- Mooncake:Kimi 生产 traffic;AMD GPU-direct RDMA(HIP dmabuf)与 ROCm wheel 补齐。
- 内存层受益:MU HBM、GPU 利用率 主题中的 KV 字节/token 上升。
70+ upstream PR
- 合入 vLLM、SGLang、TRT-LLM、Dynamo、LMCache、Mooncake、AITER 等;SemiAnalysis 与 Inferact 等协作。
- 说明 agentic inference 竞争已是 开源软件 velocity 游戏——利好 NVDA 生态锁定,也给予 AMD 若 upstream 成功则缩小 gap 的路径(对照 CUDA 护城河 全文)。
与其它 SA 笔记 / 研报交叉
- Cerebras CS-4:wafer-scale 绕开 multi-GPU composability 痛苦,但 AgentX 类 trace 是否适配另论。
- SpaceX–MSFT compute:hyperscaler 自建与 neo-cloud(CRWV、NBIS、APLD)对 推理 perf/$ 敏感。
- Gemini/GCP、Google–AMD TPU:Google 双轨(TPU + GPU agent serving)。
- VRT、CEG 等:~2MW 持续 AgentX 集群的 电力与散热 约束。
- INTC:CPU 在 orchestration / 前端/router 侧角色(Dynamo frontend CPU 优化占比较大)。
展望
- SA 预告 AgentX update 很快发布——perf/$ 排名可能随两周内 PR 再度洗牌。
- Rubin NVL72、MI455X UALoE72 入库后将重跑矩阵;关注 Rubin TCO 假设是否成立。
- 投资框架:agentic inference = GPU × composable software × memory × power;单看 TFLOPS 或单模型 leaderboard 不足。
- 若 AMD 持续 upstream ATOM 优化到 vLLM/SGLang 并补齐 DCP/PCP,AMD 推理 narrative 可强化;反之 moat 稳固利好 NVDA。
风险
- 基准 proxy 于 Claude Code;其它 agent(coding vs general)trace 分布不同。
- perf/$ 时间敏感;8/21 前后排名已变。
- ATOM vs vLLM 分叉:生产 adoption 与 benchmark 最优配置不一致。
- 正确性 bug(KV placement、split-K MoE tactic)在吞吐测试中易被忽略。
- Rack-scale 优势在 AgentX 上被 router / missing wide kernel 稀释。
参考
- SemiAnalysis — AgentX InferenceX v3(2026-08-24)
- GitHub:AgentX / InferenceX(Apache 2.0;请以原文为准)
- 关联笔记:TileRT InferenceX · CUDA 护城河 · 开源模型追赶 · Vera/Rubin TCO · Cerebras CS-4 · Gemini/GCP
- 关联研报:NVDA · AMD · MU · AVGO · GOOGL · MSFT · GPU 利用率 · CRWV · NBIS
评论
登录后评论