NVDA, AMD, GOOGL, MSFT, AMZN, MU · US

AgentX InferenceX v3: CUDA moat in agentic inference? | SemiAnalysis note

Apache 2.0 AgentX: $3M dataset, 1M+ context, 393 Claude Code traces, 95%+ KV reuse; ~2MW/1000+ chips full-stack test. DCP/PCP composability favors NVDA; post Aug 21 B200 perf/$ passed MI355X. See TileRT, open models note, AMD.

Published Updated Open interactive reader

Download MarkdownExport PDF

Cite a section with a deep link, e.g. /en/r/sa-agentx-inferencex-v3-2026#thesis

As-of 2026-09-25 (weekly refresh; equities aligned to §A. Missing series are N/A/null. Not investment advice.)

AgentX InferenceX v3: Does CUDA moat hold in agentic inference?

Source: SemiAnalysis — AgentX InferenceX v3: Does CUDA Moat Hold up in Agentic Inferencing? (2026-08-24). Investor Research digest; not a full republish.

Thesis

SemiAnalysis releases AgentX InferenceX v3—an Apache 2.0 open agentic coding benchmark with a $3M dataset, 1M+ token context, 393 Claude Code traces (95%+ KV cache reuse patterns), run on ~2MW, 1000+ chips (GB300/GB200 NVL72, MI355X, B200/B300, H200, etc.) across routing → KV offload → vLLM/SGLang/TRT-LLM/ATOM.

Core finding: agentic long-context workloads shift competition from single kernels to composability; DCP/PCP (context parallelism) remains unsupported on all vLLM AMD backends, reinforcing the CUDA moat in agentic settings. After 2026-08-21 NVIDIA vLLM optimizations, B200 perf/$ passed MI355X, but the race stays close.

Prior work TileRT InferenceX focused on tile/kernel efficiency; v3 adds multiturn, sub-agent, KV lifecycle production-proxy workloads, tied to frontier open weights in open models catching up (Kimi K3, GLM-5.2, etc.).

AgentX design & hardware matrix

Dataset & trace shape

  • 393 Claude Code session traces; P90 ISL 317k; 95%+ KV cache hit rate (multiturn prefix reuse).
  • $3M open dataset; exercises sub-agent traffic, cold prefill spikes, session stickiness absent from single-turn benchmarks.
  • Aligns with workload shapes implied by Meta infra culture and agent products at MSFT and GOOGL.

Hardware & perf/$ snapshot

Dimension SA highlights
MI355X ATOM can beat B200 vLLM on e2e in spots; production still mostly vLLM/SGLang—ATOM missing features limit adoption beyond one Alibaba ads BU
B200/B300 Pre Aug 21 MI355X SGLang ~matched B200 vLLM perf/$; post Aug 21 Inferact+NVIDIA vLLM opts put B200 perf/$ ahead of MI355X
GB300/GB200 NVL72 Dynamo TRT-LLM / vLLM + PD disagg + wide-EP; TTFT more sensitive to subagent cold prefills at high concurrency
H200 Competitive perf/$ on DeepSeek v4 at low concurrency; HBM limits high-throughput scenarios
Rubin / TPU / MI455X Rubin arriving that month; TPU & UALoE72 later in year—see Vera/Rubin NVL72 TCO and Google–AMD TPU

Frontier open-weight models

  • Kimi K3 (~2.8T total params): needs wide EP/TP/PP; early B200 PP + spec decode incompatibility hurt vs MI355X; MI355X vLLM broken week one on realistic long multiturn.
  • MiniMax M3, DeepSeek v4, GLM-5.2: expose wide EP, DCP, MSA indexer, hybrid KV combo needs—consistent with agentic era in open models catching up.

CUDA moat & software stack

DCP/PCP & composability

  • PCP: shard prefill queries; reduces giant-prompt spikes.
  • DCP: shard KV for decode; helps memory-BW-bound steps.
  • SA: DCP/PCP from NVIDIA Research, part of CUDA moat; vLLM matrix shows all AMD backends unsupported.
  • Agentic serving needs DCP to compose with prefix cache, chunked prefill, FP8 KV, MTP, speculative decoding—AMD lag is in feature composition, not one kernel score.

Routing & Dynamo

  • Dynamo router cost scales with number and length of live prefixes—can bottleneck before GPUs on agentic traces.
  • PR series: batched KV matching (+22.2% median throughput @ c512), arena ownership / request leases (-23.7% replay time), agentic router preset (p95 TTFT -43.1%, etc.).
  • Rack-scale GB200/300 underperform vs single-node on AgentX: router limits + missing wide EP/DCP kernels + higher TCO.

KV offload ecosystem

  • vLLM connectors: LMCache, Mooncake; silent wrong KV placement (127,500-token test: 2/128 → 128/128 needles)—throughput benchmarks can be fast and wrong.
  • LMCache: chunked load anti-deadlock; DCP-aware CPU offload; AMD path via Triton vs flashinfer, hipFile vs cuFile, prebuilt gfx942/gfx950 wheels.
  • Mooncake: Kimi production; AMD GPU-direct RDMA (HIP dmabuf) and ROCm wheels.
  • Memory beneficiaries: MU HBM; rising KV bytes/token in GPU utilization theme.

70+ upstream PRs

  • Landed across vLLM, SGLang, TRT-LLM, Dynamo, LMCache, Mooncake, AITER—agentic inference is an open-source velocity game.
  • Favors NVDA ecosystem lock-in; AMD can close gap if upstream succeeds (see CUDA moat full note).

Cross-links to other SA notes & research

Outlook

  • SA teases an AgentX update soon—perf/$ ranks may reshuffle within weeks of new PRs.
  • Rubin NVL72, MI455X UALoE72 entries will rerun the matrix; watch Rubin TCO assumptions.
  • Framework: agentic inference = GPU × composable software × memory × power; TFLOPS or leaderboard alone insufficient.
  • If AMD upstream ATOM wins into vLLM/SGLang and ships DCP/PCP, AMD inference narrative strengthens; else moat supports NVDA.

Risks

  • Benchmark proxies Claude Code; other agent trace distributions differ.
  • perf/$ time-sensitive; rankings changed across Aug 21.
  • ATOM vs vLLM fork: production configs ≠ benchmark winners.
  • Correctness bugs (KV placement, bad split-K MoE tactics) hide in throughput tests.
  • Rack-scale advantage diluted by router / missing wide kernels on AgentX.

References

Disclaimer: For research information only. Not investment advice or a recommendation to buy or sell.

Comments

Sign in to comment

Loading…