NVDA, AMD, GOOGL, MSFT, AMZN, MU · US
AgentX InferenceX v3: CUDA moat in agentic inference? | SemiAnalysis note
Apache 2.0 AgentX: $3M dataset, 1M+ context, 393 Claude Code traces, 95%+ KV reuse; ~2MW/1000+ chips full-stack test. DCP/PCP composability favors NVDA; post Aug 21 B200 perf/$ passed MI355X. See TileRT, open models note, AMD.
Cite a section with a deep link, e.g. /en/r/sa-agentx-inferencex-v3-2026#thesis
As-of 2026-09-25 (weekly refresh; equities aligned to §A. Missing series are N/A/null. Not investment advice.)
AgentX InferenceX v3: Does CUDA moat hold in agentic inference?
Source: SemiAnalysis — AgentX InferenceX v3: Does CUDA Moat Hold up in Agentic Inferencing? (2026-08-24). Investor Research digest; not a full republish.
Thesis
SemiAnalysis releases AgentX InferenceX v3—an Apache 2.0 open agentic coding benchmark with a $3M dataset, 1M+ token context, 393 Claude Code traces (95%+ KV cache reuse patterns), run on ~2MW, 1000+ chips (GB300/GB200 NVL72, MI355X, B200/B300, H200, etc.) across routing → KV offload → vLLM/SGLang/TRT-LLM/ATOM.
Core finding: agentic long-context workloads shift competition from single kernels to composability; DCP/PCP (context parallelism) remains unsupported on all vLLM AMD backends, reinforcing the CUDA moat in agentic settings. After 2026-08-21 NVIDIA vLLM optimizations, B200 perf/$ passed MI355X, but the race stays close.
Prior work TileRT InferenceX focused on tile/kernel efficiency; v3 adds multiturn, sub-agent, KV lifecycle production-proxy workloads, tied to frontier open weights in open models catching up (Kimi K3, GLM-5.2, etc.).
AgentX design & hardware matrix
Dataset & trace shape
- 393 Claude Code session traces; P90 ISL 317k; 95%+ KV cache hit rate (multiturn prefix reuse).
- $3M open dataset; exercises sub-agent traffic, cold prefill spikes, session stickiness absent from single-turn benchmarks.
- Aligns with workload shapes implied by Meta infra culture and agent products at MSFT and GOOGL.
Hardware & perf/$ snapshot
| Dimension | SA highlights |
|---|---|
| MI355X | ATOM can beat B200 vLLM on e2e in spots; production still mostly vLLM/SGLang—ATOM missing features limit adoption beyond one Alibaba ads BU |
| B200/B300 | Pre Aug 21 MI355X SGLang ~matched B200 vLLM perf/$; post Aug 21 Inferact+NVIDIA vLLM opts put B200 perf/$ ahead of MI355X |
| GB300/GB200 NVL72 | Dynamo TRT-LLM / vLLM + PD disagg + wide-EP; TTFT more sensitive to subagent cold prefills at high concurrency |
| H200 | Competitive perf/$ on DeepSeek v4 at low concurrency; HBM limits high-throughput scenarios |
| Rubin / TPU / MI455X | Rubin arriving that month; TPU & UALoE72 later in year—see Vera/Rubin NVL72 TCO and Google–AMD TPU |
Frontier open-weight models
- Kimi K3 (~2.8T total params): needs wide EP/TP/PP; early B200 PP + spec decode incompatibility hurt vs MI355X; MI355X vLLM broken week one on realistic long multiturn.
- MiniMax M3, DeepSeek v4, GLM-5.2: expose wide EP, DCP, MSA indexer, hybrid KV combo needs—consistent with agentic era in open models catching up.
CUDA moat & software stack
DCP/PCP & composability
- PCP: shard prefill queries; reduces giant-prompt spikes.
- DCP: shard KV for decode; helps memory-BW-bound steps.
- SA: DCP/PCP from NVIDIA Research, part of CUDA moat; vLLM matrix shows all AMD backends unsupported.
- Agentic serving needs DCP to compose with prefix cache, chunked prefill, FP8 KV, MTP, speculative decoding—AMD lag is in feature composition, not one kernel score.
Routing & Dynamo
- Dynamo router cost scales with number and length of live prefixes—can bottleneck before GPUs on agentic traces.
- PR series: batched KV matching (+22.2% median throughput @ c512), arena ownership / request leases (-23.7% replay time), agentic router preset (p95 TTFT -43.1%, etc.).
- Rack-scale GB200/300 underperform vs single-node on AgentX: router limits + missing wide EP/DCP kernels + higher TCO.
KV offload ecosystem
- vLLM connectors: LMCache, Mooncake; silent wrong KV placement (127,500-token test: 2/128 → 128/128 needles)—throughput benchmarks can be fast and wrong.
- LMCache: chunked load anti-deadlock; DCP-aware CPU offload; AMD path via Triton vs flashinfer, hipFile vs cuFile, prebuilt gfx942/gfx950 wheels.
- Mooncake: Kimi production; AMD GPU-direct RDMA (HIP dmabuf) and ROCm wheels.
- Memory beneficiaries: MU HBM; rising KV bytes/token in GPU utilization theme.
70+ upstream PRs
- Landed across vLLM, SGLang, TRT-LLM, Dynamo, LMCache, Mooncake, AITER—agentic inference is an open-source velocity game.
- Favors NVDA ecosystem lock-in; AMD can close gap if upstream succeeds (see CUDA moat full note).
Cross-links to other SA notes & research
- Cerebras CS-4: wafer-scale sidesteps multi-GPU composability pain; AgentX trace fit TBD.
- SpaceX–MSFT compute & neo-cloud (CRWV, NBIS, APLD): sensitive to inference perf/$.
- Gemini/GCP, Google–AMD TPU: Google dual track (TPU + GPU agents).
- VRT, CEG: ~2MW AgentX cluster power/cooling constraints.
- INTC: CPU role in orchestration / Dynamo frontend (meaningful frontend CPU tuning).
Outlook
- SA teases an AgentX update soon—perf/$ ranks may reshuffle within weeks of new PRs.
- Rubin NVL72, MI455X UALoE72 entries will rerun the matrix; watch Rubin TCO assumptions.
- Framework: agentic inference = GPU × composable software × memory × power; TFLOPS or leaderboard alone insufficient.
- If AMD upstream ATOM wins into vLLM/SGLang and ships DCP/PCP, AMD inference narrative strengthens; else moat supports NVDA.
Risks
- Benchmark proxies Claude Code; other agent trace distributions differ.
- perf/$ time-sensitive; rankings changed across Aug 21.
- ATOM vs vLLM fork: production configs ≠ benchmark winners.
- Correctness bugs (KV placement, bad split-K MoE tactics) hide in throughput tests.
- Rack-scale advantage diluted by router / missing wide kernels on AgentX.
References
- SemiAnalysis — AgentX InferenceX v3 (2026-08-24)
- GitHub: AgentX / InferenceX (Apache 2.0; verify upstream)
- Related notes: TileRT InferenceX · CUDA moat · Open models catching up · Vera/Rubin TCO · Cerebras CS-4 · Gemini/GCP
- Related research: NVDA · AMD · MU · AVGO · GOOGL · MSFT · GPU utilization · CRWV · NBIS
Comments
Sign in to comment