NVDA, AMD · US

TileRT: ultra-high interactivity on NVIDIA GPUs? | SemiAnalysis note

Fast modes show users pay for latency. TileRT statically compiles decode into a persistent kernel; on InferenceX B200 can reach hundreds of tok/s/user—materially faster interactivity iso-cost vs classic engines. See CS-4, Rubin TCO, AgentX v3.

Published Updated Open interactive reader

Download MarkdownExport PDF

Cite a section with a deep link, e.g. /en/r/sa-tilert-inferencex-2026#thesis

Snapshot

Date
2026-08-10
基准场景
InferenceX / GLM5 等
B200 互动性量级
最高约 ~500 tok/s/user(文中)
相对传统引擎
约 2–3× 互动性(场景依赖)
As of 2026-09-25

As-of 2026-09-25 (weekly refresh; equities aligned to §A. Missing series are N/A/null. Not investment advice.)

Structured research note on a SemiAnalysis piece (summary + investable mapping)—not a reprint. Defer to the original for detail.

Thesis

Fast modes show users pay for latency. TileRT statically compiles decode into a persistent kernel; on InferenceX B200 can reach hundreds of tok/s/user—materially faster interactivity iso-cost vs classic engines. See CS-4, Rubin TCO.

Analysis

Why it matters

Paid fast modes prove latency premium. Labs evaluate Cerebras/Groq etc., so NVIDIA software must answer whether GPU fleets can reach the same interactivity frontier.

What TileRT does

Statically compiles the decode graph into a persistent kernel, maximizing overlap of compute/memory/comms. On InferenceX, a B200 decode server can hit ~500 tok/s/user in cited cases (~2–3× vs classic engines depending on workload). Edge is mainly decode tail; TTFT is fine, not magical.

Ecosystem

Competes with vLLM/SGLang/Dynamo and specialty silicon. AMD MI455X / TPUv7 will land in InferenceX. Read iso-cost/iso-power and agentic traces—not peak tok/s alone—using AgentX InferenceX v3 (393 Claude Code traces, 95%+ KV reuse) as the production-proxy benchmark, alongside frontier open weights in open models catching up.

Implications

Angle Implication Site map
NVDA 软件护城河 抬高 GPU 互动性上限,延缓份额流失到专用硅 NVDA · Rubin
专用推理硅 仍可能在极低 batch 占优,但差距可被软件压缩 CS-4
AMD / 开源栈 SGLang 等需持续对标 AMD

Outlook

Watch InferenceX verified Rubin/TPU/MI455X numbers, TileRT production adoption, Fast-tier pricing, vs CS-4/Groq on same loads; AgentX v3 extends the benchmark from single-node interactivity to multiturn agentic full-stack composability.

Risks

References

  1. 原文(SemiAnalysis): Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX
  2. InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper
  3. 对照 CS-4: Cerebras CS-4
  4. AgentX v3 (agentic full stack): AgentX InferenceX v3 · original

Not investment advice. Copyright remains with SemiAnalysis / authors; this is an index note.

Disclaimer: For research information only. Not investment advice or a recommendation to buy or sell.

Comments

Sign in to comment

Loading…