NVDA, AMD, GOOGL, META, MSFT · US

Are Open Models Catching Up? Three eras | SemiAnalysis note

SemiAnalysis: scaling / reasoning / agentic eras; open weights close gap in ~half the time each era (Kimi K2.6 vs Opus 4.5 ~4.8 mo), but benchmarks ≠ product. See AgentX v3, CUDA moat, NVDA.

Published Updated Open interactive reader

Download MarkdownExport PDF

Cite a section with a deep link, e.g. /en/r/sa-open-models-catching-up-2026#thesis

As-of 2026-09-25 (weekly refresh; equities aligned to §A. Missing series are N/A/null. Not investment advice.)

Are Open Models Catching Up? Three eras of convergence

Source: SemiAnalysis — Are Open Models Catching Up? (2026-08-21). Investor Research digest; not a full republish.

Thesis

SemiAnalysis frames LLM progress as scaling → reasoning → agentic eras and finds open weights close the gap to the closed frontier in roughly half the time each era—while benchmark scores still diverge from daily product quality; the team still prefers Anthropic Claude for routine work.

Agentic inference stack reality is captured in the companion AgentX InferenceX v3 release: once workloads shift to long-context, multiturn, sub-agent traces, hardware perf/$ and software composability explain investability better than leaderboards (see CUDA moat and TileRT InferenceX).

Three eras & convergence

Era 1 — Scaling

  • Llama 3.1 405B vs GPT-4o: ~6 months gap (SA framing).
  • Open models catch up quickly on dense scaling, but data, compute, and post-training resources remain concentrated in closed labs.

Era 2 — Reasoning

  • DeepSeek R1 vs OpenAI o1: ~6 months.
  • Competition shifts to RL data, verifiers, and inference-time compute—parallel to CPU industry orchestration narratives and ISA landscape edge inference themes.

Era 3 — Agentic

  • Kimi K2.6 vs Claude Opus 4.5: ~4.8 months.
  • GLM-5.2 vs GPT-5.2: ~6 months.
  • Production traits—trace shape, KV reuse, routing—are quantified in AgentX v3 (e.g. P90 ISL 317k, 95%+ KV hit rate).

Product vs benchmark caveat

  • SA still rates Claude higher for daily coding/agent tasks; open models may match public leaderboards yet lag on tool stability and long-session coherence.
  • RL hill-climbing risk: open community may overfit public evals—contrasts with closed labs’ private eval + product loops (Meta superintelligence, Gemini/GCP).

Compute & stack implications

Training vs inference

Cloud & distribution

  • Closed APIs (GOOGL Gemini, MSFT Copilot, AMZN Bedrock) still own distribution and billing; open weights compress model-layer margin and raise infra / neo-cloud value (CRWV, NBIS, IREN).

China open vs Western closed

  • Kimi, GLM, DeepSeek, Qwen accelerate convergence—consistent with AgentX MI355X / Kimi K3 / MiniMax M3 matrix; export controls remain friction for Western cloud adoption of certain weights.

Outlook

  • Gap may keep narrowing in the agentic era, but whether half-life per era holds depends on closed labs widening private data and product RL advantages.
  • Track three signals: (1) open leaderboards, (2) production workloads like AgentX, (3) inference stack PR velocity (CUDA moat, TileRT).
  • INTC and CPU players may gain in edge/local agents as small open models improve; datacenter agentic battleground remains GPU + full stack.

Risks

  • Benchmark vs product divergence; RL overfitting public evals.
  • Closed labs retain top post-training compute and user feedback loops.
  • Geopolitics, copyright, weight export limits.
  • Open acceleration commoditizes the model layer (infra up, undifferentiated API down).

References

Disclaimer: For research information only. Not investment advice or a recommendation to buy or sell.

Comments

Sign in to comment

Loading…