Andrew Mercer
on this page

The LLM Ecosystem in Late 2026

Snapshot date: 8 October 2026. This field moves on a weekly cadence. Model names, prices and "latest" claims below were verified against public sources in early October 2026 — treat anything here older than a month or two as suspect and re-check vendor pricing pages before committing to a design.

TL;DR

  • Five companies set the frontier for closed models: Anthropic (Claude), OpenAI (GPT), Google DeepMind (Gemini), xAI/SpaceX (Grok) and Meta (Muse). Each now ships a tiered family (flagship / balanced / fast) rather than a single "best" model.
  • Open-weight models are close behind — DeepSeek V4, Alibaba Qwen 3.8, Moonshot Kimi K3, Thinking Machines Inkling, Mistral, Google Gemma 4 and Meta Muse Glimmer. The best of them are within a few months of the closed frontier, but the largest ones need datacentre hardware, not a workstation.
  • The competition has shifted from "smartest model" to "right model at the right price." Prices for mid-tier models roughly halved over the summer of 2026.
  • Regulation is now part of the release process. Frontier launches in 2026 have gone through government review, export-control suspensions, and gated "cyber" variants available only to vetted defenders.
  • Agents are the main use case. Long-horizon coding, computer use and tool calling (via MCP) are what labs optimise and benchmark for.

Table of contents

  1. How LLMs work (the 10-minute version)
  2. The ecosystem in layers
  3. Closed vs open weights
  4. The frontier labs and their models
  5. The open-weight labs
  6. Pricing comparison
  7. How the models actually differ
  8. Running models yourself
  9. Access, gateways and routing
  10. Agents, protocols and coding tools
  11. Regulation and safety gating
  12. Benchmarks — and why to distrust them
  13. Choosing a model: a decision guide
  14. Glossary
  15. Sources

1. How LLMs work (the 10-minute version)

Tokens and next-token prediction

An LLM is a neural network trained to predict the next token (a word fragment — roughly ¾ of an English word on average) given all the tokens before it. Everything else — chat, coding, reasoning, tool use — is built on top of that one capability. Pricing, context limits and speed are all measured in tokens.

The transformer

Nearly every production LLM is a decoder-only transformer. Its core mechanism, attention, lets every token look at every earlier token to decide what is relevant. Attention cost grows with the square of the sequence length, which is why long context is expensive and why labs invest heavily in tricks such as sparse attention, compressed KV caches and sliding windows.

Dense vs Mixture-of-Experts (MoE)

  • Dense models use every parameter for every token (e.g. a 27B dense model does 27B parameters of work per token).
  • MoE models contain many "expert" sub-networks and a router that activates only a few per token. A model may have trillions of total parameters but only tens of billions active. This gives large-model knowledge at small-model compute cost — but all the weights still have to sit in memory.

Almost every frontier-scale open model in 2026 is MoE. Some examples: DeepSeek V4-Pro is 1.6T total / 49B active, and Kimi K3 is 2.8T total / 104B active.

Training stages

Stage What happens Why it matters
Pre-training Next-token prediction over tens of trillions of tokens of text, code, images, audio and video Gives the model knowledge and raw capability. The most expensive step. Sets the knowledge cutoff.
Supervised fine-tuning (SFT) Train on curated instruction/response examples Turns a text predictor into an assistant
Preference tuning (RLHF / RLAIF / constitutional methods) Reward the model for responses humans or AI judges prefer Tone, helpfulness, harmlessness, refusals
Reinforcement learning with verifiable rewards RL on tasks with checkable answers — unit tests, maths, terminal tasks The main driver of 2025–2026 gains in coding and reasoning
Distillation Train smaller models to imitate a larger one How "Flash/Luna/Haiku" tiers inherit flagship behaviour cheaply

Reasoning ("thinking") models

Since 2024, most models can generate hidden or visible chain-of-thought tokens before answering. In 2026 this is the default: models use adaptive thinking and expose an effort control (low → max). More effort means better answers on hard problems but more output tokens, more latency and more cost. Reasoning tokens are billed as output tokens.

Context windows

The context window is the total number of tokens the model can consider at once (prompt + history + tool results + output). 1M tokens is now table stakes for flagship and mid-tier models from Anthropic, OpenAI, Google, DeepSeek and Moonshot. Having a 1M window doesn't mean the model uses all of it equally well — recall degrades with distance and cost scales with length, which is why prompt caching (cheap re-reads of an unchanged prefix) matters so much for agents.

Multimodality

Most frontier models accept text, images and often audio/video as input; most still produce text output. Image, video and voice generation usually come from separate models (e.g. OpenAI's GPT Image, Google's Nano Banana and Veo/Omni lines, xAI's Grok Imagine).

Tool use and agents

Models emit structured tool calls (JSON) that a harness executes — run a shell command, read a file, search the web, click a button — and feed the result back. An agent is just this loop running for many steps toward a goal. Almost all of the 2026 model competition is about how reliably this loop works over long horizons.


2. The ecosystem in layers

┌──────────────────────────────────────────────────────────────────┐
│ APPLICATIONS & AGENTS  ChatGPT, Claude apps, Gemini app, Meta AI, │
│                        coding agents (Claude Code, Codex, Cursor, │
│                        OpenCode, Antigravity, Kiro, Vibe…)        │
├──────────────────────────────────────────────────────────────────┤
│ FRAMEWORKS & PROTOCOLS MCP, ACP, A2A, agent SDKs, LangGraph,      │
│                        LlamaIndex, evals & observability          │
├──────────────────────────────────────────────────────────────────┤
│ GATEWAYS & ROUTERS     OpenRouter, LiteLLM, cloud AI gateways,    │
│                        OpenAI-compatible endpoints                │
├──────────────────────────────────────────────────────────────────┤
│ HOSTING / INFERENCE    First-party APIs, AWS Bedrock, Google      │
│                        Vertex, Microsoft Foundry, Together,       │
│                        Fireworks, Groq…; self-hosted vLLM/SGLang  │
├──────────────────────────────────────────────────────────────────┤
│ MODELS                 Closed: Claude, GPT, Gemini, Grok, Muse    │
│                        Open:   DeepSeek, Qwen, Kimi, Inkling,     │
│                                Mistral, Gemma, gpt-oss, Glimmer   │
├──────────────────────────────────────────────────────────────────┤
│ COMPUTE                NVIDIA GPUs, Google TPUs, AWS Trainium,    │
│                        hyperscaler & neocloud datacentres         │
└──────────────────────────────────────────────────────────────────┘

A useful mental model: the model is a component, not the product. The same Claude, GPT or Qwen model behaves very differently depending on the harness, system prompt, tools and context management around it.


3. Closed vs open weights

Closed (proprietary) Open weights
What you get API access only Downloadable model weights
Examples Claude, GPT-6, Gemini, Grok, Muse Spark DeepSeek V4, Qwen 3.8, Kimi K3, Inkling, Mistral Small 4 / Large 3, Gemma 4, gpt-oss, Muse Glimmer
Best raw capability Yes — still the frontier Typically a few months behind
Data control Data leaves your environment (subject to vendor terms) Can run fully air-gapped
Customisation Prompting, some hosted fine-tuning Full fine-tuning, quantisation, distillation
Cost model Per-token Hardware + ops (or per-token from third-party hosts)
Safety controls Vendor-enforced classifiers and policies Your responsibility

"Open" is a spectrum

  • Open weights ≠ open source. Almost no frontier "open" model releases its training data or full training code. "Open weights" is the accurate term.
  • Licences vary. Permissive (MIT, Apache 2.0) vs custom licences with usage or scale restrictions. Kimi K3 uses a Modified MIT licence; Mistral uses Apache 2.0 for some models and proprietary terms for others. Always read the licence.
  • Open ≠ runnable at home. Kimi K3 in 4-bit precision still needs roughly 1.4TB of fast memory, and Moonshot recommends at least 64 accelerators. For trillion-parameter MoE models, "open" in practice means "any cloud can host it," not "you can run it on a workstation."

The gap

Through 2026, blind arena evaluations have put the best open models (Kimi K3 in particular, on front-end coding) level with or ahead of leading US closed models on specific tasks. Commentators estimate the open-to-closed gap has narrowed from roughly six to nine months down to three to five.


4. The frontier labs and their models

Anthropic — Claude

Positioning: Safety-focused lab; strongest reputation in agentic coding, long-horizon tasks and enterprise. Products: Claude apps, Claude Code, the Claude API, and availability on AWS Bedrock, Google Cloud Vertex AI and Microsoft Foundry.

Current lineup (Claude 5 / 5.5 generation):

Model Released Role API price (in / out per 1M) Notes
Claude Mythos 5.1 2026 Top "Mythos" tier Restricted Not publicly available; used by a small number of trusted organisations through Anthropic's Project Glasswing
Claude Fable 5.1 1 Sep 2026 Public Mythos-tier model See pricing page (Fable 5 launched at $10 / $50) Same underlying model as Mythos 5.1 with extra safeguards for biology, cybersecurity and LLM R&D. Long-horizon reasoning and agentic coding
Claude Opus 5.5 22 Sep 2026 Flagship workhorse $4 / $20 Cheaper and faster than Opus 5; Anthropic says ~40% lower typical cost. Fast mode available at a premium
Claude Sonnet 5.5 28 Sep 2026 Balanced default $2 / $10 1M context, 128K output, adaptive thinking on by default
Claude Haiku 5.5 After Sonnet 5.5 Fast / cheap tier See pricing page High-volume, cost-sensitive work; first Haiku in the 5.x generation

API IDs follow the pattern claude-opus-5-5, claude-sonnet-5-5, claude-haiku-5-5, claude-fable-5-1.

Notable 2026 events:

  • Claude Fable 5 and Mythos 5 shipped on 9 June 2026. On 12 June Anthropic suspended access to both to comply with US Department of Commerce export controls. The controls were lifted on 30 June and access was restored on 1 July 2026. It was the first time a frontier model was pulled from general availability by government order and then restored.
  • Opus 5 (24 July) arrived at $5 / $25 with 1M context, adaptive thinking by default and a five-level effort setting.
  • Sonnet 5.5 introduced cyber-safety fallbacks (flagged security requests fall back to an older model), the same pattern Fable used.

Strengths: Agentic coding, terminal and computer use, long-context coherence, writing quality, prompt-injection resistance. Trade-offs: No open weights; Claude subscriptions (Pro/Max) are for Anthropic's own apps — third-party harnesses generally need API billing.


OpenAI — GPT

Positioning: Largest consumer footprint (ChatGPT) and broadest product surface: ChatGPT, ChatGPT Work (multi-hour agent), Codex (coding agent), the Responses API, realtime voice, GPT Image, Sora. Distributed through its own API and Azure / Microsoft Foundry, among others.

Naming changed in 2026. OpenAI moved to "generation number + tier name": the number is the generation, the name is the job.

Model Released Role API price (in / out per 1M)
GPT-6 Astra 3–4 Sep 2026 Flagship $10 / $50
GPT-6 Sol 22 Sep 2026 Mid tier (default for coding agents) $2 / $10
GPT-6 Luna 22 Sep 2026 Cheap / high-volume $0.10 / $0.50
GPT-5.6 Sol / Terra / Luna 9 Jul 2026 Previous generation Sol launched at $5 / $30
gpt-oss-120b / gpt-oss-20b Aug 2025 Open-weight reasoning models Self-host

Watch the naming trap: GPT-5.6 Sol was a flagship; GPT-6 Sol is the mid tier (its price sits where GPT-5.6 Terra was). There is no GPT-6 Terra. When a benchmark or price just says "Sol," check the generation.

Notable 2026 events:

  • GPT-5.6 first shipped on 26 June to about twenty government-vetted organisations and went broad only after a Commerce Department review.
  • GPT-6 Astra was likewise rolled out first to a limited set of organisations, positioned as able to operate a whole computer autonomously, and released with a mechanism to shut it down if it is used beyond its intended scope.
  • GPT-5.6 shared a 1M-token context window, 128K max output and a February 2026 knowledge cutoff across all three tiers.

Strengths: Breadth of product, tool orchestration, strong maths/science, huge ecosystem of OpenAI-compatible tooling. Trade-offs: Naming churn; vendor benchmark claims have been disputed (independent evaluators flagged benchmark gaming on one GPT-5.6 claim).


Google DeepMind — Gemini

Positioning: The most vertically integrated player — own chips (TPUs), own cloud (Vertex AI), and distribution through Search, Android, Workspace and the Gemini app (900M+ monthly users as of I/O 2026). Developer surfaces: Google AI Studio, the Gemini API, Vertex AI, and the Antigravity agentic IDE.

Current lineup:

Model Released Role Price (in / out per 1M)
Gemini 3.1 Pro Feb 2026 Current Pro flagship See pricing page
Gemini 3.8 Flash 2 Sep 2026 Default workhorse $0.75 / $3.75 introductory through 31 Dec 2026, then $1.50 / $7.50
Gemini 3.8 Flash Cyber 2 Sep 2026 Vulnerability finding/fixing Gated — Fairwind Program only
Gemini 3.5 Flash-Lite Jul 2026 Cheapest tier —
Gemma 4 Apr 2026 Open-weight family Self-host (e.g. Gemma 4 12B runs on a 16GB laptop)

The story of 2026: Google has shipped Flash models at a very fast cadence (3.6, 3.7 and 3.8 Flash within about six weeks) while Gemini 3.5 Pro was repeatedly delayed. Google confirmed Gemini 4 pre-training started on 21 July 2026; in late September DeepMind leadership said it was in post-training and would ship well before the end of the year. As of this writing it is not yet released.

Strengths: Price/performance on Flash, native multimodality (video, audio), 1M context, Google ecosystem integration. Trade-offs: Flagship (Pro) tier has lagged competitors in 2026; Flash 3.8 uses noticeably more output tokens per task than 3.7, so per-token price isn't per-task price.


xAI (now part of SpaceX) — Grok

Positioning: Consumer chat integrated with X, plus Tesla in-car assistant; API available directly and via Microsoft Foundry. xAI now operates under SpaceX ("SpaceXAI").

Lineup: Grok 4.6 (around 12 Aug 2026; 500K context, configurable reasoning, aimed at coding, agentic tasks and knowledge work) and Grok 4.7, which Musk said was trained with additional SpaceX engineering data and claimed at ~2.1T parameters; it was reported as released on 21 September 2026. Grok 5 is positioned as the next generational jump. Separate media models include Grok Imagine (video) and Grok Voice.

Strengths: Real-time X data, aggressive cost positioning, voice. Trade-offs: Release dates and specs frequently come from founder posts before official documentation; verify against xAI's own model list.


Meta — Muse (formerly Llama)

Positioning: Meta reset its AI strategy after Llama 4 was poorly received. Meta Superintelligence Labs (led by Alexandr Wang) released Muse Spark on 8 April 2026 — a natively multimodal reasoning model that Meta claims matches Llama 4 Maverick capability with over 10× less compute. Muse Spark is proprietary (a break from Llama's open weights) and powers Meta AI across Facebook, Instagram, WhatsApp and Meta's glasses.

Meta has continued releasing open models alongside: Muse Glimmer (late Aug 2026) is a 30B open-weight multimodal model aimed at local agents, coding and private visual tasks.

Strengths: Distribution to billions of users; efficient small models. Trade-offs: Flagship is closed; Muse Spark initially US-only and tied to Meta accounts.


Others worth knowing

  • Microsoft — Major distributor (Azure / Microsoft Foundry hosts OpenAI, Anthropic, xAI, Mistral and open models) and builds its own smaller models.
  • Amazon — Bedrock hosts Anthropic, OpenAI open-weight and third-party models; also invests heavily in OpenAI and Anthropic and builds Nova models and Trainium chips.
  • Apple — Integrates Gemini into Siri alongside on-device models.
  • Cohere, AI21, Reka — Enterprise-focused labs (RAG, multilingual, private deployment).

5. The open-weight labs

Lab Country Model (2026) Size Licence Notes
DeepSeek China V4-Pro, V4-Flash (preview 24 Apr 2026) Pro 1.6T / 49B active; Flash 284B / 13B active MIT 1M context; sparse/compressed attention cuts per-token compute and KV cache sharply vs V3.2; API speaks both OpenAI and Anthropic protocols. Flash ~$0.14 / $0.28, Pro ~$1.74 / $3.48 on the official API
Alibaba (Qwen) China Qwen 3.8 (Sep 2026) 27B dense up to 2.4T MoE multimodal Open weights on Hugging Face 27B fits on one 80GB GPU or two consumer GPUs quantised; Qwen3.8-Max is the hosted flagship
Moonshot AI (Kimi) China Kimi K3 (API 16 Jul, weights 26 Jul 2026) 2.8T / 104B active Modified MIT 1,048,576-token context; largest openly available model as of its release; strong front-end coding
Thinking Machines Lab US Inkling (15 Jul 2026) 975B / 41B active; Inkling-Small 276B / 12B active Apache 2.0 Mira Murati's lab; 1M context; pitched explicitly for enterprise fine-tuning rather than leaderboard wins
Mistral AI France Large 3 (Dec 2025), Small 4 (Mar 2026), Medium 3.5 (Apr 2026), Devstral 2 Large 3 ≈ 675B; Small 4 119B MoE Apache 2.0 for Large 3 / Small 4; Medium is API European/sovereign option; Small 4 unifies reasoning, vision and coding; Vibe coding CLI
Google US Gemma 4 (Apr 2026) Down to 12B Gemma licence Laptop-class local model built from Gemini research
OpenAI US gpt-oss-120b / 20b (Aug 2025) 120B / 20B MoE Apache 2.0 First OpenAI open weights since GPT-2
Meta US Muse Glimmer (Aug 2026) 30B multimodal Meta licence Local multimodal agents
Z.ai (GLM), MiniMax China GLM, MiniMax M-series Large MoE Mostly open Competitive coding models; listed companies whose share prices moved sharply on Kimi K3's release

Why the Chinese labs matter: DeepSeek, Qwen and Kimi consistently publish frontier-adjacent weights at very low API prices, exerting strong downward pressure on proprietary token pricing. Many enterprises now split workloads: routine high-volume reasoning on self-hosted or third-party-hosted open models, closed APIs for the hardest or most ambiguous work.


6. Pricing comparison

Standard list prices, USD per million tokens, short context, no batch/caching discounts. Verify before use — promotional rates and long-context surcharges are common.

Tier Model Input Output
Flagship GPT-6 Astra $10.00 $50.00
Flagship Claude Opus 5.5 $4.00 $20.00
Mid Claude Sonnet 5.5 $2.00 $10.00
Mid GPT-6 Sol $2.00 $10.00
Fast Gemini 3.8 Flash (intro to 31 Dec 2026) $0.75 $3.75
Open, hosted DeepSeek V4-Pro ~$1.74 ~$3.48
Open, hosted DeepSeek V4-Flash ~$0.14 ~$0.28
Cheapest GPT-6 Luna $0.10 $0.50

Things the sticker price hides:

  • Tokens per task varies widely. A model that thinks longer can cost more per task despite a lower per-token rate (Gemini 3.8 Flash uses ~30% more output tokens per task than 3.7 Flash at the same price). Anthropic and OpenAI now market "cost per task" improvements rather than per-token cuts.
  • Prompt caching typically cuts repeated-prefix input cost by ~90%. For agents, cache hit rate often matters more than the list price.
  • Batch / flex tiers are usually ~50% cheaper for non-urgent work.
  • Fast modes cost a premium (e.g. Opus 5.5 Fast mode is $8 / $40).
  • Long-context surcharges apply above certain thresholds on some providers.

7. How the models actually differ

7.1 Tiering is now universal

Tier Anthropic OpenAI Google Typical use
Top / gated Mythos 5.1 / Fable 5.1 GPT-6 Astra (Gemini 4, pending) Hardest research, long autonomous runs
Flagship Opus 5.5 GPT-6 Astra Gemini 3.1 Pro Complex agentic work
Balanced Sonnet 5.5 GPT-6 Sol Gemini 3.8 Flash Everyday coding and documents
Fast/cheap Haiku 5.5 GPT-6 Luna Gemini 3.5 Flash-Lite Classification, extraction, sub-agents, titles

A common production pattern: flagship for planning, a cheaper tier for the bulk of tool calls.

7.2 Dimensions that matter in practice

Dimension What varies How to evaluate
Agentic reliability How long a model can work unsupervised without derailing; tool-call accuracy Run your own multi-step tasks; Terminal-Bench, SWE-Bench Pro, OSWorld as rough guides
Reasoning effort controls Levels offered, default effort, whether thinking is visible Test at the effort level you'll actually pay for
Context Advertised window vs effective recall; cache pricing Needle-in-haystack is not enough — test real long documents
Modalities Image/audio/video input; whether output is text-only Match to workload
Latency & throughput Time-to-first-token, tokens/sec, fast modes Matters most for interactive and voice use
Safety behaviour Refusal rate, classifier fallbacks, gated domains (cyber, bio) Security research workloads may hit fallbacks or need a gated programme
Data handling Retention, training on your data, regional hosting Read enterprise terms; consider cloud-marketplace hosting for data residency
Writing style Verbosity, formatting habits, tone Subjective — sample outputs
Ecosystem SDKs, agent frameworks, IDE integrations Often decisive for teams

7.3 Lab "personalities" (generalisations)

  • Anthropic: Coding and agentic depth; careful, high-quality writing; heavy investment in safety classifiers and prompt-injection defence.
  • OpenAI: Breadth of products and modalities; strong maths/science; fastest to brand new product categories.
  • Google: Price/performance and multimodal; distribution advantage; flagship cadence slower than its Flash cadence in 2026.
  • xAI: Real-time social data, voice, cost-aggressive; announcements often precede docs.
  • Meta: Efficiency and consumer reach; strategic retreat from fully open flagships.
  • Chinese open labs: Frontier-adjacent open weights at very low prices; MoE and attention-efficiency innovation.
  • Mistral: European sovereignty, Apache-licensed open models, enterprise on-prem.

8. Running models yourself

8.1 Inference engines

Tool Best for Notes
Ollama Easiest local setup on Linux/macOS/Windows Pulls quantised models; OpenAI-compatible API on :11434
llama.cpp CPU/GPU hybrid, Apple Silicon, GGUF models Underpins many desktop tools
LM Studio Desktop GUI OpenAI-compatible server on :1234
vLLM Production GPU serving, high throughput PagedAttention, tensor parallelism, OpenAI-compatible server on :8000
SGLang High-throughput serving, structured output Strong on MoE models
TensorRT-LLM / NVIDIA NIM Maximum NVIDIA performance More setup
Hugging Face TGI HF-ecosystem deployments

8.2 Quantisation

Weights are stored at reduced precision to save memory:

Format Typical use
FP16 / BF16 Full quality, 2 bytes per parameter
FP8 Near-lossless on modern GPUs, 1 byte per parameter
4-bit (GGUF Q4_K_M, AWQ, GPTQ, EXL2) Consumer hardware; ~0.5–0.6 bytes per parameter; small quality loss

8.3 Rough VRAM math

weights_GB ≈ total_params_in_billions × bytes_per_param
total_GB   ≈ weights_GB + KV_cache + ~10–20% overhead
Model class 4-bit weights Realistic hardware
8–12B (Gemma 4 12B, Qwen 8B) ~6–8GB 12–16GB GPU or 16GB laptop
27–32B (Qwen 3.8 27B, Muse Glimmer 30B) ~18–24GB 24GB GPU, 2×16GB, or 32GB+ unified memory
120B (gpt-oss-120b) ~60–70GB 80GB GPU or 2–4 consumer GPUs
1T+ MoE (DeepSeek V4-Pro, Kimi K3) Hundreds of GB to TB+ Multi-node datacentre; use a hosted provider

Remember: MoE active parameters reduce compute, not memory. All experts must be resident.

8.4 Minimal local stack

# Ollama
ollama serve &
ollama pull qwen3:8b
curl http://127.0.0.1:11434/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3:8b","messages":[{"role":"user","content":"hello"}]}'

# vLLM (GPU host) — OpenAI-compatible, tool calling on
vllm serve Qwen/Qwen3-32B \
  --tensor-parallel-size 2 \
  --max-model-len 32768 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

Because both expose the OpenAI API shape, almost every agent and SDK can point at them by changing base_url.


9. Access, gateways and routing

Ways to reach a model

  1. First-party API — Anthropic, OpenAI, Google AI Studio/Gemini API, xAI, Mistral, DeepSeek. Newest models land here first.
  2. Cloud marketplaces — AWS Bedrock, Google Vertex AI, Microsoft Foundry. Best for enterprise billing, IAM, private networking and data residency. Claude, for instance, is available on all three.
  3. Third-party inference hosts — Together, Fireworks, Groq, Cerebras, DeepInfra, Baseten etc. Mostly open models; compete on speed and price.
  4. Gateways / routers — OpenRouter (one key, hundreds of models), LiteLLM (self-hosted proxy), cloud AI gateways. Useful for fallback, cost tracking and A/B testing.
  5. Self-hosted — vLLM/SGLang on your own GPUs.

The OpenAI-compatible API as lingua franca

The /v1/chat/completions (and increasingly /v1/responses) shape is the de facto standard. Many providers — including DeepSeek — also speak the Anthropic Messages protocol. Design your code around an abstraction (an SDK like the Vercel AI SDK, LiteLLM, or your own thin interface) so a model swap is a config change.

Subscriptions vs API

Consumer subscriptions (ChatGPT Plus/Pro, Claude Pro/Max, Google AI Pro/Ultra) are priced for use inside the vendor's own apps. Using them through third-party tools is generally restricted — for example, Claude subscriptions are no longer usable in third-party agent harnesses, which need API billing. Some third-party tools sell their own subscriptions (e.g. OpenCode Go) that bundle access to selected models.


10. Agents, protocols and coding tools

Protocols

Protocol Purpose
MCP (Model Context Protocol) Standard way to expose tools, data sources and prompts to any model/agent. Local (stdio) or remote (HTTP + OAuth). Supported by essentially every major agent and lab
ACP (Agent Client Protocol) Lets editors talk to any coding agent over stdio/JSON-RPC (so one agent can plug into many editors)
A2A (Agent-to-Agent) Google-originated protocol for agents to discover and delegate to each other
Agent Skills Folders of instructions + scripts (SKILL.md) loaded on demand; adopted across several agents
AGENTS.md Convention for project-level instructions to coding agents (build commands, conventions)

Coding agents (2026)

Agent Vendor Models Interface
Claude Code Anthropic Claude Terminal, IDE, desktop, web
Codex OpenAI GPT Terminal, IDE, cloud, ChatGPT
Antigravity Google Gemini (default) Agentic IDE
Cursor Anysphere Multi-model IDE
GitHub Copilot GitHub/Microsoft Multi-model IDE, CLI, cloud agent
Kiro AWS Multi-model incl. Claude IDE, CLI
OpenCode Anomaly (ex-SST) Any (75+ providers, local) Terminal, desktop, web — open source
Mistral Vibe Mistral Mistral (Devstral/Medium) CLI, remote agents
Cline, Roo, Aider, Goose Open source Any IDE / terminal

The trade-off is the same as with models: vendor agents are tuned end-to-end for their own model; model-agnostic agents (OpenCode, Cline, Aider) give you portability, local-model support and inspectable code at the cost of doing more integration work yourself.

Agent patterns worth knowing

  • Plan → build split: A read-only planning mode followed by an editing mode.
  • Sub-agents: Delegate searches or parallel work to cheaper models in isolated contexts.
  • Context compaction: Summarise old history when the window fills.
  • Permission systems: Allow/ask/deny rules per tool and per command pattern.
  • Sandboxing: Run agent shell commands in containers or VMs; treat any content the agent reads (web pages, issues, docs) as untrusted — prompt injection is the main security risk for agents.

11. Regulation and safety gating

2026 is the year regulation became part of the release pipeline:

  • Export controls: Anthropic's Fable 5 / Mythos 5 suspension (12 June – 1 July 2026) to comply with US Commerce Department export controls.
  • Pre-release federal review: Frontier releases from OpenAI went to small sets of vetted organisations before broad release following Commerce review. Reporting describes a US executive order (effective June 2026) establishing a de facto pre-release review of up to 30 days.
  • Gated cyber models: OpenAI (GPT-5.5-Cyber), Google (Gemini 3.8 Flash Cyber via the Fairwind Program), and Anthropic's restricted Mythos tier all limit the most capable security models to vetted defenders.
  • Classifier fallbacks: Public models increasingly route flagged cyber/bio requests to an older or more restricted model rather than refusing outright.
  • EU AI Act: Obligations for general-purpose AI model providers are phasing in; relevant to Mistral and anyone deploying in the EU.

Practical implication: Release dates slip, access can be revoked or staged, and security-research workloads may need a specific programme or an open-weight model.


12. Benchmarks — and why to distrust them

Common benchmarks in 2026:

Benchmark Measures
SWE-Bench Verified / Pro Resolving real GitHub issues
Terminal-Bench Multi-step terminal tasks
OSWorld(-Verified) Computer use in a desktop OS
DeepSWE Long-horizon software engineering
Humanity's Last Exam (HLE) Hard expert-level questions
GDPval Economically valuable knowledge work
LMArena (Elo) Blind human preference
Artificial Analysis indices Aggregated intelligence/coding scores, speed, price

Caveats:

  • Vendor-reported numbers use the vendor's own harness and effort settings. Different labs rank differently on different benchmarks — e.g. OpenAI's claimed lead for GPT-5.6 Sol on one coding index inverted on SWE-Bench Pro, where Fable 5 scored 80% vs Sol's 64.6%.
  • Benchmark gaming and contamination are real; independent evaluators (METR, Artificial Analysis, Epoch) matter.
  • Your workload is the only benchmark that counts. Keep a small private eval set of real tasks and re-run it whenever you consider switching.

13. Choosing a model: a decision guide

Is the data allowed to leave your environment?
├── No  → open weights, self-hosted (Qwen 3.8 27B, Muse Glimmer, gpt-oss,
│         Gemma 4, Mistral Small 4) or a cloud marketplace in your region
└── Yes
    ├── Hardest long-horizon/agentic work?      → Fable 5.1, Opus 5.5, GPT-6 Astra
    ├── Everyday coding & documents?            → Sonnet 5.5, GPT-6 Sol, Gemini 3.8 Flash
    ├── High volume / sub-agents / extraction?  → Haiku 5.5, GPT-6 Luna, Flash-Lite,
    │                                             DeepSeek V4-Flash
    ├── Heavy multimodal (video/audio in)?      → Gemini
    ├── Cheapest strong open model via API?     → DeepSeek V4, Kimi K3, Qwen3.8-Max
    └── EU sovereignty requirements?            → Mistral (API or self-hosted)

Practical rules:

  1. Abstract the provider. Use an OpenAI-compatible layer or SDK so switching is a config change.
  2. Route by task. Expensive model for planning/judgment, cheap model for volume.
  3. Measure cost per task, not per token. Include thinking tokens and cache hits.
  4. Keep a private eval set. Re-run it on every new release before switching.
  5. Pin model versions in production; use latest aliases only in development.
  6. Plan for availability shocks — regulatory suspensions, rate limits, deprecations. Have a fallback model configured.

14. Glossary

Term Meaning
Active parameters Parameters actually used per token in an MoE model
Adaptive thinking Model decides how much to reason per request
Context window Max tokens (input + output) per request
Distillation Training a small model to mimic a large one
Effort level User-selectable reasoning budget
Harness The software around a model: prompts, tools, loop, context management
KV cache Stored attention keys/values for previous tokens; grows with context
MoE Mixture of Experts — sparse model architecture
Open weights Downloadable model parameters (training data usually not included)
Prompt caching Discounted reuse of an unchanged prompt prefix
Quantisation Storing weights at lower precision to save memory
RLVR Reinforcement learning with verifiable rewards (tests, maths answers)
Sub-agent A delegated agent with its own context, usually cheaper
System prompt Developer instructions that frame model behaviour
Tool call Structured request from the model to run a function

15. Sources