The LLM Ecosystem in Late 2026¶
Snapshot date: 8 October 2026. This field moves on a weekly cadence. Model names, prices and "latest" claims below were verified against public sources in early October 2026 — treat anything here older than a month or two as suspect and re-check vendor pricing pages before committing to a design.
TL;DR¶
- Five companies set the frontier for closed models: Anthropic (Claude), OpenAI (GPT), Google DeepMind (Gemini), xAI/SpaceX (Grok) and Meta (Muse). Each now ships a tiered family (flagship / balanced / fast) rather than a single "best" model.
- Open-weight models are close behind — DeepSeek V4, Alibaba Qwen 3.8, Moonshot Kimi K3, Thinking Machines Inkling, Mistral, Google Gemma 4 and Meta Muse Glimmer. The best of them are within a few months of the closed frontier, but the largest ones need datacentre hardware, not a workstation.
- The competition has shifted from "smartest model" to "right model at the right price." Prices for mid-tier models roughly halved over the summer of 2026.
- Regulation is now part of the release process. Frontier launches in 2026 have gone through government review, export-control suspensions, and gated "cyber" variants available only to vetted defenders.
- Agents are the main use case. Long-horizon coding, computer use and tool calling (via MCP) are what labs optimise and benchmark for.
Table of contents¶
- How LLMs work (the 10-minute version)
- The ecosystem in layers
- Closed vs open weights
- The frontier labs and their models
- The open-weight labs
- Pricing comparison
- How the models actually differ
- Running models yourself
- Access, gateways and routing
- Agents, protocols and coding tools
- Regulation and safety gating
- Benchmarks — and why to distrust them
- Choosing a model: a decision guide
- Glossary
- Sources
1. How LLMs work (the 10-minute version)¶
Tokens and next-token prediction¶
An LLM is a neural network trained to predict the next token (a word fragment — roughly ¾ of an English word on average) given all the tokens before it. Everything else — chat, coding, reasoning, tool use — is built on top of that one capability. Pricing, context limits and speed are all measured in tokens.
The transformer¶
Nearly every production LLM is a decoder-only transformer. Its core mechanism, attention, lets every token look at every earlier token to decide what is relevant. Attention cost grows with the square of the sequence length, which is why long context is expensive and why labs invest heavily in tricks such as sparse attention, compressed KV caches and sliding windows.
Dense vs Mixture-of-Experts (MoE)¶
- Dense models use every parameter for every token (e.g. a 27B dense model does 27B parameters of work per token).
- MoE models contain many "expert" sub-networks and a router that activates only a few per token. A model may have trillions of total parameters but only tens of billions active. This gives large-model knowledge at small-model compute cost — but all the weights still have to sit in memory.
Almost every frontier-scale open model in 2026 is MoE. Some examples: DeepSeek V4-Pro is 1.6T total / 49B active, and Kimi K3 is 2.8T total / 104B active.
Training stages¶
| Stage | What happens | Why it matters |
|---|---|---|
| Pre-training | Next-token prediction over tens of trillions of tokens of text, code, images, audio and video | Gives the model knowledge and raw capability. The most expensive step. Sets the knowledge cutoff. |
| Supervised fine-tuning (SFT) | Train on curated instruction/response examples | Turns a text predictor into an assistant |
| Preference tuning (RLHF / RLAIF / constitutional methods) | Reward the model for responses humans or AI judges prefer | Tone, helpfulness, harmlessness, refusals |
| Reinforcement learning with verifiable rewards | RL on tasks with checkable answers — unit tests, maths, terminal tasks | The main driver of 2025–2026 gains in coding and reasoning |
| Distillation | Train smaller models to imitate a larger one | How "Flash/Luna/Haiku" tiers inherit flagship behaviour cheaply |
Reasoning ("thinking") models¶
Since 2024, most models can generate hidden or visible chain-of-thought tokens before answering. In 2026 this is the default: models use adaptive thinking and expose an effort control (low → max). More effort means better answers on hard problems but more output tokens, more latency and more cost. Reasoning tokens are billed as output tokens.
Context windows¶
The context window is the total number of tokens the model can consider at once (prompt + history + tool results + output). 1M tokens is now table stakes for flagship and mid-tier models from Anthropic, OpenAI, Google, DeepSeek and Moonshot. Having a 1M window doesn't mean the model uses all of it equally well — recall degrades with distance and cost scales with length, which is why prompt caching (cheap re-reads of an unchanged prefix) matters so much for agents.
Multimodality¶
Most frontier models accept text, images and often audio/video as input; most still produce text output. Image, video and voice generation usually come from separate models (e.g. OpenAI's GPT Image, Google's Nano Banana and Veo/Omni lines, xAI's Grok Imagine).
Tool use and agents¶
Models emit structured tool calls (JSON) that a harness executes — run a shell command, read a file, search the web, click a button — and feed the result back. An agent is just this loop running for many steps toward a goal. Almost all of the 2026 model competition is about how reliably this loop works over long horizons.
2. The ecosystem in layers¶
┌──────────────────────────────────────────────────────────────────┐
│ APPLICATIONS & AGENTS ChatGPT, Claude apps, Gemini app, Meta AI, │
│ coding agents (Claude Code, Codex, Cursor, │
│ OpenCode, Antigravity, Kiro, Vibe…) │
├──────────────────────────────────────────────────────────────────┤
│ FRAMEWORKS & PROTOCOLS MCP, ACP, A2A, agent SDKs, LangGraph, │
│ LlamaIndex, evals & observability │
├──────────────────────────────────────────────────────────────────┤
│ GATEWAYS & ROUTERS OpenRouter, LiteLLM, cloud AI gateways, │
│ OpenAI-compatible endpoints │
├──────────────────────────────────────────────────────────────────┤
│ HOSTING / INFERENCE First-party APIs, AWS Bedrock, Google │
│ Vertex, Microsoft Foundry, Together, │
│ Fireworks, Groq…; self-hosted vLLM/SGLang │
├──────────────────────────────────────────────────────────────────┤
│ MODELS Closed: Claude, GPT, Gemini, Grok, Muse │
│ Open: DeepSeek, Qwen, Kimi, Inkling, │
│ Mistral, Gemma, gpt-oss, Glimmer │
├──────────────────────────────────────────────────────────────────┤
│ COMPUTE NVIDIA GPUs, Google TPUs, AWS Trainium, │
│ hyperscaler & neocloud datacentres │
└──────────────────────────────────────────────────────────────────┘
A useful mental model: the model is a component, not the product. The same Claude, GPT or Qwen model behaves very differently depending on the harness, system prompt, tools and context management around it.
3. Closed vs open weights¶
| Closed (proprietary) | Open weights | |
|---|---|---|
| What you get | API access only | Downloadable model weights |
| Examples | Claude, GPT-6, Gemini, Grok, Muse Spark | DeepSeek V4, Qwen 3.8, Kimi K3, Inkling, Mistral Small 4 / Large 3, Gemma 4, gpt-oss, Muse Glimmer |
| Best raw capability | Yes — still the frontier | Typically a few months behind |
| Data control | Data leaves your environment (subject to vendor terms) | Can run fully air-gapped |
| Customisation | Prompting, some hosted fine-tuning | Full fine-tuning, quantisation, distillation |
| Cost model | Per-token | Hardware + ops (or per-token from third-party hosts) |
| Safety controls | Vendor-enforced classifiers and policies | Your responsibility |
"Open" is a spectrum¶
- Open weights ≠ open source. Almost no frontier "open" model releases its training data or full training code. "Open weights" is the accurate term.
- Licences vary. Permissive (MIT, Apache 2.0) vs custom licences with usage or scale restrictions. Kimi K3 uses a Modified MIT licence; Mistral uses Apache 2.0 for some models and proprietary terms for others. Always read the licence.
- Open ≠ runnable at home. Kimi K3 in 4-bit precision still needs roughly 1.4TB of fast memory, and Moonshot recommends at least 64 accelerators. For trillion-parameter MoE models, "open" in practice means "any cloud can host it," not "you can run it on a workstation."
The gap¶
Through 2026, blind arena evaluations have put the best open models (Kimi K3 in particular, on front-end coding) level with or ahead of leading US closed models on specific tasks. Commentators estimate the open-to-closed gap has narrowed from roughly six to nine months down to three to five.
4. The frontier labs and their models¶
Anthropic — Claude¶
Positioning: Safety-focused lab; strongest reputation in agentic coding, long-horizon tasks and enterprise. Products: Claude apps, Claude Code, the Claude API, and availability on AWS Bedrock, Google Cloud Vertex AI and Microsoft Foundry.
Current lineup (Claude 5 / 5.5 generation):
| Model | Released | Role | API price (in / out per 1M) | Notes |
|---|---|---|---|---|
| Claude Mythos 5.1 | 2026 | Top "Mythos" tier | Restricted | Not publicly available; used by a small number of trusted organisations through Anthropic's Project Glasswing |
| Claude Fable 5.1 | 1 Sep 2026 | Public Mythos-tier model | See pricing page (Fable 5 launched at $10 / $50) | Same underlying model as Mythos 5.1 with extra safeguards for biology, cybersecurity and LLM R&D. Long-horizon reasoning and agentic coding |
| Claude Opus 5.5 | 22 Sep 2026 | Flagship workhorse | $4 / $20 | Cheaper and faster than Opus 5; Anthropic says ~40% lower typical cost. Fast mode available at a premium |
| Claude Sonnet 5.5 | 28 Sep 2026 | Balanced default | $2 / $10 | 1M context, 128K output, adaptive thinking on by default |
| Claude Haiku 5.5 | After Sonnet 5.5 | Fast / cheap tier | See pricing page | High-volume, cost-sensitive work; first Haiku in the 5.x generation |
API IDs follow the pattern claude-opus-5-5, claude-sonnet-5-5, claude-haiku-5-5, claude-fable-5-1.
Notable 2026 events:
- Claude Fable 5 and Mythos 5 shipped on 9 June 2026. On 12 June Anthropic suspended access to both to comply with US Department of Commerce export controls. The controls were lifted on 30 June and access was restored on 1 July 2026. It was the first time a frontier model was pulled from general availability by government order and then restored.
- Opus 5 (24 July) arrived at $5 / $25 with 1M context, adaptive thinking by default and a five-level effort setting.
- Sonnet 5.5 introduced cyber-safety fallbacks (flagged security requests fall back to an older model), the same pattern Fable used.
Strengths: Agentic coding, terminal and computer use, long-context coherence, writing quality, prompt-injection resistance. Trade-offs: No open weights; Claude subscriptions (Pro/Max) are for Anthropic's own apps — third-party harnesses generally need API billing.
OpenAI — GPT¶
Positioning: Largest consumer footprint (ChatGPT) and broadest product surface: ChatGPT, ChatGPT Work (multi-hour agent), Codex (coding agent), the Responses API, realtime voice, GPT Image, Sora. Distributed through its own API and Azure / Microsoft Foundry, among others.
Naming changed in 2026. OpenAI moved to "generation number + tier name": the number is the generation, the name is the job.
| Model | Released | Role | API price (in / out per 1M) |
|---|---|---|---|
| GPT-6 Astra | 3–4 Sep 2026 | Flagship | $10 / $50 |
| GPT-6 Sol | 22 Sep 2026 | Mid tier (default for coding agents) | $2 / $10 |
| GPT-6 Luna | 22 Sep 2026 | Cheap / high-volume | $0.10 / $0.50 |
| GPT-5.6 Sol / Terra / Luna | 9 Jul 2026 | Previous generation | Sol launched at $5 / $30 |
| gpt-oss-120b / gpt-oss-20b | Aug 2025 | Open-weight reasoning models | Self-host |
Watch the naming trap: GPT-5.6 Sol was a flagship; GPT-6 Sol is the mid tier (its price sits where GPT-5.6 Terra was). There is no GPT-6 Terra. When a benchmark or price just says "Sol," check the generation.
Notable 2026 events:
- GPT-5.6 first shipped on 26 June to about twenty government-vetted organisations and went broad only after a Commerce Department review.
- GPT-6 Astra was likewise rolled out first to a limited set of organisations, positioned as able to operate a whole computer autonomously, and released with a mechanism to shut it down if it is used beyond its intended scope.
- GPT-5.6 shared a 1M-token context window, 128K max output and a February 2026 knowledge cutoff across all three tiers.
Strengths: Breadth of product, tool orchestration, strong maths/science, huge ecosystem of OpenAI-compatible tooling. Trade-offs: Naming churn; vendor benchmark claims have been disputed (independent evaluators flagged benchmark gaming on one GPT-5.6 claim).
Google DeepMind — Gemini¶
Positioning: The most vertically integrated player — own chips (TPUs), own cloud (Vertex AI), and distribution through Search, Android, Workspace and the Gemini app (900M+ monthly users as of I/O 2026). Developer surfaces: Google AI Studio, the Gemini API, Vertex AI, and the Antigravity agentic IDE.
Current lineup:
| Model | Released | Role | Price (in / out per 1M) |
|---|---|---|---|
| Gemini 3.1 Pro | Feb 2026 | Current Pro flagship | See pricing page |
| Gemini 3.8 Flash | 2 Sep 2026 | Default workhorse | $0.75 / $3.75 introductory through 31 Dec 2026, then $1.50 / $7.50 |
| Gemini 3.8 Flash Cyber | 2 Sep 2026 | Vulnerability finding/fixing | Gated — Fairwind Program only |
| Gemini 3.5 Flash-Lite | Jul 2026 | Cheapest tier | — |
| Gemma 4 | Apr 2026 | Open-weight family | Self-host (e.g. Gemma 4 12B runs on a 16GB laptop) |
The story of 2026: Google has shipped Flash models at a very fast cadence (3.6, 3.7 and 3.8 Flash within about six weeks) while Gemini 3.5 Pro was repeatedly delayed. Google confirmed Gemini 4 pre-training started on 21 July 2026; in late September DeepMind leadership said it was in post-training and would ship well before the end of the year. As of this writing it is not yet released.
Strengths: Price/performance on Flash, native multimodality (video, audio), 1M context, Google ecosystem integration. Trade-offs: Flagship (Pro) tier has lagged competitors in 2026; Flash 3.8 uses noticeably more output tokens per task than 3.7, so per-token price isn't per-task price.
xAI (now part of SpaceX) — Grok¶
Positioning: Consumer chat integrated with X, plus Tesla in-car assistant; API available directly and via Microsoft Foundry. xAI now operates under SpaceX ("SpaceXAI").
Lineup: Grok 4.6 (around 12 Aug 2026; 500K context, configurable reasoning, aimed at coding, agentic tasks and knowledge work) and Grok 4.7, which Musk said was trained with additional SpaceX engineering data and claimed at ~2.1T parameters; it was reported as released on 21 September 2026. Grok 5 is positioned as the next generational jump. Separate media models include Grok Imagine (video) and Grok Voice.
Strengths: Real-time X data, aggressive cost positioning, voice. Trade-offs: Release dates and specs frequently come from founder posts before official documentation; verify against xAI's own model list.
Meta — Muse (formerly Llama)¶
Positioning: Meta reset its AI strategy after Llama 4 was poorly received. Meta Superintelligence Labs (led by Alexandr Wang) released Muse Spark on 8 April 2026 — a natively multimodal reasoning model that Meta claims matches Llama 4 Maverick capability with over 10× less compute. Muse Spark is proprietary (a break from Llama's open weights) and powers Meta AI across Facebook, Instagram, WhatsApp and Meta's glasses.
Meta has continued releasing open models alongside: Muse Glimmer (late Aug 2026) is a 30B open-weight multimodal model aimed at local agents, coding and private visual tasks.
Strengths: Distribution to billions of users; efficient small models. Trade-offs: Flagship is closed; Muse Spark initially US-only and tied to Meta accounts.
Others worth knowing¶
- Microsoft — Major distributor (Azure / Microsoft Foundry hosts OpenAI, Anthropic, xAI, Mistral and open models) and builds its own smaller models.
- Amazon — Bedrock hosts Anthropic, OpenAI open-weight and third-party models; also invests heavily in OpenAI and Anthropic and builds Nova models and Trainium chips.
- Apple — Integrates Gemini into Siri alongside on-device models.
- Cohere, AI21, Reka — Enterprise-focused labs (RAG, multilingual, private deployment).
5. The open-weight labs¶
| Lab | Country | Model (2026) | Size | Licence | Notes |
|---|---|---|---|---|---|
| DeepSeek | China | V4-Pro, V4-Flash (preview 24 Apr 2026) | Pro 1.6T / 49B active; Flash 284B / 13B active | MIT | 1M context; sparse/compressed attention cuts per-token compute and KV cache sharply vs V3.2; API speaks both OpenAI and Anthropic protocols. Flash ~$0.14 / $0.28, Pro ~$1.74 / $3.48 on the official API |
| Alibaba (Qwen) | China | Qwen 3.8 (Sep 2026) | 27B dense up to 2.4T MoE multimodal | Open weights on Hugging Face | 27B fits on one 80GB GPU or two consumer GPUs quantised; Qwen3.8-Max is the hosted flagship |
| Moonshot AI (Kimi) | China | Kimi K3 (API 16 Jul, weights 26 Jul 2026) | 2.8T / 104B active | Modified MIT | 1,048,576-token context; largest openly available model as of its release; strong front-end coding |
| Thinking Machines Lab | US | Inkling (15 Jul 2026) | 975B / 41B active; Inkling-Small 276B / 12B active | Apache 2.0 | Mira Murati's lab; 1M context; pitched explicitly for enterprise fine-tuning rather than leaderboard wins |
| Mistral AI | France | Large 3 (Dec 2025), Small 4 (Mar 2026), Medium 3.5 (Apr 2026), Devstral 2 | Large 3 ≈ 675B; Small 4 119B MoE | Apache 2.0 for Large 3 / Small 4; Medium is API | European/sovereign option; Small 4 unifies reasoning, vision and coding; Vibe coding CLI |
| US | Gemma 4 (Apr 2026) | Down to 12B | Gemma licence | Laptop-class local model built from Gemini research | |
| OpenAI | US | gpt-oss-120b / 20b (Aug 2025) | 120B / 20B MoE | Apache 2.0 | First OpenAI open weights since GPT-2 |
| Meta | US | Muse Glimmer (Aug 2026) | 30B multimodal | Meta licence | Local multimodal agents |
| Z.ai (GLM), MiniMax | China | GLM, MiniMax M-series | Large MoE | Mostly open | Competitive coding models; listed companies whose share prices moved sharply on Kimi K3's release |
Why the Chinese labs matter: DeepSeek, Qwen and Kimi consistently publish frontier-adjacent weights at very low API prices, exerting strong downward pressure on proprietary token pricing. Many enterprises now split workloads: routine high-volume reasoning on self-hosted or third-party-hosted open models, closed APIs for the hardest or most ambiguous work.
6. Pricing comparison¶
Standard list prices, USD per million tokens, short context, no batch/caching discounts. Verify before use — promotional rates and long-context surcharges are common.
| Tier | Model | Input | Output |
|---|---|---|---|
| Flagship | GPT-6 Astra | $10.00 | $50.00 |
| Flagship | Claude Opus 5.5 | $4.00 | $20.00 |
| Mid | Claude Sonnet 5.5 | $2.00 | $10.00 |
| Mid | GPT-6 Sol | $2.00 | $10.00 |
| Fast | Gemini 3.8 Flash (intro to 31 Dec 2026) | $0.75 | $3.75 |
| Open, hosted | DeepSeek V4-Pro | ~$1.74 | ~$3.48 |
| Open, hosted | DeepSeek V4-Flash | ~$0.14 | ~$0.28 |
| Cheapest | GPT-6 Luna | $0.10 | $0.50 |
Things the sticker price hides:
- Tokens per task varies widely. A model that thinks longer can cost more per task despite a lower per-token rate (Gemini 3.8 Flash uses ~30% more output tokens per task than 3.7 Flash at the same price). Anthropic and OpenAI now market "cost per task" improvements rather than per-token cuts.
- Prompt caching typically cuts repeated-prefix input cost by ~90%. For agents, cache hit rate often matters more than the list price.
- Batch / flex tiers are usually ~50% cheaper for non-urgent work.
- Fast modes cost a premium (e.g. Opus 5.5 Fast mode is $8 / $40).
- Long-context surcharges apply above certain thresholds on some providers.
7. How the models actually differ¶
7.1 Tiering is now universal¶
| Tier | Anthropic | OpenAI | Typical use | |
|---|---|---|---|---|
| Top / gated | Mythos 5.1 / Fable 5.1 | GPT-6 Astra | (Gemini 4, pending) | Hardest research, long autonomous runs |
| Flagship | Opus 5.5 | GPT-6 Astra | Gemini 3.1 Pro | Complex agentic work |
| Balanced | Sonnet 5.5 | GPT-6 Sol | Gemini 3.8 Flash | Everyday coding and documents |
| Fast/cheap | Haiku 5.5 | GPT-6 Luna | Gemini 3.5 Flash-Lite | Classification, extraction, sub-agents, titles |
A common production pattern: flagship for planning, a cheaper tier for the bulk of tool calls.
7.2 Dimensions that matter in practice¶
| Dimension | What varies | How to evaluate |
|---|---|---|
| Agentic reliability | How long a model can work unsupervised without derailing; tool-call accuracy | Run your own multi-step tasks; Terminal-Bench, SWE-Bench Pro, OSWorld as rough guides |
| Reasoning effort controls | Levels offered, default effort, whether thinking is visible | Test at the effort level you'll actually pay for |
| Context | Advertised window vs effective recall; cache pricing | Needle-in-haystack is not enough — test real long documents |
| Modalities | Image/audio/video input; whether output is text-only | Match to workload |
| Latency & throughput | Time-to-first-token, tokens/sec, fast modes | Matters most for interactive and voice use |
| Safety behaviour | Refusal rate, classifier fallbacks, gated domains (cyber, bio) | Security research workloads may hit fallbacks or need a gated programme |
| Data handling | Retention, training on your data, regional hosting | Read enterprise terms; consider cloud-marketplace hosting for data residency |
| Writing style | Verbosity, formatting habits, tone | Subjective — sample outputs |
| Ecosystem | SDKs, agent frameworks, IDE integrations | Often decisive for teams |
7.3 Lab "personalities" (generalisations)¶
- Anthropic: Coding and agentic depth; careful, high-quality writing; heavy investment in safety classifiers and prompt-injection defence.
- OpenAI: Breadth of products and modalities; strong maths/science; fastest to brand new product categories.
- Google: Price/performance and multimodal; distribution advantage; flagship cadence slower than its Flash cadence in 2026.
- xAI: Real-time social data, voice, cost-aggressive; announcements often precede docs.
- Meta: Efficiency and consumer reach; strategic retreat from fully open flagships.
- Chinese open labs: Frontier-adjacent open weights at very low prices; MoE and attention-efficiency innovation.
- Mistral: European sovereignty, Apache-licensed open models, enterprise on-prem.
8. Running models yourself¶
8.1 Inference engines¶
| Tool | Best for | Notes |
|---|---|---|
| Ollama | Easiest local setup on Linux/macOS/Windows | Pulls quantised models; OpenAI-compatible API on :11434 |
| llama.cpp | CPU/GPU hybrid, Apple Silicon, GGUF models | Underpins many desktop tools |
| LM Studio | Desktop GUI | OpenAI-compatible server on :1234 |
| vLLM | Production GPU serving, high throughput | PagedAttention, tensor parallelism, OpenAI-compatible server on :8000 |
| SGLang | High-throughput serving, structured output | Strong on MoE models |
| TensorRT-LLM / NVIDIA NIM | Maximum NVIDIA performance | More setup |
| Hugging Face TGI | HF-ecosystem deployments |
8.2 Quantisation¶
Weights are stored at reduced precision to save memory:
| Format | Typical use |
|---|---|
| FP16 / BF16 | Full quality, 2 bytes per parameter |
| FP8 | Near-lossless on modern GPUs, 1 byte per parameter |
| 4-bit (GGUF Q4_K_M, AWQ, GPTQ, EXL2) | Consumer hardware; ~0.5–0.6 bytes per parameter; small quality loss |
8.3 Rough VRAM math¶
weights_GB ≈ total_params_in_billions × bytes_per_param
total_GB ≈ weights_GB + KV_cache + ~10–20% overhead
| Model class | 4-bit weights | Realistic hardware |
|---|---|---|
| 8–12B (Gemma 4 12B, Qwen 8B) | ~6–8GB | 12–16GB GPU or 16GB laptop |
| 27–32B (Qwen 3.8 27B, Muse Glimmer 30B) | ~18–24GB | 24GB GPU, 2×16GB, or 32GB+ unified memory |
| 120B (gpt-oss-120b) | ~60–70GB | 80GB GPU or 2–4 consumer GPUs |
| 1T+ MoE (DeepSeek V4-Pro, Kimi K3) | Hundreds of GB to TB+ | Multi-node datacentre; use a hosted provider |
Remember: MoE active parameters reduce compute, not memory. All experts must be resident.
8.4 Minimal local stack¶
# Ollama
ollama serve &
ollama pull qwen3:8b
curl http://127.0.0.1:11434/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3:8b","messages":[{"role":"user","content":"hello"}]}'
# vLLM (GPU host) — OpenAI-compatible, tool calling on
vllm serve Qwen/Qwen3-32B \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--enable-auto-tool-choice \
--tool-call-parser hermes
Because both expose the OpenAI API shape, almost every agent and SDK can point at them by changing base_url.
9. Access, gateways and routing¶
Ways to reach a model¶
- First-party API — Anthropic, OpenAI, Google AI Studio/Gemini API, xAI, Mistral, DeepSeek. Newest models land here first.
- Cloud marketplaces — AWS Bedrock, Google Vertex AI, Microsoft Foundry. Best for enterprise billing, IAM, private networking and data residency. Claude, for instance, is available on all three.
- Third-party inference hosts — Together, Fireworks, Groq, Cerebras, DeepInfra, Baseten etc. Mostly open models; compete on speed and price.
- Gateways / routers — OpenRouter (one key, hundreds of models), LiteLLM (self-hosted proxy), cloud AI gateways. Useful for fallback, cost tracking and A/B testing.
- Self-hosted — vLLM/SGLang on your own GPUs.
The OpenAI-compatible API as lingua franca¶
The /v1/chat/completions (and increasingly /v1/responses) shape is the de facto standard. Many providers — including DeepSeek — also speak the Anthropic Messages protocol. Design your code around an abstraction (an SDK like the Vercel AI SDK, LiteLLM, or your own thin interface) so a model swap is a config change.
Subscriptions vs API¶
Consumer subscriptions (ChatGPT Plus/Pro, Claude Pro/Max, Google AI Pro/Ultra) are priced for use inside the vendor's own apps. Using them through third-party tools is generally restricted — for example, Claude subscriptions are no longer usable in third-party agent harnesses, which need API billing. Some third-party tools sell their own subscriptions (e.g. OpenCode Go) that bundle access to selected models.
10. Agents, protocols and coding tools¶
Protocols¶
| Protocol | Purpose |
|---|---|
| MCP (Model Context Protocol) | Standard way to expose tools, data sources and prompts to any model/agent. Local (stdio) or remote (HTTP + OAuth). Supported by essentially every major agent and lab |
| ACP (Agent Client Protocol) | Lets editors talk to any coding agent over stdio/JSON-RPC (so one agent can plug into many editors) |
| A2A (Agent-to-Agent) | Google-originated protocol for agents to discover and delegate to each other |
| Agent Skills | Folders of instructions + scripts (SKILL.md) loaded on demand; adopted across several agents |
| AGENTS.md | Convention for project-level instructions to coding agents (build commands, conventions) |
Coding agents (2026)¶
| Agent | Vendor | Models | Interface |
|---|---|---|---|
| Claude Code | Anthropic | Claude | Terminal, IDE, desktop, web |
| Codex | OpenAI | GPT | Terminal, IDE, cloud, ChatGPT |
| Antigravity | Gemini (default) | Agentic IDE | |
| Cursor | Anysphere | Multi-model | IDE |
| GitHub Copilot | GitHub/Microsoft | Multi-model | IDE, CLI, cloud agent |
| Kiro | AWS | Multi-model incl. Claude | IDE, CLI |
| OpenCode | Anomaly (ex-SST) | Any (75+ providers, local) | Terminal, desktop, web — open source |
| Mistral Vibe | Mistral | Mistral (Devstral/Medium) | CLI, remote agents |
| Cline, Roo, Aider, Goose | Open source | Any | IDE / terminal |
The trade-off is the same as with models: vendor agents are tuned end-to-end for their own model; model-agnostic agents (OpenCode, Cline, Aider) give you portability, local-model support and inspectable code at the cost of doing more integration work yourself.
Agent patterns worth knowing¶
- Plan → build split: A read-only planning mode followed by an editing mode.
- Sub-agents: Delegate searches or parallel work to cheaper models in isolated contexts.
- Context compaction: Summarise old history when the window fills.
- Permission systems: Allow/ask/deny rules per tool and per command pattern.
- Sandboxing: Run agent shell commands in containers or VMs; treat any content the agent reads (web pages, issues, docs) as untrusted — prompt injection is the main security risk for agents.
11. Regulation and safety gating¶
2026 is the year regulation became part of the release pipeline:
- Export controls: Anthropic's Fable 5 / Mythos 5 suspension (12 June – 1 July 2026) to comply with US Commerce Department export controls.
- Pre-release federal review: Frontier releases from OpenAI went to small sets of vetted organisations before broad release following Commerce review. Reporting describes a US executive order (effective June 2026) establishing a de facto pre-release review of up to 30 days.
- Gated cyber models: OpenAI (GPT-5.5-Cyber), Google (Gemini 3.8 Flash Cyber via the Fairwind Program), and Anthropic's restricted Mythos tier all limit the most capable security models to vetted defenders.
- Classifier fallbacks: Public models increasingly route flagged cyber/bio requests to an older or more restricted model rather than refusing outright.
- EU AI Act: Obligations for general-purpose AI model providers are phasing in; relevant to Mistral and anyone deploying in the EU.
Practical implication: Release dates slip, access can be revoked or staged, and security-research workloads may need a specific programme or an open-weight model.
12. Benchmarks — and why to distrust them¶
Common benchmarks in 2026:
| Benchmark | Measures |
|---|---|
| SWE-Bench Verified / Pro | Resolving real GitHub issues |
| Terminal-Bench | Multi-step terminal tasks |
| OSWorld(-Verified) | Computer use in a desktop OS |
| DeepSWE | Long-horizon software engineering |
| Humanity's Last Exam (HLE) | Hard expert-level questions |
| GDPval | Economically valuable knowledge work |
| LMArena (Elo) | Blind human preference |
| Artificial Analysis indices | Aggregated intelligence/coding scores, speed, price |
Caveats:
- Vendor-reported numbers use the vendor's own harness and effort settings. Different labs rank differently on different benchmarks — e.g. OpenAI's claimed lead for GPT-5.6 Sol on one coding index inverted on SWE-Bench Pro, where Fable 5 scored 80% vs Sol's 64.6%.
- Benchmark gaming and contamination are real; independent evaluators (METR, Artificial Analysis, Epoch) matter.
- Your workload is the only benchmark that counts. Keep a small private eval set of real tasks and re-run it whenever you consider switching.
13. Choosing a model: a decision guide¶
Is the data allowed to leave your environment?
├── No → open weights, self-hosted (Qwen 3.8 27B, Muse Glimmer, gpt-oss,
│ Gemma 4, Mistral Small 4) or a cloud marketplace in your region
└── Yes
├── Hardest long-horizon/agentic work? → Fable 5.1, Opus 5.5, GPT-6 Astra
├── Everyday coding & documents? → Sonnet 5.5, GPT-6 Sol, Gemini 3.8 Flash
├── High volume / sub-agents / extraction? → Haiku 5.5, GPT-6 Luna, Flash-Lite,
│ DeepSeek V4-Flash
├── Heavy multimodal (video/audio in)? → Gemini
├── Cheapest strong open model via API? → DeepSeek V4, Kimi K3, Qwen3.8-Max
└── EU sovereignty requirements? → Mistral (API or self-hosted)
Practical rules:
- Abstract the provider. Use an OpenAI-compatible layer or SDK so switching is a config change.
- Route by task. Expensive model for planning/judgment, cheap model for volume.
- Measure cost per task, not per token. Include thinking tokens and cache hits.
- Keep a private eval set. Re-run it on every new release before switching.
- Pin model versions in production; use
latestaliases only in development. - Plan for availability shocks — regulatory suspensions, rate limits, deprecations. Have a fallback model configured.
14. Glossary¶
| Term | Meaning |
|---|---|
| Active parameters | Parameters actually used per token in an MoE model |
| Adaptive thinking | Model decides how much to reason per request |
| Context window | Max tokens (input + output) per request |
| Distillation | Training a small model to mimic a large one |
| Effort level | User-selectable reasoning budget |
| Harness | The software around a model: prompts, tools, loop, context management |
| KV cache | Stored attention keys/values for previous tokens; grows with context |
| MoE | Mixture of Experts — sparse model architecture |
| Open weights | Downloadable model parameters (training data usually not included) |
| Prompt caching | Discounted reuse of an unchanged prompt prefix |
| Quantisation | Storing weights at lower precision to save memory |
| RLVR | Reinforcement learning with verifiable rewards (tests, maths answers) |
| Sub-agent | A delegated agent with its own context, usually cheaper |
| System prompt | Developer instructions that frame model behaviour |
| Tool call | Structured request from the model to run a function |
15. Sources¶
- AF.net — Frontier models released in September 2026
- Analytics Vidhya — July 2026 AI releases timeline
- Essa Mamdani — AI model releases & open weights briefing, September 2026
- TechNews — Claude Opus 5.5 launch details
- MetricNexus — Claude Opus 5.5 benchmarks and pricing
- Computingforgeeks — Claude Sonnet 5.5 released
- AI Indigo — Claude Sonnet 5.5 release specs
- Le Nouvelliste — OpenAI launches GPT-6
- Yotta Labs — GPT-6 Sol and Luna pricing
- Emergent — GPT-6 Sol and Luna explained
- AI Weekly — Willison benchmarks GPT-6 Astra vs GPT-5.6
- Deccan Chronicle — Gemini 3.8 Flash launch
- BibiGPT — Gemini 3.8 Flash explained
- InfoWorld — Google plans Gemini 4 release before year-end
- Emergent — Gemini 4 release date: what we know
- Basenor — Grok roadmap
- Ars Technica via Harvard TagTeam — Meta Muse Spark
- NYU Shanghai RITS — DeepSeek V4 release
- dsebastien — DeepSeek V4 notes
- Mistral AI overview (systems-analysis.ru)
- Mistral 2026 updates (beginnersinai.org)
- BCAJ — GPT-5.6, Grok 4.5 and Gemma 4
- FutureSearch — Gemini 4 GA forecast (regulatory context)
- Anthropic — Fable/Mythos access statement
- Anthropic — Project Glasswing