AI engineer testing long-context LLM retrieval accuracy across multiple monitors

Every Frontier LLM Hit 1 Million Tokens in 2026: Most Still Can’t Use Half of It

Mon, Aug 24, 2026

In one tense demo, our app fed an entire corporate wiki to a 1-million-token-capable model, expecting it to recall a hidden figure in the ocean of text. Instead, our agent confidently gave the wrong answer.

As a veteran AI engineer, I have seen this pattern repeatedly: every major AI vendor raced to a "1M-token context window" milestone in 2026, but the headline number did not deliver what many teams assumed. This article breaks down what Claude, Gemini, GPT, and other frontier models actually released, then examines MRCR v2 benchmark data showing that retrieval accuracy can fall well before 1 million tokens.

The practical question is not how much text a model accepts. It is how much of that text the model can use reliably, and how that gap should influence architecture decisions. Refonte Learning develops the architectural judgment and systems thinking AI engineers need to evaluate claims like these.

The Number Every AI Vendor Is Suddenly Bragging About

Nearly every frontier LLM vendor now advertises a "1 million token" context window. In early 2026, it became the hottest marketing number, but the dates, endpoints, and output limits still matter.

  • Anthropic Claude Opus 4.8, May 28, 2026: The official Claude Opus 4.8 announcement describes general availability at standard Opus pricing. The model supports a 1,000,000-token context window and up to 128K output tokens.

  • Anthropic Claude Sonnet 5, June 30, 2026: Anthropic's platform release notes list a 1M-token context window and 128K maximum output. The same notes confirm that the $2 input and $10 output price per million tokens became permanent after Anthropic canceled the planned increase to $3 and $15.

  • Google Gemini 3.1 Pro, February 19, 2026: Google's Gemini 3.1 Pro model card lists a 1,000,000-token input context and 64K maximum output. Some third-party pages cite 2M, but Google's own documentation supports 1M for the Pro model.

  • OpenAI GPT-5.5, April 23-24, 2026: The GPT-5.5 launch announcement lists a 1M-token API window and 128K maximum output, while the Codex product uses a 400K context window.

  • Further expansion: Amazon Bedrock later documented 1M-token support for GPT-5.6 Sol, Terra, and Luna. Meta's Llama 4 Scout makes the more extreme 10M-token claim, which belongs in a different reliability category from a production API specification.

Each vendor promoted the threshold as a breakthrough, as though capacity alone were a superpower. The real story is how well a model handles information at those lengths, and the answer remains: not reliably enough to trust the headline number by itself.

What "1 Million Tokens" Actually Means in Practice

One million tokens sounds like enough to fit an encyclopedia, but the practical limit is more fragile. Using the rough rule of four characters per token, 1M tokens is about 4 million characters, or roughly hundreds of book pages. In theory, a model with a 1M window can accept that much content in one request.

  • Accuracy degrades: As more tokens enter the prompt, relevant information competes with more distractors. A model may retrieve one fact at 32K tokens and then lose track of several related facts at 512K or 1M.

  • Costs scale with traffic: Even when a provider makes long context affordable per request, repeatedly filling a million-token window can be materially more expensive than sending a targeted prompt. That is a workload-economics question, not proof of capability.

  • Engineering complexity increases: A 1M-token prompt is not an invitation to dump every available document into one request. Production systems still need chunking, compression, retrieval, history management, and relevance controls.

Put simply, a "1M-token window" is a raw capacity specification, not a guarantee of meaningful performance. What matters is whether the model can use the extra context for multi-step reasoning and factual retrieval without losing precision.

Claude Opus 4.8: Anthropic's Default 1M-Token Window

Anthropic made long context a default production option in its Opus line with Claude Opus 4.8. The rollout matters because customers no longer need a separate long-context mode on the relevant endpoints.

  • Context window: 1,000,000 tokens by default.

  • Maximum output: Up to 128K tokens in one response.

  • Standard pricing: $5 per million input tokens and $25 per million output tokens. Fast mode is priced at $10 input and $50 output per million tokens.

  • Positioning: Anthropic presents Opus 4.8 for demanding coding, agentic, analysis, and data workloads that may need sustained context.

The specification is substantial, but it still does not answer the question that matters in production: how accurately can the model retrieve and combine facts near the far end of that window?

The Price Increase That Got Canceled

Claude Sonnet 5 arrived with the same 1M-token context window and introductory pricing of $2 per million input tokens and $10 per million output tokens. Anthropic had planned to increase those rates to $3 and $15 on September 1, 2026, but canceled the increase on August 10, making the introductory rates standard.

The rollback shows how competitive long-context pricing became once several vendors crossed the same threshold. Cheap capacity is still useful only when the model can retrieve the right information from it.

Gemini 3.1 Pro and the 1M-vs-2M Confusion

Google announced Gemini 3.1 Pro on February 19, 2026. Its official model card lists a 1,000,000-token input context and 64K maximum output, and Google's launch materials describe a massive 1M-token input context.

Some media reports and aggregator pages cite 2 million tokens for the Gemini 3.1 Pro context length. Google's own documentation does not support that figure for the Pro model, so the publication-safe number is 1M. The 2M claim appears to mix future, research, or differently named variants with the documented Pro product.

  • Verified for Gemini 3.1 Pro: 1M input tokens and 64K output tokens.

  • Unverified for Gemini 3.1 Pro: 2M input tokens in third-party summaries.

  • Engineering takeaway: Check the exact model card and endpoint rather than copying an aggregator's headline figure.

Gemini 3.1 Pro therefore joins Claude at the 1M threshold. That puts Google near the center of the 2026 context race, but it does not settle the question of effective retrieval at the maximum length.

GPT-5.5 and OpenAI's Version of the Same Bet

OpenAI released GPT-5.5 on April 23, 2026, and updated the launch on April 24 when API availability began. The API supports a 1M-token context window with up to 128K output tokens, while GPT-5.5 in Codex is limited to 400K.

The distinction is not cosmetic. It means "the GPT-5.5 context window" changes by product surface. In July and August 2026, the GPT-5.6 Sol, Terra, and Luna family reached the same 1M class through Amazon Bedrock, reinforcing OpenAI's commitment to the large-window tier.

  • API: Approximately 1M input context, with model documentation commonly expressing the limit as about 1.05M total tokens.

  • Codex: 400K context for the coding product.

  • Output: Up to 128K tokens on the API model.

OpenAI also reported that GPT-5.5 improved over GPT-5.4 on several coding and agentic benchmarks while using fewer tokens. That efficiency claim is useful, but it is separate from whether the model can retrieve multiple facts reliably across the full window.

Why Codex Gets a Smaller Window

The smaller Codex cap likely reflects a product trade-off among task fit, latency, memory, and interaction design rather than a universal model limitation. Code workflows often benefit more from finding the relevant files and symbols than from feeding an entire repository into every turn.

The split is a practical reminder: an advertised limit may vary across ChatGPT, the API, Codex, cloud platforms, and account tiers. Always check the documentation for the endpoint you will actually deploy.

Five Frontier Models, One Quarter, One Threshold

By mid-2026, the 1M threshold had become the common comparison point for frontier releases. The pattern was not an officially declared race; it emerged from several vendors publishing similar capacity claims within a compressed release cycle.

  • Gemini 3.1 Pro: 1M tokens, announced in February 2026.

  • GPT-5.5: 1M-class API context, announced in April 2026.

  • Claude Opus 4.8: 1M tokens, generally available in May 2026.

  • Claude Sonnet 5: 1M tokens, launched in June 2026.

  • GPT-5.6 Sol, Terra, and Luna: 1M-token support documented on Amazon Bedrock in August 2026.

The clustering is easier to see against a broad survey such as The AI Model Landscape 2026, which maps models across many dimensions. This article is deliberately narrower: the important comparison is not who printed the largest context number, but who preserved retrieval accuracy as the prompt approached that limit.

The Benchmark That Undercuts the Marketing: MRCR v2

MRCR v2, or Multi-Round Coreference Resolution, is a multi-needle retrieval test. It inserts several similarly formatted facts into a long body of distractor text, then asks the model to recover a specified item. The eight-needle variant is difficult because the model must distinguish and track several targets rather than find one obvious string.

A cross-model MRCR v2 benchmark compilation published on March 15, 2026, reported a striking spread at the long end: Claude Opus 4.6 at about 76%, GPT-5.4 at 36.6% in the 512K-to-1M range, and Gemini 3 Pro at about 24.5% at 1M. The compilation draws on vendor results and third-party testing under different settings, so the values are convergent evidence rather than one perfectly controlled official paper.

OpenAI's GPT-5.4 launch benchmark independently reports the 36.6% figure for the 512K-to-1M MRCR v2 range. Across the broader comparisons, multi-fact retrieval commonly loses 30 to 60 percentage points somewhere after roughly 200K tokens. That is the measurable gap between advertised capacity and effective context.

  • Advertised window: The number of tokens an endpoint accepts.

  • Effective window: The amount of context a model can use at the accuracy your application requires.

  • MRCR lesson: Two models with the same nominal window can have radically different retrieval reliability.

Why Accuracy Drops Off a Cliff Past 200K Tokens

LLMs are not perfectly indexed databases. The RULER long-context benchmark was designed precisely because simple single-needle tests can overstate long-context understanding; it adds multiple needles, aggregation, and multi-hop tracing, and finds large performance drops as context length and task complexity increase.

As the context grows, distant information competes with more distractors, attention becomes less selective, and the model must preserve more intermediate associations. For many multi-fact workloads, a practical reliable range may be closer to 200K-to-400K even when the endpoint accepts 1M. The exact boundary depends on the model, prompt structure, retrieval pattern, and error tolerance.

Why a Bigger Window Doesn't Mean Better Retrieval

A massive context window is valuable only when the model can use it. MRCR v2 shows that adding tokens does not produce a proportional increase in useful task performance, especially when questions require several facts, references, or reasoning steps.

This is why retrieval and indexing remain central to production architecture. A 1M-token model does not make a vector database obsolete; it can still omit or hallucinate facts when forced to read everything at once. The Vector Database Shakeout covers the infrastructure landscape, while the long-context evidence here explains why that infrastructure still matters.

Single-needle tests can look close to perfect even at very long lengths, which makes them attractive marketing material. More demanding tests tell a different story. The NoLiMa benchmark removes easy literal matches between questions and evidence, and it finds substantial degradation as context grows.

  • Capacity is not indexing: The model does not maintain a deterministic lookup table over every token.

  • More context can add noise: Irrelevant passages can reduce the salience of the evidence you care about.

  • Retrieval remains a reliability control: Selecting a smaller, relevant evidence set often improves accuracy and auditability.

The Extremes: 2M, 10M Tokens, and What They're Actually For

A few model families advertise limits beyond the 1M tier. Those figures are useful for research and specialized workloads, but they should not be treated as proof of reliable reasoning across the full sequence.

  • Meta Llama 4 Scout: The official Llama 4 Scout model card claims a 10M-token context window. The model is open-weight, so deployment quality depends heavily on the serving stack, quantization, memory configuration, and evaluation method.

  • Google variants: Some third-party pages attach 2M-token figures to Gemini research or future variants, but Gemini 3.1 Pro's official documentation remains at 1M.

  • xAI Grok: Aggregators have circulated 2M-token claims for some Grok variants, while xAI's current model and pricing documentation lists different context limits by model. Treat the 2M figure as unverified unless the exact endpoint documentation supports it.

Extreme windows can be useful for experiments such as whole-repository analysis, long multimodal records, or broad document review. Until independent tests show stable multi-fact retrieval near the limit, they remain capacity demonstrations rather than guaranteed effective context.

What This Means for How You Architect an LLM Application

Do not design around context size alone. Start from the task: if the application must track many facts, preserve citations, or reason across several documents, plan for retrieval, chunking, and staged reasoning regardless of the model's nominal window.

  • Document question answering: Use a vector store or search engine to retrieve a small set of relevant chunks instead of dumping an entire corpus into every prompt.

  • Large codebases: Combine static analysis, symbol search, dependency graphs, and repository tools with model calls on the files that matter.

  • Long conversations: Prune, summarize, or compress older turns, and preserve critical state in structured memory rather than relying on the raw transcript alone.

The goal is localized context with high relevance. The AI Model Development and Optimization and Scaling AI Systems modules in Refonte Learning's curriculum build the judgment needed to choose a model, prompt strategy, and system architecture for a workload. The curriculum does not claim to teach this specific MRCR v2 result or every 2026 release.

When You Still Need Retrieval Instead of Raw Context

Use retrieval first when the task contains many facts, requires citations, changes frequently, or must produce an auditable evidence trail. A legal or scientific system might retrieve the five strongest passages and let the model reason over those, rather than asking it to search hundreds of thousands of unfiltered tokens internally.

I still use the standard pattern of retriever, encoder or search layer, and LLM for high-stakes knowledge work. A larger window can reduce the number of retrieval rounds or preserve more neighboring evidence, but it does not eliminate the need to select relevant information.

Cost Is a Separate Problem From Capability

Cost matters, but it is a separate axis from retrieval reliability. A model can accept a million tokens and still fail to use the last several hundred thousand accurately; a cheaper model can reduce spend without fixing that capability gap.

  • Test capability first: Establish the smallest context length that meets your accuracy target.

  • Price the real workload: Include repeated calls, cache behavior, output tokens, retries, retrieval, and concurrency.

  • Avoid false precision: A circulating comparison placed a 1M-token request at roughly $0.14 on a DeepSeek flash model and roughly $10 on an Anthropic flagship. The methodology and assumptions were not transparent, so treat it as a loose illustration, not a purchasing benchmark.

The architectural point is simple: paying to fill the window is wasteful if the model no longer retrieves reliably near the end of it. Optimize cost only after you have measured the effective context your application can use.

Common Mistakes Engineers Make Trusting Context Length Alone

These are the failure patterns I see when teams treat a context specification as a reliability guarantee:

  • Shipping without retrieval tests: The team points the model at a huge corpus and hopes it "just works," then discovers hallucinations or missing facts in production.

  • Misreading the specification: Engineers assume 1M applies across every product surface, even when the API, chat product, Codex endpoint, cloud host, or account tier uses a different limit.

  • Ignoring prompt overhead: System messages, tool definitions, conversation history, retrieved evidence, and model-generated reasoning all consume the available window.

  • Assuming more context always helps: Multi-hop accuracy can be higher at 128K than at 1M because additional distractors dilute the evidence and increase coordination difficulty.

Testing Your Own Retrieval Accuracy Before You Ship

Simulate the workload before launch. Inject eight factual needles, or a domain equivalent, at varying context lengths; measure exact retrieval, citation fidelity, refusal behavior, latency, and cost; then plot accuracy against tokens. A 40-point drop between 50K and 500K is a design signal, not a minor benchmark curiosity.

Use the result to choose smaller chunks, add retrieval, change the model, or narrow the supported workflow. LLM Evaluation Pipelines for AI Engineers covers broader evaluation-pipeline practices; this article's narrower contribution is the context-retrieval failure mode that the pipeline must measure.

AI Engineer Salaries in 2026

The market value of this judgment shows up in compensation data, although salary databases measure different populations and should not be treated as interchangeable.

The difference is about $43,000. It likely reflects methodology and title mix: ZipRecruiter aggregates job-posting and third-party data across a broad market, while Glassdoor's total-pay estimate can include more senior or specialized roles and additional compensation. The divergence is a reason to inspect the source, not to average the two numbers blindly.

Building This Judgment: The Refonte Learning AI Engineering Program

What ties the article together is architectural judgment: knowing when a specification is meaningful, how to test a model independently, and how to adapt a system when the benchmark exposes a gap.

  • AI Model Development and Optimization: Builds the habit of comparing models against workload-specific evidence rather than a marketing headline.

  • Data Engineering for AI: Supports the pipelines, indexing, and data-quality controls that make retrieval dependable.

  • Scaling AI Systems: Develops the systems perspective needed to balance context, latency, reliability, and infrastructure.

  • AI Ethics and Governance: Reinforces the need to validate model behavior before relying on it in high-stakes workflows.

The eight-module curriculum is mentored by Dr. John Anderson, a Senior AI Engineer with 17 years of experience. It does not name context windows, the MRCR benchmark, structured outputs, or the specific 2026 releases discussed here. The honest connection is that the program builds the model-evaluation and architectural reasoning needed to assess claims like these.

The context-window race taught a useful production lesson: the longest advertised context wins attention, but reliable retrieval wins customer trust. Benchmark the exact workload, keep retrieval infrastructure where it improves accuracy, and treat effective context as an empirical property rather than a vendor promise.

For structured training, mentorship, and real-world AI engineering practice, explore the Refonte Learning AI Engineering Program.