Published 27 September 2026 · Every context window, output limit, price and overflow behaviour below was read from OpenAI, Anthropic, Google, xAI and DeepSeek documentation on 26 September 2026.
GPT-4o has a 128,000-token context window and can write at most 16,384 tokens in a single reply, with the reply counted inside the same 128,000. What happens when a conversation grows past that depends on where it is running: the OpenAI API rejects the request with a 400 error unless you have switched on truncation (which drops the oldest items) or compaction (which replaces older turns with a compressed summary), while ChatGPT, which retired GPT-4o on 13 February 2026, caps every plan well below any model's full window and eventually tells you to start a new chat.
GPT-4o is now a legacy model, and the number that decides how long your chat survives is rarely the model's own ceiling. The current flagships from OpenAI, Anthropic and Google all sit at or just above one million tokens, yet a ChatGPT Plus user on the Instant model gets 54K, and quality starts to slip long before any window is full. The rest of this piece covers the current numbers for every major model, what each API and each chat app does at the limit, what a full window costs, and why none of it carries a single fact into tomorrow's session.
GPT-4o's context window, and where GPT-4o stands in September 2026
OpenAI's GPT-4o model page lists three numbers that still hold: a 128,000-token context window, a 16,384-token maximum output, and a knowledge cutoff of 1 October 2023. The context window is shared: OpenAI's conversation state guide defines it as "the total tokens that can be used for both input and output tokens (and for some models, reasoning tokens)", so a prompt of 120,000 tokens leaves room for at most 8,000 tokens of answer, not 16,384.
GPT-4o is gone from ChatGPT and still present in the API. OpenAI's help center records that GPT-4o, GPT-4.1, GPT-4.1 mini, o4-mini and GPT-5 (Instant and Thinking) were retired from ChatGPT on 13 February 2026, with old conversations moved to newer equivalents, and that these models "will continue to be available through the OpenAI API". The API is narrowing too: the deprecations page schedules the original gpt-4o-2024-05-13 snapshot for shutdown on 23 October 2026, with gpt-5.6-sol named as the substitute, and the chatgpt-4o-latest alias was already removed on 17 February 2026. If you are reading this because an application pinned to GPT-4o is hitting its 128K ceiling, the practical answer is that every current OpenAI flagship carries roughly eight times the window.
Context windows of the major models in September 2026
Every figure in this table comes from the vendor's own model page, not from a comparison site, because several of the comparison tables ranking for this query today list windows that the vendors do not.
| Model | Vendor | Context window (tokens) | Max output (tokens) | Notes |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 1.05M | 128K | Flagship, $10 input / $50 output per million |
| GPT-6 Sol | OpenAI | 1.05M | 128K | $2 / $10 |
| GPT-6 Luna | OpenAI | 1.05M | 128K | $0.10 / $0.50 |
| GPT-4o | OpenAI | 128,000 | 16,384 | Legacy; retired from ChatGPT 13 Feb 2026 |
| Claude Fable 5.1 | Anthropic | 1M | 128K | $10 / $50 |
| Claude Opus 5.5 | Anthropic | 1M | 128K | $4 / $20 |
| Claude Sonnet 5 | Anthropic | 1M | 128K | $2 / $10 |
| Claude Haiku 4.5 | Anthropic | 200K | 64K | $1 / $5 |
| Gemini 3.1 Pro (preview) | 1,048,576 | 65,536 | $2 input up to 200K, $4 above | |
| Gemini 3.8 Flash | 1,048,576 | 65,536 | Current stable Flash | |
| Grok 4.7 | xAI | 500,000 | Not listed on the model page | $2 / $6; different rates above 200K |
| DeepSeek V4.1 Flash and V4 Pro | DeepSeek | 1M | 384K | Thinking mode on by default |
Sources: OpenAI models, Anthropic models overview, Gemini 3.1 Pro and Gemini 3.8 Flash, Grok 4.7, DeepSeek, all read 26 September 2026. Claude pricing is covered line by line in our Claude API pricing guide.
Two things about that table catch people out. Everything you send counts against the window, not only the chat: Anthropic's context window documentation lists the system prompt, every message including tool results, images and documents, the tool definitions themselves, and the model's output including extended thinking, and on Claude a single request can carry at most 600 images or PDF pages (100 on 200K models), so a document-heavy request can hit a request size limit before it hits the token limit. The second is scale: Google's Gemini Apps help page puts one million tokens at roughly 1,500 pages of text or 30,000 lines of code, which is why a 1M window feels infinite in a demo and finite in an agent that reads tool output all day.
What happens when an API request exceeds the context window
The model never decides this. A transformer is trained on sequences up to a fixed length, and pushing positions past that length "may lead to catastrophically high attention scores that completely ruin the self-attention mechanism", as the authors of the Position Interpolation paper put it, so every provider enforces the ceiling in the serving layer, and each one enforces it differently.
| Surface | Input already too large | Reply would run past the window | Built-in way to keep going |
|---|---|---|---|
| OpenAI Responses API | 400 error by default (truncation: "disabled"); with truncation: "auto" the oldest items are dropped silently | Output may be truncated | Server-side compaction via context_management and compact_threshold, or the standalone /responses/compact endpoint |
| Anthropic Messages API | 400 invalid_request_error, "prompt is too long", on every model | Claude 4.5 and newer accept the request and stop with stop_reason: "model_context_window_exceeded"; older models return a validation error | Server-side compaction (beta), on demand or at a token threshold |
| Google Gemini API | 400: "The input token count (...) exceeds the maximum number of tokens allowed (1048576)" | Not documented on the pages read | None documented on the model or long-context pages |
OpenAI's behaviour is set by the truncation parameter in the Responses API reference: the default fails the request with a 400, and auto drops "items from the beginning of the conversation", which usually means the instructions you wrote first. For long conversations OpenAI's documentation steers developers toward compaction instead, where the server replaces earlier turns with an encrypted compaction item once the rendered token count crosses a threshold you set; that item is opaque and "not intended to be human-interpretable", so you cannot inspect what was kept.
Anthropic separates the two failure cases cleanly. An oversized prompt fails up front; an oversized reply on a current model returns HTTP 200 with a partial answer and a stop reason, which means code that only catches exceptions will ship truncated answers to users without noticing. Anthropic's compaction writes a readable summary on the server and can keep the most recent turns word for word. Google's error text comes from a public Gemini CLI issue opened in October 2025, where a 2,551,556-token request was refused against a 1,048,576 limit.
What happens when a chat exceeds the window in ChatGPT, Claude and Gemini
Chat apps do not give you the model's full window, and each app handles the edge differently.
| App and plan | Context window in the app |
|---|---|
| ChatGPT Free | 27K (Instant); reasoning "varies" |
| ChatGPT Go and Plus | 54K (Instant); 256K (Reasoning) |
| ChatGPT Pro | 128K (Instant); 400K (Reasoning) |
| Claude paid plans | Up to 1M on the newest models; 500K or 200K on others |
| Gemini without a plan | 32K |
| Gemini AI Plus | 128K |
| Gemini AI Pro and AI Ultra | 1M |
Sources: ChatGPT pricing, Claude usage and length limits, Gemini Apps limits, read 26 September 2026.
ChatGPT. OpenAI publishes the per-plan windows above but does not document what ChatGPT does as a conversation approaches them. What users report, including in a thread on OpenAI's own developer forum opened on 27 August 2026, is a hard stop with the message "You've reached the maximum length for this conversation, but you can keep talking by starting a new chat." Whether ChatGPT trims or summarises older turns before that point is not published, and pages that describe its internal trimming in detail are guessing. The gap between 54K in the Plus app and 1.05M for GPT-6 over the API is the whole reason a long ChatGPT chat can fail while the same model on the API would not.
Claude. Anthropic documents the behaviour. With code execution turned on, "when your conversation approaches the context window limit, Claude summarizes earlier messages to continue the conversation", the full history stays available for reference, and longer conversations consume more of your usage limit. Anthropic's API documentation adds that chat interfaces such as claude.ai "can also manage the context window on a rolling 'first in, first out' basis", which is the other name for dropping the oldest turns.
Gemini. Google's help page is candid about the failure mode rather than the mechanism: exceeding the window "could lead to responses that don't take into account all the content provided or miss connections or details". There is no error; the answer simply stops reflecting part of what you gave it.
Chats degrade before they hit the wall
Running out of window is the visible failure. The quieter one starts much earlier. Lost in the Middle, by Nelson Liu and colleagues, found that performance "is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models". Chroma's context rot study tested 18 models with task difficulty held constant and found that "model performance degrades as input length increases, often in surprising and non-uniform ways". Anthropic now writes the same thing into its own documentation: "As token count grows, accuracy and recall degrade."
So the advertised window is a capacity, not a promise of equal attention. The effective length at which a given model stays reliable on your task is not published for the current generation, and no benchmark read for this piece reports it for GPT-6, Opus 5.5 or Gemini 3.1 Pro; treat it as something to measure on your own data. For memory specifically there is a number: LongMemEval found commercial chat assistants and long-context LLMs "showing a 30% accuracy drop on memorizing information across sustained interactions".
What a full context window costs
A full window is billed on every call that carries it. At list input prices read today, one 1M-token prompt costs about $10 on GPT-6 Astra or Claude Fable 5.1, $4 on Claude Opus 5.5, $2 on GPT-6 Sol or Claude Sonnet 5, and $4 on Gemini 3.1 Pro, where Google's pricing doubles the input rate from $2 to $4 once a prompt passes 200K tokens. xAI also charges different rates for Grok 4.7 above 200K. Anthropic, by contrast, bills the full 1M window at its standard rates.
Prompt caching lowers the price of a repeated prefix but does not free any space: in Anthropic's words, "cached prompt prefixes still occupy the context window: prompt caching changes what you pay for those tokens, not whether they count." And because an API call is stateless, a chat that re-sends its history every turn pays for that history again on every turn; our Agentic Context Management paper shows that naive accumulation "grows token cost quadratically in conversation length". We worked through the arithmetic, including where caching stops being enough, in how to reduce LLM token costs in long conversations.
Truncate, summarise or compact: choosing an overflow policy
If you build on an API, you choose what happens at the limit, and the three options fail differently. Truncation (OpenAI's auto mode, or FIFO in your own code) is free and silent, and it removes the oldest turns first, which is where the user usually stated their constraints. Summarisation in your own code keeps the gist, costs an extra model call, and drifts when summaries get summarised again. Server-side compaction from OpenAI or Anthropic runs the summarisation for you; Anthropic's version can keep recent turns verbatim and shows you the summary, while OpenAI's returns an opaque item. Our paper's finding is that crude summarisation "buys linear cost at the price of an accuracy cliff, and only validated compaction achieves linear cost with preserved fidelity", which is why whatever you pick needs a check that the facts survived the squeeze.
Whichever policy you pick, the handling code has two halves. Before sending, count tokens against a budget you set below the ceiling, leaving room for the reply and for reasoning tokens; Anthropic's Models API returns max_input_tokens and max_tokens per model so the limit does not have to be hard-coded. After sending, catch the 400 on OpenAI, Gemini and Anthropic, and on Claude also check stop_reason for model_context_window_exceeded, because that failure arrives as a success. For agents, design against your own token budget rather than the model's maximum: the model's window is the point where things break, not the point where they work best.
What to do when your own chat hits the limit
For someone using ChatGPT, Claude or Gemini rather than building on them, the options are practical ones. Ask the assistant to summarise the conversation, then paste that summary into a new chat, which resets the window to near zero while keeping the thread. Move reference material into a project: Anthropic's help page notes that Claude projects use retrieval, "only loading relevant content into the context window". On Claude, turning extended thinking off or lowering the effort level for routine work leaves more room, since thinking tokens count. Upgrading changes the ceiling too; a ChatGPT Pro Instant chat gets 128K against 54K on Plus. None of this fixes the underlying problem, which is that the next chat starts empty.
A bigger context window is still not memory
A context window is per request. It holds what you sent on this call, and when the call ends nothing is kept unless your application sends it again. That is the definition, not a limitation of 2026 models, and it is why a 1M window changes how long one conversation can run while changing nothing about whether the agent knows the user next week. We make the full case on our bigger window is not a memory page, and the companion piece to this one explains the difference between short-term and long-term memory in LLMs, including what "context memory" means.
Maximem Synap is the layer we built for the part the window cannot do. Inside a conversation, Synap's Agentic Compaction keeps the working context bounded: it triggers at 3,000 tokens, 10 messages or 5 minutes idle, compresses to a 1,500-token target while keeping the last three conversation pairs verbatim, and returns a validation score and a preserved-facts count on every pass so you know the compression kept the signal. Across conversations, Synap extracts the facts worth keeping as they happen, and when the user returns the agent asks for context and gets a ranked block within a token budget (2,000 tokens by default) instead of the whole transcript. Reads are pre-fetched into a cache inside your own process, at an asserted P75 under 15 ms in conversation. The window then becomes a place to reason over the right few thousand tokens, and the bill stops tracking the length of the relationship.
Frequently asked questions
What is GPT-4o's context window?
GPT-4o has a 128,000-token context window with a maximum output of 16,384 tokens, according to OpenAI's model page read on 26 September 2026. Input and output share the window, so a 120,000-token prompt leaves at most 8,000 tokens for the reply. Its knowledge cutoff is 1 October 2023.
Is GPT-4o still available in 2026?
GPT-4o was retired from ChatGPT on 13 February 2026 and remains available through the OpenAI API. OpenAI has scheduled the original gpt-4o-2024-05-13 snapshot for shutdown on 23 October 2026 and names gpt-5.6-sol as its substitute. The current OpenAI flagships, GPT-6 Astra, Sol and Luna, each have a 1.05M-token window.
What happens when an API request exceeds the context window?
On the OpenAI Responses API the request fails with a 400 error by default, or the oldest items are dropped if truncation is set to auto. On the Anthropic API an oversized prompt returns a 400 "prompt is too long"; on Claude 4.5 and newer a reply that runs out of room stops with stop_reason: "model_context_window_exceeded". On the Gemini API an oversized input returns a 400 naming the token count and the limit.
What happens when a ChatGPT conversation gets too long?
ChatGPT caps each plan below the model's API window, for example 54K tokens on Plus with the Instant model. Users report that long chats end with "You've reached the maximum length for this conversation, but you can keep talking by starting a new chat." OpenAI does not document whether ChatGPT trims or summarises older turns before that point. Claude, by contrast, summarises earlier messages automatically when code execution is enabled.
Which LLM has the largest context window in September 2026?
Among the vendor documentation read on 26 September 2026, OpenAI's GPT-6 models list 1.05M tokens, Google's Gemini 3.1 Pro and 3.8 Flash list 1,048,576, and Anthropic's Claude Fable 5.1, Opus 5.5 and Sonnet 5 list 1M, as do DeepSeek's V4 models. Grok 4.7 lists 500,000. Quality still degrades as a window fills, so the largest window is not automatically the most reliable one.
Does a bigger context window replace long-term memory?
No. A context window only holds what is sent on the current request and is empty again on the next session unless the application re-sends it, at full token cost each time. Long-term memory stores what matters outside the model and retrieves a small relevant slice into the window when needed, which is what Maximem Synap does for AI agents.
Sources, retrieved 26 September 2026: OpenAI, GPT-4o model; OpenAI, models; OpenAI, deprecations; OpenAI, conversation state; OpenAI, compaction; OpenAI, Responses API reference; OpenAI Help Center, retiring GPT-4o in ChatGPT; ChatGPT pricing; OpenAI developer forum, maximum length thread; Anthropic, models overview; Anthropic, context windows; Anthropic, compaction; Claude Help Center, usage and length limits; Google, Gemini 3.1 Pro; Google, Gemini 3.8 Flash; Google, Gemini pricing; Google, Gemini Apps limits; Gemini CLI issue 11248; xAI, Grok 4.7; DeepSeek, models and pricing; Liu et al., Lost in the Middle; Chroma, Context Rot; Chen et al., Position Interpolation; Wu et al., LongMemEval; Maximem, Agentic Context Management.



