Models / OpenAI/ GPT-5.4

GPT-5.4

OpenAI · released Mar 5, 2026

Input: text, images and documents. Output: text.InputOutput
Type
Closed
Input
$2.50
Output
$15.00
Cached
$0.25

List price · per 1M tokens · OpenAI at 1.1M context · machine-readable source ↗

Our take

Written Sep 2, 2026

GPT-5.4 is OpenAI's flagship text model that can handle up to 1.05 million tokens in a single request, accepting text, images and files. It scores near the top on mathematics benchmarks and offers a wide spread of provider choices, though its agentic abilities lag far behind its standard coding performance.

Who should pick it

Pick this for long-document work at over one million tokens, mathematics-heavy tasks where it scores 94.15%, or budget batch work through the cheapest tier. Choose the fastest provider for high-throughput needs at 151 tokens per second. Skip it if you need strong agentic coding, autonomous tool use, or the ability to run weights locally.

The case for it

  • Extremely large request limit of 1.05 million tokens for long-document and multimodal inputs.
  • Top-tier mathematics performance at 94.15% on refreshed LiveBench tests.
  • Wide provider choice with a fourfold price spread for flexibility.
  • Highest measured throughput at 151 tokens per second on one provider.

The case against it

  • Agentic coding lags 23.7 points behind its standard coding score.
  • Arena agent scores are uniformly low across all sub-measures, near floor.
  • Proprietary weights with no disclosed licence; you cannot self-host.
00

How good is it?

A general-purpose assistant for everyday questions, drafting and editing, coding, and calling tools to carry out requests.

Good at
  • answering everyday questionsArena Text (overall) · 35th of 168
  • drafting, rewriting and editing textArena Creative Writing · 41st of 168
  • writing and completing codeArena Coding · 38th of 168
  • calling tools to carry out requestsArena Agent · Tool use · 9th of 55

EverydayGeneral questions and everyday reasoning

4 of 5

Arena Text (overall)35th of 168 · 1465

Arena Hard Prompts 34th of 168Arena Maths 40th of 163LiveBench Data Analysis 10th of 58LiveBench Mathematics 15th of 58LiveBench Reasoning 21st of 58

Also on this board: 1475 (Sep 25, 2026). Read the pair, not the higher one.

CodingWriting and fixing code on its own

4 of 5

Arena Coding38th of 168 · 1511

Arena Code (WebDev) 64th of 95LiveBench Coding 32nd of 58

Also on this board: 1521 (Sep 25, 2026). Read the pair, not the higher one.

AgenticPlanning, calling tools, staying on task

2.5 of 5

Arena Agent26th of 55 · 0.004

LiveBench Agentic Coding 29th of 58

WritingDrafting and rewriting prose

3 of 5

Arena Creative Writing41st of 168 · 1434

LiveBench Language 21st of 58

Also on this board: 1443 (Sep 25, 2026). Read the pair, not the higher one.

How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one9th of 55
Steerabilitydoes what it was asked, and changes course when told20th of 55
Recoverygets back on track after a command fails21st of 55
Task outcomefinishes what the session set out to do34th of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
Arena Instruction Following 34th of 168LiveBench 16th of 58LiveBench Instruction Following 25th of 58Arena Agent · Tool use 9th of 55Arena Agent · Steerability 20th of 55Arena Agent · Recovery 21st of 55Arena Agent · Task outcome 34th of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
77.97source ↗
53.84source ↗
77.54source ↗
79.31source ↗
70.22source ↗
82.63source ↗
94.15source ↗
88.12source ↗
0.004source ↗
0.03source ↗
0.026source ↗
−0.035source ↗
0.004source ↗
1511source ↗
1434source ↗
1487source ↗
1460source ↗
1465source ↗
1394source ↗
01

Where to rent it

Prices checked 4 hours ago — each listing carries its own date.

Cheapest published offer

OpenAI, direct

The lab is the cheapest at this context. The strip above and this offer are the same one, compared at 1.1M of context. One cheaper row below is outside that comparison: a non-standard pricing tier.

per 1M tokens
$2.50 in / $15.00 out
Context served
1.1M
Throughput
~35 tok/s
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenAIflex tierThrough OpenRouter$1.25 / $7.50checked 4 hours ago1.1M128K max reply3 tok/sNoYesunknown periodUnknown
Microsoft Azure AIThrough OpenRouter$2.50 / $15.00checked 4 hours ago1.1M128K max reply52 tok/sNoNoConfirmed
OpenRouterOpenRouter's own listing$2.50 / $15.00checked 4 hours ago1.1Mnot measuredUnknownUnknownUnknown
OpenAIDirect$2.50 / $15.00checked 4 hours ago1.1M128K max reply35 tok/sNoYesunknown periodUnknown
Microsoft Azure AIeuThrough OpenRouter$2.75 / $16.50checked 4 hours ago1.1M128K max reply18 tok/sNoNoConfirmed
Amazon Bedrockus-east-1Through OpenRouter$2.75 / $16.50checked 4 hours ago1.1M128K max reply15 tok/sNoNoUnknown
Microsoft Azure AIusThrough OpenRouter$2.75 / $16.50checked 4 hours ago1.1M128K max reply22 tok/sNoNoConfirmed
OpenAIfast tierThrough OpenRouter$5.00 / $30.00checked 4 hours ago1.1M128K max reply52 tok/sNoYesunknown periodUnknown

Across the 8 listings we hold: 7 say they do not train on prompts, 0 say they do and 1 does not say. 3 appear in the zero-retention registry we check; the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
OpenAIflexThrough OpenRouter✓✓✓
Microsoft Azure AIThrough OpenRouter✓✓✓
OpenRouterOpenRouter's own listing✓✓✓
OpenAIDirect✓✓✓
Microsoft Azure AIeuThrough OpenRouter✓✓✓
Amazon Bedrockus-east-1Through OpenRouter✓✓✓
Microsoft Azure AIusThrough OpenRouter✓✓✓
OpenAIfastThrough OpenRouter✓✓✓

Tool calling: 8 of 8 listings say yes. JSON output: 8 of 8 listings say yes. Strict schema: 8 of 8 listings say yes.

02

Models people weigh against GPT-5.4

03

When we formed this view

Recent changes

Sep 25, 2026BenchmarkScored 0.004 via High on Arena Agent
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.03 via High on Arena Agent · Recovery
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.026 via High on Arena Agent · Steerability
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.035 via High on Arena Agent · Task outcome
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.004 via High on Arena Agent · Tool use
What movedleaderboard
Sep 25, 2026BenchmarkScored 1511 on Arena Coding
What movedleaderboard
Sep 25, 2026BenchmarkScored 1434 on Arena Creative Writing
What movedleaderboard
Sep 25, 2026BenchmarkScored 1487 on Arena Hard Prompts
What movedleaderboard
Sep 25, 2026BenchmarkScored 1460 on Arena Instruction Following
What movedleaderboard
Sep 25, 2026BenchmarkScored 1460 on Arena Maths
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 8 listings does not say whether it trains on prompts.
  • We hold no batch or off-peak rate for any of its listings.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text, images and documents in, text out
Catalogue slug
openai-gpt-5-4

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us