Models / OpenAI/ GPT-5.5

GPT-5.5

OpenAI · released Apr 24, 2026

Input: text, images and documents. Output: text.InputOutput
Type
Closed
Input
$5.00
Output
$30.00
Cached
$0.50

List price · per 1M tokens · OpenAI at 1.1M context · machine-readable source ↗

Our take

Written Sep 30, 2026

GPT-5.5 is a hosted model with no download, so using it means choosing a host. Its measured results are strongest on data analysis and set-piece code, and weakest when an agent has to finish a multi-step task on its own.

Who should pick it

Reach for it on table and event-ordering work, where it places 2nd of 58 on LiveBench Data Analysis as of 25 Jun 2026, and on set-piece code generation, where it places 8th of 58 on LiveBench Coding as of 25 Jun 2026. Long documents need not be split up first, and it accepts text, images and files. Skip it if you need fixes landed in an existing codebase without review, or if you need an agent to finish a multi-step task on its own.

The case for it

  • 2nd of 58 on LiveBench Data Analysis as of 25 Jun 2026, a board of table and event-ordering tasks, so the result covers structured-data work rather than open-ended reasoning.
  • 8th of 58 on LiveBench Coding as of 25 Jun 2026, which covers code generation and completion rather than fixing issues in an existing project.
  • 2nd of 55 on Arena Agent · Tool use as of 25 Sep 2026, a score for calling the right tool and not inventing one, so it reflects tool selection rather than whether the task was finished.
  • 5th of 42 on SWE-bench Verified as of 19 Feb 2026, a result for the model inside that harness, measuring the share of real GitHub issues resolved end-to-end.

The case against it

  • 27th of 58 on LiveBench Agentic Coding as of 25 Jun 2026, a board run inside an agent harness, so it trails its own 8th of 58 on LiveBench Coding.
  • 35th of 55 on Arena Agent · Task outcome as of 25 Sep 2026, a score for finishing the task the session set out to do, against its 2nd of 55 on Arena Agent · Tool use.
  • We list no download for it, so using it means choosing a host, and the cheapest listed rate sits well under what the other hosts charge.
00

How good is it?

A general-purpose text model for everyday questions, drafting and coding, and for calling tools to carry out requests.

Good at
  • getting answers to everyday questionsArena Text (overall) · 20th of 168
  • drafts, rewrites and editingArena Creative Writing · 20th of 168
  • writing and completing codeArena Coding · 32nd of 168
  • calling tools to carry out requestsArena Agent · Tool use · 2nd of 55

EverydayGeneral questions and everyday reasoning

4 of 5

Arena Text (overall)20th of 168 · 1477

Arena Hard Prompts 23rd of 168Arena Maths 11th of 163LiveBench Data Analysis 2nd of 58LiveBench Mathematics 9th of 58LiveBench Reasoning 11th of 58

Also on this board: 1481 (Sep 25, 2026). Read the pair, not the higher one.

CodingWriting and fixing code on its own

4 of 5

Arena Coding32nd of 168 · 1513

Arena Code (WebDev) 37th of 95LiveBench Coding 8th of 58

Also on this board: 1519 (Sep 25, 2026). Read the pair, not the higher one.

AgenticPlanning, calling tools, staying on task

2.5 of 5

Arena Agent20th of 55 · 0.018

LiveBench Agentic Coding 27th of 58SWE-bench Verified 5th of 42

Also on this board: 0.056 (Sep 5, 2026), 0.044 (Sep 25, 2026). Read the pair, not the higher one.

WritingDrafting and rewriting prose

3.5 of 5

Arena Creative Writing20th of 168 · 1456

LiveBench Language 7th of 58

Also on this board: 1452 (Sep 25, 2026). Read the pair, not the higher one.

How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one2nd of 55
Steerabilitydoes what it was asked, and changes course when told19th of 55
Recoverygets back on track after a command fails17th of 55
Task outcomefinishes what the session set out to do35th of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
Arena Instruction Following 18th of 168LiveBench 8th of 58LiveBench Instruction Following 23rd of 58Arena Agent · Tool use 2nd of 55Arena Agent · Recovery 17th of 55Arena Agent · Steerability 19th of 55Arena Agent · Task outcome 35th of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model21 scoresEvery figure we hold, from 21 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
80.19source ↗
53.99source ↗
82.15source ↗
81.58source ↗
70.73source ↗
87.36source ↗
95.86source ↗
89.65source ↗
0.018source ↗
0.055source ↗
0.027source ↗
−0.04source ↗
0.008source ↗
1513source ↗
1456source ↗
1499source ↗
1500source ↗
1477source ↗
1512source ↗
74.4source ↗
01

Where to rent it

Prices checked 4 hours ago — each listing carries its own date.

Cheapest published offer

OpenAI, direct

The lab is the cheapest at this context. The strip above and this offer are the same one, compared at 1.1M of context. One cheaper row below is outside that comparison: a non-standard pricing tier.

per 1M tokens
$5.00 in / $30.00 out
Context served
1.1M
Throughput
~39 tok/s
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenAIflex tierThrough OpenRouter$2.50 / $15.00checked 4 hours ago1.1M128K max reply99 tok/sNoYesunknown periodUnknown
OpenRouterOpenRouter's own listing$5.00 / $30.00checked 4 hours ago1.1Mnot measuredUnknownUnknownUnknown
OpenAIDirect$5.00 / $30.00checked 4 hours ago1.1M128K max reply39 tok/sNoYesunknown periodUnknown
Microsoft Azure AIThrough OpenRouter$5.00 / $30.00checked 4 hours ago1.1M128K max reply63 tok/sNoNoConfirmed
Amazon Bedrockus-east-1Through OpenRouter$5.50 / $33.00checked 4 hours ago1.1M128K max reply92 tok/sNoNoUnknown
Microsoft Azure AIeuThrough OpenRouter$5.50 / $33.00checked 4 hours ago1.1M128K max reply30 tok/sNoNoConfirmed
Microsoft Azure AIusThrough OpenRouter$5.50 / $33.00checked 4 hours ago1.1M128K max reply34 tok/sNoNoConfirmed
OpenAIfast tierThrough OpenRouter$12.50 / $75.00checked 4 hours ago1.1M128K max reply80 tok/sNoYesunknown periodUnknown

Across the 8 listings we hold: 7 say they do not train on prompts, 0 say they do and 1 does not say. 3 appear in the zero-retention registry we check; the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
OpenAIflexThrough OpenRouter✓✓✓
OpenRouterOpenRouter's own listing✓✓✓
OpenAIDirect✓✓✓
Microsoft Azure AIThrough OpenRouter✓✓✓
Amazon Bedrockus-east-1Through OpenRouter✓✓✓
Microsoft Azure AIeuThrough OpenRouter✓✓✓
Microsoft Azure AIusThrough OpenRouter✓✓✓
OpenAIfastThrough OpenRouter✓✓✓

Tool calling: 8 of 8 listings say yes. JSON output: 8 of 8 listings say yes. Strict schema: 8 of 8 listings say yes.

02

Models people weigh against GPT-5.5

03

When we formed this view

Recent changes

Sep 25, 2026BenchmarkScored 0.018 on Arena Agent
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.055 on Arena Agent · Recovery
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.027 on Arena Agent · Steerability
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.04 on Arena Agent · Task outcome
What movedleaderboard
Sep 25, 2026BenchmarkScored 1513 on Arena Coding
What movedleaderboard
Sep 25, 2026BenchmarkScored 1456 on Arena Creative Writing
What movedleaderboard
Sep 25, 2026BenchmarkScored 1499 on Arena Hard Prompts
What movedleaderboard
Sep 25, 2026BenchmarkScored 1475 on Arena Instruction Following
What movedleaderboard
Sep 25, 2026BenchmarkScored 1500 on Arena Maths
What movedleaderboard
Sep 25, 2026BenchmarkScored 1477 on Arena Text (overall)
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 8 listings does not say whether it trains on prompts.
  • We hold no batch or off-peak rate for any of its listings.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text, images and documents in, text out
Catalogue slug
openai-gpt-5-5

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us