Models / OpenAI/ o3

o3

OpenAI · released Apr 16, 2025

Input: text, images and documents. Output: text.InputOutput
Type
Proprietary
Input
$2.00
Output
$8.00
Cached
$0.50

List price · per 1M tokens · OpenAI at 200K context · source ↗

Our take

Written Aug 3, 2026

o3 is a hosted-only reasoning model from OpenAI that handles text, images and files up to 200,000 tokens in a single request. It scores strongly on coding benchmarks but shows a clear gap between its specialist performance and its creative writing scores.

Who should pick it

Choose this for complex coding tasks where measured benchmark scores matter, or for long-document and multi-file work that needs the full 200,000-token reach. It suits instruction-following workflows and situations where identical pricing across both tracked providers removes comparison friction. Skip it if you need downloadable weights, if creative writing quality is central, or if you want the cheapest output rates among reasoning models.

The case for it

  • Strong measured coding performance: 84.7% on LiveCodeBench and an Arena Coding Elo of about 1459.
  • 200,000-token request limit supports long-document and multi-file tasks.
  • Identical pricing on both tracked providers, so provider choice does not create cost surprises.
  • Known throughput of 71 tokens per second on OpenAI's own infrastructure.

The case against it

  • Creative writing is its weakest arena score, 76 points below its own coding mark on the same leaderboard batch.
  • Overall text Elo sits below both its coding and maths scores, suggesting general chat is not its strength.
  • Output costs four times the input rate, a steep multiplier for a proprietary-only model.
00

How good is it?

IntelligencePuzzles, maths, exam questions

3 of 5

Arena Text (overall)54th of 143 · 1430.9

Arena Hard Prompts 60th of 143Arena Maths 35th of 139

CodingWriting and fixing code on its own

3 of 5

Arena Coding65th of 143 · 1459.4

LiveCodeBench 2nd of 14 via High

AgenticPlanning, calling tools, staying on task

Scored, not ratedSWE-bench Verified · 22nd of 39 · 58.4via mini-SWE-agent

o3 is not on Arena Agent (IPS), which is where the rating would come from, so there is no rating here. It is on SWE-bench Verified, in 22nd of 39 with 58.4.

WritingWe do not rate this

Scored, not ratedThe placings are on the right.

Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where o3 placed and give it no mark out of five.

Arena Creative Writing 65th of 143 · 1383.1
Also scored, on boards we give no mark for
Arena Instruction Following 68th of 143

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
84.7via Highindependentsource ↗
1459.4independentsource ↗
1440.1independentsource ↗
1447.1independentsource ↗
1430.9independentsource ↗
58.4via mini-SWE-agentindependentsource ↗
01

Or rent it from someone else

Cheapest published offer

Why this differs from the header. The strip above quotes OpenAI's own list price. This is the cheapest live offer at the widest standard context we hold, whoever is serving it — a reseller undercutting a lab is ordinary commerce, not an error.

per 1M tokens
$2.00 in / $8.00 out
Context served
200K
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouter$2.00 / $8.00200Knot measuredUnknownUnknownUnknown
OpenAI$2.00 / $8.00200K100K out71 tok/sNoYesunknown periodUnknown

Across the 2 listings we hold: 1 say they do not train on prompts, 0 say they do and 1 do not say. 0 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

Supported
Not supported
Not published
host gave no parameter list
API features per host
ProviderTool callingJSON outputStrict schema
OpenRouter
OpenAI

Tool calling: 2 of 2 listings say yes. JSON output: 2 of 2 listings say yes. Strict schema: 2 of 2 listings say yes.

02

Models people weigh against o3

Other sizes and versionso1
Succeedso1

Compare them field by field →

03

When we formed this view

Dates behind this page

Aug 3, 2026BenchmarkScored 84.7 via High on LiveCodeBenchleaderboard
Aug 2, 2026BenchmarkScored 1459.4 on Arena Codingleaderboard
Aug 2, 2026BenchmarkScored 1383.1 on Arena Creative Writingleaderboard
Aug 2, 2026BenchmarkScored 1440.1 on Arena Hard Promptsleaderboard
Aug 2, 2026BenchmarkScored 1401.9 on Arena Instruction Followingleaderboard
Aug 2, 2026BenchmarkScored 1447.1 on Arena Mathsleaderboard
Aug 2, 2026BenchmarkScored 1430.9 on Arena Text (overall)leaderboard
Jul 26, 2026ListedListed on LLMapfirst indexed by our pipeline
Jul 26, 2025BenchmarkScored 58.4 via mini-SWE-agent on SWE-bench Verifiedleaderboard
Apr 16, 2025Announcedo3 announced by OpenAI

Prices last checked 9d ago

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 2 listings do not say whether they train on prompts.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

Commercial API terms. We hold no licence record for this model, so there is nothing to summarise here.

Identifiers

Modality record
text+image+file->text
Catalogue slug
openai-o3

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us