Models / OpenAI/ GPT-4

GPT-4

OpenAI · released May 28, 2023

Input: text. Output: text.InputOutput
Type
Closed
Input
$30.00
Output
$60.00
Cached
None held

List price · per 1M tokens · OpenAI at 8K context · machine-readable source ↗

Our take

Written Sep 2, 2026

GPT-4 is OpenAI's 2023 flagship text model, now a legacy tier with a modest 8,191-token request limit and identical pricing across all three tracked hosts. Its measured coding and chat scores sit in the lower-middle of current options, making it a holdover for existing integrations rather than a fresh pick.

Who should pick it

Keep this for legacy integrations still pinned to the GPT-4 API contract, or basic coding assistance where resolving about one in five GitHub issues end-to-end is sufficient. Skip it if you need long-document work, multimodal input, price competition between hosts, or current-generation quality.

The case for it

  • Fastest output among its own tracked hosts at 21 tokens per second on OpenAI, versus 15 on Azure.
  • Arena Coding score of 1329.53 is its strongest measured split, though this does not lift the overall score.

The case against it

  • Request limit of 8,191 tokens is far below current standards for long-document work, with no extended variant offered.
  • Resolves only 22.4% of GitHub issues end-to-end, suggesting limited real-world coding autonomy.
  • Identical pricing across all three offers means no competitive pressure on cost.
00

How good is it?

An older general-purpose text model from OpenAI, now behind most rivals at everyday questions, drafting and coding.

Less good at
  • answering everyday questionsArena Text (overall) · 154th of 168
  • drafting and editing textArena Creative Writing · 143rd of 168
  • writing and completing codeArena Coding · 150th of 168

EverydayGeneral questions and everyday reasoning

1 of 5

Arena Text (overall)154th of 168 · 1288

Arena Hard Prompts 148th of 168Arena Maths 140th of 163

CodingWriting and fixing code on its own

1 of 5

Arena Coding150th of 168 · 1330

Arena Coding is the only board that has scored it for this.

AgenticPlanning, calling tools, staying on task

Scored, not ratedSWE-bench Verified · 38th of 42 · 22.4

Not yet scored on Arena Agent. It is on SWE-bench Verified, in 38th of 42 with 22.4.

WritingDrafting and rewriting prose

1 of 5

Arena Creative Writing143rd of 168 · 1267

Arena Creative Writing is the only board that has scored it for this.

Other boards it appears on
Arena Instruction Following 146th of 168

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done.

Every published score for this model7 scoresEvery figure we hold, from 7 boards, with who ran it and a link to the source — including the boards no rating above is built on.
1330source ↗
1267source ↗
1306source ↗
1284source ↗
1288source ↗
22.4source ↗
01

Where to rent it

Prices checked 4 hours ago — each listing carries its own date.

Cheapest published offer

OpenAI, direct

The lab is also the cheapest we hold. The strip above and this offer are the same one, so nothing on this page undercuts OpenAI on 8K of context.

per 1M tokens
$30.00 in / $60.00 out
Context served
8K
Throughput
~30 tok/s
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouterOpenRouter's own listing$30.00 / $60.00checked 4 hours ago8Knot measuredUnknownUnknownUnknown
Microsoft Azure AIThrough OpenRouter$30.00 / $60.00checked 4 hours ago8K4K max reply83 tok/sNoNoConfirmed
OpenAIDirect$30.00 / $60.00checked 4 hours ago8K4K max reply30 tok/sNoYesunknown periodUnknown

Across the 3 listings we hold: 2 say they do not train on prompts, 0 say they do and 1 does not say. 1 appears in the zero-retention registry we check; the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
OpenRouterOpenRouter's own listing✓✓✓
Microsoft Azure AIThrough OpenRouter✓✓✗
OpenAIDirect✓✓✓

Tool calling: 3 of 3 listings say yes. JSON output: 3 of 3 listings say yes. Strict schema: 2 of 3 listings say yes, 1 says no.

02

Models people weigh against GPT-4

03

When we formed this view

Recent changes

Sep 25, 2026BenchmarkScored 1330 on Arena Coding
What movedleaderboard
Sep 25, 2026BenchmarkScored 1267 on Arena Creative Writing
What movedleaderboard
Sep 25, 2026BenchmarkScored 1306 on Arena Hard Prompts
What movedleaderboard
Sep 25, 2026BenchmarkScored 1285 on Arena Instruction Following
What movedleaderboard
Sep 25, 2026BenchmarkScored 1284 on Arena Maths
What movedleaderboard
Sep 25, 2026BenchmarkScored 1288 on Arena Text (overall)
What movedleaderboard
Jul 26, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline
Apr 2, 2024BenchmarkScored 22.4 via SWE-agent on SWE-bench Verified
What movedleaderboard
May 28, 2023AnnouncedGPT-4 announced by OpenAI

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 3 listings does not say whether it trains on prompts.
  • We hold no cached-input rate for any of its listings.
  • We hold no batch or off-peak rate for any of its listings.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text in, text out
Catalogue slug
openai-gpt-4

Machine-readable model card (omc.json) →

Explore furtherWhere to run GPT-4
Something wrong on this page? Tell us