Gemma 4 31B
Google · released Mar 11, 2026 · google/gemma-4-31B-it
- Type
- Open weightsApache License 2.0
- Params
- 31.3B
- Context
- 262K
about 197K words of context
Our take
Written Sep 2, 2026Gemma 4 is a 31.3-billion-parameter text, image and video model from Google with a permissive Apache licence and a quarter-million-token request limit. Its measured coding skill outpaces its general text score, though it struggles on agent tasks and its speed varies sharply by provider.
Pick this for coding work where its measured score is strongest, or for long-context multimodal tasks with a licence that allows commercial use and redistribution. Use it when you want 23 provider options to shop between. Skip it if you need reliable agentic behaviour, web-development coding, or guaranteed fast throughput without checking each host.
The case for it
- Coding is its standout measured skill, 47.9 points above its general text score.
- Apache 2.0 licence with 23 tracked offers across seven-plus providers.
- 262,144-token request limit, large for its parameter class.
- Maths and hard-prompt scores are competitive with its coding peak.
The case against it
- Agent tasks score negatively across every measured subdimension, including recovery and tool use.
- Web-development coding sits 136 points below its general coding score.
- Throughput varies 2.8× across cheapest offers, so speed is not portable between hosts.
How good is it?
An open text model for writing and everyday questions, though multi-step tasks and tool calls are not its strength.
- carrying out multi-step tasks for youArena Agent · 55th of 55
- calling tools to carry out requestsArena Agent · Tool use · 55th of 55
- changing course when you give new instructionsArena Agent · Steerability · 51st of 55
- getting back on track after a step failsArena Agent · Recovery · 55th of 55
EverydayGeneral questions and everyday reasoning
Arena Text (overall)48th of 168 · 1453
CodingWriting and fixing code on its own
Arena Coding49th of 168 · 1502
AgenticPlanning, calling tools, staying on task
Arena Agent55th of 55 · −0.237
Arena Agent is the only board that has scored it for this.
WritingDrafting and rewriting prose
Arena Creative Writing50th of 168 · 1419
Arena Creative Writing is the only board that has scored it for this.
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model12 scoresEvery figure we hold, from 12 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Loads, but do not expect an assistant. Some of the weights sit in ordinary system memory, which is far slower than the card.19.7 GB of weights, plus 3.8 GB for the software that runs it and the smallest conversation it can hold, comes to 23.5 GB against the 22.8 GB this 24 GB device leaves free.
Comfortable fit
GeForce RTX 5090 · 32 GB
Room to spare. 7.3 GB spare means a 10% error in the size would not change the answer.
Apple M1 Pro (16-core GPU) · 32 GB
Borderline fit on an estimated size. It leaves 0.5 GB spare on a size we calculated rather than measured, and a 10% error either way would change the answer.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 4 hours and 4 days ago — each listing carries its own date.
- per 1M tokens
- $0.10 in / $0.33 out
- Context served
- 262K
- Throughput
- ~11 tok/s
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| DekaLLMThrough OpenRouter | $0.10 / $0.33checked 16 hours ago | 262K236K max reply | 11 tok/s | No | No | Confirmed |
| DeepInfraturbo tierfp4Through OpenRouter | $0.090 / $0.34checked 4 hours ago | 262K16K max reply | 23 tok/s | No | No | Unknown |
| OpenRouterOpenRouter's own listing | $0.090 / $0.34checked 4 hours ago | 262K | not measured | Unknown | Unknown | Unknown |
| CoreWeavefp4Through OpenRouter | $0.10 / $0.34checked 4 hours ago | 262K236K max reply | 28 tok/s | No | No | Confirmed |
| Venice AIfp4Through OpenRouter | $0.12 / $0.36checked 4 hours ago | 256K8K max reply | 18 tok/s | No | No | Confirmed |
| Chutesfp4Through OpenRouter | $0.12 / $0.37checked 22 hours ago | 131K66K max reply | 19 tok/s | No | Yesunknown period | Unknown |
| DeepInfrafp8Direct and through OpenRouter | $0.15 / $0.40checked 4 hours ago | 262K16K max reply through OpenRouter | 10 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Novita AIbf16Direct and through OpenRouter | $0.14 / $0.40checked 4 hours ago directchecked 4 days ago through OpenRouter | 262K131K max reply through OpenRouter | 8 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Crusoebf16Through OpenRouter | $0.14 / $0.40checked 4 hours ago | 262K262K max reply | 14 tok/s | No | No | Confirmed |
| Parasailfp8Through OpenRouter | $0.15 / $0.40checked 4 hours ago | 262K236K max reply | 17 tok/s | No | No | Confirmed |
| FriendliThrough OpenRouter | $0.14 / $0.40checked 4 hours ago | 262K8K max reply | 50 tok/s | No | Yesunknown period | Unknown |
| SiliconFlowfp8Through OpenRouter | $0.75 / $1.00checked 10 hours ago | 262K236K max reply | 38 tok/s | No | No | Confirmed |
| ModelRunfp4Through OpenRouter | $0.75 / $1.00checked 4 hours ago | 262K236K max reply | 126 tok/s | No | No | Confirmed |
| SambaNovaDirect | $0.38 / $1.15checked 4 hours ago | 131K | not measured | Unknown | Unknown | Unknown |
| Io NetThrough OpenRouter | $0.38 / $1.15checked 4 hours ago | 262K16K max reply | 27 tok/s | No | No | Confirmed |
| SambaNovaThrough OpenRouter | $0.38 / $1.15checked 10 hours ago | 262K236K max reply | 28 tok/s | No | No | Confirmed |
Across the 16 listings we hold: 14 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 2 do not say. 11 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| DekaLLMThrough OpenRouter | ✓ | ✓ | ✓ |
| DeepInfraturbo · fp4Through OpenRouter | ✗ | ✓ | ✓ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| CoreWeavefp4Through OpenRouter | ✓ | ✓ | ✓ |
| Venice AIfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Chutesfp4Through OpenRouter | ✓ | ✓ | ✓ |
| DeepInfrafp8Direct and through OpenRouter | ✓ | ✓ | ✓ |
| Novita AIbf16Direct and through OpenRouter | ✓ | ✓ | ✓ |
| Crusoebf16Through OpenRouter | ✓ | ✓ | ✓ |
| Parasailfp8Through OpenRouter | ✓ | ✓ | ✓ |
| FriendliThrough OpenRouter | ✓ | ✓ | ✓ |
| SiliconFlowfp8Through OpenRouter | ✓ | ✓ | ✓ |
| ModelRunfp4Through OpenRouter | ✓ | ✓ | ✓ |
| SambaNovaDirect | |||
| Io NetThrough OpenRouter | ✓ | ✓ | ✗ |
| SambaNovaThrough OpenRouter | ✓ | ✗ | ✗ |
Tool calling: 14 of 16 listings say yes, 1 says no, 1 publishes no parameter list. JSON output: 14 of 16 listings say yes, 1 says no, 1 publishes no parameter list. Strict schema: 13 of 16 listings say yes, 2 say no, 1 publishes no parameter list.
Models people weigh against Gemma 4 31B
When we formed this view
Recent changes
What moved
Gemma 4 31B moved on 2 hosts: DekaLLM: input +67% ($0.060 → $0.100 per 1M tokens); DeepInfra through OpenRouter (fp8): input +15% ($0.13 → $0.15 per 1M tokens), output +5% ($0.38 → $0.40 per 1M tokens); DeepInfra's own listing (fp8): input +15% ($0.13 → $0.15 per 1M tokens), output +5% ($0.38 → $0.40 per 1M tokens)What moved
input −11% ($0.090 → $0.080 per 1M tokens), output −12% ($0.34 → $0.30 per 1M tokens)What moved
input +477% ($0.13 → $0.75 per 1M tokens), output +150% ($0.40 → $1.00 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- 1 of 16 listings publishes no parameter list, so what its API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 16 listings do not say whether they train on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsApache License 2.0, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
Apache License 2.0
Fully permissive: commercial use, redistribution, and derivatives allowed. Requires attribution and a copy of the license. Includes an express patent grant.
Identifiers
- Hugging Face
- google/gemma-4-31B-it
- Architecture
- Dense
- Takes in, gives back
- Text, images and video in, text out
- Catalogue slug
- google-gemma-4-31b