Gemma 4 26B A4B
Google · released Mar 11, 2026 · google/gemma-4-26B-A4B-it
- Type
- Open weightsApache License 2.0
- Params
- 25.8B
- Context
- 262K
4B active per word · about 197K words of context
Our take
Written Aug 4, 2026Google's Gemma is a mid-size downloadable language model with a permissive Apache licence. It uses an efficient mixture-of-experts design that activates only a few billion parameters per token, making it the efficiency pick of the mid-size class.
Make this your default local mid-size pick: measured chat quality, a permissive licence and an efficient design. It is also cheap to rent for multimodal work, and runs at the edge on Cloudflare Workers AI. Skip it if you need frontier-level answers or the widest choice of hosts.
The case for it
- Measured chat quality unusually close to models many times larger.
- Apache 2.0 licence allows commercial use, fine-tuning and redistribution.
- Only about four billion active parameters per token from a 26.5-billion total, making it efficient for local use.
The case against it
- Trails the frontier on measured chat quality.
How good is it?
EverydayGeneral questions and everyday reasoning
Arena Text (overall)64th of 168 · 1438
CodingWriting and fixing code on its own
Arena Coding66th of 168 · 1483
AgenticPlanning, calling tools, staying on task
Not yet scored on Arena Agent.
WritingDrafting and rewriting prose
Arena Creative Writing70th of 168 · 1399
Arena Creative Writing is the only board that has scored it for this.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done.
Every published score for this model7 scoresEvery figure we hold, from 7 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
Comfortable fit
GeForce RTX 4090 · 24 GB
Room to spare. 4.3 GB spare means a 10% error in the size would not change the answer.
GeForce RTX 5090 · 32 GB
Room to spare. 12.3 GB spare means a 10% error in the size would not change the answer.
Apple M1 Pro (16-core GPU) · 32 GB
Room to spare. 5.5 GB spare means a 10% error in the size would not change the answer.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 4 hours and 16 hours ago — each listing carries its own date.
Cheapest of the 2 listings we can compare like for like — at 262K of context, out of 14 in the table below. One cheaper row there is outside that comparison: a different context length.
- per 1M tokens
- $0.076 in / $0.26 out
- Context served
- 262K
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| DarkbloomThrough OpenRouter | $0.042 / $0.22checked 4 hours ago | 131K33K max reply | 41 tok/s | No | Yesunknown period | Unknown |
| OpenRouterOpenRouter's own listing | $0.076 / $0.26checked 4 hours ago | 262K | not measured | Unknown | Unknown | Unknown |
| NextBitbf16Through OpenRouter | $0.076 / $0.26checked 4 hours ago | 262K236K max reply | 60 tok/s | No | No | Confirmed |
| Cloudflare Workers AIThrough OpenRouter | $0.10 / $0.30checked 4 hours ago | 256K230K max reply | 49 tok/s | No | Yesunknown period | Unknown |
| CoreWeavebf16Through OpenRouter | $0.10 / $0.30checked 4 hours ago | 262K236K max reply | 74 tok/s | No | No | Confirmed |
| MakoraThrough OpenRouter | $0.080 / $0.32checked 4 hours ago | 256K128K max reply | 57 tok/s | No | No | Confirmed |
| DekaLLMbf16Through OpenRouter | $0.060 / $0.33checked 4 hours ago | 262K236K max reply | 73 tok/s | No | No | Confirmed |
| DeepInfrafp8Direct and through OpenRouter | $0.070 / $0.34checked 4 hours ago | 262K16K max reply through OpenRouter | 22 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| SiliconFlowfp8Through OpenRouter | $0.14 / $0.40checked 4 hours ago | 262K236K max reply | 28 tok/s | No | No | Confirmed |
| Venice AIbf16Through OpenRouter | $0.13 / $0.40checked 4 hours ago | 256K8K max reply | 26 tok/s | No | No | Confirmed |
| Novita AIbf16Direct and through OpenRouter | $0.13 / $0.40checked 4 hours ago | 262K131K max reply through OpenRouter | 32 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Parasailbf16Through OpenRouter | $0.13 / $0.40checked 4 hours ago | 262K236K max reply | 24 tok/s | No | No | Confirmed |
| Io Netbf16Through OpenRouter | $0.15 / $0.50checked 10 hours ago | 262K236K max reply | 49 tok/s | No | No | Confirmed |
| Google Vertex AIglobalThrough OpenRouter | $0.15 / $0.60checked 16 hours ago | 262K236K max reply | 23 tok/s | No | No | Confirmed |
Across the 14 listings we hold: 13 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 1 does not say. 11 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| DarkbloomThrough OpenRouter | ✓ | ✓ | ✓ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| NextBitbf16Through OpenRouter | ✓ | ✓ | ✓ |
| Cloudflare Workers AIThrough OpenRouter | ✓ | ✓ | ✗ |
| CoreWeavebf16Through OpenRouter | ✗ | ✓ | ✓ |
| MakoraThrough OpenRouter | ✓ | ✓ | ✗ |
| DekaLLMbf16Through OpenRouter | ✓ | ✗ | ✗ |
| DeepInfrafp8Direct and through OpenRouter | ✓ | ✓ | ✓ |
| SiliconFlowfp8Through OpenRouter | ✗ | ✓ | ✓ |
| Venice AIbf16Through OpenRouter | ✓ | ✓ | ✓ |
| Novita AIbf16Direct and through OpenRouter | ✓ | ✓ | ✗ |
| Parasailbf16Through OpenRouter | ✗ | ✓ | ✓ |
| Io Netbf16Through OpenRouter | ✓ | ✗ | ✗ |
| Google Vertex AIglobalThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 11 of 14 listings say yes, 3 say no. JSON output: 12 of 14 listings say yes, 2 say no. Strict schema: 9 of 14 listings say yes, 5 say no.
Models people weigh against Gemma 4 26B A4B
When we formed this view
Recent changes
What moved
input −33% ($0.090 → $0.060 per 1M tokens), output −33% ($0.30 → $0.20 per 1M tokens), cache read −30% ($0.050 → $0.035 per 1M tokens)What moved
Gemma 4 26B A4B moved on 2 hosts: NextBit: input −25% ($0.0900 → $0.0675 per 1M tokens), output −25% ($0.300 → $0.225 per 1M tokens), cache read −25% ($0.0500 → $0.0375 per 1M tokens); Makora: input −20% ($0.100 → $0.080 per 1M tokens), output −6% ($0.34 → $0.32 per 1M tokens), cache read −6% ($0.034 → $0.032 per 1M tokens)What moved
input −10% ($0.100 → $0.090 per 1M tokens), output −25% ($0.40 → $0.30 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 14 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsApache License 2.0, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
Apache License 2.0
Fully permissive: commercial use, redistribution, and derivatives allowed. Requires attribution and a copy of the license. Includes an express patent grant.
Identifiers
- Hugging Face
- google/gemma-4-26B-A4B-it
- Architecture
- Mixture of experts
- Takes in, gives back
- Text, images and video in, text out
- Catalogue slug
- google-gemma-4-26b-a4b