Qwen3.6 35B A3B
Qwen · released Apr 15, 2026 · Qwen/Qwen3.6-35B-A3B
- Type
- Open weightsApache License 2.0
- Params
- 36B
- Context
- 262K
3B active per word · about 197K words of context
Our take
Written Sep 2, 2026Qwen 3.6 is a multimodal model with a mixture-of-experts design that keeps only 3 billion parameters active for each word while drawing on 36 billion total. It carries a permissive Apache licence and is widely available from thirteen hosts, though no benchmark scores have been published yet.
Pick this for self-hosted or local deployment where low active-parameter cost matters, or when you need the cheapest input rate among its tracked hosts. Choose the higher-throughput hosts — Atlas Cloud, IO Net, Venice or CoreWeave — when speed matters more than output price. Skip it if you need verified quality scores, or if you want the cheapest host and the fastest host to be the same.
The case for it
- Only 3 billion parameters active per word from 36 billion total — a 12:1 ratio that keeps inference light.
- Apache 2.0 licence allows commercial use, fine-tuning and redistribution.
- Thirteen hosted offers with a fivefold spread on input cost, and several hosts above 110 tokens per second.
The case against it
- No benchmark scores in our data — chat, coding, reasoning and multimodal quality are all unverified.
- The cheapest host is not the fastest: Darkbloom has the lowest input cost but only 26 tokens per second, while Atlas Cloud reaches 165 tokens per second at a higher output rate.
How good is it?
We hold no score for this model.
So there is no figure here for everyday use, coding, agent work or writing. That is a gap in our data, not a low score.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Loads, but do not expect an assistant. Some of the weights sit in ordinary system memory, which is far slower than the card.22.7 GB of weights, plus 2.1 GB for the software that runs it and the smallest conversation it can hold, comes to 24.8 GB against the 22.8 GB this 24 GB device leaves free.
Comfortable fit
GeForce RTX 5090 · 32 GB
Room to spare. 6.1 GB spare means a 10% error in the size would not change the answer.
Apple M3 Pro (18-core GPU) · 36 GB
Borderline fit on an estimated size. It leaves 2.3 GB spare on a size we calculated rather than measured, and a 10% error either way would change the answer.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 4 hours and 10 hours ago — each listing carries its own date.
Cheapest of the 3 listings we can compare like for like — at 262K of context, out of 11 in the table below. 3 cheaper rows there are outside that comparison: a different quantisation.
- per 1M tokens
- $0.15 in / $1.00 out
- Context served
- 262K
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| Darkbloomfp4Through OpenRouter | $0.050 / $0.70checked 4 hours ago | 262K33K max reply | 33 tok/s | No | Yesunknown period | Unknown |
| AkashMLfp8Through OpenRouter | $0.10 / $0.90checked 4 hours ago | 262K236K max reply | 51 tok/s | No | No | Confirmed |
| DeepInfrafp8Direct and through OpenRouter | $0.10 / $0.95checked 4 hours ago directchecked 10 hours ago through OpenRouter | 262K16K max reply through OpenRouter | 30 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| OpenRouterOpenRouter's own listing | $0.15 / $1.00checked 4 hours ago | 262K | not measured | Unknown | Unknown | Unknown |
| Venice AIfp8Through OpenRouter | $0.10 / $1.00checked 4 hours ago | 256K66K max reply | 122 tok/s | No | No | Confirmed |
| Parasailfp8Through OpenRouter | $0.15 / $1.00checked 4 hours ago | 262K236K max reply | 29 tok/s | No | No | Confirmed |
| AtlasCloudfp8Through OpenRouter | $0.19 / $1.11checked 4 hours ago | 262K66K max reply | 121 tok/s | No | Yesunknown period | Unknown |
| CoreWeavefp8Through OpenRouter | $0.25 / $1.25checked 4 hours ago | 262K236K max reply | 131 tok/s | No | No | Confirmed |
| PhalaThrough OpenRouter | $0.20 / $1.27checked 4 hours ago | 262K236K max reply | 97 tok/s | No | No | Confirmed |
| Novita AIDirect | $0.25 / $1.49checked 4 hours ago | 262K | not measured | Unknown | Unknown | Unknown |
| SiliconFlowfp8Through OpenRouter | $0.24 / $1.80checked 4 hours ago | 262K236K max reply | 81 tok/s | No | No | Confirmed |
Across the 11 listings we hold: 9 say they do not train on prompts (1 of them only through OpenRouter), 0 say they do and 2 do not say. 7 appear in the zero-retention registry we check (1 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| Darkbloomfp4Through OpenRouter | ✓ | ✓ | ✓ |
| AkashMLfp8Through OpenRouter | ✓ | ✓ | ✓ |
| DeepInfrafp8Direct and through OpenRouter | ✓ | ✓ | ✓ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Venice AIfp8Through OpenRouter | ✓ | ✗ | ✗ |
| Parasailfp8Through OpenRouter | ✓ | ✓ | ✓ |
| AtlasCloudfp8Through OpenRouter | ✓ | ✓ | ✗ |
| CoreWeavefp8Through OpenRouter | ✗ | ✓ | ✓ |
| PhalaThrough OpenRouter | ✓ | ✗ | ✓ |
| Novita AIDirect | |||
| SiliconFlowfp8Through OpenRouter | ✗ | ✓ | ✓ |
Tool calling: 8 of 11 listings say yes, 2 say no, 1 publishes no parameter list. JSON output: 8 of 11 listings say yes, 2 say no, 1 publishes no parameter list. Strict schema: 8 of 11 listings say yes, 2 say no, 1 publishes no parameter list.
Models people weigh against Qwen3.6 35B A3B
When we formed this view
Recent changes
What moved
input −33% ($0.15 → $0.10 per 1M tokens)What moved
input +43% ($0.133 → $0.190 per 1M tokens), output +5% ($0.94 → $0.99 per 1M tokens), cache read +49% ($0.0665 → $0.0990 per 1M tokens)What moved
input −23% ($0.13 → $0.10 per 1M tokens)What moved
input +20% ($0.20 → $0.24 per 1M tokens), output +12% ($1.60 → $1.80 per 1M tokens)What moved
input −26% ($0.19 → $0.14 per 1M tokens), output −17% ($1.19 → $0.99 per 1M tokens), cache read −22% ($0.090 → $0.070 per 1M tokens)What moved
Qwen3.6 35B A3B moved on 2 hosts: AkashML: input −29% ($0.14 → $0.10 per 1M tokens), output −10% ($1.00 → $0.90 per 1M tokens); Darkbloom: input −29% ($0.070 → $0.050 per 1M tokens)What moved
input +2% ($0.098 → $0.100 per 1M tokens), output +5% ($0.95 → $1.00 per 1M tokens)What moved
input −34% ($0.29 → $0.19 per 1M tokens), output −37% ($1.89 → $1.19 per 1M tokens), cache read −18% ($0.110 → $0.090 per 1M tokens)What moved
Qwen3.6 35B A3B moved on 2 hosts: Io Net: input −11% ($0.325 → $0.290 per 1M tokens), output −10% ($2.09 → $1.89 per 1M tokens), cache read −32% ($0.162 → $0.110 per 1M tokens); Venice: input −2% ($0.100 → $0.098 per 1M tokens), output −5% ($1.00 → $0.95 per 1M tokens)What moved
input +48% ($0.219 → $0.325 per 1M tokens), output +72% ($1.215 → $2.090 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- No independent board has scored it, so we hold no quality figures at all.
- 1 of 11 listings publishes no parameter list, so what its API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 11 listings do not say whether they train on prompts, and 1 answers only through OpenRouter, not for its own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsApache License 2.0, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
Apache License 2.0
Fully permissive: commercial use, redistribution, and derivatives allowed. Requires attribution and a copy of the license. Includes an express patent grant.
Identifiers
- Hugging Face
- Qwen/Qwen3.6-35B-A3B
- Architecture
- Mixture of experts
- Takes in, gives back
- Text, images and video in, text out
- Catalogue slug
- qwen-qwen3-6-35b-a3b