Qwen3.8 Flash
Qwen · released Aug 26, 2026
- Type
- Closed
- Input
- None held
- Output
- None held
- Cached
- None held
We don't hold a list price for this model yet · hosted only — we have no record of published weights
Our take
Written Sep 27, 2026Qwen3.8 Flash is a hosted-only model with a lopsided profile: it follows instructions very well and builds web apps well, but it is a weak agent. We list no download for it, so using it means choosing a host.
Use it for constrained rewriting and format-following work, where it places 4th of 51 on LiveBench Instruction Following as of 25 Jun 2026, or for building web apps, where it places 9th of 95 on Arena Code (WebDev) as of 25 Sep 2026. Skip it if you need an agent that reliably calls the right tool, if you want to run the model yourself, or if you need licence terms you can build a product on.
The case for it
- Among the strongest here at following instructions: 77.11% on LiveBench Instruction Following, 4th of 51 as of 25 Jun 2026, a board of constrained-rewriting tasks rather than open-ended judgement.
- Strong on set-piece maths and reasoning: 87.38% on LiveBench Reasoning and 85.82% on LiveBench Mathematics as of 25 Jun 2026, competition-style questions rather than reasoning about your own codebase.
- Good at building web apps from a prompt: 9th of 95 on Arena Code (WebDev) as of 25 Sep 2026, a board of human votes that records which result people preferred, not whether the app was correct.
- Pictures and clips go in with the question, and the request capacity takes a long report or a stack of documents beside it, though reliable recall across all of it is unverified in our data.
The case against it
- Weak at agentic tool use: 45th of 55 on Arena Agent · Tool use as of 25 Sep 2026, a board scoring whether the model calls the right tool and does not invent one, so it is a poor fit for tool-calling pipelines.
- Mid-to-lower field on general coding and language: 38th of 51 on LiveBench Coding and 37th of 51 on LiveBench Language as of 25 Jun 2026, against its 4th of 51 on Instruction Following.
- We list no download for it, so using it means choosing a host, and no licence is supplied, so nothing here tells you what you may do with its output.
How good is it?
A text model for chat and everyday questions, though it trails most models at calling tools to carry out requests.
- calling tools to carry out requestsArena Agent · Tool use · 45th of 55
EverydayGeneral questions and everyday reasoning
Not yet scored on Arena Text (overall). It is on LiveBench Reasoning, in 24th of 58 with 87.38.
CodingWriting and fixing code on its own
Not yet scored on Arena Coding. It is on Arena Code (WebDev), in 9th of 95 with 1636.
AgenticPlanning, calling tools, staying on task
Arena Agent29th of 55 · −0.005
WritingDrafting and rewriting prose
Not yet scored on Arena Creative Writing. It is on LiveBench Language, in 43rd of 58 with 74.64.
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model14 scoresEvery figure we hold, from 14 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to rent it
Prices checked 4 hours ago — each listing carries its own date.
- per 1M tokens
- $0.11 in / $0.38 out
- Context served
- 1M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| DeepInfraDirect | $0.11 / $0.38checked 4 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| OpenRouterOpenRouter's own listing | $0.15 / $0.47checked 4 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| Novita AIDirect | $0.15 / $0.47checked 4 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| Alibaba CloudThrough OpenRouter | $0.15 / $0.47checked 4 hours ago | 1M131K max reply | 45 tok/s | No | Yesunknown period | Unknown |
Across the 4 listings we hold: 1 says it does not train on prompts, 0 say they do and 3 do not say. 0 appear in the zero-retention registry we check; the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| DeepInfraDirect | |||
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Novita AIDirect | |||
| Alibaba CloudThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 2 of 4 listings say yes, 2 publish no parameter list. JSON output: 2 of 4 listings say yes, 2 publish no parameter list. Strict schema: 2 of 4 listings say yes, 2 publish no parameter list.
Models people weigh against Qwen3.8 Flash
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- 2 of 4 listings publish no parameter list, so what their API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 3 of 4 listings do not say whether they train on prompts.
- We don't hold a list price for this model yet — the gap is ours, not the lab's.
- We hold no batch or off-peak rate for any of its listings.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.
Identifiers
- Takes in, gives back
- Text, images and video in, text out
- Catalogue slug
- qwen-qwen3-8-flash