Alibaba's Qwen3.8 Max makes a clear jump in benchmark performance, but the numbers also show why model rankings are not only about the top score. According to Artificial Analysis, the new model moves closer to leading systems while becoming more expensive per task than its predecessor.
A Big Step Up On The Intelligence Index
Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index. That is a 10-point jump over Qwen3.7 Max (46), which makes the upgrade meaningful on the benchmark used in the source.
That score puts Qwen3.8 Max on par with Claude Opus 4.8. It also places the model ahead of GLM-5.2 (51), while still behind Kimi K3 (57).
The gap to Kimi K3 is small on the headline score: one point. But the comparison becomes more complicated once cost is included, because Kimi K3 also runs 25 percent cheaper.
For users comparing AI models, this is the central tension. Qwen3.8 Max looks much stronger than Qwen3.7 Max, but the model that scores slightly higher also has the cost advantage in the reported comparison.
Work-Related Tasks Show A Different Strength
On GDPval-AA, a benchmark for work-related tasks, Qwen3.8 Max shows an especially large gain. The model jumps 468 Elo points to 1,739, moving past Kimi K3 (1,685).
Only Claude Opus 5 (1,852) scores higher on that benchmark. That makes GDPval-AA the area where Qwen3.8 Max appears particularly competitive in the source data.
The result matters because GDPval-AA focuses on work-related tasks rather than a broad overall score. A model can look different depending on whether the test emphasizes general intelligence, long-context retrieval, knowledge reliability, or task execution.
Still, the way Qwen3.8 Max reaches that GDPval-AA score has consequences. The source reports that it needs 64 steps per task instead of 14. The input token load also grew 15x because the full conversation history is resent to the model at each step.
That means the stronger work-task score comes with heavier execution. More steps and more input tokens can make a model feel more deliberate, but they also affect speed and cost.
Lower Token Prices Do Not Mean Lower Task Costs
Alibaba reduced token prices for Qwen3.8 Max compared with the earlier pricing stated in the source. Input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25.
On paper, those are price reductions. In practice, the model's greater task workload changes the economics.
A single task in the Artificial Analysis Intelligence Index now costs $1.14 for Qwen3.8 Max. That is more than double Qwen3.7 Max ($0.53).
The comparison with other models is also important:
- Qwen3.8 Max costs $1.14 per task on the Intelligence Index.
- Kimi K3 scores one point higher at $0.86 per task.
- GLM-5.2 comes in at $0.57 per task.
This is why the price-to-performance picture is weaker than the token price cuts might suggest. If a model needs far more steps and a much larger amount of input, cheaper tokens alone may not reduce the final cost of getting a task done.
Regressions Cloud The Upgrade
The source also reports regressions compared with the previous Qwen version. AA-LCR dropped 2 points. That test checks whether a model can correctly pull together information from very long texts.
AA-Omniscience fell 10 points. That benchmark measures whether a model answers knowledge questions correctly or honestly admits it does not know.
The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent. In plain terms, Qwen3.8 Max guesses far more often instead of saying it does not know.
That matters because benchmark gains are not interchangeable. A model can improve on one major index while becoming less reliable in a specific behavior that users care about. For knowledge-heavy work, a higher tendency to guess can be a serious drawback even when other scores improve.
The Practical Read
Qwen3.8 Max is a stronger model than Qwen3.7 Max by the measures reported in the source. It rises sharply on the Artificial Analysis Intelligence Index, matches Claude Opus 4.8 there, and performs strongly on GDPval-AA.
But the ranking is not a clean win. Kimi K3 remains ahead on the Intelligence Index with 57, costs $0.86 per task compared with Qwen3.8 Max at $1.14, and is described as 25 percent cheaper.
The result is a model that looks more capable but less efficient. Qwen3.8 Max works more thoroughly, uses more steps, consumes far more input tokens in the described test setup, and costs more per Intelligence Index task than Qwen3.7 Max.
For teams evaluating Qwen3.8 Max, the takeaway from the source data is simple: the upgrade is real, but so are the tradeoffs. The model's best results come with slower, heavier, and more expensive task execution, while Kimi K3 still holds the edge on the overall score and cost comparison provided by Artificial Analysis.