llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Alibaba Cloud / Qwen Team model product

Qwen3 235B A22B Thinking

Qwen3-235B-A22B-Thinking-2507 is a state-of-the-art thinking-enabled Mixture-of-Experts (MoE) model with 235B total parameters (22B activated).

Updated Aug 17, 2026. Default version: Qwen3-235B-A22B-Thinking-2507

Compare
LLMBoard Score48.0Qwen3-235B-A22B-Thinking-2507
Coverage100%24 benchmark families
Context window262.1KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Compare
  • Similar models
  • About
  • FAQ

Qwen3 235B A22B Thinking Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Qwen3-235B-A22B-Thinking-2507 LLMBoard score breakdown

Qwen3 235B A22B Thinking Benchmark Results

Benchmark scores for Qwen3-235B-A22B-Thinking-2507.

25 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkCFEvalScore2,134 pointsRank01Participants2Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMulti-IFScore80.6%Rank01Participants23Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkWritingBenchScore88.3%Rank01Participants15Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkLiveBench 20241125Score78.4%Rank02Participants14Percentile92.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkArena-Hard v2Score79.7%Rank03Participants16Percentile86.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkCreative Writing v3Score86.1%Rank03Participants13Percentile83.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkBFCL-v3Score71.9%Rank06Participants19Percentile72.2%EvidenceCEvaluatedAug 17, 2026
BenchmarkOJBenchScore32.5%Rank06Participants9Percentile37.5%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ReduxScore93.8%Rank07Participants48Percentile87.2%EvidenceCEvaluatedAug 17, 2026
BenchmarkIncludeScore81.0%Rank08Participants31Percentile76.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ProXScore81.0%Rank08Participants32Percentile77.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkPolyMATHScore60.1%Rank08Participants23Percentile68.2%EvidenceCEvaluatedAug 17, 2026
BenchmarkHMMT25Score83.9%Rank11Participants25Percentile58.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkSuperGPQAScore64.9%Rank11Participants34Percentile69.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkTau2 AirlineScore58.0%Rank15Participants23Percentile36.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkTau2 RetailScore71.9%Rank16Participants26Percentile40.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkTAU-bench AirlineScore46.0%Rank17Participants23Percentile27.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkTAU-bench RetailScore67.8%Rank17Participants25Percentile33.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkLiveCodeBench v6Score74.1%Rank27Participants56Percentile52.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkIFEvalScore87.8%Rank28Participants67Percentile59.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ProScore84.4%Rank28Participants134Percentile79.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkTau2 TelecomScore45.6%Rank31Participants35Percentile11.8%EvidenceCEvaluatedAug 17, 2026
BenchmarkAIME 2025Score92.3%Rank37Participants115Percentile68.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkHumanity's Last ExamScore18.2%Rank66Participants99Percentile33.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkGPQAScore81.1%Rank84Participants239Percentile65.1%EvidenceCEvaluatedAug 17, 2026

Qwen3 235B A22B Thinking Arena Results

Preference and agent-evaluation results for the default version.

2 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
ArenatextCategoryoverallRank114Rating / score1413.9Votes9,017ObservationsN/AResult dateAug 12, 2026
Arenatext style controlCategoryoverallRank140Rating / score1399.0Votes9,017ObservationsN/AResult dateAug 12, 2026

Qwen3 235B A22B Thinking Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
From $0.20 input, $0.60 output per 1M via submodel
Tracked offerings
11
10 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProvideriFlowProvider model IDqwen3-235b-a22b-thinking-2507RegionglobalInput / 1MN/AOutput / 1MN/AContext256KUpdatedAug 17, 2026
ProviderModelScopeProvider model IDQwen/Qwen3-235B-A22B-Thinking-2507RegionglobalInput / 1MN/AOutput / 1MN/AContext262.1KUpdatedAug 17, 2026
ProvidersubmodelProvider model IDQwen/Qwen3-235B-A22B-Thinking-2507RegionglobalInput / 1M$0.20Output / 1M$0.60Context262.1KUpdatedAug 17, 2026
ProviderOpenRouterProvider model IDqwen/qwen3-235b-a22b-thinking-2507RegionglobalInput / 1M$0.23Output / 1M$2.3Context262.1KUpdatedAug 17, 2026
ProviderHugging FaceProvider model IDQwen/Qwen3-235B-A22B-Thinking-2507RegionglobalInput / 1M$0.30Output / 1M$3Context262.1KUpdatedAug 17, 2026
ProviderJiekou.AIProvider model IDqwen/qwen3-235b-a22b-thinking-2507RegionglobalInput / 1M$0.30Output / 1M$3Context131.1KUpdatedAug 17, 2026
ProviderNovitaAIProvider model IDqwen/qwen3-235b-a22b-thinking-2507RegionglobalInput / 1M$0.30Output / 1M$3Context131.1KUpdatedAug 17, 2026
ProviderLLM GatewayProvider model IDqwen3-235b-a22b-thinking-2507RegionglobalInput / 1M$0.30Output / 1M$3Context262KUpdatedAug 17, 2026
ProviderVercel AI GatewayProvider model IDalibaba/qwen3-235b-a22b-thinkingRegionglobalInput / 1M$0.40Output / 1M$4Context131.1KUpdatedAug 17, 2026
ProviderVenice AIProvider model IDqwen3-235b-a22b-thinking-2507RegionglobalInput / 1M$0.45Output / 1M$3.5Context128KUpdatedAug 17, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Qwen3 235B A22B Thinking Runtime Performance

Provider-specific output speed and catalog latency for Qwen3-235B-A22B-Thinking-2507. Runtime does not affect the capability score.

No runtime data

No provider-specific speed or latency record is linked to the default version yet.

Browse runtime rankings

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Qwen3 235B A22B Thinking Specifications

Technical details for the model's default version.

Version
Qwen3-235B-A22B-Thinking-2507
Released
Jul 25, 2025
Knowledge cutoff
Unknown
Parameters
235B
Context window
262.1K
Max output
131.1K
Inputs
text
Outputs
text
Open weights
No
License
Apache 2.0

Qwen3 235B A22B Thinking Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 rows
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionQwen3-235B-A22B-Thinking-2507ReleasedJul 25, 2025LLMBoard48.0Parameters235BContext262.1KMax output131.1KOpen weightsNoLicenseApache 2.0

Qwen3 235B A22B Thinking vs nearby models

Open a comparison with the three ranked models immediately above and below this model.

Qwen3 235B A22B ThinkingvsGLM 4.5Qwen3 235B A22B ThinkingvsGPT-5.4-nanoQwen3 235B A22B ThinkingvsMistral Medium 3.5Qwen3 235B A22B ThinkingvsGrok 3 MiniQwen3 235B A22B ThinkingvsMAI Code 1.1 FlashQwen3 235B A22B ThinkingvsGrok 4.1

Models similar to Qwen3 235B A22B Thinking

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#113-2.7
AC

Qwen3 VL 235B A22B Thinking

Alibaba Cloud / Qwen Team

45.3 LLMBoard

DetailsCompare
#95+3.6
AC

Qwen3.5 35B A3B

Alibaba Cloud / Qwen Team

51.6 LLMBoard

DetailsCompare
#126-5.3
AC

Qwen3 VL 235B A22B

Alibaba Cloud / Qwen Team

42.7 LLMBoard

DetailsCompare
#128-6.1
AC

Qwen3 235B A22B

Alibaba Cloud / Qwen Team

41.9 LLMBoard

DetailsCompare
#82+6.8
AC

Qwen3.6 35B A3B

Alibaba Cloud / Qwen Team

54.8 LLMBoard

DetailsCompare
#132-7.1
AC

Qwen3.5 9B

Alibaba Cloud / Qwen Team

40.9 LLMBoard

DetailsCompare

What is Qwen3 235B A22B Thinking?

Key information about Qwen3 235B A22B Thinking and its available data.

Qwen3-235B-A22B-Thinking-2507 is a state-of-the-art thinking-enabled Mixture-of-Experts (MoE) model with 235B total parameters (22B activated). It features 94 layers, 128 experts (8 activated), and supports 262K native context length.

This version delivers significantly improved reasoning performance, achieving state-of-the-art results among open-source thinking models on logical reasoning, mathematics, science, coding, and academic benchmarks.

Key enhancements include markedly better general capabilities (instruction following, tool usage, text generation), enhanced 256K long-context understanding, and increased thinking depth. The model supports only thinking mode with automatic <think> tag inclusion.

Data as of 2026-08-17.

FAQ

Common questions about Qwen3 235B A22B Thinking.

When was Qwen3 235B A22B Thinking released?

Qwen3 235B A22B Thinking's default version was released on Jul 25, 2025.

How much does Qwen3 235B A22B Thinking cost?

No official standard PAYG price is currently available for Qwen3 235B A22B Thinking. The lowest tracked third-party offer starts at $0.20 input and $0.60 output via submodel.

Who created Qwen3 235B A22B Thinking?

Qwen3 235B A22B Thinking was created by Alibaba Cloud / Qwen Team.

What is the context window for Qwen3 235B A22B Thinking?

The default version has a 262.1K token context window.

Is Qwen3 235B A22B Thinking open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Qwen3 235B A22B Thinking?

11 provider offerings are linked to the default version.

What models should I compare Qwen3 235B A22B Thinking with?

Nearby ranked alternatives include GLM 4.5, GPT-5.4-nano, Mistral Medium 3.5.