llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Alibaba Cloud / Qwen Team model product

Qwen3 VL 235B A22B Thinking

Qwen3-VL-235B-A22B-Thinking is the most powerful vision-language model in the Qwen series, featuring 236B parameters with MoE architecture for reasoning-enhanced multimodal understanding.

Updated Aug 17, 2026. Default version: Qwen3 VL 235B A22B Thinking

Compare
LLMBoard Score45.3Qwen3 VL 235B A22B Thinking
Coverage100%58 benchmark families
Context window262.1KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Compare
  • Similar models
  • About
  • FAQ

Qwen3 VL 235B A22B Thinking Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Qwen3 VL 235B A22B Thinking LLMBoard score breakdown

Qwen3 VL 235B A22B Thinking Benchmark Results

Benchmark scores for Qwen3 VL 235B A22B Thinking.

30 of 67 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkARKitScenesScore0.537 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkLiveBench 20241125Score79.6%Rank01Participants14Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMathVerse-MiniScore0.85 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMIABenchScore0.927 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMMUvalScore80.6%Rank01Participants4Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkObjectronScore0.712 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkOCRBench-V2 (zh)Score63.5%Rank01Participants11Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkOSWorld-GScore0.683 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkRoboSpatialHomeScore0.739 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkSIFOScore0.773 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkSIFO-MultiturnScore0.711 pointsRank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkZebraLogicScore97.3%Rank01Participants8Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkCharadesSTAScore63.5%Rank02Participants12Percentile90.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkDesign2CodeScore0.934 pointsRank02Participants2Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkInfoVQAtestScore89.5%Rank02Participants12Percentile90.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkMuirBenchScore80.1%Rank02Participants12Percentile90.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkRefSpatialBenchScore0.699 pointsRank02Participants6Percentile80.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkEmbSpatialBenchScore0.843 pointsRank03Participants8Percentile71.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkRefCOCO-avgScore0.924 pointsRank03Participants9Percentile75.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkSUNRGBDScore0.349 pointsRank03Participants4Percentile33.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkVisuLogicScore0.344 pointsRank03Participants3Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkWritingBenchScore86.7%Rank03Participants15Percentile85.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkDocVQAtestScore96.5%Rank04Participants11Percentile70.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkHypersimScore0.11 pointsRank04Participants4Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMulti-IFScore79.1%Rank04Participants23Percentile86.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkOCRBench-V2 (en)Score66.8%Rank04Participants14Percentile76.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkScreenSpotScore95.4%Rank04Participants16Percentile80.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkCC-OCRScore81.5%Rank05Participants18Percentile76.5%EvidenceCEvaluatedAug 17, 2026
BenchmarkCreative Writing v3Score85.7%Rank05Participants13Percentile66.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLongBench-DocScore0.562 pointsRank05Participants5Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLUScore90.6%Rank05Participants101Percentile96.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkZEROBench-SubScore0.277 pointsRank05Participants5Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkCountBenchScore0.937 pointsRank06Participants7Percentile16.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkHallusion BenchScore66.7%Rank06Participants18Percentile70.6%EvidenceCEvaluatedAug 17, 2026
BenchmarkMM-MT-BenchScore8.5 pointsRank06Participants17Percentile68.8%EvidenceCEvaluatedAug 17, 2026
BenchmarkVideoMME w/o sub.Score79.0%Rank06Participants10Percentile44.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkBFCL-v3Score71.9%Rank07Participants19Percentile66.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMBench-V1.1Score90.6%Rank07Participants20Percentile68.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMStarScore78.7%Rank07Participants24Percentile73.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkMathVista-MiniScore85.8%Rank08Participants24Percentile69.6%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ReduxScore93.7%Rank08Participants48Percentile85.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkBLINKScore67.1%Rank09Participants15Percentile42.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkMLVUScore83.8%Rank09Participants10Percentile11.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkSimpleVQAScore0.613 pointsRank09Participants14Percentile38.5%EvidenceCEvaluatedAug 17, 2026
BenchmarkZEROBenchScore0.04 pointsRank09Participants9Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkIncludeScore80.0%Rank10Participants31Percentile70.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ProXScore80.6%Rank10Participants32Percentile71.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkODinWScore43.2%Rank10Participants16Percentile40.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkOSWorldScore38.1%Rank11Participants20Percentile47.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkRealWorldQAScore81.3%Rank12Participants29Percentile60.7%EvidenceCEvaluatedAug 17, 2026
BenchmarkHMMT25Score77.4%Rank13Participants25Percentile50.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkLVBenchScore63.6%Rank13Participants25Percentile50.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkOCRBenchScore87.5%Rank13Participants24Percentile47.8%EvidenceCEvaluatedAug 17, 2026
BenchmarkSuperGPQAScore64.3%Rank13Participants34Percentile63.6%EvidenceCEvaluatedAug 17, 2026
BenchmarkAI2DScore89.2%Rank14Participants33Percentile59.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkERQAScore52.5%Rank14Participants24Percentile43.5%EvidenceCEvaluatedAug 17, 2026
BenchmarkScreenSpot ProScore61.8%Rank14Participants25Percentile45.8%EvidenceCEvaluatedAug 17, 2026
BenchmarkMathVisionScore74.6%Rank15Participants33Percentile56.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkVideoMMMUScore80.0%Rank17Participants26Percentile36.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkSimpleQAScore44.4%Rank18Participants47Percentile63.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkIFEvalScore88.2%Rank27Participants67Percentile60.6%EvidenceCEvaluatedAug 17, 2026
BenchmarkLiveCodeBench v6Score70.1%Rank31Participants56Percentile45.5%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ProScore83.8%Rank31Participants134Percentile77.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkCharXiv-RScore66.1%Rank36Participants51Percentile30.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMMU-ProScore69.3%Rank38Participants68Percentile44.8%EvidenceCEvaluatedAug 17, 2026
BenchmarkAIME 2025Score89.7%Rank49Participants115Percentile57.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkHumanity's Last ExamScore13.6%Rank81Participants99Percentile18.4%EvidenceCEvaluatedAug 17, 2026

Qwen3 VL 235B A22B Thinking Arena Results

Preference and agent-evaluation results for the default version.

4 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
ArenavisionCategoryoverallRank70Rating / score1208.2Votes2,347ObservationsN/AResult dateAug 6, 2026
Arenavision style controlCategoryoverallRank77Rating / score1189.8Votes2,347ObservationsN/AResult dateAug 6, 2026
ArenatextCategoryoverallRank131Rating / score1400.7Votes7,987ObservationsN/AResult dateAug 12, 2026
Arenatext style controlCategoryoverallRank145Rating / score1395.2Votes7,987ObservationsN/AResult dateAug 12, 2026

Qwen3 VL 235B A22B Thinking Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
From $0.40 input, $4 output per 1M via OpenRouter
Tracked offerings
4
4 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProviderOpenRouterProvider model IDqwen/qwen3-vl-235b-a22b-thinkingRegionglobalInput / 1M$0.40Output / 1M$4Context131.1KUpdatedAug 17, 2026
ProviderNanoGPTProvider model IDqwen3-vl-235b-a22b-thinkingRegionglobalInput / 1M$0.50Output / 1M$6Context32.8KUpdatedAug 17, 2026
ProviderNovitaAIProvider model IDqwen/qwen3-vl-235b-a22b-thinkingRegionglobalInput / 1M$0.98Output / 1M$3.95Context131.1KUpdatedAug 17, 2026
ProviderLLM GatewayProvider model IDqwen3-vl-235b-a22b-thinkingRegionglobalInput / 1M$0.98Output / 1M$3.95Context131.1KUpdatedAug 17, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Qwen3 VL 235B A22B Thinking Runtime Performance

Provider-specific output speed and catalog latency for Qwen3 VL 235B A22B Thinking. Runtime does not affect the capability score.

No runtime data

No provider-specific speed or latency record is linked to the default version yet.

Browse runtime rankings

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Qwen3 VL 235B A22B Thinking Specifications

Technical details for the model's default version.

Version
Qwen3 VL 235B A22B Thinking
Released
Sep 22, 2025
Knowledge cutoff
Unknown
Parameters
236B
Context window
262.1K
Max output
262.1K
Inputs
image, text, video
Outputs
text
Open weights
No
License
Apache 2.0

Qwen3 VL 235B A22B Thinking Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 rows
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionQwen3 VL 235B A22B ThinkingReleasedSep 22, 2025LLMBoard45.3Parameters236BContext262.1KMax output262.1KOpen weightsNoLicenseApache 2.0

Qwen3 VL 235B A22B Thinking vs nearby models

Open a comparison with the three ranked models immediately above and below this model.

Qwen3 VL 235B A22B ThinkingvsMiMo V2.5 ProQwen3 VL 235B A22B ThinkingvsGPT-OSS-20BQwen3 VL 235B A22B ThinkingvsClaude Opus 4Qwen3 VL 235B A22B ThinkingvsClaude Haiku 4.5Qwen3 VL 235B A22B ThinkingvsGPT-5-miniQwen3 VL 235B A22B Thinkingvso4 mini

Models similar to Qwen3 VL 235B A22B Thinking

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#126-2.6
AC

Qwen3 VL 235B A22B

Alibaba Cloud / Qwen Team

42.7 LLMBoard

DetailsCompare
#106+2.7
AC

Qwen3 235B A22B Thinking

Alibaba Cloud / Qwen Team

48.0 LLMBoard

DetailsCompare
#128-3.4
AC

Qwen3 235B A22B

Alibaba Cloud / Qwen Team

41.9 LLMBoard

DetailsCompare
#132-4.4
AC

Qwen3.5 9B

Alibaba Cloud / Qwen Team

40.9 LLMBoard

DetailsCompare
#134-4.9
AC

Qwen3 Next 80B A3B Thinking

Alibaba Cloud / Qwen Team

40.4 LLMBoard

DetailsCompare
#135-5.0
AC

Qwen3 Max

Alibaba Cloud / Qwen Team

40.2 LLMBoard

DetailsCompare

What is Qwen3 VL 235B A22B Thinking?

Key information about Qwen3 VL 235B A22B Thinking and its available data.

Qwen3-VL-235B-A22B-Thinking is the most powerful vision-language model in the Qwen series, featuring 236B parameters with MoE architecture for reasoning-enhanced multimodal understanding.

io/HTML/CSS/JS from images/videos), Advanced Spatial Perception (2D grounding and 3D grounding for spatial reasoning and embodied AI), Long Context & Video Understanding (native 256K context expandable to 1M, handles hours-long video with second-level indexing), Enhanced Multimodal Reasoning (excels in STEM/Math with causal analysis), Upgraded Visual Recognition (celebrities, anime, products, landmarks, flora/fauna), and Expanded OCR (32 languages, robust in low light/blur/tilt).

Architecture innovations include Interleaved-MRoPE for positional embeddings, DeepStack for multi-level ViT feature fusion, and Text-Timestamp Alignment for precise video temporal modeling.

Data as of 2026-08-17.

FAQ

Common questions about Qwen3 VL 235B A22B Thinking.

When was Qwen3 VL 235B A22B Thinking released?

Qwen3 VL 235B A22B Thinking's default version was released on Sep 22, 2025.

How much does Qwen3 VL 235B A22B Thinking cost?

No official standard PAYG price is currently available for Qwen3 VL 235B A22B Thinking. The lowest tracked third-party offer starts at $0.40 input and $4 output via OpenRouter.

Who created Qwen3 VL 235B A22B Thinking?

Qwen3 VL 235B A22B Thinking was created by Alibaba Cloud / Qwen Team.

What is the context window for Qwen3 VL 235B A22B Thinking?

The default version has a 262.1K token context window.

Is Qwen3 VL 235B A22B Thinking open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Qwen3 VL 235B A22B Thinking?

4 provider offerings are linked to the default version.

What models should I compare Qwen3 VL 235B A22B Thinking with?

Nearby ranked alternatives include MiMo V2.5 Pro, GPT-OSS-20B, Claude Opus 4.