Google model product
Gemma 3 12B is a 12-billion-parameter vision-language model from Google, handling text and image input and generating text output.
Updated Aug 12, 2026. Default version: Gemma 3 12B
Technical details for the model's default version.
This profile uses the model's current scored version. Arena ratings and prices are shown separately.
Benchmark scores for Gemma 3 12B.
| VQAv2 (val) | 71.6% | 01 | 3 | 100.0% | C | |
| ECLeKTic | 10.3% | 03 | 8 | 71.4% | C | |
| Bird-SQL (dev) | 47.9% | 04 | 7 | 50.0% | C | |
| HiddenMath | 54.5% | 04 | 13 | 75.0% | C | |
| Natural2Code | 80.7% | 04 | 8 | 57.1% | C | |
| BIG-Bench Hard | 85.7% | 06 | 21 | 75.0% | C | |
| FACTS Grounding | 75.8% | 06 | 13 | 58.3% | C | |
| Global-MMLU-Lite | 69.5% | 07 | 14 | 53.9% | C | |
| BIG-Bench Extra Hard | 16.3% | 08 | 11 | 30.0% | C | |
| InfoVQA | 64.9% | 08 | 9 | 12.5% | C | |
| MMMU (val) | 59.6% | 10 | 11 | 10.0% | C | |
| TextVQA | 67.7% | 13 | 15 | 14.3% | C | |
| MATH | 83.8% | 14 | 71 | 81.4% | C | |
| WMT24++ | 51.6% | 15 | 23 | 36.4% | C | |
| GSM8k | 94.4% | 16 | 48 | 68.1% | C | |
| MBPP | 73.0% | 21 | 33 | 37.5% | C | |
| DocVQA | 87.1% | 22 | 26 | 16.0% | C | |
| IFEval | 88.9% | 22 | 65 | 67.2% | C | |
| MathVista-Mini | 62.9% | 22 | 23 | 4.5% | C | |
| ChartQA | 75.7% | 23 | 24 | 4.3% | C | |
| AI2D | 84.2% | 24 | 32 | 25.8% | C | |
| HumanEval | 85.4% | 34 | 66 | 49.2% | C | |
| SimpleQA | 6.3% | 42 | 46 | 8.9% | C | |
| LiveCodeBench | 24.6% | 65 | 73 | 11.1% | C | |
| MMLU-Pro | 60.6% | 103 | 129 | 20.3% | C | |
| GPQA | 40.9% | 203 | 234 | 13.3% | C |
Preference and agent-evaluation results for the default version.
| text | overall | 194 | 1334.2 | 3,829 | N/A | |
| text style control | overall | 206 | 1341.9 | 3,829 | N/A |
Provider-specific output speed and catalog latency for Gemma 3 12B. Runtime does not affect the capability score.
| DeepInfra | 33 tok/s | 0.2 s | 131.1K | 131.1K |
Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.
Official vendor API pricing appears first, followed by individual provider offers.
| Helicone | gemma-3-12b-it | global | $0.05 | $0.10 | 131.1K | |
| NovitaAI | google/gemma-3-12b-it | global | $0.05 | $0.10 | 131.1K | |
| OpenRouter | google/gemma-3-12b-it | global | $0.05 | $0.15 | 131.1K | |
| Neon | gemma-3-12b | global | $0.15 | $0.50 | 131.1K |
Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.
Available versions of this model. The score column identifies the version used in the overall ranking.
| Gemma 3 12B | 11.2 | 12B | 131.1K | 131.1K | No | Gemma |
Key information about Gemma 3 12B and its available data.
Gemma 3 12B is a 12-billion-parameter vision-language model from Google, handling text and image input and generating text output. It features a 128K context window, multilingual support, and open weights. Suitable for question answering, summarization, reasoning, and image understanding tasks.
Data as of 2026-08-11.
Open a comparison with the three ranked models immediately above and below this model.
Recommendations prioritize the same model type and family, then the closest LLMBoard score.
Common questions about Gemma 3 12B.
Gemma 3 12B's default version was released on Mar 12, 2025.
No official standard PAYG price is currently available for Gemma 3 12B. The lowest tracked third-party offer starts at $0.05 input and $0.10 output via Helicone.
Gemma 3 12B was created by Google.
The default version has a 131.1K token context window.
No. The default version is not marked as having publicly available weights.
4 provider offerings are linked to the default version.
Nearby ranked alternatives include Qwen2.5 VL 7B, DeepSeek-R1-Distill-Llama, Gemma 4 E2B.