QwenQwen

Qwen3 32B

Dense 32B model with hybrid thinking

Qwen3 32B is a dense model that supports both standard and thinking (chain-of-thought) modes, allowing users to trade latency for reasoning depth depending on the task at hand.

Try Qwen3 32B

Specifications

Context window128k tokens
Input price$0.060 / 1M tokens
Output price$0.180 / 1M tokens
Credits per query3 cr
Released2025-04
LicenseApache 2.0

Features

Tool useJSON modeFunction callingCode

Benchmark Scores

What do these mean?

Sourced from official model cards and academic papers. Higher is better.

PhD-level science & reasoning

Mathematical olympiad problems

Chatbot Arena human preference ranking

Python coding — pass@1 rate

87%

Competition-level mathematics

90%

General knowledge across 57 academic domains

85%

Multi-turn instruction following (out of 10)

Real-world GitHub issue resolution rate

Run locally

Download and run this model on your own hardware — no API key or internet required after setup.

Min VRAM18 GB
Rec. VRAM20 GB
Best VRAM34 GB
Parameters32B
Speed~9 tok/s
Context32K tokens

Ollama

Terminal

Run this command in your terminal. Ollama downloads and manages the model locally.

ollama pull qwen3:32b