Qwen models on W&B Inference
Qwen3 235B A22B Thinking-2507 inference overview
Price per 1M tokens
$0.10 (input)
$0.10 (output)
Parameters
22B (active)
235B (total)
Context window
262K
Release date
July 2025
Qwen3 235B A22B Thinking-2507 inference details
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass. This “thinking-only” variant enhances structured logical reasoning, mathematics, science, and long-form generation, showing strong benchmark performance across AIME, SuperGPQA, LiveCodeBench, and MMLU-Redux. It enforces a special reasoning mode () and is designed for high-token outputs in challenging domains. The model is instruction-tuned and excels at step-by-step reasoning, tool use, agentic workflows, and multilingual tasks. This release represents the most capable open-source variant in the Qwen3-235B series, surpassing many closed models in structured reasoning use cases as of July 2025.
Created by: Alibaba
License: apache-2.0
🤗 model card: Qwen3-235B-A22B-Thinking-2507
Qwen3 235B A22B-2507 inference overview
Price per 1M tokens
$0.10 (input)
$0.10 (output)
Parameters
22B (active)
235B (total)
Context window
262K
Release date
July 2025
Qwen3 235B A22B-2507 inference details
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following, logical reasoning, math, code, and tool usage. The model does not implement “thinking mode” (
Created by: Alibaba
License: apache-2.0
🤗 model card: Qwen3-235B-A22B-Instruct-2507
Qwen3 Coder 480B A35B inference overview
Price per 1M tokens
$1.00 (input)
$1.50 (output)
Parameters
35B (active)
480B (total)
Context window
262K
Release date
July 2025
Qwen3 Coder 480B A35B inference details
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories. The model features 480 billion total parameters, with 35 billion active per forward pass (8 out of 160 experts).
Created by: Alibaba
License: apache-2.0
🤗 model card: Qwen3-Coder-480B-A35B-Instruct