Step3 is a multimodal reasoning model released by StepFun. It uses a Mixture‑of‑Experts (MoE) architecture with 321 billion total parameters and 38 billion activation parameters. The model follows an end‑to‑end design that reduces decoding cost while delivering top‑tier performance on vision‑language reasoning tasks. Thanks to the combined use of Multi‑Head Factorized Attention (MFA) and Attention‑FFN Decoupling (AFD), Step3 remains highly efficient on both flagship and low‑end accelerators. During pre‑training, it processed over 20 trillion text tokens and 4 trillion image‑text mixed tokens, covering more than ten languages. On benchmarks for mathematics, code, and multimodal tasks, Step3 consistently outperforms other open‑source models.
← Models
stepfun-ai/step3
stepfun-ai/step3 Compare
Compare pricing, specifications, performance, and benchmarks for up to four models.
+ Add model · 1/4
stepfun-ai/step3
StepFun · → text
Input$1.10 /M
Output$2.75 /M
Pick a second model to start comparing.
Popular comparisons
Related model match-ups readers also look at.
