Citation and evidence

Step-3

14 min full readUpdated 21 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

Chinese AILarge Language ModelsMixture of ExpertsMultimodal AI

Cite this article

Compare and use this model

Step-3 has an article. Its structured comparison facts are awaiting review. Browse the reviewed catalog or suggest a sourced update.

Step-3 is an open-weight large multimodal mixture of experts (MoE) model from StepFun, the Shanghai-based Chinese artificial intelligence startup also known as Jieyue Xingchen. It is a vision-language model with about 321 billion total parameters and 38 billion parameters activated per token, and it is distinguished less by raw benchmark leadership than by its central design goal: minimizing the cost of inference decoding so that a frontier-scale model can be served cheaply at high throughput.[1][2] StepFun pursued this goal through a model-system co-design that pairs two techniques, Multi-Matrix Factorization Attention (MFA) and Attention-FFN Disaggregation (AFD), to reduce attention cost and raise GPU utilization. StepFun unveiled the model in Shanghai on 25 July 2025, on the eve of the 2025 World Artificial Intelligence Conference (WAIC).[7] The accompanying research paper, titled "Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding," was submitted to arXiv on 25 July 2025, and the model weights were open-sourced on 31 July 2025 under the Apache License 2.0.[1][3][4][8]

Overview

StepFun described Step-3 as its new-generation main foundation model and as its first full-size, natively multimodal reasoning model.[7] The model accepts both image and text inputs and is aimed at multimodal AI reasoning tasks, including mathematics, science, and code, alongside general visual understanding.[5][7] In its technical report, StepFun compared Step-3's decoding costs with those of other recent models, including DeepSeek-V3, Kimi K2, Qwen3-235B-A22B, Qwen3-32B, Llama 4 Maverick, MiniMax M1, ERNIE 4.5, and Pangu Pro MoE.[1]

The defining thesis of the project, captured in the paper's title, is that a large model need not be expensive to run. Rather than shrinking the model to cut serving costs, StepFun argued that decoding cost is governed by the interaction of three factors, attention arithmetic intensity, MoE sparsity, and the way attention and feed-forward computation are placed on hardware, and that a model co-designed around the economics of real accelerators can activate more parameters per token than rivals while still costing less to serve.[1] Step-3 activates 38 billion parameters per token, more than DeepSeek-V3 or Qwen3 MoE 235B activate, yet StepFun reports lower theoretical decoding cost on the hardware it studied.[1]

StepFun

StepFun (Shanghai Jieyue Xingchen Intelligent Technology Co., Ltd.) was registered on 6 April 2023, with Jiang Daxin, a former Microsoft vice president who had led work on Bing search and natural language processing, as executive director.[11] Jiang is the company's founder and chief executive.[7][10] The company has been described as one of China's "AI Tiger" startups, a group of well-funded large-model developers, and its investors have included Tencent, Qiming Venture Partners, and Shanghai state-owned capital.[6][10]

StepFun has emphasized multimodal foundation models across text, image, audio, and video. At the World Artificial Intelligence Conference in July 2024 it launched the official version of Step-2, a trillion-parameter MoE language model, together with the Step-1.5V multimodal model and the Step-1X image-generation model.[12] In February 2025 it and Geely Auto Group open-sourced the Step-Video-T2V text-to-video model and the Step-Audio speech model.[13] Step-3 followed in July 2025 as the company's next-generation foundation model.[7]

Architecture

Step-3 is built on a sparse mixture of experts transformer design. According to the technical report, the vision-language model totals about 321 billion parameters; the language-model component comprises 316 billion parameters with 38 billion activated for each token, and there is an additional vision encoder of about 5 billion parameters that handles image inputs.[1] The released model card lists 61 layers (5 of them dense), a hidden dimension of 7,168, a maximum context length of 65,536 tokens, and a reuse of the DeepSeek-V3 tokenizer.[2] The technical report specifies that MoE layers replace every feed-forward layer except the first four and the last.[1]

The MoE feed-forward layers use 48 routed experts with 3 experts selected per token plus 1 shared expert, a design the report says was inspired by DeepSeekMoE.[1][2] StepFun chose this sparsity deliberately: the report argues that overly sparse MoE models run inefficiently on high-compute accelerators such as the H800, and says Step-3 uses a sparsity of around 0.08 (including the shared expert).[1] The model card distributes weights in both bf16 and block-FP8 formats and recommends serving through inference engines such as vLLM and SGLang.[2]

StepFun's model blog says that Step-3's 5-billion-parameter vision encoder is built on EVA-CLIP 5B, and that its image features are compressed by two successive 2D convolutional layers that reduce the visual token grid to one-sixteenth of its original size before the tokens are merged with text tokens.[5] During pretraining the model processed more than 20 trillion text tokens across more than ten languages, of which 3.7 trillion high-quality tokens were reserved for the annealing phases, plus 4 trillion image-text mixed tokens, per StepFun.[5] The same blog says Step-3 can run on eight 48 GB GPUs while processing up to 800,000 tokens in total, a figure it defines as batch size times length and calculates assuming int8 quantization of non-attention parameters; this is a serving-capacity figure, not the model's per-request context limit, which the model card gives as 65,536 tokens.[2][5]

Multi-Matrix Factorization Attention (MFA)

Multi-Matrix Factorization Attention is the attention mechanism at the core of Step-3's efficiency design. StepFun released MFA as a new attention architecture in late 2024, in a paper first posted in December 2024.[1][9] MFA applies low-rank matrix factorization to the query-key circuit, which lets StepFun scale both the number and the dimensionality of attention heads in a parameter-efficient way while keeping the KV cache small.[1][9] In Step-3, 64 query heads share a single key head and a single value head, all with a dimension of 256, and the query is down-projected from 7,168 to a low-rank dimension of 2,048 before being up-projected again.[1][2] StepFun states that this design reduces both KV-cache size and attention compute while preserving attention expressiveness, and reports that Step-3 uses 22 percent of DeepSeek-V3's per-token attention cost.[5]

The report frames MFA in terms of arithmetic intensity, the number of operations performed per byte of KV cache read. It gives Step-3's MFA an arithmetic intensity of 128 (with 8-bit KV), compared with 512 for DeepSeek-V3's multi-head latent attention and 32 for Qwen3 MoE's grouped-query attention, and argues that 128 sits closer to the compute-to-bandwidth ratios of cheaper accelerators such as the A800 and Ascend 910B.[1]

Attention-FFN Disaggregation (AFD)

Attention-FFN Disaggregation is a distributed-inference system, rather than a change to the model weights, that decouples the attention layers and the feed-forward (FFN) layers into separate, specialized subsystems running on different sets of GPUs.[1] Because attention and the MoE feed-forward layers have very different compute and memory profiles, executing them together forces compromises in batching and hardware utilization. By disaggregating the two, AFD lets each subsystem be sized and scheduled independently; the report lists among AFD's advantages that it keeps an ideal batch size for the FFN part, scales attention instances with context length, and allows heterogeneous hardware to be used to reduce decoding costs further.[1] StepFun's serving system runs attention, FFN, and communication as a three-stage pipeline with a target of 50 milliseconds per output token, and it uses StepMesh, a communication library for AFD built on GPUDirect RDMA that StepFun open-sourced.[1][15] The paper presents AFD as the system half of a co-design in which MFA and the MoE sparsity pattern are the model half.

Decoding-cost efficiency

The central contribution of Step-3 is its analysis of decoding cost, the cost of generating output tokens, which dominates serving expense for reasoning and long-output workloads. StepFun reports a theoretical decoding-cost analysis across four accelerators, NVIDIA H800, H20, and A800 and Huawei Ascend 910B, expressed in US dollars per million decoded tokens.[1] These per-token cost figures are theoretical estimates derived from the model and system design, not list prices; the comparisons should be read as such.

In that analysis, StepFun reports that Step-3 has lower theoretical decoding cost than both DeepSeek-V3 and Qwen3 MoE 235B, with the advantage widening at longer context. At an 8K context (using AFD on H800 and H20), the paper cites about 0.055 USD per million decoded tokens for Step-3 versus 0.068 for DeepSeek-V3 (with expert parallelism on H800) and 0.062 for Qwen3 MoE 235B; at 32K context the figures are 0.129 for Step-3 versus 0.211 for DeepSeek-V3 and 0.193 for Qwen3 MoE 235B.[1] By these figures, Step-3's cost is about 19 to 39 percent lower than DeepSeek-V3's and about 11 to 33 percent lower than Qwen3 MoE 235B's over those context lengths. StepFun emphasizes that Step-3 attains this lower cost despite activating more parameters per token than either comparison model, which it presents as evidence that hardware-aligned attention arithmetic intensity, MoE sparsity, and AFD jointly drive cost-effectiveness.[1]

Beyond the theoretical analysis, StepFun reports a measured result. On Hopper GPUs, with a 4,096-token average context, FP8 matrix multiplication, no multi-token prediction, and a 50-millisecond time-per-output-token target (at least 20 tokens per second), Step-3 reached a long-term average of 3,910 tokens per GPU per second and 4,039 in a peak minute with FP8 attention, around 74 percent higher than the 2,324 tokens per GPU per second that DeepSeek published from profiling DeepSeek-V3 at 4,096 context on H800.[1] The 4K deployment used two attention instances and two FFN instances, 32 GPUs in total.[1] StepFun's launch materials rounded the gap to about 70 percent.[8] All of these figures are StepFun's own.

Benchmarks

StepFun's launch materials said Step-3 set state-of-the-art results among open-source multimodal reasoning models on MMMU, MathVision, SimpleVQA, AIME 2025, and LiveCodeBench.[7] In StepFun's published comparison table, Step-3 has the highest score among the open-source vision-language models listed (Llama 4 Maverick, ERNIE 4.5-thinking, QvQ-72B-Preview, GLM-4.1V-thinking, and MiMo-VL) on every benchmark except GPQA-Diamond, where ERNIE 4.5-thinking is shown at 76.8.[5] In the same table, the open-source reasoning model DeepSeek R1-0528 scores higher than Step-3 on all five text benchmarks, and proprietary models such as OpenAI o3 and Gemini 2.5 Pro score higher on MMMU, MATH-Vision, AIME 2025, GPQA-Diamond, and LiveCodeBench.[5] Many competitor figures in the table are marked as reproduced by StepFun under the same settings. The following Step-3 scores are drawn from that table.

BenchmarkStep-3 score (StepFun-reported)
MMMU (multimodal understanding)74.2
MATH-Vision64.8
SimpleVQA62.2
HallusionBench64.2
ZeroBench (subquestions)23.0
DynaMath50.1
AIME 2025 (math)82.9
GPQA-Diamond (science)73.0
LiveCodeBench (Aug 2024 to May 2025)67.1
HMMT 2025 (math)70.0
CNMO 2024 (math)83.7

Expanded article table

Source: StepFun model blog and model card.[2][5] As with all vendor-reported benchmarks, these results reflect the developer's own testing conditions and should be treated with appropriate caution.

Specifications

AttributeDetail
DeveloperStepFun (Jieyue Xingchen), Shanghai, China
Model typeMultimodal (vision-language) mixture of experts
Total parametersAbout 321 billion (vision-language model); language model 316 billion
Active parametersAbout 38 billion per token
Vision encoderAbout 5 billion parameters (built on EVA-CLIP 5B)
Experts48 routed (3 active per token) plus 1 shared
Layers61 (5 dense)
Hidden size7,168
AttentionMFA: 64 query heads, head dimension 256, low-rank query dimension 2,048
Context length65,536 tokens
TokenizerDeepSeek-V3 tokenizer
Precision formatsbf16, block-FP8
Key techniquesMulti-Matrix Factorization Attention (MFA); Attention-FFN Disaggregation (AFD)
Pretraining dataMore than 20T text tokens, 4T image-text tokens (StepFun-reported)
Announced25 July 2025
Open-weight release31 July 2025
LicenseApache License 2.0

Expanded article table

Availability and significance

Step-3 was released as an open-weight model on 31 July 2025 under the permissive Apache License 2.0, which covers both the code repository and the model weights.[2][8] The weights are distributed on Hugging Face (as stepfun-ai/step3, with a block-FP8 version as stepfun-ai/step3-fp8) and on ModelScope, while the GitHub repository stepfun-ai/Step3 hosts the code, deployment guide, and system technical report.[2][16] The release was accompanied by support for the vLLM and SGLang serving frameworks, and StepFun offers the model through an OpenAI-compatible API on its developer platform.[2]

At the launch, StepFun also announced a "Model-Chip Ecosystem Innovation Alliance" with close to ten Chinese chip and infrastructure companies. Its first members included Huawei Ascend, MetaX, Biren, Enflame, Iluvatar CoreX, Infinigence, Cambricon, Moore Threads, and SiliconFlow. StepFun said Huawei Ascend chips were the first to run Step-3, with MetaX, Iluvatar CoreX, and Enflame also running it in initial form, and claimed that, by theoretical analysis, Step-3's inference efficiency on Chinese chips could reach up to 300 percent of DeepSeek-R1's.[7] TechNode described the alliance as an effort to optimize Step-3 for domestic chips.[14]

StepFun's model blog listed several known issues. During scaling it observed what it called a "dead expert" phenomenon, in which some experts become effectively inactive because their output weight norms vanish during training, even though tokens are still routed to them. It also said Step-3 was underoptimized for "vibe coding," and that prolonged multimodal reasoning training showed a trade-off in which visual-perception accuracy fell as textual reasoning improved.[5]

The significance of Step-3 lies primarily in its argument that frontier-scale capability and low serving cost are not in tension. By foregrounding decoding economics as a first-class design objective and co-designing the model (MFA, MoE sparsity) with the serving system (AFD) around the arithmetic of real accelerators, StepFun offered a counterpoint to the assumption that cheaper inference requires smaller models.[1] The work also fits the broader 2025 trend of Chinese laboratories releasing capable open-weight MoE models, alongside DeepSeek-V3, Qwen, Kimi K2, and MiniMax, and it extended StepFun's earlier Step-1, Step-2, and Step-1X line into a more efficiency-focused flagship. Its quantitative cost and benchmark claims originate with StepFun.

Successors

Step 3.5 Flash, the next model in the line, used a different architecture from Step-3 (see table). Its technical report calls Step-3 the company's "prior work" and revisits the dead-expert problem Step-3 reported, and the Step 3.5 Flash model card estimates decoding cost with a method it describes as similar to, but more accurate than, the Step-3 paper's.[17][21]

ModelReleaseSizeNotes
Step 3.5 FlashFebruary 2026196B total, about 11B active (MoE)45 layers, 288 routed experts (top-8) plus 1 shared; interleaved 3:1 sliding-window and full attention; 256K context; 3-way multi-token prediction. Apache 2.0.[21][17]
Step 3.7 Flash29 May 2026198B total (196B language backbone plus 1.8B vision encoder), about 11B activeVision-language model; 256K context; low, medium, and high reasoning levels. Apache 2.0.[18][19]
Step 5 Preview20 September 2026600B total, 27B active (MoE)StepFun's new flagship for agentic work; 1M context and vision input; open weights promised for 15 October 2026.[20]

Expanded article table

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22Wang, B., et al. (StepFun). "Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding." arXiv:2507.19427, 25 July 2025. arxiv.org/...2507.19427
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10"stepfun-ai/step3." Hugging Face model card. huggingface.co/...step3
  3. ^StepFun. "Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding" (v1 PDF), arXiv, 25 July 2025. arxiv.org/...2507.19427v1
  4. ^StepFun (@StepFun_ai). "Step 3 will be open-sourced on July 31st! Coming Soon." X, 26 July 2025. x.com/...1948954102127624531
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9"Step3: Cost-Effective Multimodal Intelligence." StepFun official research page, 31 July 2025. stepfun.ai/...step3 (archived copy: web.archive.org/...step3)
  6. ^"StepFun." Wikipedia. en.wikipedia.org/...StepFun
  7. ^1 ^2 ^3 ^4 ^5 ^6 ^7Leiphone (雷峰网). "WAIC 2025|阶跃发布新一代基模 Step 3:原生多模态,推理效率行业领先" (WAIC 2025: StepFun releases new-generation foundation model Step 3, natively multimodal with industry-leading inference efficiency), 25 July 2025. m.leiphone.com/...5x7SokHjBus8LCUY
  8. ^1 ^2 ^3StepFun (@StepFun_ai). "Announcing Step 3: Our latest open-source multimodal reasoning model is here!" X, 31 July 2025. x.com/...1950912271565385770
  9. ^1 ^2Hu, J., Li, H., Zhang, Y., Wang, Z., Zhou, S., Zhang, X., Shum, H.-Y., Jiang, D. "Multi-matrix Factorization Attention." arXiv:2412.19255, 26 December 2024. arxiv.org/...2412.19255
  10. ^1 ^2The Paper (澎湃新闻). "阶跃星辰完成数亿美元B轮融资,核心投资方包括上海国投和旗下基金" (StepFun completes Series B of several hundred million dollars; core investors include Shanghai State-owned Capital Investment and its funds), 23 December 2024. m.thepaper.cn/newsDetail_forward_29726642
  11. ^Leiphone (雷峰网). "独家丨前微软 NLP 大牛姜大昕创立新公司「阶跃星辰」" (Exclusive: former Microsoft NLP expert Jiang Daxin founds new company StepFun), 12 December 2023. leiphone.com/...mJcRRqbgIh6UkVuk
  12. ^CNR (央广网). "阶跃星辰亮相 WAIC 2024:首发"万亿"和"多模"大模型,全面升级产品生态" (StepFun at WAIC 2024: debuts trillion-parameter and multimodal models), 4 July 2024. tech.cnr.cn/...t20240704_526778089.shtml
  13. ^China Daily. "StepFun and Geely Auto open-source large models to global developers," 19 February 2025. chinadaily.com.cn/...WS67b58a09a310c240449d61da
  14. ^TechNode. "Chinese AI chipmakers join forces with StepFun to counter Nvidia's return to China," 30 July 2025. technode.com/...to-counter-nvidias-return-to-china
  15. ^"StepMesh: A High-Performance, Low-Latency Communication Library for Attention-FFN Disaggregation." GitHub, stepfun-ai/StepMesh. github.com/...StepMesh
  16. ^"stepfun-ai/Step3." GitHub repository. github.com/...Step3
  17. ^1 ^2StepFun. "Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters." arXiv:2602.10604, 11 February 2026. arxiv.org/...2602.10604
  18. ^StepFun. "Step 3.7 Flash: A high-efficiency Flash model for real-world agents," 29 May 2026. static.stepfun.com/...step-3.7-flash
  19. ^"stepfun-ai/Step-3.7-Flash." Hugging Face model card. huggingface.co/...Step-3.7-Flash
  20. ^StepFun (@StepFun_ai). "Introducing Step 5 Preview: Advancing the Pareto Frontier." X, 20 September 2026. x.com/...2101510462685003786
  21. ^1 ^2"stepfun-ai/Step-3.5-Flash." Hugging Face model card. huggingface.co/...Step-3.5-Flash

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

3 revisions · v4 · 2,831 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: xg07 independent adversarial verification 2026-09-23 (V9); writer audit (benchmark table) + 3 minor fixed

Cite this page: AI Wiki. "Step-3." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/step_3

Suggest edit