Citation and evidence

StepFun

24 min full readUpdated 62 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI CompaniesChinese AILarge Language ModelsOpen Source AI

Cite this article

StepFun (Chinese: 阶跃星辰, pinyin: Jiēyuè Xīngchén) is a Shanghai-based Chinese artificial intelligence company that builds the "Step" series of large language models and multimodal foundation models. Its legal name is 上海阶跃星辰智能科技股份有限公司 (Shanghai Jieyue Xingchen Intelligent Technology Co., Ltd.). The company was registered on April 6, 2023, with Jiang Daxin, a former Microsoft vice president and chief scientist of Microsoft's Software Technology Center Asia, as executive director. [15][6] Chinese media and investors count it among the country's "Six Little Tigers" (六小虎) of large-model startups, alongside Zhipu AI, Moonshot AI, MiniMax, Baichuan Intelligence, and 01.AI. [10][22]

StepFun first drew attention with Step-2, a trillion-parameter mixture-of-experts (MoE) language model whose preview in March 2024 was described as the first trillion-parameter model shown by a Chinese startup, and with a multimodal strategy spanning language, vision, video, and audio. [19][50] In January 2026 it closed a Series B+ round of more than RMB 5 billion (about $717 million), which The Paper reported was the largest single financing in China's large-model sector over the previous 12 months, and named Yin Qi, co-founder of Megvii, as chairman. [4][16] In May 2026 it was reported to have raised nearly $2.5 billion more ahead of a planned Hong Kong listing, and on September 20, 2026 it introduced Step 5 Preview, a 600-billion-parameter MoE model. [21][35]

History

When was StepFun founded?

StepFun was registered on April 6, 2023, according to business-registry data cited by Leiphone, with Jiang Daxin as executive director and manager and Zhu Yibo as supervisor. [15] It is headquartered in Shanghai's Xuhui District. [17] A December 2024 profile republished by Sina Finance said the release of ChatGPT had a deep effect on Jiang and prompted him to leave Microsoft, where he had worked for 16 years, to start the company. [19][1] The core founding team included Zhu Yibo, who led systems, and Jiao Binxing, who led data; both had worked at Microsoft. [1][19][26]

The same profile said the team began training models in July 2023, finished Step-1, a language model with more than 100 billion parameters, about two months later, and completed Step-1V, a 100-billion-parameter multimodal model, in November 2023. [19] StepFun raised early funding from HongShan (formerly Sequoia China), Qiming Venture Partners, and IDG Capital in an angel round, followed by a 2024 Series A with 5Y Capital, Shunwei Capital, and Lenovo Capital, according to the 21st Century Business Herald. [20]

2024: Step-2 and rapid model releases

StepFun showed a preview of Step-2 at the Global Developer Pioneer Conference on March 23, 2024. [19] By June 2024 it had launched Step-1V and two consumer apps: Yuewen, a ChatGPT-like assistant, and Maopaoya, a character-chat product with game features. [1] At the World Artificial Intelligence Conference (WAIC) in Shanghai in July 2024, StepFun launched three models: [50]

  • Step-2: the official version of its trillion-parameter language model, built on an MoE architecture.
  • Step-1.5V: a multimodal understanding model.
  • Step-1X: an image generation model.

At WAIC, Jiang said: "Climbing the peak of AGI, 'trillion parameters' and 'multi-modal fusion' are indispensable. The scale of trillion parameters is the basic threshold for achieving AGI; multi-modal large models are the only way to AGI." [50]

In November 2024, MarkTechPost reported that Step-2 ranked fifth on the LiveBench benchmark, behind models from OpenAI and Google, with scores of 86.57 in instruction following, 58.67 in reasoning, and 54.86 in data analysis. [3] Chinese coverage in December 2024 described it as first among Chinese models on that leaderboard. [18] By December 2024 the company said it had released 11 self-developed foundation models in the previous 10 months, covering language, image understanding, image generation, video generation, and speech. [17][18]

2025: Open-source push and automotive partnerships

On February 18, 2025, StepFun and Geely Auto Group jointly announced the open-sourcing of two models: Step-Video-T2V, a 30-billion-parameter text-to-video model that generates videos of up to 204 frames, and Step-Audio, a 130-billion-parameter speech interaction model. [7][38][39] Both were made available in the Yuewen app the same day. [7] Yicai reported that StepFun had worked with Geely on application scenario design, model evaluation, and engineering. [51] StepFun published the Step-Video-T2V weights under the MIT license and Step-Audio under Apache 2.0. [38][39] On February 16, 2025, Yuewen also added DeepSeek-R1. [53]

In April 2025, StepFun released Step1X-Edit, an open-source image editing model that uses a multimodal LLM to interpret editing instructions and a diffusion decoder to produce the edited image. StepFun said it approaches the performance of GPT-4o and Gemini 2 Flash on its own GEdit-Bench benchmark; the weights are Apache 2.0. [40]

On July 25, 2025, on the eve of WAIC in Shanghai, StepFun unveiled Step-3, a multimodal reasoning model with 321 billion total parameters and 38 billion active parameters, and, together with Geely and Qianli Technology, launched a preview of Agent OS, a smart-cockpit operating system. [12][31] Geely presented the Galaxy M9 as the first vehicle with a "human-like AI agent" powered by StepFun's end-to-end voice model. [12] Step-3 introduced two techniques: Multi-Matrix Factorization Attention (MFA), which StepFun says uses 22% of DeepSeek V3's per-token attention cost, and Attention-FFN disaggregation (AFD), which runs attention and feed-forward layers on separate groups of accelerators. [8][9] StepFun published the weights under Apache 2.0 on July 31, 2025. [8][31] At the same conference StepFun launched an alliance with Chinese AI chipmakers including Huawei Ascend, Cambricon, Biren, and Moore Threads to optimize Step-3 for domestic chips. [49]

Later in 2025 the company released the Step-Audio 2 mini speech model (August), NextStep-1 image generation (August), the Step-GUI series of on-device GUI agents (November), Step-Audio-EditX (November), and Step-DeepResearch (December). [41][42][5][44][46]

2026: New chairman, Flash models, and IPO preparation

On January 26, 2026, StepFun announced its Series B+ round of more than RMB 5 billion (about $717 million), and the same day appointed Yin Qi as chairman, responsible for overall strategy and technical direction. [4][16] Yin co-founded Megvii, one of China's early computer-vision unicorns, and since 2024 has been chairman of Qianli Technology, a Geely-backed smart-driving company formerly known as Lifan Technology. [6][14] StepFun said Yin would chair both companies and drive cooperation between them on its "AI plus terminals" strategy. [6] Caixin reported that the round exceeded the proceeds raised by Zhipu AI (HK$4.173 billion, about $535 million) and MiniMax (HK$4.596 billion) in their Hong Kong listings earlier that month. [4]

In February 2026, StepFun released Step 3.5 Flash, an open-weight MoE language model with 196 billion total parameters and 11 billion active parameters per token, under Apache 2.0. [13][32] On February 25, 2026, Bloomberg reported that StepFun was considering a Hong Kong IPO that could raise about $500 million. [11] On April 2, 2026, the company converted from a limited liability company into a joint-stock company, and Yin Qi was registered as chairman. [20] Reports in April said it was also unwinding its Cayman Islands offshore structure ahead of the listing. [20]

In May 2026, Yicai reported that StepFun was completing a funding round of almost $2.5 billion with supply-chain investors including Huaqin Technology, Longcheer, OmniVision, and ZTE; StepFun did not comment. [21] On May 29, 2026, it released and open-sourced Step 3.7 Flash, a vision-language successor to Step 3.5 Flash. [34][33] In June 2026, Chinese financial media reported, citing people familiar with the matter, that StepFun had prepared or confidentially submitted a Hong Kong listing application and that its main investors had proposed a valuation of up to $12 billion. [23][24]

On September 20, 2026, StepFun introduced Step 5 Preview as its new flagship for agentic work, a 600-billion-parameter MoE model with 27 billion active parameters, a 1 million-token context window, and vision input, available through its API with open weights promised for October 15, 2026. [35]

Who runs StepFun?

NameRoleBackground
Jiang Daxin (姜大昕)Founder and CEOPhD in computer science from the University at Buffalo; assistant professor at Nanyang Technological University. Joined Microsoft Research Asia in 2007, moved to the Software Technology Center Asia in 2011, became a Microsoft global partner and the center's deputy head and chief scientist in 2017, and was promoted to vice president in March 2023. [15] Led work on Bing, Cortana, Azure cognitive services, and natural language understanding for Microsoft 365. [1]
Zhu Yibo (朱亦博)Co-founder and CTOPhD from UC Santa Barbara, bachelor's degree from Tsinghua University, Microsoft Research PhD Fellowship (2015). Began his career as a researcher at Microsoft Research, then was a director at ByteDance responsible for AI infrastructure, and was a technical lead for GPU products on Google Cloud before co-founding StepFun; KrASIA and Yicai identify him as CTO. [26][16][5][6]
Zhang Xiangyu (张祥雨)Chief ScientistCo-author of ResNet ("Deep Residual Learning for Image Recognition"), which a 2025 Nature analysis ranked as the most-cited paper of the twenty-first century. [30][28][60] Co-author of "Delving Deep into Rectifiers," which won the 2025 Helmholtz Prize. [29]
Jiao Binxing (焦斌星)Co-founder, head of dataGraduated from the University of Science and Technology of China through its joint PhD program with Microsoft Research Asia. Later led a core search team for Microsoft's Bing search engine. [27]
Yin Qi (印奇)Chairman (since January 2026)Co-founder of Megvii; chairman of Qianli Technology since 2024. [6][16] Registered as StepFun's chairman in April 2026. [20]

Expanded article table

What models has StepFun released?

StepFun has released models across language, vision, video, audio, and multimodal domains. The table below summarizes the major releases. Benchmark figures are StepFun's own measurements unless noted.

ModelReleaseTypeParametersKey details
Step-12023Language100B+First model; completed about two months after training began in July 2023. [19]
Step-1VCompleted November 2023Multimodal100B+Multimodal model for image understanding. [19][1]
Step-1.5VJuly 2024MultimodalNot disclosedUpgraded multimodal model, launched at WAIC 2024. [50]
Step-1XJuly 2024Image generationNot disclosedLaunched at WAIC 2024. [50]
Step-2July 2024 (preview March 2024)Language1T+ (MoE)Ranked fifth on LiveBench (November 2024). [3][19][50]
GOT-OCR 2.0September 2024OCR580MUnified end-to-end OCR model for text, tables, charts, formulas, sheet music, and geometric shapes. Apache 2.0. [37]
Step-Video-T2VFebruary 2025Video generation30BUp to 204 frames; video VAE with 16x16 spatial and 8x temporal compression. MIT license. [38]
Step-AudioFebruary 2025Speech interaction130BUnified speech understanding and generation; control over dialects, emotions, singing, and rap. Apache 2.0. [39]
Step-Audio-TTS-3BFebruary 2025Text-to-speech3BDistilled from Step-Audio; StepFun describes it as the first TTS model trained on a large-scale synthetic dataset and able to generate rap and humming. [39][47]
Step-Video-TI2VMarch 2025Image-to-videoBased on T2VAdds image-conditioned generation to Step-Video-T2V. MIT license. [48]
Step1X-EditApril 2025Image editingNot disclosedMultimodal LLM plus diffusion decoder. Apache 2.0. [40]
Step-3July 2025Multimodal reasoning321B total, 38B active (MoE)316B language model plus vision encoder; 65,536-token maximum context; MFA and AFD. Apache 2.0. [8][31]
Step-Audio 2 miniAugust 2025Speech-to-speech8BEnd-to-end audio understanding and speech conversation. Apache 2.0. [41][61]
NextStep-1August 2025Image generation14BAutoregressive generation with continuous image tokens and a 157M flow-matching head. ICLR 2026 Oral. [42][43]
Step-Audio-EditXNovember 2025Audio editing3BLLM-based reinforcement-learning model for editing emotion, speaking style, and paralinguistics. [44]
Step3-VL-10BJanuary 2026Vision-language10BStepFun says it rivals or surpasses open models 10 to 20 times its size. Apache 2.0. [45]
Step 3.5 FlashFebruary 2026Language reasoning196B total, 11B active (MoE)256K context; 3-way multi-token prediction. Apache 2.0. [13][32]
Step 3.7 FlashMay 2026Vision-language198B total (196B language + 1.8B vision), about 11B active256K context; three reasoning levels; up to 400 tokens per second. Apache 2.0. [33][34]
Step 5 PreviewSeptember 2026Multimodal (agentic)600B total, 27B active (MoE)1M context and vision input; open weights promised for October 15, 2026. [35]

Expanded article table

Step 3.5 Flash and Step 3.7 Flash

Step 3.5 Flash uses a 3:1 ratio of sliding-window to full-attention layers to support a 256K-token context window, and 3-way multi-token prediction for generation throughput of 100 to 300 tokens per second in typical use, peaking at 350 tokens per second for single-stream coding. [13] In StepFun's blog, it scored 97.3 on AIME 2025 and 96.2 on HMMT 2025 (average of February and November) in standard settings; with Python code execution, it scored 99.8 on AIME 2025 and 98.0 on HMMT November 2025. StepFun also reported 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0. [13] The company said it runs locally on hardware such as the Mac Studio M4 Max and NVIDIA DGX Spark. [13]

Step 3.7 Flash combines the 196B language backbone with a 1.8B vision encoder, activates about 11 billion parameters per token, and offers low, medium, and high reasoning levels. StepFun reported a score of 56.3 on SWE-Bench Pro and priced API access at $0.20 per million input tokens (cache miss) and $1.15 per million output tokens. [33]

Step 5 Preview

Step 5 Preview is StepFun's flagship for software engineering and professional knowledge work, with particular emphasis on finance. [35] Leiphone reported that on September 20, 2026 it ranked second among open models on the Artificial Analysis Intelligence Index. [36] A separate article covers the model in detail.

Products and platform

StepFun AI app (formerly Yuewen)

Yuewen (跃问) is StepFun's consumer AI assistant, launched by mid-2024. [1] Its "拍照问" visual search feature was the first in China to be integrated with the iPhone 16 Camera Control button, according to a December 2024 report. [18] The app now appears in Apple's China App Store as "阶跃AI" (StepFun AI), on the same listing (app ID 6502382318) that previously carried the Yuewen name. [52] The listing describes it as an agent for chat and task execution that uses Step 3.7 Flash, with a "Step Agent" assistant that accepts instructions across devices and can connect to a Feishu (Lark) bot. [52] StepFun also offers StepClaw, a desktop agent built on the open-source OpenClaw framework. [54]

StepFun Open Platform

StepFun operates a developer platform at platform.stepfun.com, and platform.stepfun.ai for international users, with OpenAI-compatible APIs. [31][33] In March 2026 it launched Step Plan, a monthly subscription for coding tools and agents, in four tiers from $6.99 to $99 per month. [55] Step Plan is now billed in monthly credits and provides an endpoint for Claude Code and the Anthropic SDK. [56] Step 3.5 Flash and Step 3.7 Flash are also available through OpenRouter. [32][33]

How is StepFun funded?

RoundDateAmountInvestors
AngelBefore 2024 (date not disclosed)Not disclosedHongShan (Sequoia China), Qiming Venture Partners, IDG Capital [20]
Series A2024Not disclosed5Y Capital, Shunwei Capital, Lenovo Capital [20]
Series BDecember 2024"Several hundred million dollars"Led by Fortera Capital, the private equity arm of Shanghai State-owned Capital Investment; Tencent, 5Y Capital, Qiming Venture Partners [2][17]
Series B+January 2026More than RMB 5 billion (about $717 million)New: Shanghai State-owned Capital Investment Leading Fund, China Life Private Equity, Pudong Venture Capital, Xuhui Capital, Wuxi Liangxi Fund, Xiamen ITG, Huaqin Technology. Returning: Tencent, Qiming, 5Y Capital [5][16]
Pre-IPO (reported)May 2026Nearly $2.5 billionHuaqin Technology, Longcheer, OmniVision, ZTE; Hong Kong Investment Corporation reported among shareholders [21][22]

Expanded article table

In June 2024, media reported that StepFun was raising a round at a valuation of about $2 billion. [19] The 21st Century Business Herald reported a further 2025 financing of possibly more than $500 million, alongside a strategic partnership with Shanghai State-owned Capital Investment. [20] In February 2026, Caijing reported that the pre-IPO round would close in two tranches at pre-money valuations of about $4 billion and $5 billion to $6 billion. [25] Sina Finance reported that the company raised more than RMB 20 billion in total in the four months to May 2026, and that its main investors had proposed an IPO valuation of up to $12 billion. [24][23] StepFun's 2025 revenue was close to RMB 500 million, and it projected about RMB 1.2 billion for 2026, according to Caijing. [25]

Strategic partnerships

Automotive: Geely and Qianli Technology

Geely describes StepFun as a strategic partner in its technology ecosystem. [7] The companies co-developed Agent OS with Qianli Technology, and the Geely Galaxy M9 ships with a cockpit agent that runs on StepFun's end-to-end voice model. [12][5] The Paper reported that the Galaxy M9 sold nearly 40,000 units in its first three months on the market, and that StepFun expected its models to be installed in more than one million vehicles in 2026. [16] Yin Qi's dual role as chairman of StepFun and Qianli Technology links the two companies more closely. [6]

Smartphones and devices

StepFun has worked with the smartphone makers Honor, OPPO, and ZTE. [4][18] In January 2026 a company insider told The Paper that about 60% of China's leading phone brands had partnered with StepFun, that its models were installed on more than 42 million devices, and that they served nearly 20 million people a day. [16] Huaqin Technology, a smartphone original design manufacturer, invested in the Series B+ round. [5]

Other partnerships

StepFun formed Caiyue Xingchen with Cailian Press, which launched the financial model Finstep and the consumer wealth assistant "Xiao Caishen," and worked with Guotai Junan on a securities-industry multimodal model. [18][21] It has also signed content partnerships with China Online and CNKI. [18]

How does StepFun differ from other Chinese AI startups?

StepFun has consistently emphasized multimodal AI. By December 2024 it had released models for language, image understanding, image generation, video generation, and speech, and it has continued with audio, vision-language, and image-editing models. [17][18]

Scaling law philosophy

The company has publicly tied its strategy to the scaling law hypothesis, which holds that model performance improves as model size, training data, and compute increase. In June 2024 Zhu Yibo said: "Computing power, systems, data, and algorithms are the cores in the pursuit of the scaling law." [1] Jiang has described trillion-parameter scale as the "basic threshold" for AGI. [50]

Architectural innovations

With Step-3, StepFun introduced two architectural innovations:

  • Multi-Matrix Factorization Attention (MFA): an attention mechanism that scales the number and dimension of attention heads through low-rank factorization of the query-key circuit, reducing KV cache use. [62] StepFun says Step-3 uses 22% of DeepSeek V3's per-token attention cost. [8]
  • Attention-FFN Disaggregation (AFD): a distributed inference system that places attention layers and feed-forward layers on separate subsystems. [9]

In its technical report, StepFun said Step-3 reached a peak decoding throughput of 4,039 tokens per GPU per second on Hopper GPUs (4K context, FP8, 50 ms per-token latency target), compared with 2,324 tokens per GPU per second for DeepSeek-V3 in the same setup; StepFun's blog described this as around 70% higher. [9][8]

Is StepFun open source?

StepFun has released many models under Apache 2.0 or MIT licenses, including Step-3, Step 3.5 Flash, Step 3.7 Flash, Step-Video-T2V, Step-Audio, Step1X-Edit, GOT-OCR 2.0, Step-Audio 2 mini, NextStep-1, and Step3-VL-10B. [31][32][33][38][39][40][37][61][42][45] Its code is hosted at github.com/stepfun-ai and its weights at huggingface.co/stepfun-ai. Step 5 Preview launched as an API-only model, with weights promised for October 2026. [35]

China's Six Little Tigers

The "Six Tigers" is a label Chinese media and investors apply to six large-model startups, each valued at more than $1 billion by 2024. [10] The roster used in most reports is Zhipu AI, MiniMax, Baichuan, Moonshot AI, StepFun, and 01.AI. [10][22]

CompanyFoundedFounderNotable backingPrimary focus
Zhipu AIJune 2019Zhang Peng (CEO)Tencent, Alibaba, Meituan [10]GLM language models
MiniMax2021Yan JunjieAlibaba [10]Consumer apps (Talkie), video generation (Hailuo) [10]
Baichuan IntelligenceMarch 2023Wang XiaochuanAlibaba, Tencent, Xiaomi [10]Language models [10]
Moonshot AIMarch 2023Yang ZhilinAlibaba [58]Long-context models, Kimi [10]
StepFunApril 2023Jiang DaxinTencent, Qiming, Shanghai state funds [5]Multimodal models, AI on devices and vehicles
01.AIJuly 2023Kai-Fu LeeAlibaba Cloud [10]Yi language models

Expanded article table

By 2026 the group's paths had diverged. Zhipu AI and MiniMax listed in Hong Kong in January 2026, and StepFun moved toward a listing of its own. [4][23]

Company operations

StepFun's registered address is on the 30th floor of No. 701 Yunjin Road, Xuhui District, Shanghai. [57] In an interview in early 2026, Jiang said the company had more than 500 employees, nearly 80% of them in algorithm and technical roles. [22] The name 阶跃星辰 combines 阶跃 ("step" or "leap") and 星辰 ("stars"). The company's websites are stepfun.com and stepfun.ai.

Research contributions

Beyond its commercial models, StepFun publishes research papers and open-source tools.

GOT-OCR 2.0

In September 2024, StepFun researchers published GOT (General OCR Theory), a 580-million-parameter end-to-end OCR model that treats all artificial optical signals, including plain text, math and molecular formulas, tables, charts, sheet music, and geometric shapes, as "characters." It pairs a high-compression encoder with a long-context decoder and can output plain or formatted results such as Markdown, TikZ, and SMILES. [37] The weights are Apache 2.0. [37]

NextStep-1

NextStep-1 is a 14-billion-parameter autoregressive image generation model paired with a 157-million-parameter flow-matching head, trained with next-token prediction on discrete text tokens and continuous image tokens rather than vector-quantized image tokens. [42] It was selected for an Oral presentation at ICLR 2026. [43] NextStep-1.1, released in December 2025, improved image quality through extended training and flow-based reinforcement learning. [59]

Step-DeepResearch

Step-DeepResearch, described in a December 2025 technical report, is a 32-billion-parameter end-to-end research agent trained through agentic mid-training, supervised fine-tuning, and reinforcement learning. StepFun reported a score of 61.4% on Scale AI's Research Rubrics and introduced ADR-Bench for Chinese-language deep research tasks. [46]

Competition

Within China, StepFun competes with the other five tigers and with large technology companies including Baidu, Alibaba (Qwen), ByteDance (Doubao), and DeepSeek. Its approach centers on "AI plus terminals": deploying models in phones, cars, and other devices rather than relying only on cloud APIs and a chatbot. [6][16]

U.S. export controls restrict Chinese AI companies' access to the most advanced Nvidia chips; in 2024 Zhu Yibo called the resulting challenges "manageable." [1] StepFun designed Step-3 to run efficiently on both flagship and lower-end accelerators, and its 2025 alliance with Chinese chipmakers aimed to optimize the model for domestic hardware. [8][49]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8South China Morning Post, "Shanghai AI start-up founded by ex-Microsoft engineers bets on 'scaling law' to boost AI capabilities," June 2024. scmp.com/...bets-scaling-law-boost-ai-capabilities
  2. ^SiliconANGLE, "Chinese AI model maker Stepfun raises hundreds of millions in Series B funding," December 26, 2024. siliconangle.com/...reds-millions-series-b-funding
  3. ^1 ^2MarkTechPost, "Chinese AGI Startup 'StepFun' Developed 'Step-2': A New Trillion-Parameter MoE Architecture Model Ranking 5th on Livebench," November 20, 2024. marktechpost.com/...model-ranking-5th-on-livebench
  4. ^1 ^2 ^3 ^4 ^5Caixin Global, "StepFun Raises $717 Million, Outpacing Newly Listed AI Rivals," January 26, 2026. caixinglobal.com/...wly-listed-ai-rivals-102408206
  5. ^1 ^2 ^3 ^4 ^5 ^6KrASIA, "China's investors double down on AI frontrunners as StepFun raises RMB 5 billion," January 26, 2026. kr-asia.com/...ers-as-stepfun-raises-rmb-5-billion
  6. ^1 ^2 ^3 ^4 ^5 ^6 ^7Yicai Global, "Chinese AI Firm Stepfun Raises USD719 Mln for Model Development, AI Agent Rollout," January 26, 2026. yicaiglobal.com/...development-terminal-agents-use
  7. ^1 ^2 ^3China Daily, "StepFun and Geely Auto open-source large models to global developers," February 19, 2025. chinadaily.com.cn/...WS67b58a09a310c240449d61da
  8. ^1 ^2 ^3 ^4 ^5 ^6StepFun, "Step3: Cost-Effective Multimodal Intelligence," July 31, 2025. stepfun.ai/...step3
  9. ^1 ^2 ^3StepFun, "Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding," arXiv:2507.19427, July 2025. arxiv.org/...2507.19427
  10. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10TechNode, "Meet China's top six AI unicorns: who are leading the wave of AI in China," January 9, 2025. technode.com/...re-leading-the-wave-of-ai-in-china
  11. ^Bloomberg (@business), "StepFun is considering an initial public offering in Hong Kong that may raise about $500 million," X, February 25, 2026. x.com/...2026608790771007731
  12. ^1 ^2 ^3Geely Auto Group, "Geely Auto Group Teams Up with StepFun for a Joint Showcase at the 2025 World Artificial Intelligence Conference," Business Wire via Yahoo Finance, July 31, 2025. finance.yahoo.com/...group-teams-stepfun-101400475
  13. ^1 ^2 ^3 ^4 ^5StepFun, "Step 3.5 Flash: Fast Enough to Think. Reliable Enough to Act.," February 2026. static.stepfun.com/...step-3.5-flash
  14. ^AsiaTechDaily, "Stepfun's $719M Raise Highlights China's Shift From AI Models to Real-World Deployment," January 2026. asiatechdaily.com/...dels-to-real-world-deployment
  15. ^1 ^2 ^3Leiphone (雷峰网), "独家丨前微软 NLP 大牛姜大昕创立新公司「阶跃星辰」" (Exclusive: former Microsoft NLP expert Jiang Daxin founds new company StepFun), December 2023. leiphone.com/...mJcRRqbgIh6UkVuk
  16. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8The Paper (澎湃新闻), "阶跃星辰完成超50亿元B+轮融资,印奇出任董事长" (StepFun completes B+ round of more than 5 billion yuan; Yin Qi becomes chairman), January 26, 2026. m.thepaper.cn/newsDetail_forward_32463044
  17. ^1 ^2 ^3 ^4The Paper (澎湃新闻), "阶跃星辰完成数亿美元B轮融资,核心投资方包括上海国投和旗下基金" (StepFun completes B round of several hundred million dollars), December 23, 2024. m.thepaper.cn/newsDetail_forward_29726642
  18. ^1 ^2 ^3 ^4 ^5 ^6 ^7Sina Tech (新浪科技), "AI 独角兽阶跃星辰完成数亿美元融资,国产 AI 六小龙迈入决赛圈" (AI unicorn StepFun raises several hundred million dollars), December 23, 2024. finance.sina.com.cn/...doc-ineamnxt2689626.shtml
  19. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Sina Finance (新浪财经), "最神秘的大模型独角兽,刚刚融了几亿美元" (The most mysterious large-model unicorn just raised several hundred million dollars), December 24, 2024. finance.sina.com.cn/...doc-ineapwes8461488.shtml
  20. ^1 ^2 ^3 ^4 ^5 ^6 ^721st Century Business Herald (21世纪经济报道), "AI独角兽阶跃星辰,加速赴港IPO?" (AI unicorn StepFun accelerating toward a Hong Kong IPO?), April 15, 2026. 21jingji.com/...675da063971325cad90b6beee7a06f8b
  21. ^1 ^2 ^3 ^4Yicai Global, "StepFun to Raise Nearly USD2.5 Billion as Chinese AI Startup Advances Hong Kong IPO," May 8, 2026. yicaiglobal.com/...-startup-advances-hong-kong-ipo
  22. ^1 ^2 ^3 ^4The Standard (Hong Kong), "Stepfun, China's AI Six Tigers, finishes new US$2.5b funding round for HK IPO," May 2026. thestandard.com.hk/...25b-funding-round-for-HK-IPO
  23. ^1 ^2 ^3Sina Finance (新浪财经), "AI'六小虎'之一阶跃星辰,据传已秘密递表港交所,估值达120亿美元" (StepFun reportedly files confidentially with HKEX at up to $12 billion valuation), June 10, 2026. finance.sina.com.cn/...doc-iniaxkrt1220200.shtml
  24. ^1 ^2Sina Finance (新浪财经), "传阶跃星辰最快今日提交香港IPO申请" (StepFun said to submit Hong Kong IPO application as early as today), June 8, 2026. finance.sina.com.cn/...doc-iniatatk5351000.shtml
  25. ^1 ^2Sina Finance (新浪财经), citing Caijing, "阶跃星辰计划年内港股上市,2025年收入约5亿元" (StepFun plans Hong Kong listing this year; 2025 revenue about 500 million yuan), February 27, 2026. finance.sina.com.cn/...doc-inhpfhts3829710.shtml
  26. ^1 ^2Yibo Zhu, personal homepage (accessed September 2026). yibozhu.com
  27. ^Sina Finance (新浪财经), "和DeepSeek一样强的公司,上海也潜伏一家" (Shanghai also has a company as strong as DeepSeek), February 16, 2025. finance.sina.com.cn/...doc-ineksafe5136606.shtml
  28. ^Nature, "Exclusive: the most-cited papers of the twenty-first century," April 15, 2025. nature.com/...d41586-025-01125-9
  29. ^IEEE Computer Society TCPAMI, "The Helmholtz Prize" (accessed September 2026). tc.computer.org/...the-helmholtz-prize
  30. ^He, Zhang, Ren, Sun, "Deep Residual Learning for Image Recognition," arXiv:1512.03385. arxiv.org/...1512.03385
  31. ^1 ^2 ^3 ^4 ^5StepFun, "step3" model card, Hugging Face. huggingface.co/...step3
  32. ^1 ^2 ^3 ^4StepFun, "Step-3.5-Flash" model card, Hugging Face. huggingface.co/...Step-3.5-Flash
  33. ^1 ^2 ^3 ^4 ^5 ^6StepFun, "Step-3.7-Flash" model card, Hugging Face. huggingface.co/...Step-3.7-Flash
  34. ^1 ^2StepFun, "Step 3.7 Flash: A high-efficiency Flash model for real-world agents," May 29, 2026. static.stepfun.com/...step-3.7-flash
  35. ^1 ^2 ^3 ^4 ^5StepFun (@StepFun_ai), "Introducing Step 5 Preview: Advancing the Pareto Frontier," X, September 20, 2026. x.com/...2101510462685003786
  36. ^Leiphone (雷峰网), "一天升一位!阶跃新旗舰 Step 5 Preview 晋级 AA 全球开源前二" (StepFun's new flagship Step 5 Preview rises to second among open models on Artificial Analysis), September 2026. leiphone.com/...XYtQue8lpJrEsmvS
  37. ^1 ^2 ^3 ^4Wei et al., "General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model," arXiv:2409.01704, September 2024. arxiv.org/...2409.01704
  38. ^1 ^2 ^3 ^4Ma et al., "Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model," arXiv:2502.10248, February 2025. arxiv.org/...2502.10248
  39. ^1 ^2 ^3 ^4 ^5Huang et al., "Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction," arXiv:2502.11946, February 2025. arxiv.org/...2502.11946
  40. ^1 ^2 ^3Liu et al., "Step1X-Edit: A Practical Framework for General Image Editing," arXiv:2504.17761, April 2025. arxiv.org/...2504.17761
  41. ^1 ^2Wu et al., "Step-Audio 2 Technical Report," arXiv:2507.16632 (v3 adds Step-Audio 2 mini), 2025. arxiv.org/...2507.16632
  42. ^1 ^2 ^3 ^4NextStep Team, "NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale," arXiv:2508.10711, August 2025. arxiv.org/...2508.10711
  43. ^1 ^2StepFun, "NextStep-1" GitHub repository. github.com/...NextStep-1
  44. ^1 ^2StepFun, "Step-Audio-EditX" model card, Hugging Face. huggingface.co/...Step-Audio-EditX
  45. ^1 ^2StepFun, "Step3-VL-10B" model card, Hugging Face. huggingface.co/...Step3-VL-10B
  46. ^1 ^2StepFun, "Step-DeepResearch Technical Report," arXiv:2512.20491, December 2025. arxiv.org/...2512.20491
  47. ^StepFun, "Step-Audio-TTS-3B" model card, Hugging Face. huggingface.co/...Step-Audio-TTS-3B
  48. ^StepFun, "stepvideo-ti2v" model card, Hugging Face. huggingface.co/...stepvideo-ti2v
  49. ^1 ^2TechNode, "Chinese AI chipmakers join forces with StepFun to counter Nvidia's return to China," July 30, 2025. technode.com/...to-counter-nvidias-return-to-china
  50. ^1 ^2 ^3 ^4 ^5 ^6 ^7Pandaily, "StepFun Releases Three Large Models of the Step Series," July 5, 2024. pandaily.com/...ee-large-models-of-the-step-series
  51. ^Yicai Global, "China's Geely and Stepfun Join Open-Source AI Trend With Two Models," February 18, 2025. yicaiglobal.com/...source-ai-trend-with-two-models
  52. ^1 ^2Apple App Store (China), "阶跃AI-阶跃星辰AI助手" (StepFun AI assistant) listing, app ID 6502382318, earlier titled "阶跃AI-原跃问APP" (accessed September 2026). apps.apple.com/...id6502382318
  53. ^Sina Finance (新浪财经), citing TMTPost, "中国AI变局:腾讯、百度接入DeepSeek模型,字节反思,'大模型六虎'加速分化" (Tencent and Baidu adopt DeepSeek; the six tigers diverge), February 17, 2025. finance.sina.com.cn/...doc-inektywk3579817.shtml
  54. ^TMTPost, "StepFun Unveils Desktop Version of StepClaw," 2026. en.tmtpost.com/...7921747
  55. ^StepFun (@StepFun_ai), "Step Plan is now live on StepFun Open Platform," X, March 22, 2026. x.com/...2035815197584023835
  56. ^StepFun, "Step Plan Overview," StepFun documentation (accessed September 2026). platform.stepfun.ai/...overview
  57. ^Qichamao (企查猫), "上海阶跃星辰智能科技股份有限公司" business registration record (accessed September 2026). qichamao.com/...5b055ee7a047f4f68ffc7a320a3e33a5
  58. ^Gulf News (Bloomberg), "Alibaba leads record deal to create $2.5 billion China AI firm," February 2024. gulfnews.com/...-billion-china-ai-firm-1.101305657
  59. ^StepFun, "NextStep-1.1" model card, Hugging Face. huggingface.co/...NextStep-1.1
  60. ^Slashdot, "The Most-Cited Papers of the Twenty-First Century," April 18, 2025. science.slashdot.org/...f-the-twenty-first-century
  61. ^1 ^2StepFun, "Step-Audio-2-mini" model card, Hugging Face. huggingface.co/...Step-Audio-2-mini
  62. ^Hu et al., "Multi-matrix Factorization Attention," arXiv:2412.19255, December 2024. arxiv.org/...2412.19255

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

13 revisions · v14 · 4,888 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: xg07 independent adversarial verification 2026-09-23 (V1); writer audit + 2 material + 8 minor fixed

Cite this page: AI Wiki. "StepFun." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/stepfun

Suggest edit