StepFun
StepFun (Chinese: 阶跃星辰, pinyin: Jiēyuè Xīngchén) is a Shanghai-based Chinese artificial intelligence company that builds the "Step" series of large language models and multimodal foundation models. Its legal name is 上海阶跃星辰智能科技股份有限公司 (Shanghai Jieyue Xingchen Intelligent Technology Co., Ltd.). The company was registered on April 6, 2023, with Jiang Daxin, a former Microsoft vice president and chief scientist of Microsoft's Software Technology Center Asia, as executive director. [15][6] Chinese media and investors count it among the country's "Six Little Tigers" (六小虎) of large-model startups, alongside Zhipu AI, Moonshot AI, MiniMax, Baichuan Intelligence, and 01.AI. [10][22]
StepFun first drew attention with Step-2, a trillion-parameter mixture-of-experts (MoE) language model whose preview in March 2024 was described as the first trillion-parameter model shown by a Chinese startup, and with a multimodal strategy spanning language, vision, video, and audio. [19][50] In January 2026 it closed a Series B+ round of more than RMB 5 billion (about $717 million), which The Paper reported was the largest single financing in China's large-model sector over the previous 12 months, and named Yin Qi, co-founder of Megvii, as chairman. [4][16] In May 2026 it was reported to have raised nearly $2.5 billion more ahead of a planned Hong Kong listing, and on September 20, 2026 it introduced Step 5 Preview, a 600-billion-parameter MoE model. [21][35]
History
When was StepFun founded?
StepFun was registered on April 6, 2023, according to business-registry data cited by Leiphone, with Jiang Daxin as executive director and manager and Zhu Yibo as supervisor. [15] It is headquartered in Shanghai's Xuhui District. [17] A December 2024 profile republished by Sina Finance said the release of ChatGPT had a deep effect on Jiang and prompted him to leave Microsoft, where he had worked for 16 years, to start the company. [19][1] The core founding team included Zhu Yibo, who led systems, and Jiao Binxing, who led data; both had worked at Microsoft. [1][19][26]
The same profile said the team began training models in July 2023, finished Step-1, a language model with more than 100 billion parameters, about two months later, and completed Step-1V, a 100-billion-parameter multimodal model, in November 2023. [19] StepFun raised early funding from HongShan (formerly Sequoia China), Qiming Venture Partners, and IDG Capital in an angel round, followed by a 2024 Series A with 5Y Capital, Shunwei Capital, and Lenovo Capital, according to the 21st Century Business Herald. [20]
2024: Step-2 and rapid model releases
StepFun showed a preview of Step-2 at the Global Developer Pioneer Conference on March 23, 2024. [19] By June 2024 it had launched Step-1V and two consumer apps: Yuewen, a ChatGPT-like assistant, and Maopaoya, a character-chat product with game features. [1] At the World Artificial Intelligence Conference (WAIC) in Shanghai in July 2024, StepFun launched three models: [50]
- Step-2: the official version of its trillion-parameter language model, built on an MoE architecture.
- Step-1.5V: a multimodal understanding model.
- Step-1X: an image generation model.
At WAIC, Jiang said: "Climbing the peak of AGI, 'trillion parameters' and 'multi-modal fusion' are indispensable. The scale of trillion parameters is the basic threshold for achieving AGI; multi-modal large models are the only way to AGI." [50]
In November 2024, MarkTechPost reported that Step-2 ranked fifth on the LiveBench benchmark, behind models from OpenAI and Google, with scores of 86.57 in instruction following, 58.67 in reasoning, and 54.86 in data analysis. [3] Chinese coverage in December 2024 described it as first among Chinese models on that leaderboard. [18] By December 2024 the company said it had released 11 self-developed foundation models in the previous 10 months, covering language, image understanding, image generation, video generation, and speech. [17][18]
2025: Open-source push and automotive partnerships
On February 18, 2025, StepFun and Geely Auto Group jointly announced the open-sourcing of two models: Step-Video-T2V, a 30-billion-parameter text-to-video model that generates videos of up to 204 frames, and Step-Audio, a 130-billion-parameter speech interaction model. [7][38][39] Both were made available in the Yuewen app the same day. [7] Yicai reported that StepFun had worked with Geely on application scenario design, model evaluation, and engineering. [51] StepFun published the Step-Video-T2V weights under the MIT license and Step-Audio under Apache 2.0. [38][39] On February 16, 2025, Yuewen also added DeepSeek-R1. [53]
In April 2025, StepFun released Step1X-Edit, an open-source image editing model that uses a multimodal LLM to interpret editing instructions and a diffusion decoder to produce the edited image. StepFun said it approaches the performance of GPT-4o and Gemini 2 Flash on its own GEdit-Bench benchmark; the weights are Apache 2.0. [40]
On July 25, 2025, on the eve of WAIC in Shanghai, StepFun unveiled Step-3, a multimodal reasoning model with 321 billion total parameters and 38 billion active parameters, and, together with Geely and Qianli Technology, launched a preview of Agent OS, a smart-cockpit operating system. [12][31] Geely presented the Galaxy M9 as the first vehicle with a "human-like AI agent" powered by StepFun's end-to-end voice model. [12] Step-3 introduced two techniques: Multi-Matrix Factorization Attention (MFA), which StepFun says uses 22% of DeepSeek V3's per-token attention cost, and Attention-FFN disaggregation (AFD), which runs attention and feed-forward layers on separate groups of accelerators. [8][9] StepFun published the weights under Apache 2.0 on July 31, 2025. [8][31] At the same conference StepFun launched an alliance with Chinese AI chipmakers including Huawei Ascend, Cambricon, Biren, and Moore Threads to optimize Step-3 for domestic chips. [49]
Later in 2025 the company released the Step-Audio 2 mini speech model (August), NextStep-1 image generation (August), the Step-GUI series of on-device GUI agents (November), Step-Audio-EditX (November), and Step-DeepResearch (December). [41][42][5][44][46]
2026: New chairman, Flash models, and IPO preparation
On January 26, 2026, StepFun announced its Series B+ round of more than RMB 5 billion (about $717 million), and the same day appointed Yin Qi as chairman, responsible for overall strategy and technical direction. [4][16] Yin co-founded Megvii, one of China's early computer-vision unicorns, and since 2024 has been chairman of Qianli Technology, a Geely-backed smart-driving company formerly known as Lifan Technology. [6][14] StepFun said Yin would chair both companies and drive cooperation between them on its "AI plus terminals" strategy. [6] Caixin reported that the round exceeded the proceeds raised by Zhipu AI (HK$4.173 billion, about $535 million) and MiniMax (HK$4.596 billion) in their Hong Kong listings earlier that month. [4]
In February 2026, StepFun released Step 3.5 Flash, an open-weight MoE language model with 196 billion total parameters and 11 billion active parameters per token, under Apache 2.0. [13][32] On February 25, 2026, Bloomberg reported that StepFun was considering a Hong Kong IPO that could raise about $500 million. [11] On April 2, 2026, the company converted from a limited liability company into a joint-stock company, and Yin Qi was registered as chairman. [20] Reports in April said it was also unwinding its Cayman Islands offshore structure ahead of the listing. [20]
In May 2026, Yicai reported that StepFun was completing a funding round of almost $2.5 billion with supply-chain investors including Huaqin Technology, Longcheer, OmniVision, and ZTE; StepFun did not comment. [21] On May 29, 2026, it released and open-sourced Step 3.7 Flash, a vision-language successor to Step 3.5 Flash. [34][33] In June 2026, Chinese financial media reported, citing people familiar with the matter, that StepFun had prepared or confidentially submitted a Hong Kong listing application and that its main investors had proposed a valuation of up to $12 billion. [23][24]
On September 20, 2026, StepFun introduced Step 5 Preview as its new flagship for agentic work, a 600-billion-parameter MoE model with 27 billion active parameters, a 1 million-token context window, and vision input, available through its API with open weights promised for October 15, 2026. [35]
Who runs StepFun?
| Name | Role | Background |
|---|---|---|
| Jiang Daxin (姜大昕) | Founder and CEO | PhD in computer science from the University at Buffalo; assistant professor at Nanyang Technological University. Joined Microsoft Research Asia in 2007, moved to the Software Technology Center Asia in 2011, became a Microsoft global partner and the center's deputy head and chief scientist in 2017, and was promoted to vice president in March 2023. [15] Led work on Bing, Cortana, Azure cognitive services, and natural language understanding for Microsoft 365. [1] |
| Zhu Yibo (朱亦博) | Co-founder and CTO | PhD from UC Santa Barbara, bachelor's degree from Tsinghua University, Microsoft Research PhD Fellowship (2015). Began his career as a researcher at Microsoft Research, then was a director at ByteDance responsible for AI infrastructure, and was a technical lead for GPU products on Google Cloud before co-founding StepFun; KrASIA and Yicai identify him as CTO. [26][16][5][6] |
| Zhang Xiangyu (张祥雨) | Chief Scientist | Co-author of ResNet ("Deep Residual Learning for Image Recognition"), which a 2025 Nature analysis ranked as the most-cited paper of the twenty-first century. [30][28][60] Co-author of "Delving Deep into Rectifiers," which won the 2025 Helmholtz Prize. [29] |
| Jiao Binxing (焦斌星) | Co-founder, head of data | Graduated from the University of Science and Technology of China through its joint PhD program with Microsoft Research Asia. Later led a core search team for Microsoft's Bing search engine. [27] |
| Yin Qi (印奇) | Chairman (since January 2026) | Co-founder of Megvii; chairman of Qianli Technology since 2024. [6][16] Registered as StepFun's chairman in April 2026. [20] |
What models has StepFun released?
StepFun has released models across language, vision, video, audio, and multimodal domains. The table below summarizes the major releases. Benchmark figures are StepFun's own measurements unless noted.
| Model | Release | Type | Parameters | Key details |
|---|---|---|---|---|
| Step-1 | 2023 | Language | 100B+ | First model; completed about two months after training began in July 2023. [19] |
| Step-1V | Completed November 2023 | Multimodal | 100B+ | Multimodal model for image understanding. [19][1] |
| Step-1.5V | July 2024 | Multimodal | Not disclosed | Upgraded multimodal model, launched at WAIC 2024. [50] |
| Step-1X | July 2024 | Image generation | Not disclosed | Launched at WAIC 2024. [50] |
| Step-2 | July 2024 (preview March 2024) | Language | 1T+ (MoE) | Ranked fifth on LiveBench (November 2024). [3][19][50] |
| GOT-OCR 2.0 | September 2024 | OCR | 580M | Unified end-to-end OCR model for text, tables, charts, formulas, sheet music, and geometric shapes. Apache 2.0. [37] |
| Step-Video-T2V | February 2025 | Video generation | 30B | Up to 204 frames; video VAE with 16x16 spatial and 8x temporal compression. MIT license. [38] |
| Step-Audio | February 2025 | Speech interaction | 130B | Unified speech understanding and generation; control over dialects, emotions, singing, and rap. Apache 2.0. [39] |
| Step-Audio-TTS-3B | February 2025 | Text-to-speech | 3B | Distilled from Step-Audio; StepFun describes it as the first TTS model trained on a large-scale synthetic dataset and able to generate rap and humming. [39][47] |
| Step-Video-TI2V | March 2025 | Image-to-video | Based on T2V | Adds image-conditioned generation to Step-Video-T2V. MIT license. [48] |
| Step1X-Edit | April 2025 | Image editing | Not disclosed | Multimodal LLM plus diffusion decoder. Apache 2.0. [40] |
| Step-3 | July 2025 | Multimodal reasoning | 321B total, 38B active (MoE) | 316B language model plus vision encoder; 65,536-token maximum context; MFA and AFD. Apache 2.0. [8][31] |
| Step-Audio 2 mini | August 2025 | Speech-to-speech | 8B | End-to-end audio understanding and speech conversation. Apache 2.0. [41][61] |
| NextStep-1 | August 2025 | Image generation | 14B | Autoregressive generation with continuous image tokens and a 157M flow-matching head. ICLR 2026 Oral. [42][43] |
| Step-Audio-EditX | November 2025 | Audio editing | 3B | LLM-based reinforcement-learning model for editing emotion, speaking style, and paralinguistics. [44] |
| Step3-VL-10B | January 2026 | Vision-language | 10B | StepFun says it rivals or surpasses open models 10 to 20 times its size. Apache 2.0. [45] |
| Step 3.5 Flash | February 2026 | Language reasoning | 196B total, 11B active (MoE) | 256K context; 3-way multi-token prediction. Apache 2.0. [13][32] |
| Step 3.7 Flash | May 2026 | Vision-language | 198B total (196B language + 1.8B vision), about 11B active | 256K context; three reasoning levels; up to 400 tokens per second. Apache 2.0. [33][34] |
| Step 5 Preview | September 2026 | Multimodal (agentic) | 600B total, 27B active (MoE) | 1M context and vision input; open weights promised for October 15, 2026. [35] |
Step 3.5 Flash and Step 3.7 Flash
Step 3.5 Flash uses a 3:1 ratio of sliding-window to full-attention layers to support a 256K-token context window, and 3-way multi-token prediction for generation throughput of 100 to 300 tokens per second in typical use, peaking at 350 tokens per second for single-stream coding. [13] In StepFun's blog, it scored 97.3 on AIME 2025 and 96.2 on HMMT 2025 (average of February and November) in standard settings; with Python code execution, it scored 99.8 on AIME 2025 and 98.0 on HMMT November 2025. StepFun also reported 74.4% on SWE-bench Verified and 51.0% on Terminal-Bench 2.0. [13] The company said it runs locally on hardware such as the Mac Studio M4 Max and NVIDIA DGX Spark. [13]
Step 3.7 Flash combines the 196B language backbone with a 1.8B vision encoder, activates about 11 billion parameters per token, and offers low, medium, and high reasoning levels. StepFun reported a score of 56.3 on SWE-Bench Pro and priced API access at $0.20 per million input tokens (cache miss) and $1.15 per million output tokens. [33]
Step 5 Preview
Step 5 Preview is StepFun's flagship for software engineering and professional knowledge work, with particular emphasis on finance. [35] Leiphone reported that on September 20, 2026 it ranked second among open models on the Artificial Analysis Intelligence Index. [36] A separate article covers the model in detail.
Products and platform
StepFun AI app (formerly Yuewen)
Yuewen (跃问) is StepFun's consumer AI assistant, launched by mid-2024. [1] Its "拍照问" visual search feature was the first in China to be integrated with the iPhone 16 Camera Control button, according to a December 2024 report. [18] The app now appears in Apple's China App Store as "阶跃AI" (StepFun AI), on the same listing (app ID 6502382318) that previously carried the Yuewen name. [52] The listing describes it as an agent for chat and task execution that uses Step 3.7 Flash, with a "Step Agent" assistant that accepts instructions across devices and can connect to a Feishu (Lark) bot. [52] StepFun also offers StepClaw, a desktop agent built on the open-source OpenClaw framework. [54]
StepFun Open Platform
StepFun operates a developer platform at platform.stepfun.com, and platform.stepfun.ai for international users, with OpenAI-compatible APIs. [31][33] In March 2026 it launched Step Plan, a monthly subscription for coding tools and agents, in four tiers from $6.99 to $99 per month. [55] Step Plan is now billed in monthly credits and provides an endpoint for Claude Code and the Anthropic SDK. [56] Step 3.5 Flash and Step 3.7 Flash are also available through OpenRouter. [32][33]
How is StepFun funded?
| Round | Date | Amount | Investors |
|---|---|---|---|
| Angel | Before 2024 (date not disclosed) | Not disclosed | HongShan (Sequoia China), Qiming Venture Partners, IDG Capital [20] |
| Series A | 2024 | Not disclosed | 5Y Capital, Shunwei Capital, Lenovo Capital [20] |
| Series B | December 2024 | "Several hundred million dollars" | Led by Fortera Capital, the private equity arm of Shanghai State-owned Capital Investment; Tencent, 5Y Capital, Qiming Venture Partners [2][17] |
| Series B+ | January 2026 | More than RMB 5 billion (about $717 million) | New: Shanghai State-owned Capital Investment Leading Fund, China Life Private Equity, Pudong Venture Capital, Xuhui Capital, Wuxi Liangxi Fund, Xiamen ITG, Huaqin Technology. Returning: Tencent, Qiming, 5Y Capital [5][16] |
| Pre-IPO (reported) | May 2026 | Nearly $2.5 billion | Huaqin Technology, Longcheer, OmniVision, ZTE; Hong Kong Investment Corporation reported among shareholders [21][22] |
In June 2024, media reported that StepFun was raising a round at a valuation of about $2 billion. [19] The 21st Century Business Herald reported a further 2025 financing of possibly more than $500 million, alongside a strategic partnership with Shanghai State-owned Capital Investment. [20] In February 2026, Caijing reported that the pre-IPO round would close in two tranches at pre-money valuations of about $4 billion and $5 billion to $6 billion. [25] Sina Finance reported that the company raised more than RMB 20 billion in total in the four months to May 2026, and that its main investors had proposed an IPO valuation of up to $12 billion. [24][23] StepFun's 2025 revenue was close to RMB 500 million, and it projected about RMB 1.2 billion for 2026, according to Caijing. [25]
Strategic partnerships
Automotive: Geely and Qianli Technology
Geely describes StepFun as a strategic partner in its technology ecosystem. [7] The companies co-developed Agent OS with Qianli Technology, and the Geely Galaxy M9 ships with a cockpit agent that runs on StepFun's end-to-end voice model. [12][5] The Paper reported that the Galaxy M9 sold nearly 40,000 units in its first three months on the market, and that StepFun expected its models to be installed in more than one million vehicles in 2026. [16] Yin Qi's dual role as chairman of StepFun and Qianli Technology links the two companies more closely. [6]
Smartphones and devices
StepFun has worked with the smartphone makers Honor, OPPO, and ZTE. [4][18] In January 2026 a company insider told The Paper that about 60% of China's leading phone brands had partnered with StepFun, that its models were installed on more than 42 million devices, and that they served nearly 20 million people a day. [16] Huaqin Technology, a smartphone original design manufacturer, invested in the Series B+ round. [5]
Other partnerships
StepFun formed Caiyue Xingchen with Cailian Press, which launched the financial model Finstep and the consumer wealth assistant "Xiao Caishen," and worked with Guotai Junan on a securities-industry multimodal model. [18][21] It has also signed content partnerships with China Online and CNKI. [18]
How does StepFun differ from other Chinese AI startups?
StepFun has consistently emphasized multimodal AI. By December 2024 it had released models for language, image understanding, image generation, video generation, and speech, and it has continued with audio, vision-language, and image-editing models. [17][18]
Scaling law philosophy
The company has publicly tied its strategy to the scaling law hypothesis, which holds that model performance improves as model size, training data, and compute increase. In June 2024 Zhu Yibo said: "Computing power, systems, data, and algorithms are the cores in the pursuit of the scaling law." [1] Jiang has described trillion-parameter scale as the "basic threshold" for AGI. [50]
Architectural innovations
With Step-3, StepFun introduced two architectural innovations:
- Multi-Matrix Factorization Attention (MFA): an attention mechanism that scales the number and dimension of attention heads through low-rank factorization of the query-key circuit, reducing KV cache use. [62] StepFun says Step-3 uses 22% of DeepSeek V3's per-token attention cost. [8]
- Attention-FFN Disaggregation (AFD): a distributed inference system that places attention layers and feed-forward layers on separate subsystems. [9]
In its technical report, StepFun said Step-3 reached a peak decoding throughput of 4,039 tokens per GPU per second on Hopper GPUs (4K context, FP8, 50 ms per-token latency target), compared with 2,324 tokens per GPU per second for DeepSeek-V3 in the same setup; StepFun's blog described this as around 70% higher. [9][8]
Is StepFun open source?
StepFun has released many models under Apache 2.0 or MIT licenses, including Step-3, Step 3.5 Flash, Step 3.7 Flash, Step-Video-T2V, Step-Audio, Step1X-Edit, GOT-OCR 2.0, Step-Audio 2 mini, NextStep-1, and Step3-VL-10B. [31][32][33][38][39][40][37][61][42][45] Its code is hosted at github.com/stepfun-ai and its weights at huggingface.co/stepfun-ai. Step 5 Preview launched as an API-only model, with weights promised for October 2026. [35]
China's Six Little Tigers
The "Six Tigers" is a label Chinese media and investors apply to six large-model startups, each valued at more than $1 billion by 2024. [10] The roster used in most reports is Zhipu AI, MiniMax, Baichuan, Moonshot AI, StepFun, and 01.AI. [10][22]
| Company | Founded | Founder | Notable backing | Primary focus |
|---|---|---|---|---|
| Zhipu AI | June 2019 | Zhang Peng (CEO) | Tencent, Alibaba, Meituan [10] | GLM language models |
| MiniMax | 2021 | Yan Junjie | Alibaba [10] | Consumer apps (Talkie), video generation (Hailuo) [10] |
| Baichuan Intelligence | March 2023 | Wang Xiaochuan | Alibaba, Tencent, Xiaomi [10] | Language models [10] |
| Moonshot AI | March 2023 | Yang Zhilin | Alibaba [58] | Long-context models, Kimi [10] |
| StepFun | April 2023 | Jiang Daxin | Tencent, Qiming, Shanghai state funds [5] | Multimodal models, AI on devices and vehicles |
| 01.AI | July 2023 | Kai-Fu Lee | Alibaba Cloud [10] | Yi language models |
By 2026 the group's paths had diverged. Zhipu AI and MiniMax listed in Hong Kong in January 2026, and StepFun moved toward a listing of its own. [4][23]
Company operations
StepFun's registered address is on the 30th floor of No. 701 Yunjin Road, Xuhui District, Shanghai. [57] In an interview in early 2026, Jiang said the company had more than 500 employees, nearly 80% of them in algorithm and technical roles. [22] The name 阶跃星辰 combines 阶跃 ("step" or "leap") and 星辰 ("stars"). The company's websites are stepfun.com and stepfun.ai.
Research contributions
Beyond its commercial models, StepFun publishes research papers and open-source tools.
GOT-OCR 2.0
In September 2024, StepFun researchers published GOT (General OCR Theory), a 580-million-parameter end-to-end OCR model that treats all artificial optical signals, including plain text, math and molecular formulas, tables, charts, sheet music, and geometric shapes, as "characters." It pairs a high-compression encoder with a long-context decoder and can output plain or formatted results such as Markdown, TikZ, and SMILES. [37] The weights are Apache 2.0. [37]
NextStep-1
NextStep-1 is a 14-billion-parameter autoregressive image generation model paired with a 157-million-parameter flow-matching head, trained with next-token prediction on discrete text tokens and continuous image tokens rather than vector-quantized image tokens. [42] It was selected for an Oral presentation at ICLR 2026. [43] NextStep-1.1, released in December 2025, improved image quality through extended training and flow-based reinforcement learning. [59]
Step-DeepResearch
Step-DeepResearch, described in a December 2025 technical report, is a 32-billion-parameter end-to-end research agent trained through agentic mid-training, supervised fine-tuning, and reinforcement learning. StepFun reported a score of 61.4% on Scale AI's Research Rubrics and introduced ADR-Bench for Chinese-language deep research tasks. [46]
Competition
Within China, StepFun competes with the other five tigers and with large technology companies including Baidu, Alibaba (Qwen), ByteDance (Doubao), and DeepSeek. Its approach centers on "AI plus terminals": deploying models in phones, cars, and other devices rather than relying only on cloud APIs and a chatbot. [6][16]
U.S. export controls restrict Chinese AI companies' access to the most advanced Nvidia chips; in 2024 Zhu Yibo called the resulting challenges "manageable." [1] StepFun designed Step-3 to run efficiently on both flagship and lower-end accelerators, and its 2025 alliance with Chinese chipmakers aimed to optimize the model for domestic hardware. [8][49]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8South China Morning Post, "Shanghai AI start-up founded by ex-Microsoft engineers bets on 'scaling law' to boost AI capabilities," June 2024. scmp.com/...bets-scaling-law-boost-ai-capabilities
- ^SiliconANGLE, "Chinese AI model maker Stepfun raises hundreds of millions in Series B funding," December 26, 2024. siliconangle.com/...reds-millions-series-b-funding
- ^1 ^2MarkTechPost, "Chinese AGI Startup 'StepFun' Developed 'Step-2': A New Trillion-Parameter MoE Architecture Model Ranking 5th on Livebench," November 20, 2024. marktechpost.com/...model-ranking-5th-on-livebench
- ^1 ^2 ^3 ^4 ^5Caixin Global, "StepFun Raises $717 Million, Outpacing Newly Listed AI Rivals," January 26, 2026. caixinglobal.com/...wly-listed-ai-rivals-102408206
- ^1 ^2 ^3 ^4 ^5 ^6KrASIA, "China's investors double down on AI frontrunners as StepFun raises RMB 5 billion," January 26, 2026. kr-asia.com/...ers-as-stepfun-raises-rmb-5-billion
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Yicai Global, "Chinese AI Firm Stepfun Raises USD719 Mln for Model Development, AI Agent Rollout," January 26, 2026. yicaiglobal.com/...development-terminal-agents-use
- ^1 ^2 ^3China Daily, "StepFun and Geely Auto open-source large models to global developers," February 19, 2025. chinadaily.com.cn/...WS67b58a09a310c240449d61da
- ^1 ^2 ^3 ^4 ^5 ^6StepFun, "Step3: Cost-Effective Multimodal Intelligence," July 31, 2025. stepfun.ai/...step3
- ^1 ^2 ^3StepFun, "Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding," arXiv:2507.19427, July 2025. arxiv.org/...2507.19427
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10TechNode, "Meet China's top six AI unicorns: who are leading the wave of AI in China," January 9, 2025. technode.com/...re-leading-the-wave-of-ai-in-china
- ^Bloomberg (@business), "StepFun is considering an initial public offering in Hong Kong that may raise about $500 million," X, February 25, 2026. x.com/...2026608790771007731
- ^1 ^2 ^3Geely Auto Group, "Geely Auto Group Teams Up with StepFun for a Joint Showcase at the 2025 World Artificial Intelligence Conference," Business Wire via Yahoo Finance, July 31, 2025. finance.yahoo.com/...group-teams-stepfun-101400475
- ^1 ^2 ^3 ^4 ^5StepFun, "Step 3.5 Flash: Fast Enough to Think. Reliable Enough to Act.," February 2026. static.stepfun.com/...step-3.5-flash
- ^AsiaTechDaily, "Stepfun's $719M Raise Highlights China's Shift From AI Models to Real-World Deployment," January 2026. asiatechdaily.com/...dels-to-real-world-deployment
- ^1 ^2 ^3Leiphone (雷峰网), "独家丨前微软 NLP 大牛姜大昕创立新公司「阶跃星辰」" (Exclusive: former Microsoft NLP expert Jiang Daxin founds new company StepFun), December 2023. leiphone.com/...mJcRRqbgIh6UkVuk
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8The Paper (澎湃新闻), "阶跃星辰完成超50亿元B+轮融资,印奇出任董事长" (StepFun completes B+ round of more than 5 billion yuan; Yin Qi becomes chairman), January 26, 2026. m.thepaper.cn/newsDetail_forward_32463044
- ^1 ^2 ^3 ^4The Paper (澎湃新闻), "阶跃星辰完成数亿美元B轮融资,核心投资方包括上海国投和旗下基金" (StepFun completes B round of several hundred million dollars), December 23, 2024. m.thepaper.cn/newsDetail_forward_29726642
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Sina Tech (新浪科技), "AI 独角兽阶跃星辰完成数亿美元融资,国产 AI 六小龙迈入决赛圈" (AI unicorn StepFun raises several hundred million dollars), December 23, 2024. finance.sina.com.cn/...doc-ineamnxt2689626.shtml
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Sina Finance (新浪财经), "最神秘的大模型独角兽,刚刚融了几亿美元" (The most mysterious large-model unicorn just raised several hundred million dollars), December 24, 2024. finance.sina.com.cn/...doc-ineapwes8461488.shtml
- ^1 ^2 ^3 ^4 ^5 ^6 ^721st Century Business Herald (21世纪经济报道), "AI独角兽阶跃星辰,加速赴港IPO?" (AI unicorn StepFun accelerating toward a Hong Kong IPO?), April 15, 2026. 21jingji.com/...675da063971325cad90b6beee7a06f8b
- ^1 ^2 ^3 ^4Yicai Global, "StepFun to Raise Nearly USD2.5 Billion as Chinese AI Startup Advances Hong Kong IPO," May 8, 2026. yicaiglobal.com/...-startup-advances-hong-kong-ipo
- ^1 ^2 ^3 ^4The Standard (Hong Kong), "Stepfun, China's AI Six Tigers, finishes new US$2.5b funding round for HK IPO," May 2026. thestandard.com.hk/...25b-funding-round-for-HK-IPO
- ^1 ^2 ^3Sina Finance (新浪财经), "AI'六小虎'之一阶跃星辰,据传已秘密递表港交所,估值达120亿美元" (StepFun reportedly files confidentially with HKEX at up to $12 billion valuation), June 10, 2026. finance.sina.com.cn/...doc-iniaxkrt1220200.shtml
- ^1 ^2Sina Finance (新浪财经), "传阶跃星辰最快今日提交香港IPO申请" (StepFun said to submit Hong Kong IPO application as early as today), June 8, 2026. finance.sina.com.cn/...doc-iniatatk5351000.shtml
- ^1 ^2Sina Finance (新浪财经), citing Caijing, "阶跃星辰计划年内港股上市,2025年收入约5亿元" (StepFun plans Hong Kong listing this year; 2025 revenue about 500 million yuan), February 27, 2026. finance.sina.com.cn/...doc-inhpfhts3829710.shtml
- ^1 ^2Yibo Zhu, personal homepage (accessed September 2026). yibozhu.com
- ^Sina Finance (新浪财经), "和DeepSeek一样强的公司,上海也潜伏一家" (Shanghai also has a company as strong as DeepSeek), February 16, 2025. finance.sina.com.cn/...doc-ineksafe5136606.shtml
- ^Nature, "Exclusive: the most-cited papers of the twenty-first century," April 15, 2025. nature.com/...d41586-025-01125-9
- ^IEEE Computer Society TCPAMI, "The Helmholtz Prize" (accessed September 2026). tc.computer.org/...the-helmholtz-prize
- ^He, Zhang, Ren, Sun, "Deep Residual Learning for Image Recognition," arXiv:1512.03385. arxiv.org/...1512.03385
- ^1 ^2 ^3 ^4 ^5StepFun, "step3" model card, Hugging Face. huggingface.co/...step3
- ^1 ^2 ^3 ^4StepFun, "Step-3.5-Flash" model card, Hugging Face. huggingface.co/...Step-3.5-Flash
- ^1 ^2 ^3 ^4 ^5 ^6StepFun, "Step-3.7-Flash" model card, Hugging Face. huggingface.co/...Step-3.7-Flash
- ^1 ^2StepFun, "Step 3.7 Flash: A high-efficiency Flash model for real-world agents," May 29, 2026. static.stepfun.com/...step-3.7-flash
- ^1 ^2 ^3 ^4 ^5StepFun (@StepFun_ai), "Introducing Step 5 Preview: Advancing the Pareto Frontier," X, September 20, 2026. x.com/...2101510462685003786
- ^Leiphone (雷峰网), "一天升一位!阶跃新旗舰 Step 5 Preview 晋级 AA 全球开源前二" (StepFun's new flagship Step 5 Preview rises to second among open models on Artificial Analysis), September 2026. leiphone.com/...XYtQue8lpJrEsmvS
- ^1 ^2 ^3 ^4Wei et al., "General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model," arXiv:2409.01704, September 2024. arxiv.org/...2409.01704
- ^1 ^2 ^3 ^4Ma et al., "Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model," arXiv:2502.10248, February 2025. arxiv.org/...2502.10248
- ^1 ^2 ^3 ^4 ^5Huang et al., "Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction," arXiv:2502.11946, February 2025. arxiv.org/...2502.11946
- ^1 ^2 ^3Liu et al., "Step1X-Edit: A Practical Framework for General Image Editing," arXiv:2504.17761, April 2025. arxiv.org/...2504.17761
- ^1 ^2Wu et al., "Step-Audio 2 Technical Report," arXiv:2507.16632 (v3 adds Step-Audio 2 mini), 2025. arxiv.org/...2507.16632
- ^1 ^2 ^3 ^4NextStep Team, "NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale," arXiv:2508.10711, August 2025. arxiv.org/...2508.10711
- ^1 ^2StepFun, "NextStep-1" GitHub repository. github.com/...NextStep-1
- ^1 ^2StepFun, "Step-Audio-EditX" model card, Hugging Face. huggingface.co/...Step-Audio-EditX
- ^1 ^2StepFun, "Step3-VL-10B" model card, Hugging Face. huggingface.co/...Step3-VL-10B
- ^1 ^2StepFun, "Step-DeepResearch Technical Report," arXiv:2512.20491, December 2025. arxiv.org/...2512.20491
- ^StepFun, "Step-Audio-TTS-3B" model card, Hugging Face. huggingface.co/...Step-Audio-TTS-3B
- ^StepFun, "stepvideo-ti2v" model card, Hugging Face. huggingface.co/...stepvideo-ti2v
- ^1 ^2TechNode, "Chinese AI chipmakers join forces with StepFun to counter Nvidia's return to China," July 30, 2025. technode.com/...to-counter-nvidias-return-to-china
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Pandaily, "StepFun Releases Three Large Models of the Step Series," July 5, 2024. pandaily.com/...ee-large-models-of-the-step-series
- ^Yicai Global, "China's Geely and Stepfun Join Open-Source AI Trend With Two Models," February 18, 2025. yicaiglobal.com/...source-ai-trend-with-two-models
- ^1 ^2Apple App Store (China), "阶跃AI-阶跃星辰AI助手" (StepFun AI assistant) listing, app ID 6502382318, earlier titled "阶跃AI-原跃问APP" (accessed September 2026). apps.apple.com/...id6502382318
- ^Sina Finance (新浪财经), citing TMTPost, "中国AI变局:腾讯、百度接入DeepSeek模型,字节反思,'大模型六虎'加速分化" (Tencent and Baidu adopt DeepSeek; the six tigers diverge), February 17, 2025. finance.sina.com.cn/...doc-inektywk3579817.shtml
- ^TMTPost, "StepFun Unveils Desktop Version of StepClaw," 2026. en.tmtpost.com/...7921747
- ^StepFun (@StepFun_ai), "Step Plan is now live on StepFun Open Platform," X, March 22, 2026. x.com/...2035815197584023835
- ^StepFun, "Step Plan Overview," StepFun documentation (accessed September 2026). platform.stepfun.ai/...overview
- ^Qichamao (企查猫), "上海阶跃星辰智能科技股份有限公司" business registration record (accessed September 2026). qichamao.com/...5b055ee7a047f4f68ffc7a320a3e33a5
- ^Gulf News (Bloomberg), "Alibaba leads record deal to create $2.5 billion China AI firm," February 2024. gulfnews.com/...-billion-china-ai-firm-1.101305657
- ^StepFun, "NextStep-1.1" model card, Hugging Face. huggingface.co/...NextStep-1.1
- ^Slashdot, "The Most-Cited Papers of the Twenty-First Century," April 18, 2025. science.slashdot.org/...f-the-twenty-first-century
- ^1 ^2StepFun, "Step-Audio-2-mini" model card, Hugging Face. huggingface.co/...Step-Audio-2-mini
- ^Hu et al., "Multi-matrix Factorization Attention," arXiv:2412.19255, December 2024. arxiv.org/...2412.19255
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
13 revisions · v14 · 4,888 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: xg07 independent adversarial verification 2026-09-23 (V1); writer audit + 2 material + 8 minor fixed
Cite this page: AI Wiki. "StepFun." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/stepfun