Sand.ai
Sand.ai is an artificial-intelligence company that develops video-generation models, research software, and commercial creation tools. Independent reporting traces the company to Beijing in early 2024 and identifies computer-vision researcher Cao Yue as its founder.[2][3] Its public work includes the MAGI model family, the daVinci-MagiHuman audio-video project, MagiAttention, MagiCompiler, the VidMuse creation service, and a paid generation API.[1][6]
The company combines open-weight research releases with hosted products. Its MAGI-1 and MAGI-2 Preview repositories contain model weights and inference code, while VidMuse and the Sand.ai API are commercial services with separate terms.[8][12][15][16] This distinction matters because Sand.ai uses open-source AI in its company positioning, but it has not published every model's complete training data and training pipeline.
36Kr and Yicai reported that Sand.ai completed two financing rounds totaling more than $100 million during roughly three months in 2026.[2][3] They did not report a public valuation or complete round-by-round terms. The company's claim that VidMuse passed a $10 million annual recurring revenue run rate is also a management-reported annualization, not audited annual revenue.[1][4][5]
Identity and leadership
Sand.ai's website uses the punctuated name "Sand.ai." The site describes the business as specializing in AI video generation, and 36Kr calls it a video-model and product company founded in January 2024.[1][2] Some corporate databases use a 2023 registration date, but the available sources do not establish that the legal registration and the operating launch occurred on the same date. "Early 2024" is the better-supported founding description.
Cao Yue, who publishes academically as Yue Cao, founded the company after work at Microsoft Research Asia, the Beijing Academy of Artificial Intelligence, and the Light Years Beyond startup.[1][2][6] He was a coauthor, not the sole author, of Swin Transformer. The official ICCV 2021 awards page lists that paper as the conference's Best Paper, also called the Marr Prize, and names Yue Cao among its authors.[7]
The corporate geography has two layers in current public materials. Chinese business coverage describes a Beijing-founded startup.[2][3] Sand.ai's privacy policy identifies SandAI Pte. Limited as the operator of its online services and says SandAI is headquartered in Singapore.[18] The reviewed pages do not explain the ownership or contractual relationship between the Beijing startup and the Singapore service company, so they should not be treated as interchangeable legal names without further documentation.
Development timeline
| Date | Event | Evidence status |
|---|---|---|
| Early 2024 | Sand.ai began operating as a video-generation company | Reported by 36Kr and other Chinese business outlets [2][6] |
| April 21, 2025 | MAGI-1 weights and inference code were released | Public repository and release history [8] |
| May 19, 2025 | The MAGI-1 technical report was posted to arXiv | Developer preprint, not an identified peer-reviewed paper [9] |
| January 2026 | VidMuse launched as a music-video creation agent | Reported by 36Kr; product remains live [2][15] |
| March 23, 2026 | daVinci-MagiHuman audio-video preprint was posted | Developer collaboration preprint [11] |
| April-June 2026 | Media reported two financing rounds totaling more than $100 million | Reported financing, with valuation undisclosed [2][3][4] |
| August 5, 2026 | MAGI-2 Preview report and artifacts were published | Public weights, code, configs, and report [12][13] |
Models and research
Sand.ai's model program is rooted in computer vision and multimodal generation. Its releases have moved from chunk-based video continuation toward synchronized audio-video output and a sparse model with many stored parameters. The technical papers and repositories establish that these systems exist. Most capability comparisons, however, come from the developers' own evaluations.
MAGI-1
MAGI-1 predicts video as a sequence of fixed-length chunks. Each chunk is denoised as a unit, while generation of later chunks is conditioned on earlier ones. The result combines an autoregressive model across time with a Diffusion Transformer inside each chunk. Calling it simply an autoregressive model can obscure this hybrid design.[8][9]
The repository provides 24-billion and 4.5-billion parameter variants, along with distilled and quantized versions. Its documented modes include text-to-video, image-to-video, and video continuation. The code and model repository are labeled Apache License 2.0.[8] Larger variants require multi-GPU reference configurations, while the published 4.5B configurations target a single 24 GB GPU. These are repository compatibility statements, not guarantees for every operating system or runtime.
The developer report claimed leading physical-consistency results at release.[9] Google DeepMind's current Physics-IQ repository gives the original MAGI-1 a score of 56.0 for multiframe video continuation and 30.2 for image-to-video. As of August 15, 2026, Cosmos3-Super was higher in both comparable unreranked categories, at 59.7 and 43.8.[10] Descriptions that MAGI-1 remained the benchmark leader were therefore stale by that date. Physics-IQ accepts developer-submitted results through pull requests, so its table is a useful common benchmark record rather than a blinded independent audit.
Audio-video models
Sand.ai and SII-GAIR collaborated on daVinci-MagiHuman, a single-stream transformer that places text, video, and audio representations in one sequence and generates synchronized video and audio. The March 2026 manuscript reports human-performance, speech, and synchronization results and links a public model stack.[11] Those performance values are author evaluations from a preprint; an independent replication was not identified.
The company also markets GAGA-1 as a hosted human-performance model or service. Public descriptions connect GAGA-1 with synchronized character speech and motion, but the exact correspondence between the hosted service version and every daVinci-MagiHuman checkpoint is not fully documented. A company article should not merge the product and research names into one artifact.
MAGI-2 Preview
MAGI-2 Preview is a public research release dated August 5, 2026. Its main model has about 114 billion stored parameters and, according to Sand.ai, activates about 6 billion per token through a fine-grained mixture of experts. It places text, audio, and video tokens in a single transformer stream and generates ten-second videos with an audio track.[12][13]
The package includes weights, inference code, configuration files, and a Docker image. It does not include the training dataset or training code. It also lacks a quantitative benchmark table, human preference study, controlled ablation, throughput benchmark, or published safety evaluation.[12][13] The parameter counts document architecture, not output quality. Sand.ai's priority claims about a "world first" 100B-scale video MoE remain company claims rather than independently established historical findings.
Systems software
MagiAttention is an Apache-licensed distributed attention implementation designed for long sequences and heterogeneous attention masks. MagiCompiler is a public compilation framework for model training and inference.[14] Both support the company's model work and are usable outside a single checkpoint. Their public repositories establish implementation availability, while their performance charts remain developer-run benchmarks.
Sand.ai sometimes describes video generation as a path toward a world model. That term is an aspiration in this context. Producing plausible video or scoring well on Physics-IQ does not by itself show that a model has learned a general, action-conditioned representation of the physical world.
Products and business model
VidMuse is Sand.ai's main publicly identified commercial product. Its site offers music-video and advertising workflows, including a browser product, a command-line interface, and an agent skill.[15] Company material describes a workflow that analyzes an uploaded song, plans shots, generates assets, assembles clips, and renders a finished project.
The company's own model research is not the only possible source of VidMuse output. In an April 2026 interview, founder Cao Yue said the product team first used the best available models to test product-market fit, then brought selected stages onto Sand.ai models when doing so improved quality, cost, or margin.[5] That statement rules out a simple claim that VidMuse is powered exclusively by MAGI or GAGA checkpoints.
Sand.ai also operates a generation API. Its documentation describes accounts linked to the Magi service, separately purchased API credits, bearer-token authentication, and generation requests that can combine an image with timed text conditions.[16] The documentation establishes an active commercial interface but does not disclose customer numbers, usage volume, gross margin, or service-level reliability.
Financing and reported revenue
In April 2026, 36Kr reported from unnamed sources that Sand.ai had completed a financing round of about $50 million.[4] In June, 36Kr and Yicai reported that two rounds completed within roughly three months totaled more than $100 million.[2][3] Reported participants included Look Capital, Lollapalooza Capital, Jiukun Venture Capital, Matrix Partners China, MSA Capital, Sinovation Ventures, Xianghe Capital, Source Code Capital, CAS Star, Capital Today, IDG, and Baidu Ventures. Xinghan Capital was named as financial adviser.[2][3]
This reporting supports the aggregate and investor list as media-reported facts. It does not support describing the event as one $100 million round, assigning a valuation, or calculating an exact lifetime total by adding partially disclosed earlier rounds. No public filing, cash schedule, dilution table, or valuation document was identified.
The revenue claim needs a similar qualification. Sand.ai's current About page says VidMuse exceeded $10 million in annual recurring revenue within two months of launch.[1] A 36Kr news report repeated that figure, and a published interview quoted VidMuse product lead Zhang Zihe making the same claim.[4][5] ARR annualizes a short run rate. It is not the same as recognized annual revenue, cash receipts, or profit. Sand.ai has not publicly disclosed audited statements, subscriber count, churn, refunds, net revenue retention, customer concentration, or gross margin. Yicai specifically reported that Cao Yue did not disclose whether the company was profitable.[3]
The claim is therefore best written as "Sand.ai reported a $10 million ARR run rate," not "Sand.ai earned $10 million." Claims that VidMuse reached the threshold faster than every competing product lack a published comparison method and are omitted.
Licensing and hosted-service data use
Sand.ai's public repositories demonstrate a substantial open-release program. MAGI-1, MagiAttention, and the top-level MAGI-2 Preview code and model repositories carry Apache-2.0 labels.[8][13][14] That does not make the entire company or every hosted product open source. Training data are not published for the model family, complete training code is absent from MAGI-2 Preview, and bundled components can have their own terms.
Hosted services are governed by separate contracts. Sand.ai's terms say users keep ownership of inputs and receive any company rights in their outputs, to the extent such rights exist. The same terms give SandAI broad licenses to host and operate user content and say inputs, outputs, and interactions may be used to train and improve its models unless otherwise agreed.[17] Users remain responsible for having rights to uploaded material and for the legality of outputs.
The privacy policy says the services collect prompts, images, video, outputs, account details, device and usage information, and transaction history. It permits sharing with service providers, including third-party AI providers, and says data may be transferred to Singapore or other jurisdictions. The service is not directed to children under 16.[18] Publicly shared outputs may be indexed or copied by other users and search engines.
These disclosures do not show that every input is used for training, but they mean users should not assume a private, local-only workflow when using the hosted products. Conversely, the Apache-licensed local checkpoints do not automatically inherit the hosted platform's monitoring, use restrictions, or content controls.
Evaluation and limitations
Sand.ai has released enough code and weights for external inspection and deployment, but evidence differs by project. MAGI-1 has a long technical report and entries on a third-party benchmark repository. daVinci-MagiHuman has a developer preprint. MAGI-2 Preview has downloadable artifacts but no quantitative capability evaluation.[9][10][11][12]
No company-wide training-data manifest, model safety report, provenance or watermark evaluation, or independent audit was identified. Hosted terms prohibit unlawful, infringing, non-consensual, and other harmful content and allow moderation, but contract language is not evidence that a model cannot generate such material.[17] Open checkpoints may be deployed without the hosted service's controls.
The company's technical direction is also not proof of general intelligence. Autoregressive chunking, synchronized audio-video generation, sparse scaling, and distributed attention are specific engineering approaches. Benchmark rankings can change as new systems appear, and product run-rate claims can change with retention and pricing. The most durable facts about Sand.ai are its founder, public artifacts, documented product interfaces, and independently reported financing, with performance and financial claims kept in their original attributed context.
References
- ^Sand.ai, *About Sand.ai*, accessed August 15, 2026. sand.ai/about-us
- ^Deng Yongyi, *Exclusive: Sand.ai's Cao Yue secures over USD 100M in financing: Why video is the most critical path to the world model*, 36Kr, June 29, 2026. 36kr.com/...3873965241931014
- ^Lu Qian, *Su Hua and Wang Huiwen bet on video-generation startup*, Yicai, June 22, 2026. yicai.com/...103240296
- ^36Kr, *VidMuse ARR exceeds $10 million and Sand.ai completes a financing round of approximately $50 million*, April 7, 2026. eu.36kr.com/...3756106572022274
- ^Source Code Capital, *Sand.ai completes a financing round of approximately $50 million and pursues a product-plus-model strategy*, including a GeekPark interview with Cao Yue and Zhang Zihe, April 14, 2026. sohu.com/...1009353758_701018
- ^QbitAI, *Tsinghua scholarship winner's team releases open video-generation AI and a 61-page technical report*, The Paper, April 23, 2025. m.thepaper.cn/newsDetail_forward_30699931
- ^ICCV 2021, *ICCV 2021 Paper Awards*, accessed August 15, 2026. iccv2021.thecvf.com/iccv-2021-paper-awards
- ^SandAI-org, *MAGI-1: Autoregressive Video Generation at Scale*, GitHub repository, accessed August 15, 2026. github.com/...MAGI-1
- ^Sand.ai et al., *MAGI-1: Autoregressive Video Generation at Scale*, arXiv:2505.13211, May 19, 2025. arxiv.org/...2505.13211
- ^Google DeepMind, *Physics-IQ benchmark leaderboard and evaluation code*, accessed August 15, 2026. github.com/...physics-iq-benchmark
- ^SII-GAIR and Sand.ai, *Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model*, arXiv:2603.21986, March 23, 2026. arxiv.org/...2603.21986
- ^Sand.ai, *MAGI-2 Preview: Scaling Video Generation Models Efficiently*, August 5, 2026. sand.ai/...magi-2-preview
- ^SandAI-org, *MAGI-2-preview*, GitHub repository, accessed August 15, 2026. github.com/...MAGI-2-preview
- ^SandAI-org, *Sand.ai GitHub organization and MagiAttention repository*, accessed August 15, 2026. github.com/SandAI-org
- ^VidMuse, *AI Music Video and Product Ads Generator*, accessed August 15, 2026. vidmuse.ai/en
- ^Sand.ai, *API Platform documentation*, accessed August 15, 2026. platform.sand.ai/docs
- ^Sand.ai, *Terms of Service*, accessed August 15, 2026. sand.ai/terms-of-service
- ^Sand.ai, *Privacy Policy*, accessed August 15, 2026. sand.ai/privacy-policy
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 2,283 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Cite this page: AI Wiki. "Sand.ai." aiwiki.ai, updated 15 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/sand_ai