Sora

RawGraph

Sora was a text-to-video generation model developed by OpenAI. OpenAI revealed the research model on February 15, 2024 and granted early access to red teamers, visual artists, designers, and filmmakers. The company released a faster product version, Sora Turbo, to eligible ChatGPT Plus and Pro subscribers on December 9, 2024.[1][5]

Sora generated silent video from text and could also accept images or video as conditioning input. OpenAI's research report described it as a diffusion model whose transformer operated on spacetime patches of compressed visual data. The report demonstrated generation at varied durations, resolutions, and aspect ratios, including samples up to one minute, but did not disclose the model's parameter count, full architecture, training compute, or complete dataset.[2]

The first-generation Sora product was superseded by Sora 2 in September 2025. OpenAI discontinued the Sora web and app experiences on April 26, 2026. As of the July 28, 2026 research cutoff for this article, the company said the Sora API would be discontinued on September 24, 2026.[13][15]

Scope and naming

This article covers the model previewed as Sora in February 2024 and the Sora Turbo product released that December. It does not treat every later use of the Sora name as the same model. Sora 2 added synchronized audio and was deployed through a distinct social Sora app; both have separate articles.[13]

"Sora" can therefore refer to a research model, the faster Sora Turbo model, or the product through which Sora models were offered. OpenAI did not publish a dated model identifier or a technical comparison that would allow every internal change between the February preview and December product to be reconstructed. Claims about an exact parameter count, proprietary model size, or undocumented intermediate version are not established by the public technical record.[2][5][6]

Release history

Research preview

OpenAI introduced Sora as a research preview rather than a public service. Its launch page said the model could make videos up to one minute while maintaining visual quality and following a prompt. It also showed image-conditioned generation and video extension. The examples on the February page were described as direct model outputs without modification, but they were a selected showcase rather than a random or independently audited sample.[1]

The accompanying technical report framed the work as an experiment in large-scale visual generative modeling. It compared samples produced at base, four-times, and 32-times training compute with fixed inputs and seeds. Sample quality improved in that qualitative comparison. OpenAI did not publish a numerical scaling law or enough implementation detail to reproduce the training run, so the comparison supports a limited claim about the displayed samples rather than a general performance guarantee.[2]

OpenAI expanded early access to creative professionals during 2024. A March showcase included work by filmmakers, designers, and artists, including the short film Air Head by the production company shy kids. OpenAI stated that these videos were edited by the artists and that participants could modify generated material. The finished works therefore demonstrated Sora as one component in a production workflow, not unedited one-prompt output.[4]

Sora Turbo launch

Sora Turbo launched on December 9, 2024 at sora.com for eligible ChatGPT Plus and Pro users. It could generate video in widescreen, vertical, or square formats at up to 1080p and 20 seconds. Its interface included Storyboard, and users could extend, remix, or blend their own assets. They could start from text or upload an image or video where the product's rules permitted it.[5]

At launch, Plus users were offered up to 50 priority videos at 480p each month, or fewer videos at 720p. Pro users received ten times as much usage, higher resolution, and longer duration. These were launch entitlements, not permanent model specifications. OpenAI initially excluded users under 18, ChatGPT Team, Enterprise, and Edu accounts, and users in the United Kingdom, Switzerland, and the European Economic Area.[5]

Demand briefly exceeded available service capacity. The Associated Press reported that OpenAI temporarily disabled new Sora account creation during the release surge. It also reported that most users could not generate people at launch while OpenAI monitored misuse, with a limited group of testers receiving that capability.[7]

Architecture and disclosed training

Sora represented videos in a compressed latent space. A compression network mapped raw video into a representation compressed across space and time, and a decoder mapped generated latent data back to pixels. The compressed representation was divided into spacetime patches, which served as tokens for a transformer. During diffusion training, the model learned to predict clean patches from noisy patches while conditioned on information such as a text prompt.[2]

This design belongs to the broader diffusion transformer lineage. Peebles and Xie's 2023 DiT paper had shown that a transformer could replace the conventional U-Net backbone in latent image diffusion and that increasing forward-pass compute improved image-generation results in their experiments. Sora extended the patch-and-transformer approach to joint training on images and videos of varied sizes and durations. The DiT paper provides architectural lineage, but it does not disclose Sora's proprietary layer count, parameter layout, or training recipe.[2][3]

OpenAI said it trained the research system jointly on images and videos at their native aspect ratios and resolutions. It used recaptioning, in which a descriptive captioner produced detailed training captions, and used a GPT model at inference time to expand short user prompts. The company reported that native-size training improved framing in an internal comparison with a square-cropped model. That comparison was vendor-reported and qualitative.[2]

The December system card grouped model data into selected publicly available data, proprietary data from partnerships, and human feedback. It did not enumerate all sources, give a total number of videos or images, or state how much each category contributed. OpenAI has public content partnerships, but the system card's examples of partners do not prove that every named partner's assets were used to train Sora. More specific claims about particular libraries or percentages are therefore unwarranted.[6]

Capabilities

Generation and editing modes

The research model accepted several forms of conditioning. Text prompts specified subjects, settings, motion, and style. A still image could establish the first frame and visual content. A video input could be extended forward or backward, changed through an SDEdit-style process, or connected to another video by generating an intermediate sequence. OpenAI also demonstrated producing a loop by extending a video in both temporal directions and arranging the result.[2]

Training across varied dimensions let the research model generate landscape, portrait, and other aspect ratios without first forcing every source into a square crop. OpenAI displayed 1920 by 1080 and 1080 by 1920 examples and said the same model could prototype at lower resolution before generating at a larger size. Sora could also generate still images; the technical report showed images up to 2048 by 2048 pixels. These research demonstrations should not be confused with the 20-second, up-to-1080p limits of the December Sora Turbo product.[2][5]

The model sometimes maintained characters, scene layout, or simulated camera motion over comparatively long clips. OpenAI presented those behaviors as emerging from scale and argued that they could be steps toward general-purpose simulators. The report also said they occurred "often, though not always." Describing Sora as a proven world model would go beyond the evidence: the report offered selected qualitative examples, and OpenAI explicitly documented failures in physical and causal consistency.[2]

Independent evaluation

Early independent analysis was constrained because researchers could inspect only the public clips chosen by OpenAI. A February 2024 preprint collected ten representative Sora clips from public sources and used six in its quantitative comparison. The authors generated comparison videos with Pika and Runway Gen-2 from the first frame of each Sora clip and the same text prompt. Their geometry-based measures and 3D reconstructions favored Sora, suggesting stronger frame-to-frame geometric consistency in that small sample.[8]

The study did not establish general physical understanding. Its Sora sample was small and curated, and Sora's longer clips provided more views for reconstruction than the competing systems. The paper itself attributed part of Sora's advantage to longer duration and additional camera information. Its result is best read as evidence about geometric consistency in selected public examples, not as a comprehensive benchmark of the unreleased model.[8]

A broader peer-reviewed evaluation reached a more critical result. The Physics-IQ study, published at WACV 2026, assembled 396 eight-second real videos covering 66 scenarios in solid mechanics, fluid dynamics, optics, thermodynamics, and magnetism. The study evaluated Sora in image-to-video mode: it provided a manually selected "switch frame", the last frame of a three-second conditioning clip, together with a human-written caption, and asked Sora to generate a five-second continuation. Sora did not receive the full conditioning clip. Scores were normalized so that the variation between two real recordings of a scenario corresponded to 100.[9]

Sora scored 10.0 on the composite Physics-IQ measure; the best system in the study scored 29.5. The authors found no statistically significant relationship between their visual-realism measure and physical-understanding score. The benchmark was intentionally difficult and used proxy metrics that penalized hallucinated objects, camera movement, and shot changes. Sora produced camera or shot changes more often than several other systems, so the number should not be interpreted as a universal percentage of physical facts understood. It nevertheless supports the narrower conclusion that realistic-looking output did not reliably preserve the tested physical processes.[9]

Limitations

OpenAI's February materials identified recurring failures in basic physics, causal order, spatial directions, and temporal consistency. A generated person might take a bite from a cookie without leaving a bite mark, objects might appear spontaneously, animals or people might duplicate, and a scene could confuse left with right. Complex interactions and precise camera trajectories were particularly difficult.[1][2]

Sora could also produce implausible state changes or merge entities. The technical report showed a glass shattering while remaining intact, a chair duplicating, and a basketball passing through a hoop before exploding. Long videos could become incoherent, especially when objects were occluded or moved outside the frame. These examples qualify OpenAI's claims about object persistence and 3D consistency.[2]

Sora Turbo did not remove those problems. At public launch, OpenAI said it still generated unrealistic physics and struggled with complex actions over long durations. Prompt adherence also did not guarantee factual or safe output. A visually convincing clip could depict an event that never happened, which made provenance and restrictions on human likeness central deployment concerns.[5][6]

Safety and provenance

Before public release, OpenAI said external red teamers in nine countries tested more than 15,000 generations from September through December 2024. The company iterated on mitigations after testers identified adversarial prompts and gaps. This was a vendor-organized exercise whose system card documented methods and internal evaluations; it was not an independent certification that the service could not be misused.[6]

The product used multiple layers of moderation. OpenAI described classifiers applied to text, uploaded media, and generated output; specialized filters for nudity, minors, violence, deceptive content, and likeness misuse; and textual blocklists. Some videos could be blocked before delivery. Restrictions on depicting people were deliberately stricter at launch, and the general public initially could not upload images or videos containing people.[5][6][7]

OpenAI said every Sora Turbo video included C2PA metadata and that visible watermarks were applied by default. It also documented an internal reverse-video search tool and other detection systems. These measures could help identify origin, but they did not make generated media impossible to crop, re-encode, or mislabel. The system card presented provenance as one part of a broader safety stack rather than a guarantee against deepfakes.[5][6]

Creative use and reception

Sora's preview prompted both experimentation and concern in film and advertising. In February 2024, filmmaker Tyler Perry said he put a planned $800 million expansion of his Atlanta studio on hold after viewing the demonstrations. The plan would have added 12 sound stages. His statement documented one producer's business decision and concern about jobs; it does not show that Sora caused an industry-wide reduction in employment.[11]

In June 2024, Toys "R" Us and agency Native Foreign released a brand film built around Sora-generated imagery. Reporting said human scriptwriters and visual-effects artists were still needed to create usable generations and assemble the finished film. Toys "R" Us said the production used corrective visual effects and an original score composed by Aaron Marsh; it named Kim Miller Olko as executive producer and Native Foreign's Nik Kleverov as director. The example showed that Sora could contribute production material before public release, while also illustrating the difference between raw model output and a finished commercial.[12]

Relations with some early-access artists were contentious. In November 2024, a group of artists protesting what they described as inadequate compensation published access to Sora through a public interface. The Washington Post reported that OpenAI suspended access after the leak. The protest showed disagreement over the terms of creative testing; it did not establish that every early-access artist shared the organizers' position.[10]

Successor and discontinuation

OpenAI released Sora 2 on September 30, 2025 and described it as more controllable and more physically accurate than its prior systems, with synchronized dialogue and sound effects. Those comparisons were OpenAI's launch claims. Sora 2 was a successor rather than a feature update that should be retroactively assigned to first-generation Sora, which generated silent video.[13]

In March 2026, OpenAI announced that it would close the Sora social app. The Associated Press reported the announcement and noted public concerns about nonconsensual imagery and realistic synthetic media, but OpenAI's brief notice did not provide a definitive reason for the shutdown. Claims that a particular compute cost, revenue figure, engagement number, or copyright dispute caused the decision remain speculative unless OpenAI documents that explanation.[14]

OpenAI's support notice says the web and app experiences were discontinued on April 26, 2026. It directs users to a sunset page to export their content and says associated data will be permanently deleted after any final export window. The same notice schedules API discontinuation for September 24, 2026. Because that date was still in the future at this article's July 28 research cutoff, it is a scheduled event rather than a completed one.[15]

See also

References

  1. ^OpenAI. "Sora: Creating video from text." February 15, 2024. openai.com/...sora
  2. ^Brooks, T., et al. "Video generation models as world simulators." OpenAI, February 15, 2024. openai.com/...eneration-models-as-world-simulators
  3. ^Peebles, W., and Xie, S. "Scalable Diffusion Models with Transformers." *Proceedings of the IEEE/CVF International Conference on Computer Vision*, 2023. openaccess.thecvf.com/...nsformers_ICCV_2023_paper
  4. ^OpenAI. "Sora first impressions." March 25, 2024. openai.com/...sora-first-impressions
  5. ^OpenAI. "Sora is here." December 9, 2024. openai.com/...sora-is-here
  6. ^OpenAI. "Sora System Card." December 9, 2024. openai.com/...sora-system-card
  7. ^O'Brien, M. "OpenAI releases AI video generator Sora but limits how it depicts people." Associated Press, December 10, 2024. apnews.com/...deo-214d578d048f39c9c7b327f870dc6df8
  8. ^Li, X., et al. "Sora Generates Videos with Stunning Geometrical Consistency." arXiv:2402.17403, February 27, 2024. arxiv.org/...2402.17403
  9. ^Motamed, S., et al. "Do Generative Video Models Understand Physical Principles?" *Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision*, 2026. openaccess.thecvf.com/...rinciples_WACV_2026_paper
  10. ^Tiku, N. "OpenAI hits pause on video model Sora after artists leak access in protest." *The Washington Post*, November 26, 2024. washingtonpost.com/...-video-model-artists-protest
  11. ^Edwards, B. "Tyler Perry puts $800 million studio expansion on hold because of OpenAI's Sora." *Ars Technica*, February 23, 2024. arstechnica.com/...wood-warning-over-ai-video-tech
  12. ^Edwards, B. "Toys 'R' Us riles critics with 'first-ever' AI-generated commercial using Sora." *Ars Technica*, June 26, 2024. arstechnica.com/...generated-commercial-using-sora
  13. ^OpenAI. "Sora 2 is here." September 30, 2025. openai.com/...sora-2
  14. ^O'Brien, M. "OpenAI is shutting down Sora, its experimental AI video app." Associated Press, March 24, 2026. apnews.com/...c60de960536923f33edc04b92ddbe1cd
  15. ^OpenAI Help Center. "What to know about the Sora discontinuation." help.openai.com/...-about-the-sora-discontinuation

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

12 revisions · v13 · 2,598 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent 2026-07-28 fact-check: 15 explicit first-party, scholarly, peer-reviewed, and dated independent-reporting sources; 45 resolved citation calls; 15 canonical published internal targets; and all 15 cited source groups independently rechecked. Root replayed all 151 corrected-parent checksums and 37 root checksums, inspected desktop and mobile contact sheets representing all 32 captures, and verified first-generation scope, research-preview and public-release chronology, disclosed architecture and training boundaries, launch limits, independent evaluation protocols, limitations, safety, provenance, creative use, successor, and discontinuation. Exactly five source-precision corrections establish the early-access cohort, launch-tool wording, Physics-IQ switch-frame protocol, Toys R Us human roles, and reverse-video-search terminology, with no other semantic change. Root inspected all 22 preservation mappings and explicitly authorized the 2,598-word candidate as a 59.76459656187083-percent protected-shorter replacement for the 6,457-word baseline. The coordinated infobox correction separates the preview and release, bounds first-generation modalities and disclosed training, identifies Sora 2, and distinguishes completed web/app discontinuation from the scheduled API date. The terminal guard made one SELECT-only call, zero writes, passed all 29 checks, and confirmed the exact live-and-stamped Claude 3.5 Sonnet predecessor. Verification follows only after exact infobox and article postflights.

Cite this page: AI Wiki. "Sora." aiwiki.ai, updated 31 Jul 2026, fact-checked 31 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/sora

Suggest edit