Ox Alpha
Ox Alpha was the anonymous preview alias for GLM-5.3-Flash, a natively multimodal large language model developed by Z.ai.
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 301-360 of 408 articles
Ox Alpha was the anonymous preview alias for GLM-5.3-Flash, a natively multimodal large language model developed by Z.ai.
Phi-4 is a 14-billion-parameter small language model developed by Microsoft Research and released in December 2024, designed to match or beat models several times its size on reasoning tasks by training…
Phi-4-reasoning is a 14 billion parameter open weight reasoning model released by Microsoft Research on April 30, 2025.
Phi-4-mini is a 3.8 billion parameter open weight small language model released by Microsoft on February 26, 2025, under the permissive MIT license.
Phi-4-mini-flash-reasoning is a 3.8 billion parameter open weight reasoning model released by Microsoft in July 2025.
Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…
Pika 2.5 is a generative video model developed by Pika Labs, the San Francisco based AI video startup co-founded in April 2023 by Stanford AI Lab dropouts Demi Guo and Chenlin Meng.
Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.
Pixtral Large is a 124-billion-parameter multimodal (vision-language) large language model released by Mistral AI on November 18, 2024.
Question answering (QA) models are natural language processing systems that take a natural-language question as input and return a natural-language answer, optionally grounded in a supplied passage, document…
Qwen-Image is an open-weight image-generation foundation model released by Alibaba's Qwen team in August 2025.
Qwen-Image-3.0 is a text-to-image foundation model announced by Alibaba's Qwen team on July 21, 2026, as the third generation of the Qwen-Image series .
Qwen3 is the third-generation family of large language models developed by the Qwen Team at Alibaba Cloud (also known as Tongyi Qianwen lab).
Qwen3-Max is the flagship large language model in Alibaba's Qwen series and the first Qwen model to cross one trillion parameters, released in preview on September 5, 2025 and formally launched at the Apsara…
Qwen3-Next is an efficiency-focused large language model and model architecture released in September 2025 by the Qwen team at Alibaba Cloud.
Qwen3.8 is the name Qwen uses for a model generation that includes hosted services and downloadable checkpoints.
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
RFM-1 (Robotics Foundation Model 1) is an 8 billion parameter multimodal transformer for robotic manipulation announced by Covariant on March 11, 2024 at the MODEX 2024 trade show in Atlanta.
RT-2 (Robotic Transformer 2) is a vision-language-action model developed by Google DeepMind that enables robots to execute novel tasks by transferring knowledge from internet-scale vision-language pretraining…
Recraft V3 is a text-to-image generation model developed by Recraft AI and released on October 30, 2024, that became the first model to reach the number-one position on the Artificial Analysis Text-to-Image…
Reka Core is a frontier class multimodal foundation model developed by Reka AI, a research and product company founded in 2022 by former scientists from DeepMind, Google Brain, Meta FAIR, and Baidu.
Reka Edge is a 7-billion-parameter multimodal language model developed by Reka AI, introduced in April 2024 as the smallest member of the company's first publicly described model family.
Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.
Ring-1T is an open-weight reasoning model released in October 2025 by inclusionAI, the open-source research initiative associated with Ant Group
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
A robot foundation model is a large-scale machine learning model, typically based on the transformer architecture, that is pre-trained on broad, diverse datasets of robot interactions and then adapted to a…
Robotics models are machine learning systems that give robots the ability to perceive their surroundings, plan actions, and execute motor control.
Rodin Gen-2 is a generative artificial intelligence model for producing three dimensional assets from text prompts or reference images, developed by the Shanghai based company Deemos and delivered through the…
Runway Act-Two is a generative motion capture and character animation model developed by Runway, publicly introduced on July 15, 2025.
Runway Aleph is an in-context AI video editing model from Runway that transforms and edits an existing video clip from a plain-text instruction, performing a wide range of tasks in a single model: adding…
Runway Gen-4 is the fourth-generation video generation model developed by Runway (company), announced and released on March 31, 2025.
Runwayml/stable-diffusion-v1-5 is the Hugging Face repository name of the Stable Diffusion v1.5 checkpoint, a text-to-image latent diffusion model published on October 20
SAM 2 (Segment Anything Model 2) is a promptable visual segmentation model for both images and video developed by Meta AI and released on 29 July 2024.
Sapiens is a family of human-centric computer vision foundation models developed by Meta (Reality Labs), introduced in 2024 and presented as an oral paper at the European Conference on Computer Vision (ECCV)…
SciBERT is a BERT-based language model pretrained from scratch on a large corpus of scientific papers, built by the Allen Institute for AI (AI2).
Seedance is the family of foundation video generation models built by the Seed team at ByteDance, the Chinese internet company that owns TikTok and Douyin.
Seedance 2.0 is a multimodal video generation model developed by ByteDance, released in February 2026 as the second major version of the company's Seedance line.
Seedance 2.5 is a proprietary AI video generation model released by ByteDance on July 31, 2026. It jointly generates audio and video from combinations of text, image, video, and audio inputs.
Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.
Seedream 5.0 is a text-to-image generation model developed by ByteDance, released in February 2026 as the fifth major version of the company's Seedream line.
Segment Anything Model (SAM) is a promptable image segmentation foundation model released by Meta AI on April 5, 2023 that lets users "cut out" any object in an image with a single click, box, or mask prompt…
Sentence similarity models are machine learning systems that map natural language sentences to fixed-length numerical vectors (sentence embeddings) so that semantically related sentences sit close together in…
Sentence-BERT (SBERT) is a modification of the pretrained BERT transformer network that produces semantically meaningful, fixed-size sentence embeddings comparable with simple cosine similarity.
Sentence-transformers/all-MiniLM-L6-v2 model
sentence-transformers/all-mpnet-base-v2
Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.
Shap-E is a conditional generative model for 3D assets developed by OpenAI, introduced in the paper "Shap-E: Generating Conditional 3D Implicit Functions" by Heewoo Jun and Alex Nichol, submitted to arXiv on…
Skild AI is a robotics artificial intelligence company building a general-purpose foundation model for physical embodiments
Skywork R1V is a family of open-weight multimodal vision-language models built for chain-of-thought reasoning, developed by Skywork AI
SmolLM is a family of small, fully open language models released by Hugging Face on July 16, 2024 in three sizes, 135 million, 360 million, and 1.7 billion parameters, all trained on a curated open dataset…
SmolLM 2 is a family of compact open-weight language models released by Hugging Face on November 1, 2024.
SmolLM 3 is a fully open 3 billion parameter language model released by Hugging Face on July 8, 2025, trained on 11.2 trillion tokens and designed as a small, multilingual, long-context reasoner.
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…
Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025.
State is a machine learning "virtual cell" model developed by the Arc Institute, a nonprofit biomedical research organization, to predict how cells respond to perturbations such as drugs, cytokines, and…
Step-3 is an open-weight large multimodal mixture of experts (MoE) model released in July 2025 by StepFun, the Shanghai-based Chinese artificial intelligence startup also known as Jieyue Xingchen.
Summarization models are natural language processing systems that condense a source document, set of documents, or dialogue into a shorter version that preserves the most important information.
Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…