Chinese AI

Explore Chinese AI through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Multimodal AI

Articles that also belong to these categories. Counts cover all of Chinese AI.

Showing 1-27 of 27 articles

Baidu ERNIE

Baidu ERNIE (Enhanced Representation through Knowledge Integration) is the family of large language and multimodal foundation models built by the Chinese technology company Baidu, spanning the original 2019…

Large Language ModelsMultimodal AI

CogAgent

CogAgent is an open visual language model built to act as a graphical user interface (GUI) agent: given a screenshot and a natural-language goal, it predicts the next on-screen action, such as where to click…

AI AgentsMultimodal AI

DeepSeek Janus

DeepSeek Janus is a family of open-weight unified multimodal models from Chinese AI lab DeepSeek that perform both image understanding and text-to-image generation in a single autoregressive Transformer.

AI ModelsMultimodal AI

DeepSeek-OCR

DeepSeek-OCR is an open-source optical character recognition (OCR) and document-understanding system released by DeepSeek on 20 October 2025 that pioneers a contexts optical compression paradigm: it encodes…

Computer VisionMultimodal AI

Doubao Seed 1.6

Doubao Seed 1.6 is a family of general-purpose foundation models developed by the ByteDance Seed research team and released through Volcano Engine on 11 June 2025 at the company's Force Original Power…

AI ModelsLarge Language Models

ERNIE 5.0

ERNIE 5.0 is a natively omni-modal foundation model from Baidu, unveiled at the company's annual Baidu World 2025 conference in Beijing on 13 November 2025 as the flagship in the ERNIE line at launch.

Large Language ModelsMultimodal AI

HY-World 2.0

HY-World 2.0 (also written HunyuanWorld 2.0 or Hunyuan World Model 2.0) is an open multimodal 3D world model from Tencent's Hunyuan team, with a technical report and first code release published on April 16

Multimodal AIWorld Models

InternVL

InternVL is a family of open-source multimodal large language models developed by the OpenGVLab research group at the Shanghai Artificial Intelligence Laboratory in collaboration with academic partners…

Multimodal AIOpen Source AI

MiniCPM-V

MiniCPM-V is a family of open-weights multimodal large language models published by OpenBMB, the shared open-source brand of Tsinghua University's Natural Language Processing lab (THUNLP) and the Beijing…

Multimodal AIOpen Source AI

Qwen-VL

Qwen-VL is the first family of open vision-language (multimodal) models from the Qwen team at Alibaba Cloud, able to take images, text, and bounding boxes as input and produce text and bounding boxes as output.

Multimodal AIOpen Source AI

Qwen2-VL

Qwen2-VL is a family of open-weight vision-language models released by the Qwen team at Alibaba Cloud between August and September 2024, in 2B, 7B, and 72B Instruct sizes.

Multimodal AIOpen Source AI

Qwen2.5-VL

Qwen2.5-VL is a series of open-weight vision-language models released on 26 January 2025 by the Qwen team at Alibaba (Alibaba Cloud), succeeding the earlier Qwen2-VL family.

Multimodal AIOpen Source AI

Qwen3-Omni

Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…

Large Language ModelsMultimodal AI

Skywork-R1V

Skywork-R1V is an open-weight family of multimodal reasoning models released by Skywork AI, the AGI and AIGC division of Beijing Kunlun Tech Co., Ltd. (Kunlun Wanwei).

Multimodal AIReasoning Models

UI-TARS

UI-TARS is a native graphical user interface (GUI) agent model developed by ByteDance through its Seed research team.

AI AgentsMultimodal AI