Data & Datasets

Explore Data & Datasets through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Large Language Models

Articles that also belong to these categories. Counts cover all of Data & Datasets.

Showing 1-11 of 11 articles

Cosmopedia

Cosmopedia is an open synthetic pretraining dataset released by Hugging Face in February 2024, made up of textbooks, blog posts, stories, and WikiHow-style articles written entirely by a large language model.

Large Language ModelsOpen Source AI

Dolma

Dolma is an open three-trillion-token English pretraining corpus released by the Allen Institute for AI (AI2) to power its fully open OLMo language models and to let researchers study how training data shapes…

Large Language ModelsOpen Source AI

RefinedWeb

RefinedWeb is a large-scale English pretraining dataset for large language models, built from filtered and deduplicated Common Crawl web data alone and released in June 2023 by the Technology Innovation…

Large Language ModelsOpen Source AI

SlimPajama

SlimPajama is a 627-billion-token English-language pre-training corpus for large language models, produced by Cerebras Systems in collaboration with the Opentensor Foundation by extensively cleaning and…

Large Language ModelsOpen Source AI

TxT360

TxT360 is an open large-scale pretraining corpus for large language models, released in October 2024 by the LLM360 project, a collaboration led by Petuum and the Mohamed bin Zayed University of Artificial…

Large Language ModelsOpen Source AI

UltraChat

UltraChat is a large-scale synthetic multi-turn instructional conversation dataset released in May 2023 by the OpenBMB group at Tsinghua University, comprising approximately 1.5 million dialogues generated by…

Chinese AILarge Language Models