Large Language Models

Explore language models, how they work, and the techniques used to build applications with them.

Explore articles

Reset filters
Browse subtopics: AI Hardware

Articles that also belong to these categories. Counts cover all of Large Language Models.

Showing 1-4 of 4 articles

ChipNeMo

ChipNeMo is a research project and a family of domain-adapted large language models developed by Nvidia to assist with industrial semiconductor and chip-design tasks.

AI HardwareNVIDIA

Gemini Nano

Gemini Nano is the smallest and most efficient variant of Google's Gemini family of multimodal large language models, designed to run directly on phones and other edge hardware instead of in cloud data centers.

AI HardwareGoogle

Huawei AI

Huawei's AI strategy is to build a fully self-sufficient, vertically integrated AI stack inside China, spanning custom Ascend accelerators, the CANN software layer, the MindSpore framework, the Pangu family of…

AI CompaniesAI Hardware

Tokens per second

Tokens per second (TPS) is a key performance metric for measuring the speed of large language model (LLM) inference. It quantifies how many tokens a model can generate or process in one second.

AI BenchmarksAI Hardware