AI Accelerator
An AI accelerator is hardware designed or configured to execute artificial intelligence and machine learning workloads more efficiently than a general-purpose processor executing the same workload without…
Explore AI Inference through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Inference.
Showing 1-30 of 30 articles
An AI accelerator is hardware designed or configured to execute artificial intelligence and machine learning workloads more efficiently than a general-purpose processor executing the same workload without…
The AMD Instinct MI350P is a PCIe add-in card built on AMD's CDNA 4 architecture, introduced on May 7, 2026 as the third member of the Instinct MI350 series and the first AMD Instinct product in a conventional…
AWS Graviton is a family of Arm-based server processors designed by Amazon Web Services for use in its own cloud computing fleet.
AWS Inferentia is a family of custom application specific integrated circuits (ASICs) designed by Amazon Web Services for machine learning inference in the cloud, built to deliver, in AWS's words, "high…
The Edge TPU is a small application-specific integrated circuit (ASIC) designed by Google to run machine learning inference on low-power devices.
Sohu is a transformer-specialized application-specific integrated circuit (ASIC) built by Etched, a Silicon Valley AI hardware startup founded in 2022 by Harvard dropouts Gavin Uberti, Chris Zhu, and Robert…
FP4 (4-bit floating point) is a numerical format that stores a real number in just 4 bits, the smallest floating-point type in mainstream use for deep learning.
A field-programmable gate array (FPGA) is an integrated circuit whose logic functions and internal wiring are set by the customer after the chip has been manufactured, and can be reset later.
Google TPU 8i is an eighth-generation Tensor Processing Unit from Google, built specifically for AI inference rather than model training.
Groq hardware is a family of artificial-intelligence accelerators and multi-chip systems built around a statically scheduled streaming architecture.
InferenceX, launched in October 2025 under the name InferenceMAX, is an open-source benchmark that continuously measures large language model inference performance across AI accelerators and serving software.
Intel Crescent Island is a data-center GPU from Intel built for artificial-intelligence inference workloads.
MTIA (Meta Training and Inference Accelerator) is a family of custom silicon chips that Meta designs for use in its own data centers rather than for sale.
MediaTek Inc. is a Taiwanese fabless semiconductor company headquartered in Hsinchu, Taiwan.
NVIDIA Groq 3 LPX is a rack-scale inference accelerator that NVIDIA introduced at GTC 2026, built around 256 Groq Language Processing Units and designed to sit beside Vera Rubin NVL72 racks as a dedicated…
NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.
NVIDIA Rubin CPX is a class of GPU announced by NVIDIA on September 9, 2025, purpose-built to accelerate the compute-heavy "context" phase of large-model inference.
A neural processing unit (NPU) is a processor, or a block inside a larger chip, designed specifically to run neural network inference at high throughput and low power.
On-device AI is the practice of running machine learning models on the phone, laptop, watch, or embedded board a person is actually using, instead of sending the input to a remote data center.
Positron AI is an American semiconductor startup headquartered in Reno, Nevada, that designs and manufactures purpose-built hardware for transformer inference.
Qualcomm AI200 is a rack-scale data-center accelerator for artificial intelligence inference, announced by Qualcomm on 27 October 2025 and slated for commercial availability in 2026 .
Qualcomm AI250 is a planned data-center artificial intelligence inference accelerator and rack-scale system announced by Qualcomm in late October 2025.
REBEL-Quad is a chiplet-based AI inference accelerator developed by Rebellions, a South Korean AI-chip company.
SRAM, or static random-access memory, is a semiconductor memory that stores each bit in a latch built from cross-coupled inverters.
Snapdragon AR1 is a family of Qualcomm system-on-chip platforms built specifically for smart glasses.
Tenstorrent Galaxy Blackhole is an AI inference server built by Tenstorrent, the fabless semiconductor company led by chief executive Jim Keller.
Xiaomi XRING O100 is a dedicated AI accelerator designed by Xiaomi for local large-model inference on consumer devices.
d-Matrix is a privately held American semiconductor company headquartered in Santa Clara, California that builds accelerators, I/O cards and software for AI inference in data centers.
Corsair is the first commercial AI accelerator product from d-Matrix, a Silicon Valley AI inference hardware startup based in Santa Clara, California.
Raptor is the second-generation AI inference accelerator from d-Matrix, a Santa Clara semiconductor startup, and the first commercial chip built on the company's 3D stacked digital in-memory compute technology