CLIP Score
CLIP Score (also written CLIPScore or CLIP-S) is a reference-free automatic evaluation metric that measures how well a text caption matches an image, computed as the rescaled cosine similarity of the image and…
Explore Computer Vision through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Computer Vision.
Showing 1-13 of 13 articles
CLIP Score (also written CLIPScore or CLIP-S) is a reference-free automatic evaluation metric that measures how well a text caption matches an image, computed as the rescaled cosine similarity of the image and…
ControlNet is a neural network architecture that adds spatial and structural control to large pretrained text-to-image diffusion models.
CycleGAN (Cycle-Consistent Generative Adversarial Network) is a deep learning architecture for unpaired image-to-image translation.
The Frechet Inception Distance (FID) is the standard metric for measuring the quality of images produced by generative models: it computes the Frechet distance between two multivariate Gaussian distributions…
GenEval is an object-focused benchmark for evaluating how well text-to-image models follow the content of a prompt.
Grok Imagine is a generative media product from xAI, the company founded by Elon Musk.
Ideogram 3.0 is a text-to-image generation model released by Ideogram on March 26, 2025.
Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.
Photoroom is an AI-powered photo editing platform headquartered in Paris, France, specializing in background removal, product photography, and generative image editing.
Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.
StyleGAN is a family of style-based generative adversarial network (GAN) architectures developed by NVIDIA Research for high-quality unconditional image synthesis
Topaz Labs is a software company headquartered in Dallas, Texas, that develops AI-powered tools for photo and video enhancement, best known for Gigapixel, Photo AI, Video AI, and its Bloom and Project…
pix2pix is a supervised image-to-image translation method that learns a mapping between two visual domains from aligned input-output image pairs.