BGE (BAAI General Embedding)
BGE (BAAI General Embedding) is a family of open-source text embedding and reranking models from the Beijing Academy of Artificial Intelligence (BAAI), first released in August 2023 and distributed through the…
Explore Information Retrieval through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Information Retrieval.
Showing 1-12 of 12 articles
BGE (BAAI General Embedding) is a family of open-source text embedding and reranking models from the Beijing Academy of Artificial Intelligence (BAAI), first released in August 2023 and distributed through the…
EmbeddingGemma is an open text embedding model from Google, released in September 2025, that turns text into dense numeric vectors for search, retrieval, classification, and clustering.
FAISS (Facebook AI Similarity Search) is an open-source library from Meta for efficient similarity search and clustering of dense vectors
GraphRAG is a graph-based approach to retrieval-augmented generation developed by Microsoft Research, first described publicly on February 13, 2024 and formalized in the paper "From Local to Global: A Graph…
Haystack is an open-source AI orchestration framework developed by deepset, a Berlin-based company, for building production-ready natural language processing (NLP), retrieval-augmented generation (RAG), and AI…
Jina Embeddings v3 is a multilingual text embedding model released by Jina AI on September 18, 2024, with 570 million parameters, support for 89 languages, an 8,192 token context window, and a stack of…
LanceDB is an open-source, developer-friendly vector database and multimodal lakehouse built on the Lance columnar storage format, designed to store vector embeddings, images, video, audio, and structured…
LlamaIndex is an open-source data framework for building large language model (LLM) applications, with a particular focus on retrieval-augmented generation (RAG) and document processing.
MTEB, short for Massive Text Embedding Benchmark, is the standard public leaderboard for evaluating text embedding models across many task types at once.
Qwen3 Embedding is a family of open text embedding and reranking models released by Alibaba's Qwen team in June 2025.
Vespa is an open-source big-data serving engine that combines vector search, lexical search, and structured search inside a single query, with real-time indexing and machine-learned ranking executed on the…
WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.