The MTEB embedding leaderboard
Listed inEmbedding Models on Hugging FaceEmbeddingson
Benchmark scores across retrieval, clustering, and reranking tasks. Useful for a shortlist, dangerous as a final answer.
Turning text, images, and audio into vectors — the substrate for search, RAG, and classification.
14 articles
Listed inEmbedding Models on Hugging FaceEmbeddingson
Benchmark scores across retrieval, clustering, and reranking tasks. Useful for a shortlist, dangerous as a final answer.
Listed inSentence TransformersEmbeddingson
The library most self-hosted embedding pipelines are built on: pooling, training, and inference.
Listed inOpenAI Embeddings APIEmbeddingson
text-embedding-3, dimension truncation, and batching economics.
Listed inAnomaly DetectionEmbeddingson
Flagging outliers by distance from the dense region of your corpus.
Listed inEmbedding Models on Hugging FaceEmbeddingson
Using MTEB leaderboards to pick a model without over-trusting them.
Listed inRecommendation SystemsEmbeddingson
Using vector similarity to surface related items without hand-built rules.
Listed inJina EmbeddingsEmbeddingson
Long-context and multilingual open embedding models.
Listed inWhat are Embeddings?Embeddingson
Dense vectors that put semantically similar things near each other.
Listed inCohere EmbedEmbeddingson
Multilingual embeddings plus a reranker that often beats a bigger index.
Listed inEmbedding ModelsEmbeddingson
Dimensions, context limits, and why you can never mix models in one index.
Listed inData ClassificationEmbeddingson
Labelling at scale by embedding once and comparing against class centroids.
Listed inSentence TransformersEmbeddingson
The open library behind most self-hosted embedding pipelines.
Listed inSemantic SearchEmbeddingson
Retrieving by meaning instead of keywords, and where it still loses to BM25.
Listed inGemini EmbeddingEmbeddingson
Google's embedding models and the task-type hints that tune them.