Evaluation

Knowing whether a change helped — the only thing separating engineering from vibes.

10 articles

Ragas documentation

Listed inRagasEvaluationon

Faithfulness, answer relevance, and context precision — how the metrics are computed and what they miss.

External
Ragas · docs.ragas.io
#evaluation
#tooling

OpenAI Evals

Listed inLLM EvaluationsEvaluationon

The open-source eval framework and registry — a good template for structuring your own graded test set.

External
GitHub · github.com
#evaluation
#tooling

Ragas

Listed inRagasEvaluationon

A RAG-specific eval framework for faithfulness and context precision.

Intermediate5 minDraft
#tooling

Human Evals

Listed inHuman EvalsEvaluationon

Rubrics, annotator agreement, and when humans are the only ground truth.

Intermediate6 minDraft
#methods