A retrieval-augmented generation (RAG) assistant answers from your own content. Whether it's ready depends on two things: whether retrieval finds the right passages, and whether the answer stays true to them.
Retrieval
Context recall asks whether retrieval found the passages that hold the answer. Context precision asks whether those passages were ranked above irrelevant ones. Low recall means answers are missing information; low precision means the model is distracted by noise.
Generation
Faithfulness, also called groundedness, is the share of an answer's claims supported by the retrieved context. Response relevancy is whether the answer addresses the question asked. Evaluation tools such as Ragas compute these against a test set.
The test set
The measures are only as good as the questions. A few hundred questions written by subject-matter experts, with reference answers, covers the common cases and the awkward ones. Add adversarial prompts to test prompt injection, which tops the OWASP Top 10 for LLM Applications.
After launch
Keep measuring. User ratings, corrections and hand-offs to people show where answers fall short, and each fix goes back into the test set before the next release.
