Labsco
github logo

phoenix-evals

โœ“ Officialโ˜… 36,202

by github ยท part of github/awesome-copilot

Build and run evaluators for AI/LLM applications using Phoenix.

๐Ÿ”ฅ๐Ÿ”ฅFreeQuick setup
๐Ÿงฉ One of 7 skills in the github/awesome-copilot package โ€” works on its own, and pairs well with its siblings.

This is the playbook your agent receives when the skill activates โ€” you don't need to read it to use the skill, but it's here to audit before installing.

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

Workflows

Starting Fresh: observe-tracing-setup โ†’ error-analysis โ†’ axial-coding โ†’ evaluators-overview

Building Evaluator: fundamentals โ†’ common-mistakes-python โ†’ evaluators-{code|llm}-{python|typescript} โ†’ validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag โ†’ evaluators-code-* (retrieval) โ†’ evaluators-llm-* (faithfulness)

Production: production-overview โ†’ production-guardrails โ†’ production-continuous

Reference Categories

PrefixDescription
fundamentals-*Types, scores, anti-patterns
observe-*Tracing, sampling
error-analysis-*Finding failures
axial-coding-*Categorizing failures
evaluators-*Code, LLM, RAG evaluators
experiments-*Datasets, running experiments
validation-*Validating evaluator accuracy against human labels
production-*CI/CD, monitoring

Key Principles

PrincipleAction
Error analysis firstCan't automate what you haven't observed
Custom > genericBuild from your failures
Code firstDeterministic before LLM
Validate judges>80% TPR/TNR
Binary > LikertPass/fail, not 1-5