1. Golden set
20 cases you never change
2. LLM as judge
a model grades against a rubric
3. Rubric scoring
one number per dimension
4. Trajectory eval
grade the path, not the answerstep 4 called the wrong tool
5. Tool unit tests
test the hands, not the brain
6. Regression suite
did the new prompt break turn 4
7. A/B in prod
real traffic, split live
8. Human review
sample it, do not read it all
9. Shadow run
the candidate runs, nobody sees itsame input, unseen output
10. Red team
try to break it before they dojailbreakexfilprompt injecttool abuse
1 got through
agent