Session Outline
A good evaluation system can help you catch regressions, report and set goals, and create your roadmap for your AI agent. This session delves into practical strategies for tackling these challenges, from using scalable LLM judges to leveraging conversation-level metrics for deeper insights.
Key Takeaways
- Building a representative metric
- Building a representative evaluation dataset
- Building an evaluator (human, LLM judge, or a code-based judge)