Speaker: Tony Zhang – Senior Manager, GenAI Senior Data Scientist, Manulife
Description
The session will explore the use of synthetic data and LLM-based techniques to access the performance. I will discuss how synthetic data, generated by LLMs, can emulate real-world interactions to enable comprehensive and rigorous model evaluation.
Key Takeaways:
– Synthetic data generation: Techniques for producing synthetic data
– LLM-as-a-Judge framework: How models can score outputs on various dimensions.
– Human evaluation in the loop: Incorporate human evaluation for deeper insights
– Metrics in RAG evaluation
– Potential challenges and risks