Have a request for an upcoming news/science story? Submit a Request

[RCAC Workshop]How Do We Know an LLM Is Good? Practical Evaluation of Large Language Models

📅 Date: October 16th, 2026 ⏰ Time: 11AM-12PM 💻 Location: Virtual 🏫 Instructor: Mihir Ahlawat

Large language models can produce responses that appear useful, but evaluating their quality in a consistent and repeatable way can be challenging. This session will introduce practical approaches for evaluating LLMs and LLM-powered applications, including common metrics, task-specific evaluation, rubric-based scoring, and LLM-as-a-judge approaches. The session will also discuss how evaluations can help compare models, prompts, and application changes over time.

Who should attend

Students, researchers, and developers who use or build applications with large language models and want to better understand how to measure their performance and reliability.

What you'll learn

Participants will learn the fundamentals of LLM evaluation, common evaluation approaches and metrics, and how evaluations can be used to systematically compare and improve LLM-based systems.

Level

Introductory to intermediate. Basic familiarity with large language models is helpful but not required. The session will primarily be lecture-based, with examples and demonstrations as appropriate.

🔗 Register now: LINK

Originally posted: