# Evaluations and Reliability MLOps Lab

> In this lab, Zach walks through the ML Ops for LangChain agents, focusing on the new Tests section and how to run them in CI with nightly evals and PR critical runs. He shows how PyTest filters critical tests with marks, and warn that eval tests call OpenAI and can burn tokens fast, so they are gat…

- Web page: https://www.dataexpert.io/lesson/endtoendaiapplicationsday1lab-mar26-p737-l2154
- Program: [AIExpert](https://www.dataexpert.io/program/ai-expert)
- Module: Week 5: End-to-End AI Applications
- Access: Requires enrollment in AIExpert
- Length: 58 min video
- Skills: Python, MLOps, Git, APIs
- Academy: DataExpert.io Academy

## About this lesson

In this lab, Zach walks through the ML Ops for LangChain agents, focusing on the new Tests section and how to run them in CI with nightly evals and PR critical runs. He shows how PyTest filters critical tests with marks, and warn that eval tests call OpenAI and can burn tokens fast, so they are gated. He explains how seeded cases and judge rubrics work, and highlights that judge quality can miss real artifact quality unless we pass the DAG properly. Finally, he demonstrates DSPy based autoprompt optimization.

## Navigation

- Previous lesson: [Evaluations and Reliability MLOps Lecture](https://www.dataexpert.io/lesson/endtoendaiapplicationsday1lecture-mar26-p737-l2153.md)
- Next lesson: [How to deploy AI Agents using Claude managed Agents Lecture](https://www.dataexpert.io/lesson/endtoendaiapplicationsday2lecture-mar26-p737-l2155.md)
