What's new
Warez.Ge

This is a sample guest message. Register a free account today to become a member! Once signed in, you'll be able to participate on this site by adding your own topics and posts, as well as connect with other members through your own private inbox!

Testing LLMs With DeepEval, Promptfoo, RAG & CI/CD

voska89

Moderator
Staff member
Top Poster Of Month
3584f582e09d00b266b4b258f9b4a56d.webp

Testing LLMs With DeepEval, Promptfoo, RAG & CI/CD
Published 9/2026
Created by Madhulika Mitra
MP4 | Video: h264, 1920x1080 | Audio: AAC, 44.1 KHz, 2 Ch
Level: Intermediate | Genre: eLearning | Language: English | Duration: 23 Lectures ( 3h 59m ) | Size: 2 GB​

Evaluate RAG apps, red-team for security, and automate quality gates in CI/CD - become job-ready in AI QA
What you'll learn

⚡ Explain why unit, integration, and E2E tests are not enough for generative AI
⚡ Test RAG end to end - retrieval quality and grounded generation along with agentic multi flow testing
⚡ Deepeval evaluation metrics , Promptfoo security testing
⚡ Put evaluations in CI/CD with thresholds, quality gates, and cost control
Requirements

❗ Python programming expertise and Testing fundamentals
Description

AI/LLM Testing Mastery: From DeepEval to Production CI/CD
Traditional tests can be green while your chatbot is still wrong, biased, or jailbroken. This course teaches you how to test LLM applications the way production teams actually have to: with metrics, judges, adversarial checks, and quality gates.
You will learn to evaluate correctness, relevance, faithfulness, hallucination, toxicity, bias, and security. You will test RAG pipelines and AI agents, run prompt-injection and jailbreak cases, and wire the whole suite into GitHub Actions so a bad answer can block a deploy.
We use DeepEval and Promptfoo, local models with Ollama (no API key required to learn), and finish with a portfolio-ready capstone: a support chatbot, an evaluation suite, and a CI pipeline.
You will learn how to
✨ Explain why unit, integration, and E2E tests are not enough for generative AI
✨ Score free-form answers with LLM-as-judge instead of exact string matches
✨ Test RAG end to end - retrieval quality and grounded generation
✨ Test tool selection, multi-step agents, and memory
✨ Red-team chatbots for injection, jailbreaks, and data leaks
✨ Put evaluations in CI/CD with thresholds, quality gates, and cost control
✨ Debug a failing eval: was the chatbot wrong, or was the judge wrong?
Includes hands-on labs, homework after each module, and a GitHub project.
Who this course is for

⭐ QA engineers, SDETs, and testers moving into AI quality roles
Homepage
Code:
https://www.udemy.com/course/testing-llms-with-deepeval-promptfoo-rag-cicd

Recommend Download Link Hight Speed | Please Say Thanks Keep Topic Live ​
No Password - Links are Interchangeable​
 

Users who are viewing this thread

Back
Top