Advertisements

AI Agent Assurance: Evals, Hallucinations & Monitoring

Advertisements
Prove your AI agents work — evals, hallucination taxonomies, oversight & observability
1
1/5
(41) Ratings
7 students
Created by Dr. Amar Massoud
Advertisements

What you'll learn

  • Write behavior and boundary statements that are observable, bounded and falsifiable
  • Rank what to test hardest using consequence tiers and structural likelihood signals
  • Classify any agent failure with a six-mechanism hallucination taxonomy
  • Design golden-set, rubric, trajectory and adversarial evaluations — and know where each misleads
  • Instrument LLM-as-judge for position, verbosity and self-preference bias
  • Build a test-case library with traceable coverage and defensible pass/warn/fail thresholds
  • Design human review as a real control under EU AI Act Article 14
  • Specify monitoring signals, cost guards and drift detection for a production agent
  • Triage an agent incident to the layer that actually broke
  • Make and document a defensible go/no-go release decision, and write the assurance report
This course includes:
5.5 total hours on-demand video
0 articles
0 downloadable resources
55 lessons
Full lifetime access
Access on mobile and TV
Certificate of completion
Advertisements

Course content

Requirements

  • Working familiarity with what an AI agent is — tools, autonomy, multi-step tasks
  • No coding required. Every artifact is a document, matrix or checklist a non-engineer can own
  • A spreadsheet application (Excel, Google Sheets or equivalent) for the templates

Description

This course contains the use of artificial intelligence.

“How do you know your AI agent works?”

Not “has anyone complained” — how do you know. Most teams answer that question with a demo. A demo is not evidence, and the moment an auditor, a regulator or an executive asks for evidence, the gap becomes obvious.

This course closes that gap. You will not learn how to build an agent. You will learn how to prove one works, catch it failing, and sign off a release you can defend — using nine artifacts you build as you go, none of which require writing a line of code.

Why agents break every testing habit you have

A test suite assumes the same input gives the same output. An agent samples from a distribution. It assumes a finite set of paths — an agent with six tools over ten steps has a path space you cannot enumerate. It assumes failures announce themselves. An agent that invents a policy clause returns a beautifully formatted, confident, completely wrong answer with an HTTP 200. Nothing goes red.

Governance says what should be true. Assurance proves what is true. This course is about the second one.

What makes this different

  • Hallucination is not one thing. Six distinct mechanisms — fabricated fact, invented citation, tool-result confabulation, ungrounded retrieval, false capability claim, grounding decay — each with its own detection signal and its own kind of test. Three are demonstrated live, so you watch the failure happen rather than reading about it.

  • LLM-as-judge has measured biases. Position bias swings win rates by 10–15 points on slot order alone. Verbosity bias inflates by 15–30 points — and verbosity correlates with hedging, which means your judge can actively reward the hallucination you are hunting.

  • Telemetry cannot tell you whether an answer was right. The OpenTelemetry GenAI conventions deliberately stop short of output evaluation. That is a category boundary, not a tooling gap — and knowing where observability stops is what tells you what else to build.

The AI Agent Assurance Pack

You finish with nine artifacts built for an agent you actually care about: an agent risk assessment, a behavior and boundary specification, a test-case and evaluation library, a hallucination taxonomy, a human-oversight matrix, a monitoring dashboard design, an agent incident log, a release-readiness checklist, and a governance and assurance report.

Every one is worked end to end against Halden Insurance Group, a mid-market insurer running three agents with no evidence for any of them — including a production incident where every answer was grounded, cited, traceable, and wrong for six weeks.

Anchored to the frameworks you will be assessed against

NIST AI RMF (the Measure function), ISO/IEC 42001 Clause 9, EU AI Act Article 14 on human oversight, and the OWASP Top 10 for LLM Applications — not as theory, but as the reason each artifact takes the shape it does.

Who this is for

AI governance professionals, QA leads, risk managers, internal auditors and AI product owners. No coding required. Engineers are welcome and will find the vocabulary useful for defending their own evaluation work.

Who this course is for:

  • AI governance and compliance professionals who must evidence that agents work
  • QA leads and test managers whose suites stop working on non-deterministic systems
  • Risk managers and internal auditors assessing AI against ISO/IEC 42001 or the EU AI Act
  • AI product owners accountable for a release decision they cannot currently justify
  • Engineers who want the governance vocabulary to defend their own evaluation work
Advertisements
F4606698B0EE4E36921C
Advertisements
Advertisements
Free Online Courses with Certificates
Logo
Register New Account