This course contains the use of artificial intelligence.
“How do you know your AI agent works?”
Not “has anyone complained” — how do you know. Most teams answer that question with a demo. A demo is not evidence, and the moment an auditor, a regulator or an executive asks for evidence, the gap becomes obvious.
This course closes that gap. You will not learn how to build an agent. You will learn how to prove one works, catch it failing, and sign off a release you can defend — using nine artifacts you build as you go, none of which require writing a line of code.
Why agents break every testing habit you have
A test suite assumes the same input gives the same output. An agent samples from a distribution. It assumes a finite set of paths — an agent with six tools over ten steps has a path space you cannot enumerate. It assumes failures announce themselves. An agent that invents a policy clause returns a beautifully formatted, confident, completely wrong answer with an HTTP 200. Nothing goes red.
Governance says what should be true. Assurance proves what is true. This course is about the second one.
What makes this different
-
Hallucination is not one thing. Six distinct mechanisms — fabricated fact, invented citation, tool-result confabulation, ungrounded retrieval, false capability claim, grounding decay — each with its own detection signal and its own kind of test. Three are demonstrated live, so you watch the failure happen rather than reading about it.
-
LLM-as-judge has measured biases. Position bias swings win rates by 10–15 points on slot order alone. Verbosity bias inflates by 15–30 points — and verbosity correlates with hedging, which means your judge can actively reward the hallucination you are hunting.
-
Telemetry cannot tell you whether an answer was right. The OpenTelemetry GenAI conventions deliberately stop short of output evaluation. That is a category boundary, not a tooling gap — and knowing where observability stops is what tells you what else to build.
The AI Agent Assurance Pack
You finish with nine artifacts built for an agent you actually care about: an agent risk assessment, a behavior and boundary specification, a test-case and evaluation library, a hallucination taxonomy, a human-oversight matrix, a monitoring dashboard design, an agent incident log, a release-readiness checklist, and a governance and assurance report.
Every one is worked end to end against Halden Insurance Group, a mid-market insurer running three agents with no evidence for any of them — including a production incident where every answer was grounded, cited, traceable, and wrong for six weeks.
Anchored to the frameworks you will be assessed against
NIST AI RMF (the Measure function), ISO/IEC 42001 Clause 9, EU AI Act Article 14 on human oversight, and the OWASP Top 10 for LLM Applications — not as theory, but as the reason each artifact takes the shape it does.
Who this is for
AI governance professionals, QA leads, risk managers, internal auditors and AI product owners. No coding required. Engineers are welcome and will find the vocabulary useful for defending their own evaluation work.








