Test and quality-assure AI agents

Move from “it seems to work” to a repeatable quality decision. You define measurable quality, representative cases, blocking errors, human and automated graders, regression checks and a documented release decision.
The course has nine parts. You decide what quality means for the task, risk-classify failures, build a representative reference dataset covering normal, edge, missing, conflicting and untrusted cases, and write a rubric with six criteria.
You combine deterministic checks with human review, test difficult cases and stopping behaviour, compare versions to find regressions, and run a test lab with five cases plus a separate decision. The final test has eight questions and seven correct answers give you the certificate.
What you will build
AI agent
You will learn to: Test, measure and quality-assure an AI agent before live use
What you will be able to do
- ✓Build a test suite with representative cases
- ✓Measure quality in a comparable way
- ✓Set limits for when a human must review
- ✓Follow up on quality over time
Course content
- 1Quality is a decided measure1 lessons
- 2Risk-class failures and scope the tests1 lessons
- 3Build a representative reference dataset1 lessons
- 4Write a rubric1 lessons
- 5Combine automatic and human graders1 lessons
- 6Test hard cases and stopping behaviour1 lessons
- 7Regression and version comparison1 lessons
- 8Run the test lab and decide1 lessons
- 9Final assignment and knowledge test1 lessons
1 495 kr
Planned price
Not available for purchase yet
Try the free mini course firstSwedish version of this courseSource-checked content
The course is reviewed against official documentation and established standards. The organisations below are examples of the material the content builds on.
- OpenAI
- Anthropic
- NIST
- NIST
- OpenAI
- Last checked:
- 19 September 2026
- Next scheduled review:
- 19 December 2026