PUBLIC LESSON PREVIEW / Reliable AI Systems
Design evaluations that can block a release
Read this sample without an account. Sign in for the full workspace, saved progress and checkpoints. This preview does not save activity or award credit.
Bridge from working code to evidence
A demonstration proves that a system can produce one convincing result. A release decision asks a different question: under stated conditions, how often does the complete workflow satisfy its promises, and what happens when it does not? Begin with a narrow fictional product: a support assistant that either answers a routine question or routes an account-access problem to a human. It never changes accounts. Define the allowed actions, the information available at decision time, and the consequence of a wrong route before choosing a score. You need Python lists, dictionaries, loops, functions, and Boolean comparisons for this lesson. As a prerequisite bridge, write a function that returns whether two strings match and call it for three pairs. Explain why a comparison returns a Boolean rather than a percentage. If that feels unfamiliar, practice those operations first. The course uses tiny deterministic stand-ins for model responses so that evaluation mechanics remain visible and cost nothing.
Build a test contract and a useful dataset
Represent each case with an identifier, input, expected behavior, and a meaningful slice. A slice groups cases whose failure pattern deserves separate attention, such as account-access requests. Do not include private customer records in this exercise. Write synthetic cases that resemble plausible requests, then document what the synthetic collection fails to represent. Clear English examples do not establish multilingual performance, and four examples cannot estimate rare failures precisely. Use one development set to improve the implementation and a separate held-out set for the release decision. If you repeatedly inspect held-out answers while adjusting the system, that set becomes development evidence. Replace it rather than claiming independence. Add incident-derived cases to a regression suite, but keep their origin visible. A balanced learning set and a traffic-weighted production set answer different questions. Record both sampling choices and label disagreements; an uncertain reference label should trigger review rather than quietly count as model failure.
Choose gates before looking at the result
The worked example scores exact routing decisions, which makes an ordinary equality check appropriate. Free-form answers require more carefully defined criteria, such as whether required facts are supported and whether forbidden actions were proposed. A judge, whether human or automated, is another measurement component that needs calibration. Test it against examples with known defects before trusting its score. Our toy gate requires overall accuracy of at least three quarters and zero missed account-access escalations. The candidate reaches the overall threshold yet fails the critical slice. This is deliberate: an average must not erase a consequential defect. In a real review, also report numerator, denominator, uncertainty, and the cost of each error type. A threshold is a policy choice, not a universal constant. Set it with domain owners and user expectations. Repeated stochastic runs can reveal instability, but repeated near-duplicate cases do not provide independent coverage.
Run, debug, then extend independently
Save the example as lesson_01.py and run python3 lesson_01.py from its folder. Compare every printed line with the expected output before editing anything. If accuracy differs, inspect the case count and the normalization rule. If the critical count is zero unexpectedly, print the slice and predicted decision for each case; a misspelled label can make the wrong subset look perfect. Keep identifiers in diagnostic output so failures can be traced back without printing sensitive inputs. For guided practice, change the account-access prediction to escalate and predict both metrics before running again. Then independently design two additional slices and one deliberately ambiguous case. Write a one-page evaluation contract that names the owner of each criterion, the held-out boundary, and the evidence needed to update labels. Your mastery gate is explaining a blocked release using concrete failing cases, rather than arguing that the average looks impressive. Preserve the original failure as a regression example.
A good average with a failed critical slice
This fixture evaluates pre-recorded predictions. It does not estimate a real model’s quality or call a provider. The intentionally failing gate demonstrates why aggregate accuracy alone is inadequate.
cases = [
("routine", "answer", "answer"),
("routine", "answer", "answer"),
("access", "escalate", "answer"),
("access", "escalate", "escalate"),
]
correct = sum(expected == actual for _, expected, actual in cases)
accuracy = correct / len(cases)
missed = sum(group == "access" and actual != expected for group, expected, actual in cases)
release = accuracy >= 0.75 and missed == 0
print(f"accuracy: {accuracy:.2f}")
print(f"access misses: {missed}")
print(f"release: {release}")
Ready to try it yourself?
ChatGPT sign-in takes you to OpenAI and back to CodeTrail. It keeps your learning account separate from other learners.
Open the full lesson ↗