IOanyT Innovations

Lesson 5 of 10 · 9 min read

Evaluation for reliability

In one paragraph

Evaluation for reliability means measuring an AI-assisted workflow against real cases with agreed correct answers, running each case more than once to check consistency, sorting the errors by kind, and repeating the whole test after every change so that regressions are caught before users see them.

In this lesson

  • Build a golden dataset for one workflow
  • Measure both accuracy and consistency, and know why you need both
  • Use error analysis and regression evaluation to make improvements stick

The video version of this lesson is in production. The full lesson is below.

Lesson 4 said every step moves up the ladder only after it has proven itself on your data. This lesson is about what “proven” means in practice. None of it is specific to IOanyT; it is the general discipline any team should apply before trusting an AI-assisted workflow.

Step 1: build a golden dataset

A golden dataset is a set of real cases from the workflow, each with an answer your own experts agree is correct.

  • Use real cases, not invented ones. Pull them from past tickets, applications, documents or calls, removing personal data where needed.
  • Include the hard ones. Ask the people who do the work which cases trip them up. Those matter more than another hundred easy ones.
  • Agree the answers. Where experts disagree, settle it before testing. If your own team can’t agree what’s correct, no system can be measured against it, and you’ve found a policy question, not a technology one.

Step 2: measure accuracy and consistency

Run the system over every case and compare its output with the agreed answers. That’s accuracy.

Then run every case again, ideally several times, and check whether the outputs match each other. That’s consistency. A system can score well on accuracy in one run and still give a different answer to the same case on the next. For any decision people rely on, you need both numbers.

A compiled (deterministic) component will show perfect consistency by construction. The interesting measurements are accuracy, and the consistency of any steps that still involve a model at run time.

Step 3: analyse the errors

A single score hides what to fix. Group the mistakes by kind, for example:

  • misread input (a date, an amount, a name)
  • missing rule (a case the logic never anticipated)
  • wrong rule (the logic disagrees with policy)
  • ambiguous case (experts would hesitate too)
  • refusal or hand-off that should, or shouldn’t, have happened

Each group has a different fix: better input checks, a new rule, a corrected rule, a hand-off to a person, or a policy decision. The biggest group is usually the best next fix.

Step 4: fix, then re-run everything

After every change (a new rule, a different model, a prompt edit, a provider update), re-run the whole golden dataset. This is regression evaluation. Fixes in one place often break behaviour somewhere else, and the only way to know is to look.

Grow the dataset over time. Every time production surfaces a new kind of mistake, add the case. The test set becomes a record of everything the system has learned not to get wrong.

Agree the bar before you test

Decide in advance what result counts as good enough, based on what an error costs in that workflow. A drafting assistant with a reviewer behind it can tolerate far more than a step that sets prices or approves applications. Setting the bar after seeing the results is how weak systems get approved.

Try it yourself: a 20-case starter set

  1. Pick one workflow and collect 20 real past cases, including at least 5 that experts call “tricky”.
  2. Have two experts record the correct answer for each, independently, then reconcile disagreements.
  3. Run your current tool or process over all 20, twice.
  4. Count: how many were right, how many changed between runs, and what kinds of mistakes appeared.

You now have a baseline, which is more than most AI pilots ever measure.

What comes next

Reliability is one side of the case for compiled AI; cost is the other. Lesson 6 works through the cost math of compiled versus runtime models.

Measure before you trust General practice for evaluating any AI-assisted workflow, repeated on every change. 1: Collect — real cases from, the workflow. 2: Label — your experts agree, the right answers. 3: Run — the system on every, case, more than once. 4: Analyse — group the errors, by kind. 5: Fix & re-run — change, then repeat, the whole set. IOANYT ACADEMY / LESSON 5 Measure before you trust General practice for evaluating any AI-assisted workflow, repeated on every change. 01 Collect real cases fromthe workflow 02 Label your experts agreethe right answers 03 Run the system on everycase, more than once 04 Analyse group the errorsby kind 05 Fix & re-run change, then repeatthe whole set the same test set runs after every change, so regressions show up before users see them

Key takeaways

  • A golden dataset is real cases with answers your own experts agree are correct.
  • Accuracy and consistency are different. A system can be right on average and still give different answers to the same case.
  • Error analysis (grouping mistakes by kind) tells you what to fix; a single score doesn't.
  • Regression evaluation re-runs the whole set after every change, because fixing one thing often breaks another.

Check yourself

Pick an answer, then open the card to compare.

  1. 1. What makes a dataset 'golden'?

    • A.It was generated by the most expensive model
    • B.It contains real cases with answers your own experts agree are correct
    • C.It is very large
    Show the answer

    B. It contains real cases with answers your own experts agree are correct Golden means trusted answers on realistic cases. Size helps, but agreement and realism come first.

  2. 2. A system scores 95% accuracy on a test run once. What important question is still unanswered?

    • A.Whether it gives the same answers when the same cases are run again
    • B.Whether the interface looks good
    • C.Whether the test was fast
    Show the answer

    A. Whether it gives the same answers when the same cases are run again Accuracy from a single run says nothing about consistency. Run each case more than once.

  3. 3. Why re-run the whole test set after every change?

    • A.To make the reports longer
    • B.Because a fix in one place can quietly break behaviour elsewhere
    • C.Because the old results expire
    Show the answer

    B. Because a fix in one place can quietly break behaviour elsewhere That is regression evaluation: it catches behaviour that got worse before users do.

Common questions

How many cases does a golden dataset need?

Enough to cover the common cases and the known hard ones. Many teams start with a few dozen to a few hundred and grow it every time production surfaces a new kind of mistake.

What score is good enough?

It depends on what an error costs in that workflow, and should be agreed before testing, not after seeing the results. A drafting tool and a credit decision deserve very different bars.

Can a model grade another model's output?

It can help triage at scale, but its judgements need checking against human-labelled cases too. Treat it as another system to evaluate, not as ground truth.