Lesson 1 of 10 · 8 min read
Why "probably right" isn't good enough
In one paragraph
Run-to-run variance is the difference between a system's outputs when exactly the same input is processed more than once. Language models produce it by design, which is harmless when a person reviews the output and risky when a customer, an auditor or a decision relies on the answer being the same every time.
In this lesson
- Explain what run-to-run variance is and why language models produce it
- Tell the difference between variance that is harmless and variance that is business risk
- Run a simple consistency test on an AI tool your team already uses
The video version of this lesson is in production. The full lesson is below.
Ask a good AI assistant a question and you get a fluent, confident answer. Ask it again in a fresh chat and you may get a different one: different wording, sometimes different facts. Most of the time nobody notices, because most of the time nobody asks twice.
This lesson is about the moment that stops being harmless.
What run-to-run variance is
A language model writes its answer one word at a time, choosing each word from a list of likely candidates. Some randomness in that choice is deliberate: it is what makes the output read naturally instead of robotically. The side effect is that the same input can produce different outputs on different runs. That is run-to-run variance.
Even when the randomness is turned down to its minimum (a setting called temperature 0), production systems can still drift. Requests are batched together on shared hardware, providers update their models, and the context sent with each question shifts slightly. The deterministic AI explainer covers the mechanics. For this lesson, the point is simple: a language model is not guaranteed to give the same answer twice.
When variance is harmless, and when it is risk
Variance is not a defect in itself. What matters is what happens to the answer next.
| How the answer is used | Does variance matter? |
|---|---|
| Drafting, brainstorming, summarising for yourself | Rarely: variety can even help |
| Internal work a colleague reviews before it goes anywhere | A little: the reviewer catches it |
| An answer a customer reads and relies on | Yes: two customers can be told two different things |
| A price, a refund, an eligibility or credit decision | Critically: money and fairness depend on consistency |
The line is crossed when someone acts on the answer without a person checking it first. Below that line, a varying answer is a style issue. Above it, it is a business decision being made differently each time.
You own what your AI says
In February 2024, a Canadian tribunal decided Moffatt v. Air Canada. The airline’s website chatbot had told a customer he could apply for a bereavement fare discount after travelling. That wasn’t the airline’s actual policy. Air Canada argued, in effect, that the chatbot was a separate legal entity responsible for its own actions.
The tribunal rejected that. It held that the chatbot was part of Air Canada’s website and that “it makes no difference whether the information comes from a static page or a chatbot”. The airline was ordered to pay the customer CA$812.02 in total.
The amount was small. The principle wasn’t: a business is responsible for what its AI tells customers. If two customers can get two different answers, the business owns both.
The audit problem
There is a second, quieter risk. When a customer, a regulator or your own leadership asks “why did the system decide this?”, the honest answer has to come from running the same case again and seeing the same result.
If the decision came from a fresh model generation, running it again may produce a different outcome. Then nobody can show how the original decision was made. For anything regulated (credit, insurance, employment, eligibility), that is a serious gap.
This is part of why the market is cooling on unconstrained AI. Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls.
Try it yourself: a ten-minute consistency test
You don’t need engineers for this. Pick one question your team really asks an AI tool, ideally one where the answer matters: a policy question, a pricing rule, a summary of a contract clause.
- Open a fresh chat and ask the question exactly as written. Copy the answer into a spreadsheet.
- Repeat 10 to 20 times, each time in a new chat, with the identical wording.
- For each answer, mark the parts a customer or decision would rely on: numbers, yes/no conclusions, deadlines, eligibility.
- Count how many distinct versions of those important parts you got.
If the important parts are identical every time, that workflow is consistent today (re-test after any tool or model change). If they differ even once, you’ve found a place where “probably right” is being treated as right.
What comes next
The fix isn’t to stop using AI. It is to move the decision itself out of the model’s hands while keeping the model where it helps. The next lesson defines deterministic AI, the term for doing exactly that.
Key takeaways
- Language models generate each answer fresh, so the same question can get different answers.
- Variance is fine when a person reviews the output; it becomes risk once someone acts on the answer.
- A business owns what its AI tells customers. In Moffatt v. Air Canada, the tribunal rejected the argument that a chatbot was a separate legal entity.
- If you can't reproduce a decision, you can't explain it to a customer or an auditor.
Check yourself
Pick an answer, then open the card to compare.
-
1. A team uses an AI tool to draft first versions of marketing emails, which an editor rewrites before sending. How worried should they be about run-to-run variance?
- A.Very worried: every variation is a compliance risk
- B.Not very: a person reviews and rewrites every output
- C.It depends only on the temperature setting
Show the answer
B. Not very: a person reviews and rewrites every output Variance only becomes risk when someone relies on the raw output. Here an editor sits between the model and the customer.
-
2. Which of these is the strongest sign that variance has become business risk?
- A.The tool is used by many employees
- B.The answers are long
- C.Customers or decisions act on the answer without review
Show the answer
C. Customers or decisions act on the answer without review Risk appears when an answer turns into a refund, a price, an eligibility decision or advice that someone acts on.
-
3. Why does reproducibility matter for audits?
- A.Auditors prefer shorter answers
- B.If a decision can't be reproduced, nobody can show how it was made
- C.Reproducible systems never contain errors
Show the answer
B. If a decision can't be reproduced, nobody can show how it was made An auditor asks why a decision was made. If running the same case again gives a different result, there is no reliable answer.
Common questions
Doesn't setting temperature to 0 fix this?
It reduces variation but doesn't guarantee identical answers in production. Batching on shared hardware, provider-side model updates and small changes in context can all shift the output. Lesson 2 covers why.
Is variance always bad?
No. For brainstorming, drafting and summarising, variety is often useful. The problem is variance in places where people rely on the answer being the same.
How many times should I run the same question to test consistency?
Ten to twenty runs in fresh sessions is enough to see whether answers drift. If the important parts of the answer differ even once, treat the workflow as inconsistent.