IOanyT Innovations

Lesson 3 of 10 · 8 min read

Compiled AI: LLMs at build time, code at run time

In one paragraph

Compiled AI is a way of building software in which large language models write, test and harden the logic at build time, and deterministic code runs it in production. The model does the engineering; reviewed, versioned code does the executing, so the same input produces the same output every time.

In this lesson

  • Describe the build-time and run-time halves of compiled AI
  • Walk through how a policy becomes tested, versioned code
  • Recognise which workflows are good candidates for compiling

The video version of this lesson is in production. The full lesson is below.

Lesson 2 ended with a puzzle: if the model isn’t making decisions at run time, where does the AI go? The answer has a name: compiled AI.

Two halves, two jobs

The word “compiled” comes from programming. A compiler turns code a person writes into a form a computer runs, once, before the program ships. Compiled AI borrows the idea: the expensive, flexible thinking happens once, before release, and what runs afterwards is fast, fixed and checkable.

At build time, the model does the engineering. It reads the material a human expert would: written policies, past cases, edge cases, regulations. It drafts the logic as code and rules, and it drafts the tests that check that logic. Engineers review every change, and the tests must pass before anything ships.

At run time, the code does the executing. When a customer’s case arrives, the system runs the reviewed logic. No model improvises a new answer. Ask the same thing a thousand times and you get the same result a thousand times, and every result traces back to a specific version of the code.

Walking through an example

Take a common decision: whether a customer qualifies for a refund. Today it might be handled by staff reading a policy document, or by a chatbot answering from that document. The compiled approach looks like this:

  1. Read. A model reads the refund policy, a sample of past decisions and the known exceptions.
  2. Draft. It writes the eligibility logic as code, for example time limits, product categories and proof requirements, plus a set of test cases, including tricky ones a human expert flagged.
  3. Check. The tests run. Where the drafted logic disagrees with how experts decided past cases, a person resolves it: either the logic is wrong, or the past decision was.
  4. Review and release. An engineer reviews the code, the tests pass, and version 1 ships.
  5. Run. Every refund request is decided by version 1, identically, with a record of which rule applied.

When the policy changes, steps 1–4 repeat and version 2 ships. Nothing changes silently.

What the research says

An April 2026 research preprint, Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation, tested this pattern. Its authors reported that generating code once broke even with calling a model at run time after about 17 transactions in their setup, and used 57 times fewer tokens at 1,000 transactions. These are the authors’ own results on their workloads, not a guarantee for yours; lesson 6 shows how to work out your own numbers.

Good candidates and poor ones

Compiling pays off when three things are true:

  • The rules are stable enough to write down. They change monthly or quarterly, not hourly.
  • There is real volume. The same kind of decision happens hundreds or thousands of times.
  • A wrong answer costs something. Money, fairness, a regulator’s attention or a customer’s trust.

It’s a poor fit for open-ended creative work, for inputs that change shape every day, and for personal productivity tasks where nobody relies on an identical answer tomorrow.

Try it yourself: the write-it-down test

Pick a decision your team makes repeatedly. Ask the person who knows it best to write the rules in plain English on one page.

  • If they can, and two experts reading the page would decide the same cases the same way, the decision is a strong candidate for compiling.
  • If the page keeps growing with “it depends”, note the exceptions. Some become rules; the truly ambiguous ones become cases a person handles, which is still a deterministic design: the code decides when to hand over.

What comes next

Not every step can be fully compiled. Lesson 4 introduces the Graduated Exposure Ladder, a way to decide how much model belongs at run time, step by step.

Compile once, run many times The model works before release. The code works for every customer after it. BUILD TIME: Read, Draft, Review. RUN TIME: Execute, Repeat, Trace. Both meet at: A versioned system. IOANYT ACADEMY / LESSON 3 Compile once, run many times The model works before release. The code works for every customer after it. BUILD TIME Read policies, examples, edge cases Draft code, rules and tests Review engineers approve every change RUN TIME Execute the reviewed logic, nothing else Repeat same input, same output Trace each result to a version A versioned system NOT A PROMPT a change in the rules is a new release, tested before it ships

Key takeaways

  • Compiled AI moves the model's work to build time, where it can be reviewed and tested.
  • What reaches customers is ordinary versioned software: repeatable, traceable, testable.
  • Changes to rules ship as new releases, so you always know which version made which decision.
  • The best candidates are decisions with stable rules, real volume and a cost when they go wrong.

Check yourself

Pick an answer, then open the card to compare.

  1. 1. In compiled AI, when does the language model do most of its work?

    • A.On every customer request
    • B.At build time, before release
    • C.Only when a customer complains
    Show the answer

    B. At build time, before release The model reads, drafts and tests at build time. At run time, the reviewed code executes.

  2. 2. A refund policy changes. What happens in a compiled system?

    • A.The model starts answering differently on its own
    • B.The logic is updated, re-tested and released as a new version
    • C.Nothing; compiled systems can't change
    Show the answer

    B. The logic is updated, re-tested and released as a new version Changes are ordinary software releases. That is what keeps every decision traceable to a version.

  3. 3. Which workflow is the best candidate for compiling?

    • A.Brainstorming campaign ideas
    • B.Deciding refund eligibility thousands of times a month under a written policy
    • C.Summarising meeting notes for yourself
    Show the answer

    B. Deciding refund eligibility thousands of times a month under a written policy Stable rules, real volume and a cost when wrong. The other two benefit from a live model's flexibility.

Common questions

Do we lose the model's flexibility?

At run time, deliberately, on the decision path. You keep flexibility where it helps, such as drafting or reading messy inputs, in bounded roles. Lesson 4 covers how to decide.

Who reviews the code the model writes?

Engineers. Model-drafted code, rules and tests go through the same review and testing as any other change before release.

Is this the same as a rules engine?

The output can look similar: explicit, testable logic. The difference is how it is built. Models do much of the drafting and testing, so logic is faster to create and easier to update.