IOanyT Innovations

Lesson 6 of 10 · 8 min read

Cost math: compiled vs runtime LLM

In one paragraph

The cost comparison between runtime and compiled AI is a break-even question: a runtime model costs a little on every transaction, while compiled logic costs more up front to build and test but almost nothing per transaction afterwards. Volume, how often the rules change, and latency decide which is cheaper for a given workflow.

In this lesson

  • Write down the two cost shapes and the break-even formula
  • Explain why token-only comparisons understate the real decision
  • Weigh volume, change rate, latency and stakes together

The video version of this lesson is in production. The full lesson is below.

Reliability is the main reason to compile the decision path. Cost is often the reason it gets approved. This lesson gives you a simple way to compare the two approaches for any workflow, using your own numbers.

Two cost shapes

Runtime model. Every transaction calls a model, so you pay tokens and compute on every case. The total grows in a straight line with volume: twice the cases, twice the bill.

Compiled logic. You pay up front, for the engineering to build, review and test the logic (with models doing much of the drafting). After that, each case costs ordinary compute, which is close to nothing per transaction. The total is almost flat as volume grows.

The two lines cross at the break-even volume.

The break-even formula

Break-even volume = up-front build cost ÷ (runtime cost per case − compiled cost per case)

An illustrative example, with numbers chosen for easy arithmetic rather than taken from any real project: if building and testing the compiled version costs 10,000, a model call costs 0.02 per case, and compiled execution costs effectively nothing, break-even is 10,000 ÷ 0.02 = 500,000 cases. Past that volume, every case is cheaper compiled; below it, the runtime model is cheaper on cost alone.

Use your own figures: real build estimates, your provider’s actual per-call cost at your prompt sizes, and your real monthly volume.

Why token-only figures look so different

An April 2026 research preprint on compiled AI reported break-even after about 17 transactions and 57 times fewer tokens at 1,000 transactions. That comparison counts tokens: the one-off tokens used to generate the code against the tokens a model uses on every call. It is a fair technical measure, and it shows how quickly runtime token costs add up.

For a business decision, also count the engineering time to review, test and maintain the compiled logic. That’s why a full-cost break-even sits much higher than a token-only one. Both numbers are honest; they answer different questions.

Four factors, weighed together

FactorFavours runtime modelFavours compiled
VolumeLow, occasional useHigh, steady volume
Change rateRules change dailyRules change monthly or quarterly
LatencySeconds are acceptableReal-time or high throughput
StakesA person reviews every outputCustomers, money or regulators rely on the answer

Stakes often decide it before cost does: if a decision has to be repeatable and explainable, a runtime model in the decision path is the wrong design at any price.

The costs people forget

  • Evaluation and monitoring. Both approaches need a golden dataset and regression runs (lesson 5). Runtime models also need watching for drift when providers update them.
  • Re-compiling. Each rule change means a new build-and-test cycle. Frequent change raises the compiled line.
  • Incidents. A wrong answer that reaches customers has a cost (refunds, complaints, regulators) that rarely appears in a per-call price.

Try it yourself: a five-minute estimate

For one workflow, write down:

  1. Monthly volume of decisions.
  2. Your current cost per decision with a runtime model (from your provider’s bill ÷ volume).
  3. A rough estimate to build and test a compiled version.
  4. How often the rules change per year.

Divide (3) by (2) to get break-even volume, and compare it with (1) × 12. Then check the stakes column. If the decision must be repeatable, the cost question is about when, not whether.

What comes next

The next three lessons turn from engineering to accountability: who is liable for what an AI says (lesson 7), what regulators have penalised (lesson 8), and what the EU AI Act requires (lesson 9).

Two ways to pay for intelligence Runtime models charge on every transaction. Compiled logic is paid for once, at build time. RUNTIME MODEL: Per call, Grows with volume, Model latency. COMPILED: Up front, Near-flat, Code speed. Both meet at: Break-even volume. IOANYT ACADEMY / LESSON 6 Two ways to pay for intelligence Runtime models charge on every transaction. Compiled logic is paid for once, at build time. RUNTIME MODEL Per call tokens on every transaction Grows with volume twice the cases, twice the bill Model latency a network call per decision COMPILED Up front engineering to build and test Near-flat ordinary compute per case Code speed no model in the path Break-even volume WHERE THE LINES CROSS the right answer depends on your volume, change rate and stakes

Key takeaways

  • Runtime models charge per transaction; compiled logic is mostly paid for once.
  • Break-even volume = up-front build cost ÷ (runtime cost per case − compiled cost per case).
  • Token-only comparisons (like the research figure of about 17 transactions) leave out engineering; include it for a real decision.
  • Fast-changing rules raise the cost of compiling; high volume and high stakes favour it.

Check yourself

Pick an answer, then open the card to compare.

  1. 1. Which cost shape describes a runtime model?

    • A.Large up-front cost, near-zero per transaction
    • B.Small cost on every transaction, growing with volume
    • C.A fixed monthly fee regardless of use
    Show the answer

    B. Small cost on every transaction, growing with volume Every call uses tokens and compute, so the bill rises with the number of cases.

  2. 2. Building and testing a compiled workflow costs 10,000; a model call costs 0.02 per case and compiled execution about nothing. Roughly where is break-even?

    • A.About 5,000 cases
    • B.About 500,000 cases
    • C.About 50 cases
    Show the answer

    B. About 500,000 cases 10,000 ÷ (0.02 − 0) = 500,000 cases. These are illustrative numbers; use your own.

  3. 3. Which factor makes compiling LESS attractive?

    • A.Very high volume
    • B.Rules that change every day
    • C.Decisions with legal consequences
    Show the answer

    B. Rules that change every day Each rule change means re-building and re-testing. Daily change raises the up-front cost again and again.

Common questions

The research said break-even at about 17 transactions. Why is your example in the hundreds of thousands?

The research compared tokens: the one-off tokens to generate code against the tokens a model uses per call. It didn't include engineering time for review and testing. Both views are useful; a business decision needs the second.

Doesn't model pricing keep falling?

It has fallen a lot, which lowers the runtime line. But prices can also rise or change structure. Compiled logic doesn't depend on a model's price at run time, which acts as a hedge.

What about latency?

A runtime model adds a network round trip and generation time to every decision. Compiled logic runs at ordinary code speed. For real-time or high-throughput paths, that alone can decide it.