Agent Loop Foundry.
A dark operations room at night, walls of rack-mounted equipment; a worn industrial robot sits at the console, reading the glowing dashboard.

Foundry Field Reports

The AI Conversion Engine

Every company can buy the same models. The only durable advantage left is how fast you convert new capability into a measured, adopted, cost-optimized part of how the business runs.

Report 0112 sections18 min read
FIG. 00Operations, after hours. The loop never sleeps.

Why most AI programs stall, and how to build one that compounds.

Contents (12)
  1. 01The measurement gap.
  2. 02The race is for the fastest conversion.
  3. 03Instrument one: knowing when AI is good enough.
  4. 04Instrument two: doing it repeatably.
  5. 05Where the revenue comes from.
  6. 06Inside the engine: ten capabilities.
  7. 07Scoring the engine.
  8. 08Inside the task scorecard.
  9. 09Three scores, not one.
  10. 10The gates that override the math.
  11. 11Where tasks saturate, and where the frontier keeps winning.
  12. 12Two worked examples.

01

The measurement gap.

Almost every executive team is running AI experiments right now. Very few can point to the revenue.

The technology works fine in most cases. The real failure is that most organizations have no way to measure whether it's working. A pilot saves someone two hours a week, but nobody tracked how long the task took before, nobody knows what those hours were redeployed into, and six months later the CFO asks a fair question: what did we actually get? At most companies, nobody actually knows.

“Productivity gains” is where AI investment goes to disappear.

That measurement gap is the central reason AI programs stall, not a side issue. When you can't measure impact, prioritization becomes guesswork, scaling stalls, failing bets don't get killed, and the case for the next dollar of investment falls apart. The experiments pile up. The transformation never arrives.

A tall stack of printed reports on a desk under a hard lamp, the top one open and covered in handwritten marks, a calculator and red pen beside it.
FIG. 01 The pilots ran. Nobody measured what they changed.

02

The race is for the fastest conversion.

Every company now has access to roughly the same AI models. Model access is not an advantage and never will be.

What can be an advantage is speed of conversion: how quickly your organization turns a new AI capability into a measured, adopted, cost-optimized part of how the business runs. This is the loop the best operators are building.

Discoverfrontier modelMeasureevalsHarnessmake it a processAdoptchange the workDownshiftcheaper modelReinvestfund the nextshortest loopwins

Discover with frontier models · measure until reliable · harness into a process · drive adoption · downshift to cheaper models · reinvest the savings.

Run it once and you get a useful automation. Run it continuously and you get a compounding advantage: every cycle lowers costs, frees capacity, generates data, and funds the next cycle. The companies that win will be the ones with the shortest loop, not the flashiest demo.

The framework has two instruments: a task-level scorecard that tells you where AI is ready to create value right now, and a company-level readiness score that tells you whether your organization can capture that value repeatedly. One finds the money. The other builds the machine that keeps finding it.

03

Instrument one: knowing when AI is good enough.

Every AI-automated task follows the same arc. At first it isn't good enough. Then it improves fast. Eventually it crosses a threshold where additional intelligence stops adding business value.

good-enough thresholdmore intelligence,~no more valueBUSINESS VALUEMODEL CAPABILITY & COST →

The s-curve threshold: past the good-enough bar, every extra dollar of intelligence is margin donated back.

Consider a system that routes inquiries correctly 97% of the time, with the rest escalated to a person. Paying ten times more for a premium model that reaches 98.5% buys almost nothing — the escalation path already catches the difference. The scorecard evaluates each candidate on three questions.

The three questions

  • Is it ready? Repetitive, predictable tasks with clear right answers, recoverable errors, and cheap review cross the threshold early. Ambiguous, high-stakes work doesn't, and shouldn't be forced.
  • Is it worth it? Volume, labor cost, and proximity to revenue set priority. A perfectly automatable task nobody does often deserves no engineering time.
  • Will the investment last? Infrastructure built on business logic gets more valuable as models improve. Infrastructure built on model babysitting gets deleted.

One distinction wastes real money when it's confused: a task can be critical to revenue and still not need the most expensive AI. Product tagging drives conversion, but once tags are 97% accurate and mistakes are cheap to fix, premium intelligence adds nothing. What matters is whether more intelligence changes the outcome.

Importance sets the priority. Quality-sensitivity sets the model budget.

04

Instrument two: doing it repeatably.

Finding one good automation is easy. Building an organization that converts capability into results over and over is the hard part — and where the compounding lives.

The Capability Conversion Readiness score measures the organization across ten dimensions, weighted by how load-bearing each is. Four are worth naming, because they're where programs actually die: measurement (can you tell if it improved or regressed this week?), adoption (a perfect system employees route around is worth zero), speed (a six-month cycle can't capture the window between newly possible and commoditized), and economics (cost per task, savings, revenue attributed — tracked, not guessed).

Evals15
Harness15
Traces12
Adoption12
Task portfolio10
Routing10
Sensing8
Velocity8
Governance5
Economics5

load-bearing — a score of 2 or below here caps the whole total

Evals and the harness carry the most weight; every other dimension depends on them.

This system runs at the speed of its weakest link.

The score deliberately penalizes bottlenecks rather than averaging them away, because an average hides exactly the weakness that will stall the program. Brilliant technology with no adoption fails. Enthusiastic adoption with no measurement is dangerous. Great measurement with slow deployment misses the window.

05

Where the revenue comes from.

Cost savings are the visible part of this framework. They are not the real story. What matters more is what the loop does to revenue.

It redeploys your best people toward growth. Every routine task moved to cheap, reliable automation returns measured capacity that can be pointed at selling, serving, and building. It compresses the cycles that drive revenue — lead response, proposals, onboarding, support resolution. It reserves premium intelligence for premium problems. It funds its own expansion, turning AI from a cost center into a flywheel. And it compounds: every cycle creates data and infrastructure that make the next one faster.

A vintage industrial punch clock on a factory wall beside a full rack of paper time cards, one card inserted mid-punch.
FIG. 02 Capacity, measured — not assumed.

The questions to put to your team

  • Which workflows are close to revenue, high-volume, and measurable — and what does each cost per task today?
  • For our current AI systems: is anyone actually using them, and how do we know?
  • Which AI work has crossed the good-enough line where we're overpaying for premium models?
  • How long does it take us to go from “newly possible” to “running in production and used”?
  • Where is the measurement that would let us answer any of the above with a number instead of an anecdote?

06

Inside the engine: ten capabilities.

The engine runs ten capabilities in the order they operate across one cycle. Weakness early in the chain breaks things later, in ways that are hard to diagnose unless you know where to look.

Close-up of a large industrial machine's control section, rows of gauges, dials, and heavy levers in brushed worn metal.
FIG. 03 The engine, mid-cycle. Ten parts, one machine.

Sensing — where frontier capability appears.

Weak sensing means you find out about new capability the way everyone else does: a vendor launch, a competitor's product, an employee on a personal account. Strong companies track model releases like a trading desk tracks a market and test new models against their own tasks as routine. A vendor's benchmark measures something else entirely; the only test that matters is your own task.

The task portfolio — automation as an investment book.

A living, ranked map of every candidate workflow — economic ranking, saturation score, risk classification, named ownership, refresh cadence — beats a scattering of pilots nobody compares. Ten disconnected pilots produce ten stories and no mechanism for killing a bet that already failed twice.

Evals — the measurement layer everything depends on.

Task-specific test suites that score a system against your own real cases. Evals are the dimension most likely to become the central bottleneck: without them you can't safely downshift, can't diagnose stalled adoption, can't prove ROI. Ground truth, regression testing, grader calibration, edge-case coverage, saturation detection, and freshness are the parts that make an eval trustworthy rather than a vibe.

A company that cannot answer these questions isn't being cautious. It's operating blind.

The harness — turning a model into a business process.

The software and process layer wrapped around a raw model: model abstraction, versioning, safe integration and rule enforcement, escalation and permissioning, observability, rollback and state. A model without a harness is a demo. The payoff shows up the second time — the first workflow is built from nothing; the second ships in a fraction of the time on the same substrate.

Industrial pipework and bundled cable conduits running along a concrete wall, valves and junction boxes.
FIG. 04 The plumbing that turns a model into a process.

Traces — how the system learns.

The recorded history of each pass through the system — input, output, model, cost, human edits, whether the output was accepted or overridden, the failure mode, the eventual outcome. Traces become new eval cases, training data, and the evidence that justifies moving a segment to cheaper infrastructure. What separates a learning loop from a log is feedback latency: closing the loop in days, not quarters.

A dark data room, server racks with rows of small indicator lights; a worn industrial robot stands in the aisle among the cables.
FIG. 05 Every production run is either evidence or a wasted opportunity.

Routing — the profit lever.

Which model handles which piece of work, decided systematically: frontier models for novel or high-value cases, mid-tier for the normal run, small or open models for saturated work, deterministic rules where a model was never needed, humans for the genuinely ambiguous. This is where the cost savings actually get realized.

A railway switching yard at dusk, many steel tracks diverging through junction points and signals.
FIG. 06 Every task on a frontier model past its threshold is margin handed back.

Adoption

where value is won or lost

The point where a working system becomes how people do their jobs — or a tool beside the real workflow. Placement, incentive, managerial backing, and workflow redesign decide it. The quiet tell of failure: a shadow workflow running the real work through email.

Velocity

shipping faster than the frontier moves

One cycle-time number: capability observed → evaluated → harnessed → adopted → downshifted. A six-month deployment cycle is structurally incapable of capturing the window before an advantage is arbitraged away.

Governance

lanes, not gates

Risk tiers and approval lanes so a low-stakes internal tool and a customer-facing decision never run the same gauntlet; auditability, data controls, a real incident process. Strong governance speeds deployment up — every team knows what's allowed before they start.

Economics

closing the loop

Cost per task, return attribution, the frontier premium, downshift savings — tracked well enough to make real capital-allocation decisions. Without economics, nothing else can be justified when budgets tighten, and the reinvestment step never happens.

07

Scoring the engine.

Everything above turns into a single number — and the arithmetic itself carries a message: weakness in the wrong place caps the whole score, no matter how well everything else performs.

ScoreWhat it means
0No meaningful capability exists.
1Ad hoc — dependent on individual champions; if they leave, it leaves.
2Isolated pilots exist, but nothing repeatable.
3A repeatable process exists in some teams.
4The capability is institutionalized across important workflows.
5A compounding system: clear ownership, active measurement, a running improvement loop.
Each of the ten dimensions is scored 0–5 against these anchors, then weighted.

A bottleneck penalty sits on top of the weighted total, and some failures cap it outright. Eval maturity at 1 or below caps the total at 50. Adoption at 1 or below caps it at 60; the harness at 1 or below, at 55. The loop runs at the speed of its weakest load-bearing link.

0–25

Experimentation

26–50

Pilot

51–70

Operational

71–85

Compounding

86–100

AI-native OS

The resulting number sorts a company into one of five stages.

Model access is equal for everyone. Conversion speed is the only durable advantage left standing.

08

Inside the task scorecard.

Every rollout asks the same question of every task: is this good enough yet, and good enough for what. The right object of analysis is never a task in isolation.

Task × Company Context × Model-and-Harness × Business Consequence.

The same task — ticket classification, contract review, product tagging — can sit on opposite sides of the automation line at two companies, or at the same company a year apart, because the context changed even though the task description didn't. Eight entities make up the full picture; skipping any one is where a scorecard exercise goes wrong.

EntityWhat it governs
Company ContextThe environment: business model, customer tolerance, regulatory exposure, release velocity.
Business ProcessWhere the task sits: revenue proximity, failure visibility, workflow maturity.
TaskThe work itself: repetitiveness, ambiguity, determinism, tool dependency.
Model OptionThe candidate: capability, cost, reliability, deployment control, privacy fit.
HarnessEverything wrapped around the model — the entity that keeps a threshold crossed.
Error SurfaceHow bad a failure is: severity, recoverability, detectability, time to detection.
Evaluation SystemWhether you can even know a threshold was crossed: ground truth, agreement, freshness.
Economic OutcomeWhat it's worth: volume, revenue impact, frontier premium, cheap-model savings.

Two systems, one question.

The real comparison is not one model against another. System A pairs a frontier model with a light harness. System B pairs a cheaper model with a stronger harness. The question is whether System B has reached the business quality bar where System A's extra intelligence stops changing the result that matters.

ConditionRequirement
Quality barSystem B meets or exceeds the minimum acceptable quality for the task.
Value vs. costSystem A's incremental value is smaller than its incremental cost, latency, complexity, and vendor risk combined.
A task sits in the saturation zone once both conditions hold at once.

Prioritization instrument

Rate a real task while you read.

The Task Scorecard runs this exact framework — three scores, five gates, the same weights — on one task you pick. Rate it in a few minutes and get a live verdict.

Open the Task Scorecard

09

Three scores, not one.

The framework keeps three scores separate on purpose. Collapsing them into a single number early hides exactly the information an executive needs.

Saturation Readiness — can this safely move from frontier to a cheaper model plus harness? Economic Priority — even if automatable, is it worth anyone's time? Harness Durability — will the investment survive the next model release, or get erased by it? A task can be perfectly ready and not worth doing; enormously valuable and nowhere near ready; ready and valuable and built on a harness one upgrade will make worthless.

Read together, three scores make a decision. Read alone, three arguments that never resolve.

SaturationPriorityDurabilityRecommended strategy
HighHighHighMove to a cheaper model plus harness; use frontier for exceptions.
HighHighLowAutomate, but rebuild the harness around durable assets, not brittle prompts.
HighLowAnyAutomate only if the effort is genuinely low.
MediumHighHighModel routing: cheap model first, frontier escalation.
MediumHighLowUse frontier while building evals, logs, and better workflow structure.
LowHighHighKeep the frontier model; use it as a scout and data generator.
LowHighLowRisky; prototype and build evals before committing.
LowLowAnyDeprioritize.
Reading the three scores together — the fastest way to turn three numbers into a decision.

10

The gates that override the math.

A good weighted score is not a license to skip judgment. Five gates sit above the scoring model. A task that fails one does not get to downshift, no matter how attractive its numbers.

The five gates

  • Minimum acceptable quality. The cheaper system must clear the SLA the business actually requires.
  • Error severity. Major legal, financial, safety, or reputational harm demands frontier, human review, or both.
  • Evaluation confidence. If it can't be reliably evaluated, the next investment is better evals, not a better model.
  • Distribution drift. If inputs change quickly, today's saturation may be temporary.
  • Hidden failure risk. If failures are quiet, delayed, or customer-discovered, safeguards must exceed the score.
Illustrative SLA (varies by company and task)Example threshold
Extraction accuracy≥ 98%
Critical-field accuracy≥ 99.5%
Hallucination rate≤ 0.5%
Escalation accuracy≥ 95%
Human override rate≤ 10%
Illustrations, not universal standards — a fintech and a meeting summarizer set entirely different numbers.

11

Where tasks saturate, and where the frontier keeps winning.

Some task types saturate early almost regardless of company. Others keep rewarding frontier capability well past the point where most work has commoditized.

Saturates earlyWhy
ClassificationClear labels and objective evals.
ExtractionGround truth can be checked directly.
Support triageHigh volume, recoverable, easy escalation.
Invoice & receipt processingStructured, repetitive, measurable.
Knowledge-base QAGrounded in retrieval and citations.
Frontier keeps winningWhy
Novel strategyNo stable ground truth exists.
Complex legal reasoningHigh consequence and ambiguity.
Long-horizon agentsPlanning failures compound over steps.
Executive decision supportSmall quality differences matter a lot.
Creative directionTaste and originality may be the product.

A task belongs in the first table because it can be constrained and scored — not because it matters less to the business.

12

Two worked examples.

Two tasks show how the three scores and the gates resolve in practice.

Worked example

Customer-support ticket classification

Saturation Readinessstrong
82
Economic Priorityhigh
76
Harness Durabilitystrong
84

Repetitive, easy to evaluate, recoverable, high volume. All three scores clear a strong bar — a clear downshift candidate. A cheap model handles routine classification by default; deterministic rules enforce tier and SLA logic; uncertain cases escalate to frontier or a human; corrections feed the eval set; weekly runs decide whether more traffic can shift.

Worked example

Enterprise contract redlining

Saturation Readinessnot saturated
35
Economic Priorityhigh
78
Harness Durabilitymoderate
72

Ambiguous, high-consequence, quality-sensitive — the saturation score stays low even though the task matters enormously. Keep the frontier model, paired with retrieval, strict permissions, mandatory human approval, and a strong audit trail.

A harness built for a task that is not yet saturated is not a wasted investment. It is what makes the eventual downshift possible.

Your turn

Score your own highest-volume task.

You've seen how a routine task and a high-stakes one land on opposite ends of the same scale. Run the numbers on a task from your own queue — same three scores, same five gates, a verdict in minutes.

Open the Task Scorecard

The models get better for everyone, at the same time.

That improvement is free to your competitors too. The only durable advantage is being the organization that converts it fastest — and can prove it in numbers.