
Foundry Field Reports
The AI Conversion Engine
Every company can buy the same models. The only durable advantage left is how fast you convert new capability into a measured, adopted, cost-optimized part of how the business runs.
Why most AI programs stall, and how to build one that compounds.
Contents (12)
- 01The measurement gap.
- 02The race is for the fastest conversion.
- 03Instrument one: knowing when AI is good enough.
- 04Instrument two: doing it repeatably.
- 05Where the revenue comes from.
- 06Inside the engine: ten capabilities.
- 07Scoring the engine.
- 08Inside the task scorecard.
- 09Three scores, not one.
- 10The gates that override the math.
- 11Where tasks saturate, and where the frontier keeps winning.
- 12Two worked examples.
01
The measurement gap.
Almost every executive team is running AI experiments right now. Very few can point to the revenue.
The technology works fine in most cases. The real failure is that most organizations have no way to measure whether it's working. A pilot saves someone two hours a week, but nobody tracked how long the task took before, nobody knows what those hours were redeployed into, and six months later the CFO asks a fair question: what did we actually get? At most companies, nobody actually knows.
“Productivity gains” is where AI investment goes to disappear.
That measurement gap is the central reason AI programs stall, not a side issue. When you can't measure impact, prioritization becomes guesswork, scaling stalls, failing bets don't get killed, and the case for the next dollar of investment falls apart. The experiments pile up. The transformation never arrives.

02
The race is for the fastest conversion.
Every company now has access to roughly the same AI models. Model access is not an advantage and never will be.
What can be an advantage is speed of conversion: how quickly your organization turns a new AI capability into a measured, adopted, cost-optimized part of how the business runs. This is the loop the best operators are building.
Discover with frontier models · measure until reliable · harness into a process · drive adoption · downshift to cheaper models · reinvest the savings.
Run it once and you get a useful automation. Run it continuously and you get a compounding advantage: every cycle lowers costs, frees capacity, generates data, and funds the next cycle. The companies that win will be the ones with the shortest loop, not the flashiest demo.
The framework has two instruments: a task-level scorecard that tells you where AI is ready to create value right now, and a company-level readiness score that tells you whether your organization can capture that value repeatedly. One finds the money. The other builds the machine that keeps finding it.
03
Instrument one: knowing when AI is good enough.
Every AI-automated task follows the same arc. At first it isn't good enough. Then it improves fast. Eventually it crosses a threshold where additional intelligence stops adding business value.
The s-curve threshold: past the good-enough bar, every extra dollar of intelligence is margin donated back.
Consider a system that routes inquiries correctly 97% of the time, with the rest escalated to a person. Paying ten times more for a premium model that reaches 98.5% buys almost nothing — the escalation path already catches the difference. The scorecard evaluates each candidate on three questions.
The three questions
- Is it ready? Repetitive, predictable tasks with clear right answers, recoverable errors, and cheap review cross the threshold early. Ambiguous, high-stakes work doesn't, and shouldn't be forced.
- Is it worth it? Volume, labor cost, and proximity to revenue set priority. A perfectly automatable task nobody does often deserves no engineering time.
- Will the investment last? Infrastructure built on business logic gets more valuable as models improve. Infrastructure built on model babysitting gets deleted.
One distinction wastes real money when it's confused: a task can be critical to revenue and still not need the most expensive AI. Product tagging drives conversion, but once tags are 97% accurate and mistakes are cheap to fix, premium intelligence adds nothing. What matters is whether more intelligence changes the outcome.
Importance sets the priority. Quality-sensitivity sets the model budget.
04
Instrument two: doing it repeatably.
Finding one good automation is easy. Building an organization that converts capability into results over and over is the hard part — and where the compounding lives.
The Capability Conversion Readiness score measures the organization across ten dimensions, weighted by how load-bearing each is. Four are worth naming, because they're where programs actually die: measurement (can you tell if it improved or regressed this week?), adoption (a perfect system employees route around is worth zero), speed (a six-month cycle can't capture the window between newly possible and commoditized), and economics (cost per task, savings, revenue attributed — tracked, not guessed).
load-bearing — a score of 2 or below here caps the whole total
Evals and the harness carry the most weight; every other dimension depends on them.
This system runs at the speed of its weakest link.
The score deliberately penalizes bottlenecks rather than averaging them away, because an average hides exactly the weakness that will stall the program. Brilliant technology with no adoption fails. Enthusiastic adoption with no measurement is dangerous. Great measurement with slow deployment misses the window.
05
Where the revenue comes from.
Cost savings are the visible part of this framework. They are not the real story. What matters more is what the loop does to revenue.
It redeploys your best people toward growth. Every routine task moved to cheap, reliable automation returns measured capacity that can be pointed at selling, serving, and building. It compresses the cycles that drive revenue — lead response, proposals, onboarding, support resolution. It reserves premium intelligence for premium problems. It funds its own expansion, turning AI from a cost center into a flywheel. And it compounds: every cycle creates data and infrastructure that make the next one faster.

The questions to put to your team
- Which workflows are close to revenue, high-volume, and measurable — and what does each cost per task today?
- For our current AI systems: is anyone actually using them, and how do we know?
- Which AI work has crossed the good-enough line where we're overpaying for premium models?
- How long does it take us to go from “newly possible” to “running in production and used”?
- Where is the measurement that would let us answer any of the above with a number instead of an anecdote?
06
Inside the engine: ten capabilities.
The engine runs ten capabilities in the order they operate across one cycle. Weakness early in the chain breaks things later, in ways that are hard to diagnose unless you know where to look.

Sensing — where frontier capability appears.
Weak sensing means you find out about new capability the way everyone else does: a vendor launch, a competitor's product, an employee on a personal account. Strong companies track model releases like a trading desk tracks a market and test new models against their own tasks as routine. A vendor's benchmark measures something else entirely; the only test that matters is your own task.
The task portfolio — automation as an investment book.
A living, ranked map of every candidate workflow — economic ranking, saturation score, risk classification, named ownership, refresh cadence — beats a scattering of pilots nobody compares. Ten disconnected pilots produce ten stories and no mechanism for killing a bet that already failed twice.
Evals — the measurement layer everything depends on.
Task-specific test suites that score a system against your own real cases. Evals are the dimension most likely to become the central bottleneck: without them you can't safely downshift, can't diagnose stalled adoption, can't prove ROI. Ground truth, regression testing, grader calibration, edge-case coverage, saturation detection, and freshness are the parts that make an eval trustworthy rather than a vibe.
A company that cannot answer these questions isn't being cautious. It's operating blind.
The harness — turning a model into a business process.
The software and process layer wrapped around a raw model: model abstraction, versioning, safe integration and rule enforcement, escalation and permissioning, observability, rollback and state. A model without a harness is a demo. The payoff shows up the second time — the first workflow is built from nothing; the second ships in a fraction of the time on the same substrate.

Traces — how the system learns.
The recorded history of each pass through the system — input, output, model, cost, human edits, whether the output was accepted or overridden, the failure mode, the eventual outcome. Traces become new eval cases, training data, and the evidence that justifies moving a segment to cheaper infrastructure. What separates a learning loop from a log is feedback latency: closing the loop in days, not quarters.

Routing — the profit lever.
Which model handles which piece of work, decided systematically: frontier models for novel or high-value cases, mid-tier for the normal run, small or open models for saturated work, deterministic rules where a model was never needed, humans for the genuinely ambiguous. This is where the cost savings actually get realized.

Adoption
where value is won or lost
The point where a working system becomes how people do their jobs — or a tool beside the real workflow. Placement, incentive, managerial backing, and workflow redesign decide it. The quiet tell of failure: a shadow workflow running the real work through email.
Velocity
shipping faster than the frontier moves
One cycle-time number: capability observed → evaluated → harnessed → adopted → downshifted. A six-month deployment cycle is structurally incapable of capturing the window before an advantage is arbitraged away.
Governance
lanes, not gates
Risk tiers and approval lanes so a low-stakes internal tool and a customer-facing decision never run the same gauntlet; auditability, data controls, a real incident process. Strong governance speeds deployment up — every team knows what's allowed before they start.
Economics
closing the loop
Cost per task, return attribution, the frontier premium, downshift savings — tracked well enough to make real capital-allocation decisions. Without economics, nothing else can be justified when budgets tighten, and the reinvestment step never happens.
07
Scoring the engine.
Everything above turns into a single number — and the arithmetic itself carries a message: weakness in the wrong place caps the whole score, no matter how well everything else performs.
| Score | What it means |
|---|---|
| 0 | No meaningful capability exists. |
| 1 | Ad hoc — dependent on individual champions; if they leave, it leaves. |
| 2 | Isolated pilots exist, but nothing repeatable. |
| 3 | A repeatable process exists in some teams. |
| 4 | The capability is institutionalized across important workflows. |
| 5 | A compounding system: clear ownership, active measurement, a running improvement loop. |
A bottleneck penalty sits on top of the weighted total, and some failures cap it outright. Eval maturity at 1 or below caps the total at 50. Adoption at 1 or below caps it at 60; the harness at 1 or below, at 55. The loop runs at the speed of its weakest load-bearing link.
0–25
Experimentation
26–50
Pilot
51–70
Operational
71–85
Compounding
86–100
AI-native OS
The resulting number sorts a company into one of five stages.
Model access is equal for everyone. Conversion speed is the only durable advantage left standing.
08
Inside the task scorecard.
Every rollout asks the same question of every task: is this good enough yet, and good enough for what. The right object of analysis is never a task in isolation.
Task × Company Context × Model-and-Harness × Business Consequence.
The same task — ticket classification, contract review, product tagging — can sit on opposite sides of the automation line at two companies, or at the same company a year apart, because the context changed even though the task description didn't. Eight entities make up the full picture; skipping any one is where a scorecard exercise goes wrong.
| Entity | What it governs |
|---|---|
| Company Context | The environment: business model, customer tolerance, regulatory exposure, release velocity. |
| Business Process | Where the task sits: revenue proximity, failure visibility, workflow maturity. |
| Task | The work itself: repetitiveness, ambiguity, determinism, tool dependency. |
| Model Option | The candidate: capability, cost, reliability, deployment control, privacy fit. |
| Harness | Everything wrapped around the model — the entity that keeps a threshold crossed. |
| Error Surface | How bad a failure is: severity, recoverability, detectability, time to detection. |
| Evaluation System | Whether you can even know a threshold was crossed: ground truth, agreement, freshness. |
| Economic Outcome | What it's worth: volume, revenue impact, frontier premium, cheap-model savings. |
Two systems, one question.
The real comparison is not one model against another. System A pairs a frontier model with a light harness. System B pairs a cheaper model with a stronger harness. The question is whether System B has reached the business quality bar where System A's extra intelligence stops changing the result that matters.
| Condition | Requirement |
|---|---|
| Quality bar | System B meets or exceeds the minimum acceptable quality for the task. |
| Value vs. cost | System A's incremental value is smaller than its incremental cost, latency, complexity, and vendor risk combined. |
Prioritization instrument
Rate a real task while you read.
The Task Scorecard runs this exact framework — three scores, five gates, the same weights — on one task you pick. Rate it in a few minutes and get a live verdict.
Open the Task Scorecard09
Three scores, not one.
The framework keeps three scores separate on purpose. Collapsing them into a single number early hides exactly the information an executive needs.
Saturation Readiness — can this safely move from frontier to a cheaper model plus harness? Economic Priority — even if automatable, is it worth anyone's time? Harness Durability — will the investment survive the next model release, or get erased by it? A task can be perfectly ready and not worth doing; enormously valuable and nowhere near ready; ready and valuable and built on a harness one upgrade will make worthless.
Read together, three scores make a decision. Read alone, three arguments that never resolve.
| Saturation | Priority | Durability | Recommended strategy |
|---|---|---|---|
| High | High | High | Move to a cheaper model plus harness; use frontier for exceptions. |
| High | High | Low | Automate, but rebuild the harness around durable assets, not brittle prompts. |
| High | Low | Any | Automate only if the effort is genuinely low. |
| Medium | High | High | Model routing: cheap model first, frontier escalation. |
| Medium | High | Low | Use frontier while building evals, logs, and better workflow structure. |
| Low | High | High | Keep the frontier model; use it as a scout and data generator. |
| Low | High | Low | Risky; prototype and build evals before committing. |
| Low | Low | Any | Deprioritize. |
10
The gates that override the math.
A good weighted score is not a license to skip judgment. Five gates sit above the scoring model. A task that fails one does not get to downshift, no matter how attractive its numbers.
The five gates
- Minimum acceptable quality. The cheaper system must clear the SLA the business actually requires.
- Error severity. Major legal, financial, safety, or reputational harm demands frontier, human review, or both.
- Evaluation confidence. If it can't be reliably evaluated, the next investment is better evals, not a better model.
- Distribution drift. If inputs change quickly, today's saturation may be temporary.
- Hidden failure risk. If failures are quiet, delayed, or customer-discovered, safeguards must exceed the score.
| Illustrative SLA (varies by company and task) | Example threshold |
|---|---|
| Extraction accuracy | ≥ 98% |
| Critical-field accuracy | ≥ 99.5% |
| Hallucination rate | ≤ 0.5% |
| Escalation accuracy | ≥ 95% |
| Human override rate | ≤ 10% |
11
Where tasks saturate, and where the frontier keeps winning.
Some task types saturate early almost regardless of company. Others keep rewarding frontier capability well past the point where most work has commoditized.
| Saturates early | Why |
|---|---|
| Classification | Clear labels and objective evals. |
| Extraction | Ground truth can be checked directly. |
| Support triage | High volume, recoverable, easy escalation. |
| Invoice & receipt processing | Structured, repetitive, measurable. |
| Knowledge-base QA | Grounded in retrieval and citations. |
| Frontier keeps winning | Why |
|---|---|
| Novel strategy | No stable ground truth exists. |
| Complex legal reasoning | High consequence and ambiguity. |
| Long-horizon agents | Planning failures compound over steps. |
| Executive decision support | Small quality differences matter a lot. |
| Creative direction | Taste and originality may be the product. |
A task belongs in the first table because it can be constrained and scored — not because it matters less to the business.
12
Two worked examples.
Two tasks show how the three scores and the gates resolve in practice.
Worked example
Customer-support ticket classification
Repetitive, easy to evaluate, recoverable, high volume. All three scores clear a strong bar — a clear downshift candidate. A cheap model handles routine classification by default; deterministic rules enforce tier and SLA logic; uncertain cases escalate to frontier or a human; corrections feed the eval set; weekly runs decide whether more traffic can shift.
Worked example
Enterprise contract redlining
Ambiguous, high-consequence, quality-sensitive — the saturation score stays low even though the task matters enormously. Keep the frontier model, paired with retrieval, strict permissions, mandatory human approval, and a strong audit trail.
A harness built for a task that is not yet saturated is not a wasted investment. It is what makes the eventual downshift possible.
Your turn
Score your own highest-volume task.
You've seen how a routine task and a high-stakes one land on opposite ends of the same scale. Run the numbers on a task from your own queue — same three scores, same five gates, a verdict in minutes.
Open the Task ScorecardThe models get better for everyone, at the same time.
That improvement is free to your competitors too. The only durable advantage is being the organization that converts it fastest — and can prove it in numbers.