A named owner puts the quality bar, the dealbreaker failures, and the operating constraints in writing before anything is tested.
Independent AI unit economics for software products
Your AI model mix is a gross margin decision.
AI Economics Lab tests which requests and workflow stages actually need premium models, which can route down safely, and what happens to product margin after retries, validation, rework, and human review are counted.
See how Benchmark #1 works →Stephen McDaniel. I build and ship production AI systems and previously led analytics and product work at Tableau, Netflix, and SAS. The quality bar is fixed before testing, and the fee does not change when the expensive model wins.
That is an illustration, not a savings claim. If your workload is materially smaller or the route cannot plausibly change the economics, use the free scorecard first.
The commercial front door
$10K to answer a product routing question with evidence.
Two weeks, 3 to 5 repeatable product workloads. Determine which requests should run on the default model, which should route down, when to escalate, and what the decision does to accepted-output economics.
AI Economics Audit
Find where premium inference genuinely changes the accepted outcome, where lower-cost routes are already good enough, and where a cheap route becomes expensive once the full workflow is counted.
Your current setup, cheaper alternatives, premium models, and a cheap-model-plus-checker cascade, all judged against the same bar.
Model and tool spend, checkers, retries, rework, human review, and the work time it actually takes to reach an answer you can use.
Default route, when to escalate, what to fall back to, whether premium is justified, the annual numbers, and the hours your team gets back.
- A written product quality bar and dealbreaker list, signed off by a named owner
- A side-by-side route comparison with CPAO for each option
- A default route, an escalation rule, and a fallback
- Annualized product AI cost impact and review/rework implications
- The trigger that tells you when to re-test
- Every number reproducible from the run record
- One repeatable job, judged by one acceptance test, owned by one person. Fifty pipelines doing the same job against the same bar count as one workload, not fifty.
- Examples: text-to-SQL · classification · extraction · AI search or RAG responses · agent stages · request routing · customer-facing summaries.
- Not a workload: an entire product, a whole department, or "all our AI". If you are unsure how yours splits, use Work with us and I will scope it before you pay anything.
No production access needed. Representative, de-identified samples only. Any commercial relationship with a benchmarked vendor is disclosed with the relevant research.
For AI, analytics, BI, data and software vendors
Three places inference economics quietly attacks product margin.
The problem is rarely the provider price alone. It is buying more intelligence than a request needs, paying again when a cheap route fails downstream, and leaving routing unchanged after the market moves.
Premium everywhere
One expensive default model handles routine and difficult requests alike, even when a lower-cost route clears the same customer-visible quality bar.
Cheap on the invoice, expensive in the workflow
Retries, checkers, tool calls, validation, support burden, and human correction can erase the apparent savings from a lower-priced model.
The right route goes stale
Model releases, provider changes, prices, latency, and product requirements can change the economically correct route long after the original architecture decision.
Enterprise AI and analytics teams: the same method also applies when accepted-workflow cost and human review are material, but the Q1 site is intentionally optimized for software vendors.
Evidence, after the offer
The research discipline is a closing asset, not the opening pitch.
Benchmark #1 results are not published yet. Until they are, buyers can inspect the rules that prevent result-shopping. Once results publish, each finding states the economic consequence and the routing decision it changes.
Workloads, acceptance criteria, holdout structure, model-lock rules, and publication conditions are set before testing.
02 · IndependenceThe expensive model does not pay more.Flat fees only, no percentage of model spend or savings, no purchased placement, and relevant vendor relationships are disclosed.
03 · CorrectionsReproducible errors produce visible revisions.Implementation, version, pricing, scoring, and methodology errors can be corrected without erasing the history.
Three outcomes are equally valid.
These are decision patterns, not Benchmark #1 results. No numeric result or winner appears here until the evidence exists.
A lower-cost route clears the frozen quality bar.
- Finding
- Cheaper route meets acceptance and critical-failure requirements.
- Economic consequence
- Premium spend is not buying additional accepted outcomes on this workload.
- Decision
- Route down by default, subject to the documented operating constraints.
Premium intelligence materially changes the accepted outcome.
- Finding
- Lower-cost routes miss the threshold or create unacceptable failures, review, or rework.
- Economic consequence
- The higher model bill is justified by better accepted-workflow economics or materially lower risk.
- Decision
- Keep premium on this workload or on the cases where it demonstrably earns the price.
No single model wins every stage or case.
- Finding
- A default-plus-checker, cascade, or escalation path outperforms one-model-for-everything.
- Economic consequence
- Most work can run cheaply while difficult or high-risk cases receive more intelligence.
- Decision
- Define the default, escalation condition, fallback, cost ceiling, and re-test trigger.
Benchmark #1
Named models. Real analytical workloads. The full cost of an acceptable output.
Classification, extraction, and SQL generation across 8 to 10 model candidates. The quality bar is frozen before model runs, public and held-out results are reported separately, and the published result ends in a routing implication.
AI Analyst Benchmark #1
Which AI workflows actually earn their price on analyst work?
Classification
Acceptance, critical failures, retries, checking, and economics on high-volume decision work.
Extraction
Structured extraction quality, failure severity, escalation cases, validation, and complete accepted-output cost.
SQL generation
Executable-query acceptance, reasoning failures, retries, rework, and when stronger models earn the premium.
Planned model slate and lock rules
Planned scope: 8 to 10 candidates spanning frontier, mid-tier, small, and open-weight options where practical. Immediately before the first run, commercial availability is reverified, exact API model IDs and versions are locked, the quality bar and scoring rules are frozen, and provider pricing is snapshotted. The published benchmark names every model actually tested.
If the result is not defensible by the target date, it publishes later. Credibility outranks calendar compliance.
How the holdout works. Before any model runs, the task set is split about 70% public and 30% held back, both drawn from the same pool. The held-back part is never shown to anyone, is never picked after the results are in, and rotates over time so no vendor can simply train against the test. Details →
Independent by design. Flat fees only. AI Economics Lab never takes a percentage of AI spend or savings, so the commercial incentive does not change when a premium model wins. Relevant vendor relationships are disclosed with the benchmark. Standards → Corrections →
Named models, real numbers, and the failures included. Always free, never gated. One email when there is something worth reading.
Kit will ask you to confirm your email. No sales drip. Unsubscribe any time.
Cost per Acceptable Outcome
Frame the AI decision before model price or prestige frames it for you.
The CPAO Scorecard makes your team define what acceptable means, add up the full cost of each route, and pick the cheapest option that clears the bar. It is the fastest way to see whether a paid benchmark is even necessary.
It is an interactive web page, not a download. It opens immediately after signup. You will also receive new AI Economics Lab benchmarks and major research updates; Kit will ask you to confirm your email. Unsubscribe any time. Forwarding the tool link is fine.
After the Audit, only when justified
Two follow-ons, one for depth and one for keeping the answer current.
The $10K Audit remains the default first paid step. Follow on only when the decision is unusually consequential or when model and price changes can make the validated route stale.
Benchmark Sprint
Deeper commissioned multi-model, multi-workload evidence for a product launch, provider migration, or other high-stakes decision.
Routing Policy
Maintain default route, threshold, escalation, fallback, premium justification, and re-test triggers as models and economics change.
Work with us
Bring me the workload you are least sure about, or the question blocking the decision.
Send the workload or question first. I will reply personally with a straight fit or no-fit answer and the smallest sensible next step.
Useful things to include: what the workload does, roughly what you spend on it now, and what a bad answer costs you. Use the Work with us form and I will reply personally.
Work with us. Send enough context for a useful reply.