Independent AI unit economics for software products

Your AI model mix is a gross margin decision.

AI Economics Lab tests which requests and workflow stages actually need premium models, which can route down safely, and what happens to product margin after retries, validation, rework, and human review are counted.

See how Benchmark #1 works →

Stephen McDaniel. I build and ship production AI systems and previously led analytics and product work at Tableau, Netflix, and SAS. The quality bar is fixed before testing, and the fee does not change when the expensive model wins.

When a $10K Audit is economically sensibleAt $35K/month of all-in product AI cost, every 10% route-cost difference is $3.5K/month.

That is an illustration, not a savings claim. If your workload is materially smaller or the route cannot plausibly change the economics, use the free scorecard first.

The commercial front door

$10K to answer a product routing question with evidence.

Two weeks, 3 to 5 repeatable product workloads. Determine which requests should run on the default model, which should route down, when to escalate, and what the decision does to accepted-output economics.

Start here

AI Economics Audit

Find where premium inference genuinely changes the accepted outcome, where lower-cost routes are already good enough, and where a cheap route becomes expensive once the full workflow is counted.

$10,000
$10,000Your decision first · Two weeks · 3 to 5 workloads · Remote
Step 01Write down what good enough means

A named owner puts the quality bar, the dealbreaker failures, and the operating constraints in writing before anything is tested.

Step 02Test routes, not just models

Your current setup, cheaper alternatives, premium models, and a cheap-model-plus-checker cascade, all judged against the same bar.

Step 03Add up the real cost and time

Model and tool spend, checkers, retries, rework, human review, and the work time it actually takes to reach an answer you can use.

Step 04Make the routing call

Default route, when to escalate, what to fall back to, whether premium is justified, the annual numbers, and the hours your team gets back.

What you leave with
  • A written product quality bar and dealbreaker list, signed off by a named owner
  • A side-by-side route comparison with CPAO for each option
  • A default route, an escalation rule, and a fallback
  • Annualized product AI cost impact and review/rework implications
  • The trigger that tells you when to re-test
  • Every number reproducible from the run record
Illustrative structure, not a client resultWhat the decision page looks like
No fabricated savings or case-study numbers.
Acceptance barWritten threshold + dealbreaker failures
Routes comparedCurrent, cheaper, premium, cascade
Economic resultCPAO + review/rework + time to acceptance
DecisionDefault + escalation + fallback + re-test trigger
What counts as one workload
  • One repeatable job, judged by one acceptance test, owned by one person. Fifty pipelines doing the same job against the same bar count as one workload, not fifty.
  • Examples: text-to-SQL · classification · extraction · AI search or RAG responses · agent stages · request routing · customer-facing summaries.
  • Not a workload: an entire product, a whole department, or "all our AI". If you are unsure how yours splits, use Work with us and I will scope it before you pay anything.

No production access needed. Representative, de-identified samples only. Any commercial relationship with a benchmarked vendor is disclosed with the relevant research.

The Audit buys a decision, not promised savings. If the measured evidence says your current route is already the right one, or that the alternatives are too close to justify a change, that is the answer I deliver.

For AI, analytics, BI, data and software vendors

Three places inference economics quietly attacks product margin.

The problem is rarely the provider price alone. It is buying more intelligence than a request needs, paying again when a cheap route fails downstream, and leaving routing unchanged after the market moves.

01

Premium everywhere

One expensive default model handles routine and difficult requests alike, even when a lower-cost route clears the same customer-visible quality bar.

02

Cheap on the invoice, expensive in the workflow

Retries, checkers, tool calls, validation, support burden, and human correction can erase the apparent savings from a lower-priced model.

03

The right route goes stale

Model releases, provider changes, prices, latency, and product requirements can change the economically correct route long after the original architecture decision.

Strong buying triggers:product launchgross-margin pressuremodel migrationprovider changerouting redesignquality or rework problems

Enterprise AI and analytics teams: the same method also applies when accepted-workflow cost and human review are material, but the Q1 site is intentionally optimized for software vendors.

Evidence, after the offer

The research discipline is a closing asset, not the opening pitch.

Benchmark #1 results are not published yet. Until they are, buyers can inspect the rules that prevent result-shopping. Once results publish, each finding states the economic consequence and the routing decision it changes.

How published findings will be packaged

Three outcomes are equally valid.

These are decision patterns, not Benchmark #1 results. No numeric result or winner appears here until the evidence exists.

Spend less

A lower-cost route clears the frozen quality bar.

Finding
Cheaper route meets acceptance and critical-failure requirements.
Economic consequence
Premium spend is not buying additional accepted outcomes on this workload.
Decision
Route down by default, subject to the documented operating constraints.
Spend more

Premium intelligence materially changes the accepted outcome.

Finding
Lower-cost routes miss the threshold or create unacceptable failures, review, or rework.
Economic consequence
The higher model bill is justified by better accepted-workflow economics or materially lower risk.
Decision
Keep premium on this workload or on the cases where it demonstrably earns the price.
Route intelligently

No single model wins every stage or case.

Finding
A default-plus-checker, cascade, or escalation path outperforms one-model-for-everything.
Economic consequence
Most work can run cheaply while difficult or high-risk cases receive more intelligence.
Decision
Define the default, escalation condition, fallback, cost ceiling, and re-test trigger.
The publishing rule:Finding → economic consequence → decision. If a benchmark number does not change a product, routing, spend, quality, or workflow decision, it is supporting detail, not the headline.Work with us: Run the same decision process on my workloads →

Benchmark #1

Named models. Real analytical workloads. The full cost of an acceptable output.

Classification, extraction, and SQL generation across 8 to 10 model candidates. The quality bar is frozen before model runs, public and held-out results are reported separately, and the published result ends in a routing implication.

AI Analyst Benchmark #1

Which AI workflows actually earn their price on analyst work?

Pre-registered
WORKLOAD 01

Classification

Acceptance, critical failures, retries, checking, and economics on high-volume decision work.

WORKLOAD 02

Extraction

Structured extraction quality, failure severity, escalation cases, validation, and complete accepted-output cost.

WORKLOAD 03

SQL generation

Executable-query acceptance, reasoning failures, retries, rework, and when stronger models earn the premium.

Bar set firstQuality threshold and dealbreaker failures frozen before any model runs.
8 to 10 model candidatesFrontier, mid-tier, small, and open-weight options where practical. Exact IDs lock before testing.
3 route strategiesCheaper model, premium model, and cheaper model with a checking step.
Full cost and timeAI, tools, retries, rework, checking, human review, and time to a usable answer.
Benchmark #1 pre-registrationRead the full research standards →
Plan postedAugust 10, 2026
Target publishSeptember 2026*
Public / holdoutAbout 70% / 30%, same task pool
Planned model slate and lock rules

Planned scope: 8 to 10 candidates spanning frontier, mid-tier, small, and open-weight options where practical. Immediately before the first run, commercial availability is reverified, exact API model IDs and versions are locked, the quality bar and scoring rules are frozen, and provider pricing is snapshotted. The published benchmark names every model actually tested.

If the result is not defensible by the target date, it publishes later. Credibility outranks calendar compliance.

How the holdout works. Before any model runs, the task set is split about 70% public and 30% held back, both drawn from the same pool. The held-back part is never shown to anyone, is never picked after the results are in, and rotates over time so no vendor can simply train against the test. Details →

Independent by design. Flat fees only. AI Economics Lab never takes a percentage of AI spend or savings, so the commercial incentive does not change when a premium model wins. Relevant vendor relationships are disclosed with the benchmark. Standards → Corrections →

Want this run on your own workloads?Work with us: See the $10K Audit →
Get each benchmark the day it publishes.

Named models, real numbers, and the failures included. Always free, never gated. One email when there is something worth reading.

    Kit will ask you to confirm your email. No sales drip. Unsubscribe any time.

    Free working tool

    Cost per Acceptable Outcome

    Frame the AI decision before model price or prestige frames it for you.

    The CPAO Scorecard makes your team define what acceptable means, add up the full cost of each route, and pick the cheapest option that clears the bar. It is the fastest way to see whether a paid benchmark is even necessary.

      It is an interactive web page, not a download. It opens immediately after signup. You will also receive new AI Economics Lab benchmarks and major research updates; Kit will ask you to confirm your email. Unsubscribe any time. Forwarding the tool link is fine.

      AI Economics Lab
      Use the lowest-CPAO option that clears the agreed quality bar and every operating constraint.
      Define acceptableCompare candidatesCalculate CPAOMake the call
      Interactive · printable · built for the decision meeting

      After the Audit, only when justified

      Two follow-ons, one for depth and one for keeping the answer current.

      The $10K Audit remains the default first paid step. Follow on only when the decision is unusually consequential or when model and price changes can make the validated route stale.

      $35K

      Benchmark Sprint

      Deeper commissioned multi-model, multi-workload evidence for a product launch, provider migration, or other high-stakes decision.

      $20K to $35K/year

      Routing Policy

      Maintain default route, threshold, escalation, fallback, premium justification, and re-test triggers as models and economics change.

      Work with us

      Bring me the workload you are least sure about, or the question blocking the decision.

      Send the workload or question first. I will reply personally with a straight fit or no-fit answer and the smallest sensible next step.

      Useful things to include: what the workload does, roughly what you spend on it now, and what a bad answer costs you. Use the Work with us form and I will reply personally.

      Work with us. Send enough context for a useful reply.

      I reply personally, usually within one business day.