Research standards

Threshold first. Price second. Route by evidence.

Every benchmark I publish is public evidence, so the rules behind it should be public too. This page says exactly how the evidence gets made, how you can challenge it, how corrections work, and what keeps it independent. If I break one of these rules, hold me to it.

Pre-registration

AI Analyst Benchmark #1

The first benchmark stays deliberately narrow: classification and SQL generation. I compare single models against a cheap-model-plus-checker setup, judged on Cost per Acceptable Outcome. Narrow and defensible beats broad and shaky.

Plan postedAugust 10, 2026
Target publicationSeptember 2026, only if the result is defensible
WorkloadsClassification; SQL generation
Public / holdoutAbout 70% / 30%, both drawn from the same task pool
Primary thresholdsLocked per workload before any result is known
Threshold sensitivityReported at 90%, 95%, and 99% wherever the scoring scale makes that meaningful
Split manifest hashPosted here before the first model runs
Pricing snapshotProvider pricing pages captured and dated immediately before the run

Planned model candidates as of August 11, 2026, not yet the tested slate: eight candidates across OpenAI, Anthropic, Google, and Mistral, spanning frontier, mid-tier, and lower-cost options. Immediately before testing starts I reverify commercial availability and then lock the exact API model IDs and versions, the scoring rules, what counts as acceptable, the dealbreaker failures, the task split, and the price snapshot. The published benchmark names every model actually tested.

If a model becomes unavailable or something in the build forces a change before the run, I log the change and the reason. The held-back task set itself is never published.

Methodology

How I run a benchmark

  • Real tasks first. I build the task set before I know which model would look good on it.
  • Set the bar before testing. Scoring, the minimum quality that counts as acceptable, and the dealbreaker failures are all fixed before any result comes in.
  • Compare routes, not just model names. A single model and a multi-step workflow can both enter, as long as they answer the same question.
  • Count the whole bill. Model and API spend, tools and checkers, retries, rework, and human review all count when they affect the decision.
  • Report what breaks. Failure types and how bad they are get published next to the pass rates, not buried.
  • Keep public and held-back results separate. The two never get blended into one flattering headline number.
  • Keep the receipts. Model IDs and versions, provider, settings, build details, and dated pricing sources are all kept with the benchmark.

Holdout policy

Some of the evidence stays unseen.

  • Before testing, the task set is split about 70% public and 30% held back, both drawn from the same pool.
  • The split is timestamped, and a hash of it is published before the first model runs, so nobody can claim I reshuffled it later.
  • The held-back set is never shown to vendors, sponsors, clients, or the public.
  • Public and held-back scores are always published separately.
  • When a benchmark is refreshed, only part of the held-back set moves into public, and fresh unseen tasks replace it.

Corrections and right of reply

If I get it wrong, I publish the fix.

  • Send me a reproducible challenge on a build error, the wrong model or version, a pricing mistake, inconsistent scoring, or a flaw in the method, and I will work through it with you.
  • Not liking the result is not a reason for a rerun. Show me the error.
  • When a real error is confirmed, I publish the original number, the corrected number, the date, and the reason. The old version stays visible.
DateBenchmarkRevision
2026-08-10Benchmark #1Pre-registration posted. No benchmark result published yet.

Independence and disclosure

The expensive model does not pay me more.

  • I never take a percentage of your model spend or your savings. Flat fees only, so a cheap recommendation costs me nothing.
  • No payment buys placement, a better score, a buried result, or editorial control.
  • If I have a commercial relationship with a vendor in a benchmark, it is disclosed on that benchmark, not hidden in a policy page.
  • Commissioned work is clearly labeled and kept visually separate from research I started myself.
  • A sponsor can correct a factual detail. A sponsor cannot edit a conclusion or see the held-back tasks.
  • The quality bar and the scoring never change after results are in, in either direction.
  • About YakData. I run an implementation practice, YakData, under the same parent company. It keeps me working in production, which makes the research better. Where YakData work overlaps with a benchmarked vendor or platform, that overlap is disclosed on the benchmark itself.