Score your AI feature before release

For the founder or CTO whose AI feature works in the demo and breaks in production. I build a test harness around your feature, fix the inputs and the scoring rule before the first run, and hand you the measured baseline, the failure modes, the token spend and a written plan for what to repair first. The harness stays in your repository and your engineers can re-run it.

Acceptance criterion, written into the scope

The harness runs on your machine and reproduces the baseline I report, on the same fixed inputs.

public repo, tag v0.20.0-public

22 measured skills, each published with a trigger F1 and an activation rate at Wilson 95 % confidence intervals, n = 30

Eval harness report output listing per-skill trigger F1 and activation rates with confidence intervals

The bar this work has to clear

You approve the inputs, the scoring rule and this list before the first run. An eval harness is a test suite for model behavior: fixed inputs, a scoring rule, and results anyone can re-run. Pre-registered means the specification is committed to git before any result is collected, so the bar cannot move after the fact.

  • The harness runs on your machine and reproduces the baseline I report, on the same fixed inputs.
  • The scoring rule and the input set are committed before the first scored run, with the git history showing that order.
  • Every failing case is kept with its input and its output, grouped into a named failure mode.
  • The harness, the results and the fix plan are handed over in full, yours to the last line.

What you know at the end that you do not know now

  • A number for how often the feature is right

    The same inputs every run, one scoring rule agreed with you, and a result anyone on your team can reproduce.

    • Inputs drawn from your real traffic or written with you, then frozen
    • One scoring rule, written down before any result is collected
    • Each score published with the interval it could move in, so nothing rests on one lucky run
    • The whole run reproducible on your machine, not only on mine
  • The list of ways it fails, ranked

    Every failed case is kept with its input and its output, grouped into failure modes and counted, so the repair work has an order.

    • Failure modes named and counted, not described in the abstract
    • The failing cases kept beside the harness as regression tests
    • Refusals, hallucinated facts, format breaks and timeouts counted separately
    • The cases where the feature was right for the wrong reason, flagged
  • What the feature costs to run

    Token spend is read per flow and per call, so the price of an answer is a number in the report instead of a surprise on the invoice.

    • Cost per call, broken down by flow
    • Where the prompt spends tokens it does not need
    • The cost effect of each fix in the plan, estimated before you pay for it
    • A cheaper model tested against the same harness when one is worth testing
  • A fix plan you can hand to your own engineers

    The plan is ordered by what the numbers say to repair first, and each item names the score it should move.

    • Each fix tied to the failure mode it addresses
    • The expected effect on the baseline, stated before the work
    • What I would not change, and why
    • The plan is yours to run in-house or to hand back to me

What lands in your repository

  • The eval harness, in your repository, with the input set and the scoring rule committed before the first scored run.
  • The baseline report: the score per flow, each with the interval it could move in.
  • The failure-mode list, counted and ranked, with the failing cases kept as regression tests.
  • The token-spend read: cost per call and per flow, and where the prompt wastes tokens.
  • The written fix plan, ordered by the numbers, with what I would not change and why.
  • A walkthrough session with the engineers who will re-run it.

What I would not build

  • A score for a feature that does not exist yet
  • A benchmark against another company's product
  • An evaluation nobody intends to act on
  • A scoring rule nobody can approve, so the criterion has no owner
  • Work in crypto, gambling or adult

Who this is for

You are a founder, a CTO or a head of product. The AI feature demos well, and nobody on your side can say how often it is right. Support forwards the bad answers one at a time. Somebody asks what it costs per user and the answer is a shrug. Shipping it feels like a coin flip, so it does not ship. This work replaces the shrug with a number, on your feature, with the rule that produced the number written down first.

How the work runs

We start by agreeing what counts as a correct answer. That sounds obvious and it is where most of the value is: I write the flows, the inputs and the scoring rule into a short spec, you approve it, and it goes into git before a single result exists. That order is the whole point. A bar written after the results is a bar that moved.

Then I build the harness around your feature and run it. Every case is kept: the input, the output, the score. Failures get grouped into named modes and counted, so "it hallucinates sometimes" becomes a rate you can argue about. The same run reads the token spend per flow, which is usually where the surprise lives.

You get the baseline, the failure list and a fix plan ordered by the numbers. The harness is in your repository, and your own engineers re-run it after each change. If you then want me to do the repairs, that is a separate scope, agreed the same way.

Best fit

Best fit if the feature already runs somewhere, even badly, and someone on your side can say what a correct answer looks like.

What was measured on the last builds

DirectiveForge is my own eval harness, MIT licensed and public. The specification was committed before the results and the git history shows it. 22 skills each carry a published score for how reliably they fire, with Wilson 95 % confidence intervals at 30 trials, and one regression is published next to the gains. The repair pattern is the same one I run for clients: the failing grade goes out first, then the cause, then the re-run on the unchanged test. A link check that raised 13 false alarms now raises none. Read the DirectiveForge case.

At Dr.Acula I built the model orchestration and the backend. Token spend fell 45 % in one month, read off the model billing dashboard. The app is live in the App Store; my active build phase there ran from January to June 2025. Read the Dr.Acula case.

If you are hiring

I am also open to a full-time remote role. If that is what you are staffing, send me the details.

The cases behind these numbers

Tell me which feature you cannot vouch for

Send the flow and what a correct answer looks like to you. I come back with the scope, the scoring rule and what you own at the end.