Score your AI feature before release
For the founder or CTO whose AI feature works in the demo and breaks in production. I build a test harness around your feature, fix the inputs and the scoring rule before the first run, and hand you the measured baseline, the failure modes, the token spend and a written plan for what to repair first. The harness stays in your repository and your engineers can re-run it.
The harness runs on your machine and reproduces the baseline I report, on the same fixed inputs.
22 measured skills, each published with a trigger F1 and an activation rate at Wilson 95 % confidence intervals, n = 30

The bar this work has to clear
You approve the inputs, the scoring rule and this list before the first run. An eval harness is a test suite for model behavior: fixed inputs, a scoring rule, and results anyone can re-run. Pre-registered means the specification is committed to git before any result is collected, so the bar cannot move after the fact.
- The harness runs on your machine and reproduces the baseline I report, on the same fixed inputs.
- The scoring rule and the input set are committed before the first scored run, with the git history showing that order.
- Every failing case is kept with its input and its output, grouped into a named failure mode.
- The harness, the results and the fix plan are handed over in full, yours to the last line.
What you know at the end that you do not know now
A number for how often the feature is right
The list of ways it fails, ranked
What the feature costs to run
A fix plan you can hand to your own engineers
What lands in your repository
- The eval harness, in your repository, with the input set and the scoring rule committed before the first scored run.
- The baseline report: the score per flow, each with the interval it could move in.
- The failure-mode list, counted and ranked, with the failing cases kept as regression tests.
- The token-spend read: cost per call and per flow, and where the prompt wastes tokens.
- The written fix plan, ordered by the numbers, with what I would not change and why.
- A walkthrough session with the engineers who will re-run it.
What I would not build
Who this is for
You are a founder, a CTO or a head of product. The AI feature demos well, and nobody on your side can say how often it is right. Support forwards the bad answers one at a time. Somebody asks what it costs per user and the answer is a shrug. Shipping it feels like a coin flip, so it does not ship. This work replaces the shrug with a number, on your feature, with the rule that produced the number written down first.
How the work runs
We start by agreeing what counts as a correct answer. That sounds obvious and it is where most of the value is: I write the flows, the inputs and the scoring rule into a short spec, you approve it, and it goes into git before a single result exists. That order is the whole point. A bar written after the results is a bar that moved.
Then I build the harness around your feature and run it. Every case is kept: the input, the output, the score. Failures get grouped into named modes and counted, so "it hallucinates sometimes" becomes a rate you can argue about. The same run reads the token spend per flow, which is usually where the surprise lives.
You get the baseline, the failure list and a fix plan ordered by the numbers. The harness is in your repository, and your own engineers re-run it after each change. If you then want me to do the repairs, that is a separate scope, agreed the same way.
Best fit
Best fit if the feature already runs somewhere, even badly, and someone on your side can say what a correct answer looks like.
What was measured on the last builds
DirectiveForge is my own eval harness, MIT licensed and public. The specification was committed before the results and the git history shows it. 22 skills each carry a published score for how reliably they fire, with Wilson 95 % confidence intervals at 30 trials, and one regression is published next to the gains. The repair pattern is the same one I run for clients: the failing grade goes out first, then the cause, then the re-run on the unchanged test. A link check that raised 13 false alarms now raises none. Read the DirectiveForge case.
At Dr.Acula I built the model orchestration and the backend. Token spend fell 45 % in one month, read off the model billing dashboard. The app is live in the App Store; my active build phase there ran from January to June 2025. Read the Dr.Acula case.
If you are hiring
I am also open to a full-time remote role. If that is what you are staffing, send me the details.
The cases behind these numbers
A link check that raised 13 false alarms now raises none, on the same unchanged test
Token spend down 45 % in one month at Dr.Acula, read off the model billing dashboard
Tell me which feature you cannot vouch for
Send the flow and what a correct answer looks like to you. I come back with the scope, the scoring rule and what you own at the end.