Fixed-scope work, scored before it ships.
I build AI products and score them before release. The eval harness is public and pre-registered.
Where to start, and the bar each one has to clear
For a feature that works in the demo and breaks in production
Score your AI feature before release
The harness runs on your machine and reproduces the baseline I report, on the same fixed inputs.

A link check that raised 13 false alarms now raises none, on the same unchanged test
I build a harness on your feature and hand back the measured baseline, the ranked failure modes and a written fix plan.
For a team with a working demo and no production path
One AI feature, shipped and scored
Happy-path success at or above 95 % on the flows named in the scope.

1,000+ sites generated and build time down 40 %, on the AI website builder I led at Name.am
One AI feature shipped to your users, with checks that stop bad answers and the eval report at handover.
For a founder who needs a product built, not a team to manage
The whole product, one owner
The gates named in the scope block a merge that fails them, and they pass on the build that ships.

1,059 commits since March 2026 on a commerce platform I own and still run
I build your web or AI product end to end, including payments and the content layer your team edits without me.
Before you write
What would you not build?
Crypto and gambling front ends, adult content, and open-ended internal admin systems where the scope gets discovered as we go. I also pass on work where nobody on your side can approve a scope, because then the acceptance criterion has no owner.
How is a scope written?
We talk once about what you are trying to ship and the deadline behind it. Then I write the scope: what gets built, what is out, the criterion it has to meet, and what you own at the end. You approve it in writing before any code. A number quoted before that conversation would be a placeholder.
Who owns the code?
You do. The code is yours to the last line, with a clean handoff: the repository, the environment variables documented, and a walkthrough of how it runs and how to deploy it.
How is the work measured?
Every project carries an acceptance criterion written into the scope before the build. For an AI feature that is 95 % or better on the happy paths, meaning the flows we listed in the scope. For an eval audit it is that the harness re-runs on your machine and reproduces the baseline I reported, on the same fixed inputs. For a whole product it is that the checks named in the scope block a merge that fails them and pass on the build that ships. I run the measurement, commit the result, and you can re-run it.
What happens if the criterion is not met?
The work is not done. I keep going until the bar is met, or we renegotiate the scope in writing. The criterion is a target agreed before the build. It is a standard for the thing I deliver, never a promise about your rankings or your revenue.
Do you work with an existing codebase?
Yes. The eval audit is written for a feature already running, and an AI feature goes into the stack you have. I read the repository before I scope anything, and I say so when a rewrite is the honest answer instead of a patch.