AI Engineer & Full-Stack Product Engineer

I build AI products and score them before release.

The eval harness is public and pre-registered. Open to a full-time remote role or a fixed-scope project.

pre-registered eval harness, public repo

0.33 to 0.83

0.330.19 to 0.510.830.66 to 0.93

n = 30

The DirectiveForge README delta table at the public tag v0.20.0-public: the skill activation rate 0.3333 before the fix and 0.8333 after, each with its Wilson 95 % confidence interval over 30 runs, above the grades line and the disclosed regression.

How often one skill, inversion, fired when it should, before the fix and after. A skill is a rule the coding agent runs when the input matches it, and each side is 30 runs on the same frozen test. In the harness: the inversion activation rate with Wilson 95% confidence intervals.

directiveforge/directiveforge, tag v0.20.0-public

For founders and CTOs at SaaS and AI companies whose feature works in the demo and breaks in production.

My own kit shipped with a defect of that kind, and the repair is on the record: a link check that raised 13 false alarms now raises nonepublic repo, DirectiveForge, measured on the same unchanged test before and after. Fixed scope, with the acceptance criteria written before the build. The code is yours to the last line, clean handoff.