Open to a full-time remote role or a fixed-scope project.

Fixed-scope work, scored before it ships.

I build AI products and score them before release. The eval harness is public and pre-registered.

Hiring for a full-time remote role

Fixed-scope work

Where to start, and the bar each one has to clear

The three stand alone, and they run in order: measure what you have, ship one feature and measure it, or hand over the product. The scope and the bar are approved before the first commit.

For a feature that works in the demo and breaks in production

Score your AI feature before release

Acceptance criterion, written into the scope

The harness runs on your machine and reproduces the baseline I report, on the same fixed inputs.

Eval harness report output listing per-skill trigger F1 and activation rates with confidence intervals

public repo, DirectiveForge

A link check that raised 13 false alarms now raises none, on the same unchanged test

I build a harness on your feature and hand back the measured baseline, the ranked failure modes and a written fix plan.

  • The harness itself, in your repository, with the inputs and the scoring rule fixed before the first run
  • The measured baseline, each score published with the interval it could move in, so nothing rests on a single lucky run
  • The failure-mode list, ranked by how often each one fired, with the token spend behind them
  • A written fix plan, ordered by what the numbers say to repair first

For a team with a working demo and no production path

One AI feature, shipped and scored

Acceptance criterion, written into the scope

Happy-path success at or above 95 % on the flows named in the scope.

The name.am home page: the domain search box, the price list per top-level domain, and the mobile app badge, in Armenian.

Name.am's figures, private

1,000+ sites generated and build time down 40 %, on the AI website builder I led at Name.am

One AI feature shipped to your users, with checks that stop bad answers and the eval report at handover.

  • Scope, flows and the event plan written and approved before the first commit
  • Guardrails, the rules that block unsafe or wrong output before a user sees it
  • Named analytics events in the stack you already run, so the feature stays countable after launch
  • The eval report at handover, on fixed inputs and one scoring rule your team can re-run

For a founder who needs a product built, not a team to manage

The whole product, one owner

Acceptance criterion, written into the scope

The gates named in the scope block a merge that fails them, and they pass on the build that ships.

Public donation form with donation frequency, preset amounts in Armenian drams, and a custom amount field

private repo, commit count

1,059 commits since March 2026 on a commerce platform I own and still run

I build your web or AI product end to end, including payments and the content layer your team edits without me.

  • One scope, written first, carrying the acceptance criterion and a named approver on your side
  • The build in your repository: schema, API, interface, payments, and the content layer your team edits without me
  • Lighthouse budgets, accessibility checks and structured-data validation wired into CI, so a later commit cannot quietly undo them
  • A handoff document and a walkthrough with the engineers who will run it

An eval harness is a test suite for model behavior: fixed inputs, a scoring rule, and results anyone can re-run.

Before you write

The questions that come up before a first message, answered here rather than one reply at a time.

What would you not build?

Crypto and gambling front ends, adult content, and open-ended internal admin systems where the scope gets discovered as we go. I also pass on work where nobody on your side can approve a scope, because then the acceptance criterion has no owner.

How is a scope written?

We talk once about what you are trying to ship and the deadline behind it. Then I write the scope: what gets built, what is out, the criterion it has to meet, and what you own at the end. You approve it in writing before any code. A number quoted before that conversation would be a placeholder.

Who owns the code?

You do. The code is yours to the last line, with a clean handoff: the repository, the environment variables documented, and a walkthrough of how it runs and how to deploy it.

How is the work measured?

Every project carries an acceptance criterion written into the scope before the build. For an AI feature that is 95 % or better on the happy paths, meaning the flows we listed in the scope. For an eval audit it is that the harness re-runs on your machine and reproduces the baseline I reported, on the same fixed inputs. For a whole product it is that the checks named in the scope block a merge that fails them and pass on the build that ships. I run the measurement, commit the result, and you can re-run it.

What happens if the criterion is not met?

The work is not done. I keep going until the bar is met, or we renegotiate the scope in writing. The criterion is a target agreed before the build. It is a standard for the thing I deliver, never a promise about your rankings or your revenue.

Do you work with an existing codebase?

Yes. The eval audit is written for a feature already running, and an AI feature goes into the stack you have. I read the repository before I scope anything, and I say so when a rewrite is the honest answer instead of a patch.

Start with the scope

Send what you are trying to ship and the deadline behind it. I come back with the scope, the bar it has to clear and what you own at the end. You approve that in writing before any code is written. If you are not sure which of the three fits, say what breaks today and I will tell you.

The code is yours to the last line, with a clean handoff.