AI Engineer & Full-Stack Product Engineer

Samvel Avagyan

I build AI products and score them before release. The eval harness is public and pre-registered.

Open to a full-time remote role or a fixed-scope project.

Yerevan, ArmeniaRemote, UTC+4

dated roles, this page

7 years on the web, 3 shipping LLM products

the work history below

case studies, this site

5+ AI products in production

the case studies

public repo, tag v0.20.0-public

Skill activation 0.33 to 0.83, n = 30

the repository at the tag

What I build, and how it gets measured

I'm Samvel Avagyan. I build AI products end to end and score them before release. I work from Yerevan, Armenia, on UTC+4, through Samwell.dev, the studio that has carried my client work since August 2024dated roles, this page. Seven years of that work has been on the web, the last three shipping LLM products.

The scoring is public. DirectiveForge is my eval harness for AI workflows: MIT licensed, pre-registered so the spec is committed before the results, and readable at github.com/directiveforge/directiveforge. It measures 22 skillspublic repo, tag v0.20.0-public on trigger F1, which counts how often a skill fires when it should and stays quiet when it should not, and it publishes a Wilson 95% confidence interval for each one at n = 30. One regression is published next to the gains.

The same habit runs through the client work. I built Teach For Armenia a donation platform on Fastify and PostgreSQL, with AmeriaBank vPOS 3.1 and AES-256-GCM encryption at rest. It runs in production. The Name.am builder I led generated more than 1,000 sitesName.am case study and took 40% off build time. At Dr.Acula, token spend fell 45%billing dashboard, private in one month on the model billing dashboard, and the landing page was rebuilt to a performance budget.

The same work is available as an eval audit of a feature you already ship, one AI feature shipped and scored and a whole product built end to end. The project case studies carry each number with the source it came from.

View the work history

See all projects

What each tool was used on, and what was counted

Every line names the system the tool ran on and the count taken there, with the date of the count. The link opens the case study that carries the same figure.

What the services are written in

TypeScriptNode.jsNestJS

Interface and quality gates

Next.jsnext-intlCore Web VitalsTechnical SEO / JSON-LD

Payments, stored data and the content clients edit

FastifyPostgreSQLDockerPayment integrationSanity v5

Model work and the checks on it

Evals and LLMOpsAnthropic ClaudeDeepSeekAgents and orchestrationMCPAnti-hallucination gatesPrompt engineeringSSE streamingClaude Code / Cursor workflow engineering

Also in production use: JavaScript, Python, React, Tailwind CSS, SCSS, Redux, Framer Motion, Three.js and React Three Fiber, Playwright, Axe and WCAG 2.2 AA, Express, Redis, MongoDB, Vercel, Railway, Medusa v2, Payload CMS v3, OpenAI, Gemini, RAG and embeddings, and model routing.

Work history

Employment

  • Built the AI website generation platform from scratch, including a modular prompt system
  • Integrated the DeepSeek API with SSE streaming so pages arrive token by token
  • Generated 1,000+ sites and took 40% off build timeName.am case study
Explore the AI website builder

Clients via Samwell.dev

Samwell.dev is my studio; I led each of these engagements, with a team of up to four.

  • Ran SEO and generative-engine optimization on the existing sites from March 2025
  • Built the new Teach For Armenia website: 16 pages, 2 locales, 61 Sanity schemasproject repo, commit a86a314
  • Built the donation platform solo on Fastify and PostgreSQL with AmeriaBank vPOS 3.1 and AES-256-GCM at rest, launched in 2026 and running in production
  • Built the Armenia Education Initiative site on Payload CMS v3: 12 EN/HY pages and a 5-step application wizardclient site, private
Read the Teach For Armenia case study

Own products

  • Created an open-source (MIT) AI-workflow generator for Claude Code and Cursor
  • Built a pre-registered eval harness measuring trigger F1 and skill activation with Wilson 95% CIs at n = 30public repo, tag v0.20.0-public
  • Calibrated an LLM-as-judge against a human answer key (18/20 exact) before trusting its verdictspublic repo, tag v0.20.0-public
  • Published the before and after deltas, including one disclosed regression (F1 0.9091 to 0.8889)public repo, tag v0.20.0-public
Read the DirectiveForge case study

Education

  • Completed the first year; currently enrolled.own record

Open to a full-time remote role or a fixed-scope project. Remote, UTC+4.