Skip to content

ResearchNote

AFHL: Agent-first, Human-last engineering

A way of building software in which agents plan, implement, verify, and record the work, and a person comes in last to try the real thing and judge its UI and UX. Based on the ten days in which we started five apps.

  • Engineering
  • Agents

AFHL (Agent-first, Human-last engineering) is our default way of building apps. Agents plan, implement, verify, review their own work, and keep the record. A person comes in at the end, uses the app on a real device, and judges the UI and UX.

Why a person comes last

In a small company, the scarcest resource is human attention. Agents can read and write code, run builds, and review their own changes. What they can’t settle is whether something feels right in the hand. That takes a person actually using it. So we spend human attention on that final judgment. At Our World, it falls to our founder, who has a background in design.

Two modes

  • Iteration (the default). Agents focus on the experience: the UI, how it feels to use, how quickly it responds, and how smoothly it moves. Broad cleanup and exhaustive testing are put off, and anything deferred goes into a refinement backlog.
  • Refinement (only when asked). Agents clean up the design, the code, and the tests without changing the agreed experience. When that’s done, work returns to Iteration.

We separated the two because asking for both at once tends to leave both half-done.

Check-ins along the way

Human-last doesn’t mean the person stays away until the end. Before building a new screen, the agent shows a preview image and the person confirms the direction.

A person trying the app never replaces the agent’s own checks. Reports keep “the build passed” separate from “a person used it and confirmed it,” and anything not yet checked is reported as unchecked.

In practice

Over ten days, from September 19 to 28, 2026, we started five apps this way. The first one we plan to release is SnapCap, which takes a photo with your iPhone’s camera and puts it on your Mac’s screen in one step.

The bigger change was not speed but where the human time went. Our founder no longer reviews code day to day, and instead spends that time using the apps and telling the agents what to fix.

Limits

Judging the experience comes down to one person’s eye, and we haven’t yet worked out how to share that standard as the team grows. Deferred tests and cleanup can also pile up, and deciding when to switch to Refinement is still a human call.

Related

Version history

  1. v1.0Sep 30, 2026Added what we learned from starting five apps in ten days, and marked the note stable.
  2. v0.2Sep 9, 2026Split the process into two modes, Iteration and Refinement.
  3. v0.1Aug 18, 2026First version: how work is divided between agents and a person.

This document’s URL will not change. When citing it, please include the version number.

All research