Back to blog

September 30, 2026 · 6 min read

Polished does not mean correct

How I build with AI. Research first, a goal I keep rewriting, fast laps, and a cleanup pass at the end. Tokens go to intent, not retries.

AI workflowProcessClaude Code

Building fast with AI is the easy part. The hard part is making sure what gets built is what was actually needed. So I spend tokens on intent and context instead of retries. I stay at the one step only I can do. And I change the goal as I learn.

It's the design process, with an AI doing the building.

Polished does not mean correct

Left: a pixelated photo of Barack Obama. Right: the AI's sharpened result, a clear, realistic face of a different man.
The Illusion. PULSE face upscaler (Duke University, 2020), given a pixelated photo of Barack Obama.

This model was asked to sharpen a blurry photo. It returned a clean, convincing face of the wrong person. It measured one thing, "does this look like a real face," and it passed perfectly. It had no idea who it was supposed to be, so it filled the gap with its defaults.

Autonomous loops do the same thing to software. They keep rerunning until a check passes and the output looks polished. Polish is easy to measure. Intent is not.

The illusion only works on someone who doesn't know the goal. So the person who knows the goal stays in the loop.

Two ways to spend tokens

Loop until it passesIntent first (how I work)
The goalFixed at the startRewritten as I learn
Done meansA check the machine can run passesIt does what the written goal says
Tokens go toRetries, and rereading the codebase every passContext once, then small targeted changes
Wrong is caughtOnly if the check happens to measure itAt "Try it," by the person who knows the goal
When it's doneStops when the check passesA cleanup pass removes what the laps left behind
You end up withPolished and plausible, sometimes the wrong thingRough at first, then the right thing

Loops have their place. When a machine can check "done," like tests passing on a big refactor, let it run. Most design work isn't like that.

The loop

I leadAI leadsOne proposes, the other decidesDEFINE matching design phase
PROJECT CONTEXTEvery session starts cold and reads these firstCLAUDE.mdspec.mdconventions.mdimplementation.mdskills/evals/changelog.mdSmall on purpose. Skills load only when needed.1 · MEDEFINEDefine goalCriteria + out of scope2 · AI → MEIDEATEPlanCheap to redirect here3 · AIPROTOTYPEBuildThe approved plan only4 · AIQACheckTests + build must pass8 · AIRELEASEShipVersion, changelog, push7 · AIDOCUMENTRecordBuilt + how checked6 · ME → AISYSTEMCompoundEach fix becomes a rule5 · METESTTry itWhat did I learn?approveall passfixesnext featurefails: AI fixes and rerunsLearned something: rewrite the goal, back to step 1goalreadslogrules, skillschangelogSTART · MEDISCOVERResearchUsers + requirementsgoal metall checks passFINISH · AI → MEClean upCut what isn't neededMake it run fastTidy the codeTighten the designREFINELaunch
Research sets the first goal. The red loop is automated. The orange loop is me: I use it, learn, rewrite the goal, and go around again.

Research comes first: users, requirements, a first goal. Then one feature at a time goes around the loop.

  1. Define goal. I write what I want, the criteria, and what's out of scope. On later laps I rewrite it with what I learned.
  2. Plan. The AI proposes an approach before writing code. A wrong plan costs 30 seconds. Wrong code costs an hour.
  3. Build. The approved plan, nothing more.
  4. Check. The AI runs the tests and the build before it says it's done. If something fails, it fixes it and reruns. Broken work never costs my attention.
  5. Try it. I use it the way a real person would. This is where I learn what the goal should have been.
  6. Compound. Every correction becomes a rule or a skill, so the same mistake never costs twice.
  7. Record. The AI logs what got built and how it was checked.
  8. Ship. Version it, write what changed, push.

When the goal is met and every check passes, the work leaves the loop for one cleanup pass. Cut what isn't needed. Make it run fast. Tidy the code and the design. Then launch.

It's the design process

Designers already work this way. Testing changes the problem, not just the solution. I run the same loop with an AI building each prototype. The difference is speed, so I learn more per day and the goal sharpens faster.

StepDesign phaseWhat it means here
ResearchDiscoverTalk to users, gather requirements, write the first goal.
1 Define goalDefineFrame the problem. On later laps, reframe it with what testing taught me.
2 PlanIdeateThe AI proposes an approach. I pick or redirect before anything gets built.
3 BuildPrototypeA working prototype, at the fidelity the question needs.
4 CheckQAAutomated checks catch broken work before it reaches a person.
5 Try itTestUse it against the real goal.
6 CompoundSystemPatterns that worked become rules and skills, like components in a design system.
7 RecordDocumentDecisions and state, written down so the next session starts caught up.
8 ShipReleaseVersion it, write what changed, put it in front of people.
Clean upRefineCut what isn't needed, make it run well, tidy the code and the design.

The context files

Every AI session starts cold. A handful of small files at the project root catch it up, so it doesn't have to read the whole codebase.

FileWritten byJobWhy it saves tokens
CLAUDE.mdMeStanding orders: how to use the other filesUnder ~100 lines. Points to files instead of repeating them.
spec.mdMeThe goal, audience, decisions. Rewritten as I learn.The goal is written once, so it never has to be re-explained.
implementation.mdAIWhat's built and how it was checkedThe agent reads one page instead of the codebase.
conventions.mdMeHouse style, grown from correctionsA mistake gets paid for once.
skills/BothReusable know-howLoads only when the task needs it.
evals/MeGood and bad examples for AI featuresWritten by someone who knows the goal, so polish alone can't pass.
changelog.mdAIWhat changed, written for peopleCatching up costs one read.

Two habits that keep it honest

The audit. Every few features, the AI compares the spec, the implementation log and the actual code, then runs the checks. It reports gaps and stale rules. It doesn't fix them. I decide which side is wrong, because I'm the one who knows the goal.

Evals. Whenever AI writes something a customer will see, I keep 10 to 20 real inputs with examples of good and bad output, and run them every time a prompt or skill changes. And I design for when it's wrong: drafts instead of sends, a person reviews before anything goes live, undo, and showing where an answer came from.

In one breath

Anyone can loop an AI until the result looks polished. I spend tokens on learning what we're actually trying to build, and I let the goal change when I learn it.