September 30, 2026 · 6 min read
Polished does not mean correct
How I build with AI. Research first, a goal I keep rewriting, fast laps, and a cleanup pass at the end. Tokens go to intent, not retries.
Building fast with AI is the easy part. The hard part is making sure what gets built is what was actually needed. So I spend tokens on intent and context instead of retries. I stay at the one step only I can do. And I change the goal as I learn.
It's the design process, with an AI doing the building.
Polished does not mean correct

This model was asked to sharpen a blurry photo. It returned a clean, convincing face of the wrong person. It measured one thing, "does this look like a real face," and it passed perfectly. It had no idea who it was supposed to be, so it filled the gap with its defaults.
Autonomous loops do the same thing to software. They keep rerunning until a check passes and the output looks polished. Polish is easy to measure. Intent is not.
The illusion only works on someone who doesn't know the goal. So the person who knows the goal stays in the loop.
Two ways to spend tokens
| Loop until it passes | Intent first (how I work) | |
|---|---|---|
| The goal | Fixed at the start | Rewritten as I learn |
| Done means | A check the machine can run passes | It does what the written goal says |
| Tokens go to | Retries, and rereading the codebase every pass | Context once, then small targeted changes |
| Wrong is caught | Only if the check happens to measure it | At "Try it," by the person who knows the goal |
| When it's done | Stops when the check passes | A cleanup pass removes what the laps left behind |
| You end up with | Polished and plausible, sometimes the wrong thing | Rough at first, then the right thing |
Loops have their place. When a machine can check "done," like tests passing on a big refactor, let it run. Most design work isn't like that.
The loop
Research comes first: users, requirements, a first goal. Then one feature at a time goes around the loop.
- Define goal. I write what I want, the criteria, and what's out of scope. On later laps I rewrite it with what I learned.
- Plan. The AI proposes an approach before writing code. A wrong plan costs 30 seconds. Wrong code costs an hour.
- Build. The approved plan, nothing more.
- Check. The AI runs the tests and the build before it says it's done. If something fails, it fixes it and reruns. Broken work never costs my attention.
- Try it. I use it the way a real person would. This is where I learn what the goal should have been.
- Compound. Every correction becomes a rule or a skill, so the same mistake never costs twice.
- Record. The AI logs what got built and how it was checked.
- Ship. Version it, write what changed, push.
When the goal is met and every check passes, the work leaves the loop for one cleanup pass. Cut what isn't needed. Make it run fast. Tidy the code and the design. Then launch.
It's the design process
Designers already work this way. Testing changes the problem, not just the solution. I run the same loop with an AI building each prototype. The difference is speed, so I learn more per day and the goal sharpens faster.
| Step | Design phase | What it means here |
|---|---|---|
| Research | Discover | Talk to users, gather requirements, write the first goal. |
| 1 Define goal | Define | Frame the problem. On later laps, reframe it with what testing taught me. |
| 2 Plan | Ideate | The AI proposes an approach. I pick or redirect before anything gets built. |
| 3 Build | Prototype | A working prototype, at the fidelity the question needs. |
| 4 Check | QA | Automated checks catch broken work before it reaches a person. |
| 5 Try it | Test | Use it against the real goal. |
| 6 Compound | System | Patterns that worked become rules and skills, like components in a design system. |
| 7 Record | Document | Decisions and state, written down so the next session starts caught up. |
| 8 Ship | Release | Version it, write what changed, put it in front of people. |
| Clean up | Refine | Cut what isn't needed, make it run well, tidy the code and the design. |
The context files
Every AI session starts cold. A handful of small files at the project root catch it up, so it doesn't have to read the whole codebase.
| File | Written by | Job | Why it saves tokens |
|---|---|---|---|
| CLAUDE.md | Me | Standing orders: how to use the other files | Under ~100 lines. Points to files instead of repeating them. |
| spec.md | Me | The goal, audience, decisions. Rewritten as I learn. | The goal is written once, so it never has to be re-explained. |
| implementation.md | AI | What's built and how it was checked | The agent reads one page instead of the codebase. |
| conventions.md | Me | House style, grown from corrections | A mistake gets paid for once. |
| skills/ | Both | Reusable know-how | Loads only when the task needs it. |
| evals/ | Me | Good and bad examples for AI features | Written by someone who knows the goal, so polish alone can't pass. |
| changelog.md | AI | What changed, written for people | Catching up costs one read. |
Two habits that keep it honest
The audit. Every few features, the AI compares the spec, the implementation log and the actual code, then runs the checks. It reports gaps and stale rules. It doesn't fix them. I decide which side is wrong, because I'm the one who knows the goal.
Evals. Whenever AI writes something a customer will see, I keep 10 to 20 real inputs with examples of good and bad output, and run them every time a prompt or skill changes. And I design for when it's wrong: drafts instead of sends, a person reviews before anything goes live, undo, and showing where an answer came from.
In one breath
Anyone can loop an AI until the result looks polished. I spend tokens on learning what we're actually trying to build, and I let the goal change when I learn it.