HOYALABA PERSONAL RESEARCH NOTEBOOK
/

Engineering Notes / 2026-09-14 / 8 MIN

Prompt, plan, verify: how I hand coding work to an AI

The hard part of AI-assisted coding has moved from writing code to checking it. This is the routine I use to keep that checking small: structure the request, review the plan, then test the result myself.

In the previous part of this series, we set up Git: commit, break something on purpose, and roll it back with git restore. Once you can undo anything, you can afford to experiment. So this post gets to the main event: how to actually hand work to an AI.

(This is part 3 of my Korean series Coding in the Age of AI. Parts 1 and 2, on setting up the tools and on Git, are on Tistory for now.)

1. The bottleneck has moved

Before AI, the hard part was writing the code. Now the hard part is knowing whether the code is right.

That isn't just a feeling. In Sonar's 2026 State of Code survey, 96% of developers said they don't fully trust AI-generated code, yet only 48% said they always verify it before committing. Faros AI's engineering telemetry shows the other side: on teams with heavy AI adoption, pull-request review time rose 91%, and incidents per pull request rose by roughly 243%.

So almost nobody fully trusts the output, only about half of us check it, and incidents are climbing in the gap between the two.

Why does AI code take so long to review? My first guess was sheer volume. But there's another reason:

When the syntax is perfect but the logic is wrong, your eyes slide right past it. In the same Sonar survey, 38% of developers said reviewing AI code takes more effort than reviewing a colleague's. Human mistakes follow patterns you learn to spot over time. AI mistakes don't.

So my approach has three parts:

  1. Ask well, so there's less to check in the first place.
  2. Review the plan, so mistakes get caught before any code exists.
  3. Verify the result, for whatever still slips through.

Catching a problem after the code is written is expensive. Catching it in the plan costs almost nothing.

(A side note: I recently watched engineers from a large tech company describe how they verify work now that LLMs write so much of their code. Reviewing line by line had become very hard, so they lean on testing as many inputs and situations as they can. The bigger the project, the more I suspect everyone ends up there.)

2. Four parts of a good request

Why does the same model give wildly different results to different people? Look at this request:

Build me an expense tracker.

Now the AI has to decide everything for you: web app or Python script, where the data lives, what the screen looks like. And if you don't like its choices, you start over.

So I always include four things:

  • Purpose: who it's for and when they'll use it
  • Features: what it needs to do, one item at a time
  • Constraints: file format, tech to use, things to avoid
  • Deliverable: what I want back right now (for example, "show me a plan before any code")

The same expense tracker then becomes:

Purpose: log my own daily spending. Features: enter an amount and a category; show today's total. Constraint: a single index.html file. Deliverable: a plan before any code.

Now I get what I actually wanted, and when it misses, I can see exactly which part it missed. "Just make it good" is a shortcut to trouble.

(Tools like Claude and Codex have gotten better at asking clarifying questions on their own. Still, wouldn't you rather be the kind of boss who knows how to delegate?)

One more habit: ask for small pieces. Not "build an app with login, photo uploads, and payments," but "first, just let me add a to-do item." Check that it works, then move on to the next feature.

If that sounds familiar, it's the same idea as keeping branches short from the Git post. Big chunks are risky everywhere. Small pieces mean that when something breaks, you know where; the plan stays short enough to actually read and fix; and you burn through less of your usage limit.

This isn't about prompt-engineering tricks. It's the difference between saying "get me from Seoul to Busan" and saying "take the KTX to Suwon, stop for some of the famous galbi, catch a bus to Chungju, and then head down to Busan." (And if you're building something small and don't want to think that hard, you honestly don't need to.)

3. Plan mode: the part that matters most

Picture a contractor renovating your apartment. If they knock down a wall without a word and then announce "all done," it's too late. If they show you the drawing first, one sentence fixes everything: "Keep that wall."

Changes on paper are free. Changes after construction are expensive. Code works the same way, and that's why plan mode exists: the agent stops before touching any code, shows you its plan, and waits for your approval.

The major tools all have some version of it, under slightly different names:

  • Claude Code: Plan Mode. It explores the codebase and proposes a structured plan.
  • Codex CLI: plan mode. It proposes changes, waits for approval, and asks when something is ambiguous.
  • Antigravity: Implementation Plan. It pauses before editing and asks for approval, and you can edit the plan directly.

You don't even need the feature. Adding "show me a plan before you write any code" to your prompt gets you most of the way there. That's why I think of it as a habit rather than a tool feature.

When a plan comes back, I read it with three questions:

  • Which files will it touch? If there's a file I don't recognize, I ask why it's needed.
  • Is everything I asked for there? Is anything missing? Is anything I didn't ask for sneaking in?
  • Is anything risky? For deleting files, installing programs, or changing settings, I always ask for the reason.

Remember the rule from part 1, "never approve a command you don't understand"? Today it grows into this:

Approving means "I take responsibility for this." If an item is unclear, just ask: "Explain item 3 so that a beginner could follow it."

So please read the plan before you approve it, even briefly. At the very least, make sure your assistant understood what you asked and is about to do exactly that.

One cost tip: the plan is where the real thinking happens, so it's economical to plan with a more capable model on a deeper reasoning setting and then execute with a lighter one. For example, I'll plan with Claude Opus 5 on Ultracode and carry out the edits with Claude Sonnet 5 on high.

4. Don't trust "Done!"

AI reports its results confidently, and it's exactly as confident when it's wrong. That makes sense when you think about it: if the student taking the test also grades it, you can't trust the score.

You don't need to read every line of code. Anyone can open the result and judge whether it does what they wanted. I verify in three steps:

  1. Open it yourself, in your own browser rather than the AI's preview.
  2. Click through everything you asked for, in order.
  3. Try to break it: save an empty field, paste a very long string, enter strange values.

When you find a problem, describe the symptom precisely. "Clicking the done checkbox doesn't change anything" is a hundred times more useful than "it's broken." It's the same principle as print() debugging from the last post.

The final step is turning the habit into a procedure by slotting it into the commit cycle from part 2:

Before asking the AI  → git commit
When the plan arrives → read → revise → approve
When the result lands → run it yourself and check
If it works           → git commit -m "Add X (checked by hand)"
If it doesn't         → git diff, then git restore .

Glance at git diff before git restore ., so you don't throw away anything you meant to keep.

5. The three ways AI fails

Finally, damage control. AI tends to fail in three ways:

  • Hallucination: it describes features or files that don't exist. Run the code and see for yourself.
  • Overconfidence: it says "Perfect!" when the thing doesn't actually work. You do the grading.
  • Overreach: you asked it to change a button color, and it rebuilt the whole structure. Commit first, and restore if needed.

The nice part is that all three are handled by what we already have: verify, commit, restore.

Sometimes a conversation just spins its wheels. The AI repeats the same mistake, forgets what you said earlier, and makes things worse with every fix. At that point it's faster to open a new conversation and summarize the situation in three lines. The longer a conversation runs, the more the AI keeps leaning on its own earlier failed attempts, so starting fresh often beats trying to rescue a long thread.

One word on usage limits. Free tiers usually charge for how much work the AI does, not how many times you ask. One vague, sprawling request can use up more than five small, clear ones. The lesson is the same as everything above: be specific, go small, plan first. Seen that way, the limit is less a constraint than good training.

Today's eight rules

  1. Structure the request: purpose, features, constraints, deliverable.
  2. Ask for one feature at a time.
  3. Get the plan before the code.
  4. Read, revise, then approve. Don't just accept whatever you're given.
  5. If you don't understand something, ask before approving.
  6. Don't trust "Done!" Open it and check.
  7. If it breaks, roll back: commit, diff, restore.
  8. If it spins, start fresh: a new conversation and a three-line summary.

Try this for a day and you'll notice one annoyance right away: every new conversation starts with explaining your project from scratch. The stack, the scale, what not to do... The next post solves that with a single file called AGENTS.md.

Sources