Building with AI, and building AI in
The loop I run with Claude Code and Codex, and the rules for AI features people can trust.
Meer Habib
Senior Mobile Engineer · Chittagong
There are two ways AI shows up in my work. It helps me build software, and it's part of the software I build. They need different habits, so this piece is in two halves.
Part one: building with an agent
Most of my code now starts in Claude Code or Codex. That hasn't made the job smaller. It has moved it: less typing, more deciding, and a lot more checking.
- you decide
- the agent does
- shared
Give it what a new teammate would need
An agent is a very fast engineer who joined this morning. It doesn't know your conventions, your decisions, or the thing you tried last month that didn't work. So I write those down, in the repo, where it reads them first:
- A short instructions file at the root: how to run, test and build; naming and folder conventions; what never to touch.
- Phase plans for anything bigger than an afternoon: the flow, the states, the acceptance criteria.
- A decision log. When I decide something, it gets a number and a sentence. Mend's log has dozens of them, like the one that made consent a server-side check instead of a screen.
The working agreement in Mend's repo is five lines: draft one phase, review the whole flow, record the approval, build only what was approved, verify against the criteria. Agents follow it well because it's short and it's written down.
Small diffs you can prove
One change per turn of the loop, small enough to read in a minute. And every change has to prove itself before I look at it:
- the type checker passes,
- the tests pass, or new ones exist,
- the app runs, and for UI, there's a screenshot of the simulator.
Agents are very good at saying something works. Make them show it.
Read every diff
I read everything before it's committed. I read slowest, and sometimes rewrite by hand, anything touching:
- authentication and permissions,
- payments and subscriptions,
- deleting or migrating real user data,
- what leaves the device, and when.
Where it's strong, and where it isn't
It's excellent at first drafts of screens, refactors that have tests around them, migrations, reading unfamiliar code, writing tests and docs, and the hundred small tasks around a release.
It's weak at knowing what the product is for, at taste, at knowing what to leave out, and at performance claims nobody measured. Those stay with me.
Part two: AI inside the product
One job, done well
The AI features that work are narrow. Summarise this call. Draft a reply. What breed is this dog? A narrow job can be tested, priced and trusted. "Ask me anything" usually can't.
Ask for structure, then check it
Don't parse paragraphs. Ask for data that matches a schema, validate it, and have a plan for when it's wrong:
import { generateObject } from "ai";
import { z } from "zod";
const Summary = z.object({
title: z.string().max(80),
actions: z.array(z.object({ text: z.string(), owner: z.string().optional() })).max(10),
});
const { object } = await generateObject({ model, schema: Summary, prompt: transcript });The model comes from one place in the codebase, so changing providers is a one-line change. Dev Partner is wired to more than one provider this way.
Say what leaves the phone
If words, photos or voice go to a model, tell the user before it happens, in plain language, and make "no" a real answer. I wrote up how Mend does this.
Price it per user
Every call has a cost and a delay. Work out the cost per active user per month before launch, not after. Stream long answers so the first words arrive fast. Cache anything that's asked twice.
Keep a small test set
Save twenty real inputs, with what a good answer looks like. Run them every time you change a prompt or a model. It takes minutes and it catches the regressions that users would have found for you.
The short version
- Write down context the agent can't guess.
- One small diff at a time, and make it prove itself.
- Read everything; read auth, payments and data twice.
- In the product: one job, structured output, honest consent, known cost, a test set.
And for keeping your balance while all of this changes every month, see riding the AI roller coaster.
Building something like this?
Booking new projects for Q4. Replies within 24h.