A case study in engineering judgment
How to think about
a production feature
A two-week review as a case study
Watch first · read later
01The situation
Sept 2
Sept 17
review session another product reuses a piece
One feature. Two weeks of review.
“Does the code work?”
The real question: will this still be the right shape
when the tenth product arrives?
02The feature
Dine-in
Live chat · a user asks
Shared kitchen
one agent engine
Overnight catering
Smart Recommendations
⏱ scheduled job · nightly
“Why did spend jump?”
answer
Pause 4 targets
$400 spent · 0 sales
Apply
Chat opened
Pause 4 targets?
Approve
Live chat is dine-in: a user asks, the agent answers.
Smart Recommendations is overnight catering:
a scheduled job writes insight cards.
Apply opens a chat with the change ready.
Same kitchen, two ways to order.
03The twelve-question habit
Before writing code, walk the twelve questions.
The last one matters most:
what would the tenth product need from this?
04Kernel vs shells
Agent loop
think → tools → think
Model adapter
talks to the model
Smart
Recommendations
prompt · tools · job limits
Live chat
prompt · tools · a human
watching
Save messages
Telemetry
The kitchen is shared: one agent loop, one model adapter.
Each product is a shell: its own prompt, tools and limits.
Saving and telemetry listen from outside.
They are not the loop’s guts.
05Reuse is a caller test
shared/
Extracted ≠ shared
Moving files into a shared folder is not reuse.
CLAIM
Implementer
“I can migrate chat later.”
Verbal confidence.
Not yet evidence.
✓
a second caller, actually calling
Reuse is a caller test: a second caller must actually call it.
Another product took only the adapter first.
The loop can come later.
06One contract, two directions
One validation class
one shape for every change
Apply in chat
a human approves
Allow-list in · what the AI may change
Translate out · change → Apply payload
One definition, used both ways:
allow-list in, translate out.
the validation class — opened live in review
INPUT · what the agent recommended
{ target: "keyword 1182",
bid: 1.20 → 0.90 }
↓
OUTPUT · what Apply receives
interaction: update_bid(keyword 1182, 0.90)
In review, walk the file:
show the real input, then the real output.
07Thin edge, fat owner
Host
the seeder
seats the table
Manager
the agent service · runs the restaurant
✓ creates the chat
✓ adds the opening messages
✓ adds the interaction
✓ records telemetry
Waiter
conversation manager
already writes messages
reuse
Keep the edge thin. The host seats the table;
the owner runs the restaurant.
Is this insight still true?
Needs review
decided by the insight service, the owner of the insight
Whoever owns the noun owns its freshness:
old vs proposed vs current.
DRAFT · HIDDEN FROM USERS
PUBLISHED · VISIBLE
● chat created
● messages added
● interaction added
● freshness stored
Writes that can fail halfway: draft first,
publish when every step is done.
08Review judgment
Decide yourself
Bring options
Core owner in the room
If/else or a separate class?
Merge before evals exist?
Build an LLM gateway now?
Change how every chat behaves
Sort every decision: decide it, bring options,
or bring the owner.
APPLY PATH WANTS TO
make bid & state optional
Core chat
every existing chat
needs the chat owner
Touching every chat needs the chat owner in the room.
save(chat = optional)
updates an object in memory
one name · two behaviors
An optional flag that forks behavior
is really two functions.
09Leave room for ten
If it only works for this feature, it’s a fork
wearing shared-code clothes.
✓Ship the adapter now
→Share the loop later
✕Build a giant LLM gateway today
Leave room for ten. Build for the ones you have.
10Eight ideas to keep
The shape of the thinking, in eight lines.
This video was the map.
The essay is the territory.
production-thinking-from-reviews.pages.dev
One sitting · concepts first, then the stories that taught them