THE AGENT READS IT BACK AS EDITABLE CHIPS — CORRECT-ME BEATS GUESSING.
GLASS BOX →LEGIBLE CERTAINTY →CONSENT →RECOVERABLE →IF IT'S IRIS, THE AGENT IS ACTING →GLASS BOX →LEGIBLE CERTAINTY →CONSENT →RECOVERABLE →IF IT'S IRIS, THE AGENT IS ACTING →
( 01 — the problem )
Agents that shop for you are useless if you can't trust them to.
Hand a shopping list to an AI and two things immediately go wrong. It makes choices you'd never make — swapping your brand, blowing your budget, guessing when it should ask. And it does all of it silently, so you only find out at checkout.
Every real decision in this project pushed on the same question: how does an agent earn enough trust to act on your behalf — and give it back the moment it's unsure? The answer isn't a smarter model. It's an interface where the work is visible, confidence is a signal you can read, and the human is never more than one tap from the wheel.
( 02 — the rules )
Four rules the whole design obeys.
GLASS BOX
Show the work
Every pick can answer "why?" — reasoning, sources, and a flag to tell the agent when its reasoning is wrong.
LEGIBLE CERTAINTY
Confidence is a signal
A three-dot scale , never a false-precision percentage. Low confidence changes the UI, not just a number.
CONSENT
The human holds the wheel
Nothing is purchased until you approve, and approval is a deliberate act — you slide to place the order.
RECOVERABLE
Every edge state teaches
Ambiguity, out-of-stock, over-budget — each resolves in ≤2 taps and becomes a preference the agent remembers.
( 03 — one mission, end to end )
Delegate, watch, decide, learn.
The happy path is four screens. Iris marks every moment the agent is acting or reasoning, so you can scan the trust model before reading a word.
New mission
The agent's read-back
Weekly restock$120 capLock · coffee
Deliver tomorrow AMtap any chip to edit
Delegate to agent
01 INTENT COMPOSERNatural language becomes editable chips. Correct the agent before it starts.
Weekly restock
Comparing prices…
Oat milk 2-pack beats singles by $1.20
Locking Peet's YOUR RULE
$83.40 of $120
Pause · redirect
02 AGENT WORKINGA live status rail and reasoning stream. Pause or redirect at any time.
Your cart is ready
16 auto-picked · 4 need your OK
Spindrift Lime replaces La Croix SUB
Oat milk 2-pack
Review 4 items
03 CART PROPOSALConfidence per item. The few things needing a human are the primary action.
Mission complete
$112.80 · $7.20 under cap
What I learned
Whole-bean OK when 15%+ offnew memory · editable
Saved $6.30 vs last week
Save as template
04 REASONING RECEIPTLogs the reasoning, not just prices. New memory chips are editable on the spot.
HI-FI FRAMES EXPORTED FROM FIGMA (S2 · S4 · S6 · S12) — SCHEMATICS STAY AS FALLBACK. THE RULE HOLDS EITHER WAY: IF IT'S IRIS, THE AGENT IS ACTING.
( 04 — the flow, before the pixels )
Seven screens for the happy path. Five for everything else.
Anyone can design the screen where the agent succeeds. The interesting work is the other five — the moments where delegation is either earned or thrown away. I wireframed the whole system before a single hi-fi pixel, and every edge state resolves in two taps or fewer, explains itself, and feeds agent memory.
HAPPY PATH — 7 SCREENS
S1 · HomeA live mission surfaces its own progress. You never have to open it to know it's behaving.
S2 · IntentNatural language becomes editable chips. The agent shows its read before acting on it.
S3 · GuardrailsCap, locks, rules, and three approval postures. Delegation is configured, not assumed.
S4 · WorkingStep rail, reasoning stream, running budget. Nothing is purchased until you approve.
S6 · Cart22 items triaged by confidence rather than dumped in a list.
S7 · Why this?Per-item reasoning — and a “this reasoning is wrong” path that writes back to memory.
S11 · ApprovalA decision summary, not just a total. Slide to order — friction where it's worth paying for.
EDGE STATES — WHERE DELEGATION IS WON OR LOST
E1 · AmbiguityExactly one question, the tradeoff on each option, and an offer to remember the answer.
E2 · Out of stockA proposal, never a silent swap. Accept, pick different, or drop — plus a permanent lock.
E3 · Cap hitThree costed paths, under a standing promise: I never raise your cap without asking.
E4 · Low confidenceThe agent says it doesn't know and hands the decision back. “Choose for me anyway” stays available.
S12 · ReceiptWhat it did, why, and what it learned — with edit and delete on every memory.
LO-FI AND HI-FI SHARE ONE TOKEN SET — THE WIREFRAMES ARE THE SAME SYSTEM IN A WIREFRAME MODE, NOT A DIFFERENT PALETTE.
( 05 — the judgment layer )
Three decisions I'd defend in a review.
D-01 · INTENT, NOT A SEARCH BOX
The agent reads your sentence back — as chips you can edit
The riskiest moment is the start: if the agent misreads the goal, everything downstream is wrong. So instead of a chat box that hides its interpretation, the Intent Composer parses the sentence into visible, editable chips — budget, locks, quantities, timing. You fix the read before delegating. Correcting one chip is cheaper than unwinding a whole cart.
PARSED FROM "…NEVER SUBSTITUTE MY COFFEE"
Weekly restockBudget · $120 capNever substitute · Peet'sDeliver tomorrow AM
D-02 · WHEN THE AGENT ISN'T SURE, IT SAYS SO
Low confidence hands over the wheel — with the honest comparison
You've bought three detergent brands in three months; there's no pattern to safely guess from. Rather than pick and hope, the agent flags low confidence and lays out the real trade-offs — most-repurchased, best eco-match, cheapest per load — and lets you choose. Your pick becomes a memory chip, so it never has to ask twice. Admitting uncertainty is the trustworthy behavior.
LOW CONFIDENCE · YOU DECIDE
Tide Pods 42ctyour most repurchased$12.99
Seventh Gen 45ozmatches your eco picks$10.49
Kirkland 85ctcheapest per load$9.99
D-03 · THE CAP IS NEVER RAISED SILENTLY
Over budget resolves as a choice between honest paths
When the cart lands $8.40 over the $120 cap, the agent doesn't quietly overspend or gut the list. It offers three paths — swap name brands for taste-matched store brands, drop the low-priority items, or raise the cap just this once — and recommends one, transparently. The guardrail the person set is treated as a promise, not a suggestion.
$8.40 OVER — PICK A PATH
A · Swap 3 brands for store brand−$9.10 · taste-match 4.4+REC
B · Drop 2 low-priority items−$8.98 · snacks rated "sometimes"
C · Raise cap to $130this mission only
( 06 — not a screen, a system )
The Agentic Commerce Kit.
Ten components, built on design tokens in DTCG format so they cross from Figma to code without translation. The interesting problem in an AI-native design system isn't the pixels — it's governance: how much of a layout an agent is allowed to assemble on its own. Every component declares its tier.
T1 · AGENT-COMPOSED
Structure guaranteed, content generative
The agent assembles these from data — confidence badges, status steps, reasoning rows, intent chips.
COLOUR, SPACE, AND TYPE LIVE AS VARIABLES ( KIT · COLOR / KIT · SPACE & RADIUS ) AND SHIP AS A DTCG TOKEN FILE. THE AGENT GROUP IS RESERVED: IF IT'S IRIS, THE AGENT IS ACTING OR REASONING — THE RULE HOLDS IN FIGMA, IN CODE, AND ON THIS PAGE.
( 07 — the number this is built to move )
Does showing the reasoning actually build trust?
The whole thesis rests on one testable claim: that seeing the agent's reasoning raises willingness to let it act. So the study measures a trust rating before and after a participant sees the reasoning receipt, run unmoderated in Maze.
n=12
UNMODERATED MAZE PARTICIPANTS — THE SWEET SPOT FOR CATCHING ~85% OF USABILITY ISSUES
Δ trust
PRIMARY METRIC: 1–7 TRUST RATING, MEASURED BEFORE AND AFTER THE REASONING RECEIPT. RESULT PENDING THE TEST RUN
2 fixes
COMMITTED TO SHIPPING ONE ITERATION ON THE TWO HIGHEST-FRICTION MOMENTS THE TEST SURFACES
Honest by design: the metrics above are the plan, not invented results. The Maze test plan is written and ready to field; this section fills in with real numbers the moment it's run — and the iteration it prompts is part of the story, not hidden.
( see it working )
Ninety seconds, end to end.
The full mission running: one sentence in, the agent reading it back as editable chips, the reasoning streaming in the glass box, the low-confidence call handed back, and the receipt that logs what it learned.
Run the full mission end to end — delegate an intent, watch the agent reason, decide the calls it hands back, and read the receipt. Or open the Figma file to see the screens and the component kit up close.