( product design case study — agentic commerce )

I design intent-based experiences people can trust.

An agentic grocery app where the AI does the shopping and the human keeps the wheel — the whole design is a study in delegating without going blind.

THE PERSON TYPES ONE SENTENCE:
"Restock my kitchen for the week. Under $120. Never substitute my coffee."
THE AGENT READS IT BACK AS EDITABLE CHIPS — CORRECT-ME BEATS GUESSING.
( 01 — the problem )

Agents that shop for you are useless if you can't trust them to.

Hand a shopping list to an AI and two things immediately go wrong. It makes choices you'd never make — swapping your brand, blowing your budget, guessing when it should ask. And it does all of it silently, so you only find out at checkout.

Every real decision in this project pushed on the same question: how does an agent earn enough trust to act on your behalf — and give it back the moment it's unsure? The answer isn't a smarter model. It's an interface where the work is visible, confidence is a signal you can read, and the human is never more than one tap from the wheel.

( 02 — the rules )

Four rules the whole design obeys.

GLASS BOX

Show the work

Every pick can answer "why?" — reasoning, sources, and a flag to tell the agent when its reasoning is wrong.

LEGIBLE CERTAINTY

Confidence is a signal

A three-dot scale , never a false-precision percentage. Low confidence changes the UI, not just a number.

CONSENT

The human holds the wheel

Nothing is purchased until you approve, and approval is a deliberate act — you slide to place the order.

RECOVERABLE

Every edge state teaches

Ambiguity, out-of-stock, over-budget — each resolves in ≤2 taps and becomes a preference the agent remembers.

( 03 — one mission, end to end )

Delegate, watch, decide, learn.

The happy path is four screens. Iris marks every moment the agent is acting or reasoning, so you can scan the trust model before reading a word.

Intent Composer — the agent's read-back of your sentence as editable chips 01 INTENT COMPOSER Natural language becomes editable chips. Correct the agent before it starts.
Agent Working — the reasoning stream, visible while it shops 02 AGENT WORKING A live status rail and reasoning stream. Pause or redirect at any time.
Cart Proposal — confidence dots and the decision handed back 03 CART PROPOSAL Confidence per item. The few things needing a human are the primary action.
Reasoning Receipt — what it learned, what it saved, what you can edit 04 REASONING RECEIPT Logs the reasoning, not just prices. New memory chips are editable on the spot.
HI-FI FRAMES EXPORTED FROM FIGMA (S2 · S4 · S6 · S12) — SCHEMATICS STAY AS FALLBACK. THE RULE HOLDS EITHER WAY: IF IT'S IRIS, THE AGENT IS ACTING.
( 04 — the flow, before the pixels )

Seven screens for the happy path. Five for everything else.

Anyone can design the screen where the agent succeeds. The interesting work is the other five — the moments where delegation is either earned or thrown away. I wireframed the whole system before a single hi-fi pixel, and every edge state resolves in two taps or fewer, explains itself, and feeds agent memory.

HAPPY PATH — 7 SCREENS
S1 · Home wireframe
S1 · HomeA live mission surfaces its own progress. You never have to open it to know it's behaving.
S2 · Intent wireframe
S2 · IntentNatural language becomes editable chips. The agent shows its read before acting on it.
S3 · Guardrails wireframe
S3 · GuardrailsCap, locks, rules, and three approval postures. Delegation is configured, not assumed.
S4 · Working wireframe
S4 · WorkingStep rail, reasoning stream, running budget. Nothing is purchased until you approve.
S6 · Cart wireframe
S6 · Cart22 items triaged by confidence rather than dumped in a list.
S7 · Why this? wireframe
S7 · Why this?Per-item reasoning — and a “this reasoning is wrong” path that writes back to memory.
S11 · Approval wireframe
S11 · ApprovalA decision summary, not just a total. Slide to order — friction where it's worth paying for.
EDGE STATES — WHERE DELEGATION IS WON OR LOST
E1 · Ambiguity wireframe
E1 · AmbiguityExactly one question, the tradeoff on each option, and an offer to remember the answer.
E2 · Out of stock wireframe
E2 · Out of stockA proposal, never a silent swap. Accept, pick different, or drop — plus a permanent lock.
E3 · Cap hit wireframe
E3 · Cap hitThree costed paths, under a standing promise: I never raise your cap without asking.
E4 · Low confidence wireframe
E4 · Low confidenceThe agent says it doesn't know and hands the decision back. “Choose for me anyway” stays available.
S12 · Receipt wireframe
S12 · ReceiptWhat it did, why, and what it learned — with edit and delete on every memory.
LO-FI AND HI-FI SHARE ONE TOKEN SET — THE WIREFRAMES ARE THE SAME SYSTEM IN A WIREFRAME MODE, NOT A DIFFERENT PALETTE.
( 05 — the judgment layer )

Three decisions I'd defend in a review.

D-01 · INTENT, NOT A SEARCH BOX

The agent reads your sentence back — as chips you can edit

The riskiest moment is the start: if the agent misreads the goal, everything downstream is wrong. So instead of a chat box that hides its interpretation, the Intent Composer parses the sentence into visible, editable chips — budget, locks, quantities, timing. You fix the read before delegating. Correcting one chip is cheaper than unwinding a whole cart.

PARSED FROM "…NEVER SUBSTITUTE MY COFFEE"
Weekly restock Budget · $120 cap Never substitute · Peet's Deliver tomorrow AM
D-02 · WHEN THE AGENT ISN'T SURE, IT SAYS SO

Low confidence hands over the wheel — with the honest comparison

You've bought three detergent brands in three months; there's no pattern to safely guess from. Rather than pick and hope, the agent flags low confidence and lays out the real trade-offs — most-repurchased, best eco-match, cheapest per load — and lets you choose. Your pick becomes a memory chip, so it never has to ask twice. Admitting uncertainty is the trustworthy behavior.

LOW CONFIDENCE · YOU DECIDE  
Tide Pods 42ctyour most repurchased$12.99
Seventh Gen 45ozmatches your eco picks$10.49
Kirkland 85ctcheapest per load$9.99
D-03 · THE CAP IS NEVER RAISED SILENTLY

Over budget resolves as a choice between honest paths

When the cart lands $8.40 over the $120 cap, the agent doesn't quietly overspend or gut the list. It offers three paths — swap name brands for taste-matched store brands, drop the low-priority items, or raise the cap just this once — and recommends one, transparently. The guardrail the person set is treated as a promise, not a suggestion.

$8.40 OVER — PICK A PATH
A · Swap 3 brands for store brand−$9.10 · taste-match 4.4+REC
B · Drop 2 low-priority items−$8.98 · snacks rated "sometimes"
C · Raise cap to $130this mission only
( 06 — not a screen, a system )

The Agentic Commerce Kit.

Ten components, built on design tokens in DTCG format so they cross from Figma to code without translation. The interesting problem in an AI-native design system isn't the pixels — it's governance: how much of a layout an agent is allowed to assemble on its own. Every component declares its tier.

T1 · AGENT-COMPOSED

Structure guaranteed, content generative

The agent assembles these from data — confidence badges, status steps, reasoning rows, intent chips.

T2 · AGENT-POPULATED

Human frame, agent values

Fixed layouts the agent fills — buttons, guardrail chips, memory chips, substitution cards.

T3 · HUMAN-ONLY

Money and consent moments

The approval slider, interrupt sheets. The agent can never assemble these. A person lays them out, on purpose.

T1 CONFIDENCE BADGE T1 STATUS STEP T1 REASONING ROW T1 INTENT CHIP T2 BUTTON T2 GUARDRAIL CHIP T2 MEMORY CHIP T2 SUBSTITUTION CARD T3 APPROVAL SLIDER T3 INTERRUPT SHEET

COLOUR, SPACE, AND TYPE LIVE AS VARIABLES ( KIT · COLOR / KIT · SPACE & RADIUS ) AND SHIP AS A DTCG TOKEN FILE. THE AGENT GROUP IS RESERVED: IF IT'S IRIS, THE AGENT IS ACTING OR REASONING — THE RULE HOLDS IN FIGMA, IN CODE, AND ON THIS PAGE.

( 07 — the number this is built to move )

Does showing the reasoning actually build trust?

The whole thesis rests on one testable claim: that seeing the agent's reasoning raises willingness to let it act. So the study measures a trust rating before and after a participant sees the reasoning receipt, run unmoderated in Maze.

n=12
UNMODERATED MAZE PARTICIPANTS — THE SWEET SPOT FOR CATCHING ~85% OF USABILITY ISSUES
Δ trust
PRIMARY METRIC: 1–7 TRUST RATING, MEASURED BEFORE AND AFTER THE REASONING RECEIPT. RESULT PENDING THE TEST RUN
2 fixes
COMMITTED TO SHIPPING ONE ITERATION ON THE TWO HIGHEST-FRICTION MOMENTS THE TEST SURFACES
Honest by design: the metrics above are the plan, not invented results. The Maze test plan is written and ready to field; this section fills in with real numbers the moment it's run — and the iteration it prompts is part of the story, not hidden.
( see it working )

Ninety seconds, end to end.

The full mission running: one sentence in, the agent reading it back as editable chips, the reasoning streaming in the glass box, the low-confidence call handed back, and the receipt that logs what it learned.

▶ Watch the film Try the prototype ↗
( 08 — see it working )

The agent shops. You keep the wheel.

Run the full mission end to end — delegate an intent, watch the agent reason, decide the calls it hands back, and read the receipt. Or open the Figma file to see the screens and the component kit up close.

Try the live prototype ↗ Open the Figma file ↗ ← Back to all work