( technical program management — AI governance & evaluation )

The gate said GO. I recommended against scaling on it.

Governing an AI adoption decision in a regulated collections operation — seven phases, a pilot that passed its own rubric, and an analysis that showed the result couldn't be trusted yet.

0.87
RUBRIC SAID GO — I SAID NOT YET
±4.06
SE OF THE DIFFERENCE · CI CROSSED ZERO
18pt
RANKING VS CHANNEL ADHERENCE GAP
7
PHASES · 17 STORIES · 18 ARTIFACTS
0
GUARDRAIL VIOLATIONS
12pt
BASELINE SWING THAT RESHAPED THE VERDICT
( 01 — the problem )

The floor didn't need more agents. It needed a better answer to who first, and how.

A collections operation works roughly 6,000 delinquent accounts a month across aging buckets. Which accounts get contacted, in what order, through which channel — all of it decided by agent judgment, static queue order, and habit.

The cost is invisible until you measure it. High-risk accounts wait while agents work stale queue positions. Effort goes where outreach changes nothing. Channel and timing are guesswork. Performance swings without anyone able to separate skill from queue luck.

AI-assisted triage is an obvious candidate. It's also the kind of thing an organization either adopts deliberately — with governance, evaluation, and guardrails — or absorbs later, ad hoc and ungoverned. Cadence existed to make it the first one.

( 02 — the method )

Metrics before vendors. Weights locked in advance.

I defined seven KPIs and locked them before looking at a single vendor, then built a baseline from a 1,200-account, six-month dataset: cure rate 42.83%, RPC 36.81%, time-to-first-touch 4.98 days, roll rate 27.25%.

The baseline did two jobs. It set the before-picture — you cannot claim improvement without a measured starting point. And it doubled as a data-quality probe, surfacing 42 records (3.5%) with internally contradictory fields: contacts logged with no touch date, accounts flagged cured with no cure date. Caught by cross-field validation.

It also surfaced something I wasn't looking for. Monthly cure rate swung from 36.5% to 49.0% — a 12-point spread with no trend. I logged it as a finding. It came back later and reshaped the program's entire conclusion.

Then the build/buy/extend decision, scored against eight weighted criteria: Extend 3.72 · Build 3.60 · Buy 3.32. Buy ranked last despite scoring highest on raw capability — because black-box explainability and poor fit with our data quality are heavily weighted when you operate under guardrails. A capability-first evaluation picks Buy and walks into a compliance wall.

( 03 — the gate )

A gate that can only say yes is theater.

The evaluation rubric was built and locked before the pilot ran — seven weighted dimensions with pre-set targets, feeding logic that outputs GO / ADJUST / NO-GO, with NO-GO pre-agreed in writing as a legitimate, non-failure result.

And guardrail compliance sits outside the weighted score, as a hard gate. Because a weighted score treats everything as tradeable, and some things must not be. I tested it: one simulated violation slams the verdict to NO-GO regardless of a 0.87 score.

ADVISORY ONLY

No AI-triggered actions

Scoring recommends; humans act. Per-score explanations required.

OVERRIDES FREE

Always available, always logged

Never penalized — a 23% override rate is judgment, not non-compliance.

HARD GATE

Guardrails outside the score

Any violation forces NO-GO regardless of every other dimension.

( 04 — the finding )

The gate said GO. I recommended against scaling on it.

The pilot looked like a win: cure rate +5.00 points over control, time-to-first-touch 1.41 days faster, roll rate 6.67 points better, zero guardrail violations, rubric 0.87 → GO.

Then I actually analyzed it. First I validated the design — the control cohort reproduced the historical baseline within half a point on four of five KPIs, so the matched-cohort comparison rests on solid ground. Proven, not assumed.

Then I broke the headline. Measured weekly, the lift ran −12, +2, +2, +4, +12, +22 — a 34-point spread, with the pilot losing by 12 points in week two. Read top to bottom it looks like a learning curve. It's also exactly what noise looks like when you want a pattern.

So I quantified it. Standard error of the difference between two proportions: 4.06 points. The 95% confidence interval on that +5.00-point lift: −2.95% to +12.95%. It contains zero. The lift is consistent with the capability doing nothing at all.

The rubric's 0.87 records that targets were met. It doesn't encode confidence. Those two things diverged, and nobody would have noticed if I hadn't looked. The pilot was underpowered relative to its own baseline's volatility — and the fault was mine, in Phase 4, where I sized the cohort from operational convenience while the 12-point variance finding was already sitting in my Phase 1 model. The information was in hand and I didn't use it.

F — ADHERENCE SPLIT

Agents took the what and declined the how

Followed the ranking ("work this account first"): 70.21%. Followed the channel/timing rec ("reach them by SMS at 6pm"): 52.48%. An 18-point gap. The cure gain came from the half they adopted; the efficiency gain lived entirely in the half they didn't — and a single "adoption: 70%" number hid it completely.

F — DENOMINATOR

A denominator decided the verdict

Adoption against all pilot accounts reads 66% — fails. Against accounts that actually received a recommendation, 70.21% — passes. The difference is 18 accounts from a two-day scoring outage where there was no guidance to follow. Counting those as non-adoption measures uptime and blames the floor for it.

F — MARGIN

A pass on paper, a coin flip in fact

Even at 70.21% against a 70% target, that's a 0.21-point margin against a ±2.7-point standard error.

( 05 — the recommendation )

Scale to Wave 1 only. Not the floor.

The gate said GO and I declined to take it at face value. Wave 1 doubles the cohort and becomes the confirmatory sample the pilot wasn't; the full floor waits for Wave 2. If the effect materially weakens at n=12, the program returns to ADJUST. Explicitly not recommended: accelerating on the strength of a 0.87.

Attached conditions: split adoption into ranking-adherence and recommendation-adherence; add roll rate to the rubric — a 6.67-point improvement the rubric never asked for and arguably the most economically meaningful result in the study; require an override reason in the console, closing the 6.7% logged silently; guardrails stay a hard gate at every wave.

And I reopened my own Phase 2 decision with evidence. Stay on Extend. The binding constraint was never model sophistication — it was agent adherence at 52%. A more capable model, bought or built, would have produced better recommendations agents were already declining to follow.

( see the program )

Three minutes, baseline to verdict.

The locked KPIs and the noise I found in them, the weighted decision matrix, the guardrail-gated pilot — and the confidence interval that made me recommend against scaling on my own GO.

▶ Watch the walkthrough
( 06 — what I'd do differently )

Let the baseline's variance size the experiment.

This is the one. A target is not a design. I knew the metric swung 12 points before I designed the pilot, and I still sized the cohort by what the floor could spare. Everything downstream — the interval crossing zero, the coin-flip adoption pass — traces to that.

Measure each behavior separately. If a capability asks people to change more than one thing, one adoption number will hide the failure. Measure wider than the rubric. Roll rate wasn't a dimension and delivered the best result in the study.

The program's defining choice was made in Phase 1 and held to the end: triage, not autonomy. Advisory scoring, human-in-the-loop, explainability, the hard gate, overrides as data — all of it descends from that one decision. And when the gate said go, the honest answer was not yet.