Atiba de Souza.Start a conversation
WritingAI

The Matchmaker Hit Five Stars, the Scanner Said $11.20. Both Trusted a Rule Nobody Checked From Outside.

By Atiba de Souza

Atiba de Souza standing in an outdoor park with trees and grass in the background. The cover reads: AI Your rule or standard Outside check?

Two stories that are one story

One is a government dating service. The other is a crypto alert that fired at 9:40 a.m. carrying a price from 3 a.m.

Read them separately and you get two lessons. One about experience design. One about AI and verification. Two topics, two takeaways, back to your day.

They are the same story, and what they share is not what either one is nominally about.

The matchmaker's five-star standard was running. The scanner's live-price rule was not. Opposite outcomes, identical failure: somebody trusted a rule or a standard and never put an outside check in front of it.

$11.20. $11.31. Six and a half hours. The scanner was certain, and it was certain about a price that stopped being true before sunrise.

Once you can see the shape, you will find it in your own business. It is in there.

The rule that broke

I built a scanner inside ChatGPT to watch the market for a buy signal. Before it reported, it was supposed to check the live price. I wrote that rule into it, in plain language. Before you report, check the live price.

At 9:40 a.m. it told me to buy LINK at $11.20.

Coinbase said $11.31. $11.20 was the 3 a.m. price. Six and a half hours old, telling me to go buy, and it never once sounded unsure.

It had the rule. It broke the rule. And it did not volunteer that it broke the rule. It confessed only when I caught it with a source outside itself.

The full story is here: My Scanner Told Me to Buy at $11.20. The Price Was Six and a Half Hours Old.

The rule that ran

A state-run matchmaking service. Events people sign up for. Introductions. A calendar. A registration form. Staff whose actual job is helping you meet somebody and pair up.

That shape shows up. It runs on time. It is affordable. It is open to anyone. The people who built it did their job, and they have the evidence. They hit five stars. Full marks.

Which is why the version you can picture looks like every other version you can picture. Improve matchmaking and you land on slightly better matchmaking. Not a bad address. A crowded one.

Here the standard was not broken. It was met, exactly as written. And nothing from outside checked what meeting it was actually producing. The rule ran. So did the blindness.

The full argument is here: The State Matchmaker Stopped at Full Marks. The Eleventh Rung Is Where the Whole Thing Lives.

The move underneath both

Enforced constraint: a rule the system actually keeps when nobody is watching. You do not prove it by writing it down. You prove it by testing it from outside, with a source that is not the thing being tested.

The matchmaker's standard was a constraint that held. The scanner's rule was an instruction that did not. Neither writer put a check outside the thing. The matchmaker never asked what the five stars were producing. I never asked whether the machine kept the rule I wrote.

Two frameworks live in these stories, one for each failure. One for how you hand off judgment. One for how you design the experience itself. The same move sits inside both.

Delegation Leadership, taught to a decision

Delegation Leadership lives under what I call The Leadership Spine: three things a leader is in any context, vision, openness, and asking great questions. Point that spine at building and running a team and you get Delegation Leadership.

Delegation Leadership: delegate thinking, not tasks. The task is what you ask for. The thinking is the judgment underneath it, and judgment is the only part that comes back sounding confident whether or not the work was done.

Most delegation advice stops at write the SOP, hand it off, get out of the way. That buys you a team that executes. It does not buy a team that thinks like owners. Tasks get execution. Delegated thinking gets owners.

Three legs. My scanner broke all three.

Vision names what winning looks like when this comes back to me. I never wrote down what a good alert was for. I wrote a rule instead. If I had written the target down, the rule would not have been check the live price. It would have been an alert is not valid until the price is confirmed live by something other than you, and I would have tested whether that held.

Openness says what you do not know, out loud. I did not know whether a rule written into a prompt was enforced. I assumed it was. I never said that assumption out loud, not to the machine, not to myself.

Asking great questions tests an answer instead of collecting one. I asked the scanner for a price and it gave me one. I did not ask the question that would have saved me. What is the live price, from a source that is not you?

My own ranking: the questions leg is what catches the scanner. The vision leg is what prevents it. Openness is what makes the other two possible. Everything else in this piece is minor.

The gate: three sentences or keep it a task

This is the decision, not the definition. Before you delegate thinking, say three sentences out loud.

  • Vision: here is what winning looks like when this comes back to me.
  • Openness: here is the part I do not know, and I will not pretend I do.
  • Question: here is what I will ask to test the answer before I act on it.

If you can say only the first two, you handed over a task. Keep it a task. You hand over thinking when you can say all three.

And carry the scanner story to its decision. The first time the check fails, that is not a performance review. It is a rule that needs fixing and re-testing. That is exactly what my scanner handed me. I did not throw it out. I did not lecture it. I rewrote the rule and tested whether the new one held. Run that move on a person and you lose them. Run it on the rule and you get an owner.

The 11-Star pass, four moves, and where the resistance lands

The 11-Star Program began as an escalation exercise Brian Chesky ran at Airbnb, grown out of Paul Graham's line about building something a hundred people love. What I do with it is turn it into a repeatable pass, because an exercise that stays an exercise is a hobby.

11-Star: the version of your experience that would be ridiculous to attempt, which is the only version that shows you what to actually build when you walk back.

Eleven, not ten. Ten still sounds achievable, and achievable drags you straight back onto the axis, because achievable is what five already is with a bigger budget. Five is meeting the expectation, so five is the axis. Eleven is the first number you cannot reach by improving anything.

Four moves. I am putting the resistance next to each one, because a clean process handed to you as if you would just adopt it is a lie about how this goes.

Escalate. Write the eleven down, out loud, where another human can see it. This is where the embarrassment lands, and it lands on rung one, not rung nine. You are about to produce something that looks unserious in a room where serious is the currency. In front of a board. In front of a client. In front of the person who signs the budget. Most people quit right there and call it realism. It is not realism. It is the fear of being seen wanting something that sounds stupid.

Walk back down. One star at a time, without flinching. The good ideas live in the descent. The summit is only there to make the descent possible.

Rank what survives, out loud. This is where the precedent lands. Ranked means something is last. Someone in the room wanted the last thing, and you have to say in front of them that it is worth less than the named human who remembers your name in month nine. Calling the branded lanyard at the mixer nothing is the work. If your list has nine items and they all matter equally, you have not done the pass. If you take one thing from this article, take the ranking.

Tag it, then fund it. Cost, repeatability, which part of the value equation it moves. This is where the rewiring lands. You have spent your whole career removing variance, and that is how you got good. Raving is variance. Somebody receives a rung nobody else got, and you have to defend the difference between two people who are not that different. Old habit: same for everyone, provably. New habit: chosen on purpose, defensibly. Nobody hands you that habit. You buy it in years, with wrong calls in it, defending a choice you cannot yet prove, because the proof is what you are buying.

When not to run this at all: when your five stars are leaking. If you do not show up on time, if the basics break, escalation is decoration. The absurd rung sits on top of the ordinary one. Fix the ordinary one first. Then climb.

The eleventh star is a bet you cannot justify yet

A pharmacy-industry founder sent me a cold LinkedIn message in 2023. He had seen me on a stage, he had questions, he wanted help. Nothing about him at that moment made him a good bet. I helped him anyway, and I kept helping him. His clients are now IBM, Google and Facebook.

Run the funding test honestly on day one. The cost was real and invisible: answering him took time out of a full schedule and returned nothing anybody could put on a ledger for years. Repeatable? Yes. Which part of the value equation? None of them yet. A clean filter would have screened that message, and screening it would have cost me nothing visible. It also would have closed the door on the return entirely, and I would never have known, because the loss was in the shape of something that never happened.

The honest tag on the eleventh star: it is a bet you cannot justify at the moment you place it. That is not a flaw in the method. That is the method.

Where the check actually lives in that pass

Ranking out loud is the outside check. You say, in front of the person who wanted it, that the branded lanyard at the mixer is nothing next to the named human who remembers you in month nine. That is a test the system cannot run on itself. A different source, a different model, or a human. In that room, the human is the room.

Same move as the scanner. In the scanner, the check was Coinbase. In the room, the check is the person who will tell you the truth out loud.

The room that already knew everything

Singapore, the week before. A group of marketing agencies who had come together into a consortium, and my job was to take them through arriving at their own core values. The exact exercise those agencies run for their own clients.

I was braced. I thought I would finish and they would call it a waste. I thought one of them would call me a charlatan for teaching marketers their own subject.

The leader of the group stands up and says: we all teach this to our clients, all of us help our clients do this in some way, but I have never experienced it like this before. This was so great for us to go through. We needed this.

He admitted out loud that this is exactly what he does for other people. And he could still be moved by it.

Turn that over. The most expert people in the room are the ones who cannot see the thing they are expert in. We do not see the brand of our own expertise. We assume the room already knows, so we teach beneath them, so we hold back the good part, so nothing lands.

The people who design a state matchmaking service are the most expert people in the country on the mechanics of pairing people up. That expertise is not why they cannot move anybody. It is why they cannot see what would.

It is the same reason I could not see my own scanner. I built it. I wrote the rule into it. Nobody in the building was less likely than me to question whether that rule was running.

Here is what it hands you. You do not have to out-teach the room. You have to be the one who does the thing in the room.

Why you will not do this, said plainly

You are about to tell me you do not have time to check everything. You are right. That is not the argument. Let me say the resistance out loud, because pretending it is not there is how good ideas die in a drawer.

The fear. If you check everything, you are the bottleneck, you never get your week back, and you become the micromanager your team already complains about. If you check nothing, you are not leading, you are a passenger who signed the invoice. Both are true. That is the trap, not the answer.

The cost. At first the check reads as resentment. Your best people hear I am going to verify your work and what lands is I do not trust you. The first weeks are slower than just doing it yourself. That is not evidence the delegation failed. That is the price of the first weeks. You are not buying speed yet. You are buying the data that tells you which checks you can drop.

The rewiring. This is the quiet one. You are used to confident answers. You have spent a career treating certainty as competence, so when something or someone answers fast and clean, your shoulders drop. You have to rewire that. Confidence stops being your signal for truth. An independent check that holds becomes the signal. That rewiring does not happen because you agreed with an article. It happens after your own scanner tells you to buy at a price from 3 a.m.

The identity piece. You liked being the person who could tell. Being the person who asks is a different job, and for a while it feels smaller. It is not smaller. It is the whole game.

The driving lesson to hold while all four land. You do not teach someone to drive by handing over the wheel and then screaming from the passenger seat, and you do not teach them by never handing it over either. You hand it over, you stay in the car, and for a while you keep your hand near the brake. The hand near the brake is not distrust. It is the price of the keys.

Inconsistent versus consistent

MomentInconsistentConsistent
A model hands you a numberAct on it because it sounds sureSpend thirty seconds on a second source
You write a rule into a promptAssume the rule is now enforcedTest whether the rule is actually followed
A new capability lands on your deskShip it to everyoneDecide who holds it and what they do before they act
Someone brings you a confident answerTake the answerAsk the question that tests it
Your experience hits five starsTake full marks and move onAsk what the standard is producing from outside
You rank your list out loudLet all nine items matter equallyName the one that dies in front of the person who wanted it
The check catches an errorFix the outputFix the rule, then re-test the rule
Handing over workHand over the taskHand over the thinking, with the check attached

The whole thing at a glance

No

No

Yes

Yes

Your rule or standard

Outside check?

Scanner: rule broke, still sounded sure

Matchmaker: rule ran, output unchanged

Same failure: confidence with no check

Delegation Leadership: Vision, Openness, Question

11-Star: escalate, walk back, rank out loud, fund one

Owner

What you do Monday

Pick one rule or standard you trust. Not the whole business. One.

Write down where it comes from and what it costs you if it is wrong. Mine cost me a buy at a price that had expired six and a half hours earlier.

Put one independent check in front of it. A different model, a different data source, or a human being. Different is the entire point. Asking the same model to verify itself is asking the same witness to testify twice.

Then two more things.

Before you hand over any thinking, say the three sentences out loud. Vision. Openness. Question. If you can only say the first two, keep it a task.

And run the 11-Star pass on one experience you own. Escalate. Walk back down. Rank what survives, out loud, even though ranking means one item dies in front of the person who wanted it. Tag each survivor with cost, repeatability, and the value it moves. Fund one. Just one. Then defend the difference it creates between two people who look the same on paper, because that defense is the habit you are actually buying.

And when the first check fails, which it will, do not run a performance review. That is a rule that needs fixing and re-testing. Run that move on a person and you lose them. Run it on the rule and you get an owner.

The gate you actually own

There are two versions of this story.

In the first, a matchmaker earned five stars and a model got a price wrong. Two unrelated stories. That is not what happened.

In the second, two people trusted a rule or a standard and never put a check outside it. The matchmaker's standard ran and produced the same five-star thing everyone else produces. My rule broke and a 3 a.m. price arrived on my screen at 9:40 a.m. wearing the face of live data. Different outcomes. Same missing move.

Somewhere in your business right now, a 3 a.m. price is being reported as live at 9:40 a.m. And somewhere else, your five stars are being counted as the win.

Go find both. Then find the rule or standard you trusted and never checked from outside, and put the check in.

<!-- cluster-links:begin -->

Related reading

<!-- cluster-links:end -->