Atiba de Souza.Start a conversation
WritingAI

Your AI Doesn't Need To Think. It Needs To Decide.

By Atiba de Souza

Atiba de Souza wearing a blue cap and a black shirt, smiling and giving a thumbs up with both hands. The cover reads: AI Same decision, over and over, and the answer Does it need reasoning across messy context

The alert that should have been impossible

Ten minutes before an eight hour training was supposed to start, an alert fired. It came from a crypto scanner I built inside ChatGPT on purpose, because I wanted to watch how it would handle its own rules.

LINK, it said. Time to buy. $11.20.

I opened Coinbase. Coinbase said $11.31. So I went back and told it what I was seeing. It said, oh yes, it is no longer time to buy, my data was stale. So I double checked. $11.20 was the price at 3 a.m.

Six and a half hours late. It was 9:40 in the morning. The number was from 3 a.m. And it told me to buy, sounding completely certain, with a written rule sitting inside it that said the one thing it needed to do: check the live price before you report.

Script: a learned answer to a familiar situation, run fast enough that you stop noticing you are running it.

Not the cost. The cost is the easy part. The interesting part is the difference between asking a model to follow a rule and building a model that does nothing else.

Scripts are how everything moves

People run scripts all day, and it is not a flaw. Ellen Langer's research on mindlessness found that a small request with an empty reason ("because I have to make copies") got granted almost as often as the same request with a real reason. Not because the reason was good. Because the word "because" ran the script.

Judee Burgoon's work on violated expectations gives you the useful half: when a script gets interrupted, attention shows up. The interruption is the opening.

Then Friestad and Wright's research on persuasion knowledge gives you the line every marketer should sit with: naming a script weakens it. Once a person can see the pattern, or see the tactic being used on them, they judge it instead of following it.

Now read that back and swap the subject. A decision model is a script, built on purpose, running fast, and it does not mind being named. That is the entire shift.

What a decision model actually is

Most people hear those two words and picture a smaller, cheaper chatbot.

Stop picturing every AI as something that thinks and writes. Some are built to make an instant yes or no decision and nothing else.

That is the whole idea. A yes or no question does not need a model that reasons and writes. Sending a simple decision to a large language model wastes money and time, and it does something worse than that: it hides the decision inside a story.

Because here is what happens when the call gets made by a big reasoning model. You see reasoning, but it is a story told after the fact. The rule you wrote is not a property of the thing. It is a sentence in the room this time, and next time the room is different.

Ranking the parts, so you do not spend your attention evenly

  1. The routing decision matters most. Which decisions get a fast model that only decides, and which get a reasoning model that thinks. Get this wrong and nothing else on this page saves you.
  2. The check matters second. A fast decision with nothing checking it is a faster way to be confidently wrong. My scanner taught me that at 9:40 in the morning.
  3. And if you take one thing from this, take this instead of the ranked list: the reward sentence. A model built only to decide is only as good as the sentence you trained it against. Most teams cannot write that sentence about their own funnel. That gap is the real project, and it is not a technical one.

I am putting that third item last on the list and first in your attention on purpose. The ranking is tidy. The rewiring is not.

Where the cost actually lands

You have a prompt library. Your team spent months on it. The agent stack works well enough. Now somebody is telling you the rules were never rules.

That is the real cost. Not the budget. The admission.

There is a second cost, and it is the one that stops teams cold. Building a decision model is a different skill than prompting one. You have to know what you want decided, stated in a form a reward signal can measure. That is a sentence about your business, not about AI, and it is usually the first time anyone has tried to write it down.

So name what you have to rewire: from I tell the model what to do to I decide what should be decided, and I measure whether it was. It sounds small on the page and it feels enormous on a Monday.

And the fear underneath it is legitimate. You are about to hand a fast, confident decision maker the front door. Which is exactly why the check gets built before the model, not after.

The honest use, applied to a decision engine

Every pressure tactic rests on something true about people. We never just do its opposite. We find the truth it rests on and use that thing in the open. Three parts, and then the test.

1. The human truth it relies on. Most decisions are yes or no, and they should be fast. A person deciding whether to click has zero patience for your deliberation, and neither does a system running at scale.

2. The corruption that makes it harmful. The rule stays hidden instead of published. Confidence gets used as a stand in for correctness. A prompt gets treated as a constraint. The call runs at machine speed and nobody ever measures whether it was right.

3. The honest use. The same truth, used openly. The decision that repeats is the decision that gets a model built only to decide. The rule is published. The call gets logged, and a second independent check stands next to it.

The test: it still works if the person can see exactly what you are doing.

Run your decision engine through that test. Does it still work if your customer can read the rule? If the answer is no, you have not built a decision engine. You have built a tactic with better latency.

The mistake I already made in public

Two and a half weeks and about a thousand dollars of tokens into one build inside Claude, going in circles. I finally said it out loud: I am back where I was two and a half weeks ago. This is a cycle.

So I had it write up the spec of everything we had done. Then I took that spec into ChatGPT and laid out the situation, the spec, the problem, and the outcome I wanted. ChatGPT came back with ideas we had never thought of. When I carried them back to Claude, Claude called them novel.

The second model was not smarter. The first model was not broken. The thread was. And the thing that actually moved between the two systems was the spec.

That spec is the artifact. It is the difference between a prompt that says "check the live price" and a model built around checking the live price. The prompt is a request. The spec is a decision, written down, that a second thing can be measured against.

Which is the same move as the reward sentence. You cannot train a decision you have not described.

Where the rule actually lives, and what it costs you to put it there

Reasoning model asked to decideModel built only to decide
What you hand itA prompt with the rule in it, this timeThe decision, and nothing else
What happens when the data conflictsIt narrates its way around the rule, with confidenceIt answers the question it was built for
How consistent it isIt varies. Same rule, different day, different answerThe rule is the model. Consistency is the point
Where the cost sitsEvery single call pays for the reasoning you did not needPaid up front, in the work of writing a measurable reward
What it does to the person on the other endThey cannot see the rule, only the outputThey can be told the rule, because the rule is the product
What it costs you to get thereLow. You write a prompt this afternoonHigh. You have to write the reward sentence, and you have to mean it

Look at the last row. That is the honest tension in everything above. The cheap column is cheap for a reason, and the expensive column is expensive for the same reason. I am not going to pretend the trade is free.

Pair it, or do not ship it

The rule from the crypto story, in one line: pair every AI generated output with a check run by a different model or a human, and never treat an instruction written into a prompt as an enforced constraint. Move the repeated yes or no decision onto a model built for it, and keep the check standing next to it, because "built for it" is still not "guaranteed."

Two and a half weeks and a thousand dollars is what the loop costs when nothing is checking the loop. Six and a half hours late is what confidence costs when nobody is checking the price.

Yes

No

No

Yes

Same decision, over and over, and the answer is yes or no

Does it need reasoning across messy context

Keep the reasoning model. Pay for the thinking, because you need it

Build a model that only decides. Write the reward sentence first

Human truth: scripts run fast, and fast is correct here

Corruption: the rule stays hidden, confidence stands in for correctness, a prompt gets treated as a constraint

Honest use: publish the rule, log the call, put an independent check next to it

It still works if they can see exactly what you are doing

You built a tactic with better latency

Ship it. Then keep checking

When NOT to do this

  • Not for a decision that happens a handful of times a month. The training cost never pays back and the sample is too small to learn from.
  • Not when the decision genuinely needs reasoning across messy context. That is the case for the big model. Routing it to a fast one is penny wise and pound foolish.
  • Not when you cannot yet say out loud what a correct decision looks like. A reward signal cannot be built on a sentence you refuse to write.
  • And never as a marketing claim. "We run decision models" is not a value proposition. It is a vendor badge.

Two things to do this week. Pick the one decision in your funnel that repeats, gets a yes or no, and currently goes to a model that reasons and writes. Then write the sentence that describes what a correct decision looks like, in a form you could measure. If you cannot write that sentence, you have just found the actual project, and it was never about the model.

<!-- cluster-links:begin -->

Related reading

<!-- cluster-links:end -->

Questions

What people ask next

Okay, but how do you actually tell which decisions are simple enough to hand to a pure decision model versus the ones that still need a reasoning model? The article says this call matters most but doesn't say how to make it.
Start with whether the decision repeats. A decision that repeats is the one that gets a model built only to decide, because repetition is what lets you write a reward sentence and measure it. If you cannot state the decision as a sentence a reward signal could score, that is your answer too: it is not ready for a decision model yet, and it stays with reasoning or a person until it is.
If naming a tactic weakens it, and the whole pitch here is publishing the rule in the open, what stops people from gaming the rule once they can see it?
The test is whether the decision engine still works when the person can see exactly what you are doing. A tactic only survives being hidden; an honest decision engine is built so the rule holds up in the open, with a published rule and an independent check sitting next to the call. If seeing the rule lets someone break it, that is the signal you built a tactic with better latency, not a real decision engine.
Who writes that 'measurable reward sentence' in practice when different people in the company don't even agree on what the right decision should be?
The article is honest that most teams cannot write that sentence about their own funnel, and that is the real project, not a technical one. It has to be written by whoever owns the outcome the decision is supposed to produce, not whoever owns the prompt. If people disagree on what the right decision is, that disagreement is the actual work to surface before any model gets built, not something to route around.
The 'check' that's supposed to sit next to the fast decision - what does that actually look like day to day, and who owns catching it when it's wrong?
It looks like a second, independent check built before the model goes live, not bolted on after, with every call logged so the rule and the outcome can both be reviewed. The article does not hand this to one named role, but it insists the check cannot be the model checking itself, since that is exactly what failed at 9:40 in the morning. Ownership follows the rule: whoever published the rule is accountable for the check that verifies it.
Doesn't stripping out the reasoning step also strip out the one place where a model might have caught something was stale or off, like what happened with the scanner?
The scanner case actually shows the opposite: the reasoning model had the rule written inside it and still narrated past it with total confidence instead of catching the stale data. The catch came from checking against Coinbase, outside the model, not from the model's own reasoning. That is the argument for building the check before the model rather than trusting a reasoning step to notice its own mistake.
Once you've paid the upfront cost of writing the spec and building the decision model, how do you know if the rule itself needs to change later, since nothing is narrating its logic anymore?
You do not rely on narration for this, you rely on the logged calls and the independent check sitting next to the model. If the check starts disagreeing with the decision model more often, that gap is the signal the reward sentence no longer matches reality. The rule is the model now, so revisiting the rule means going back to the sentence and rewriting it, the same work it took to build it the first time.