Atiba de Souza.Start a conversation
WritingAI

AI Coding Made CI the Bottleneck. The Bottleneck Was Always You.

By Atiba de Souza

An editorial cover image for a piece on AI Coding Made CI the Bottleneck. The Bottleneck Was Always You. The cover reads: AI AI Coding Made CI the Bottleneck. The Bottleneck Was Always You.

Your pipeline didn't get slower. Your output got faster.

Six months ago your CI was invisible. It ran. It went green. Nobody thought about it, because nothing was waiting on it.

Then the models started writing code. Now there is a queue. Pull requests stacked four deep. A red build nobody has touched since Tuesday. Three people in a thread asking who owns the pipeline.

And every conversation in the building is about the same thing: how do we make CI faster.

Wrong question. It is also the exact question every owner-dependency story starts with, including mine.

Your CI did not become your bottleneck. Your CI became the first place your company has to say yes or no, and your company has only ever had one person who says yes or no. The pipeline is not the problem. The pipeline is the report.

The thing we are actually talking about

Organizational Independence: building a company that can think without you. A stronger organization. A freer you.

That is the whole game, stated as an outcome. Build to replace yourself.

Here is why it belongs in a conversation about CI. Your pipeline is the cheapest thinking-without-you machine you already own. It runs at 3am. It runs while you are on a plane. It runs when you are sick of the whole thing. If it cannot decide, you have not built a pipeline. You have built an expensive notification system that wakes you up.

Three questions, in order, and CI lives in the last one

Organizational Independence is three questions in sequence, with a learning loop that returns to the start. All of it sits inside the Brain. All of it stands on Trust.

The Brain

learning loop:
a red build becomes a change to the definition

1 Alignment
Vision / Org chart / Theory of Constraints

2 Movement
Values / Frameworks / Culture

3 Measurement
Start past the end / Everyone has a KPI

Trust: the floor it all stands on

1 · Alignment: do we know where we're going and what we own?

Vision is a clear and compelling future. The org chart is who owns what thinking. The Theory of Constraints says focus on the Big 3.

For CI, alignment collapses into one sentence: what does green mean?

Not "the tests pass." What does green mean about the product, the customer, the risk you are willing to carry. If three people on your team would give three different answers, you do not have a pipeline. You have a mood ring.

Alignment is also where you own the constraint. Your CI has a Big 3, the checks carrying the actual weight, and a long tail of ceremony nobody would miss. If you cannot name which is which, you will spend your next month making the ceremony faster.

This is the heaviest of the three questions. Everything downstream is decided here. If you only fix one thing after reading this, fix this one.

2 · Movement: can they decide and act without me?

Values are the principles that guide decisions. Frameworks are simple ways to think and act. Culture is a team that improves together.

This is where you find out whether the pipeline belongs to the company or to you. CI goes red at 4:55 on a Friday. What happens? The values answer exists so nobody has to text you.

Now the part nobody says out loud. The hard thing about movement is not writing the rule. It is living under it. The second the rule is written down, you stop being the exception to it. You have been the exception for years, and being the exception is a real job with real rewards. It is also the job you are trying to give away.

The fear underneath is not about CI. It is that if they can decide without you, you find out you were never the reason it worked. Sit with that one, because it is the actual obstacle. Your team is not refusing to decide. You never handed them the decision.

3 · Measurement: can we tell if we were right?

Start past the end. Define success, then work backward. Everyone has a KPI, which means clear signals for progress.

For a pipeline, that means two concrete things.

First: write the sentence about what green means before you touch a check, then work backward to the smallest set of checks that can prove it. Most teams do this in reverse. They inherit a pile of checks and then try to remember what the checks were for.

Second: every check has a name on it. Not the team's name. A person's name, and a signal that tells that person whether the check is doing its job. What it protects. How it fails. When it last caught something real. That is what a KPI looks like for a pipeline.

And this is where you pay. Measurement costs you a week that ships no feature and wins no applause. It costs you the moment in front of your team where the honest answer is "I don't know, what does the check say," and you have to let that be the final word.

The rewiring is one habit. Old habit: the number I care about is how many builds are green. New habit: the number I care about is how many calls this pipeline made that I did not have to make. Same pipeline. Completely different company a year from now.

Then the loop closes

Every red build is a lesson about what green should have meant, and that lesson goes back into alignment. Not into a postmortem document nobody opens. Into the definition, the ownership, the Big 3.

Say a build goes red on Friday on a check everyone agrees was a false alarm. The fast lesson is "fix the test." The real lesson is a question that belongs in alignment: did green ever mean what that check was testing? Some weeks the answer is no, and the check comes out of the pipeline or gets promoted into the Big 3. Some weeks the answer is yes, and the sentence gets a word added to it. Either way the change lands in the definition, not in the log.

That is why this is a loop and not a checklist.

Your pipeline is a vending machine

Here is the picture that makes it click.

Most people use CI the way most people use AI. Put in a coin, press a slot code, get an item.

Commit goes in, green light comes out.

The vending machine still works. It will give you the Snickers bar. It is a floor, not a fault. But a green light is an output, not a judgment, and getting an output faster is not the same as getting a better output. That is the trap of the last four years. It is how chat AI was introduced, and it is why nearly everyone still asks for things one line at a time.

"Did the tests pass" is a slot code. "Assume this is wrong, go prove it against the real system, and tell me what I did not ask" is a thinking partner. Same machine. Completely different posture.

What happened when I kept reading every plan myself

I was building features with Claude Code. Every feature got a written plan first, and I read every plan before it ran. That was the system. Me at the gate.

It worked, right up until it didn't.

Sometimes I would catch a mistake and change the plan. Then the plans got too complex and I did not even understand what they were saying. And here is the part that should worry you: even the ones I did correct executed with a ton of errors it had to go back and fix.

Read that twice. The reviews that worked did not work.

My instinct was to read harder. Slow down. Ask for shorter plans. That is what everybody does, and it is the same move as making your pipeline faster.

What actually fixed it was a second agent whose only job is to assume every line is wrong and go prove it against the live database. Invented table names. Invented column names. All of it caught by something that did not trust the plan and did not need me to understand it.

Almost 70% fewer errors. Not from reading the plan harder. From putting a reviewer in the pipeline that did not need me to understand it.

I did not get better at being the reviewer. I built a reviewer. You cannot scale your judgment by working harder inside it. You scale it by putting the judgment somewhere that runs without you.

Your CI is that same machine. Which is why "we need a faster pipeline" is the wrong sentence.

The thousand-dollar loop, and your recurring red build

Two and a half weeks and about a thousand dollars of tokens into one build. Going in circles. I finally said it out loud: we are back where we were two and a half weeks ago. This is a cycle and we are not getting anywhere.

The way out was not more effort in the same place. I had the model write the spec of everything done so far, took that spec to a rival model, and it came back with ideas the first one called novel.

Your pipeline has the same failure mode. Same check red for the third week running. Same fix applied again. Same green. Same red. That is not a bug. That is a loop with no exit, and the exit is not another hour in the thread. When a working thread starts repeating itself, get the full spec of the work so far and take it somewhere else before you spend another hour in it.

Inconsistent vs consistent

Owner-dependent CIOrganizationally independent CI
What green meansWhatever you decide it means that dayOne sentence everyone on the team can finish
Where the judgment livesIn your headIn the check
Who the pipeline reports toYou, in a thread, at 11pmA person by name, with a signal
When a build goes redYou open the plan and read itThe check proves the failure against the live system
When it breaks againAnother review, by youAnother check, owned by someone
What you measureHow many builds are greenHow many calls it made without you
You take two weeks offThe queue becomes yours to dig outThe pipeline keeps deciding

What this costs you, and why you will resist it

Nobody tells you this part. You will hate the first month.

Your pipeline gets worse before it gets better. It flags things you would have waved through in three seconds. It goes red on a Friday for a reason you disagree with. And your hand moves toward the file on its own, because reading the diff is what you know, and handing the call to a check feels like lowering the bar.

The cost is real. A week of work that ships no feature. A skill you are proud of, deliberately set down. A meeting where you say the sentence out loud and let it stand.

The rewiring is one reflex. Old reflex: red build, open the file, fix it. New reflex: red build, ask which check should have caught this and why it didn't, then strengthen the check. Same ten minutes out of your day. Completely different company a year from now.

That is the trade. Less of you in the loop, and a pipeline that can carry your judgment at 3am without waking you up.

When not to do this

Do not touch the pipeline if you cannot say, in one sentence, what green means. That is an alignment problem wearing a CI costume, and speeding up a pipeline that measures the wrong thing just gets you to the wrong answer faster. Write the sentence first.

Do not do this if the pipeline is genuinely unreliable at the plumbing level. Flaky infrastructure is a real job. Fix the plumbing, then come back and fix the thinking. Do not dress a broken machine up as a leadership problem.

And do not start with the checks. The checks are the cheapest part and the most tempting, because they feel like progress. The expensive part is Movement. That is the one you will want to skip, because it requires you to stop being the answer.

Then close the loop. Take the last red build you personally rescued and ask the question you have been avoiding: which check should have caught this, and why didn't it exist? Give that check a name. Then let it be right without you.

A stronger organization. A freer you.

<!-- cluster-links:begin -->

Related reading

<!-- cluster-links:end -->

Questions

What people ask next

If I actually let the team decide without me and they get it wrong, do I step back in, or does that just prove I was right to be the gatekeeper all along?
Getting it wrong isn't proof you were right to gatekeep, it's a signal that belongs back in alignment, not in your hands. The question isn't 'was the decision correct,' it's 'did green ever mean what we said it meant, and does the definition or ownership need to change.' Stepping back in as the fixer just re-installs you as the exception you were trying to remove.
How do I actually tell which checks are the real load-bearing ones versus the ceremony, without that judgment call becoming its own bottleneck?
You do it once, in alignment, not continuously: write the sentence for what green actually means about the product and the risk you're willing to carry, then work backward to the smallest set of checks that can prove that sentence. Anything that doesn't trace back to that sentence is ceremony, and you'll know because nobody could tell you what it protects or when it last caught something real.
Once I stop being the exception and nobody needs to text me anymore, what's actually left for me to do?
Alignment stays yours: the vision, the ownership of the Big 3, and rewriting the sentence for what green means when a red build teaches you it was wrong. That's the heaviest of the three questions and it never stops needing you, it just stops needing you at 4:55 on a Friday.
Won't 'how many calls this made without me' just turn into another number people quietly learn to game?
It resists gaming better than a green-count because it isn't one number, it's a name on every check plus a real signal: what the check protects, how it fails, and when it last caught something. A gamed check shows up fast under that scrutiny because someone has to answer 'what does the check say' with an honest 'I don't know,' and that answer is allowed to stand.
Is this really specific to CI, or should I be running this same three-question loop on every other process where I'm still secretly the approval gate?
CI is just the first place the question showed up loud enough to notice. The three questions and the loop describe Organizational Independence generally, so anywhere you're still the one saying yes or no, alignment on what a decision means, movement on whether people can act without you, and measurement on whether they were right, applies the same way.