I asked Claude Code for a plan a few weeks ago. Live demos for a talk I am giving at GTC in October, the sort that have to work on stage in front of people. Rough was fine and I said so: give me a direction, tell me what you are not sure about, and we can spike the gaps. Twenty minutes later I had a list of things to work through (with Claude, little did it know!). The things that needed to be built, the order they needed to happen in, a few unknowns named explicitly as things somebody (again Claude ;) ) would have to go and check. And an estimate. Two days.
Nobody in that exchange said the word “points”. Certainly not me. There is no story point mode. It handed me a number in human time because human time is the only unit it has. A question teams have argued about for fifteen years got settled for me by a piece of software that was not consulting anybody.
It was also wrong. I had agreed a partial version for a date a week later, and it landed in six hours. Most of those were spent steering between meetings, which is more often my reality than I would like. But it is wrong in an interesting direction. A year ago the same question came back as a week. The estimates are shrinking, and my guess is that the models have started pricing in the fact that an agent is going to do a good deal of the work.
None of which is about the estimate. The part that matters is the twenty minutes. Getting to a breakdown used to be expensive enough that we built a ceremony to guarantee it happened. That is no longer true. Which does not mean I am about to tell you to stop pointing, whatever the headline suggests. I manage managers, and what a team does in its planning session is not my call. I will come back to that.
What interests me is what it was doing while nobody was looking, and what has to take that job over.
What the ceremony was for
Story points were sold as complexity-relative sizing, which was never the job. The job was to force a conversation that would otherwise not happen. What does this touch. What do we not know yet. Where does it get hard. The number was the outcome. The conversation was the actually important thing. Though we didn’t often acknowledge that given the obsession with velocity numbers. I once told a founder I could show them ever increasing velocity numbers, but there was no guarantee any of it had any business impact. That blew their minds.
So the conversation was worth having. The ritual we built around it is another matter. I never liked planning poker and I never believed it earned its place. The person doing the work should size it and present the number, and should not have to defend it to the nth degree. Their manager should check the number, because that is the job. Asking eight people to vote on a story most of them have not read is not collective intelligence. It is a meeting.
The expensive parts got cheap
All of this is from my own work rather than my teams’.
The breakdown. Surface area, dependency order, the unknowns made known. It is not perfect, but it is close enough to work from, and closer than what most teams I have watched get to in a room with a whiteboard.
The spike. Minutes, or an afternoon, against the week or three it used to take. And part of why points existed at all was to put a box around spikes, because engineers are bad at time-boxing themselves and always have been. Exploring is fun! I get it. When the spike costs an afternoon, the box stops mattering very much.
The obvious objection is that AI plans are often wrong. I do not think that is the right description. They do not know what they have not been told, and the ones that come back thin come back thin because the context was thin. Steering them is a skill. That is an argument for reading the plan properly and treating it as a draft, not an argument for keeping story points. It does mean the list of unknowns is only as good as the context you gave it, which matters later, when I start holding people to that list.
There will still be expensive surprises. Some things you do not learn until production, and no amount of planning ever changed that. What has changed is that you know what has already been tried. Write the decision down, note what you attempted and why you abandoned it, and revisit it when something shifts. That is what ADRs were always for, and they have got more useful now that trying something costs so little.
Which of the other jobs were real
If the conversation was the point, and the conversation is now cheap, the question is what else the ceremony was doing, and whether those jobs were as real as we tell ourselves. Some of them were.
Calibration across the team. Real, and oversold. It mattered, and it mattered most to the engineers who need to be heard on everything the team is doing, which is a legitimate need and not one I would dismiss. But the business only ever wanted rough delivery timelines. We dressed that up as a calibration exercise, and then handed back a unitless number instead of an answer. That is a large part of why the business got so angry with engineering about story points in the first place.
The shield. Real, and it never worked. Points are unitless by design so that a number cannot be converted into a commitment on its way up the organisation. I have used that shield. As an engineer I occasionally made it a good deal bigger than it needed to be, precisely so that nobody could pin a date on me. And it did not work, because every manager I know, including me, takes the number and applies a fudge factor based on what they know about that engineer or that team. The shield bought vagueness. It never bought protection.
A forcing function for the quiet engineer. Mostly not true, in my experience. The loud voices won the estimation arguments the same way they won every other argument, and half the room had zoned out by the third story. Getting a quiet engineer’s objection into the open is an engineering manager’s job. You do it by asking who does not think this will work, or by admitting you are not sure yourself so that it is safe for somebody else to say it too. That is management, not a ceremony.
A licence to say “I don’t know yet”. No. If a team needs a ritual to make that sentence sayable, the ritual is not the thing that needs fixing. “I don’t know, and I am going to find out” should always have been an acceptable answer.
Read those back and none of them is really about estimation. Getting a quiet voice heard, admitting you do not know something, and holding a shield up against your own management are all workarounds for not trusting the people above you. The ceremony was where we put them because there was nowhere else to put them. Hold that thought, because none of it is fixed by a faster breakdown.
What replaces the shield
Which leaves the actual problem, and it is mine as much as anyone’s. I need commitment on dates from my engineering managers. The business has always needed rough delivery timelines. Points were partly a shield against exactly that, and if the shield is going, something has to take its place.
The reflex answer is the iron triangle: fix the date and flex the scope, or the reverse. I do not think that is the answer any more, because the build is no longer what sets the date. Flexing scope was a way of managing uncertainty in the build, and it assumed the build dominates the schedule. When the plan says two days and the work takes three hours, trading scope barely moves the date. You are negotiating the wrong variable.
What sets a date now is how long it takes to get agreement from the people who have to give it, how long verification and integration take, and what you find when you go looking. None of those are scope.
Where the uncertainty gathers around a small number of specific things, list them and put a date on each. Points hid the uncertainty inside the unit, which was the whole trick: a number nobody could convert because nobody could interpret it. This does the opposite. Not “three days” but three days if the payment provider behaves like its documentation, and we will know by Tuesday.
That is a commitment I can take to somebody, and one I can be held to fairly, because it arrives carrying the thing that would prove it wrong. It only became affordable in the last couple of years. Getting to a list of unknowns used to need the two-day spike; now it takes twenty minutes.
Where you can cut the work small instead, do that, and the conditions stop mattering very much. That option used to be available on far less work than it is now. Slicing something into weekly drops only helps if a week of building produces something you can show, and for a great deal of work it did not. The cheap build is what changed that.
The fudge factor was always the interface
Somebody always absorbs the uncertainty. Points hid the engineer’s share of it from the business entirely; conditions move it to the manager, which is where it already sat.
I said above that every manager applies a fudge factor. I have always treated that as a slightly grubby habit, the thing you do not put in the process document. I think that is wrong. It is the translation layer between conditional information inside the team and a committed date outside it. We just never called it anything.
What changes is the quality of what I am fudging against. Judgement applied to “three days if the provider behaves, known Tuesday” is a different act from judgement applied to an eight.
What the customer gets
None of this means handing a customer a list of conditions. They will not thank you for it and it is not what they asked for. They get a date somebody is willing to own. It is a better date because the conditions were worked out before it was given, not discovered afterwards.
For the demos I cut the work small instead of listing conditions. I agreed a partial for a date this week. It landed in six hours against a two-day estimate. The two after it go out weekly. Nobody on the other end is tracking a condition, because there is nothing to track. The increment is a week and the thing either turns up or it does not.
That is the shape I would reach for first, wherever the work allows it: a cadence of drops small enough that a miss shows up immediately, instead of a single date at the end. What the customer needs is no surprises. They are planning rollouts and training and internal comms around you, and certainty was never really on offer. Small drops give them that for free. Where the work will not slice, conditions are how you get there instead, and you find out on Tuesday rather than in week three.
The obvious way this goes wrong
Tuesday, unless X. Then on Tuesday, but Y. Then Friday, but Z. Every instalment perfectly defensible and the thing never lands. Uncertainty is no longer hidden in the unit; it is hidden in an unending supply of legitimate reasons. This is the strongest objection to any of it, and it is also where conditions earn their place, for one reason. They can be wrong. “This is an eight” cannot be wrong, because there is nothing in it to check. “Tuesday unless X” is right or wrong by Tuesday.
Which gives you something specific to hold a manager to: was the condition on the list? If Tuesday’s reason is Y, and nobody named Y, that is not a condition resolving. It is a new unknown appearing, which means the breakdown missed it. Once is ordinary. Three times in three weeks is either a breakdown nobody is doing properly, or genuinely unfamiliar territory where you should be spiking rather than committing.
It also makes the fudge factor measurable, because you can look back and see how often a manager’s stated conditions held. Points never let you do that. And it tells you when to stop re-forecasting and change mode instead. After the second or third unnamed condition, stop taking conditional dates on that piece of work. Put a hard box round it and ship what exists on Friday, or accept that it is discovery and stop pretending it has a date at all.
Things to watch out for
Work that is never finished. A product surface that keeps evolving has no terminal state. There is nothing to put a date on, so drops are all you have: this is how often we ship, this is what is next, this is what we have stopped doing. That was equally true before, and it is where the ceremony was at its most performative. Pointing an endless stream of work produces a velocity number that measures nothing. If the business is demanding a completion date for something that has no completion, the estimation method is not the problem.
Somebody also has to keep the list. If nobody tracks which conditions were named, you get the excuse machine with better vocabulary. That is worse than points, because now it sounds rigorous. This falls on management in a different way.
And none of this survives in an organisation that punishes a missed date regardless of whether the stated condition held. I suspect that describes a lot of organisations, which makes it the strongest argument against everything above. In that place the shield was entirely rational. Naming a condition hands somebody a more specific thing to be angry about. Making an estimate checkable where trust is low gives management a stick it did not previously have. Remember what those ceremony jobs turned out to be. If you remove the workaround without repairing the thing it was working around, a team will have built the next one by Christmas. That is a trust problem, and no estimation method has ever solved one.
Not my call, and not yours either
I manage managers. My engineering managers run their own process, and I have no intention of taking that away from them. I wanted to say this, because there is a version of this post that ends with a director reading it on a Sunday and telling four teams to stop pointing on Monday. That director is making precisely the mistake of the one who mandated pointing in the first place: deciding, from a level up, what a team needs in order to think.
The job at my altitude is to make sure a team can say out loud what the ceremony was doing for them, so that if they drop it they are doing so intentionally rather than leaving a hole and finding out about it later. Some of them will keep it, and that is a legitimate answer.
That is most of what empowering a manager means, and it is less comfortable than it sounds. It means letting an EM keep a ritual I think is finished, because they can tell me what it does for their team and I cannot. My leverage is the quality of the question I ask, not the decision I make.
Process where you need it
You owe scrum nothing. You owe kanban nothing. You certainly owe SAFe nothing. Have process where you need it and only where you need it. Performative process should not exist. Not less of it, not watched more carefully. It should not be there at all. That is an ideal and I know it, because the stuff accumulates anyway, in every organisation, without anybody deciding to add it. Nobody schedules a ceremony intending it to become theatre. It gets there by outliving the reason it was introduced.
So the point here is not about making one purge. It is noticing. And the noticing does not stop at planning. Estimation, stand-ups, retros, review gates: the whole shape of how a team has agreed to work deserves the same question. Some of it will survive easily. The things with a human element, retrospectives most of all, are doing something an agent has no view on. Whether they should still run at the same cadence, in the same shape, with the same people, is a bigger question than this post.
Where this leaves estimating
Story points were not a good answer. They were an okay answer to a real problem, and getting to the breakdown used to be expensive enough to justify them. It is not any more, so they are a worse answer than they were.
I am not against estimating. The business needs to know roughly when things land, it has always needed that, and pretending otherwise is how engineering lost the argument in the first place. What I want gone is the performance around it.
If your teams are going to drop the ceremony, make sure they can name what it was doing for them first. Then go and find the thing that does that job better.