Story points or hour bands: estimating agency work

They answer different questions. Most estimation arguments are really about which question is being asked.

Story points measure relative size and become a schedule only when multiplied by measured velocity. Hour bands communicate cost and commitment to a client. Agencies need both: points for internal sequencing, bands with written assumptions for the commercial conversation. Most estimation disputes are two people answering different questions with the same number.

They answer different questions

A story point is a statement about relative size. This piece of work is about twice the one we did last sprint. It says nothing about calendar time until you multiply it by measured velocity.

An hour band is a statement about cost and commitment. This will take between forty and sixty hours, assuming these things are true.

Confusing them produces the classic failure: a client hears "eight points" as "eight days", or a team converts hours to points by dividing, which discards the only property points have.

Atlassian's estimation guidance describes points as a relative measure of effort, complexity and risk rather than duration, which is the part most teams drop. Risk is doing real work in that definition: two tasks of identical effort can carry different point values because one of them touches something nobody understands.

Why points work internally

Estimating relative size is a thing humans do reasonably well. Estimating absolute duration is a thing humans do badly, consistently, and with confidence.

Points also decouple the estimate from who does the work. A task is the same size whether your fastest developer or your newest one picks it up; velocity absorbs the difference over a few sprints.

The catch is that points are meaningless without measured velocity, and velocity takes two or three sprints to establish. A team quoting points before that has a vocabulary, not an estimation method.

Why the scale skips numbers

Fibonacci spacing exists because precision is false at larger sizes. The gap between 8 and 13 reflects genuine uncertainty; a 10 or 11 pretends to a confidence nobody has. The intended reading is proportional: Atlassian describes a 5 as roughly twice as complex as a 3 and about half as difficult as an 8. If your team cannot make that sentence true about its own numbers, the scale is decorative.

Write the acceptance criteria first

One precondition is skipped more than any other. A vague item cannot be sized, because its scope is unbounded. "Improve system performance" has no honest estimate; "reduce catalogue page load to under two seconds" does. Write the acceptance criteria first, then point it, in that order.

Why clients need bands

A client cannot fund a point. They can fund a range with stated assumptions, and they can decide against a range that is too wide.

Bands also make the uncertainty visible, which is the honest position. A single number implies a confidence that almost never exists on technical work, and it gets treated as a commitment regardless of the caveats attached verbally.

The essential part is the assumptions. A band without written assumptions is a number that will be renegotiated the first time content arrives late or feedback comes from three stakeholders separately.

Band width is information, not weakness. A band of forty to sixty hours says something specific about how much you know; a band of forty to a hundred and forty says you are quoting a discovery phase and should charge for it separately.

Which to use when

Which to use when
SituationUseWhy
Sprint planningStory pointsRelative sizing against measured velocity
Client proposalHour bandsFundable, with assumptions attached
Fixed-fee quoteDecomposed hoursMargin depends on the decomposition
Unknown scopeDiscovery phaseNeither; price the investigation
Change requestHours deltaComparable to the original commitment

The fourth row is the one that saves the most money. Where a line of work has no precedent, do not buffer it: carve it out as a separate discovery phase with its own small fixed fee.

Clients accept "we will spend two days establishing this, then price it" far more readily than a large number with no visible basis. Buffers, by contrast, get negotiated away first because they look like padding, which removes the cover and leaves the risk.

The conversion, done honestly

Here is the arithmetic, because "convert through velocity" is usually said and rarely shown. Take the last six sprints of completed points. Discard nothing, including the bad sprint, because the bad sprint is data. You now have a low, a median and a high.

Divide the estimated points by the high velocity to get your optimistic hours, and by the low velocity to get your pessimistic hours. That range is the band, and it is derived rather than felt. If the band is wider than the client will tolerate, the honest responses are to decompose further or to sell a discovery phase, not to narrow it by asserting.

The capacity correction

Then apply the second correction, which agencies forget more often than the first. Velocity is throughput and capacity is availability, and Microsoft's Azure Boards documentation keeps them separate for good reason: capacity accounts for the hours actually available after holidays, leave and non-working days. A team with a stable velocity still delivers less in a month containing two statutory holidays and a conference.

The agency drag, counted once

In an agency there is a third deduction on top of both. The people doing the work also attend pitches, answer other clients, and cover holidays. A velocity measured across a normal quarter already contains that drag, which is exactly why measured velocity beats a theoretical hours calculation. Do not deduct it twice.

The failure mode nobody names

Once velocity exists as a number, someone in management will want it to go up. That is the moment the method breaks, and it breaks quietly.

Both major tool vendors warn about it in their own documentation. Microsoft's guidance is to use velocity to determine capacity and not to confuse it with a key performance indicator. Atlassian is blunter about the comparison problem: velocity is team-specific and is not a measure for comparing the performance of different teams.

The mechanism is simple and completely rational from the team's side. Points are self-assigned, so a team asked to increase velocity increases the points, not the output. The chart improves, the estimates stop meaning anything, and the first person to notice is a client whose project is late.

Two guardrails are enough. Velocity is never reported outside the delivery team as a performance figure, and it is never compared between teams. Use it for the one thing it is for: converting a size into a range.

Running both without contradiction

Decompose the work until each line resembles something the team has actually done. Estimate those lines in points for sequencing. Convert to hours through measured velocity for the commercial conversation, and state the assumptions the conversion depends on.

Keep the conversion one-directional. Points to hours through velocity is a calculation. Hours to points by division is a habit that destroys the only useful property points have.

Then hold the line in delivery. A written change record (what changed, the effort delta, the decision) sent the same day, with the option to absorb, defer or re-price. Most agencies have a change control clause and have never issued a change order, which is where fixed-fee margin actually goes.

Quote the change in the same unit as the original commitment. A client who was sold a sixty-hour band and is told a change is "three points" has been handed a translation problem, and translation problems get resolved in the client's favour.

When estimation is the symptom

If the same estimate is wrong the same way twice, the problem is not the technique. It is that retrospectives are not happening, so nothing is being learned from the variance.

If nobody can say where a fixed-fee project went over, the problem is that work was not decomposed finely enough to locate the overrun.

If estimates are produced by whoever sold the work rather than whoever will build it, no technique will help. That is a delivery structure problem, and it is the gap a fractional technical project manager is brought in to close: the role, described as a job rather than a title.

A clinical research organisation had a related version of this: the development team understood the target state and could not express it as a plan their VP sponsor could fund or track. Decomposing the program into 159 tracked work items with acceptance criteria, sequenced with dependencies made explicit, is what made it fundable: the estimate was downstream of the decomposition, not a substitute for it.

Start a conversationMore insights