A Team at 95% Busy Waits 3.4 Times Longer Than at 85%
In the simplest queue model, average wait grows as utilization divided by one minus utilization. I work the numbers by hand, add variability, and test the rule against its best critique.
Here is the claim, with plain numbers. In a queue with one server and random arrivals, the average wait grows as , where is utilization. At 85% busy the factor is 5.67. At 95% busy it is 19. The wait is 3.35 times longer, so "about 3.4". A target of "fully busy" does not signal efficiency. In this model it produces the long waits that customers feel.
I will derive this from the textbook formula, add variability with Kingman's approximation, and then give the best objection its full weight. Part of the "85% rule" folklore is wrong, and I will say which part. The curve is right. The magic number is not.
One slow box
Draw the process as boxes: arrive, wait, work, done. Only "work" has a server. Every item that arrives when the server is busy waits.
Take round numbers. A person handles one kind of request. Each takes 1 hour on average. Arrivals are random, and the service time is random with the same spread as the arrival gaps (the M/M/1 case). Then the mean time an item waits before work starts is [2]:
Here is the mean service time (1 hour) and is arrival rate divided by service rate. I computed every number below by hand from this formula. I did not use the Lab.
| Utilization | Factor ρ/(1-ρ) | Wait before work | Total time in system |
|---|---|---|---|
| 50% | 1.00 | 1.0 h | 2 h |
| 70% | 2.33 | 2.3 h | 3.3 h |
| 80% | 4.00 | 4.0 h | 5 h |
| 85% | 5.67 | 5.7 h | 6.7 h |
| 90% | 9.00 | 9.0 h | 10 h |
| 95% | 19.0 | 19 h | 20 h |
| 99% | 99.0 | 99 h | 100 h |
Total time in system is . Read the rows. The work itself takes 1 hour in every row. Only the waiting changes. At 95%, an item spends 19 hours waiting for 1 hour of work. Where does it wait? In the line in front of one busy person.
The step from 85% to 95% looks small. It is 10 points. It multiplies the wait by 3.35. The step from 50% to 70% is 20 points and multiplies it by only 2.3. Each extra point costs more than the last. That is the shape of the curve, and it is why I dislike full utilization as a goal.
Little's law shows the same thing as a headcount
Little's law says work in system equals rate times time: [1]. John Little proved in 1961 that it holds over a long observation period [1]. It needs few assumptions, so it works as a check on any process.
Apply it. At 85%, the arrival rate is 0.85 items per hour and the time in system is 6.67 hours. So items in the system. At 95%, the rate is 0.95 per hour and the time is 20 hours. So items.
The number of items in the system equals the same factor, . That is not an accident. It is the queue seen from the inside. At 95% busy, a manager who walks the floor sees 19 open items for one person. At 85%, the manager sees under 6. Same person, same task, 10 points of load.
The same law explains why a backlog hurts even when the team "keeps up". If the backlog is 19 items and the team finishes 0.95 per hour, a new item waits about 20 hours. The queue is the delay.
Variability changes the curve, and a manager can control it
The M/M/1 model sets the spread of arrivals and service equal to that of the exponential distribution. Real work differs. Kingman's approximation covers the general case [2]:
Here is the coefficient of variation of arrival gaps and is that of service times. For the M/M/1 case both equal 1, and the middle term is 1 [2]. The formula is an approximation for one server. It is known to be accurate mainly near saturation [2]. So treat the next numbers as estimates.
Take a team with steady arrivals of the same random type () but two kinds of service work.
- Standard work. Almost every item takes the same time: . The middle term is . At 95%, the wait is hours.
- Mixed work. Some items take 10 minutes, some take 6 hours: . The middle term is . At 85%, the wait is hours.
Read that again. The team with mixed work at 85% waits longer than the team with standard work at 95%. Utilization is only one of the three terms. Variability is another. A manager who cuts the spread of item sizes can get more wait reduction than a manager who adds a little capacity.
This is where lean and queueing thinking agree. Small batches, standard work and level release all lower and . They move the curve down. A utilization target moves a team along the curve, and usually in the wrong direction.
The strongest objection
The objection has three parts, and I take it seriously.
First: idle people cost money. A team at 60% pays for 40% of time that produces nothing. A finance reader will not accept that cost without a reason. This is true. The model does not say that low utilization is free. It says that high utilization has a price, paid in wait, and the price rises steeply. The right target comes from setting the cost of one hour of delay against the cost of one hour of spare capacity. I do not have a universal number for that, and nobody does.
Second: the "85% rule" is a fallacy. The best critique I found is by Nathan Proudlove. He studied the 85% bed occupancy target in hospitals and found that it came from a simulation with other conditions than the ones where people applied it [3]. For a pediatric ward with mostly emergency admissions, an 85% target meant a 33% risk that all beds were full [3]. To get that risk down to 0.1%, the ward needed about 55% average occupancy [3]. A separate paper carries the title "The 85% bed occupancy fallacy" [4]. I read only its listing, so I cite it for the title and nothing else.
I concede the point. My working title used "85% rule", and the number 85 is not a law. Small pools of servers need much lower average load, because random swings are large relative to the pool. A large pool can run hotter. The M/M/1 curve I used has one server, so it cannot show this. The lesson from Proudlove is to compute the curve for your own system, not to copy 85 from a poster [3].
Third: demand is not steady. My formulas assume that arrival and service statistics stay the same over time. Real demand has peaks, seasons and bad Mondays. This is a real limit, and I know I lean too hard on steady-state models. Two points remain. First, peaks make the problem worse. During a peak, local utilization rises, and the wait follows the steep part of the curve. Second, a steady-state formula gives the best case. If the best case already shows a 19-hour wait at 95%, the real wait will not be better.
What remains of my view after the objection? The shape holds. Wait rises faster than load as load nears 100%. The location of the knee depends on variability and on the number of servers, and a manager must find it for each process.
What follows if I am right
If the model is right, a year-end target such as "keep every person above 90% busy" will do the reverse of its intent. It will raise the time that customers wait, and it will raise the work in progress that people must track. Teams will also hide the problem. They will fill the queue, because a full queue guarantees that nobody is idle.
Three changes follow, in order of cost.
- Measure wait, not busyness. Report the time an item sits before work starts. Customers feel that number. Effort metrics do not show it.
- Cut variability before you add people. Split large items. Release work at an even pace. Standardize the common case. Each change lowers the Kingman middle term.
- Set the load target from the wait you can accept. For a single server, choose the wait, then solve for . If the item takes 1 hour and you accept a 2-hour wait, the M/M/1 factor must be 2, which gives , about 67%. A 4-hour limit gives . For a team, check with a model that has the right number of servers, then confirm with your own data.
One line names the bottleneck: the slow box is the busiest person, because the busiest box sets the queue for the whole line.
What would change my mind? I would drop the strong claim for a given team if its measured wait stayed flat while utilization rose from 85% to 95%. That outcome is possible if the pool of servers is large or the work is very regular. It would mean the team sits on the flat part of its own curve. I would like to see that data. The Lab run I have planned will plot waits at 70, 85, 95 and 99 percent, so readers can compare the hand formulas with a measured curve.