Close-up of hands reviewing business report with colorful charts and graphs on a wooden desk. Inside a reporting rebuild: why the average was the wrong number
Photo by RDNE Stock project on Pexels

Costs

Part of Can a reporting pack survive contact with its own audience?

Inside a reporting rebuild: why the average was the wrong number

Reporting rebuild worked through on invented figures: why a rising average meant a mix shift, and what replaced the headline number afterwards.

This is a worked example, not a real company. Every figure below is invented to make the arithmetic visible, and the point is the sequence of decisions rather than the numbers themselves.

The situation: a support organization publishes a monthly pack. Its headline measure is average resolution time. For four months the average has risen, four different explanations have been offered in the meeting, and nothing anyone did has moved it.

What to take away

  • The pack was not wrong. It was answering a question nobody had asked, using a summary that could not show the thing that had changed.
  • The fix cost one afternoon of arithmetic and no new tooling. What it cost politically was giving up a number people had grown attached to.
  • The measure that finally worked was less flattering and much harder to argue with, which is the usual shape of a good outcome here.

The pack as it stood

One page, one headline number, one line chart, one paragraph of commentary. The invented figures:

Month Cases closed Average resolution time, hours
One 1,000 20.0
Two 1,000 22.0
Three 1,000 25.5
Four 1,000 28.0

Volume was flat, so the natural reading was that the team had slowed down. The commentary each month said some version of that, and the interventions followed: a staffing review, a push on first contact resolution, a training session.

The first afternoon: splitting the total

The team had two case types in the source data and had never reported them apart. Splitting month one against month four gave this:

Case type Month one volume Month one average Month four volume Month four average
Standard 900 10.0 600 10.0
Escalated 100 110.0 400 55.0

Neither type had got slower. Standard cases were unchanged. Escalated cases had actually halved in duration. The overall average rose because the share of escalated cases went from a tenth to nearly half.

This is the everyday form of Simpson's paradox: every group improved, and the total went the other way, because the weights moved. The four months of interventions had been aimed at a slowdown that never happened.

The second finding: the average was the wrong summary

Even within standard cases, the mean was doing poor work. The invented distribution for month four: most standard cases closed within two hours, a long tail ran to several days, and a handful sat open for weeks awaiting a customer reply.

A mean sits wherever the tail drags it. The median and a couple of upper percentiles describe what an actual customer experiences, and they move for different reasons. A quick histogram of closure times made this obvious in a way that four months of commentary had not, because the shape of the distribution was visibly two humps rather than one.

What the pack looks like now

The rebuild was deliberately small.

  • Volume and mix on the page, above the timing measures, because mix is what moves the total.
  • Timing reported per case type, never blended.
  • Median and an upper percentile per type, with the mean dropped entirely.
  • A separate count of cases paused awaiting the customer, excluded from the timing measures and reported on its own.
  • One paragraph of commentary, written only where a measure moved outside its band.

The pack got shorter. The headline number disappeared, which was the hardest part to agree.

What failed in the rebuild

Two things, both worth recording.

The percentile choice was argued about for a fortnight. Nobody could settle which upper percentile to publish, and the discussion stalled the release twice. The resolution was to publish one and revisit after six months rather than to decide correctly in advance.

The paused-case measure was gamed within two periods. Once pausing a case removed it from the timing measures, cases got paused more readily. The counter was to report the paused count and its age distribution beside the timing measures, so a rise in pausing is visible in the same glance. That pairing should have been designed in from the start.

What generalizes

Very little of this was about reporting technique. The sequence that transfers is: split the total before explaining it, look at the distribution before choosing a summary, and expect any measure that can be avoided to be avoided.

The other transferable part is the political one. Retiring a headline number is harder than adding five better ones, and it works best when the replacement is presented as a decomposition of the old measure rather than as a repudiation of it.

Related reading on this site

The general production discipline is in building a report worth keeping, and the commentary technique used in the rebuilt pack is in writing the paragraph beside the numbers. For the analysis sequence that produced the split, see how to work through a question. The summary choice is covered in metric shapes and their traps, and the page-level version of the same errors is in nine dashboard errors.

Common questions

Is this a real case?

No. It is a composite written to make the arithmetic visible, and every number in it is invented. The pattern is common enough that the shape will be familiar even though the figures are not.

Our source data has no case type field. What then?

Then you cannot explain movements in the total, only observe them, and that limitation belongs in the commentary. It is also the strongest argument you will ever have for changing the source system.

Should we always drop the mean?

No. Keep it where the distribution is roughly symmetric and where somebody genuinely needs a total divided by a count, such as capacity planning. Drop it where a long tail is doing the talking.

How do we stop the next measure from being gamed?

Assume it will be, and name the behavior that would show it before you publish. Then report that behavior beside the measure. You will not prevent the response, but you will be able to see it.

More in Costs

Latest from Analysis Desk