Analytics metrics review card with baselines and counterfactual warning. Analytics tools metrics: the 2027 review
Photo by Intelligence Dashboard Metrics on card

Maintenance

Part of Best analytics tools for 2027: a ranked shortlist

Analytics tools metrics: the 2027 review

Analytics tooling measured honestly: the baselines to take before buying, depth rather than breadth afterwards, and the counterfactual nobody escapes.

A tool purchase gets justified with a business case and then never measured. Two years later nobody can say whether it worked, and the renewal conversation runs on impressions.

This page is about measuring an analytics tooling decision honestly: what to record before you buy, what to watch afterwards, and why most of the obvious measures are worthless.

What to take away

  • Decide the measures before the purchase, because every measure chosen afterwards will be the one that justifies the decision already made.
  • License counts and login counts are the two most reported numbers here and neither one indicates anything about value.
  • The honest problem is the counterfactual: you cannot see what the same effort would have produced without the purchase, so measure the mechanism instead of the outcome.

The measures that are usually reported, and what is wrong with them

Commonly reported What it actually tracks Why it misleads
Seats provisioned Procurement activity Grows regardless of use
Monthly logins Habit and curiosity A daily login to check one number is not depth
Dashboards created Builder productivity Creation is cheap and accumulation is a cost
Queries run Machine activity A scheduled refresh inflates this without a human present
Satisfaction survey Recent experience Answered by people who chose the tool

Every one of these rises when the estate gets worse. That is the test of a bad measure: it improves under the failure mode you most want to detect.

Table of five commonly reported analytics measures, what they track, and why they mislead (Analytics tools metrics: the 2027 review)
Every commonly reported measure rises when the estate gets worse, which is the test of a bad measure. Image: Intelligence Dashboard Metrics

What to record before the purchase

Four baselines, taken in the month before anything changes. Without them nothing afterwards is interpretable.

Checklist of four pre-purchase baselines: request queue, time to first answer, artifact use, disputed numbers (Analytics tools metrics: the 2027 review)
Without these four baselines taken in the month before anything changes, nothing afterwards is interpretable. Image: Intelligence Dashboard Metrics
  • The request queue: how many analytical requests arrive per month, and the median age of an open one.
  • Time to first answer for a standard, well-defined question, measured a handful of times.
  • The count of artifacts in use, and how many were opened in the last month.
  • The number of numbers currently in dispute, meaning the same concept calculated two ways in two places.

That last one is subjective and worth recording anyway. It is the measure most likely to move if the purchase is genuinely working, and the one nobody thinks to capture beforehand.

What to watch afterwards

Depth of use, not breadth. The interesting count is people who built or modified something that another person then used. Breadth measures distribution; depth measures whether anyone is doing work in it.

Change in the request queue. A working purchase shows up as fewer arriving requests of the kind the tool was meant to absorb, not as more delivered requests. If delivery rose and arrivals rose with it, you have bought throughput, not self-service.

Retirement rate. How many artifacts were decommissioned this quarter. An estate that only grows is accumulating maintenance, and the retirement number is the only direct evidence that anyone is looking.

Freshness and failure. Share of scheduled refreshes that completed on time, and the median time between a failure and someone noticing. This is the measure that predicts whether people will keep trusting the estate.

Definitional convergence. Track your disputed-numbers list and whether it shrinks. This is slow, subjective, and the closest thing to a real outcome measure available.

The counterfactual problem, stated plainly

You cannot know what the same money and attention would have produced elsewhere, and there is no control group. Any claim that a tool caused an improvement is a claim about a comparison that does not exist.

That does not make measurement pointless. It changes what the measurement is for. You are not proving causation; you are checking whether the mechanism the business case described is actually occurring. The case said self-service would reduce arriving requests. Did arriving requests fall? If they did not, the mechanism failed regardless of whether other things improved.

Be similarly careful with the numbers a supplier's own published results report. Those are drawn from customers who succeeded and agreed to be written about, which is survivorship bias built into the sampling frame before anyone does any arithmetic.

Measures that are worth a target, and measures that are not

Set targets on operational measures: refresh reliability, time to notice a failure, retirement rate. These are within someone's control and gaming them requires doing the actual work.

Do not set targets on adoption. A target on user counts produces mandated logins, and a target on artifacts produced produces artifacts. This is Goodhart's law at its most predictable, and adoption measures are unusually easy to satisfy without doing anything useful.

The renewal conversation

Six weeks before renewal, write one page: the mechanism the original case claimed, the four baselines, the same four measures now, and a plain statement of what you would do differently. Circulate it before anyone asks.

A team that produces this once earns an enormous amount of latitude on the next purchase, mostly because almost nobody does it.

Related reading on this site

Purchase decisions are covered in buying analytics tooling, and comparison and trial design in how to compare intelligence platforms.

For measuring artifacts rather than the tool, see what to measure about a page and an estate. Queue reading behind request measures is in business intelligence. Skeptical reading of any vendor-published result is in how to read a platform case study.

Common questions

We already bought it and took no baseline. Is it too late?

Mostly, for the before-and-after comparison. You can still take the baselines now and use them for the renewal decision and the next purchase, which is where they matter more anyway.

Our supplier provides a usage dashboard. Can we use that?

For operational measures, yes. For anything about value, no: it reports the measures that make the product look used, and depth of use is generally not among them.

How long before any of this shows a signal?

Two quarters for queue and freshness measures, a year for definitional convergence, and never for anything that requires a counterfactual. Say so up front rather than promising a six-month verdict.

What if the honest answer is that nothing changed?

Then write that down, because it is the most valuable finding you will produce. The alternative is a renewal justified by the same impressions that justified the purchase.

More in Maintenance

Latest from Analysis Desk