The Formula Hides the Argument
A game review aggregator recently tried to fix a real problem: Metacritic freezes a verdict at launch, so a patched, redeemed game carries a five-year-old number forever, no matter how much work went into fixing it after release. The proposed fix reweights old critic scores against how well they track current player sentiment, so the number drifts toward the game's present state instead of its first week. That's a genuine improvement on staleness. It does nothing about the actual fault line, which is that a 7/10 means something different depending on who typed it, and no amount of reweighting tells you whether the crowd's opinion and a critic's opinion should count the same in the first place. That's still a shrug. It just arrives with more decimal places.
Credit scoring did the same move decades earlier. Lenders started with a mess of inconsistent human judgment about who counts as creditworthy, then built a formula out of proxies: payment history, utilization, length of history, all weighted and summed to three digits. The formula is precise. The question it's standing in for, whether this specific person should get this specific loan, was never a math problem, and building a more elaborate weighting scheme around it doesn't resolve the question, it just relocates the person who has to answer it from a loan officer to whoever set the weights years earlier and is no longer in the room.
Grading curves and GDP deflators run the same trick. A rough, openly subjective call, this student did well, this economy grew, turns out to be inconsistent across raters or years, so the fix is item-response theory or a chained price index. Both genuinely solve a comparability problem. Neither one touches the prior question of what "doing well" or "growth" was supposed to mean, because that question was never arithmetic. It got a procedure attached to it, and the procedure now does the talking.
Once you notice the pattern it's everywhere a contested measurement exists: the response to "this number is wrong" is almost never "here's what we actually mean, stated plainly," it's "here's a more elaborate way to compute the same number." The elaboration doesn't answer the objection. It moves the objection into a coefficient nobody outside the methodology section ever reads. The tell is you can still ask, weighted how, chosen by whom, compared to what, and get silence, just silence wearing more digits than it used to.
Precision and validity aren't the same axis, and a number doesn't earn extra decimal places by being more correct. It earns them by coming out of a longer procedure, and longer procedures generate decimals whether or not there's anything at that resolution worth measuring. The eye reads three decimal places as rigor. The eye is guessing.
None of this makes reweighting worthless. Comparable methodology beats no methodology, a consistent proxy beats an inconsistent one, a chained deflator really does fix a real accounting error. The complaint is narrower than "formulas bad": when a measurement gets a fancier formula in direct response to a complaint about what it's actually measuring, check whether the fix answers the complaint or just outsources it to a part of the pipeline nobody argues about anymore.
The honest version of any of these numbers would state the judgment in words first, then let the arithmetic compute exactly that and nothing more. A 7/10 that says "worse than the median game we reviewed this year, better than the worst" is a claim you can argue with. A weighted composite that produces 7.3 is a claim dressed as a measurement, and it's harder to argue with precisely because it no longer sounds like an argument.