Measured versus judged: the two kinds of criteria
A spec criterion is a number nobody argues with; a rated one is a judgment with evidence. How each becomes a 0–100 score, and why the split matters for trusting a grade.
The framework has 102 criteria, and they come in two kinds that are scored in completely different ways. Knowing which kind you are looking at is the single most useful thing for deciding how much to trust a number on a scorecard, so this guide is about the split.
Measured criteria
A measured criterion — a spec, in the framework’s own vocabulary — is a real quantity with a unit: towing capacity in pounds, combined fuel economy in MPG, rear legroom in inches, base price in dollars, a 0–60 time in seconds. There are 32 of them. The scorecard records the value, the unit, and where it came from, and the reader sees the figure itself.
What the reader does not see is anyone deciding how good that figure is. The 0–100 score for a measured criterion is computed from a published normalization band for the vehicle’s segment, never authored. A band is a minimum and a maximum; the value’s position between them, as a fraction, is the score. For criteria where lower is better — price, 0–60 time, turning circle, annual fuel cost — the fraction is inverted.
The consequence is that the research is the whole job for a measured criterion. Get the value right and the score follows. Get it wrong and the score is wrong in a way anyone can check — which is why the value, the unit and the source are printed next to every one. The bands themselves are published on the report card, and Graded within its segment: why a truck's MPG is not a sedan's explains why they differ by body type.
Judged criteria
A judged criterion — rated, in the framework — is one where no honest number exists. Ride quality. Predicted reliability. How responsive the infotainment is. Whether the headlights are any good. There are 70 of them, and they are scored 0–100 directly, against written guidance that describes what an 80–100, a 50–79 and a 0–49 look like for that criterion.
A judged score has to carry its own justification, so each one is stored with two things a measured score does not need: a reasoning sentence or two explaining why it landed where it did, and an evidence field naming what that reasoning rests on — an IIHS headlight rating, a Consumer Reports reliability prediction, a manufacturer’s own equipment list, a published road test. Both are shown on the car page. A judgment with no visible basis is not published as a score; it is published as a lower-confidence score, which brings us to the third field.
Confidence
Every criterion, of either kind, carries a confidence level: HIGH when the figure or judgment rests on a primary source, MEDIUM when it rests on a reputable secondary one, LOW when the sourcing is thin, and INFERRED when it is an informed estimate with no citation — the evidence field is empty by rule in that case, so an inference can never masquerade as a sourced fact. The scorecard’s headline confidence, shown on the car page, summarizes the split across its criteria; a MEDIUM scorecard is the ordinary case — most figures are documented but not every one is straight from the manufacturer — and it is a reason to check a figure that matters to you against the source named beside it, not a reason to distrust the whole card. How the whole fleet splits across the four levels is on how we grade, with the source order the research follows.
Why the split matters
Subjects differ enormously in their mix. Efficiency & Range is 8 measured criteria out of 9: it is nearly all arithmetic over EPA figures, and its grade is as reliable as the figures. Reliability & Quality is 8 judged criteria out of 8: there is no such thing as a measured reliability, so the grade is a synthesis of published predictions, recall history and build-quality assessments, and it should be read as one. Neither is worse; they are different kinds of claim, and the framework keeps them visibly separate rather than blending them into a number that pretends to be measured.
The other consequence is about correction. If a measured value is wrong, send the right one with its source and the score recomputes itself. If a judged score seems wrong, the argument has to be with the reasoning — and it is printed, so it can be. How we grade sets out the source order and what the validator refuses to publish.