How we grade
Every vehicle on MotorGPA sits the same exam — 102 criteria across 10 subject areas — and the result is published in full: the figure behind each score, where it came from, and the reasoning that turned it into a grade. This page is the standard that work is held to.
Where the numbers come from, in order
Research follows a fixed source order, and it is enforced by the tooling rather than left to judgment. Production history and generation boundaries come from Wikipedia’s US model-year records first. Specifications come from the manufacturer’s own documentation next — model pages, spec sheets, brochure PDFs and press releases — alongside the federal and independent testers: fueleconomy.gov for EPA economy and range, nhtsa.gov for crash-test ratings, and iihs.org for the Institute’s own award status. General web search is the last resort, used to fill gaps and corroborate. A manufacturer’s published figure always outranks a secondhand one, and a figure we could not verify is marked at lower confidence rather than guessed.
Two kinds of criteria, scored differently
A spec criterion is a real measured quantity — towing in pounds, combined MPG, rear legroom in inches, base price in dollars. Its 0–100 score is computed from published normalization bands, never authored: nobody decides what 7,200 lb of towing is worth, the band does. A rated criterion is a judgment where no number exists — ride comfort, reliability outlook, how usable the infotainment actually is — and is scored 0–100 against written evidence you can read on the car’s page. Criteria that don’t apply are dropped rather than scored zero, and the category’s remaining weights are redistributed, so a sedan is never penalized for lacking a bed.
Graded against its own segment
Normalization bands are per segment. A truck’s fuel economy is graded against trucks; efficiency is graded within powertrain class, because MPG and MPGe are not the same measurement and a single fleet-wide winner would always be whichever EV posted the biggest MPGe. That is why a grade here is not comparable across those boundaries, and why the site says which pool a grade was drawn from wherever one is shown.
How a score becomes a GPA
Scoring runs on a 0–100 scale internally and maps onto the 4.0 scale on a fixed, published curve: 45 points is a 1.0, 88 is a 4.0, linear between and clamped outside. Those anchors are constants, not a curve fitted to whatever is in the catalog today — so a car’s GPA never moves because other cars were added or regraded. The report card publishes every weight, band and threshold behind that.
Written with AI assistance, and checked by tooling
The research and the written assessment for each scorecard are produced by a large language model working to the framework above, and each scorecard records which model produced it. We state that plainly because it is true and because the data says so anyway.
What that model is not permitted to do is decide the numbers. Every spec score is recomputed from the framework’s normalization bands after the fact, and every category score is recomputed as a weighted average of its items — an authored total that disagrees with its own parts is overwritten, not published. A validator then re-checks the whole catalog before every build: specs must carry a value and a unit and fall inside a sanity range, confidence must be declared, and criteria must not appear on vehicles they don’t apply to. A scorecard that fails is rolled back rather than shipped.
Which model produced the current model year’s scorecards, and how many of them carry a written verdict — 265 of 265 do; the rest show a verdict composed from the scores alone, and say so:
| Scorer | Scorecards |
|---|---|
| claude-opus-5 | 172 |
| claude-opus-4-7 | 86 |
| claude-sonnet-5 | 6 |
| gpt-5.6-sol | 1 |
How confident the figures are
Every criterion carries a confidence level: HIGH when it rests on a primary source (manufacturer, EPA, NHTSA, IIHS), MEDIUM on a reputable secondary one, LOW when the sourcing is thin, and INFERRED when it is an informed estimate with no citation — the evidence field is empty by rule in that case, so an inference can never masquerade as a sourced fact. Across the latest scorecard of every graded model:
| Confidence | Criteria | Share |
|---|---|---|
| HIGH | 9,508 | 34.0% |
| MEDIUM | 14,986 | 53.6% |
| LOW | 3,170 | 11.3% |
| INFERRED | 282 | 1.0% |
What is finished and what isn't
295 models are graded, covering 984 model years of the 2283 we track — 43%. The current model year is the priority and is effectively complete; the ten-year back catalog is filled in behind it. A model year we have not graded is labelled as such and has no page — we would rather show a gap than a guess. Nothing on this site is a projection or an estimate presented as a measurement.
Independence
No manufacturer pays for placement, no grade is for sale, and no ranking is influenced by an affiliate relationship. The site carries display advertising, which is served by a third-party network with no knowledge of or influence over the grades. The framework and the grades are open — anyone can check our arithmetic against the published bands.
Read the long version
This page is the standard in short. The guides carry the long versions — a worked example and live fleet figures for the scale, the two kinds of criteria, segment normalization, powertrain classes and the weights — plus a deep dive per subject and the patterns in the data that surprise people, including the calibration gap in Ownership Cost that this page would otherwise have to leave unexplained.
Corrections
Specs change mid-cycle, recalls land, and we get things wrong. If a figure here disagrees with the manufacturer or a federal source, that is a defect and we want it. Send the criterion, the correct figure and where it is published, through the feedback form. Verified corrections are applied to the scorecard, which recomputes the affected category and the vehicle’s GPA; each vehicle keeps a dated history of what changed.