In PRIMAP-hist, the standard compiled national historical GHG emissions dataset, the choice between two "official" tie-breaking rules for conflicting emissions sources moves industrialized (Annex I) countries' totals roughly fourteen times more than everyone else's — because Annex I countries are the ones with a second, independent source to disagree with in the first place.
Here is a reasonable guess about how emissions bookkeeping should work: rich, heavily regulated countries file rigorous inventories, so if you compare their numbers against an independent cross-check, the two should land close together. Poorer countries, with thinner statistical infrastructure, should be where the estimates wobble.
The opposite happens. In the standard compiled dataset of national historical greenhouse-gas emissions, switching between two equally "official" ways of resolving the same underlying sources changes the 40 industrialized (Annex I) countries' 2019 totals by a median of 9.1% — versus just 0.6% for the other 166 countries in the world. That is not a rounding difference. It is roughly a fourteen-fold gap in how much a single bookkeeping choice moves the number.
The explanation is not that rich countries' data is worse. It's the reverse — and it changes how you should read any single country's "final" emissions figure.
The dataset in question is PRIMAP-hist, a widely used compiled historical emissions time series running from 1750 to 2019 for essentially every country in the world, built by combining official country submissions with independent third-party estimates into one harmonized table. It is a standard input for climate-policy research and country-comparison work precisely because it turns a patchwork of incompatible national reports into a single number per country per year.
What most users don't stop to notice is that PRIMAP-hist actually ships two parallel versions of every country's number, distinguished by which upstream source wins when more than one exists. One version, code-named HISTCR, prioritizes what the country itself reported to the UN. The other, HISTTP, prioritizes independent third-party estimates instead. The dataset's own documentation recommends HISTCR as the safe default for policy analysts — specifically so their numbers match what governments officially submitted. Both versions draw from the same pool of sources; they differ only in the tie-breaking rule.
Laid out country by country, the pattern is stark. The absolute difference between the two scenario values has a median of 9.1% for Annex I countries against 0.6% for everyone else, and a two-sided Mann-Whitney test confirms the two groups are not behaving like samples from the same distribution (U=4,590.0, p=1.04×10⁻⁴). By mean, the gap looks smaller (14.7% versus 9.6%) only because a handful of large non-Annex I outliers pull the average up — the median and the rank-based test are the numbers that matter here, because they aren't distorted by a few extreme cases.
The Annex I/non-Annex I line is not a data artifact; it's a formal distinction written into the 1992 UN climate treaty. The roughly 40 Annex I countries — industrialized nations at the time the treaty was signed — must file detailed, standardized greenhouse-gas inventories every single year in a shared electronic format, reviewed under a formal UN expert process. Non-Annex I countries report far less often, on a far less standardized template, and many have historically struggled to submit anything at all.
Two of the 40 Annex I Parties: the United States and Russia.
The mechanism is visible directly in the data. The two scenario values are byte-for-byte identical for only 1 of the 40 Annex I countries — 2.5% — but for 82 of the 166 non-Annex I countries, essentially a coin flip at 49.4%. A Fisher's exact test on that 2×2 table puts the odds ratio at 38.07 (p=3.53×10⁻⁹): an Annex I country's odds of having its two versions disagree at all are about 38 times higher than a non-Annex I country's.
That number reframes everything above it. HISTCR and HISTTP are not two independent measurements of the same quantity, cross-checking each other. They are two different tie-breaking rules pointed at a shared, often single-item pool of sources. When a country has no country-reported inventory to prioritize — the norm for many non-Annex I countries — both rules fall back to the identical third-party number by construction. The "agreement" you see for half of non-Annex I countries is not corroboration between two methods. It's the absence of a second method to disagree with.
Some individual countries make the abstract statistic concrete. Iceland's two scenario values differ by 127.5% — 4,740 thousand tonnes CO2-equivalent under the country-reported version against 21,400 under the third-party version, more than a fourfold gap for an Annex I country. Norway, another Annex I country, shows a 48.0% swing. On the other side of the ledger, Georgia's two values differ by -113.1% (62,300 versus 17,300), and Barbados by -111.2% — both non-Annex I, both proof that individual outliers don't respect the group pattern even when the group median does.
That last point is worth stating plainly rather than glossing over: ranking all 206 countries by the size of their divergence and looking at just the most extreme tenth (21 countries), only 4 of them — 19% — are Annex I, almost exactly Annex I's 19.4% share of the full list. The single largest individual gaps are not concentrated among industrialized countries. The group-level finding above is about where a typical country sits, not about who owns the single biggest number — both things are true at once, and neither cancels the other out.
Repeating the comparison for 1990, 2000, 2010, and 2019 shows the gap in every single year, and it has been widening: Cliff's delta rises from 0.214 in 1990 to 0.226 in 2000 to 0.279 in 2010 to 0.383 in 2019, with every 95% confidence interval excluding zero. The Annex I median divergence itself climbs from 6.3% in 1990 to 9.1% in 2019, while the non-Annex I median never breaks 1%. This is a structural, persistent, and apparently deepening feature of how the dataset is built — not a one-year anomaly.
Two obvious objections don't survive contact with the data either. Recomputing the 2019 comparison from the no-rounding source file gives Cliff's delta of 0.395 versus 0.383 for the standard file — a difference of 0.013, meaning the published files' rounding is not driving the result. And splitting the aggregate greenhouse-gas basket into individual gases shows the effect holds separately: Cliff's delta of 0.312 for CO2 and 0.416 for CH4, both with confidence intervals that clear zero. Whatever is happening, it isn't a quirk of how carbon-dioxide-equivalents get calculated.
One more natural objection: maybe this is just a country-size effect — bigger economies have messier, more-revised statistics, and Annex I happens to include large emitters. It doesn't hold up. Across all 206 countries, the correlation between a country's total emissions and how much its two scenarios diverge is weak (Spearman ρ=0.142, p=0.042), and inside each group separately it disappears entirely — ρ=-0.208 (p=0.197) within Annex I, ρ=0.110 (p=0.158) within non-Annex I. Knowing a country's emissions volume tells you almost nothing about how much its two numbers will disagree. Knowing its Annex I status tells you a great deal.
PRIMAP-hist is not a niche academic artifact — it's one of the standard inputs feeding climate-policy research, NDC-tracking tools, and cross-country comparisons, precisely because it turns fragmented national reports into one clean number. The dataset's own documentation treats the country-reported version as the "safe" default. This analysis shows that default is not a neutral formatting choice for the countries most often the direct subject of climate accountability work.
The pattern here also isn't unique to this dataset. A separate, independent study of U.S. city-level greenhouse-gas inventories found that cities under-report their own emissions by an average of 18.3% relative to independent bottom-up estimates — a different quantity, at a different geographic scale, measured a different way, so the two numbers shouldn't be read as directly comparable. But the direction is the same: wherever a self-reported figure can be checked against an independently-sourced one, at whatever scale researchers have looked, a real gap tends to show up.
None of this says which version — country-reported or third-party — is more accurate. This is an associational finding about how the dataset is constructed, not a verdict on whose numbers are right. What it does say is narrower and more actionable: for industrialized countries, the prioritization choice can move a national total by roughly a tenth, and any analysis that silently mixes the two variants across countries — or compares countries without checking which variant each one used — is comparing numbers built on different rules without realizing it.
The next time a chart shows two countries' emissions side by side, the more interesting question may not be which one is higher. It's whether both numbers were even built the same way.