A systematic review of "data marketplace" research kept 280 of 505 records it found — and the screening ledger behind that decision reveals more about how keyword search works than about data marketplaces themselves.
Before a single data-marketplace finding could make it into the published review, 225 of 505 records the search turned up had to be thrown out — and 160 of those 225 rejections (71.1%) were tossed for one blunt reason: "not about data marketplace/market." A chi-square test confirms this isn't noise around an even split across five exclusion categories — it's a massive, overwhelming skew (χ²(4)=378.6, p=1.16e-80, Cohen's w=1.30, an unusually large effect size for social-science data).
Two of the five exclusion reasons — off-topic and "marketplace peripheral" — explain 85.8% of every rejected record. The other three reasons combined (wrong publication type, non-English, missing abstract) explain barely one in seven rejections. This wasn't a review agonizing over borderline judgment calls; it was mostly a blunt topical filter working exactly as broad keyword searches are designed to: cast wide, catch mostly garbage, and let screening sort it out.
That 71% number is the real headline of this dataset. It's not a story about data marketplaces — it's a story about what happens when you search Scopus for a two-word phrase that also matches decades of unrelated "market" research.
This screening table is the raw coding sheet behind a real, published systematic literature review — "Business Data Sharing through Data Marketplaces" (Abbas, Agahari, van de Ven, Zuiderwijk & de Reuver, JTAER, 2021) — which queried Scopus on 6 July 2020 for records matching "data market*" or "data marketplace*," then screened them down to 133 papers that actually belonged in a review of data marketplaces as a business phenomenon.
A data marketplace is, in effect, a stock exchange for data: a platform that connects organizations that have data with organizations that want to buy it. Interest in these platforms accelerated through the 2010s as companies started treating data as a monetizable strategic asset — but most data marketplaces built so far have struggled to become commercially viable businesses, which is exactly the gap the original review set out to map.
The review organizes its 133 included papers with the Service-Technology-Organization-Finance (STOF) framework and finds the literature lopsided: heavy on Technology (pricing algorithms, system architecture), thin on Organization and Finance — the harder questions of who actually buys data and how a marketplace sustains itself as a business.
None of that context comes from the numbers in this screening table directly — it's the world the numbers live in. Keep it in mind for what follows.
Across all 505 screened records spanning 33 distinct years, the share of records that get included climbs steadily with publication year — a Cochran-Armitage trend test finds Z=7.18 (p=6.90e-13), Holm-corrected p=1.38e-12, backed up by a Spearman correlation of ρ=0.24 between year and inclusion (p=3.46e-08).
Included records also skew measurably more recent than excluded ones: a median publication year of 2018 (IQR 2015-2019) for includes versus 2016 (IQR 2005-2018) for excludes (Mann-Whitney U=40305.5, p=5.28e-08). The effect size is moderate, not enormous (rank-biserial r=0.28) — an included record is newer than an excluded one about 64% of the time, not near-certainly.
Here's the structural explanation the detective work supplies: "data marketplace" only became a distinct, searchable academic term in roughly the mid-2010s. A broad lexical search for "data market*" necessarily pulls in a long tail of older records — papers about commodity markets, agricultural markets, financial market data — that happen to share the search string but were never candidates for inclusion. Older records are disproportionately off-topic not because older research was worse, but because the search term itself is young.
The p-values above describe a smooth trend, but the underlying shape is closer to a step-change. Records published before 2011 (n=89, spanning three decades) were included only 15.7% of the time. Records from 2011 onward (n=416) were included 63.9% of the time — a 4.06x jump. This is not a gradual, decade-by-decade drift; it's a hinge.
Nearly half the entire corpus — 224 of 505 records, 44.4% — was published in just the final three years before the search was run (2018-2020), with 2019 alone showing the highest single-year inclusion rate in the dataset at 71.0%. The corpus is heavily front-loaded toward the moment the search was conducted, which is exactly what you'd expect from a field whose literature is still actively accumulating.
Two independent checks were run to make sure the year trend is real and not an artifact of how the corpus was assembled. Re-running the tests on Scopus-only records — with 9 hand-picked "core" articles removed — leaves the trend essentially unchanged (Cochran-Armitage Z=6.99, p=2.76e-12 on n=496, versus Z=7.18 on the full n=505). Re-binning publication years into 5-year windows instead of single years also preserves it (Z=7.28, p=3.32e-13), with inclusion rate jumping from near-zero before 1994 to 60.3% in 2009-2014 and 63.8% in 2014-2020.
Those 9 hand-seeded articles matter because they're the opposite of a random sample: they were added directly by the review authors as known-relevant "core" papers rather than surfaced by the Scopus query, and every one of them is an Include. They're only 3.2% of all included records, but a skeptical reader is right to ask whether they're doing all the work — the Scopus-only robustness check exists precisely to answer that question, and the answer is no.
This is standard practice in careful systematic reviews: seed a small number of known-important papers to guard against a search strategy missing them, then explicitly test whether removing those seeds changes the conclusion. Here, it doesn't.
The overall shape here — a keyword search dominated by off-topic exclusions, with inclusion rate rising toward the present — is a well-known signature in PRISMA-style systematic reviews of emerging technical topics, not a peculiarity of this particular review. Broad database queries are deliberately over-inclusive, trading precision for recall, and the screening stage exists specifically to filter that noise back out.
So the honest reading of this dataset isn't "data marketplace research is exploding" — though it probably is growing, that's not what these particular numbers prove. It's "here is what the paper trail of one careful literature search looks like, and here is why any similarly-built search on any similarly young topic will produce almost exactly this shape: recency-skewed includes, an off-topic-dominated exclusion pile, and a robustness check to prove the trend isn't just a few seeded papers doing the work."
No photographs or embeddable media accompany this piece by design — the underlying objects are academic bibliographic records, which don't have a natural visual subject. The argument here is built entirely from the shape of the numbers themselves.
These are descriptive associations within one review's screening record, not properties of the data-marketplace literature at large or a causal claim about publication year. The search was run once, on 6 July 2020, so 2020's slightly lower rate than 2019 (56.6% vs. 71.0%) reflects a partial year, not a decline.