Back to paper

A 2011 table. Two jobs. 23 requirements.

One version of this malaria test treats a sick patient. The other tracks a hidden outbreak. Of the 23 requirements below, how many do you think are set differently for the two jobs?

the two use-cases agree they set a different priority

Just 4 of 23 — and all four are about identifying which parasite is present, not just whether one is.

Keep reading ↓

Data story · Target product profiles for malaria diagnostics

Two Jobs, One Test

What a 2011 malaria diagnostics table got right about a problem the world is still solving.

Source: malERA Consultative Group on Diagnoses and Diagnostics (2011), Supplementary Table 1. Licence CC BY 4.0.

Of the 23 requirements in this table, 19 are identical in priority between the two intended uses of a malaria diagnostic test. But look at requirement 11: “Ability to detect hypnozoites.” For a clinician treating a sick patient in an elimination setting, it is Not Required — the case-management column leaves it blank. For a surveillance program screening a district for hidden transmission, it is Desirable, the largest single jump in stringency anywhere in the table.

It isn’t alone. Genotyping and the ability to detect gametocytes follow the identical pattern: Not Required for the person being treated, but tagged Optional for the population being watched. These are the only three criteria in the entire 23-row table where surveillance asks for more than case management does — and all three involve detecting a parasite’s genetic or biological subtype, not simply confirming that someone is infected.

That is not an arbitrary preference. Most malaria rapid tests detect infection via a protein called HRP2 — but some strains of the parasite carry gene deletions that make them invisible to HRP2-based tests, producing false negatives. A single clinician doesn’t need to know a patient’s HRP2 status in advance to treat them; a surveillance program absolutely needs to know whether HRP2-deletion strains are spreading, because if they are, an entire region’s diagnostic strategy has to change. In April 2026, the World Health Organization prequalified new rapid tests specifically designed to catch these deletion-carrying strains — fifteen years after this table first drew the distinction.

Zero = both use-cases assign equal stringency. 19 of 23 criteria aren’t shown here because both use-cases agree.

Moussa Diagne, an entomologist with Senegal's Parasite Control Service, administers a malaria rapid diagnostic test to a patient.
Moussa Diagne, an entomologist with Senegal’s Parasite Control Service, administers a malaria rapid diagnostic test. Every requirement in this table describes a physical object like this one. Photo: Nicole Schiegg / USAID, 2010 — public domain.

The document behind the table

This table is Supplementary Table 1 from “A Research Agenda for Malaria Eradication: Diagnoses and Diagnostics,” published in 2011 by the malERA Consultative Group on Diagnoses and Diagnostics — one of seven expert panels convened after global malaria policy shifted from “control” toward “eradication.” Their job was to specify, in advance of any product existing, what a future diagnostic test would need to do.

That kind of document is called a target product profile, or TPP — a planning tool used across drug and diagnostic development to fix the bar a future product must clear before anyone builds it. Each of the 23 criteria in this table carries a priority tier: Essential (a hard minimum), Desirable (a target to aim for), or Optional (nice to have, gates nothing). The table assigns those tiers twice over, once for each of two distinct jobs a malaria test can be asked to do: case management in elimination settings, where a health worker needs an answer for the patient in front of them, and screening/surveillance at the district level, where a program needs a population-level picture of where transmission is still occurring.

The stakes behind the exercise are not abstract. The World Health Organization’s 2025 report estimates 282 million malaria cases and 610,000 deaths globally in 2024 — about 9 million more cases than the year before — with the African Region carrying 94% of cases and 95% of deaths, three-quarters of them in children under five. Since 2000, an estimated 2.3 billion cases and 14 million deaths have been averted and 47 countries plus one territory are now certified malaria-free, which is precisely the eradication trajectory this table’s authors were trying to accelerate. WHO recommends confirming every suspected case with a diagnostic test before treatment — and frames that diagnosis as serving both purposes this table separates: getting one patient the right care, and building the surveillance picture that tells a program where to act next.

A rapid diagnostic test being administered in the field.

Set the requirement text aside and just count priority tiers, and the two use-cases look almost interchangeable. Case management splits its 30 tagged requirement items exactly 50.0% Essential and 50.0% Desirable, with zero tagged Optional. Surveillance splits its 32 tagged items 43.8% Essential, 50.0% Desirable, and 6.2% Optional — the Desirable share is identical between the two, and the only visible shift is a few points moving from Essential into a newly nonzero Optional category.

Two pre-registered statistical tests confirm what that comparison suggests. A Freeman-Halton exact test on the full 2×3 table of tagged items returns p = 0.59 (effect size Cramér’s V = 0.18) — no significant difference in overall tier composition. A second test asks a sharper question directly: of the 23 individual criteria, are the two use-cases’ specific requirements identical or different about as often as chance would predict? Nine are identical, fourteen differ — a proportion of 0.61 that an exact binomial test cannot distinguish from one-half (p = 0.40).

Correcting both p-values together for multiple comparisons with the Holm-Bonferroni method leaves both at an adjusted p = 0.81. Neither test was close to significant before correction, so the correction changes nothing about the conclusion: by these two aggregate measures, the case-management and surveillance columns of this table are statistically indistinguishable.

The paradox, resolved

Lay all 23 criteria out side by side and a clean structure appears. Thirteen are Essential in both use-cases — the non-negotiable technical and operational core every malaria test must clear regardless of purpose: sensitivity, specificity, safety, ease of use, and more. Three are Desirable in both (packaging, training, cost). Three are Not Required in either. That’s 19 of 23 criteria — 83% of the table — where the two use-cases assign the exact same priority tier.

But “assigns the same tier” and “is worded identically” are not the same claim, and conflating them is exactly how a real signal disappears into a null aggregate test. Only 9 of those 19 agreeing criteria use identical requirement text; the other 10 reach the same priority tier through differently worded specifications — different time thresholds, different population descriptions, that still land on the same Essential or Desirable tier. Strip away that wording noise and the real disagreement is small and specific: just 4 criteria, 17% of the table, where the two use-cases genuinely assign a different stringency.

It’s worth being honest about what the two null tests do and don’t establish here. With only 23 criteria and 62 tagged items, both pre-registered tests are underpowered by construction, and their non-significant results should not be read as evidence the use-cases are equivalent — only that a global count-the-tiers test isn’t sensitive enough to see a signal concentrated in 4 specific rows out of 23. A test built to detect a diffuse difference across the whole table will correctly report “no difference” even when a small, concentrated, meaningful one exists.

Analytic sensitivity (parasite/µl)Technical specificationsEEEssential (both)
Diagnostic sensitivityTechnical specificationsEEEssential (both)
Analytic specificityTechnical specificationsNot required (both)
Diagnostic specificityTechnical specificationsEEEssential (both)
Temperature stabilityTechnical specificationsEEEssential (both)
Integrity of packagingTechnical specificationsEEEssential (both)
Pf predominant areasTechnical specificationsEEEssential (both)
Pf and non-Pf areasTechnical specificationsEEEssential (both)
GenotypingTechnical specificationsOMore stringent for surveillance
Ability to detect gametocytesTechnical specificationsOMore stringent for surveillance
Ability to detect hypnozoitesTechnical specificationsDMore stringent for surveillance
Packaging of tests or reagentsHealth systems and technical specificationsDDDesirable (both)
Field stability / shelf life of consumablesHealth systems and technical specificationsEEEssential (both)
Training requirementsHealth systems and technical specificationsDDDesirable (both)
Reagent requirementsHealth systems and technical specificationsEEEssential (both)
InvasivenessHealth systems and technical specificationsEEEssential (both)
Rapidity of resultsHealth systems and technical specificationsEEEssential (both)
Ease of useHealth systems and technical specificationsEEEssential (both)
CostHealth systems and technical specificationsDDDesirable (both)
SafetyHealth systems and technical specificationsEEEssential (both)
Waste disposalHealth systems and technical specificationsNot required (both)
Inter-reader reliability (clarity of result)Health systems and technical specificationsNot required (both)
Instrumentation and laboratory infrastructure requirementsHealth systems and technical specificationsEDMore stringent for case management
Two voices, 23 note-pairs — one for case management, one for surveillance. Listen for where they split.

Criterion 1 (technical specs)Criterion 23 (lab infrastructure)

The 4 criteria that don’t agree

Three of those four diverging criteria — genotyping, gametocyte detection, hypnozoite detection — sit next to each other in the table’s Technical specifications section, and all three share the same underlying logic: they’re about identifying which parasite subtype or life stage is present, not merely confirming an infection exists. All three are Not Required for case management and elevated for surveillance. This isn’t scattered noise across the table; it’s one coherent domain concept showing up three separate times.

It shows up in the tier tags too: across all 62 tagged requirement items in the whole table, the Optional tier is used exactly twice — and both uses are in the surveillance column, on two of these same three criteria (genotyping and gametocyte detection). Case management never uses Optional at all. In this table, “Optional” functions as a surveillance-specific middle ground for research-adjacent detection capabilities that individual patient care simply has no use for.

2 of 62 Optional-tier tags in the entire table — both belong to the surveillance column, both on criteria in the same divergent trio

The fourth diverging criterion runs the other way: “Instrumentation and laboratory infrastructure requirements” is Essential for case management (no external power source assumed) but only Desirable for surveillance (infrastructure assumed to already be provided at a testing site). Between the two directions, the pattern is domain-coherent rather than arbitrary: surveillance asks more of a test’s ability to characterize what it finds; case management asks more of a test’s ability to work anywhere, unassisted.

The Essential bar, checked against reality

Nearly every sensitivity and specificity threshold in the table is either identical or purpose-agnostic between the two use-cases — diagnostic sensitivity is set at Essential>95%, Desirable≥99% with the exact same wording in both columns. Diagnostic specificity is the single exception: case management asks for Essential>90%/Desirable>95%, but surveillance tightens the Essential floor to >99% specifically “in low-transmission areas.” A false positive costs more when a program is trying to confirm malaria has been pushed down to near-zero than when it’s confirming what’s already widely suspected.

Since 2008, WHO — with FIND and the US CDC — has run an independent laboratory program testing every commercially available malaria rapid test against exactly this kind of performance bar. By the program’s eighth and final round (2016–2018), 79.4% of tested products (27 of 34) met every performance criterion, up from 26.8% (11 of 41) in the first round a decade earlier. The trend line points the right direction, but even at the program’s close, roughly 1 in 5 tested products still fell short of criteria in the same family as this table’s Essential threshold.

CriterionCase managementSurveillance
Analytic sensitivity (parasite/µl)E, 100–200, D<5E = 20, D≤5
Diagnostic sensitivityE>95%, D≥99%E>95%, D≥99%
Analytic specificityNegative all pathogens, common blood disordersNegative all pathogens, common blood disorders
Diagnostic specificityE>90%, D>95%E>99% surveillance low-transmission areas, E>95% screening

The one criterion in the table with a genuinely use-case-conditional numeric threshold, highlighted.

79.4% of WHO-tested RDT products met every performance criterion by the program’s final round (2016–2018) — up from 26.8% a decade earlier
Illustration of a grid of malaria rapid diagnostic test cassettes on a laboratory bench, with one cassette glowing warmly among the rest.
Illustration: a grid of test cassettes, one singled out.AI-generated illustration

This table was written in 2011, before anyone could name the scale of the problem its surveillance column was quietly bracing for. HRP2/HRP3 gene deletions were, at the time, a documented curiosity in scattered case reports. By 2026, they are enough of an operational problem that WHO has built dedicated surveillance protocols, an international laboratory network for gene-deletion mapping, and — as of this April — a new round of prequalified tests engineered specifically to work around them.

None of this is abstract. It sits underneath 282 million cases and 610,000 deaths a year, most of them in places where the difference between a test built to treat one patient and a test built to track a hidden epidemic is not a technicality — it is the entire strategy for getting from 610,000 deaths a year to zero.

The malERA panel could not have specified HRP2/HRP3 deletion tracking as a named criterion in 2011 — the problem hadn’t yet forced itself onto the agenda. What they specified instead was a category: know your parasite, not just its presence, when your job is watching the whole population rather than treating the person in front of you. Fifteen years on, that’s the exact criterion the field is still building tools to satisfy.