RARE DISEASERESEARCH ATLAS

About

This site answers one question for each rare disease in Orphanet: does anyone appear to be working on it — in the published literature, in interventional trials, observational studies, or in gene–disease curation?

Author: Amit Bhattacharya · Instagram

Not medical advice

Everything here is derived landscape data. It is not a diagnosis, not a prognosis, and not a recommendation about care. Decisions about health should be made with qualified clinicians. Counts can be wrong.

Data sources & licences

  • Orphanet / Orphadata — rare disease nomenclature (en_product1.xml) and prevalence classes (en_product9_prev.xml). © Orphanet / INSERM. Licence: CC BY 4.0. Attribution required. Product dates in this build: 2026-06-23 07:53:50.
  • Europe PMC — publication search API (EMBL-EBI). Used for hit counts, author extraction, and yearly trends.
  • ClinicalTrials.gov — API v2 (U.S. NLM). Interventional-study totals and recruiting sample; observational and expanded-access records are not counted as clinical trials.
  • GenCC — gene–disease validity submissions export (CC0). Joined on MONDO / ORPHA identifiers.
  • Mondo Disease Ontology — is_a hierarchy for zero-publication naming-artifact detection and India NPRD umbrella (parent) matching, plus cross-references (MeSH, UMLS, OMIM, NCIT) and exact synonyms used for identifier-based matching.
  • NLM MeSH — descriptor / concept labels resolved from Mondo MeSH cross-references, unioned into both the Europe PMC and ClinicalTrials.gov queries.
  • India NPRD layer — hand-curated from the National Policy for Rare Diseases 2021 and later MoHFW/PIB lists. Last verified 2026-07-26. Editable at data/india-nprd.json. Official notified-disease counts are inconsistent across sources (recorded as both ~55 and 63 with citations).

Matching by name and by identifier

The root cause of most data errors here is matching on disease name strings — names are misspelled, differ between databases, or are hyper-specific. We reduce this by matching on structured identifiers as well as names. From Mondo we extract each disease's MeSH, UMLS, OMIM and NCIT cross-references, resolve MeSH descriptor labels, and union them into the queries. A trial that registers its condition as a broader MeSH descriptor (e.g. “Fatty Acid Oxidation Disorders” for an LCHAD-deficiency study) is then found even though no name phrase would match. Each disease page records whether a trial matched via phrase, mesh, or both.

Each Europe PMC query is the Orphanet preferred label plus Orphanet and Mondo exact synonyms, each as a quoted phrase, OR'd together and unioned with a MeSH query where a cross-reference exists. A stoplist drops terms under 5 characters, bare acronyms under 4 characters, single common English words, and hyphenated/spaced numeric-prefix + single letter generics (e.g. Poly-X, Tetra X). Dropped synonyms are logged and shown on each disease page.

Labels are normalised additively at parse time and the source value is never overwritten. A pure-alphabetic token that is not itself a corpus word but splits cleanly into two frequent words (a missing word boundary) is corrected and queried alongside the original. Suspected misspellings are detected but only flagged for human review, never silently applied — edit-distance cannot distinguish a typo from a legitimately different medical term (“Ebstein anomaly”≠“epstein”), and guessing would corrupt the very source we are trying to represent. In the current Orphanet build this detector found no genuine missing-space corruptions; the earlier examples do not exist in the source XML.

Credibility is per-signal. The publicationsDenominator excludes low-confidence / naming-artifact rows and failed publication fetches. The trialsDenominator excludes only failed ClinicalTrials.gov fetches and incomplete scans — a publication name-collision flag does not remove a disease from trial percentages. The headline counts only interventional trials, while the methodology also reports matched observational and other registered studies. Combined “thin attention” uses the intersection of both sets. Every percentage on the site names which denominator it uses.

Post-hoc rules: zero publications with GenCC Definitive/Strong, or with a Mondo parent that has substantial literature, force low confidence and exclude from publication neglect metrics. Trials use quoted phrases via ClinicalTrials.gov query.cond, then post-filtered. Only records whose study type is INTERVENTIONAL enter trial totals; observational and expanded-access records do not. Pan-disease registries (e.g. NCT01793168) are stored separately and not counted in trial totals. Percentiles compare each disease to the appropriate denominator after ingest.

Query health

Confidence reasons about ambiguity; query health reasons about whether the search itself worked. A record is broken when every strategy returns zero across both Europe PMC and ClinicalTrials.gov — that pattern almost always means a query-construction problem, not a global absence of research, so broken records are excluded from every denominator and reported separately. In this build, 43 records were excluded as broken. A record is suspect when a label correction was detected, a source fetch failed, or only one of several strategies returned hits.

Errors run in both directions

We no longer claim the “no trial” share is a lower bound. Matching produces false positives (we match a trial that belongs to another condition, which undercounts no-trial and pushes the true share higher) and false negatives (a broken query or a MeSH mismatch misses a real trial, which overcounts no-trial and pushes the true share lower). Both are present, so the headline is a bounded point estimate, not a floor.

The legacy human-reviewed reference measured trial recall at 96% and precision at 86%.

How we count trials, and why the number is uncertain

Deciding whether a trial “counts” for a rare disease is a judgment call, not a lookup. Disease names differ between medical literature, trial registries, and reference databases, and a trial may register under a broad category name while enrolling only a specific subtype — or the reverse.

We take the conservative approach: a trial counts toward the headline only when it names the specific condition. Trials registered for a broader Mondo parent category (for example, “Gaucher disease” when the page is a Gaucher subtype) are shown separately on the disease page rather than counted.

This choice matters. In this build, counting only specific-condition matches puts the share of rare diseases with no interventional trial at 59.2% (151 of 255). Treating parent-category registrations as filling a zero puts it at 49.0% (125 of 255). Both are defensible. We report the conservative figure and show you the other so you can judge. Earlier matching choices in this project landed in the mid-50s to mid-70s — that spread is itself a finding about how poorly disease naming maps between literature and trial registries.

We also exclude 43 diseases (14% of this sample) whose official names return nothing in either database — not because no research exists, but because we can't search for them reliably. Their pages say so.

Interventional trials versus all registered studies

The headline is an editorial definition, not a correction to the data: 151 of 255 diseases have no matched interventional trial testing a treatment, while 128 of 255 have no matched registered study of any type. Observational and natural-history studies are not trials, but they are genuine progress: regulators encourage them in rare disease because endpoints often cannot be designed until disease progression is understood. They are therefore shown prominently on disease pages and may offer families an actionable way to participate. Both figures exclude pan-disease registries, which remain listed separately.

Is the trial zero real, or a search failure?

The strongest single check on the headline: of the diseases with no interventional trial, 47 have substantial published literature (at least the median recent publication count). If a name were broken, Europe PMC would find nothing either — so a name that demonstrably matches papers makes the trial zero far likelier to be real than a query artifact. A further 104 have little or no literature and are the likelier query artifacts.

This is supportive evidence, not proof. Europe PMC searches abstracts and full text, while ClinicalTrials.gov matches a structured condition field using standard clinical terminology — so trial matching underperforms publication matching for a systematic reason, independent of our bugs.

Other known weaknesses: polysemous names still over-count; rare spellings and non-English literature under-count; author deduplication is imperfect; a zero often means “named differently,” not neglected.

Re-running ingestion

npm run ingest         # --limit 50 (sorted; reproducible)
npm run ingest:sample  # --sample 300 (random draw; use for neglect-rate estimates)
npm run ingest:full    # entire Orphanet set (hours)
# --resume   skip codes already in data/diseases.checkpoint.json
# --no-cache ignore .cache/ reads (monthly refresh)

Responses are cached under .cache/. Progress and failures append to ingest.log. Mid-run checkpoints go to data/diseases.checkpoint.json; the live diseases.json is published only when the target set is complete and fully scanned. Rate limiting (~3 req/sec) is in scripts/lib/http.ts.

Report an error

Every disease page links to a prefilled GitHub issue with the ORPHAcode. Corrections to the India list can also be proposed as edits to data/india-nprd.json.

Get in touch

Questions about the atlas, collaborations, corrections outside a specific disease page, or press: reach the author on LinkedIn or follow the project on Instagram.

Licence

This software is Apache-2.0. Upstream data remain under their own licences (notably Orphanet CC BY 4.0).