Tempus AI announced on September 11, 2026 that it is building a research platform containing 100,000 whole genomes linked to longitudinal clinical records. The company plans to organize the initial dataset around disease populations and patient outcomes, then pursue a longer-term goal of one million genomes. The important distinction is not simply scale. Tempus is pairing genomic sequences with the medical context needed to ask how genetic variation relates to disease progression and treatment response.
What makes the dataset multimodal
A genome alone is a large description of inherited and acquired variation. It becomes more useful for health research when it can be compared with what happened to the patient: diagnoses, therapies, imaging, pathology, laboratory results, and outcomes over time.
Tempus says researchers will be able to examine those de-identified data types together inside its existing environment and use Tempus Lens to build and validate models without moving the underlying dataset between systems. It is also designing the resource for model pre-training and post-training, rather than treating AI as an analytical layer added after collection.
A planned tenfold expansion
Whole-genome sequences paired with longitudinal clinical outcomes.
Source: Tempus, September 11, 2026. The 1 million figure is a long-term goal, not a completed dataset.
Disease-linked data can answer different questions
Large national biobanks often begin with a broad population and support many research designs. Tempus is proposing a more targeted resource built around people with particular diseases and the course of their care. That can make the dataset more useful for questions about treatment response, rare variants, progression, and patient subgroups that are difficult to study from genomic data alone.
This is adjacent to, but distinct from, DeepMind's AlphaGenome Atlas. AlphaGenome predicts how DNA changes may affect biological processes. Tempus is assembling observed genomes alongside clinical histories and outcomes. One is a model-generated map of variant effects; the other aims to become an empirical training and validation environment.

Size does not settle quality or access
Tempus's announcement establishes a target, not a completed dataset or a demonstrated medical result. Its value will depend on cohort diversity, missing-data patterns, the consistency of clinical labels, follow-up duration, and whether researchers can reproduce findings outside the Tempus environment.
De-identification is also necessary but not a complete description of governance. Genomic information is unusually persistent and identifying. The announcement does not provide detailed terms for patient consent, researcher eligibility, pricing, export controls, or how requests for data removal would work. Those questions will matter as much as the genome count for hospitals, researchers, and patients deciding whether the platform is trustworthy.
A hundred thousand genomes is a scale claim. A useful medical dataset also needs representative cohorts, reliable outcomes, and governance researchers and patients can understand.
Metir analysis
The defensible conclusion today is narrow. Tempus is attempting to build a model-ready bridge between whole-genome sequencing and longitudinal care data. If it reaches the stated scale with strong cohort quality and governance, it could support research that general-population resources are not optimized to answer. The announcement itself does not prove that those conditions will be met.
Sources:
- Tempus launches 100,000-genome AI health dataset initiative | Tempus
- Tempus announcement filed with investors | Tempus Investor Relations
- Tempus Investor Day, May 2026 | Tempus Investor Relations
Image credits
Illumina HiSeq 2000 sequencing room at BGI Hong Kong, Scotted400, via Wikimedia Commons, licensed under CC BY 3.0.