Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog

Library

Find the model or dataset for the question you have.

Ask in plain language and we will work out where the answer lives — or filter the index yourself. Almost everything here runs on our servers, from this page, with no setup.

Models
297
Datasets
245
Organisations
128

Labs, institutes and companies

Ready to run
24

Including 6 that run free and instantly

Looking for a particular value, not a particular dataset?Search schemas, rows and column statistics inside the data itself.

Things that work

Filtering

51 entries in Genomics

Showing 48 of 51

AlphaFoldDBLiteFoldAlphaFoldDB prediction index — open database of predicted protein 3D structures with confidence scores, providing structural coverage for known protein sequences at scale.Predicted StructuresBrowse
ATBAllTheBacteriaAllTheBacteria: a comprehensive collection of ~2 million bacterial genome assemblies from public sequence databases, standardized for large-scale genomic analysis.Microbial GenomicsBrowse
Bac-Corpus-dna-intergenic-sequences-high-diversityAllTheBacteriaHigh-diversity corpus of bacterial intergenic DNA sequences for training DNA language models on non-coding regulatory regions.Microbial GenomicsBrowse
Bac-Corpus-protein-sequences-high-diversityAllTheBacteriaHigh-diversity corpus of bacterial protein sequences derived from the ATB collection, filtered for maximum sequence diversity to support protein language model pretraining.Microbial GenomicsBrowse
BFDLiteFoldBFD source archive index — clustered UniProt + metagenomic protein sequences, the standard database for AlphaFold-style homology search and MSA construction.Protein SequencesBrowse
bioreason-pro-sft-reasoning-datawanglabReasoning trace dataset used to supervised-fine-tune BioReason-Pro — multimodal biological problems with rationales over genomic variants and pathway data.Biological Reasoning CorpusBrowse
carbon-pretraining-corpusHuggingFaceBioCarbon pretraining corpus — the curated DNA sequence dataset used to train the Carbon-500M / 3B / 8B genomic foundation models.Genomic Pretraining CorpusBrowse
CATHLiteFoldCATH domain classification — hierarchical classification of protein domain structures by class, architecture, topology, and homologous superfamily.Domain ClassificationBrowse
clinvar-vepHuggingFaceBioClinVar variants annotated with VEP (Variant Effect Predictor) — formatted for benchmarking genomic-LLM pathogenicity classification.Clinical VariantsBrowse
EmeraldBaytahoebioEmerald Bay — single-cell perturbation dataset of 1.8M+ transcriptomic profiles spanning 52 cell lines × 91 drug treatments (plus combinations), generated on Tahoe’s MOSAIC high-throughput platform with paired transcriptional and drug-phenotype readouts over a five-day culture.Single-Cell PerturbationBrowse
ESMAtlasLiteFoldHub mirror of the ESM Metagenomic Atlas — large-scale predicted metagenomic protein structures generated with ESMFold, packaged as webdataset shards.Predicted StructuresBrowse
ESMC-SAE-FeaturesbiohubFeature table for the 16,384-feature sparse autoencoder trained on layer 60 of ESMC-6B — supports mechanistic-interpretability research on protein language models.InterpretabilityBrowse
EvolutionaryLiteFoldPrecomputed evolutionary sequence-alignment data (MMseqs2-style) packaged for the Hub — drop-in MSA features for protein-LM and structure-prediction pipelines. ## Models (98)MSA / Training CorpusBrowse
genomic-niahHuggingFaceBioGenomic Needle-in-a-Haystack — long-context evaluation for genomic LLMs, measuring retrieval and reasoning over long DNA sequences.Long-Context BenchmarkBrowse
GOLiteFoldGene Ontology terms — structured controlled vocabulary for describing gene-product functions, biological processes, and cellular components across organisms.OntologyBrowse
GOALiteFoldGene Ontology Annotation (UniProt) — high-quality evidence-coded GO annotations for UniProtKB proteins, RNAs, and protein complexes.OntologyBrowse
HumanProteinAtlasLiteFoldHuman Protein Atlas — gene and protein expression across tissues, cells, organs, pathology, and subcellular locations for human proteins.Protein AnnotationBrowse
img_virus_plasmidwanglabCombined IMG/VR (uncultivated virus genomes) and IMG/PR (plasmids from genomes and metagenomes) catalog with rich functional, taxonomic, and ecological metadata.Microbial GenomicsBrowse
InterProLiteFoldInterPro entries — integrated protein family, domain, repeat, and functional-site signatures across multiple member databases for sequence classification.Protein AnnotationBrowse
keggwanglabKEGG pathway entries paired with variant annotations for training and evaluating multimodal biological reasoning models (used by the BioReason work).Biological ReasoningBrowse
MgnifyLiteFoldMGnify protein catalogues — microbiome-derived protein sequences from EBI’s MGnify platform with train/test splits for downstream metagenomic protein tasks.MetagenomicsBrowse
multi_species_genomesInstaDeepAIWhole-genome sequences for 850 species spanning bacteria, fungi, plants, and animals — the pre-training corpus for the Nucleotide Transformer model family.GenomicsBrowse
NCBILiteFoldNCBI RefSeq protein shard index — curated non-redundant protein subset of NCBI’s Reference Sequence collection, packaged as parquet for efficient loading.Protein SequencesBrowse
NTv3_benchmark_datasetInstaDeepAIBenchmark dataset with functional tracks and genome annotations across 7 species.GenomicsBrowse
nucleotide_transformer_downstream_tasksInstaDeepAI18 genomic prediction benchmark tasks covering histone marks, regulatory regions, splice sites, and promoter activity across human and multi-species genomes.GenomicsBrowse
OGtattabioOpen Genomes (OG) — curated genome-sequence corpus from Tatta Bio for genomic ML pretraining and benchmarking.Genomic CorpusBrowse
OMGtattabioOpen Mixed Genomes (OMG) — large mixed-organism nucleotide corpus underpinning Tatta Bio’s gLM2 genomic foundation models.Genomic CorpusBrowse
opengenome2arcinstituteCurated collection of prokaryotic and eukaryotic genomic sequences for training and benchmarking large-scale biological foundation models.GenomicsBrowse
OpenProteinSetLiteFoldOpenProteinSet — open-source MSA training corpus released by the OpenFold team (Ahdritz et al., NeurIPS 2023) reproducing AlphaFold-style training data.MSA / Training CorpusBrowse
PATHOS-PLM-EMBEDDINGSDSIMBPATHOS protein language-model embeddings — precomputed feature representations supporting pathogenicity prediction and downstream macromolecular ML on protein sequences.Protein EmbeddingsBrowse
PDBLiteFoldHub-packaged index over the Protein Data Bank mmCIF entries — the global archive of experimentally-determined 3D structures of biological macromolecules.Protein StructureBrowse
Perturb-SapiensarcinstituteLarge-scale human single-cell perturbation dataset used in the STACK foundation-model lineage — paired baseline and perturbed expression profiles for genetic perturbation screens.Single-Cell PerturbationBrowse
perturbation-benchHuggingFaceBioGenetic perturbation benchmark — evaluates how well genomic foundation models predict perturbation responses across cell states.Genomic BenchmarkBrowse
PfamLiteFoldPfam — protein family and domain annotation built from curated multiple sequence alignments and profile HMMs, mirrored to the Hub as JSONL.Protein FamiliesBrowse
plant-genomic-benchmarkInstaDeepAIPlant genomics benchmark spanning gene expression, chromatin accessibility, and agronomic trait prediction tasks across multiple crop and model plant species.Plant GenomicsBrowse
protenix-dataLiteFoldTraining data for Protenix — ByteDance’s open-source PyTorch reproduction of AlphaFold3 handling proteins, DNA, RNA, ligands, ions, and modifications.Structure PredictionBrowse
replogle-nadig-de-rhaistertahoebioDifferential-expression summary statistics for the Replogle-Nadig genome-scale CRISPR screen (4 cell lines × ~2,000 gene knockdowns: HepG2, Jurkat, K562, RPE1) — per-gene fold changes, significance, and pseudobulk deltas used to train Rhaister.CRISPR Differential ExpressionBrowse
Replogle-Nadig-PreprintarcinstituteReplogle-Nadig single-cell perturbation dataset (preprint release) — Perturb-seq screens used in the STATE single-cell embedding work for perturbation-response modelling.Single-Cell PerturbationBrowse
SE-167M-Humanarcinstitute167M human single-cell RNA expression profiles across diverse tissues and cell types, used for training STACK and SE single-cell foundation models.Single-Cell BiologyBrowse
SPIREAllTheBacteriaSearchable Planetary-scale mIcrobiome REsource: a large-scale metagenomics resource aggregating environmental microbiome samples from diverse global habitats.Microbial GenomicsBrowse
Stack-CellxGene45Marcinstitute45M curated single-cell profiles drawn from the CellxGene corpus, standardised for in-context learning and cross-study perturbation analysis.Single-Cell BiologyBrowse
State-Tahoe-FilteredarcinstituteFiltered Tahoe-100M slice used in the STATE workflow — high-quality single-cell perturbation profiles for training and benchmarking cross-study cell-state models.Single-Cell PerturbationBrowse
Tahoe-100MtahoebioGiga-scale perturbation atlas with 100M+ single-cell profiles from 50 cancer cell lines and 1,100 drugs.Single-Cell BiologyBrowse
Tahoe-x1-embeddingstahoebioPre-computed cell and gene embeddings from the Tahoe-x1 foundation model.Single-Cell BiologyBrowse
traitgymHuggingFaceBioTraitGym — multi-task benchmark for evaluating genomic foundation models on trait prediction across diverse phenotypes.Trait Prediction BenchmarkBrowse
true-cds-protein-tasksInstaDeepAICoding sequence and protein function prediction benchmark tasks.Protein TasksBrowse
UniProtKBLiteFoldProcessed UniProtKB release — comprehensive, high-quality protein sequence and functional annotation knowledgebase (Swiss-Prot + TrEMBL) reformatted for ML pipelines.Protein SequencesBrowse
UniRef50LiteFoldUniRef50 — UniProt reference cluster representatives at 50% identity, the most-aggressively clustered tier widely used for protein-LM pretraining.Protein SequencesBrowse

3 more behind this view