Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog

Library

Find the model or dataset for the question you have.

Ask in plain language and we will work out where the answer lives — or filter the index yourself. Almost everything here runs on our servers, from this page, with no setup.

Models
297
Datasets
245
Organisations
128

Labs, institutes and companies

Ready to run
24

Including 6 that run free and instantly

Looking for a particular value, not a particular dataset?Search schemas, rows and column statistics inside the data itself.

Things that work

Filtering

80 entries in Genomics

Showing 48 of 80

Ready to run

3
ReadyModel

DNABERT-2

Zhou et al., Northwestern

BPE-tokenised genomic BERT for motif, splice site and regulatory classification.

motif discovery
ReadyModel

Evo 2

Arc Institute / Stanford

Genome-scale foundation model spanning nucleotide to whole-genome context.

sequence generation
ReadyModel

Nucleotide Transformer v2

InstaDeep / Wellcome Sanger

DNA foundation model for regulatory elements, promoters and variant scoring.

regulatory prediction

Everything else indexed

45
EnformerDeepMindSequence-to-expression across a 200kb window covering distal enhancers.expression predictionOn demand
AlphaFoldDBLiteFoldAlphaFoldDB prediction index — open database of predicted protein 3D structures with confidence scores, providing structural coverage for known protein sequences at scale.Predicted StructuresBrowse
AlphaGenomegoogleGoogle DeepMind model predicting DNA regulatory features — gene expression, chromatin accessibility, and TF binding — at single-nucleotide resolution.GenomicsBrowse
ATBAllTheBacteriaAllTheBacteria: a comprehensive collection of ~2 million bacterial genome assemblies from public sequence databases, standardized for large-scale genomic analysis.Microbial GenomicsBrowse
Bac-Corpus-dna-intergenic-sequences-high-diversityAllTheBacteriaHigh-diversity corpus of bacterial intergenic DNA sequences for training DNA language models on non-coding regulatory regions.Microbial GenomicsBrowse
Bac-Corpus-protein-sequences-high-diversityAllTheBacteriaHigh-diversity corpus of bacterial protein sequences derived from the ATB collection, filtered for maximum sequence diversity to support protein language model pretraining.Microbial GenomicsBrowse
BFDLiteFoldBFD source archive index — clustered UniProt + metagenomic protein sequences, the standard database for AlphaFold-style homology search and MSA construction.Protein SequencesBrowse
bioreason-pro-sft-reasoning-datawanglabReasoning trace dataset used to supervised-fine-tune BioReason-Pro — multimodal biological problems with rationales over genomic variants and pathway data.Biological Reasoning CorpusBrowse
carbon-pretraining-corpusHuggingFaceBioCarbon pretraining corpus — the curated DNA sequence dataset used to train the Carbon-500M / 3B / 8B genomic foundation models.Genomic Pretraining CorpusBrowse
CATHLiteFoldCATH domain classification — hierarchical classification of protein domain structures by class, architecture, topology, and homologous superfamily.Domain ClassificationBrowse
CellFMMulti-institution consortium100M-cell foundation model for pan-tissue annotation and cross-dataset integration.cell annotationBrowse
clinvar-vepHuggingFaceBioClinVar variants annotated with VEP (Variant Effect Predictor) — formatted for benchmarking genomic-LLM pathogenicity classification.Clinical VariantsBrowse
EmeraldBaytahoebioEmerald Bay — single-cell perturbation dataset of 1.8M+ transcriptomic profiles spanning 52 cell lines × 91 drug treatments (plus combinations), generated on Tahoe’s MOSAIC high-throughput platform with paired transcriptional and drug-phenotype readouts over a five-day culture.Single-Cell PerturbationBrowse
ESMAtlasLiteFoldHub mirror of the ESM Metagenomic Atlas — large-scale predicted metagenomic protein structures generated with ESMFold, packaged as webdataset shards.Predicted StructuresBrowse
ESMC 300Mbiohub300M-parameter ESMC protein language model — compact variant of the ESMC family for protein-embedding workflows, fits on a single consumer GPU.Protein Language ModelBrowse
ESMC 300M (Dec 2024)biohubDecember 2024 release of the 300M-parameter ESMC protein language model — compact, widely used checkpoint for protein-embedding pipelines.Protein Language ModelBrowse
ESMC 600Mbiohub600M-parameter ESMC protein language model — provides digital representations of proteins for therapeutic engineering, variant-effect prediction, and transfer learning across protein tasks.Protein Language ModelBrowse
ESMC 600M (Dec 2024)biohubDecember 2024 release of the 600M-parameter ESMC protein language model — widely adopted checkpoint underpinning protein engineering and variant-effect workflows.Protein Language ModelBrowse
ESMC 6BbiohubFlagship 6B-parameter ESMC protein language model — state-of-the-art protein embeddings for variant-effect prediction, protein engineering, and biology research, trained on billions of protein sequences.Protein Language ModelBrowse
ESMC-SAE-FeaturesbiohubFeature table for the 16,384-feature sparse autoencoder trained on layer 60 of ESMC-6B — supports mechanistic-interpretability research on protein language models.InterpretabilityBrowse
Evo-2 40Barcinstitute40B-parameter DNA language model trained on 9.3 trillion nucleotides across all domains of life — zero-shot function prediction, variant effect scoring, and sequence generation.Genomic Foundation ModelBrowse
Evo-2 7Barcinstitute7B-parameter instruction-tuned DNA language model for gene function prediction, CRISPR guide design, and cross-species sequence analysis.Genomic Foundation ModelBrowse
EvolutionaryLiteFoldPrecomputed evolutionary sequence-alignment data (MMseqs2-style) packaged for the Hub — drop-in MSA features for protein-LM and structure-prediction pipelines. ## Models (98)MSA / Training CorpusBrowse
GeneformerTheodoris et al., Gladstone InstitutesTransformer over 30M transcriptomes for network inference and target prioritisation.network inferenceBrowse
genomic-niahHuggingFaceBioGenomic Needle-in-a-Haystack — long-context evaluation for genomic LLMs, measuring retrieval and reasoning over long DNA sequences.Long-Context BenchmarkBrowse
GOLiteFoldGene Ontology terms — structured controlled vocabulary for describing gene-product functions, biological processes, and cellular components across organisms.OntologyBrowse
GO-GPTwanglabGenerative model that predicts Gene Ontology functional annotations directly from protein sequences — bringing LLM-style decoding to functional protein characterisation.Protein Function ModelBrowse
GOALiteFoldGene Ontology Annotation (UniProt) — high-quality evidence-coded GO annotations for UniProtKB proteins, RNAs, and protein complexes.OntologyBrowse
HumanProteinAtlasLiteFoldHuman Protein Atlas — gene and protein expression across tissues, cells, organs, pathology, and subcellular locations for human proteins.Protein AnnotationBrowse
img_virus_plasmidwanglabCombined IMG/VR (uncultivated virus genomes) and IMG/PR (plasmids from genomes and metagenomes) catalog with rich functional, taxonomic, and ecological metadata.Microbial GenomicsBrowse
InterProLiteFoldInterPro entries — integrated protein family, domain, repeat, and functional-site signatures across multiple member databases for sequence classification.Protein AnnotationBrowse
keggwanglabKEGG pathway entries paired with variant annotations for training and evaluating multimodal biological reasoning models (used by the BioReason work).Biological ReasoningBrowse
LULA-1.1omtxLULA-1.1 is a lightweight, sequence-only protein-ligand binding scorer from Om Therapeutics. It takes a protein amino-acid sequence and ligand SMILES and returns a binding score. There is no structure input, docking, or folding step. ## Blog Posts (33)drug discoveryBrowse
MgnifyLiteFoldMGnify protein catalogues — microbiome-derived protein sequences from EBI’s MGnify platform with train/test splits for downstream metagenomic protein tasks.MetagenomicsBrowse
multi_species_genomesInstaDeepAIWhole-genome sequences for 850 species spanning bacteria, fungi, plants, and animals — the pre-training corpus for the Nucleotide Transformer model family.GenomicsBrowse
NCBILiteFoldNCBI RefSeq protein shard index — curated non-redundant protein subset of NCBI’s Reference Sequence collection, packaged as parquet for efficient loading.Protein SequencesBrowse
NessorecursionpharmaOpen-source, coarse-grained co-folding model that significantly accelerates binding-affinity predictions.Drug discoveryBrowse
NTv3_benchmark_datasetInstaDeepAIBenchmark dataset with functional tracks and genome annotations across 7 species.GenomicsBrowse
nucleotide_transformer_downstream_tasksInstaDeepAI18 genomic prediction benchmark tasks covering histone marks, regulatory regions, splice sites, and promoter activity across human and multi-species genomes.GenomicsBrowse
OGtattabioOpen Genomes (OG) — curated genome-sequence corpus from Tatta Bio for genomic ML pretraining and benchmarking.Genomic CorpusBrowse
OMGtattabioOpen Mixed Genomes (OMG) — large mixed-organism nucleotide corpus underpinning Tatta Bio’s gLM2 genomic foundation models.Genomic CorpusBrowse
OpenDDEaurekaresearchOpenDDE is an open-source, all-atom biomolecular foundation model that turns co-folding into a scalable engine for structure prediction, design, and optimization in drug discovery.Biomolecular foundation modelBrowse
opengenome2arcinstituteCurated collection of prokaryotic and eukaryotic genomic sequences for training and benchmarking large-scale biological foundation models.GenomicsBrowse
OpenProteinSetLiteFoldOpenProteinSet — open-source MSA training corpus released by the OpenFold team (Ahdritz et al., NeurIPS 2023) reproducing AlphaFold-style training data.MSA / Training CorpusBrowse
PATHOS-PLM-EMBEDDINGSDSIMBPATHOS protein language-model embeddings — precomputed feature representations supporting pathogenicity prediction and downstream macromolecular ML on protein sequences.Protein EmbeddingsBrowse

32 more behind this view