Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog

Library

Find the model or dataset for the question you have.

Ask in plain language and we will work out where the answer lives — or filter the index yourself. Almost everything here runs on our servers, from this page, with no setup.

Models
297
Datasets
245
Organisations
128

Labs, institutes and companies

Ready to run
24

Including 6 that run free and instantly

Looking for a particular value, not a particular dataset?Search schemas, rows and column statistics inside the data itself.

Things that work

Filtering

102 entries in Biology

Showing 48 of 102

AlphaFoldDBLiteFoldAlphaFoldDB prediction index — open database of predicted protein 3D structures with confidence scores, providing structural coverage for known protein sequences at scale.Predicted StructuresBrowse
ATBAllTheBacteriaAllTheBacteria: a comprehensive collection of ~2 million bacterial genome assemblies from public sequence databases, standardized for large-scale genomic analysis.Microbial GenomicsBrowse
Bac-Corpus-dna-intergenic-sequences-high-diversityAllTheBacteriaHigh-diversity corpus of bacterial intergenic DNA sequences for training DNA language models on non-coding regulatory regions.Microbial GenomicsBrowse
Bac-Corpus-protein-sequences-high-diversityAllTheBacteriaHigh-diversity corpus of bacterial protein sequences derived from the ATB collection, filtered for maximum sequence diversity to support protein language model pretraining.Microbial GenomicsBrowse
BEANS-ZeroEarthSpeciesProjectZero-shot bioacoustics benchmark evaluating audio-language models on species detection, classification, and captioning across diverse animal taxa.Bioacoustics BenchmarkBrowse
BFDLiteFoldBFD source archive index — clustered UniProt + metagenomic protein sequences, the standard database for AlphaFold-style homology search and MSA construction.Protein SequencesBrowse
BioMysteryBench-fullAnthropicFull BioMysteryBench evaluation set — challenging biology problems used to probe expert-level scientific reasoning in frontier models.Biology BenchmarkBrowse
BioMysteryBench-previewAnthropicPreview slice of BioMysteryBench — challenging, expert-curated biology problems for evaluating AI scientific reasoning capability.Biology BenchmarkBrowse
bioreason-pro-sft-reasoning-datawanglabReasoning trace dataset used to supervised-fine-tune BioReason-Pro — multimodal biological problems with rationales over genomic variants and pathway data.Biological Reasoning CorpusBrowse
BixBenchfuturehouseBenchmark with 205 reproducible research questions paired with data capsules for AI evaluation.Research BenchmarkBrowse
camelyon16-featuresowkinPre-extracted features from the CAMELYON16 breast cancer lymph node metastasis detection challenge, enabling efficient benchmarking of MIL methods.Computational PathologyBrowse
carbon-pretraining-corpusHuggingFaceBioCarbon pretraining corpus — the curated DNA sequence dataset used to train the Carbon-500M / 3B / 8B genomic foundation models.Genomic Pretraining CorpusBrowse
CATHLiteFoldCATH domain classification — hierarchical classification of protein domain structures by class, architecture, topology, and homologous superfamily.Domain ClassificationBrowse
clinvar-vepHuggingFaceBioClinVar variants annotated with VEP (Variant Effect Predictor) — formatted for benchmarking genomic-LLM pathogenicity classification.Clinical VariantsBrowse
CryptoCENmaomlabCryptoCEN — Cryptococcus coexpression network dataset for fungal pathogen biology and drug-target prioritisation.Coexpression NetworkBrowse
CT_DeepLesion-MedSAM2wanglabCT volumes from the DeepLesion benchmark with mask annotations restructured for training and evaluating MedSAM2, the universal medical image segmentation foundation model.Medical ImagingBrowse
DisProtLiteFoldDisProt — manually curated database of intrinsically disordered proteins and regions with experimental evidence and functional annotations.Protein AnnotationBrowse
drug-target-activityeve-bioDrug-target interaction measurements for 1,397 FDA-approved small molecule drugs.Drug DiscoveryBrowse
DynaCLR-databiohubTraining and evaluation data for DynaCLR — dynamic contrastive learning of cell embeddings from live-cell microscopy time series, including infection-status labels.MicroscopyBrowse
EmeraldBaytahoebioEmerald Bay — single-cell perturbation dataset of 1.8M+ transcriptomic profiles spanning 52 cell lines × 91 drug treatments (plus combinations), generated on Tahoe’s MOSAIC high-throughput platform with paired transcriptional and drug-phenotype readouts over a five-day culture.Single-Cell PerturbationBrowse
ESMAtlasLiteFoldHub mirror of the ESM Metagenomic Atlas — large-scale predicted metagenomic protein structures generated with ESMFold, packaged as webdataset shards.Predicted StructuresBrowse
ESMC-SAE-FeaturesbiohubFeature table for the 16,384-feature sparse autoencoder trained on layer 60 of ESMC-6B — supports mechanistic-interpretability research on protein language models.InterpretabilityBrowse
EvolutionaryLiteFoldPrecomputed evolutionary sequence-alignment data (MMseqs2-style) packaged for the Hub — drop-in MSA features for protein-LM and structure-prediction pipelines. ## Models (98)MSA / Training CorpusBrowse
FireProtDBLiteFoldFireProtDB — curated thermostability mutation data for proteins, supporting protein-engineering and stability-prediction modelling.Protein StabilityBrowse
FLIP2LiteFoldFLIP2 — second-generation Fitness Landscape Inference for Proteins benchmark (Feb 2026 bioRxiv) spanning NVIDIA, Microsoft, Caltech, and Profluent for protein-engineering eval.Protein BenchmarkBrowse
GDPa1ginkgo-datapointsAntibody developability dataset with biophysical assay data for 242 antibodies across 9 assays.Antibody DevelopabilityBrowse
GDPx1ginkgo-datapointsDRUG-seq functional genomics dataset with chemical perturbation experiments in A549 cells.Functional GenomicsBrowse
GDPx2ginkgo-datapointsDRUG-seq transcriptomic profiling across 4 primary human cell types with 85 compounds.Functional GenomicsBrowse
GDPx3ginkgo-datapointsHigh-content Cell Painting imaging dataset for AI/ML model training in drug discovery.Cell ImagingBrowse
GDPx4ginkgo-datapointsDRUG-seq transcriptomic profiling in engineered HEK293 cells with inducible gene overexpression, enabling systematic study of gene-drug interactions.Functional GenomicsBrowse
genomic-niahHuggingFaceBioGenomic Needle-in-a-Haystack — long-context evaluation for genomic LLMs, measuring retrieval and reasoning over long DNA sequences.Long-Context BenchmarkBrowse
GOLiteFoldGene Ontology terms — structured controlled vocabulary for describing gene-product functions, biological processes, and cellular components across organisms.OntologyBrowse
GOALiteFoldGene Ontology Annotation (UniProt) — high-quality evidence-coded GO annotations for UniProtKB proteins, RNAs, and protein complexes.OntologyBrowse
her2-challenge-2026owkinHER2 scoring challenge dataset with H&E-stained whole-slide images for evaluating AI-based HER2 status prediction in breast cancer.Computational PathologyBrowse
hestMahmoodLabHEST-1k — 1,276 spatial-transcriptomic profiles each linked and aligned to a Whole Slide Image (pixel size <1.15 µm/px), the largest paired histology + spatial-transcriptomics resource on the Hub.Spatial TranscriptomicsBrowse
hest-benchMahmoodLabHEST-Bench — companion benchmark to HEST-1k for evaluating spatial-transcriptomics and pathology foundation models on aligned histology / expression tasks.Pathology BenchmarkBrowse
HumanProteinAtlasLiteFoldHuman Protein Atlas — gene and protein expression across tissues, cells, organs, pathology, and subcellular locations for human proteins.Protein AnnotationBrowse
IEDBLiteFoldIEDB assay export — experimentally characterised B-cell, T-cell, and MHC-binding epitopes across infectious, allergic, autoimmune, and transplant settings.ImmunologyBrowse
img_virus_plasmidwanglabCombined IMG/VR (uncultivated virus genomes) and IMG/PR (plasmids from genomes and metagenomes) catalog with rich functional, taxonomic, and ecological metadata.Microbial GenomicsBrowse
InterProLiteFoldInterPro entries — integrated protein family, domain, repeat, and functional-site signatures across multiple member databases for sequence classification.Protein AnnotationBrowse
keggwanglabKEGG pathway entries paired with variant annotations for training and evaluating multimodal biological reasoning models (used by the BioReason work).Biological ReasoningBrowse
lab-benchfuturehouseLanguage Agent Biology Benchmark - 8 categories of scientific research tasks including cloning, figures, and protocols.Research BenchmarkBrowse
MegaScale-Tsuboyama2023LiteFoldMegaScale (Tsuboyama et al., 2023) — large-scale experimental measurements of protein stability across hundreds of thousands of mutations for stability prediction modelling.Protein StabilityBrowse
MgnifyLiteFoldMGnify protein catalogues — microbiome-derived protein sequences from EBI’s MGnify platform with train/test splits for downstream metagenomic protein tasks.MetagenomicsBrowse
Molecule3DmaomlabCurated 3D molecular structures with computed properties — supports geometric deep learning for property prediction and conformer-aware modelling.Molecular PropertiesBrowse
multi_species_genomesInstaDeepAIWhole-genome sequences for 850 species spanning bacteria, fungi, plants, and animals — the pre-training corpus for the Nucleotide Transformer model family.GenomicsBrowse
NCBILiteFoldNCBI RefSeq protein shard index — curated non-redundant protein subset of NCBI’s Reference Sequence collection, packaged as parquet for efficient loading.Protein SequencesBrowse
nct-crc-heowkinColorectal cancer tissue classification dataset with H&E-stained patches across 9 tissue classes, widely used for benchmarking pathology models.Computational PathologyBrowse

54 more behind this view