Corollary

Research

  • Ask

Library

  • Catalog

Account

  • Overview
  • Jobs
  • Usage
  • Billing
Settings
Corollary
  1. Catalog

Library

Find the model or dataset for the question you have.

Ask in plain language and we will work out where the answer lives — or filter the index yourself. Almost everything here runs on our servers, from this page, with no setup.

Models
297
Datasets
245
Organisations
128

Labs, institutes and companies

Ready to run
24

Including 6 that run free and instantly

Looking for a particular value, not a particular dataset?Search schemas, rows and column statistics inside the data itself.

Things that work

Filtering

50 entries in Benchmark

Showing 48 of 50

active_matterpolymathic-aiHigh-fidelity simulations of self-propelled particle systems for benchmarking learned PDE solvers and emergent collective behaviour models.Physics SimulationBrowse
aimo-validation-aimeAI-MOAIME I/II problems reformatted for AIMO challenge validation — 15-question integer-answer format, covering competition math at difficulty levels 5–9.Competition MathBrowse
aimo-validation-amcAI-MOAMC 10/12 competition problems reformatted for AIMO challenge validation, covering algebra, geometry, and number theory at difficulty levels 1–5.Competition MathBrowse
aimo-validation-math-level-4AI-MOLevel-4 MATH benchmark problems (pre-calculus difficulty) used for AIMO challenge validation and fine-grained model evaluation.Math ProblemsBrowse
aimo-validation-math-level-5AI-MOLevel-5 MATH benchmark problems (highest difficulty) used for AIMO challenge validation and measuring the ceiling of model mathematical reasoning.Math ProblemsBrowse
BEANS-ZeroEarthSpeciesProjectZero-shot bioacoustics benchmark evaluating audio-language models on species detection, classification, and captioning across diverse animal taxa.Bioacoustics BenchmarkBrowse
BioMysteryBench-fullAnthropicFull BioMysteryBench evaluation set — challenging biology problems used to probe expert-level scientific reasoning in frontier models.Biology BenchmarkBrowse
BioMysteryBench-previewAnthropicPreview slice of BioMysteryBench — challenging, expert-curated biology problems for evaluating AI scientific reasoning capability.Biology BenchmarkBrowse
BixBenchfuturehouseBenchmark with 205 reproducible research questions paired with data capsules for AI evaluation.Research BenchmarkBrowse
camelyon16-featuresowkinPre-extracted features from the CAMELYON16 breast cancer lymph node metastasis detection challenge, enabling efficient benchmarking of MIL methods.Computational PathologyBrowse
ChemBenchjablonkagroupManually curated benchmark of 3,000+ chemistry and materials science questions across spectroscopy, reactivity, synthesis, and property prediction for evaluating LLMs.Chemistry BenchmarkBrowse
clinvar-vepHuggingFaceBioClinVar variants annotated with VEP (Variant Effect Predictor) — formatted for benchmarking genomic-LLM pathogenicity classification.Clinical VariantsBrowse
CloudSEN12Plusisp-uv-esLarge-scale cloud detection dataset with 49,000+ Sentinel-2 patches and expert-quality cloud/shadow annotations across global biomes and seasons.Earth ObservationBrowse
CombiBenchAI-MOCombinatorics problems drawn from AMC, AIME, and olympiad competitions, formalised for benchmarking discrete-mathematics reasoning in language models.CombinatoricsBrowse
equational-theories-benchmarkSAIRfoundationFull benchmark suite of equational theory problems spanning algebraic structures, designed to evaluate formal reasoning capabilities of AI models.Mathematical ReasoningBrowse
equational-theories-selected-problemsSAIRfoundationCurated selection of equational theory problems for benchmarking LLM mathematical reasoning and automated theorem proving.Mathematical ReasoningBrowse
ether0-benchmarkfuturehouseChemistry reasoning benchmark covering SMILES-based tasks including reaction prediction, retrosynthesis, and molecular property estimation for evaluating chemistry LLMs.Chemistry BenchmarkBrowse
FLIP2LiteFoldFLIP2 — second-generation Fitness Landscape Inference for Proteins benchmark (Feb 2026 bioRxiv) spanning NVIDIA, Microsoft, Caltech, and Profluent for protein-engineering eval.Protein BenchmarkBrowse
frontierscienceopenaiFrontier science evaluation benchmark probing model capabilities on expert-level reasoning across natural sciences — designed to surface what AI systems can and cannot do at the research frontier.Scientific Reasoning BenchmarkBrowse
genomic-niahHuggingFaceBioGenomic Needle-in-a-Haystack — long-context evaluation for genomic LLMs, measuring retrieval and reasoning over long DNA sequences.Long-Context BenchmarkBrowse
GeometryLeanBenchAI-MOGeometry theorem proving problems formalised in Lean 4, covering Euclidean, affine, and metric geometry for automated reasoning evaluation.Theorem ProvingBrowse
healthbenchopenaiRealistic multi-turn health conversations graded against physician-written rubrics across multiple axes (accuracy, completeness, communication) — an open evaluation benchmark for AI assistants in medicine.Medical BenchmarkBrowse
healthbench-professionalopenaiProfessional-graded subset of HealthBench: physician evaluators score model responses to clinically realistic conversations, targeting expert-level health assessment.Medical BenchmarkBrowse
her2-challenge-2026owkinHER2 scoring challenge dataset with H&E-stained whole-slide images for evaluating AI-based HER2 status prediction in breast cancer.Computational PathologyBrowse
hest-benchMahmoodLabHEST-Bench — companion benchmark to HEST-1k for evaluating spatial-transcriptomics and pathology foundation models on aligned histology / expression tasks.Pathology BenchmarkBrowse
lab-benchfuturehouseLanguage Agent Biology Benchmark - 8 categories of scientific research tasks including cloning, figures, and protocols.Research BenchmarkBrowse
MaCBenchjablonkagroupMaterials Chemistry Benchmark — multimodal QA, multiple-choice, and visual-question-answering items for evaluating LLMs on materials and inorganic chemistry tasks.Materials Chemistry BenchmarkBrowse
MHD_64polymathic-ai3D magnetohydrodynamics turbulence simulations at 64³ resolution for training and benchmarking physics-informed neural operators.Physics SimulationBrowse
minif2f_testAI-MOTest set for miniF2F formal mathematics benchmark.Theorem ProvingBrowse
nct-crc-heowkinColorectal cancer tissue classification dataset with H&E-stained patches across 9 tissue classes, widely used for benchmarking pathology models.Computational PathologyBrowse
NTv3_benchmark_datasetInstaDeepAIBenchmark dataset with functional tracks and genome annotations across 7 species.GenomicsBrowse
nucleotide_transformer_downstream_tasksInstaDeepAI18 genomic prediction benchmark tasks covering histone marks, regulatory regions, splice sites, and promoter activity across human and multi-species genomes.GenomicsBrowse
openadmet-expansionrx-challenge-dataopenadmetFull ExpansionRx challenge dataset of RNA-targeted small-molecule compounds with measured ADMET properties for open pharmacokinetics benchmarking.Drug DiscoveryBrowse
opensr-testisp-uv-esBenchmark dataset for real-world Sentinel-2 super-resolution, with paired low/high-resolution imagery and perceptual quality metrics.Earth ObservationBrowse
Patho-BenchMahmoodLabPatho-Bench — benchmark designed to evaluate patch and slide encoder foundation models for whole-slide images across cancer subtyping, biomarker prediction, and survival tasks.Pathology BenchmarkBrowse
perturbation-benchHuggingFaceBioGenetic perturbation benchmark — evaluates how well genomic foundation models predict perturbation responses across cell states.Genomic BenchmarkBrowse
plant-genomic-benchmarkInstaDeepAIPlant genomics benchmark spanning gene expression, chromatin accessibility, and agronomic trait prediction tasks across multiple crop and model plant species.Plant GenomicsBrowse
popsiclebiohubPOPSICLE (Particle/Object Picking & Segmentation In CryoET Learning Evaluation) — multi-modal cryo-electron tomography benchmark for particle picking, localization, and segmentation across phantom, yeast, bacteria, and motor datasets.CryoET BenchmarkBrowse
principia-benchfacebookCurated benchmark of challenging STEM problems requiring multi-step reasoning, quantitative analysis, and domain knowledge across natural sciences.STEM BenchmarkBrowse
ProteinGymLiteFoldHub-packaged ProteinGym — the standard benchmark suite for evaluating protein fitness and variant-effect predictors, from the OATML / Marks lab (NeurIPS 2023).Protein BenchmarkBrowse
rayleigh_benardpolymathic-aiRayleigh–Bénard thermal convection simulations at varying Rayleigh and Prandtl numbers for benchmarking turbulence and heat transfer models.Physics SimulationBrowse
SKEMPI2LiteFoldSKEMPI v2 — manually curated benchmark of experimental changes in protein-protein binding affinity, kinetics, and thermodynamics upon mutation. ================================================================================ ## Topic: Biology (/topics/biology.md) ================================================================================ # Biology — Hugging Science > Life sciences, genomics, Binding AffinityBrowse
spiqagoogleScientific Paper Image Question Answering benchmark requiring multimodal reasoning over figures, charts, and diagrams from research papers across scientific domains.Scientific BenchmarkBrowse
surya-bench-flare-forecastingnasa-ibm-ai4scienceFull-disk solar flare forecasting dataset from NOAA GOES observations, providing multi-hour-ahead flare probability labels for heliophysics model evaluation.Solar PhysicsBrowse
TCGA-UniformTumor-8KMahmoodLabTCGA-UniformTumor-8K — region-level pan-cancer subtyping resource with 25,495 ROIs of 8,192×8,192 px tiled from TCGA whole-slide images, designed for high-resolution slide-encoder evaluation.Pathology BenchmarkBrowse
ThousandWorldsAstroAutomataThousandWorlds — benchmark for emulating exoplanet climates: 1,760 simulations across 5 GCMs, 8 planet parameters, and atmospheric variables on a 32 × 64 × 10 lat-lon-pressure grid, with three nested benchmark subsets, two evaluation protocols, and eight released baseline methods.Exoplanet Climate BenchmarkBrowse
traitgymHuggingFaceBioTraitGym — multi-task benchmark for evaluating genomic foundation models on trait prediction across diverse phenotypes.Trait Prediction BenchmarkBrowse
true-cds-protein-tasksInstaDeepAICoding sequence and protein function prediction benchmark tasks.Protein TasksBrowse

2 more behind this view