Open any dataset and read its schema, its real rows and its column statistics — then search inside it. 245 curated, and far more beyond them.
Name one and we will usually open it straight away — paste a link, an owner/name, or just describe it.
245 datasets
ADSKAILab
One million CAD-quality 3D shapes drawn from the ABC dataset — the foundation training corpus for the Make-A-Shape and WaLa generative models.
CAD Geometry Corpus
ADSKAILab
Megatron-formatted CodeParrot release used for large-scale code language-model pretraining experiments at Autodesk AI Lab. ## Models (25)
Code Pretraining
ADSKAILab
Narrative planning task set for evaluating LLM planning and reasoning over multi-step design and engineering scenarios.
Planning Benchmark
ADSKAILab
Curated 100K-example subset of Zero-To-CAD — useful for benchmarking and lightweight fine-tuning of CAD-from-image models.
CAD Vision-Language Corpus
ADSKAILab
1M paired image-and-CAD-program examples for training vision-language models that synthesise parametric CAD from images.
CAD Vision-Language Corpus
Ahmad0067
Realistic synthetic medical dialogue–SOAP note pairs generated to support training and evaluation of clinical documentation models without exposing real patient data.
Clinical NLP
AI-MO
AIME I/II problems reformatted for AIMO challenge validation — 15-question integer-answer format, covering competition math at difficulty levels 5–9.
Competition Math
AI-MO
AMC 10/12 competition problems reformatted for AIMO challenge validation, covering algebra, geometry, and number theory at difficulty levels 1–5.
Competition Math
AI-MO
Level-4 MATH benchmark problems (pre-calculus difficulty) used for AIMO challenge validation and fine-grained model evaluation.
Math Problems
AI-MO
Level-5 MATH benchmark problems (highest difficulty) used for AIMO challenge validation and measuring the ceiling of model mathematical reasoning.
Math Problems
AI-MO
Raw problem posts and discussion threads from the Art of Problem Solving forums, spanning AMC, AIME, and international olympiad competitions.
Competition Math
AI-MO
Combinatorics problems drawn from AMC, AIME, and olympiad competitions, formalised for benchmarking discrete-mathematics reasoning in language models.
Combinatorics
AI-MO
Geometry theorem proving problems formalised in Lean 4, covering Euclidean, affine, and metric geometry for automated reasoning evaluation.
Theorem Proving
AI-MO
Prompt-set for training and evaluating Kimina, a Lean 4 theorem prover that uses reinforcement learning over formal mathematical proofs.
Theorem Proving
AI-MO
Test set for miniF2F formal mathematics benchmark.
Theorem Proving
AI-MO
860K+ competition math problems from 17 sources with verified solutions — the training backbone of the gold-medal solution at the 2024 AI Mathematical Olympiad.
Math Problems
AI-MO
NuminaMath with Chain-of-Thought reasoning annotations.
Math Problems
AI-MO
Mathematical problems formalized in LEAN proof assistant.
Theorem Proving
AI-MO
NuminaMath with Tool-Integrated Reasoning annotations.
Math Problems
AI-MO
Olympiad-level mathematical problems collected from international and national competitions, formatted for training and evaluating mathematical reasoning models. ## Models (3)
Math Reasoning Corpus
AI-MO
Extended reference set of olympiad problems with verified step-by-step solutions, used for Chain-of-Thought and formal reasoning training.
Competition Math
AI-MO
Canonical reference set of international and national mathematical olympiad problems, used as the base for downstream NuminaMath training splits.
Competition Math
Aignostics
Pre-analyzed H&E whole-slide images from TCGA across breast, bladder, colorectal, liver, and lung cancers — cell-level annotations and tumour-microenvironment spatial features generated by Atlas H&E-TME.
Digital Pathology
allenai
Precomputed embeddings released alongside the OlmoEarth paper — drop-in feature set for evaluating Earth-observation foundation-model performance on downstream tasks without re-running the backbone.
Earth-Observation Embeddings