Coding datasets and repositories for AI training and evaluation
Coding datasets are curated, real-world code repositories (with full commit, pull-request, and test history) used to train and evaluate AI coding models and agents. Appen's corpus spans over 30 private, production-grade environments: over 25 million lines of code, over 65,000 pull requests, and execution-ready test coverage, held out from public indexes for uncontaminated benchmarking.
What you get in the coding datasets corpus
Repository-level coding datasets
Full repos with commit graphs and review context, not isolated snippets, for repository-scale software-engineering tasks.
Held-out coding benchmarks
Private environments never indexed publicly, so your evaluation signal isn't contaminated by training leakage.
Code training data with tests
58.8% average test coverage enables execution-based verification of model output, not string-match scoring.
Agentic coding evaluation data
Before/after PR pairs with reviewer context, natural tasks for evaluating coding agents.
Polyglot coding datasets
Over 10 languages (Ruby, JS/TS, PHP, Java/Kotlin, Python, Go, C#, C++) for cross-language generalization.
Off-the-shelf data catalog
Browse the wider off-the-shelf dataset portfolio across text, image, audio, and video.
How the corpus is built and verified
Sourced from real engineering teams across over 30 private repositories; each pull request is a reviewed before/after pair.
Quality controls: execution-based verification against existing test suites (58.8% avg coverage), decontamination checks against public benchmark sets, and provenance/licensing review before delivery.
Backed by Appen's 30 years in data for AI and a global contributor and engineering network.
Full dataset catalog
Over 30 qualifying environments
More than 30 environments clear the 100+ pull-request threshold. Identifiers withheld; full metadata available under NDA. Table columns: vertical, language, repos, pull requests, total LOC, commits.
Frequently asked questions
What are coding datasets?
How does Appen keep held-out benchmarks reliable?
How large is Appen's coding corpus?
How do we verify model output on this data?
Which languages and verticals do we cover?
How do I access the data?
Ready to benchmark with confidence?
Talk to our team about coding datasets and repositories for AI training and evaluation.