Coding datasets and repositories for AI training and evaluation

Coding datasets are curated, real-world code repositories (with full commit, pull-request, and test history) used to train and evaluate AI coding models and agents. Appen's corpus spans over 30 private, production-grade environments: over 25 million lines of code, over 65,000 pull requests, and execution-ready test coverage, held out from public indexes for uncontaminated benchmarking.

How the corpus is built and verified

Sourced from real engineering teams across over 30 private repositories; each pull request is a reviewed before/after pair.

Quality controls: execution-based verification against existing test suites (58.8% avg coverage), decontamination checks against public benchmark sets, and provenance/licensing review before delivery.

Backed by Appen's 30 years in data for AI and a global contributor and engineering network.

Off-the-shelf datasets

Full dataset catalog

Over 30 qualifying environments

More than 30 environments clear the 100+ pull-request threshold. Identifiers withheld; full metadata available under NDA. Table columns: vertical, language, repos, pull requests, total LOC, commits.

FAQ

Frequently asked questions

What are coding datasets?

We curate real-world code repositories (with commit, PR, and test history) used to train and evaluate coding models and agents.

How does Appen keep held-out benchmarks reliable?

Our environments have never been publicly indexed, so it can't leak into a model's training data, keeping evaluation signal uncontaminated.

How large is Appen's coding corpus?

Over 30 private environments, over 25 million lines of code, over 65,000 pull requests, across over 10 languages, at 58.8% average test coverage.

How do we verify model output on this data?

We run execution-based verification against each repository's existing test suite, not string matching.

Which languages and verticals do we cover?

We cover over 10 languages across SaaS, e-commerce, healthcare, fintech, HR tech, and more.

How do I access the data?

Identifiers are withheld; full metadata and samples are available under NDA. Talk to an expert.

Ready to benchmark with confidence?

Talk to our team about coding datasets and repositories for AI training and evaluation.

Talk to an expert

Contact us

Thank you for getting in touch! We appreciate you contacting Appen. One of our colleagues will get back in touch with you soon! Have a great day!