v1.2.2 · Commercial License

24,316
Synthetic Enterprise Worlds

Fully simulated software companies with codebases, team Slack conversations, injected bugs, engineering missions, and architecture decisions — purpose-built for training and benchmarking AI agents on real enterprise scenarios.

✓ 100% synthetic provenance ✓ IP audit passed 10/10 ✓ 140-test validation suite
WORLD-00AF966D · Fintech Sample
Codebase 42 files · 3 repos
Slack History 14 threads · 2 incidents
Injected Bugs 3 · critical + medium
Missions 3 · difficulty 1–5
ADRs 3 architecture records
24,316

Enterprise Worlds

Across 10 domains from fintech to biotech.

97,290

RAG Evaluation Tasks

Retrieval, reasoning, and grounding QA tasks.

10

Enterprise Domains

Fintech, healthcare, security, SaaS, and more.

100%

Fact Accuracy

Structured labels verified via 140-test suite.

What's Inside a World

Every world is a complete, self-contained enterprise software environment.

01

Source Code

Multi-repository codebases with realistic stub implementations across JavaScript, TypeScript, and Python. Each world includes 30–60 files organized by microservice boundaries.

src/services/payment/validate.js src/utils/auth/middleware.ts tests/integration/api.test.js
02

Team Slack History

Multi-participant conversations about production incidents, feature planning, and architecture debates. Threads reference specific code files and inject realistic team dynamics.

@maria: The payment validation is silently dropping
transactions over $50K. Checked the logs — no error thrown.
@raj: We need to add a threshold check before
the currency conversion step.
03

Injected Bugs + Missions

Each world contains 1–3 deliberately injected bugs across critical and medium severity. Three missions per world, with the first one targeting the injected bug for realistic debugging.

CRITICAL Missing input validation on processRefund() allows negative amounts.
04

Architecture Decisions

ADR documents explaining the rationale behind technical choices — database selection, service boundaries, caching strategies. Grounds evaluation tasks in documented decisions.

ADR-005: Adopt Kafka for payment event streaming Status: Accepted · Decides: Event sourcing pattern
05

Company Metadata

Narrative metadata per world: company description, organizational structure, engineering culture, active initiatives, and persona profiles with roles and domain expertise.

Employees: 142 · Culture: Remote-first Stack: Node.js, PostgreSQL, Kafka, Kubernetes
06

RAG Evaluation Tasks

Three task types extracted from each world: retrieval (find relevant code from Slack), reasoning (multi-step bug diagnosis), and grounding (factual QA from metadata/ADR).

Retrieval Reasoning Grounding

10 Enterprise Domains

Balanced coverage across the most demanding enterprise software verticals.

Healthcare 524 worlds
Legal 504 worlds
Security 490 worlds
SaaS 480 worlds
E-Commerce 472 worlds
Biotech 468 worlds
Energy 460 worlds
Gaming 450 worlds
Education 448 worlds
Fintech 428 worlds

Editions sampled by strategic domain importance + proportional representation.

Proven Quality.
Verified Integrity.

Every world passes a multi-layer validation pipeline before inclusion. No blind generation.

140 Automated Tests

12 test files covering label correctness, schema validity, org name uniqueness, RAG export integrity, and bug-category balance. Run before every release.

Label Correctness Audit

Two-layer audit: rule-based checks on all 97,290 tasks + automated AI-assisted verification on a stratified sample. 100% structured fact accuracy.

ChatGPT IP Audit

Independent third-party review of all documentation and sample data. Score: 10/10. IP risk: 1/10. No organizational, procedural, or trademark concerns.

Split Integrity

Deterministic train/val/test split (70/15/15) via SHA-256 hash. Zero overlap across splits. Full reproducibility — generate the same worlds with the same seed.

Human Review

50-world manual spot check confirmed structural correctness, narrative coherence, and mission-bug alignment. All documented in QA report.

Commercial License

No restrictive "research only" shackles. Full Commercial Training License — train models, deploy in production, benchmark, and publish results.

Built For

Four primary use cases validated by our early-access partners.

01

AI Agent Training

Train coding agents to debug, diagnose, and fix real-world bugs in multi-file enterprise codebases — not isolated LeetCode problems.

76.5% field accuracy 3+ files per mission
02

RAG Evaluation

Benchmark retrieval-augmented generation pipelines across 97,290 structured tasks spanning retrieval, reasoning, and factual grounding.

97,290 tasks 3 task types
03

Agent Benchmarking

Replace SWE-bench with richer, structured environments. Measure task completion, bug-fix accuracy, and architectural reasoning in simulated enterprises.

vs SWE-bench · richer 10 domains · harder
04

Fine-Tuning Data

Generate supervised fine-tuning examples from mission-code pairs, Slack-based reasoning traces, and ADR-grounded QA for domain-specific models.

72,891 grounding tasks 1–5 difficulty scale

Editions

From free samples to enterprise-wide deployment. All editions include full Commercial Training License.

Feature Evaluation Research Professional Enterprise
5 Preview Worlds
1,000 Worlds
5,000 Worlds
Full 24,316 Worlds
Commercial License
QA Report
Dataset Updates 6 months 12 months Ongoing
Priority Support
Evaluation
5 worlds · Evaluate quality before you commit.
  • 5 enterprise worlds
  • Sample from all 10 domains
  • All 5 components per world
  • Unrestricted evaluation
Download Free Sample
Professional
5,000 worlds · Built for commercial AI teams.
  • 5,000 enterprise worlds
  • Stratified domain sampling
  • 20,000+ RAG evaluation tasks
  • Full commercial training license
  • Priority support
Purchase →

Download Free Sample

5 complete enterprise worlds — code, Slack, bugs, missions, ADRs. Enter your details to receive the download link.

Please fill in all required fields and agree to the evaluation license.
Evaluation License Summary By downloading, you agree to:
  • Use the dataset for evaluation purposes only
  • No redistribution, resale, or public hosting
  • No commercial training without a paid license
  • Delete all copies if you do not purchase a license
Read full license agreement →
You've reached the download limit. Please try again later or contact us.
Please complete the security check.

We've emailed your download

Check your inbox for the download link and instructions.

  • ✓ 5 complete enterprise worlds
  • ✓ QA Report
  • ✓ Dataset Card
  • ✓ Evaluation License

Need 1,000 or 5,000 worlds?

Request Pricing →
html> >