SWE-Enterprise vs SWE-bench
SWE-bench tests isolated Python bug fixes. SWE-Enterprise adds enterprise context, multi-file reasoning, Slack history, and architecture decisions at 40x the task volume.
Head-to-Head Comparison
How SWE-Enterprise stacks up against SWE-bench.
| Dimension | SWE-Enterprise | SWE-bench |
|---|---|---|
| Total Tasks | 97,290 | 2,294 |
| Enterprise Context | Code + Slack + ADRs + Metadata | Single file only |
| Multi-File Reasoning | 3+ files per mission | Single file |
| Bug Injection | 1–3 deliberate bugs per world | Real backported bugs |
| RAG Evaluation | Retrieval + Reasoning + Grounding | None |
| Domain Coverage | 10 enterprise domains | Python only |
| Data Richness | Complete simulated companies | Git diff + test only |