SWE-Enterprise vs HumanEval
HumanEval tests 164 isolated Python function signatures. SWE-Enterprise tests real-world enterprise debugging across 97,290 multi-file tasks.
Head-to-Head Comparison
How SWE-Enterprise stacks up against HumanEval.
| Dimension | SWE-Enterprise | HumanEval |
|---|---|---|
| Total Tasks | 97,290 | 164 |
| Enterprise Context | Code + Slack + ADRs + Metadata | Single function only |
| Multi-File Reasoning | 3+ files per mission | Single function |
| Bug Complexity | Realistic multi-step bugs | Simple function errors |
| RAG Evaluation | Retrieval + Reasoning + Grounding | None |
| Domain Coverage | 10 enterprise domains | Python only |
| Scale | 24,316 worlds | 164 problems |