Want to try it first? A free, individually-audited 20-task sample (10 operations + 10 investigation) is available on the Demo Samples page.
What Enterprise Bench is
A simulated company in Enterprise Bench is not a single API — it is a federation of enterprise systems over a shared database. Salesforce, Slack, Linear/Jira, Notion, GitHub, PagerDuty, Stripe, Grafana, SonarQube, Optimizely, LaunchDarkly, Greenhouse, Lattice, and more are all mounted as MCP tool servers backed by the same coherent world: the Linear ticket Slack is discussing is the same ticket the PagerDuty incident references, the Salesforce account the email mentions is the one billing tracks in Stripe. An agent works the task with native tool calls — searching, reading, and (on write tasks) mutating records across systems — exactly as a human operator would move between browser tabs. Tasks fall into two complementary families, both set in these enterprise environments:- Top-down — operations & coordination (WRITE). A persona-driven, multi-turn request to get something done across several systems: triage a CVE across Linear and Slack, onboard a customer across Salesforce/Stripe/Calendar, run a pre-launch readiness check and start an experiment. The agent discovers context across systems, then executes coordinated writes. Graded on the final state of the world — a DB-diff against a golden delta plus an LLM rubric of goal / process criteria.
- Bottom-up — investigation & QA (READ-ONLY). An answer-mined question over the same environments — “has this feature flag ever actually been enabled?”, “are these two engineers duplicating work?”, “how many of these incidents are unresolved and are they all linked to one ticket?” The agent investigates with read-only tools and returns a grounded, cited answer. Graded on answer recall against mined ground truth (entity / fact / source recall, minus a distractor penalty).
At a glance
What’s inside
Task categories
Enterprise Bench is organized by enterprise environment: each of the 27 verticals below is a distinct simulated company, and serves as a case study for the kind of cross-system enterprise work an agent must handle. Each environment is realized as a concrete named company (theai_ml_startup world ships as Lattice AI, ecommerce_scale as Cartable, mobility_fleet as RideGrid, …), populated with its own people, teams, tickets, documents, incidents, accounts, experiments, and feature flags — so a task at Lattice AI is about INT8 quantization rollbacks and GPU quota, while one in freight_logistics is about denied-party screening and shipment latency.
Both task families instantiate across these worlds, giving broad capability coverage without re-using surface content.
Expanding the categories
The 27 environments are case studies, not a closed set. Each is a world + tool servers + task generators + reward verifiers, so the corpus grows along three axes:- More verticals — manufacturing/ERP ops, banking back-office, clinical/EHR workflows, retail POS and supply chain, telco network operations, and other industries, each with its own systems and domain content.
- More task families — beyond top-down writes and bottom-up QA: scheduled / long-running operations, approval and escalation chains, policy-and-compliance enforcement, cross-system reconciliation, and incident-response runbooks.
- Deeper system coverage — pulling more of each environment’s ~40 connected systems into tasks (HR/Workday, IaC/ArgoCD, secrets/Bitwarden, observability/Datadog/Sentry, and the wider doc and CRM surface).
Difficulty profile
Bottom-up tasks carry an explicit difficulty label and reasoning tags; the distribution skews hard by design (counts from the evaluated read-only slice):
Answer types are dominated by factual (60%) and explanation (36%), with a sharp tail of count, comparison, and list questions (≈4% combined) that prove the most punishing — they require complete, exhaustive enumeration rather than locating a single fact. Top-down tasks scale on a different axis: number of systems touched, number of coordinated writes, and the order constraints between them.
How challenging is the data
As a reference point, a frontier-scale open-weight model — Qwen3.5-397B (qwen3-5-397b) — was evaluated on the bottom-up read-only QA family: 1,932 answer-mined questions across all 27 environments, scored both by each task’s recall verifier and by an independent LLM rubric.
Headline (LLM-judge rubric, reward ∈ [0, 1]):
The global mean sits near 0.56, but the signal is in the variance.
By answer type — locating one fact is tractable; exhaustive enumeration is not:
By environment — capability is sharply uneven across verticals (selected; patched rubric mean):
†
omnichannel_retail is a small slice (n=13); treat its mean as indicative, not robust.
Why it loses points. A second-pass audit of the non-perfect answers (given the gold and the docked criteria) found the misses are overwhelmingly real agent errors, not grader noise — three failure modes dominate:
- Wrong-similar-entity (≈39%) — the model grabs a near-duplicate of the right record (the adjacent ticket, the other engineer, PR #2 instead of PR #4). The environments deliberately plant look-alikes that demand ID / assignee / status cross-verification.
- Retrieval-gives-up (≈28%) — when one tool returns thin data the model concludes “no history available” instead of trying the audit log, transition history, or another system where the fact actually lives.
- Fabrication (≈17%) — on sparse data the model invents plausible-but-unsupported evidence (a Slack message, a metric, an ID) rather than stating absence.
Trajectory length
The bottom-up tasks are read-only but genuinely investigative. The table below summarizes, for the evaluated read-only slice, the assistant turns (steps) and tool calls per rollout — shown as mean / median / p90, broken out by difficulty tier:- Tool calls (~17) far outnumber assistant turns (~9) — the agent issues several searches and reads in parallel per turn.
- Trajectory length is roughly flat across tiers — the difficulty comes from which records to find and reconcile, not from longer rollouts.
- The top-down write family runs longer still — in the demo set, 21–49 turns and 10–26 tool calls across 4–5 systems each, since it interleaves discovery with coordinated writes and multi-turn confirmation.
Tool usage
The read-only investigation family is search-dominated: the agent sweeps chat, ticketing, docs, and code, then reads the specific records it needs. Across the 1,932-task slice, tool calls concentrate in:- Chat (Slack) — ≈29% of all calls;
slack__search,slack__read_channel,slack__list_channels. - Ticketing (Jira / Linear / Asana / Zendesk) —
jira__search_tickets,jira__get_ticket,linear__search_tickets,zendesk__search_tickets. - Docs (Confluence / Notion) —
confluence__search_docs,notion__search_docs. - Code & incidents (GitHub / PagerDuty) —
github__list_pull_requests,github__search_code, plus incident lookups.
linear__link_tickets, salesforce__add_note / log_activity, Stripe customer creation, notion__add_space_member, calendar writes — and is graded by the resulting database delta, not by the answer text.