Overview
Understanding the APEX-Agents benchmark and tool ecosystem
What is APEX-Agents?
APEX-Agents (Agentic Professional Evaluations and Experiments) is a benchmark designed to evaluate AI agents on realistic professional services tasks. The benchmark covers three domains: Law, Investment Banking, and Management Consulting.
Each task places an AI agent in a simulated professional environment ("world") with access to various MCP (Model Context Protocol) tool servers. The agent must use these tools to complete real-world workflows - from drafting legal documents to analyzing financial spreadsheets.
Tasks are graded against detailed rubric criteria using automated verifiers, providing a standardized way to measure agent capabilities across professional domains.
Key Concepts
Worlds
Simulated professional environments with pre-configured file systems, documents, email accounts, calendars, and other tools. Each world represents a realistic workplace setup.
Tasks
Specific work assignments within a world. Tasks include a prompt, expected output type, gold response, and rubric criteria for automated evaluation.
MCP Servers
Tool servers that agents interact with via the Model Context Protocol. Each server provides a set of tools (file I/O, document editing, email, etc.) that agents use to complete tasks.
Rubric Criteria
Detailed evaluation criteria for each task. Each criterion has a verifier that checks whether the agent's output meets the requirement (file existence, content matching, etc.).
Tool Ecosystem
21 tools across 9 MCP servers
Core Infrastructure
Read, search, and inspect files and directories
Code Execution
Office Productivity
Create, read, and edit Word documents (.docx) with structured operations
Create, read, and edit PowerPoint presentations (.pptx) with slide operations
Create, read, and edit Excel spreadsheets (.xlsx) with cell-level operations