Contact Us

APEX-Agents, developed by Mercor, evaluates AI agents on real day-to-day professional work across Law, Investment Banking, and Management Consulting. Rather than testing on synthetic puzzles, agents face the same tasks that junior professionals handle daily.

The benchmark is 33 worlds, 480 tasks, and grading rubrics.

This explorer helps browse the benchmark data, inspect tasks and worlds, and review agent traces.

Tasks
480
Worlds
33
Avg Criteria
4.1
Domains
3
Tools
21
Rubric Criteria
1,948

Model

Score

Opus 4.6 (High)

29.8% ± 3.6%

Gemini 3 Flash (High)

24% ± 3.3%

GPT 5.2 (High)

23% ± 3.2%

Opus 4.5 (High)

18.4% ± 2.9%

Gemini 3 Pro (High)

18.4% ± 2.7%

0%
20%
40%
60%

Tasks by Domain