Cost-aware security benchmarking
Language-model agents on both sides of security operations — offensive CTF (Cybench) and defensive SOC investigation (Splunk BOTS v1) — with every result decomposed by model tokens, tool calls, dollars, and time. The July 2026 set adds seven Chinese frontier models and GPT-5.6 reruns with trusted cyber access.
Hover any chart for exact values · click a model chip below a chart to show or hide its runs · jump anywhere from All charts · pick models with the Models menu.
- Offense doesn’t predict defense. GPT-5.5 leads offense (94.1% Cybench pass@1); Claude Opus 4.8 leads defense (93.8% of BOTS points, with GPT-5.5 at 81.0%).
- Offense scales with spend. Raising the per-sample dollar budget keeps buying more solved challenges — no frontier model saturates within its run budget.
- Defense is tool discipline, not spend. The best defender is near its ceiling within ~20 tool calls per question; extra calls and dollars buy little.
- Refusals gate offense — until access. Claude Fable 5 refuses every Cybench task (0% offense, 88.4% on defense); GPT-5.6 Sol refuses 90.6% without trusted cyber access but solves 87.2% with it. Policy and access tier, not capability, set the newest US models’ scores.
Offense vs. defense
Offense scales with spend
Defense tracks tool discipline
Cybench
Autonomous capture-the-flag hacking: 39 hard-variant challenges × 3 epochs · ReAct agent with bash + Python in a Kali sandbox · fractional pass@1. GPT-5.6 models appear twice: ⊕ with and ⊘ without trusted cyber access (curves default to ⊕); MiniMax M3 is partial (111/117 sample-epochs).
Results
Cost, refusals & anatomy
Scaling curves
Splunk BOTS v1
Security-incident investigation over real Splunk logs: 31 scored questions (10,300 points) × 3 epochs · Splunk, web search + priced enrichment tools. The trusted-cyber (⊕) GPT-5.6 reruns are excluded here — they ran with a different context length and aren't comparable.
Results
Tool economics
Questions & scenarios
Scaling curves
How these numbers are made
All results come from Inspect-based agent harnesses run for three epochs per task, with
every outcome decomposed by model tokens, priced tool calls, dollars, and time, and all
charts recomputed from the audited evaluation logs. For the full experimental setup,
scoring rules, and cost-curve definitions, see
“Beyond
Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security
Agents” (Kassianik, Nelson & Singer, 2026) — arXiv:2607.15263.