Interactive results · July 2026

Cost-aware security benchmarking

Language-model agents on both sides of security operations — offensive CTF (Cybench) and defensive SOC investigation (Splunk BOTS v1) — with every result decomposed by model tokens, tool calls, dollars, and time. The July 2026 set adds seven Chinese frontier models and GPT-5.6 reruns with trusted cyber access.

Hover any chart for exact values · click a model chip below a chart to show or hide its runs · jump anywhere from All charts · pick models with the Models menu.

Offense vs. defense

Offense scales with spend

Defense tracks tool discipline

Offense · CTF

Cybench

Autonomous capture-the-flag hacking: 39 hard-variant challenges × 3 epochs · ReAct agent with bash + Python in a Kali sandbox · fractional pass@1. GPT-5.6 models appear twice: ⊕ with and ⊘ without trusted cyber access (curves default to ⊕); MiniMax M3 is partial (111/117 sample-epochs).

Results

Cost, refusals & anatomy

Scaling curves

Defense · SOC

Splunk BOTS v1

Security-incident investigation over real Splunk logs: 31 scored questions (10,300 points) × 3 epochs · Splunk, web search + priced enrichment tools. The trusted-cyber (⊕) GPT-5.6 reruns are excluded here — they ran with a different context length and aren't comparable.

Results

Tool economics

Questions & scenarios

Scaling curves

Methodology

How these numbers are made

All results come from Inspect-based agent harnesses run for three epochs per task, with every outcome decomposed by model tokens, priced tool calls, dollars, and time, and all charts recomputed from the audited evaluation logs. For the full experimental setup, scoring rules, and cost-curve definitions, see “Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents” (Kassianik, Nelson & Singer, 2026) — arXiv:2607.15263.