Windows Enterprise Intrusion
Sep 2026BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting
About BlueBench-Intrusion
The series covers three environments across three benchmarks: 001 macOS Threat Investigation, 002 Windows Enterprise Intrusion, and 003 AWS Cloud Intrusion.
About This Benchmark
The dataset contains 3.9M+ real events across 40 log tables (EDR telemetry, Windows event logs, Zeek network metadata, and IDS/EDR alerts), plus the recovered malware and a memory image. Agents get SQL query tools, binary analysis utilities, and Volatility 3, and are scored on guided detection engineering, malware analysis, and open-ended IR and TH reports. Every submitted detection rule is re-executed against the live dataset; a rule that doesn't run scores zero.
Sample Questions
Q: Write a detection rule for the WMI-driven remote execution observed in this case: the wmiprvse.exe → cmd.exe → powershell.exe lineage, with noise controls that separate attacker activity from IT automation.
A: [DuckDB SQL, re-executed against the dataset and graded on required fields, noise controls, and whether it fires]
Q: Two differently-named binaries were recovered from the gateway host. Using only artifacts inside each binary, determine whether they are the same underlying tool and explain what the rename reveals about the attacker’s masquerading approach.
A: Same open-source tunneling tool, renamed to impersonate a security product (specifics redacted)
Q: Starting from a single high-severity alert, reconstruct the incident and write a complete incident response report: scope, attack chain, affected accounts and hosts, and containment recommendations.
A: [Free-form report scored against a hidden rubric of expert-curated facts and evidence checks]
Key Findings
Results
GPT-5.6 Sol had the highest benchmark score at 80.7%. The top results among 16 evaluated models were GPT-5.6 Sol (80.7%), GLM-5.3 (78.2%), GPT-6 Astra (76.2%). These are observed scores on this benchmark.
- 1GPT-5.6 Sol
- 2GLM-5.3
- 3GPT-6 Astra
- 4Kimi K3
- 5Claude Opus 4.8
- 6Grok 4.6
- 7GLM-5.2
- 8GPT-5.6 Terra
- 9DeepSeek V4 Flash
- 10GPT-5.6 Luna
Track results
Detection Engineering: GLM-5.2 led at 95.8%. Malware Analysis: GPT-6 Astra led at 88.8%. IR Report: Kimi K3 led at 77.8%. TH Report: GLM-5.3 led at 89.7%.
Open-weight results
GLM-5.3 (78.2%) had the highest score among the 6 open-weight models evaluated. GLM-5.3 cost $1.91/task.
Cost per task
DeepSeek V4 Flash had the lowest measured cost at $0.12/task, with 62.9% benchmark score. GPT-5.6 Sol scored 80.7% at $1.88/task. Costs reflect the evaluated task mix and the stated pricing assumptions.
- 1DeepSeek V4 Flash
- 2GPT-5.6 Luna
- 3GPT-5.6 Terra
- 4GLM-5.2
- 5Claude Opus 4.8
- 6DeepSeek V4 Pro
- 7MiniMax M3
- 8Claude Sonnet 5
- 9GPT-5.6 Sol
- 10Gemini 3.6 Flash
Time per task
GPT-5.6 Terra had the lowest reported average latency at 1m 29s/task, with 68.6% benchmark score. Latency includes the work performed by the agent during evaluation.
- 1GPT-5.6 Terra
- 2GPT-5.6 Luna
- 3Claude Opus 4.8
- 4Claude Sonnet 5
- 5GPT-5.6 Sol
- 6GPT-6 Astra
- 7GLM-5.2
- 8Gemini 3.6 Flash
- 9DeepSeek V4 Flash
- 10DeepSeek V4 Pro
Reliability: score spread
GPT-6 Astra had the smallest observed score spread at 34.7 percentage points; its overall score was 76.2%. Spread is the highest minus lowest score across recorded tasks and trials. It includes differences in task difficulty and is not a measure of repeatability on one fixed task.
- 1GPT-6 Astra
- 2Kimi K3
- 3GLM-5.3
- 4Gemini 3.6 Flash
- 5GPT-5.6 Sol
- 6GPT-5.6 Terra
- 7Grok 4.6
- 8DeepSeek V4 Pro
- 9GLM-5.2
- 10GPT-5.6 Luna
Run completion
16 of 16 models completed all recorded runs without an unrecoverable error. Completion does not imply a correct answer or rule out a refusal.
- 1GPT-5.6 Sol
- 2GLM-5.3
- 3GPT-6 Astra
- 4Kimi K3
- 5Claude Opus 4.8
- 6Grok 4.6
- 7GLM-5.2
- 8GPT-5.6 Terra
- 9DeepSeek V4 Flash
- 10GPT-5.6 Luna
Measured Trade-offs
- GPT-5.6 Sol — Highest observed benchmark score: 80.7% at $1.88/task.
- DeepSeek V4 Flash — Lowest measured cost: $0.12/task, with 62.9% benchmark score.
- GLM-5.3 — Highest observed open-weight benchmark score: 78.2% at $1.91/task.
Methodology
Scoring
- Score: Equal-weight mean of detection engineering, incident response, threat hunting, and malware analysis (25% each). This replaces the earlier guided-investigation score
- Detection Rules: Re-executed against the dataset and graded on whether they fire on the attacker behavior and control noise
- Reports: Open-ended IR and TH reports scored against hidden expert rubrics
- Cost / task: USD per logical evaluated task, including repeated attempts. A bundled question set counts its individual questions; a report counts once. Prompt caching is included; judge costs are excluded
- Latency: Average elapsed time per logical evaluated task, using the export normalization for bundled question sets
- Reliability / spread: On this page, highest minus lowest score across recorded tasks and trials, in percentage points (lower is tighter). The BlueBench reliability score is 100 minus the mean of these six spreads: inverted spread, not same-task repeatability
Setup
- Logs from a real Windows Active Directory intrusion
- 3.9M+ events across 40 log tables covering five hosts: EDR telemetry, Windows event logs, Zeek network metadata, and IDS/EDR alerts
- Recovered malware artifacts and a memory image available for binary analysis and Volatility 3 memory forensics
- 40 scored tasks per model: 18 detection engineering, 18 malware analysis, and 4 open-ended investigation reports
Scoring
- Detection rules are re-executed against the live dataset; a rule that fails to run scores zero
- Analysis questions are LLM-judged against ground truth from the original intrusion
- IR and threat hunting reports are written open-ended with no section guidance, then scored against hidden expert rubrics
Controls
- Same minimal system prompt for all models, no per-model tuning
- "Thinking" mode enabled where available
- Agent loops capped at 100 iterations for guided tasks and 150 for reports
- All models given identical tool access and data
Caveats
- Results describe the current export and its recorded task mix; counts of repeated attempts can differ across models.
- Spread is the highest minus lowest score across recorded tasks and trials. Different task difficulty and question bundling affect it; it does not isolate same-task repeatability.
- Costs are USD per evaluated task, including repeated attempts. Bundled question sets use logical question counts; each report attempt is counted once. Provider prompt caching is included and judge costs are excluded.
- In a track-filtered view, the score changes but cost and latency remain the benchmark-wide averages; spread is shown only in the all-tracks view.
- Report scores depend on LLM judging. Detection queries are executed against the evidence. Close scores do not establish a statistically significant ordering.
Other Benchmarks
BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth
BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth
BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth
BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, testing detection engineering and evidence-backed investigation reports
BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering
Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving
Multiple-choice cybersecurity knowledge evaluation across 10,000 questions
Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills
Multi-label classification of MITRE ATT&CK techniques from Sigma rules
Blue team CTF scenarios testing incident response and threat hunting
AI for the blue team.
Run Cotool's harness in your environment to get real security work done