Skip to main content

Windows Enterprise Intrusion

Sep 2026

BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting

Metric
Track
View
Provider
40 tasks · 16 models · Sep 2026
RankModels (16)ScoreSpread?Highest minus lowest score across recorded tasks and trials (percentage points). Lower is less variable.$/taskLat
1
GPT-5.6 Sol
80.7%55.3pp$1.883m 51s
2
GLM-5.3
78.2%43.5pp$1.919m 34s
3
GPT-6 Astra
76.2%34.7pp$9.433m 52s
4
Kimi K3
76.0%40.6pp$1.888m 45s
5
Claude Opus 4.8
72.6%90.0pp$0.532m 58s
6
Grok 4.6
72.1%63.5pp$2.249m 48s
7
GLM-5.2
70.6%71.1pp$0.454m 18s
8
GPT-5.6 Terra
68.6%58.2pp$0.441m 29s
9
DeepSeek V4 Flash
62.9%80.6pp$0.125m 4s
10
GPT-5.6 Luna
61.0%71.2pp$0.271m 38s
Best ScoreBest Open Weight
40 tasks
·
3.9M+ log events
Task Distribution
Detection Engineering·18 tasks(25% of score)
Malware Analysis·18 tasks(25% of score)
IR Report·2 tasks(25% of score)
TH Report·2 tasks(25% of score)
Detection Engineering
Malware Analysis
IR Report
TH Report
Detection Engineering
Malware Analysis
Incident Response
Threat Hunting
Lateral Movement
Credential Access
Memory Forensics

About BlueBench-Intrusion

BlueBench-Intrusion is built on real compromises. Each benchmark replays a genuine intrusion carried out by a hands-on-keyboard attacker in a controlled lab and captured with production-grade telemetry: EDR, host and cloud audit logs, Zeek network metadata, and the alerts a SOC would actually see. Nothing is synthesized; the data is what the attacker left behind, surrounded by the routine activity of the environment around them. Agents investigate with SQL query tools over the raw logs and, where a case includes them, recovered malware and memory images. Reports are graded against hidden expert rubrics and every submitted detection rule is re-executed against the live dataset. The intrusion datasets were developed in partnership with Threat Hunting Labs.

The series covers three environments across three benchmarks: 001 macOS Threat Investigation, 002 Windows Enterprise Intrusion, and 003 AWS Cloud Intrusion.

About This Benchmark

The second entry in the BlueBench-Intrusion series is built on a real intrusion: a hands-on-keyboard attacker working through a Windows Active Directory environment in a controlled lab. The attacker sprayed RDP passwords, silently installed a remote-access tool, created a hidden admin account, dumped credentials with Mimikatz, tunneled traffic through a renamed open-source tool disguised as a Windows service, moved laterally over WMI across five hosts, and stole domain credentials from the DC via replication.
Attack path
RDP password spray
Remote-access tool install
Hidden admin account
Mimikatz credential dump
Tunneled C2 service
WMI lateral movement
DC credential replication

The dataset contains 3.9M+ real events across 40 log tables (EDR telemetry, Windows event logs, Zeek network metadata, and IDS/EDR alerts), plus the recovered malware and a memory image. Agents get SQL query tools, binary analysis utilities, and Volatility 3, and are scored on guided detection engineering, malware analysis, and open-ended IR and TH reports. Every submitted detection rule is re-executed against the live dataset; a rule that doesn't run scores zero.

Sample Questions

Detection Engineering

Q: Write a detection rule for the WMI-driven remote execution observed in this case: the wmiprvse.exe → cmd.exe → powershell.exe lineage, with noise controls that separate attacker activity from IT automation.

A: [DuckDB SQL, re-executed against the dataset and graded on required fields, noise controls, and whether it fires]

Malware Analysis

Q: Two differently-named binaries were recovered from the gateway host. Using only artifacts inside each binary, determine whether they are the same underlying tool and explain what the rename reveals about the attacker’s masquerading approach.

A: Same open-source tunneling tool, renamed to impersonate a security product (specifics redacted)

IR Report

Q: Starting from a single high-severity alert, reconstruct the incident and write a complete incident response report: scope, attack chain, affected accounts and hosts, and containment recommendations.

A: [Free-form report scored against a hidden rubric of expert-curated facts and evidence checks]

Key Findings

Results

GPT-5.6 Sol had the highest benchmark score at 80.7%. The top results among 16 evaluated models were GPT-5.6 Sol (80.7%), GLM-5.3 (78.2%), GPT-6 Astra (76.2%). These are observed scores on this benchmark.

Score by ModelTop 10 of 16
  1. 1GPT-5.6 Sol80.7%
  2. 2GLM-5.378.2%
  3. 3GPT-6 Astra76.2%
  4. 4Kimi K376.0%
  5. 5Claude Opus 4.872.6%
  6. 6Grok 4.672.1%
  7. 7GLM-5.270.6%
  8. 8GPT-5.6 Terra68.6%
  9. 9DeepSeek V4 Flash62.9%
  10. 10GPT-5.6 Luna61.0%

Track results

Detection Engineering: GLM-5.2 led at 95.8%. Malware Analysis: GPT-6 Astra led at 88.8%. IR Report: Kimi K3 led at 77.8%. TH Report: GLM-5.3 led at 89.7%.

Open-weight results

GLM-5.3 (78.2%) had the highest score among the 6 open-weight models evaluated. GLM-5.3 cost $1.91/task.

Cost per task

DeepSeek V4 Flash had the lowest measured cost at $0.12/task, with 62.9% benchmark score. GPT-5.6 Sol scored 80.7% at $1.88/task. Costs reflect the evaluated task mix and the stated pricing assumptions.

Cost per TaskTop 10 of 16
  1. 1DeepSeek V4 Flash$0.12
  2. 2GPT-5.6 Luna$0.27
  3. 3GPT-5.6 Terra$0.44
  4. 4GLM-5.2$0.45
  5. 5Claude Opus 4.8$0.53
  6. 6DeepSeek V4 Pro$0.90
  7. 7MiniMax M3$0.91
  8. 8Claude Sonnet 5$1.13
  9. 9GPT-5.6 Sol$1.88
  10. 10Gemini 3.6 Flash$1.88

Time per task

GPT-5.6 Terra had the lowest reported average latency at 1m 29s/task, with 68.6% benchmark score. Latency includes the work performed by the agent during evaluation.

Task Duration (avg)Top 10 of 16
  1. 1GPT-5.6 Terra1m 29s
  2. 2GPT-5.6 Luna1m 38s
  3. 3Claude Opus 4.82m 58s
  4. 4Claude Sonnet 53m 18s
  5. 5GPT-5.6 Sol3m 51s
  6. 6GPT-6 Astra3m 52s
  7. 7GLM-5.24m 18s
  8. 8Gemini 3.6 Flash4m 34s
  9. 9DeepSeek V4 Flash5m 4s
  10. 10DeepSeek V4 Pro6m 16s

Reliability: score spread

GPT-6 Astra had the smallest observed score spread at 34.7 percentage points; its overall score was 76.2%. Spread is the highest minus lowest score across recorded tasks and trials. It includes differences in task difficulty and is not a measure of repeatability on one fixed task.

Score spread (percentage points)Top 10 of 16
  1. 1GPT-6 Astra34.7pp
  2. 2Kimi K340.6pp
  3. 3GLM-5.343.5pp
  4. 4Gemini 3.6 Flash48.0pp
  5. 5GPT-5.6 Sol55.3pp
  6. 6GPT-5.6 Terra58.2pp
  7. 7Grok 4.663.5pp
  8. 8DeepSeek V4 Pro69.0pp
  9. 9GLM-5.271.1pp
  10. 10GPT-5.6 Luna71.2pp

Run completion

16 of 16 models completed all recorded runs without an unrecoverable error. Completion does not imply a correct answer or rule out a refusal.

Task Completion RateTop 10 of 16
  1. 1GPT-5.6 Sol100.0%
  2. 2GLM-5.3100.0%
  3. 3GPT-6 Astra100.0%
  4. 4Kimi K3100.0%
  5. 5Claude Opus 4.8100.0%
  6. 6Grok 4.6100.0%
  7. 7GLM-5.2100.0%
  8. 8GPT-5.6 Terra100.0%
  9. 9DeepSeek V4 Flash100.0%
  10. 10GPT-5.6 Luna100.0%

Measured Trade-offs

  • GPT-5.6 Sol Highest observed benchmark score: 80.7% at $1.88/task.
  • DeepSeek V4 Flash Lowest measured cost: $0.12/task, with 62.9% benchmark score.
  • GLM-5.3 Highest observed open-weight benchmark score: 78.2% at $1.91/task.

Methodology

Scoring

  • Score: Equal-weight mean of detection engineering, incident response, threat hunting, and malware analysis (25% each). This replaces the earlier guided-investigation score
  • Detection Rules: Re-executed against the dataset and graded on whether they fire on the attacker behavior and control noise
  • Reports: Open-ended IR and TH reports scored against hidden expert rubrics
  • Cost / task: USD per logical evaluated task, including repeated attempts. A bundled question set counts its individual questions; a report counts once. Prompt caching is included; judge costs are excluded
  • Latency: Average elapsed time per logical evaluated task, using the export normalization for bundled question sets
  • Reliability / spread: On this page, highest minus lowest score across recorded tasks and trials, in percentage points (lower is tighter). The BlueBench reliability score is 100 minus the mean of these six spreads: inverted spread, not same-task repeatability

Setup

  • Logs from a real Windows Active Directory intrusion
  • 3.9M+ events across 40 log tables covering five hosts: EDR telemetry, Windows event logs, Zeek network metadata, and IDS/EDR alerts
  • Recovered malware artifacts and a memory image available for binary analysis and Volatility 3 memory forensics
  • 40 scored tasks per model: 18 detection engineering, 18 malware analysis, and 4 open-ended investigation reports

Scoring

  • Detection rules are re-executed against the live dataset; a rule that fails to run scores zero
  • Analysis questions are LLM-judged against ground truth from the original intrusion
  • IR and threat hunting reports are written open-ended with no section guidance, then scored against hidden expert rubrics

Controls

  • Same minimal system prompt for all models, no per-model tuning
  • "Thinking" mode enabled where available
  • Agent loops capped at 100 iterations for guided tasks and 150 for reports
  • All models given identical tool access and data

Caveats

  • Results describe the current export and its recorded task mix; counts of repeated attempts can differ across models.
  • Spread is the highest minus lowest score across recorded tasks and trials. Different task difficulty and question bundling affect it; it does not isolate same-task repeatability.
  • Costs are USD per evaluated task, including repeated attempts. Bundled question sets use logical question counts; each report attempt is counted once. Provider prompt caching is included and judge costs are excluded.
  • In a track-filtered view, the score changes but cost and latency remain the benchmark-wide averages; spread is shown only in the all-tracks view.
  • Report scores depend on LLM judging. Detection queries are executed against the evidence. Close scores do not establish a statistically significant ordering.

Other Benchmarks

Impact & Exfiltration
Impact & Exfiltration

BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth

Simulation
Ransomware
Backup Sabotage
Data Exfiltration
Windows
Linux
3 samples · 16 models · Sep 2026
Identity & Active Directory Attacks
Identity & Active Directory Attacks

BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth

Simulation
Active Directory
Password Spray
Domain Persistence
Credential Access
Hands-on-Keyboard
Windows
4 samples · 16 models · Sep 2026
Initial Access & Command-and-Control
Initial Access & Command-and-Control

BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth

Simulation
Initial Access
Command & Control
Mail Compromise
Webshell
DGA Botnet
Windows
Linux
5 samples · 16 models · Sep 2026
AWS Cloud Intrusion
AWS Cloud Intrusion

BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, testing detection engineering and evidence-backed investigation reports

AWS
Incident Response
CloudTrail
Identity
Persistence
Exfiltration
8 samples · 16 models · Sep 2026
macOS Threat Investigation
macOS Threat Investigation

BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering

Incident Response
Threat Hunting
Detection Engineering
macOS Forensics
Credential Access
Data Exfiltration
14 samples · 16 models · Sep 2026
NYU CTF Bench
NYU CTF Bench

Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving

Reverse Engineering
Forensics
Miscellaneous
81 samples · 11 models · Feb 2026
CyberMetric
CyberMetric

Multiple-choice cybersecurity knowledge evaluation across 10,000 questions

Standards & Certifications
Network Security
Cryptography
Risk Management
Access Control
Incident Response
Application Security
Cloud Security
10180 samples · 13 models · Feb 2026
Cybench (Defensive Subset)
Cybench (Defensive Subset)

Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills

Forensics
Reverse Engineering
Miscellaneous
Hardware
18 samples · 10 models · Jan 2026
Sigma Detection Classification
Sigma Detection Classification

Multi-label classification of MITRE ATT&CK techniques from Sigma rules

Detection Engineering
MITRE ATT&CK
SIEM
Windows
Linux
Cloud
Network
Application
2733 samples · 12 models · Jan 2026
BOTSv3 Blue Team CTF
BOTSv3 Blue Team CTF

Blue team CTF scenarios testing incident response and threat hunting

Incident Response
Threat Hunting
Alert Triage
Log Analysis
Advanced Persistent Threat (APT)
Cloud Security (AWS/Azure)
51 samples · 15 models · Dec 2025

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo