LLM Penetration Testing Leaderboard

AI Security Model Performance Comparison

Comprehensive benchmarking of Large Language Models in penetration testing and vulnerability discovery. Compare detection rates, cost efficiency, and performance metrics across leading AI models in cybersecurity applications.

Started February 2025

Active Benchmarking

Real-World Testing

Community Driven

🎯 Read Our Take on this Evaluation Run

We've analyzed the data and written up our key insights from this groundbreaking run. Discover which models dominated, why cost matters more than you think, and what this means for the future of AI-powered security testing.

Read Our Analysis

~10 min read

Evaluation Run:

Single-Page-Express-App

Model Performance Vulnerability Comparison Security Finding Details

LLM Detection Rates

Comparing model detection of verified vulnerabilities

Total Cost Per Model

Cost comparison across different LLM models (US Dollars)

Cost Efficiency: Cost Per Bug Discovered

Cost per vulnerability discovered - lower is better (US Dollars)

Token Usage: Input vs Output

Comparison of input and output token consumption by model

Execution Time Analysis

Total execution time and number of LLM calls per model

Average Response Time Per Call

Average latency per LLM call - lower is faster

google-gemini-2.5

Found 9 of 9 vulnerabilities (100%)
Critical
4
High
5
Medium
0
Low
0
Tokens: 96,858 Cost: $0.2977

Cost per bug: $0.0331
Calls: 12 Total: 13m 60s

Avg per call: 70.0s

google-gemini-2.5-flash

Found 6 of 9 vulnerabilities (67%)
Critical
3
High
3
Medium
0
Low
0
Tokens: 61,719 Cost: $0.0149

Cost per bug: $0.0025
Calls: 8 Total: 2m 31s

Avg per call: 18.8s

openai-gpt-4.1

Found 6 of 9 vulnerabilities (67%)
Critical
3
High
3
Medium
0
Low
0
Tokens: 58,489 Cost: $0.1832

Cost per bug: $0.0305
Calls: 8 Total: 3m 1s

Avg per call: 22.6s

openai-o3

Found 5 of 9 vulnerabilities (56%)
Critical
3
High
2
Medium
0
Low
0
Tokens: 51,581 Cost: $0.7788

Cost per bug: $0.1558
Calls: 7 Total: 3m 12s

Avg per call: 27.4s

claude-opus-4-anthropic

Found 5 of 9 vulnerabilities (56%)
Critical
3
High
2
Medium
0
Low
0
Tokens: 50,197 Cost: $1.3301

Cost per bug: $0.2660
Calls: 7 Total: 3m 37s

Avg per call: 31.1s

claude-sonnet-4-anthropic

Found 5 of 9 vulnerabilities (56%)
Critical
3
High
2
Medium
0
Low
0
Tokens: 49,850 Cost: $0.2561

Cost per bug: $0.0512
Calls: 7 Total: 2m 33s

Avg per call: 21.9s

lmstudio-qwen3

Found 4 of 9 vulnerabilities (44%)
Critical
3
High
1
Medium
0
Low
0
Tokens: 40,753 Cost: $0.0000

Cost per bug: $0.0000
Calls: 6 Total: 7m 54s

Avg per call: 79.1s

openai-gpt-4.1-nano

Found 4 of 9 vulnerabilities (44%)
Critical
3
High
1
Medium
0
Low
0
Tokens: 43,081 Cost: $0.0066

Cost per bug: $0.0017
Calls: 6 Total: 1m 10s

Avg per call: 11.6s

Model Detection Rate by Vulnerability Category

openai-o3

Percentage of vulnerabilities detected by category

claude-opus-4-anthropic

Percentage of vulnerabilities detected by category

claude-sonnet-4-anthropic

Percentage of vulnerabilities detected by category

google-gemini-2.5-flash

Percentage of vulnerabilities detected by category

google-gemini-2.5

Percentage of vulnerabilities detected by category

lmstudio-qwen3

Percentage of vulnerabilities detected by category

openai-gpt-4.1

Percentage of vulnerabilities detected by category

openai-gpt-4.1-nano

Percentage of vulnerabilities detected by category