LLM Penetration Testing Leaderboard
AI Security Model Performance Comparison
Comprehensive benchmarking of Large Language Models in penetration testing and vulnerability discovery. Compare detection rates, cost efficiency, and performance metrics across leading AI models in cybersecurity applications.
Started February 2025
Active Benchmarking
Real-World Testing
Community Driven
🎯 Read Our Take on this Evaluation Run
We've analyzed the data and written up our key insights from this groundbreaking run. Discover which models dominated, why cost matters more than you think, and what this means for the future of AI-powered security testing.
~10 min read
Evaluation Run:
Single-Page-Express-App
Model Performance Vulnerability Comparison Security Finding Details
LLM Detection Rates
Comparing model detection of verified vulnerabilities
Total Cost Per Model
Cost comparison across different LLM models (US Dollars)
Cost Efficiency: Cost Per Bug Discovered
Cost per vulnerability discovered - lower is better (US Dollars)
Token Usage: Input vs Output
Comparison of input and output token consumption by model
Execution Time Analysis
Total execution time and number of LLM calls per model
Average Response Time Per Call
Average latency per LLM call - lower is faster
google-gemini-2.5
Found 9 of 9 vulnerabilities (100%)
Critical
4
High
5
Medium
0
Low
0
Tokens: 96,858 Cost: $0.2977
Cost per bug: $0.0331
Calls: 12 Total: 13m 60s
Avg per call: 70.0s
google-gemini-2.5-flash
Found 6 of 9 vulnerabilities (67%)
Critical
3
High
3
Medium
0
Low
0
Tokens: 61,719 Cost: $0.0149
Cost per bug: $0.0025
Calls: 8 Total: 2m 31s
Avg per call: 18.8s
openai-gpt-4.1
Found 6 of 9 vulnerabilities (67%)
Critical
3
High
3
Medium
0
Low
0
Tokens: 58,489 Cost: $0.1832
Cost per bug: $0.0305
Calls: 8 Total: 3m 1s
Avg per call: 22.6s
openai-o3
Found 5 of 9 vulnerabilities (56%)
Critical
3
High
2
Medium
0
Low
0
Tokens: 51,581 Cost: $0.7788
Cost per bug: $0.1558
Calls: 7 Total: 3m 12s
Avg per call: 27.4s
claude-opus-4-anthropic
Found 5 of 9 vulnerabilities (56%)
Critical
3
High
2
Medium
0
Low
0
Tokens: 50,197 Cost: $1.3301
Cost per bug: $0.2660
Calls: 7 Total: 3m 37s
Avg per call: 31.1s
claude-sonnet-4-anthropic
Found 5 of 9 vulnerabilities (56%)
Critical
3
High
2
Medium
0
Low
0
Tokens: 49,850 Cost: $0.2561
Cost per bug: $0.0512
Calls: 7 Total: 2m 33s
Avg per call: 21.9s
lmstudio-qwen3
Found 4 of 9 vulnerabilities (44%)
Critical
3
High
1
Medium
0
Low
0
Tokens: 40,753 Cost: $0.0000
Cost per bug: $0.0000
Calls: 6 Total: 7m 54s
Avg per call: 79.1s
openai-gpt-4.1-nano
Found 4 of 9 vulnerabilities (44%)
Critical
3
High
1
Medium
0
Low
0
Tokens: 43,081 Cost: $0.0066
Cost per bug: $0.0017
Calls: 6 Total: 1m 10s
Avg per call: 11.6s
Model Detection Rate by Vulnerability Category
openai-o3
Percentage of vulnerabilities detected by category
claude-opus-4-anthropic
Percentage of vulnerabilities detected by category
claude-sonnet-4-anthropic
Percentage of vulnerabilities detected by category
google-gemini-2.5-flash
Percentage of vulnerabilities detected by category
google-gemini-2.5
Percentage of vulnerabilities detected by category
lmstudio-qwen3
Percentage of vulnerabilities detected by category
openai-gpt-4.1
Percentage of vulnerabilities detected by category
openai-gpt-4.1-nano
Percentage of vulnerabilities detected by category