Skip to main content
WP HealthKit
Engine benchmarks

Choose how deep the audit goes

Every audit runs the same 45 deterministic scanners. The tier you pick chooses the AI engines that reason on top — and we show our workings: real plugins, real findings, published evaluation reports.

Standard

DeepSeek v4 Flash
1 tokenper audit
£0 (free tier) — included

Fast and thorough on classic vulnerability classes — nonce checks, escaping, SQLi patterns. The engine that has powered every audit since launch.

Best for: Everyday audits, CI gates, bulk scanning

Advanced

Kimi K3
2 tokensper audit
£9.98 PAYG — included in Pro and above

A reasoning engine that traces flaws across files — authorization-logic bugs, trust-boundary violations, and injection paths that pattern matching misses.

Best for: Tricky codebases, pre-submission checks, logic flaws

Premium

Kimi K3 + Claude Opus 4.7 (cross-validated)
5 tokensper audit
£24.95 PAYG — included in Pro and above

Two top models audit independently. Findings both models catch are promoted to HIGH confidence; single-model claims are marked for review. The fewest false positives money can buy.

Best for: Client-facing reports, wp.org submissions, post-incident review

The evaluation, in numbers

From our 2026-07 evaluation runs on two production plugins (~30 auditable files each) plus a seeded-vulnerability fixture corpus.

MetricStandardAdvancedPremium
Security findings (plugin A)47 — incl. an authorization-logic flaw no other model caught7+7 cross-validated
Security findings (plugin B)4 — incl. an attribute-breakout XSS4+7 cross-validated
Engines completing without parse errors2/3 (a11y parse fail, now fixed via structured outputs)3/33/3
Seeded-vulnerability recall (fixture corpus)5/55/55/5
False positives on known-clean fixture000
Typical cost per audit~£0.03~£0.40~£0.80

Methodology: identical plugin inputs per model, raw engine output (before production false-positive filters), costs at provider list prices. Sample size is two plugins — directionally useful, not statistically definitive. Full evaluation reports are linked from our changelog; the harness (scripts/eval-models.ts) is public in our repo and re-runnable by anyone.

Try all three on your own plugin

Free accounts get unlimited Standard scans plus one full AI audit every month.