Skip to main content
WP HealthKit
Engine benchmarks · published evaluation runs

One engine. Cross-validated.

Every audit runs the same 58 deterministic scanners, then the WPHK Engine reasons on top — security findings cross-validated by two independent models. And we show our workings: real plugins, real findings, published evaluation reports.

WPHK Engine · two-model security cross-validation · benchmarked monthly

Security engine · cross-validationmerging
Kimi K3
Missing capability check on admin action
admin/settings.php:88 · privilege escalation
flagged
DeepSeek v4 Flash
Missing capability check on admin action
admin/settings.php:88 · privilege escalation
flagged
2/2 agree → promoted to HIGH confidencesingle-model claims → marked for review
The engine

One engine. Every audit.

Security engine

Kimi K3 + DeepSeek v4 Flash

Two frontier models audit security independently. Findings both models corroborate are promoted to HIGH confidence; single-model claims are marked for review.

Runs on every audit

Quality & accessibility engines

DeepSeek v4 Flash

Architecture, maintainability, WordPress best practices, documentation, error handling, and WCAG 2.1 AA / EAA accessibility.

Runs on every audit

Theme & performance engines

DeepSeek v4 Flash

A seven-phase theme audit for theme ZIPs, plus optional server-load analysis when the performance engine is selected.

Themes / optional
Open benchmark

Live results live

A fixed corpus of plugins with known-flaw ground truth, run monthly against the frontier models. Click a model to see its per-fixture card; click a column to sort. Methodology and latest results are public via /api/v1/benchmark.

ModelSeeded recallClean FPsCost / auditLatencyLast run

Corpus v1.0: 8 fixtures with known vulnerabilities (SQLi, XSS, CSRF, REST permission gaps, capability checks, hardcoded secrets, LFI, nonce-missing AJAX) + 3 clean fixtures with correct code. Public API: /api/v1/benchmark

Applied, not marketed

Why these models, in this engine

Model composition isn't a marketing decision — it's the benchmark table above, applied. Models earn their slot on recall and false-positive discipline, and they can lose it the same way.

Security engine

Kimi K3 + DeepSeek v4 Flash

Two models audit security independently. Findings both corroborate get HIGH confidence; single-model claims are marked for review. Cross-validation is how the engine kills hallucinations without slowing every audit.

Quality / a11y / theme / performance

DeepSeek v4 Flash

The disciplined workhorse: 100% recall AND zero false positives on clean code, at the lowest cost on the board. When the cheapest model is also the most careful, it runs everything that isn't security.

The roster is earned

No permanent seats

A model that can't beat the incumbent on recall and FP discipline doesn't get promoted — no matter what its marketing says. Every candidate runs this corpus before it touches your audits.

Nobody gets promoted on reputation

When a new frontier model ships we run it through this page's corpus before it touches production — Claude Opus 5 (10 FPs on release week) and Claude Fable 5 (0% recall, cyber filter) are both still on the ladder as proof that nobody gets promoted on reputation.

Real-world suite

Oiko plugins, every model, pinned commits live

The seeded corpus measures planted-bug recall. This is the other half: every model against five production plugins (boost, fields, forms, guard, sessions), pinned to a commit SHA per run. Severity chips are finding counts per model per plugin — the honest picture of how output varies on real code, and what a full audit costs at each plugin size.

ModelFindings by severityTotalLatencyCost
COClaude Opus 5
0 crit1 high4 med4 low
101m 50s$0.51
G4Grok 4.5
0 crit0 high1 med4 low
82m 41s$0.13
COClaude Opus 4.7
0 crit0 high7 med4 low
131m 33s$0.44
KKKimi K3
0 crit0 high0 med3 low
51m 43s$0.19
CFClaude Fable 5
0 crit0 high0 med0 low
014s$0.62
CSClaude Sonnet 5
0 crit0 high1 med3 low
527s$0.14
G5GPT-5.2
0 crit1 high1 med1 low
444s$0.12
G5GPT-5 mini
0 crit0 high1 med1 low
21m 42s$0.03
G3Gemini 3.6 Flash
1 crit0 high1 med1 low
352s$0.10
G3Gemini 3.1 Pro Preview
0 crit1 high0 med1 low
21m 57s$0.26
DVDeepSeek v4 Pro
0 crit0 high0 med1 low
21m 34s$0.04
G3Gemini 3 Flash Preview
0 crit0 high1 med0 low
11m 12s$0.03
G4GLM 4.6 (z.ai)
0 crit0 high3 med3 low
653s$0.03
MMMiniMax M2
0 crit0 high1 med1 low
240s$0.01
QCQwen3 Coder
0 crit0 high0 med2 low
216s$0.01
G2Gemini 2.5 Pro
not in this run
DVDeepSeek v4 Flash
0 crit0 high1 med2 low
31m 22s$0.03

oiko-boost · 24 auditable files · 17 models · pinned a0d76a6

5 production plugins · every model on the ladder · pinned commit SHA per run.

Methodology

What the numbers mean

Every metric above comes from the same monthly run — same corpus, same prompts, raw engine output before our production false-positive filters.

Recall

Of the vulnerabilities we deliberately seeded into the corpus fixtures, what fraction did the model find? 100% means it caught everything — every SQL injection, XSS, open REST route, and hardcoded secret we planted. This is the headline quality number: a model with low recall is auditing blind.

Clean FPs (false positives)

Findings the model reported on fixtures that contain no bugs at all — correct, well-escaped WordPress code. Every FP is a false alarm a developer has to read, triage, and dismiss. High FP counts destroy trust in the tool, so we weight them as heavily as recall. Zero is the only good score here.

Cost (USD)

What the full corpus run cost at provider list prices, input + output tokens. This is what determines whether a model can run every audit or only makes sense as a security cross-validator. Costs on /models are shown in USD because that's how providers bill us.

Latency

Wall-clock seconds for the corpus run. Fast models (under ~10s) work for interactive uploads and CI gates; slow models (60s+) only work when the user is waiting for a deep review anyway. Reasoning models trade latency for recall — that's the roster decision in one number.

Corpus & fixtures

The benchmark runs a fixed set of 11 small plugins: 8 seeded with known vulnerabilities (SQLi, XSS, CSRF, REST permission gaps, capability checks, hardcoded secrets, file includes, nonce-missing AJAX) and 3 written correctly. Because the ground truth is known, we can measure exactly what each model catches and what it invents.

How the roster uses this

Models earn their slot in the WPHK Engine. A model that can't beat the incumbent on recall and FP discipline doesn't get promoted — no matter what its marketing says. When a new frontier model ships, we run it through this corpus and publish the results here before it touches your audits.

Try the engine on your own plugin

Free accounts get unlimited deterministic scans plus one full AI audit. Paid plans run the same WPHK Engine with 50 or 300 AI audits a month.