Can an LLM reviewer tell a real defect from a convincing false alarm?
WebFixBench is a reproducible benchmark for measuring the reliability of LLM-assisted code review on web-application changes — not just whether a model can produce a patch.
The benchmark
Review quality is more than finding bugs.
A reviewer that flags everything can look impressive on a dataset containing only defects. WebFixBench deliberately measures the failure modes that make review tooling hard to trust.
Detection
True positives and false negatives are scored against explicit per-case labels.
False positives
Four tempting clean controls make indiscriminate vulnerability reporting cost something.
Structural reliability
Malformed responses are counted explicitly and never silently treated as correct.
Confidence
Per-finding and overall confidence are stored so confidence can be compared with correctness.
Cost & latency
Runs capture wall-clock latency and provider-reported usage, with optional user-supplied pricing.
Ecosystem breakdowns
Results can be broken down across PHP, Laravel, WordPress and the supported defect categories.
Current suite
PHP web ecosystem first.
The engine is language-agnostic, but v0.1 makes a deliberately narrow claim. The current suite contains synthetic PHP web changes designed to isolate one defect at a time and keep the ground truth decidable.
Read the dataset documentation ↗Methodology
Ground truth is frozen before evaluation.
Cases are authored or accepted by a human maintainer. Model output is never used as the answer key. Matching is deterministic, versioned and reproducible.
Fixed input. A unified diff plus explicit context defines exactly what the reviewer may assume.
Human label. Expected findings are reviewed before model evaluation.
Deterministic scoring. Findings match by normalized defect type, optionally also by file — no LLM judge.
Raw runs preserved. Paid model outputs can be rescored later without making another API call.
Results
No real-model leaderboard yet.
WebFixBench v0.1 is currently a release candidate with a working evaluation harness and a small labelled dataset. Real-model results have not been published, so the correct result is still not yet measured.
A committed mock-provider report exists only to exercise the pipeline; it is not a model benchmark result.
Reproduce it
Run the whole pipeline locally.
Python 3.9+, no runtime dependencies for the core package. Paid providers are opt-in and require an explicit model ID.
git clone https://github.com/lex127/webfixbench.git
cd webfixbench
pip install -e .
webfixbench run --provider mock --out results/my-run.json
webfixbench evaluate results/my-run.json
webfixbench report results/my-run.json --out results/my-run.md
Roadmap
Small by design. Broader only when the labels can support it.
Make the baseline publishable
Freeze the remaining case, complete academic review, validate live provider adapters and add a second reader.
Increase evidential value
Grow the PHP suite, add carefully sourced real-world cases, multi-finding cases, harder controls and inter-rater agreement.
Expand context and ecosystems
Confidence calibration, context-level experiments, JavaScript/TypeScript and Python web suites.
Limitations
What v0.1 does not prove.
16 cases are small. Per-category numbers rest on only one to four cases.
The cases are synthetic. They isolate textbook patterns; they do not establish real-world review performance.
One maintainer reviewed the frozen labels. Inter-rater agreement is not yet measured.
This is not certification. A good score does not mean a model or tool is safe enough to replace human review.
People
Built for reproducible evaluation.
Open source