Copy this into your coding agent
Paste into Cursor, Codex, Claude Code, or another coding agent in your project.What the scan checks
The scan does not execute your application or produce a measured quality score.
AI review suggestions cite their evidence and remain suggestions until tested.
An AI review cannot prove that an unexecuted tool works.
An additional Evidence review, when available, checks whether cited sources
support the AI review’s claims. Its percentage is an estimate of evidence
agreement, not your application’s quality score. Missing or oversized evidence
stays unavailable; it does not become a passing check.
Potentially lower model cost
Potentially 86% lower model cost means a candidate model has a lower price for the same input and output volume. It does not mean your production bill has fallen. For example, standard text prices checked on 19 September 2026 are:
Both prices are about 86% lower for GPT-5 Mini. At an illustrative 1,000 input
and 500 output tokens per call, that is about 1.25 per 1,000 calls.
This example assumes identical usage, standard text requests, and no caching,
batch discounts, tools, hosting or negotiated rates. Different models can use
different numbers of tokens, retries or tool calls.
Sources: GPT-5.2 pricing
and GPT-5 Mini pricing.
The finding sidebar identifies both models and links to their price sources.
If input and output prices fall by different percentages, Bench shows a range.
Missing models and stale prices remain unknown, not zero. Quick-scan prices
expire after 30 days unless refreshed.
Turn an opportunity into evidence
- Bench the original setup against system-specific criteria and cases.
- Test supported candidates against the same cases.
- Compare quality, failures and estimated model cost per call.
- Validate the proposed change on independent cases before rollout.
- Use production usage after deployment to verify the real impact.

