Skip to main content
Bench connects your prompts, code and business context to repeatable evaluations. Start with a repository or uploaded prompts. Inspect discovered AI systems, run a bench, and review the failed checks and proposed improvements.
Use the setup instructions shown in your Bench account.

Run your first bench

Connect a source and evaluate a prompt.

Connect your coding agent

Use Claude Code, Cursor or Codex.

Define success

Add business rules and examples when useful.

Build a test library

Keep trusted cases and customize criteria.

What a bench does

Bench creates criteria and scenarios, measures the current prompt, then tests candidate prompt and model changes on that suite. You can inspect the cases, scores, explanations and comparison before accepting a recommendation.