Copy this into your coding agent
Paste into Cursor, Codex, Claude Code, or another coding agent in your project.1
Connect your repository
Connect GitHub and select a repository. Bench scans its code for AI systems,
prompts and tools. You can also begin with SDK events without GitHub.
2
Check your system’s business outcome
Open an AI system’s Understanding page. Add or correct what success means:
for example, “refund an eligible order exactly once” and the relevant refund
policy. Bench uses this context to make evaluation criteria relevant.
3
Confirm your SDK connection
Choose an environment, install the SDK and run a synthetic request. The step
completes after Bench receives an event. Creating a key alone is not evidence
that instrumentation works. Follow the SDK quickstart.
4
Run and inspect a Bench
Select Start benching on a system to compare prompts and models. Inspect
failed cases and the evidence behind recommendations. Use Real app testing
to inspect tests of application behavior and tool effects. See
real app testing.
5
Connect your coding agent, optionally
Connect MCP with sign-in, then ask your agent to list your Bench
systems. The checklist confirms a successful authenticated MCP request. An API
key created for an agent does not by itself confirm MCP connection.

