The benchmark for agent-friendly CLIs
Built for humans.
Ready for agents?
We put developer CLIs in a clean sandbox and let an agent take the wheel. See what works, where it stalls, and the evidence behind every score.
prisma01 / 04
Example command: npm install --global prismaChain of evidence01 Command02 Capture03 FindingEvery claim has a source.
A clean start. A real task. An inspectable result.
THE PUBLIC CATALOG
Find your next CLI.
Compare the interface. Inspect the evidence.| CLI 11 in the catalog | Protocol checks | Agent outcome | Report |
|---|---|---|---|
| auth0saas | 0/10 measured | Not run | ↗ |
| awscloud | 0/10 measured | Not run | ↗ |
| firebasesaas | 0/10 measured | Not run | ↗ |
| ghvcs | 0/10 measured | Not run | ↗ |
| neondata | 0/10 measured | Not run | ↗ |
| prismadata | 0/10 measured | Not run | ↗ |
| sentry-cliobs | 0/10 measured | Not run | ↗ |
| stripesaas | 0/10 measured | Not run | ↗ |
| supabasecloud | 0/10 measured | Not run | ↗ |
| vercelcloud | 0/10 measured | Not run | ↗ |
| wranglercloud | 0/10 measured | Not run | ↗ |
Protocol checks describe sampled commands. Agent outcomes apply to a specific task. These are experimental measurements, not overall CLI ratings.
Read the methodology ↗Don't see a CLI? Request one.
No setup. No login. We’ll review it and run a Floor eval.
We’ll review it and run a Floor eval from the operator console.