Payments
stripe
Planned pin 1.50.11. Not a score.
Workflow AX
Not measured
End-to-end workflows have not been measured.
Fix first
Ranked fails with the retained evidence and the next action. Unmeasured stays unmeasured.
No validated workflow findings are available. This is not a 0/3 verdict. Unmeasured behavior is not a pass or a failure.
Historical diagnostics
Historical readiness and cold-task trials stay separate from workflow results. They are not averaged.
Manage test products and prices
Listed in Payments. This report is not ranked against other CLIs.
- Score
- Not measured
- Verdict
- End-to-end workflows have not been measured.
- Pin
- Planned 1.50.11
- Model
- Not measured
- Date
- N/A
End-to-end workflows have not been measured.
Measured version: Not measured · Planned pin 1.50.11
Weakest dimension: Structured I/O, 30. Historical readiness and cold-task trials stay separate from workflow results. They are not averaged.
Workflow benchmark status
End-to-end workflows have not been measured.
Workflow campaign not published. The Historical diagnostics fold describes different tests and cannot stand in for an end-to-end workflow result.
Manage test products and prices
In test mode, locate the planted product, create a follow-up product and price, update its metadata, and recover from invalid price input. Repeating the request must not create duplicate objects.
Limitation: Test mode only. Does not cover Connect, Billing portals, live keys, or webhook endpoints. Live-mode keys are rejected before setup.
- Normal workflowCampaign pending
- Ambiguous or larger stateCampaign pending
- Controlled recoveryCampaign pending
The planted fault belongs to Controlled recovery only. Normal and ambiguous journeys start from a healthy fixture.
Maintainers can use the historical diagnostic’s weakest dimension as a lead. A workflow finding is published only after a complete retained journey and independently checked outcome.
Scope and reproduction
The task covers discover and prepare, read and select, preview and change, verify the result, replay safely, recover and clean up. Target-service operations must go through the evaluated CLI; a separate verifier checks state.
40 model turns · 200K total tokens · 15 minutes per journey · fresh state per trial. Shell and file tools are available inside isolation. Environment errors stay separate from task failures.