Auth
auth0
Planned pin 1.35.0. Not a score.
Workflow AX
Not measured
End-to-end workflows have not been measured.
Fix first
Ranked fails with the retained evidence and the next action. Unmeasured stays unmeasured.
No validated workflow findings are available. This is not a 0/3 verdict. Unmeasured behavior is not a pass or a failure.
Historical diagnostics
Historical readiness and cold-task trials stay separate from workflow results. They are not averaged.
Configure a disposable application
Listed in Auth. This report is not ranked against other CLIs.
- Score
- Not measured
- Verdict
- End-to-end workflows have not been measured.
- Pin
- Planned 1.35.0
- Model
- Not measured
- Date
- N/A
End-to-end workflows have not been measured.
Measured version: Not measured · Planned pin 1.35.0
Weakest dimension: Structured I/O, 30. Historical readiness and cold-task trials stay separate from workflow results. They are not averaged.
Workflow benchmark status
End-to-end workflows have not been measured.
Workflow campaign not published. The Historical diagnostics fold describes different tests and cannot stand in for an end-to-end workflow result.
Configure a disposable application
Create a non-production application in the evaluation tenant, set its callback configuration, independently verify the persisted settings, replay without duplicates, and recover from an invalid callback URL.
Limitation: Does not cover production tenants, user import, or Actions. Management API reads belong to the verifier, not as an agent shortcut for writes.
- Normal workflowCampaign pending
- Ambiguous or larger stateCampaign pending
- Controlled recoveryCampaign pending
The planted fault belongs to Controlled recovery only. Normal and ambiguous journeys start from a healthy fixture.
Maintainers can use the historical diagnostic’s weakest dimension as a lead. A workflow finding is published only after a complete retained journey and independently checked outcome.
Scope and reproduction
The task covers discover and prepare, read and select, preview and change, verify the result, replay safely, recover and clean up. Target-service operations must go through the evaluated CLI; a separate verifier checks state.
40 model turns · 200K total tokens · 15 minutes per journey · fresh state per trial. Shell and file tools are available inside isolation. Environment errors stay separate from task failures.