Spend by model (all model calls, incl. verify/reflect)
No cost data yet — run some tasks.
Model performance (as the task's acting model)
| model | tasks | success | avg steps | avg $/task | total spend |
|---|
Run the same prompt on two models and compare rows — success rate and retries fall straight out of the event journal.