Trace & eval tools
Know what an agent did after it ran.
- Step traces, latency, cost, quality scores
- Debug failures and regressions in runs
- Framework-agnostic observability
Still missing
- No shared outcome ownership across agents
- No completion gate tied to evidence
- Does not stop unfinished work from shipping