BM Watch started as internal tooling. We built it because we could not debug our own agents without it, and every customer who saw it asked whether they could have it. It is now generally available.
- Step-level traces across tool calls, retrievals, model versions, latencies and cost.
- Offline eval suites on every prompt change, plus online evals on sampled live traffic.
- Drift and regression alerts wired to the metric you actually care about, not just token counts.
- Replay: re-run any historical decision exactly as it was made.
It works with agents we built and agents you built. If it speaks OpenTelemetry, BM Watch can read it.