OpenClaw · agent evaluation
Evaluation observability
A public-safe view of deterministic release-gate evidence and optional live-model observations. The page reads adjacent redacted artifacts and never calls Gmail, a model, or the live EMAIL-13 producer.
pass^k · primary—
pass@k · secondary—
fixed gates—
seeded defects caught—
p50 / p95 logical latency—
total cost—
Optional live-model observations
These measurements are separately budgeted and non-gating. They may compare model variance and open-ended quality, but they never change the deterministic release decision above.
No optional live-trial artifact is published for this build. The deterministic evidence above remains the complete release-gate result.
recorded trials—
models observed—
Unknown outcomes—
budget used—
tokens used—
judge calibration—
Model comparison
| Model | Trials | Known / unknown cases | pass@k | pass^k | Quality | p50 / p95 | Tokens | Cost |
|---|
Judge calibration
| Dimension | Labels | Known / unknown | Agreement |
|---|
Live failure-category observations
Redacted trial observations
Trial rows contain bounded status and metric metadata only. Model output, prompt text, tool payloads, state, credentials, and session content are not published.
| Model | Scenario | Trial | Status | Success | Quality | Latency | Tokens | Cost | Failures |
|---|
Version and boundary
Baseline comparison
Reliability and resource trends
| Run | pass^k | pass@k | p50 | p95 | Retries | Tokens | Cost |
|---|
Failure-category heatmap
Scenario and family heatmap
| Family | Cases | Fixed passed | Seeded caught |
|---|
| Scenario | Family | Fixed | Seeded | Coverage |
|---|
Redacted trace timelines
Arguments, results, state payloads, email bodies, credentials, and session text are intentionally absent. Expand a scenario to inspect model → tool → state event metadata.