OpenClaw · agent evaluation

Evaluation observability

A public-safe view of deterministic release-gate evidence and optional live-model observations. The page reads adjacent redacted artifacts and never calls Gmail, a model, or the live EMAIL-13 producer.

pass^k · primary—
pass@k · secondary—
fixed gates—
seeded defects caught—
p50 / p95 logical latency—
total cost—

Optional live-model observations

These measurements are separately budgeted and non-gating. They may compare model variance and open-ended quality, but they never change the deterministic release decision above.

No optional live-trial artifact is published for this build. The deterministic evidence above remains the complete release-gate result.
recorded trials—
models observed—
Unknown outcomes—
budget used—
tokens used—
judge calibration—
Model comparison
ModelTrialsKnown / unknown casespass@kpass^kQualityp50 / p95TokensCost
Judge calibration
DimensionLabelsKnown / unknownAgreement

Live failure-category observations

Redacted trial observations

Trial rows contain bounded status and metric metadata only. Model output, prompt text, tool payloads, state, credentials, and session content are not published.

ModelScenarioTrialStatusSuccessQualityLatencyTokensCostFailures
Version and boundary
Baseline comparison

Reliability and resource trends

Runpass^kpass@kp50p95RetriesTokensCost

Failure-category heatmap

Scenario and family heatmap

FamilyCasesFixed passedSeeded caught
ScenarioFamilyFixedSeededCoverage

Redacted trace timelines

Arguments, results, state payloads, email bodies, credentials, and session text are intentionally absent. Expand a scenario to inspect model → tool → state event metadata.