Labs / AI Evaluation Explorer
AI Evaluation Explorer
prototypev0.1.0
| Model | Latency (ms) | Tokens / $ | Hallucination markers | Formatting adherence |
|---|---|---|---|---|
| claude-sonnet-5 | 820 | 42,000 | 1 | 98% |
| claude-haiku-4-5 | 310 | 96,000 | 3 | 91% |
| claude-opus-5 | 1450 | 18,000 | 0 | 99% |
Static fixture snapshot from fixtures/eval-runs.json — not live model calls.
How it works
Multi-model fixture comparison viewer — latency, token efficiency, hallucination markers, and formatting adherence across static eval runs.
Known limitations & edge cases
- Data is static fixtures in fixtures/eval-runs.json, not live model calls.
- Radar chart is unscaled across dimensions with very different ranges.
- No historical trend view yet — single snapshot per run.
Telemetry & privacy notice
This experiment collects no analytics and stores nothing server-side. Any state you see lives in this browser tab only (memory or localStorage) and clears on reset or reload.
Feedback
Spotted something broken? Open an issue on the asafarim-platform repo with the experiment slug (ai-eval-explorer) in the title.