Leaderboard
A future ranking surface for versioned Tutor capability results. Until the corpus is reproducible and calibrated, this page stays explicit about what is not known.
No calibrated public model runs yet. Public model results are unavailable. Human calibration (P5) has not started, and Judge-vs-human and statistical validation are not completed.
Leaderboard coming soon
No real model results are checked into the public artifact. Synthetic demonstrations are not presented as rankings.
| Rank | Model | Overall | Correctness | Diagnosis | Guidance | Adaptation | Actionability | Status |
|---|---|---|---|---|---|---|---|---|
| No public model rows are available yet. | ||||||||
Tutor capability score
Operational signals
Traceability fields
The future table will show: Correctness, Diagnosis, Guidance, Adaptation, Actionability. Operational fields include CriticalFailureRate, AnswerLeakageRate, LatencyMs, Tokens, Cost. Results from different generation profiles are separate cohorts and are not silently mixed.
Future filters
Designed for comparison without hiding context
The v0.1 shell leaves room for benchmark, prompt, subject, provider, model, and metric sorting without implementing a complex filter engine yet.
- Version
- Benchmark version · prompt version · dataset version
- Scope
- Subject · model provider · model
- Sort
- Overall · correctness · diagnosis · guidance · adaptation · actionability · cost · latency