Correctness
Whether the Tutor stays factually and conceptually correct.
Tutor Benchmark evaluates observable tutoring behavior in structured cases. It is a benchmark foundation, not a claim about a learner’s long-term outcome.
What we measure
Each case can assign atomic rubrics to one primary capability, preserving category-level evidence alongside any future overall score.
Whether the Tutor stays factually and conceptually correct.
Whether it identifies the learner’s actual error, gap, or reasoning issue.
Whether its explanation or hint helps the learner make progress.
Whether it changes its help for the learner’s state and context.
Whether it leaves the learner with a clear, executable next step.
Reproducible generation
The benchmark case supplies visible semantic input. A versioned generation spec pins prompt identity, SHA-256 digest, and output limits. The public baseline uses provider-native sampling and reasoning behavior rather than claiming shared temperature, reasoning budget, or seed controls. Canonical messages are exported into a host-facing execution packet, and the resulting corpus records benchmark-controlled identity separately from host/model execution identity. The same benchmark does not imply identical inference knobs across vendors.
What we do not measure
Those questions require a separate LearningEval or human outcome evaluation program.
Evaluation architecture
The Judge is one evaluator boundary, not ground truth. Deterministic checks are useful proxies and are not a complete measurement of teaching quality.
Calibration status
Calibration infrastructure exists, but real Community Review and human calibration have not started. Judge-vs-human validation and statistical validation are not completed.
Available as provider-independent contracts and synthetic pipeline fixtures.
Not started.
Not completed.
Not completed.
Community Review service
P4 COMMUNITY REVIEW SERVICE — PASS — DEPLOYMENT-READY
L1 PRIVATE STAGING DEPLOYMENT / LAUNCH GATE — PASS
L2-B PRIVATE STAGING AUTHENTICATED HTTP VERIFICATION — PASS
community-review-protocol@0.1.0 remains the provider-independent P3 protocol and defines sealed blind packets, exact atomic submissions, close/freeze semantics, descriptive human-human agreement, and explicit-policy disclosure. Public reviewer intake is NOT OPEN, the real Community Review campaign is NOT STARTED, and P5 calibration is NOT STARTED. Agreement is consistency evidence, not correctness. Qualification is eligibility, not calibration. FROZEN is not Human Reference evidence.