Developer PreviewTransparent scope

Methodology

Tutor Benchmark evaluates observable tutoring behavior in structured cases. It is a benchmark foundation, not a claim about a learner’s long-term outcome.

What we measure

Five capabilities

Each case can assign atomic rubrics to one primary capability, preserving category-level evidence alongside any future overall score.

CorrectnessDiagnosisGuidanceAdaptationActionability

Correctness

Whether the Tutor stays factually and conceptually correct.

Diagnosis

Whether it identifies the learner’s actual error, gap, or reasoning issue.

Guidance

Whether its explanation or hint helps the learner make progress.

Adaptation

Whether it changes its help for the learner’s state and context.

Actionability

Whether it leaves the learner with a clear, executable next step.

Reproducible generation

Case, spec, packet, corpus

The benchmark case supplies visible semantic input. A versioned generation spec pins prompt identity, SHA-256 digest, and output limits. The public baseline uses provider-native sampling and reasoning behavior rather than claiming shared temperature, reasoning budget, or seed controls. Canonical messages are exported into a host-facing execution packet, and the resulting corpus records benchmark-controlled identity separately from host/model execution identity. The same benchmark does not imply identical inference knobs across vendors.

What we do not measure

Do not overread a benchmark score.

  • Actual long-term learning
  • Retention
  • Transfer
  • Student satisfaction
  • Real classroom outcomes

Those questions require a separate LearningEval or human outcome evaluation program.

Evaluation architecture

From response to result

Tutor response
Deterministic evaluators
Semantic Judge
Atomic rubrics
Category aggregation
Benchmark result

The Judge is one evaluator boundary, not ground truth. Deterministic checks are useful proxies and are not a complete measurement of teaching quality.

Calibration status

Infrastructure exists; validation is not yet complete.

Calibration infrastructure exists, but real Community Review and human calibration have not started. Judge-vs-human validation and statistical validation are not completed.

Calibration infrastructure

Available as provider-independent contracts and synthetic pipeline fixtures.

P5 human calibration

Not started.

Judge-vs-human validation

Not completed.

Statistical validation

Not completed.

Community Review service

P4 is deployment-ready; public review is not open.

P4 COMMUNITY REVIEW SERVICE — PASS — DEPLOYMENT-READY
L1 PRIVATE STAGING DEPLOYMENT / LAUNCH GATE — PASS
L2-B PRIVATE STAGING AUTHENTICATED HTTP VERIFICATION — PASS

community-review-protocol@0.1.0 remains the provider-independent P3 protocol and defines sealed blind packets, exact atomic submissions, close/freeze semantics, descriptive human-human agreement, and explicit-policy disclosure. Public reviewer intake is NOT OPEN, the real Community Review campaign is NOT STARTED, and P5 calibration is NOT STARTED. Agreement is consistency evidence, not correctness. Qualification is eligibility, not calibration. FROZEN is not Human Reference evidence.

Read the Community Review protocol ↗