Evidence before deployment

Test how AI tutors
behave before
you ship them.

Teachometry turns authored learning situations into reproducible Tutor Health evaluations: inspect teaching decisions, diagnose concrete failures, and keep the evidence needed to fix and retest them.

Open source Evidence-first Developer Preview
CASE 03 / 48Upper Elementary · Mathematics

Adding Fractions

I think 1/3 + 1/4 = 2/7. Can you help me check it?
Authored learning objective

Help the student diagnose the fraction-addition mistake, explain the relevant fraction-unit idea, and guide them toward a common denominator without revealing the final answer.

Open the full case Illustrative case walkthrough
Scroll to explore

Evidence for
human-centered AI tutoring.

Five dimensions of tutoring

More than
right or wrong.

Teachometry examines observable tutoring behavior with structured rubrics and transparent evaluation. Each dimension captures a distinct aspect of a response in an authored scenario.

Explore the evaluation method

01

Did it understand the learner?

Diagnosis

Whether it identifies the learner’s actual error, gap, or reasoning issue.

Understands learner thinking

02

Did it help the learner move forward?

Guidance

Whether its explanation or hint helps the learner make progress.

Provides helpful, appropriate hints

03

Does the learner know what to do next?

Actionability

Whether it leaves the learner with a clear, executable next step.

Gives concrete next-step suggestions

04

Is the help actually correct?

Correctness

Whether the Tutor stays factually and conceptually correct.

Maintains factual and conceptual accuracy

05

Did it respond to this learner, not just any learner?

Adaptation

Whether it changes its help for the learner’s state and context.

Adjusts to learner needs and context

Response under evaluation

Measurement infrastructure for AI tutoring

Open data.
Transparent evaluation.
Observable behavior.

Explore the data
Synthetic cases
48
Public development scenarios
Authored rubrics
128
Case-specific evaluation criteria
Current dataset
0.2a.6
tutor-eval-v0.2a
Evaluator version
0.3a.4
Open and reproducible