Applications not open yetParticipate

Help improve TutorBench

TutorBench is a developing benchmark for evaluating how well AI systems teach, not only whether they produce correct answers. Some tutoring behaviors are difficult to validate with deterministic checks or an LLM Judge alone, so future Community Review will add structured human evaluation.

Applications not open yet

Why human review

Teaching quality is more than a correct answer.

TutorBench looks at whether a Tutor understands a learner's difficulty, offers a useful next step, and adapts its help to the learner's state. These behaviors depend on context and judgment, so deterministic rules or an LLM Judge alone cannot reliably cover all of them.

Early human review will help improve and validate the evaluation method. It is not automatically treated as a gold answer or Human Reference.

What participants may do

A structured review, when the program opens.

Future reviewers would work from a clear task and shared criteria. The exact materials and schedule will be announced later.

  • 01Read an AI Tutor response and the evaluation task
  • 02Judge the response using the provided criteria
  • 03Submit a structured review
  • 04Complete a short qualification step before reviewing real assignments

Future application

What we expect to ask when applications open

The first application contract is intentionally small: one contact email, a preferred review language, a short motivation, optional relevant experience, and a coarse availability category.

Applications are not open yet. This section describes a future contract, not a form.

  • 01One contact email for a future invitation
  • 02Preferred review language
  • 03A short motivation for participating
  • 04Optional relevant experience summary
  • 05Approximate availability category

How it could work

A high-level path from interest to blind review.

Participation would be invite-only at first. An application would not create a reviewer account automatically.

  1. 01Application
  2. 02Manual review
  3. 03Invitation
  4. 04Consent
  5. 05Qualification
  6. 06Blind review

Current status

Participation is not open yet.

The reviewer workflow is invite-only infrastructure under preparation. Public reviewer intake is not open, the real Community Review campaign has not started, and P5 human calibration has not started.

Public information

Open

Applications and reviewer intake

Not open

Real Community Review campaign

Not started

P5 human calibration

Not started

Evidence boundary

Human review informs the method; it does not create a gold standard.

Early human review will be used to improve and validate the evaluation method. Agreement is consistency evidence, not correctness; qualification is eligibility, not calibration.

Future announcements

Follow the benchmark for the next opening.

When participation opens, announcements will appear on the project homepage and GitHub repository. This page has no application form, waitlist, or reviewer login.

Follow the repository on GitHub ↗