Help improve TutorBench
TutorBench is a developing benchmark for evaluating how well AI systems teach, not only whether they produce correct answers. Some tutoring behaviors are difficult to validate with deterministic checks or an LLM Judge alone, so future Community Review will add structured human evaluation.
Why human review
Teaching quality is more than a correct answer.
TutorBench looks at whether a Tutor understands a learner's difficulty, offers a useful next step, and adapts its help to the learner's state. These behaviors depend on context and judgment, so deterministic rules or an LLM Judge alone cannot reliably cover all of them.
Early human review will help improve and validate the evaluation method. It is not automatically treated as a gold answer or Human Reference.
What participants may do
A structured review, when the program opens.
Future reviewers would work from a clear task and shared criteria. The exact materials and schedule will be announced later.
- 01Read an AI Tutor response and the evaluation task
- 02Judge the response using the provided criteria
- 03Submit a structured review
- 04Complete a short qualification step before reviewing real assignments
Future application
What we expect to ask when applications open
The first application contract is intentionally small: one contact email, a preferred review language, a short motivation, optional relevant experience, and a coarse availability category.
Applications are not open yet. This section describes a future contract, not a form.
- 01One contact email for a future invitation
- 02Preferred review language
- 03A short motivation for participating
- 04Optional relevant experience summary
- 05Approximate availability category
How it could work
A high-level path from interest to blind review.
Participation would be invite-only at first. An application would not create a reviewer account automatically.
- 01Application
- 02Manual review
- 03Invitation
- 04Consent
- 05Qualification
- 06Blind review
Current status
Participation is not open yet.
The reviewer workflow is invite-only infrastructure under preparation. Public reviewer intake is not open, the real Community Review campaign has not started, and P5 human calibration has not started.
Open
Not open
Not started
Not started
Evidence boundary
Human review informs the method; it does not create a gold standard.
Early human review will be used to improve and validate the evaluation method. Agreement is consistency evidence, not correctness; qualification is eligibility, not calibration.
Future announcements
Follow the benchmark for the next opening.
When participation opens, announcements will appear on the project homepage and GitHub repository. This page has no application form, waitlist, or reviewer login.