Models

A transparent catalog
of model evidence
for AI tutoring.

Public model profiles will be derived from versioned benchmark result artifacts. Each profile keeps model identity, execution context, benchmark cohort, and traceable evidence attached.

No calibrated public model runs yet. Profiles are evidence records, not rankings.

Filter models
View

0 public models · Filters become available when public model profiles exist.

Model registry

Model identity and evidence

0 public profiles

No public profiles

No public model profiles yet.

No calibrated public model runs yet.

A profile will appear only when model identity, snapshot/version, benchmark cohort, generation identity, and traceable trial records are available in a public artifact.

View Results status

Future profile contract

What a public model profile must carry

Schema labels show the evidence boundary. Dashes are placeholders, not model results.

Benchmark version
0.1
Registry context only; no model run is implied.
Dataset cohort
tutor-eval-v0.2a@0.2a.6
The cohort a future public artifact must identify.

Identity

Identity is only publishable when it resolves to a versioned artifact.

Model
—
Provider
—
Snapshot / version
—

Evaluation evidence

Scores appear only when valid result records and their evaluation context are public.

Overall
—
Correctness
—
Diagnosis
—
Guidance
—
Adaptation
—
Actionability
—
Critical-failure rate
—
Answer-leakage rate
—
Latency
—
Tokens
—
Cost
—

Traceability

Generation and trial identity keep a profile auditable and comparable.

Dataset cohort
—
Generation spec ID
—
Generation spec version
—
Prompt version
—
Prompt identity
—
Prompt SHA-256
—
Max output tokens
—
Trial runs
—

Future strengths and weaknesses may be derived from stored category metrics, failure rates, and verified result evidence—not from a free-form post-hoc AI summary.

Our approach

Comparable evidence,
not claims.

Future model profiles should be interpreted only within matched benchmark versions, dataset cohorts, generation conditions, and evaluation procedures. Teachometry measures observable tutoring behavior in structured authored cases—not long-term learning gains, retention, transfer, student satisfaction, or general classroom teaching effectiveness.

Read the methodology

Current model-publication status

Evidence before conclusions.

There are no calibrated public model runs or public profiles yet. The registry stays empty until the publication boundary is met.

Public model artifact / schema
AvailableThe public contract is defined; entries remain empty.
Calibrated public model runs
NoneNo calibrated public model runs yet.
Public model profiles
NoneProfiles require identity plus traceable evidence.
Official public rankings
NoneResults / Leaderboard remains the ranking surface.
Human calibration
Not completedNo human reference set is available.
Judge-vs-human validation
Not completedNo validation claim is made.
Statistical validation
Not completedNo statistical validation claim is made.

Looking ahead

Better evidence before more conclusions.

The registry will grow only from public, versioned artifacts that satisfy the project’s publication and traceability boundaries.