Results

Comparisons

QuestionKindVaryingRunsMadeAnswer

Scoring runs

One row per dataset, task, variant, model and the variants its head trained on, from the newest scoring run that covered it. At first each variant shows only its own head; choose "every head" to see heads scored on variants they didn't train on. Scores are on held-out examples, and hovering "score" says what each task's measure is. Skill is how far each got from chance toward the best score possible, 0 at chance and 100 at the best. The head's score is colored by what the run recorded: red when it's no better than chance, yellow when another check fails, and otherwise green as deep as its skill. โš™ marks a fine-tuned model and โš  a failed check, and hovering either says more. Pick values above a column to filter, and click a number's heading to sort.