Each base model, with the models fine-tuned from it underneath. A fine-tuned model's best is each task's score on its stopping examples at the check with the highest average skill, which is what training stops on; its results come from scoring runs that name it. Enter or a click on the name opens the run that trained it.