prima·bench
Repo ↗

Sample-level evals · eval bucket

Whole-sample phenotype.

Whole-sample classification and phenotype — grading a model on the sample as a whole.

11 evals implemented · 1 families
01

Overview

The sample bucket grades whole-sample classification and phenotype — one prediction per biological sample. 11 evals span 1 task family, each a (task × dataset-arm) pair drawn straight from the runtime registry.

02

Families

Task families present in this bucket, by eval count.

03

Catalog

Filter and sort every sample eval; expand any card for its IO contract.

11 of 11 evals · 1 family

sample_classification11

  • 9m
  • 9m
  • 9m
  • 9m
  • 9m
  • 9m
  • 9m
  • 9m
  • 9m
  • 9m
  • 9m