evaluation-validity · active

Whether the measurement made the finding

Whether a reported result describes the model or the way it was scored — the metric, the prompt format, the contamination, the baseline.

Tags: evaluation

A finding can be an artifact of how it was measured. A discontinuous scoring rule turns smooth improvement into an apparent jump; a baseline given a smaller budget loses; a benchmark in the training data flatters everyone. This capability collects claims about when a measurement stops reflecting the thing it names. It sits oddly beside the others because it is not a property of a model, but every claim in this index rests on some measurement, so it is the one capability whose failures propagate everywhere.

What counts as this capability

Scope boundary used when deciding whether a paper is really about this capability, rather than merely mentioning it.

In scope when the paper's finding is about the measurement rather than the model: showing a result changes or disappears under a different metric, prompt format, baseline budget, or decontaminated split. NOT in scope: a benchmark paper reporting how models score, which belongs to whichever capability it tests; and not general "we need better evals" position papers with no demonstration.

Claims

Techniques

Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.