judge-position-bias · active

Biased when judging other outputs

A model used as a judge favors whichever answer is shown first, longer answers, and its own outputs.

Also called: LLM-as-judge bias, position bias, verbosity bias, self-enhancement bias

Tags: evaluation, llm-as-judge

When scoring or ranking candidate answers, a capable judge gives the same verdict regardless of the order of candidates, their length, or which model produced them, and agrees with careful human raters.

Claims

Techniques

Related: Telling the user what they want to hear
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.