visual-grounding · active

Grounding answers in the image

Saying only what the image supports, rather than describing plausible objects that are not there.

Tags: knowledge, behavior

The vision-side analogue of hallucination, kept separate because the mechanisms and mitigations differ: text hallucination is addressed with retrieval and verification, while visual ungroundedness is about whether the model attends to image evidence at all, or leans on language priors about what usually appears in such a scene.

What counts as this capability

Scope boundary used when deciding whether a paper is really about this capability, rather than merely mentioning it.

In scope when the paper measures or mitigates a model asserting content the image does not support — object hallucination, invented attributes or relations, answers driven by language priors rather than pixels. NOT in scope: general recognition accuracy, and not text-only hallucination, which belongs to the hallucination capability.

Claims

Techniques

None yet.

Related: Stating false facts confidently
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.