principles.fyi · the brain · concept

representational harm

When a model's outputs reinforce unfair stereotypes about groups of people.

Representational harm is damage done by how a system depicts or describes people — perpetuating stereotypes, demeaning or erasing a group, or systematically associating certain identities with negative traits. With language models it often traces back to biases baked into the pretraining data scraped from the web, which the model then reproduces and can amplify. It is distinct from allocational harm (denying someone a concrete resource like a loan); representational harm is about the messages and associations the outputs carry. Measuring and reducing it is a central part of responsible AI work.

Appears in

Nearby in the brain