Interactive note · ICCV 2025

PRISM: Reducing Spurious Implicit Biases in Vision-Language Models

[Paper, arXiv, Code, BibTeX]
Problem

CLIP takes shortcuts. Ask it for a waterbird and it partly checks for water, so a waterbird on land gets missed.

Insight

The same shortcut lives in CLIP's text side, and an LLM can spell it out in plain sentences.

Fix

Learn one linear projection from those sentences alone. No images, no group labels. Worst-group accuracy on Waterbirds goes from 36.4% to 84.2%.

1 · The setting

A classifier that looks at the background

CLIP classifies an image by comparing its embedding with the embeddings of text prompts such as a photo of a landbird and a photo of a waterbird, and picking the closest. Nothing tells it which parts of the image should matter.

Waterbirds are usually photographed on water and landbirds on land, so the prompt a photo of a waterbird ends up pointing partly toward water. The benchmark measures what happens to the rare cases: a waterbird on land, a landbird on water. The score that matters is worst-group accuracy, the accuracy on the group the model handles worst.

The figure is a small model of CLIP's embedding space. Its horizontal axis is what we want, the bird; the vertical axis is what we don't, the background. Slide the prompts' lean toward the background and watch the boundary tilt.

1.0
0 means the prompts only describe the bird.
Embedding spacecolour: bird, fill: background
Accuracy by groupthe lowest bar is the worst group
Zero-shot CLIP, in miniature. Each dot is an image embedding. Shaded regions show which prompt each point is closest to, and the diamonds mark the two prompts. When the prompts lean on background, the boundary tilts and birds on the “wrong” background fall on the wrong side. Real CLIP has hundreds of dimensions; this toy keeps the two that tell the story.
2 · Stage 1

An LLM names the shortcut

To remove a bias you usually need to know what it is, and to have images labelled with it. PRISM needs neither. It gives an LLM the class names and asks which spurious attributes a CLIP classifier might latch onto. Then it asks for scene descriptions that mix every class with every attribute.

This works because LLMs learn which words co-occur. If waterbird and lake appear together often, the LLM knows, whether or not the link is causal. That is exactly the knowledge needed to name a shortcut.

Prompt

Provide a list of potential bias attributes associated with the following zero-shot classification using CLIP: a photo of a landbird, a photo of a waterbird.

Attributes found
Scene descriptions, one per group
Stage 1, illustrated. Pick an attribute to swap it into the template. Every class appears with every background, so the bird and the background no longer travel together. These sentences are the entire training set for stage 2. The prompt is the one from the paper's appendix; the outputs shown here are illustrative.

The paper also checks that this is the right place to look. Classify the scene descriptions themselves with CLIP's text prompts and the bias appears in text alone: worst-group accuracy is 33.1% on descriptions versus 38.3% on images. Since CLIP aligns the two modalities, a bias found in text is a bias in images too.

3 · Stage 2

Learn a projection from sentences, fix the images

PRISM learns a linear projection \(P\) of the embedding space. It is trained with the Latent-space Debiasing loss on the text embeddings of the scene descriptions, \(\mathcal{T}_{a,y}\) for attribute \(a\) and class \(y\):

\[\mathcal{L}_{\mathrm{LD}} = \underbrace{\textstyle\sum_{y,\;a\neq a'} \big(1 - \langle P\mathcal{T}_{a,y}, P\mathcal{T}_{a',y}\rangle\big)}_{\text{same bird, different background: pull together}} \;+\; \underbrace{\textstyle\sum_{a,\;y\neq y'} \max\{0,\ \langle P\mathcal{T}_{a,y}, P\mathcal{T}_{a,y'}\rangle - m\}}_{\text{same background, different bird: push apart}}\]

Then CLIP classifies as before, only in the projected space. The encoders are never fine-tuned and no image is used to learn \(P\). Press Train PRISM and watch what happens to the images.

Embedding space after \(P\)zero-shot
Worst-group accuracyon images P never saw
31%
LD loss on the sentenceswhat P is trained on
Stage 2, trained live in your browser. Gradient descent on \(\mathcal{L}_{\mathrm{LD}}\) uses only the sentence embeddings. As \(P\) learns that background words should not change meaning, it squeezes out the background axis: same-bird clusters merge, the boundary stands upright, and the worst group recovers. PRISM-mini skips training and projects out the embeddings of the attribute words directly: the purple squares mark land and water, and every point slides along the dashed line between them. It is cheaper, but in this toy the line is tilted, because the words land and water also lean toward the birds, so it removes some of the bird signal too.

The two terms of the loss do different jobs. The first makes the projection blind to background: a landbird in a forest and a landbird on a beach should look the same. The second stops it from going blind to everything: a landbird and a waterbird on the same beach must stay apart, by at least the margin \(m\). The paper finds \(m = 0.6\) works best on both benchmarks.

4 · Experiments

What the experiments show

The same picture in real CLIP

Waterbirds embeddings, CLIP ViT-L/14

CLIP embeddings: the four groups are mixed into two background-driven clusters
CLIP
CLIP with PRISM: the groups form more structured, separated clusters
CLIP with PRISM

From the paper: each colour is one (bird, background) group. Without PRISM the groups are mixed, a sign that the representation is organised by background. With PRISM the clusters are more structured and separated.

Data-free debiasing

Best worst-group accuracy without any images

CLIP ViT-L/14, zero-shot. WG is worst-group accuracy, Acc is average accuracy. Among methods that use no images, PRISM is best on both benchmarks, and on Waterbirds it also beats methods that do use images.

Method Waterbirds WG Acc CelebA WG Acc
Zero-shot CLIP 36.4 89.3 72.8 87.6
Orth-Cali 68.8 84.5 76.1 86.2
RoboShot 45.2 79.2 82.6 85.5
PRISM-mini 69.5 92.6 82.6 84.4
PRISM 84.2 93.6 84.0 86.9
FairerCLIP uses images 78.1 85.1 86.1 88.0
The LLM matters

The choice of LLM matters

CelebA worst-group accuracy when different LLMs write the scene descriptions.

Cheap to run

One matrix, no fine-tuning

The only thing learned is \(P\). CLIP's encoders stay frozen, so the debiased model remains a general zero-shot classifier, and PRISM-mini needs no optimisation at all. The main limitation is that the quality of the bias list depends on the LLM, and a linear projection may not capture highly non-linear biases.

Cite

BibTeX

@inproceedings{molahasani2025prism,
  title     = {PRISM: Reducing Spurious Implicit Biases in
               Vision-Language Models with LLM-Guided
               Embedding Projection},
  author    = {Molahasani, Mahdiyar and Motamedi, Azadeh and
               Greenspan, Michael and Kim, Il-Min and Etemad, Ali},
  booktitle = {Proceedings of the IEEE/CVF International
               Conference on Computer Vision (ICCV)},
  year      = {2025}
}

* Equal contribution. The models in §1 and §3 are a three-dimensional toy of CLIP's embedding space, built and trained in your browser. The numbers in §4 are from the paper.

← All writing