Interactive note · ICCV 2025
PRISM: Reducing Spurious Implicit Biases in Vision-Language Models
[Paper, arXiv, Code, BibTeX]CLIP takes shortcuts. Ask it for a waterbird and it partly checks for water, so a waterbird on land gets missed.
The same shortcut lives in CLIP's text side, and an LLM can spell it out in plain sentences.
Learn one linear projection from those sentences alone. No images, no group labels. Worst-group accuracy on Waterbirds goes from 36.4% to 84.2%.
A classifier that looks at the background
CLIP classifies an image by comparing its embedding with the embeddings of text prompts such as
a photo of a landbird
and a photo of a waterbird
, and picking the closest. Nothing tells it which
parts of the image should matter.
Waterbirds are usually photographed on water and landbirds on land, so the prompt
a photo of a waterbird
ends up pointing partly toward water. The benchmark measures what happens
to the rare cases: a waterbird on land, a landbird on water. The score that matters is
worst-group accuracy, the accuracy on the group the model handles worst.
The figure is a small model of CLIP's embedding space. Its horizontal axis is what we want, the bird; the vertical axis is what we don't, the background. Slide the prompts' lean toward the background and watch the boundary tilt.
An LLM names the shortcut
To remove a bias you usually need to know what it is, and to have images labelled with it. PRISM needs neither. It gives an LLM the class names and asks which spurious attributes a CLIP classifier might latch onto. Then it asks for scene descriptions that mix every class with every attribute.
This works because LLMs learn which words co-occur. If waterbird
and lake
appear together often, the
LLM knows, whether or not the link is causal. That is exactly the knowledge needed to name a shortcut.
Provide a list of potential bias attributes associated with the following zero-shot classification using CLIP:
a photo of a landbird
, a photo of a waterbird
.
The paper also checks that this is the right place to look. Classify the scene descriptions themselves with CLIP's text prompts and the bias appears in text alone: worst-group accuracy is 33.1% on descriptions versus 38.3% on images. Since CLIP aligns the two modalities, a bias found in text is a bias in images too.
Learn a projection from sentences, fix the images
PRISM learns a linear projection \(P\) of the embedding space. It is trained with the Latent-space Debiasing loss on the text embeddings of the scene descriptions, \(\mathcal{T}_{a,y}\) for attribute \(a\) and class \(y\):
Then CLIP classifies as before, only in the projected space. The encoders are never fine-tuned and no image is used to learn \(P\). Press Train PRISM and watch what happens to the images.
landand
water, and every point slides along the dashed line between them. It is cheaper, but in this toy the line is tilted, because the words
landand
wateralso lean toward the birds, so it removes some of the bird signal too.
The two terms of the loss do different jobs. The first makes the projection blind to background: a landbird in a forest and a landbird on a beach should look the same. The second stops it from going blind to everything: a landbird and a waterbird on the same beach must stay apart, by at least the margin \(m\). The paper finds \(m = 0.6\) works best on both benchmarks.
What the experiments show
Waterbirds embeddings, CLIP ViT-L/14
From the paper: each colour is one (bird, background) group. Without PRISM the groups are mixed, a sign that the representation is organised by background. With PRISM the clusters are more structured and separated.
Best worst-group accuracy without any images
CLIP ViT-L/14, zero-shot. WG is worst-group accuracy, Acc is average accuracy. Among methods that use no images, PRISM is best on both benchmarks, and on Waterbirds it also beats methods that do use images.
| Method | Waterbirds WG | Acc | CelebA WG | Acc |
|---|---|---|---|---|
| Zero-shot CLIP | 36.4 | 89.3 | 72.8 | 87.6 |
| Orth-Cali | 68.8 | 84.5 | 76.1 | 86.2 |
| RoboShot | 45.2 | 79.2 | 82.6 | 85.5 |
| PRISM-mini | 69.5 | 92.6 | 82.6 | 84.4 |
| PRISM | 84.2 | 93.6 | 84.0 | 86.9 |
| FairerCLIP uses images | 78.1 | 85.1 | 86.1 | 88.0 |
The choice of LLM matters
CelebA worst-group accuracy when different LLMs write the scene descriptions.
One matrix, no fine-tuning
The only thing learned is \(P\). CLIP's encoders stay frozen, so the debiased model remains a general zero-shot classifier, and PRISM-mini needs no optimisation at all. The main limitation is that the quality of the bias list depends on the LLM, and a linear projection may not capture highly non-linear biases.
BibTeX
@inproceedings{molahasani2025prism,
title = {PRISM: Reducing Spurious Implicit Biases in
Vision-Language Models with LLM-Guided
Embedding Projection},
author = {Molahasani, Mahdiyar and Motamedi, Azadeh and
Greenspan, Michael and Kim, Il-Min and Etemad, Ali},
booktitle = {Proceedings of the IEEE/CVF International
Conference on Computer Vision (ICCV)},
year = {2025}
}
* Equal contribution. The models in §1 and §3 are a three-dimensional toy of CLIP's embedding space, built and trained in your browser. The numbers in §4 are from the paper.