
research note
What Does the Caption Really Say? Counterfactual Phrase Intervention for Compositional Data Selection in Vision-Language Pretraining
This paper addresses a key limitation in CLIP-style vision-language pretraining data curation, where conventional pair-level filtering based on global image-caption alignment saturates and fails to…










