DEPARTMENT OF COMPUTATIONAL AND DATA SCIENCES
Ph.D. Thesis Colloquium
Speaker: Ms. Abhipsa Basu
S.R. Number: 06-18-01-10-12-20-2-19109
Title: “Understanding, Measuring, and Mitigating Bias in Visual Recognition and Generation”
Research Supervisor: Prof. Venkatesh Babu
Date & Time : September 30, 2026 (Wednesday),11:30 AM
Venue : #102, CDS Seminar Hall
ABSTRACT
Human beings form their understanding of the world through the people, places, and events they encounter, but these experiences are inherently limited and uneven, causing our perceptions to reflect harmful biases and stereotypes shaped by our surroundings. Artificial intelligence systems, trained on data collected from the world, can inherit and amplify these same biases. In discriminative systems, this can lead models to rely on spurious associations or perform poorly for underrepresented groups; in generative systems, it can result in stereotypical or narrow representations of people, places, and environments. Because the datasets used to train modern AI systems are collected at scale and are difficult to curate or control, dedicated methods are needed to understand, measure, and mitigate such biases. This thesis studies these challenges in visual AI.
The first part investigates and mitigates different forms of bias in discriminative vision systems. For image classification, we address biases arising from imbalanced data distributions and spurious correlations through an adaptive clustering-based margin loss that enables debiasing in the presence of frozen or blackbox feature extractors, and explore diffusion-based synthetic data generation for constructing more balanced training sets. Extending beyond unimodal image classification, we investigate language biases in visual question answering (VQA), where spurious question-answer correlations can cause models to ignore image content. We mitigate these biases through a learnable adaptive margin-loss framework that preserves both in-distribution and out-of-distribution performance.
The second part examines bias in text-to-image (T2I) generation from a geographic perspective. Through a large-scale crowdsourced study, we find that T2I models tend to represent everyday visual concepts using imagery associated with only a limited set of countries when geographic information is unspecified in the prompt, while many other regions remain poorly represented. Explicitly specifying a country substantially improves its representation, suggesting that models can depict diverse regions but do not necessarily choose to do so by default. We further investigate the training data underlying these models by geographically profiling captions from large-scale vision-language datasets and find that a large majority of image-caption pairs can be geographically attributed to only a few countries.
Geographical representation, however, is not only a question of how frequently a country appears, but also of how it is portrayed. We introduce GeoDiv, an interpretable framework for measuring diversity along visual and socioeconomic dimensions, revealing that countries are often represented through a limited range of characteristics. Finally, we investigate the mechanistic origins of these socioeconomic biases in the T2I pipeline and find that they are likely distributed across multiple components rather than arising from a single source.
Together, these works advance the diagnosis and mitigation of bias in modern vision systems, providing methods and frameworks for building more equitable visual AI.
ALL ARE WELCOME



