Sweta Mahajan
I am a PhD student in the Computer Vision and Machine Learning group at the Max Planck Institute for Informatics (MPI-INF), Saarbrücken, supervised by Prof. Bernt Schiele and Prof. Alexander Koller. Since November 2023, I have been part of the RTG 2853 Neuroexplicit Models of Language, Vision, and Action. My research focuses on understanding how multimodal foundation models work and making them more reliable.
Before my PhD, I was a research scholar at the Yardi School of Artificial Intelligence, Indian Institute of Technology (IIT) Delhi, working on active learning for computer vision with Prof. Chetan Arora and Prof. Parag Singla. I received my BS-MS dual degree in Mathematics & Statistics from IISER Kolkata.
News
- Aug 2026 One paper, TEVI, accepted at EMNLP (Findings) 2026. See you in Budapest!
- Jun 2026 Attending CVPR conference in Denver. Interested in multimodal learning? Drop me a Hi!
- May 2026 Two-week research visit with Prof. Anna Rohrbach, in the Multimodal Grounded Learning group, Technische Universität (TU) Darmstadt, Germany.
- May 2026 Gave a talk on ‘Understanding and Improving Multimodal Models’ at the Computer Vision and Multimodal Learning Un-Workshop, Tübingen AI Center, Tübingen, Germany.
- Apr 2026 Co-organising the Explainable Computer Vision (eXCV) Workshop at the European Conference on Computer Vision (ECCV) 2026.
- June 2025 Attending the CVPR conference at Nashville.
Show moreShow less
- May 2025 Organising the Explainable Computer Vision (eXCV) Workshop at the International Conference on Computer Vision (ICCV) 2025. Find the previous edition of the workshop here.
- Nov 2024 Invited talk at the DWS Colloquium, University of Mannheim, Germany.
- Oct 2024 Attending the European Conference on Computer Vision (ECCV), Milan, Italy.
- Sept 2024 Attending the German Conference on Pattern Recognition (GCPR), Munich, Germany.
- July 2024 Attending the ICVSS Summer School, Sicily, Italy.
- July 2024 One paper, Discover-then-Name, accepted at ECCV 2024.
- Jan 2023 Attended the Research Week with Google.
Selected Publications
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
DN-CBM inverts the usual concept bottleneck paradigm: instead of pre-selecting concepts for a task, it uses sparse autoencoders to first discover the concepts a CLIP model has learnt, then names them and trains linear probes on top, yielding performant and interpretable classifiers.
Talks
- May 2026 ‘Understanding and Improving Multimodal Models’, Multimodal AI Lab, TU Darmstadt, Germany.
- May 2026 ‘Understanding and Improving Multimodal Models’, Computer Vision and Multimodal Learning Un-Workshop, Tübingen AI Center, Germany.
- Nov 2024 ‘Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery‘, DWS Colloquium, University of Mannheim, Germany.
Academic Service
Co-organizer: eXCV Workshop (ECCV 2026), eXCV Workshop (ICCV 2025)
Reviewer: ICLR 2027, CVPR 2026, ECCV 2026
Workshop Reviewer: XAI4CV (CVPR 2024,2025), eXCV (ECCV 2024), WiCV (CVPR 2025)