← Search

Gemma Roig

15 accepted papers

2026

Temporal Slowness in Central Vision Drives Semantic Object Learning

ICLR 2026poster

Humans acquire semantic object representations from egocentric visual streams with minimal supervision. Importantly, the visual system processes with high resolution only the center of its field of view and learns similar representations for visual inputs occurring close in time. This emphasizes slo…

Cited by 0SourcecodeScholar
2025

BrainACTIV: Identifying visuo-semantic properties driving cortical selectivity using diffusion-based image manipulation

ICLR 2025poster

The human brain efficiently represents visual inputs through specialized neural populations that selectively respond to specific categories. Advancements in generative modeling have enabled data-driven discovery of neural selectivity using brain-optimized image synthesis. However, current methods in…

Cited by 0SourcePDFScholar
2025

Efficient Unsupervised Shortcut Learning Detection and Mitigation in Transformers

ICCV 2025poster

Shortcut learning, i.e., a model's reliance on undesired features not directly relevant to the task, is a major challenge that severely limits the applications of machine learning algorithms, particularly when deploying them to assist in making sensitive decisions, such as in medical diagnostics. In…

2025

One Hundred Neural Networks and Brains Watching Videos: Lessons from Alignment

ICLR 2025poster

What can we learn from comparing video models to human brains, arguably the most efficient and effective video processing systems in existence? Our work takes a step towards answering this question by performing the first large-scale benchmarking of deep video models on representational alignment to…

Cited by 0SourcePDFScholar
2025

UIDAPLE: Unsupervised Incremental Domain Adaptation through Adaptive Prompt Learning

ICASSP 2025accepted

Continual learning poses significant challenges for deep neural networks, notably catastrophic forgetting, particularly when faced with shifting data distributions that compromise previously acquired knowledge. This paper tackles these issues within the Unsupervised Incremental Domain Adaptation (UI…

Cited by 0SourceScholar
2024

Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience

ICML 2024poster

Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated. Moreover, recent critiques raise issues that question its usefulness to advance the broader goals of AI. However, it…

Cited by 4SourcePDFScholar
2023

Analyzing Vision Transformers for Image Classification in Class Embedding Space

NeurIPS 2023poster

Despite the growing use of transformer models in computer vision, a mechanistic understanding of these networks is still needed. This work introduces a method to reverse-engineer Vision Transformers trained to solve image classification tasks. Inspired by previous research in NLP, we demonstrate how…

2022

Tafsir Dataset: A Novel Multi-Task Benchmark for Named Entity Recognition and Topic Modeling in Classical Arabic Literature

COLING 2022main

Various historical languages, which used to be lingua franca of science and arts, deserve the attention of current NLP research. In this work, we take the first data-driven steps towards this research line for Classical Arabic (CA) by addressing named entity recognition (NER) and topic modeling (TM)…

Cited by 4SourcePDFScholar
2022

What Do Navigation Agents Learn About Their Environment?

CVPR 2022poster

Today's state of the art visual navigation agents typically consist of large deep learning architectures trained end to end. Such models offer little to no interpretability about the skills learned by the agent or the actions taken by it in response to its environment. While past works have explored…

Cited by 16PDFcodeScholar
2020

Duality Diagram Similarity: a generic framework for initialization selection in task transfer learning

ECCV 2020poster

In this paper, we tackle an open research question in transfer learning, which is selecting a model initialization to achieve high performance on a new task, given several pre-trained models. We propose a new highly efficient and accurate approach based on duality diagram similarity (DDS) between de…

2020

Multi-Source Open-Set Deep Adversarial Domain Adaptation

ECCV 2020poster

We introduce a novel learning paradigm based on multi-source open-set unsupervised domain adaptation (MS-OSDA). Recently, the notion of single-source open-set domain adaptation (OSDA) has drawn much attention which considers the presence of previously unseen open-set (unknown) classes in the target-…

Cited by 43SourcePDFScholar
2019

JSIS3D: Joint Semantic-Instance Segmentation of 3D Point Clouds With Multi-Task Pointwise Networks and Multi-Value Conditional Random Fields

CVPR 2019oral

Deep learning techniques have become the to-go models for most vision-related tasks on 2D images. However, their power has not been fully realised on several tasks in 3D space, e.g., 3D scene understanding. In this work, we jointly address the problems of semantic and instance segmentation of 3D poi…

Cited by 262PDFcodeScholar