← Search

Dina Katabi

33 accepted papers

2026

Physiology as Language: Translating Nocturnal Breathing to EEG

ICML 2026poster

This paper introduces a novel cross-physiology translation task: synthesizing sleep electroencephalography (EEG) from respiration signals. To address the significant complexity gap between the two modalities, we propose a waveform-conditional generative framework that preserves fine-grained respirat…

Cited by 0SourceScholar
2025

Language-Guided Image Tokenization for Generation

CVPR 2025poster

Image tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generation. However, mainstream image tokenization methods generally have limited compression rates, making high-resolution image…

Cited by 7SourcePDFScholar
2025

RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning

NeurIPS 2025poster

Reinforcement learning (RL) has recently emerged as a compelling approach for enhancing the reasoning capabilities of large language models (LLMs), where an LLM generator serves as a policy guided by a verifier (reward model). However, current RL post-training methods for LLMs typically use verifier…

Cited by 0SourcecodeScholar
2025

Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity

NeurIPS 2025poster

Knowledge Distillation (KD) aims to train a lightweight student model by transferring knowledge from a large, high-capacity teacher. Recent studies have shown that leveraging diverse teacher perspectives can significantly improve distillation performance; however, achieving such diversity typically…

Cited by 0SourceScholar
2024

Learning Vision from Models Rivals Learning Vision from Data

CVPR 2024poster

We introduce SynCLR a novel approach for learning visual representations exclusively from synthetic images without any real data. We synthesize a large dataset of image captions using LLMs then use an off-the-shelf text-to-image model to generate multiple images corresponding to each synthetic capti…

2024

Leveraging Unpaired Data for Vision-Language Generative Models via Cycle Consistency

ICLR 2024spotlight

Current vision-language generative models rely on expansive corpora of $\textit{paired}$ image-text data to attain optimal performance and generalization capabilities. However, automatically collecting such data (e.g. via large-scale web scraping) leads to low quality and poor image-text correlation…

2024

Return of Unconditional Generation: A Self-supervised Representation Generation Method

NeurIPS 2024oral

Unconditional generation -- the problem of modeling data distribution without relying on human-annotated labels -- is a long-standing and fundamental challenge in generative models, creating a potential of learning from large-scale unlabeled data. In the literature, the generation quality of an unco…

2024

Scaling Laws of Synthetic Images for Model Training ... for Now

CVPR 2024poster

Recent significant advances in text-to-image models unlock the possibility of training vision systems using synthetic images potentially overcoming the difficulty of collecting curated data at scale. It is unclear however how these models behave at scale as more synthetic data is added to the traini…

2023

Change is Hard: A Closer Look at Subpopulation Shift

ICML 2023poster

Machine learning models often perform poorly on subgroups that are underrepresented in the training data. Yet, little is understood on the variation in mechanisms that cause subpopulation shifts, and how algorithms generalize across such diverse shifts at scale. In this work, we provide a fine-grain…

2023

Improving CLIP Training with Language Rewrites

NeurIPS 2023poster

Contrastive Language-Image Pre-training (CLIP) stands as one of the most effective and scalable methods for training transferable vision models using paired image and text data. CLIP models are trained using contrastive loss, which typically relies on data augmentations to prevent overfitting and sh…

2023

MAGE: MAsked Generative Encoder To Unify Representation Learning and Image Synthesis

CVPR 2023poster

Generative modeling and representation learning are two key tasks in computer vision. However, these models are typically trained independently, which ignores the potential for each task to help the other, and leads to training and model maintenance overheads. In this work, we propose MAsked Generat…

2023

Rank-N-Contrast: Learning Continuous Representations for Regression

NeurIPS 2023spotlight

Deep regression models typically learn in an end-to-end fashion without explicitly emphasizing a regression-aware representation. Consequently, the learned representations exhibit fragmentation and fail to capture the continuous nature of sample orders, inducing suboptimal results across a wide rang…

2023

SimPer: Simple Self-Supervised Learning of Periodic Targets

ICLR 2023top-5%

From human physiology to environmental evolution, important processes in nature often exhibit meaningful and strong periodic or quasi-periodic changes. Due to their inherent label scarcity, learning useful representations for periodic tasks with limited or no supervision is of great benefit. Yet, ex…

2023

Unsupervised Object Localization with Representer Point Selection

ICCV 2023poster

We propose a novel unsupervised object localization method that allows us to explain the predictions of the model by utilizing self-supervised pre-trained models without additional finetuning. Existing unsupervised and self-supervised object localization methods often utilize class-agnostic activati…

Cited by 4PDFcodeScholar
2022

"On Multi-Domain Long-Tailed Recognition, Imbalanced Domain Generalization and Beyond"

ECCV 2022poster

"Real-world data often exhibit imbalanced label distributions. Existing studies on data imbalance focus on single-domain settings, i.e., samples are from the same data distribution. However, natural data can originate from distinct domains, where a minority class in one domain could have abundant in…

2022

Targeted Supervised Contrastive Learning for Long-Tailed Recognition

CVPR 2022poster

Real-world data often exhibits long tail distributions with heavy class imbalance, where the majority classes can dominate the training process and alter the decision boundaries of the minority classes. Recently, researchers have investigated the potential of supervised contrastive learning for long…

Cited by 249PDFcodeScholar
2022

Unsupervised Domain Generalization by Learning a Bridge Across Domains

CVPR 2022oral

The ability to generalize learned representations across significantly different visual domains, such as between real photos, clipart, paintings, and sketches, is a fundamental capacity of the human visual system. In this paper, different from most cross-domain works that utilize some (or full) sour…

Cited by 49PDFcodeScholar
2020

Harnessing Structures for Value-Based Planning and Reinforcement Learning

ICLR 2020talk

Value-based methods constitute a fundamental methodology in planning and deep reinforcement learning (RL). In this paper, we propose to exploit the underlying structures of the state-action value function, i.e., Q function, for both planning and deep RL. In particular, if the underlying system dynam…

Cited by 40SourcecodeScholar
2020

Learning Compositional Koopman Operators for Model-Based Control

ICLR 2020spotlight

Finding an embedding space for a linear approximation of a nonlinear dynamical system enables efficient system identification and control synthesis. The Koopman operator theory lays the foundation for identifying the nonlinear-to-linear coordinate transformations with data-driven methods. Recently,…

Cited by 152SourceScholar
2020

Learning Longterm Representations for Person Re-Identification Using Radio Signals

CVPR 2020poster

Person Re-Identification (ReID) aims to recognize a person-of-interest across different places and times. Existing ReID methods rely on images or videos collected using RGB cameras. They extract appearance features like clothes, shoes, hair, etc. Such features, however, can change drastically from o…

Cited by 124PDFScholar
2019

ME-Net: Towards Effective Adversarial Robustness with Matrix Estimation

ICML 2019oral

Deep neural networks are vulnerable to adversarial attacks. The literature is rich with algorithms that can easily craft successful adversarial examples. In contrast, the performance of defense techniques still lags behind. This paper proposes ME-Net, a defense method that leverages matrix estimatio…

2019

Making the Invisible Visible: Action Recognition Through Walls and Occlusions

ICCV 2019poster

Understanding people's actions and interactions typically depends on seeing them. Automating the process of action recognition from visual data has been the topic of much research in the computer vision community. But what if it is too dark, or if the person is occluded or behind a wall? In this pap…

Cited by 170PDFScholar
2019

Through-Wall Human Mesh Recovery Using Radio Signals

ICCV 2019poster

This paper presents RF-Avatar, a neural network model that can estimate 3D meshes of the human body in the presence of occlusions, baggy clothes, and bad lighting conditions. We leverage that radio frequency (RF) signals in the WiFi range traverse clothes and occlusions and bounce off the human body…

Cited by 123PDFScholar
2018

Through-Wall Human Pose Estimation Using Radio Signals

CVPR 2018poster

This paper demonstrates accurate human pose estimation through walls and occlusions. We leverage the fact that wireless signals in the WiFi frequencies traverse walls and reflect off the human body. We introduce a deep neural network approach that parses such radio signals to estimate 2D poses. Sinc…

Cited by 731SourcePDFScholar
2017

Learning Sleep Stages from Radio Signals: A Conditional Adversarial Architecture

ICML 2017poster

We focus on predicting sleep stages from radio measurements without any attached sensors on subjects. We introduce a new predictive model that combines convolutional and recurrent neural networks to extract sleep-specific subject-invariant features from RF signals and capture the temporal progressio…

Cited by 325SourcePDFScholar
2015

Guaranteeing Spoof-Resilient Multi-Robot Networks

RSS 2015poster

Multi-robot networks use wireless communication to provide wide-ranging services such as aerial surveillance and unmanned delivery. However, effective coordination between multiple robots requires trust, making them particularly vulnerable to cyber-attacks. Specifically, such networks can be gravely…

Cited by 124SourcePDFScholar