← Search

Vignesh Ramanathan

19 accepted papers

2026

Prepare for Warp Speed: Sub-Millisecond Visual Place Recognition Using Event Cameras

ICRA 2026poster

Visual Place Recognition (VPR) enables systems to identify previously visited locations within a map, a fundamental task for autonomous navigation. Prior works have developed VPR solutions using event cameras, which asynchronously measure per-pixel brightness changes with microsecond temporal resolu…

2024

Context Diffusion: In-Context Aware Image Generation

ECCV 2024poster

"We propose Context Diffusion, a diffusion-based framework that enables image generation models to learn from visual examples presented in context. Recent work tackles such in-context learning for image generation, where a query image is provided alongside context examples and text prompts. However,…

Cited by 9SourcePDFScholar
2024

EventMASK: A Frame-Free Rapid Human Instance Segmentation With Event Camera Through Constrained Mask Propagation

RA-L 2024

Human Instance Segmentation (HIS) is essential in robotics for applications such as autonomous driving and human-robot interaction, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">etc</i> . Existing HIS solutions using conventional cameras are comput

Cited by 1SourceScholar
2023

Filtering, Distillation, and Hard Negatives for Vision-Language Pre-Training

CVPR 2023poster

Vision-language models trained with contrastive learning on large-scale noisy data are becoming increasingly popular for zero-shot recognition problems. In this paper we improve the following three aspects of the contrastive pre-training pipeline: dataset noise, model initialization and the training…

2023

PACO: Parts and Attributes of Common Objects

CVPR 2023highlight

Object models are gradually progressing from predicting just category labels to providing detailed descriptions of object instances. This motivates the need for large datasets which go beyond traditional object masks and provide richer annotations such as part masks and attributes. Hence, we introdu…

2023

PartDistillation: Learning Parts From Instance Segmentation

CVPR 2023poster

We present a scalable framework to learn part segmentation from object instance labels. State-of-the-art instance segmentation models contain a surprising amount of part information. However, much of this information is hidden from plain view. For each object instance, the part information is noisy,…

2022

Event-LSTM: An Unsupervised and Asynchronous Learning-Based Representation for Event-Based Data

RA-L 2022

Event cameras are activity-driven bio-inspired vision sensors that respond asynchronously to intensity changes resulting in sparse data known as events. It has potential advantages over conventional cameras, such as high temporal resolution, low latency, and low power consumption. Given the sparse a

Cited by 16SourceScholar
2021

Weakly Supervised Instance Segmentation for Videos With Temporal Mask Consistency

CVPR 2021poster

Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of objects and (b) missing object predictions. We show that thes…

Cited by 31PDFScholar
2019

Activity Driven Weakly Supervised Object Detection

CVPR 2019poster

Weakly supervised object detection aims at reducing the amount of supervision required to train detection models. Such models are traditionally learned from images/videos labelled only with the object class and not the object bounding box. In our work, we try to leverage not only the object class la…

Cited by 40PDFScholar
2018

Exploring the Limits of Weakly Supervised Pretraining

ECCV 2018poster

State-of-the-art visual perception models for a wide range of tasks rely on supervised pretraining. ImageNet classification is the de facto pretraining task for these models. Yet, ImageNet is now nearly ten years old and is by modern standards "small". Even so, relatively little is known about the b…

2018

What Makes a Video a Video: Analyzing Temporal Information in Video Understanding Models and Datasets

CVPR 2018poster

The ability to capture temporal information has been critical to the development of video understanding models. While there have been numerous attempts at modeling motion in videos, an explicit analysis of the effect of temporal information for video understanding is still missing. In this work, we…

Cited by 181SourcePDFScholar
2017

Learning to Learn From Noisy Web Videos

CVPR 2017poster

Understanding the simultaneously very diverse and intricately fine-grained set of possible human actions is a critical open problem in computer vision. Manually labeling training videos is feasible for some action classes but doesn't scale to the full long-tailed distribution of actions. A promising…

Cited by 36PDFScholar
2016

Detecting Events and Key Actors in Multi-Person Videos

CVPR 2016oral

Multi-person event recognition is a challenging task, often with many people active in the scene but only a small subset contributing to an actual event. In this paper, we propose a model which learns to detect events in such videos while automatically "attending" to the people responsible for the e…

Cited by 284PDFScholar
2016

Social LSTM: Human Trajectory Prediction in Crowded Spaces

CVPR 2016spotlight

Humans navigate complex crowded environments based on social conventions: they respect personal space, yielding right-of-way and avoid collisions. In our work, we propose a data-driven approach to learn these human-human interactions for predicting their future trajectories. This is in contrast to t…

Cited by 4143PDFScholar
2015

Learning Semantic Relationships for Better Action Retrieval in Images

CVPR 2015poster

Human actions capture a wide variety of interactions between people and objects. As a result, the set of possible actions is extremely large and it is difficult to obtain sufficient training examples for all actions. However, we could compensate for this sparsity in supervision by leveraging the ric…

Cited by 150SourcePDFScholar