← Search

Ram Nevatia

24 accepted papers

2024

CaesarNeRF: Calibrated Semantic Representation for Few-Shot Generalizable Neural Rendering

ECCV 2024poster

"Generalizability and few-shot learning are key challenges in Neural Radiance Fields (NeRF), often due to the lack of a holistic understanding in pixel-level rendering. We introduce CaesarNeRF, an end-to-end approach that leverages scene-level CAlibratEd SemAntic Representation along with pixel-leve…

2024

Large Language Models are Good Prompt Learners for Low-Shot Image Classification

CVPR 2024poster

Low-shot image classification where training images are limited or inaccessible has benefited from recent progress on pre-trained vision-language (VL) models with strong generalizability e.g. CLIP. Prompt learning methods built with VL models generate text features from the class names that only hav…

2024

SEAS: ShapE-Aligned Supervision for Person Re-Identification

CVPR 2024poster

We introduce SEAS using ShapE-Aligned Supervision to enhance appearance-based person re-identification. When recognizing an individual's identity existing methods primarily rely on appearance which can be influenced by the background environment due to a lack of body shape awareness. Although some m…

Cited by 8SourcePDFScholar
2022

Self-Supervised Learning for Sentiment Analysis via Image-Text Matching

ICASSP 2022accepted

There is often a resemblance in the sentiment expressed in social media posts (text) and their accompanying images. In this paper, We leverage this sentiment congruence for self-supervised representation learning for sentiment analysis. By teaching the model to pair an image with its corresponding s…

Cited by 0SourceScholar
2021

SimPLE: Similar Pseudo Label Exploitation for Semi-Supervised Classification

CVPR 2021poster

A common classification task situation is where one has a large amount of data available for training, but only a small portion is annotated with class labels. The goal of semi-supervised training, in this context, is to improve classification accuracy by leverage information not only from labeled d…

Cited by 207PDFcodeScholar
2021

Visual Semantic Role Labeling for Video Understanding

CVPR 2021poster

We propose a new framework for understanding and representing related salient events in a video using visual semantic role labeling. We represent videos as a set of related events, wherein each event consists of a verb and multiple entities that fulfill various roles relevant to that event. To study…

Cited by 80PDFcodeScholar
2020

SPAN: Spatial Pyramid Attention Network for Image Manipulation Localization

ECCV 2020poster

Tehchniques for manipulating images are advancing rapidly; while these are helpful for many useful tasks, they also pose a threat to society with their ability to create believable misinformation. We present a novel, Spatial Pyramid Attention Network (SPAN) for detection and localization of multiple…

2019

Activity Driven Weakly Supervised Object Detection

CVPR 2019poster

Weakly supervised object detection aims at reducing the amount of supervision required to train detection models. Such models are traditionally learned from images/videos labelled only with the object class and not the object bounding box. In our work, we try to leverage not only the object class la…

Cited by 40PDFScholar
2019

NOTE-RCNN: NOise Tolerant Ensemble RCNN for Semi-Supervised Object Detection

ICCV 2019poster

The labeling cost of large number of bounding boxes is one of the main challenges for training modern object detectors. To reduce the dependence on expensive bounding box annotations, we propose a new semi-supervised object detection formulation, in which a few seed box level annotations and a large…

Cited by 123PDFScholar
2018

LEGO: Learning Edge With Geometry All at Once by Watching Videos

CVPR 2018poster

Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a “3D as-smooth-as-possible (3D-ASAP)” prior inside the pipeline, which enables joint estimation of edges and 3D scene, yiel…

Cited by 215SourcePDFScholar
2018

Motion-Appearance Co-Memory Networks for Video Question Answering

CVPR 2018poster

Video Question Answering (QA) is an important task in understanding video temporal structure. We observe that there are three unique attributes of video QA compared with image QA: (1) it deals with long sequences of images containing richer information not only in quantity but also in variety; (2) m…

Cited by 310SourcePDFScholar
2017

AMC: Attention guided Multi-modal Correlation Learning for Image Search

CVPR 2017poster

Given a user's query, traditional image search systems rank images according to its relevance to a single modality (e.g., image content or surrounding text). Nowadays, an increasing number of images on the Internet are available with associated meta data in rich modalities (e.g., titles, keywords, t…

Cited by 49PDFcodeScholar
2017

TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals

ICCV 2017poster

We address the problem of Temporal Action Proposal (TAP) generation. This is an important problem, as fast extraction of semantically important (e.g. human actions) segments from untrimmed videos is an important step for large-scale video analysis. To tackle this problem, we propose a novel Temporal…

Cited by 486PDFcodeScholar
2016

ProNet: Learning to Propose Object-Specific Boxes for Cascaded Neural Networks

CVPR 2016poster

This paper aims to classify and locate objects accurately and efficiently, without using bounding box annotations. It is challenging as objects in the wild could appear at arbitrary locations and in different scales. In this paper, we propose a novel classification architecture ProNet based on convo…

Cited by 80PDFScholar