← Search

C.V. Jawahar

25 accepted papers

2026

DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios

ICRA 2026poster

Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification and mitigation of potential risks. Although recent Multimodal Large Language Models (MLLMs) have shown promise on general vision–language tasks, our fi…

2026

EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer

CVPR 2026

Real-world object detection must operate in evolving environments where new classes emerge, domains shift, and unseen objects must be identified as unknown--all without accessing prior data. We introduce Evolving World Object Detection (EWOD), a paradigm coupling incremental learning, domain adaptat

Cited by 0SourceScholar
2026

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

ICRA 2026poster

Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags far behind fourwheelers, where multimodal datasets have driven major advances in Advanced Driver Assistance Systems (ADAS). To address this gap, we p…

2026

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction

ICRA 2026poster

Pedestrian intention and trajectory prediction are critical for the safe deployment of autonomous driving systems, directly influencing navigation decisions in complex traffic environments. Recent advances in large vision–language models offer a powerful new paradigm for these tasks by combining hig…

2025

A Dataset for Semantic Segmentation in the Presence of Unknowns

CVPR 2025poster

Before deployment in the real-world deep neural networks require thorough evaluation of how they handle both knowns, inputs represented in the training data, and unknowns (anomalies). This is especially important for scene understanding tasks with safety critical applications, such as in autonomous…

2024

Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach

IROS 2024poster

With the increased importance of autonomous navigation systems has come an increasing need to protect the safety of Vulnerable Road Users (VRUs) such as pedestrians. Predicting pedestrian intent is one such challenging task, where prior work predicts the binary cross/no-cross intention with a fusion…

Cited by 0SourceScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Understanding the Generalization of Pretrained Diffusion Models on Out-of-Distribution Data

AAAI 2024technical

This work tackles the important task of understanding out-of-distribution behavior in two prominent types of generative models, i.e., GANs and Diffusion models. Understanding this behavior is crucial in understanding their broader utility and risks as these systems are increasingly deployed in our d…

2022

My View Is the Best View: Procedure Learning from Egocentric Videos

ECCV 2022poster

"Procedure learning involves identifying the key-steps and determining their logical order to perform a task. Existing approaches commonly use third-person videos for learning the procedure, making the manipulated object small in appearance and often occluded by the actor, leading to significant err…

2022

New Objects on the Road? No Problem, We'll Learn Them Too

IROS 2022poster

Object detection plays an essential role in providing localization, path planning, and decision making capabilities in autonomous navigation systems. However, existing object detection models are trained and tested on a fixed number of known classes. This setting makes the object detection model dif…

Cited by 0SourceScholar
2022

TRoVE: Transforming Road Scene Datasets into Photorealistic Virtual Environments

ECCV 2022poster

"High-quality structured data with rich annotations are critical components in intelligent vehicle systems dealing with road scenes. However, data curation and annotation require intensive investments and yield low-diversity scenarios. The recently growing interest in synthetic data raises questions…

2021

Looking Farther in Parametric Scene Parsing with Ground and Aerial Imagery

ICRA 2021poster

Parametric models that represent layout in terms of scene attributes are an attractive avenue for road scene understanding in autonomous navigation. Prior works that rely only on ground imagery are limited by the narrow field of view of the camera, occlusions and perspective foreshortening. In this…

Cited by 0SourceScholar
2020

Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis

CVPR 2020poster

Humans involuntarily tend to infer parts of the conversation from lip movements when the speech is absent or corrupted by external noise. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate natural speech given only the lip movements of a speaker. Acknowledging t…

Cited by 130PDFcodeScholar
2020

Recurrent Image Annotation With Explicit Inter-Label Dependencies

ECCV 2020poster

Inspired by the success of the CNN-RNN framework in the image captioning task, several works have explored this in multi-label image annotation with the hope that the RNN followed by a CNN would encode inter-label dependencies better than using a CNN alone. To do so, for each training sample, the ea…

2020

RoadText-1K: Text Detection & Recognition Dataset for Driving Videos

ICRA 2020poster

Perceiving text is crucial to understand semantics of outdoor scenes and hence is a critical requirement to build intelligent systems for driver assistance and self-driving. Most of the existing datasets for text detection and recognition comprise still images and are mostly compiled keeping text in…

Cited by 62SourceScholar
2020

Weakly Supervised Instance Segmentation by Learning Annotation Consistent Instances

ECCV 2020poster

Recent approaches for weakly supervised instance segmentations depend on two components: (i) a pseudo label generation model that provides instances which are consistent with a given annotation; and (ii) an instance segmentation model, which is trained in a supervised manner using the pseudo labels…

Cited by 78SourcePDFScholar
2019

Improved Road Connectivity by Joint Learning of Orientation and Segmentation

CVPR 2019poster

Road network extraction from satellite images often produce fragmented road segments leading to road maps unfit for real applications. Pixel-wise classification fails to predict topologically correct and connected road masks due to the absence of connectivity supervision and difficulty in enforcing…

Cited by 237PDFScholar
2019

Scene Text Visual Question Answering

ICCV 2019poster

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visu…

Cited by 417PDFScholar
2018

Efficient Optimization for Rank-Based Loss Functions

CVPR 2018poster

The accuracy of information retrieval systems is often measured using complex loss functions such as the average precision (AP) or the normalized discounted cumulative gain (NDCG). Given a set of positive and negative samples, the parameters of a retrieval system can be estimated by minimizing these…

Cited by 62SourcePDFScholar