← Search

Yogesh S Rawat

24 accepted papers

2025

Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement

ICCV 2025poster

Clothes-Changing Re-Identification (CC-ReID) aims to recognize individuals across different locations and times, irrespective of clothing. Existing methods often rely on additional models or annotations to learn robust, clothing-invariant features, making them resource-intensive. In contrast, we exp…

2025

Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding

ICLR 2025poster

In this work, we focus on Weakly Supervised Spatio-Temporal Video Grounding (WSTVG). It is a multimodal task aimed at localizing specific subjects spatio-temporally based on textual queries without bounding box supervision. Motivated by recent advancements in multi-modal foundation models for groun…

Cited by 2SourcePDFScholar
2025

DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID

CVPR 2025poster

Clothes-changing person re-identification (CC-ReID) aims to recognize individuals under different clothing scenarios. Current CC-ReID approaches either concentrate on modeling body shape using additional modalities including silhouette, pose, and body mesh, potentially causing the model to overlook…

2025

LR0.FM: Low-Res Benchmark and Improving robustness for Zero-Shot Classification in Foundation Models

ICLR 2025poster

Visual-language foundation Models (FMs) exhibit remarkable zero-shot generalization across diverse tasks, largely attributed to extensive pre-training on largescale datasets. However, their robustness on low-resolution/pixelated (LR) images, a common challenge in real-world scenarios, remains undere…

2025

MolVision: Molecular Property Prediction with Vision Language Models

NeurIPS 2025poster

Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELF…

Cited by 0SourcecodeScholar
2025

STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding

CVPR 2025poster

In this work, we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries and no bounding box supervision. Inspired by recent advances in vision-language foundation models, we investigate their u…

Cited by 1SourcePDFScholar
2024

Navigating Hallucinations for Reasoning of Unintentional Activities

EMNLP 2024finding

In this work we present a novel task of understanding unintentional human activities in videos. We formalize this problem as a reasoning task under zero-shot scenario, where given a video of an unintentional activity we want to know why it transitioned from intentional to unintentional. We first eva…

2023

A Large-Scale Robustness Analysis of Video Action Recognition Models

CVPR 2023poster

We have seen great progress in video action recognition in recent years. There are several models based on convolutional neural network (CNN) and some recent transformer based approaches which provide top performance on existing benchmarks. In this work, we perform a large-scale robustness analysis…

Cited by 33SourcePDFScholar
2023

On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes

NeurIPS 2023poster

This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O- JHMDB consisting of synthetically controlled static/dynamic occlusions, OVIS- UCF and OVIS-JHMDB consisting of occlusions with realistic mot…

2023

Revealing the unseen: Benchmarking video action recognition under occlusion

NeurIPS 2023poster

In this work, we study the effect of occlusion on video action recognition. To facilitate this study, we propose three benchmark datasets and experiment with seven different video action recognition models. These datasets include two synthetic benchmarks, UCF-101-O and K-400-O, which enabled underst…

Cited by 1SourcePDFScholar
2022

Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action Segmentation

NeurIPS 2022accept

We propose Differentiable Temporal Logic (DTL), a model-agnostic framework that introduces temporal constraints to deep networks. DTL treats the outputs of a network as a truth assignment of a temporal logic formula, and computes a temporal logic loss reflecting the consistency between the output an…

Cited by 40SourcePDFScholar
2022

Robustness Analysis of Video-Language Models Against Visual and Language Perturbations

NeurIPS 2022accept

Joint visual and language modeling on large-scale datasets has recently shown good progress in multi-modal tasks when compared to single modal learning. However, robustness of these approaches against real-world perturbations has not been studied. In this work, we perform the first extensive robust…

2021

In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning

ICLR 2021poster

The recent research in semi-supervised learning (SSL) is mostly dominated by consistency regularization based methods which achieve strong performance. However, they heavily rely on domain-specific data augmentations, which are not easy to generate for all data modalities. Pseudo-labeling (PL) is a…

2021

Modeling Multi-Label Action Dependencies for Temporal Action Localization

CVPR 2021poster

Real world videos contain many complex actions with inherent relationships between action classes. In this work, we propose an attention-based architecture that model these action relationships for the task of temporal action localization in untrimmed videos. As opposed to previous works which lever…

Cited by 82PDFcodeScholar
2021

Reformulating Zero-shot Action Recognition for Multi-label Actions

NeurIPS 2021poster

The goal of zero-shot action recognition (ZSAR) is to classify action classes which were not previously seen during training. Traditionally, this is achieved by training a network to map, or regress, visual inputs to a semantic space where a nearest neighbor classifier is used to select the closest…

Cited by 25SourcePDFScholar
2020

A Recurrent Transformer Network for Novel View Action Synthesis

ECCV 2020poster

In this work, we address the problem of synthesizing human actions from novel views. Given an input video of an actor performing some action, we aim to synthesize a video with the same action performed from a novel view with the help of an appearance prior. We propose an end-to-end deep network to s…

2020

Multi-view Action Recognition using Cross-view Video Prediction

ECCV 2020poster

In this work, we address the problem of action recognition in a multi-view environment. Most of the existing approaches utilize pose information for multi-view action recognition. We focus on RGB modality instead and propose an unsupervised representation learning framework, which encodes the scene…

2019

CapsuleVOS: Semi-Supervised Video Object Segmentation Using Capsule Routing

ICCV 2019poster

In this work we propose a capsule-based approach for semi-supervised video object segmentation. Current video object segmentation methods are frame-based and often require optical flow to capture temporal consistency across frames which can be difficult to compute. To this end, we propose a video ba…

Cited by 84PDFcodeScholar