← Search

Shruti Vyas

8 accepted papers

2025

LR0.FM: Low-Res Benchmark and Improving robustness for Zero-Shot Classification in Foundation Models

ICLR 2025poster

Visual-language foundation Models (FMs) exhibit remarkable zero-shot generalization across diverse tasks, largely attributed to extensive pre-training on largescale datasets. However, their robustness on low-resolution/pixelated (LR) images, a common challenge in real-world scenarios, remains undere…

2025

MolVision: Molecular Property Prediction with Vision Language Models

NeurIPS 2025poster

Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELF…

Cited by 0SourcecodeScholar
2024

Semi-supervised Active Learning for Video Action Detection

AAAI 2024technical

In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as un- labeled data along with informative sample selection for ac- tion detection. Video action detection requires spatio-te…

2023

A Large-Scale Robustness Analysis of Video Action Recognition Models

CVPR 2023poster

We have seen great progress in video action recognition in recent years. There are several models based on convolutional neural network (CNN) and some recent transformer based approaches which provide top performance on existing benchmarks. In this work, we perform a large-scale robustness analysis…

Cited by 33SourcePDFScholar
2022

Robustness Analysis of Video-Language Models Against Visual and Language Perturbations

NeurIPS 2022accept

Joint visual and language modeling on large-scale datasets has recently shown good progress in multi-modal tasks when compared to single modal learning. However, robustness of these approaches against real-world perturbations has not been studied. In this work, we perform the first extensive robust…

2020

A Recurrent Transformer Network for Novel View Action Synthesis

ECCV 2020poster

In this work, we address the problem of synthesizing human actions from novel views. Given an input video of an actor performing some action, we aim to synthesize a video with the same action performed from a novel view with the help of an appearance prior. We propose an end-to-end deep network to s…

2020

Multi-view Action Recognition using Cross-view Video Prediction

ECCV 2020poster

In this work, we address the problem of action recognition in a multi-view environment. Most of the existing approaches utilize pose information for multi-view action recognition. We focus on RGB modality instead and propose an unsupervised representation learning framework, which encodes the scene…