← Search

Uttaran Bhattacharya

20 accepted papers

2026

GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling

CVPR 2026

Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward modeling, but these models struggle to capture subtle temporal dynamics. We propose a fundamentally different approach: repurposing video generative models

Cited by 0SourceScholar
2025

Evaluation and Incident Prevention in an Enterprise AI Assistant

AAAI 2025technical

Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a comprehensive framework for monitoring, benchmarking, and continuously improving such complex, multi-component systems und…

Cited by 0SourcePDFScholar
2025

HUMOTO: A 4D Dataset of Mocap Human Object Interactions

ICCV 2025poster

We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures interactions with 63 precisely modeled objects and 72 articulated…

2025

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

CVPR 2025poster

We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and t…

Cited by 1SourcePDFScholar
2025

SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction

CVPR 2025poster

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion triplets, has opened new avenues for training, yielding promising r…

2024

DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning

AAAI 2024technical

We present DanceAnyWay, a generative learning method to synthesize beat-guided dances of 3D human characters synchronized with music. Our method learns to disentangle the dance movements at the beat frames from the dance movements at all the remaining frames by operating at two hierarchical levels.…

2024

HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances

CVPR 2024poster

Text-to-image generative models can generate high-quality humans but realism is lost when generating hands. Common artifacts include irregular hand poses shapes incorrect numbers of fingers and physically implausible finger orientations. To generate images with realistic hands we propose a novel dif…

Cited by 27SourcePDFScholar
2024

Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior

ICLR 2024spotlight

Shannon and Weaver's seminal information theory divides communication into three levels: technical, semantic, and effectiveness. While the technical level deals with the accurate reconstruction of transmitted symbols, the semantic and effectiveness levels deal with the inferred meaning and its effec…

2024

TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing

ACL 2024findings

Given a source and its edited version performed based on human instructions in natural language, how do we extract the underlying edit operations, to automatically replicate similar edits on other images? This is the problem of reverse designing, and we present TAME-RD, a model to solve this problem…

Cited by 0SourcePDFScholar
2022

Learning Unseen Emotions from Gestures via Semantically-Conditioned Zero-Shot Perception with Adversarial Autoencoders

AAAI 2022technical

We present a novel generalized zero-shot algorithm to recognize perceived emotions from gestures. Our task is to map gestures to novel emotion categories not encountered in training. We introduce an adversarial autoencoder-based representation learning that correlates 3D motion-captured gesture sequ…

Cited by 17SourcePDFScholar
2021

HighlightMe: Detecting Highlights From Human-Centric Videos

ICCV 2021poster

We present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multiple observable human-centric modalities in the videos, such as poses and faces. We use an autoencoder network equipped wit…

Cited by 10PDFScholar
2020

CMetric: A Driving Behavior Measure using Centrality Functions

IROS 2020poster

We present a new measure, CMetric, to classify driver behaviors using centrality functions. Our formulation combines concepts from computational graph theory and social traffic psychology to quantify and classify the behavior of human drivers. CMetric is used to compute the probability of a vehicle…

Cited by 47SourceScholar
2020

EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's Principle

CVPR 2020poster

We present EmotiCon, a learning-based algorithm for context-aware perceived human emotion recognition from videos and images. Motivated by Frege's Context Principle from psychology, our approach combines three interpretations of context for emotion recognition. Our first interpretation is based on u…

Cited by 177PDFScholar
2020

Forecasting Trajectory and Behavior of Road-Agents Using Spectral Clustering in Graph-LSTMs

RA-L 2020

We present a novel approach for traffic forecasting in urban traffic scenarios using a combination of spectral graph analysis and deep learning. We predict both the low-level information (future trajectories) as well as the high-level information (road-agent behavior) from the extracted trajectory o

Cited by 175SourceScholar
2020

GraphRQI: Classifying Driver Behaviors Using Graph Spectrums

ICRA 2020poster

We present a novel algorithm (GraphRQI) to identify driver behaviors from road-agent trajectories. Our approach assumes that the road-agents exhibit a range of driving traits, such as aggressive or conservative driving. Moreover, these traits affect the trajectories of nearby road-agents as well as…

Cited by 30SourceScholar
2020

RoadTrack: Realtime Tracking of Road Agents in Dense and Heterogeneous Environments

ICRA 2020poster

We present a realtime tracking algorithm, Road-Track, to track heterogeneous road-agents in dense traffic videos. Our approach is designed for dense traffic scenarios that consist of different road-agents such as pedestrians, two-wheelers, cars, buses, etc. sharing the road. We use the tracking-by-d…

Cited by 11SourceScholar
2020

Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping

ECCV 2020poster

We present an autoencoder-based semi-supervised approach to classify perceived human emotions from walking styles obtained from videos or motion-captured data and represented as sequences of 3D poses. Given the motion on each joint in the pose at each time step extracted from 3D pose sequences, we h…

Cited by 62SourcePDFScholar
2019

DensePeds: Pedestrian Tracking in Dense Crowds Using Front-RVO and Sparse Features

IROS 2019poster

We present a pedestrian tracking algorithm, DensePeds, that tracks individuals in highly dense crowds (>2 pedestrians per square meter). Our approach is designed for videos captured from front-facing or elevated cameras. We present a new motion model called Front-RVO (FRVO) for predicting pedestrian…

Cited by 22SourceScholar
2019

TraPHic: Trajectory Prediction in Dense and Heterogeneous Traffic Using Weighted Interactions

CVPR 2019poster

We present a new algorithm for predicting the near-term trajectories of road agents in dense traffic videos. Our approach is designed for heterogeneous traffic, where the road agents may correspond to buses, cars, scooters, bi-cycles, or pedestrians. We model the interactions between different road…

Cited by 347PDFScholar