← Search

Mubbasir Kapadia

23 accepted papers

2026

RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation

CVPR 2026

Success in generative modeling across language, image, and video demonstrates that large, well-curated datasets are the key driver for building capable models. 3D Human motion, however, has lagged behind, constrained by an unsatisfying choice between small, high-fidelity motion capture datasets and

Cited by 0SourceScholar
2025

Cardiverse: Harnessing LLMs for Novel Card Game Prototyping

EMNLP 2025

The prototyping of computer games, particularly card games, requires extensive human effort in creative ideation and gameplay evaluation. Recent advances in Large Language Models (LLMs) offer opportunities to automate and streamline these processes. However, it remains challenging for LLMs to design

2025

Less is More: Improving Motion Diffusion Models with Sparse Keyframes

ICCV 2025poster

Recent advances in motion diffusion models have led to remarkable progress in diverse motion generation tasks, including text-to-motion synthesis.However, existing approaches represent motions as dense frame sequences, requiring the model to process redundant or less informative frames.The processin…

Cited by 0SourcePDFScholar
2025

StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion

ICCV 2025poster

We present StyleMotif, a novel Stylized Motion Latent Diffusion model, generating motion conditioned on both content and style from multiple modalities. Unlike existing approaches that either focus on generating diverse motion content or transferring style from sequences, StyleMotif seamlessly synth…

2024

Learning from Synthetic Human Group Activities

CVPR 2024poster

The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation we introduce M3…

2023

Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation Learning

IJCAI 2023poster

Digraph Representation Learning aims to learn representations for directed homogeneous graphs (digraphs). Prior work is largely constrained or has poor generalizability across tasks. Most Graph Neural Networks exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and pres…

Cited by 0SourcePDFScholar
2023

MSI: Maximize Support-Set Information for Few-Shot Segmentation

ICCV 2023poster

FSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract the information relevant to target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature exci…

Cited by 32PDFcodeScholar
2023

Procedure-Aware Pretraining for Instructional Video Understanding

CVPR 2023poster

Our goal is to learn a video representation that is useful for downstream procedure understanding tasks in instructional videos. Due to the small amount of available annotations, a key challenge in procedure understanding is to be able to extract from unlabeled videos the procedural knowledge such a…

2022

COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality

ECCV 2022poster

"Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by modeling the video as tokens that represent the multi-scale semantic concepts in the video. We propose COMPOSER, a Multi…

2022

Cross-Modal Coherence for Text-to-Image Retrieval

AAAI 2022technical

Common image-text joint understanding techniques presume that images and the associated text can universally be characterized by a single implicit model. However, co-occurring images and text can be related in qualitatively different ways, and explicitly modeling it could improve the performance of…

2022

HM: Hybrid Masking for Few-Shot Segmentation

ECCV 2022poster

"We study few-shot semantic segmentation that aims to segment a target object from a query image when provided with a few annotated support images of the target class. Several recent methods resort to a feature masking (FM) technique to discard irrelevant feature activations which eventually facilit…

2022

Harnessing Fourier Isovists and Geodesic Interaction for Long-Term Crowd Flow Prediction

IJCAI 2022poster

With the rise in popularity of short-term Human Trajectory Prediction (HTP), Long-Term Crowd Flow Prediction (LTCFP) has been proposed to forecast crowd movement in large and complex environments. However, the input representations, models, and datasets for LTCFP are currently limited. To this end,…

2022

MUSE-VAE: Multi-Scale VAE for Environment-Aware Long Term Trajectory Prediction

CVPR 2022poster

Accurate long-term trajectory prediction in complex scenes, where multiple agents (e.g., pedestrians or vehicles) interact with each other and the environment while attempting to accomplish diverse and often unknown goals, is a challenging stochastic forecasting problem. In this work, we propose MUS…

Cited by 90PDFScholar
2022

SEAN 2.0: Formalizing and Generating Social Situations for Robot Navigation

RA-L 2022

We present SEAN 2.0, an open-source system designed to advance social navigation via the training and benchmarking of navigation policies in varied social contexts. A key limitation of current social navigation research is that policies are often trained and evaluated considering only a few social c

Cited by 74SourceScholar
2021

AESOP: Abstract Encoding of Stories, Objects, and Pictures

ICCV 2021poster

Visual storytelling and story comprehension are uniquely human skills that play a central role in how we learn about and experience the world. Despite remarkable progress in recent years in synthesis of visual and textual content in isolation and learning effective joint visual-linguistic representa…

Cited by 19PDFcodeScholar
2021

Hopper: Multi-hop Transformer for Spatiotemporal Reasoning

ICLR 2021poster

This paper considers the problem of spatiotemporal object-centric reasoning in videos. Central to our approach is the notion of object permanence, i.e., the ability to reason about the location of objects as they move through the video while being occluded, contained or carried by other objects. Exi…

2020

HID: Hierarchical Multiscale Representation Learning for Information Diffusion

IJCAI 2020poster

Multiscale modeling has yielded immense success on various machine learning tasks. However, it has not been properly explored for the prominent task of information diffusion, which aims to understand how information propagates along users in online social networks. For a specific user, whether and w…

2020

Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior Knowledge

CVPR 2020poster

Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superi…

Cited by 92PDFScholar
2020

Laying the Foundations of Deep Long-Term Crowd Flow Prediction

ECCV 2020poster

Predicting the crowd behavior in complex environments is a key requirement for crowd and disaster management, architectural design, and urban planning. Given a crowd's immediate state, current approaches must be successively repeated over multiple time-steps for long-term predictions, leading to com…

2019

Semantic Graph Convolutional Networks for 3D Human Pose Regression

CVPR 2019poster

In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convolution filters and shared transformation matrix for each node. To address these limitations, we propose Semantic Graph Co…

Cited by 694PDFcodeScholar
2018

Learning to Forecast and Refine Residual Motion for Image-to-Video Generation

ECCV 2018poster

We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such problems typically train transformation networks to generate future frames conditioned on the structure sequence. Parallel wor…

Cited by 119SourcePDFScholar
2018

Show Me a Story: Towards Coherent Neural Story Illustration

CVPR 2018poster

We propose an end-to-end network for the visual illustration of a sequence of sentences forming a story. At the core of our model is the ability to model the inter-related nature of the sentences within a story, as well as the ability to learn coherence to support reference resolution. The framework…