← Search

Ehsan Adeli

43 accepted papers

2026

AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion

CVPR 2026

Reconstructing 3D human motion and human-object interactions (HOI) from Internet videos is a fundamental step toward building large-scale datasets of human behavior. Existing methods struggle to recover globally consistent 3D motion under dynamic cameras, especially for motion types underrepresented

Cited by 0SourceScholar
2026

Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)

CVPR 2026

Diffusion Magnetic Resonance Imaging (dMRI) plays a critical role in studying microstructural changes in the brain. It is, therefore, widely used in clinical practice; yet progress in learning general-purpose representations from dMRI has been limited. A key challenge is that existing deep learning

Cited by 0SourcecodeScholar
2026

ESAM++: Efficient Online 3D Perception on the Edge

CVPR 2026

Online 3D scene perception in real time is essential for robotics, AR/VR, and autonomous systems, particularly in edge computing scenarios where computational resources are limited and privacy is crucial. Recent state-of-the-art methods like EmbodiedSAM (ESAM) demonstrate the promise of online 3D pe

Cited by 0SourcecodeScholar
2026

Feed-forward Human Performance Capture via Progressive Canonical Space Updates

ICLR 2026poster

We present a feed-forward human performance capture method that renders novel views of a performer from a monocular RGB stream. A key challenge in this setting is the lack of sufficient observations, especially for unseen regions. Assuming the subject moves continuously over time, we take advantage…

Cited by 0SourceScholar
2026

Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation

ICML 2026poster

Latent diffusion models excel at generating high-quality images but lose the benefits of end-to-end modeling. They discard information during image encoding, require a separately trained decoder, and model an auxiliary distribution to the raw data. In this paper, we propose Latent Forcing, a simple …

Cited by 0SourceScholar
2026

QUANTIPHY: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models

CVPR 2026

Understanding the physical world is essential for generalist AI agents. However, it remains unclear whether state-of-the-art vision perception models (e.g., large VLMs) can perform quantitative physical reasoning tasks. Existing evaluations are predominantly VQA-based and qualitative, offering limit

Cited by 0SourcecodeScholar
2026

Spherical Leech Quantization for Visual Tokenization and Generation

CVPR 2026

Lookup-free quantization has received much attention due to its efficiency on parameters and scalability to a large codebook. In this paper, we present a unified formulation of different non-parametric quantization methods through the lens of lattice coding. The geometry of lattice codes explains th

Cited by 0SourcecodeScholar
2026

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

CVPR 2026

Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human behavior as a translation task--co-speech gesture or text-to-motion that maps a fixed utterance to motion clips--without requiring agentic decision-maki

Cited by 0SourceScholar
2025

Confounder-Free Continual Learning via Recursive Feature Normalization

ICML 2025poster

Confounders are extraneous variables that affect both the input and the target, resulting in spurious correlations and biased predictions. There are recent advances in dealing with or removing confounders in traditional models, such as metadata normalization (MDN), where the distribution of the lear…

Cited by 0SourcePDFScholar
2025

Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation

NeurIPS 2025poster

Capturing diversity is crucial in conditional and prompt-based image generation, particularly when conditions contain uncertainty that can lead to multiple plausible outputs. To generate diverse images reflecting this diversity, traditional methods often modify random seeds, making it difficult to d…

Cited by 0SourceScholar
2025

LOMM: Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

ICCV 2025poster

In this paper, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation that significantly improves long-term instance tracking. At the core of our method is Latest Object Memory (LOM), which robustly tracks and continuously updates the latest states of…

Cited by 0SourcePDFScholar
2025

Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis

CVPR 2025highlight

Scaling by training on large datasets has been shown to enhance the quality and fidelity of image generation and manipulation with diffusion models; however, such large datasets are not always accessible in medical imaging due to cost and privacy issues, which contradicts one of the main application…

Cited by 0SourcePDFScholar
2025

Re-thinking Temporal Search for Long-Form Video Understanding

CVPR 2025poster

Efficient understanding of long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms for long-form video understanding, studying a fundamental issue pertaining to all state-of-the-art (SOTA) long-context vision-language models (VLMs). In…

2025

Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation

ICCV 2025poster

Text-to-image diffusion models have seen significant development recently due to increasing availability of paired 2D data. Although a similar trend is emerging in 3D generation, the limited availability of high-quality 3D data has resulted in less competitive 3D diffusion models compared to their 2…

Cited by 0SourcePDFScholar
2025

The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

CVPR 2025poster

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for creating virtual characters that can communicate naturally in a…

Cited by 4SourcePDFScholar
2025

UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation

ICCV 2025poster

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately predicting and simulating movement from a first-person persp…

Cited by 0SourcePDFScholar
2024

Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly Detection

AAAI 2024technical

Logical anomalies (LA) refer to data violating underlying logical constraints e.g., the quantity, arrangement, or composition of components within an image. Detecting accurately such anomalies requires models to reason about various component types through segmentation. However, curation of pixel-le…

2024

H-ViT: A Hierarchical Vision Transformer for Deformable Image Registration

CVPR 2024highlight

This paper introduces a novel top-down representation approach for deformable image registration which estimates the deformation field by capturing various short- and long-range flow features at different scale levels. As a Hierarchical Vision Transformer (H-ViT) we propose a dual self-attention and…

2024

OccFusion: Rendering Occluded Humans with Generative Diffusion Priors

NeurIPS 2024poster

Existing human rendering methods require every part of the human to be fully visible throughout the input video. However, this assumption does not hold in real-life settings where obstructions are common, resulting in only partial visibility of the human. Considering this, we present OccFusion, an a…

Cited by 3SourcePDFScholar
2022

MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

NeurIPS 2022accept

Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabilities. While complex human activities are often hierarchical and compositional,…

Cited by 23SourcePDFScholar
2022

PrivHAR: Recognizing Human Actions from Privacy-Preserving Lens

ECCV 2022poster

"The accelerated use of digital cameras prompts an increasing concern about privacy and security, particularly in applications such as action recognition. In this paper, we propose an optimizing framework to provide robust visual privacy protection along the human action recognition pipeline. Our fr…

Cited by 32SourcePDFScholar
2022

Rethinking Architecture Design for Tackling Data Heterogeneity in Federated Learning

CVPR 2022poster

Federated learning is an emerging research paradigm enabling collaborative training of machine learning models among different organizations while keeping data private at each institution. Despite recent progress, there remain fundamental challenges such as the lack of convergence and the potential…

Cited by 224PDFcodeScholar
2021

3D CNNs With Adaptive Temporal Feature Resolutions

CVPR 2021poster

While state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there…

Cited by 39PDFcodeScholar
2021

Home Action Genome: Cooperative Compositional Action Understanding

CVPR 2021poster

Existing research on action recognition treats activities as monolithic events occurring in videos. Recently, the benefits of formulating actions as a combination of atomic-actions have shown promise in improving action understanding with the emergence of datasets containing such annotations, allowi…

Cited by 90PDFcodeScholar
2021

MOMA: Multi-Object Multi-Actor Activity Parsing

NeurIPS 2021poster

Complex activities often involve multiple humans utilizing different objects to complete actions (e.g., in healthcare settings, physicians, nurses, and patients interact with each other and various medical devices). Recognizing activities poses a challenge that requires a detailed understanding of a…

Cited by 32SourcePDFScholar
2021

TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild

ICCV 2021poster

Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems. Predicting body dynamics requires capturing subtle information embedded in the humans' interactions with each other an…

Cited by 64PDFScholar
2020

It is not the Journey but the Destination: Endpoint Conditioned Trajectory Prediction

ECCV 2020poster

Human trajectory forecasting with multiple socially interact-ing agents is of critical importance for autonomous navigation in human environments, e.g., for self-driving cars and social robots. In this work, we present Predicted Endpoint Conditioned Network (PECNet) for flexible human trajectory pre…

2020

Procedure Planning in Instructional Videos

ECCV 2020poster

In this paper, we study the problem of procedure planning in instructional videos, which can be seen as the first step towards enabling autonomous agents to plan for complex tasks in everyday settings such as cooking. Given the current visual observation of the world and a visual goal, we ask the qu…

Cited by 119SourcePDFScholar
2020

Socially and Contextually Aware Human Motion and Pose Forecasting

RA-L 2020

Smooth and seamless robot navigation while interacting with humans depends on predicting human movements. Forecasting such human dynamics often involves modeling human trajectories (global motion) or detailed body joint movements (local motion). Prior work typically tackled local and global human mo

Cited by 94SourceScholar
2020

Spatio-Temporal Graph for Video Captioning With Knowledge Distillation

CVPR 2020poster

Video captioning is a challenging task that requires a deep understanding of visual scenes. State-of-the-art methods generate captions using either scene-level or object-level information but without explicitly modeling object interactions. Thus, they often fail to make visually grounded predictions…

Cited by 354PDFScholar
2020

Spatiotemporal Relationship Reasoning for Pedestrian Intent Prediction

RA-L 2020

Reasoning over visual data is a desirable capability for robotics and vision-based applications. Such reasoning enables forecasting the next events or actions in videos. In recent years, various models have been developed based on convolution operations for prediction or forecasting, but they lack t

Cited by 186SourceScholar
2019

Self-Supervised Representation Learning via Neighborhood-Relational Encoding

ICCV 2019poster

In this paper, we propose a novel self-supervised representation learning by taking advantage of a neighborhood-relational encoding (NRE) among the training data. Conventional unsupervised learning methods only focused on training deep networks to understand the primitive characteristics of the visu…

Cited by 50PDFcodeScholar
2019

Unsupervised Feature Ranking and Selection Based on Autoencoders

ICASSP 2019accepted

Feature selection is one of the most important and widely-used dimension reduction techniques due to its efficiency and intractability of the results. In this paper, we propose a simple but efficient unsupervised feature ranking and selection method by exploiting the geometry of the original feature…

Cited by 0SourceScholar
2018

Adversarially Learned One-Class Classifier for Novelty Detection

CVPR 2018poster

Novelty detection is the process of identifying the observation(s) that differ in some respect from the training observations (the target class). In reality, the novelty class is often absent during training, poorly sampled or not well defined. Therefore, one-class classifiers can efficiently model…

2017

Structured prediction with short/long-range dependencies for human activity recognition from depth skeleton data

IROS 2017poster

One of the main abilities that the robots need to maintain is to efficiently communicate with people in a humanly manner. Thus, human activity recognition (HAR) would be an integral part of such a human-robot interaction system. One of the major challenges in HAR is that the individuals perform thei…

Cited by 10SourceScholar