← Search

Kris Kitani

61 accepted papers

2026

BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning

ICLR 2026poster

Building Behavioral Foundation Models (BFMs) for humanoid robots has the potential to unify diverse control tasks under a single, promptable generalist policy. However, existing approaches are either exclusively deployed on simulated humanoid characters, or specialized to specific tasks such as trac…

Cited by 0SourcecodeScholar
2026

DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction

CVPR 2026

We present DuoMo, a generative method that recovers human motion in world-space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade-off: generalizing from diverse and noisy video inputs while maintaining global

Cited by 0SourcecodeScholar
2026

DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation

CVPR 2026

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without considering social context or modeling the mutual dynamics between

Cited by 0SourceScholar
2026

Faster Vision Transformers with Adaptive Patches

ICLR 2026poster

Vision Transformers (ViTs) partition input images into uniformly sized patches regardless of their content, resulting in long input sequence lengths for high-resolution images. We present Adaptive Patch Transformers (APT), which addresses this by using multiple different patch sizes within the same…

Cited by 0SourcecodeScholar
2026

Ground Reaction Inertial Poser: Physics-based Human Motion Capture from Sparse IMUs and Insole Pressure Sensors

CVPR 2026

We propose Ground Reaction Inertial Poser (GRIP), a method that reconstructs physically plausible human motion using four wearable devices. Unlike conventional IMU-only approaches, GRIP combines IMU signals with foot pressure data to capture both body dynamics and ground interactions. Furthermore, r

Cited by 0SourceScholar
2026

Grounded Latents for Entity-Centric 4D Scene Generation

CVPR 2026

Although recent work has explored generative modeling of 3D or 4D driving scenes, most approaches operate on dense voxel-based representations, which are computationally expensive and struggle to maintain temporal or structural consistency. These methods often produce blurred or merged entities (i.e

Cited by 0SourceScholar
2026

SAM 3D Body: Robust Full-Body Human Mesh Recovery

CVPR 2026

We introduce SAM 3D Body (3DB), a promptable model for single-image full-body 3D human mesh recovery (HMR) that demonstrates state-of-the-art performance, with strong generalization and consistent accuracy in diverse in-the-wild conditions. 3DB estimates the human pose of the body, feet, and hands.

Cited by 0SourcecodeScholar
2025

ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills

RSS 2025poster

Humanoid robots hold the potential for unparalleled versatility by performing human-like, whole-body skills. However, achieving agile and coordinated whole-body motions remains a significant challenge due to the dynamics mismatch between simulation and real-world physics. Existing approaches, such a…

Cited by 15PDFcodeScholar
2025

ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling

ICCV 2025poster

Parametric body models offer expressive 3D representation of humans across a wide range of poses, shapes, and facial expressions, typically derived by learning a basis over registered 3D meshes. However, existing human mesh modeling approaches struggle to capture detailed variations across diverse b…

Cited by 0SourcePDFScholar
2025

ExpertAF: Expert Actionable Feedback from Video

CVPR 2025poster

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do differently on the user. We introduce a novel method to generate act…

Cited by 3SourcePDFScholar
2025

HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models

CVPR 2025poster

We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (e.g. text, objects, partia…

Cited by 1SourcePDFScholar
2025

Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object Detection

CVPR 2025poster

While recent advancements in camera-based 3D object detection demonstrate remarkable performance, they require thousands or even millions of human-annotated frames. This requirement significantly inhibits their deployment in various locations and sensor configurations. To address this gap, we propos…

Cited by 0SourcePDFScholar
2025

ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping

CVPR 2025poster

Robotic grasping is a cornerstone capability of embodied systems. Many methods directly output grasps from partial information without modeling the geometry of the scene, leading to suboptimal motion and even collisions. To address these issues, we introduce ZeroGrasp, a novel framework that simulta…

Cited by 0SourcePDFScholar
2024

"Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos"

ECCV 2024oral

"Goal-oriented planning, or anticipating a series of actions that transition an agent from its current state to a predefined objective, is crucial for developing intelligent assistants aiding users in daily procedural tasks. The problem presents significant challenges due to the need for comprehensi…

Cited by 2SourcePDFScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis

CVPR 2024poster

We propose G-HOP a denoising diffusion based generative prior for hand-object interactions that allows modeling both the 3D object and a human hand conditioned on the object category. To learn a 3D spatial diffusion model that can capture this joint distribution we represent the human hand via a ske…

Cited by 0SourcePDFScholar
2024

JaywalkerVR: A VR System for Collecting Safety-Critical Pedestrian-Vehicle Interactions

ICRA 2024poster

Developing autonomous vehicles that can safely interact with pedestrians requires large amounts of pedestrian and vehicle data in order to learn accurate pedestrian-vehicle interaction models. However, gathering data that include crucial but rare scenarios - such as pedestrians jaywalking into heavy…

Cited by 0SourceScholar
2024

Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation

IROS 2024poster

We present Human to Humanoid (H2O), a reinforcement learning (RL) based framework that enables real-time whole-body teleoperation of a full-sized humanoid robot with only an RGB camera. To create a large-scale retargeted motion dataset of human movements for humanoid robots, we propose a scalable "s…

Cited by 83SourceScholar
2024

Real-Time Simulated Avatar from Head-Mounted Sensors

CVPR 2024highlight

We present SimXR a method for controlling a simulated avatar from information (headset pose and cameras) obtained from AR / VR headsets. Due to the challenging viewpoint of head-mounted cameras the human body is often clipped out of view making traditional image-based egocentric pose estimation chal…

Cited by 8SourcePDFScholar
2024

Video Question Answering with Procedural Programs

ECCV 2024poster

"We propose to answer questions about videos by generating short procedural programs that solve visual subtasks to obtain a final answer. We present ˙ which uses a large language model to generate Procedural Video Querying (), such programs from an input question and an API of visual modules in the…

2024

Zero-Shot Multi-Object Scene Completion

ECCV 2024poster

"We present a 3D scene completion method that recovers the complete geometry of multiple unseen objects in complex scenes from a single RGB-D image. Despite notable advancements in single-object 3D shape completion, high-quality reconstructions in highly cluttered real-world multi-object scenes rema…

Cited by 1SourcePDFScholar
2023

Azimuth Super-Resolution for FMCW Radar in Autonomous Driving

CVPR 2023poster

We tackle the task of Azimuth (angular dimension) super-resolution for Frequency Modulated Continuous Wave (FMCW) multiple-input multiple-output (MIMO) radar. FMCW MIMO radar is widely used in autonomous driving alongside Lidar and RGB cameras. However, compared to Lidar, MIMO radar is usually of lo…

2023

Ego-Humans: An Ego-Centric 3D Multi-Human Benchmark

ICCV 2023oral

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios, which limit the generalization of computer vision algorithms…

Cited by 39PDFScholar
2023

Joint Metrics Matter: A Better Standard for Trajectory Forecasting

ICCV 2023poster

Multi-modal trajectory forecasting methods commonly evaluate using single-agent metrics (marginal metrics), such as minimum Average Displacement Error (ADE) and Final Displacement Error (FDE), which fail to capture joint performance of multiple interacting agents. Only focusing on marginal metrics c…

Cited by 13PDFcodeScholar
2023

Observation-Centric SORT: Rethinking SORT for Robust Multi-Object Tracking

CVPR 2023poster

Kalman filter (KF) based methods for multi-object tracking (MOT) make an assumption that objects move linearly. While this assumption is acceptable for very short periods of occlusion, linear estimates of motion for prolonged time can be highly inaccurate. Moreover, when there is no measurement avai…

2023

Perpetual Humanoid Control for Real-time Simulated Avatars

ICCV 2023poster

We present a physics-based humanoid controller that achieves high-fidelity motion imitation and fault-tolerant behavior in the presence of noisy input (e.g. pose estimates from video or generated from language) and unexpected falls. Our controller scales up to learning ten thousand motion clips with…

Cited by 85PDFScholar
2023

ST-MVDNet++: Improve Vehicle Detection with Lidar-Radar Geometrical Augmentation via Self-Training

ICASSP 2023accepted

We aim to improve the performance of the vehicle detection model with Lidar-Radar fusion and data augmentation. The recent works for Lidar-Radar fusion such as MVDNet or ST-MVDNet, have been proposed to have effective performance in detecting vehicles, and address the issue regarding missing modalit…

Cited by 0SourceScholar
2023

Trace and Pace: Controllable Pedestrian Animation via Guided Trajectory Diffusion

CVPR 2023poster

We introduce a method for generating realistic pedestrian trajectories and full-body animations that can be controlled to meet user-defined goals. We draw on recent advances in guided diffusion modeling to achieve test-time controllability of trajectories, which is normally only associated with rule…

Cited by 118SourcePDFScholar
2022

Cross-Domain Adaptive Teacher for Object Detection

CVPR 2022poster

We address the task of domain adaptation in object detection, where there is a domain gap between a domain with annotations (source) and a domain of interest without annotations (target). As an effective semi-supervised learning method, the teacher-student framework (a student model is supervised by…

Cited by 233PDFcodeScholar
2022

DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion

CVPR 2022poster

A typical pipeline for multi-object tracking (MOT) is to use a detector for object localization, and following re-identification (re-ID) for object association. This pipeline is partially motivated by recent progress in both object detection and re-ID, and partially motivated by biases in existing t…

Cited by 327PDFcodeScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

GLAMR: Global Occlusion-Aware Human Mesh Recovery With Dynamic Cameras

CVPR 2022oral

We present an approach for 3D global human mesh recovery from monocular videos recorded with dynamic cameras. Our approach is robust to severe and long-term occlusions and tracks human bodies even when they go outside the camera's field of view. To achieve this, we first propose a deep generative mo…

Cited by 136PDFcodeScholar
2022

Learnable Spatio-Temporal Map Embeddings for Deep Inertial Localization

IROS 2022poster

Indoor localization systems often fuse inertial odometry with map information via hand-defined methods to reduce odometry drift, but such methods are sensitive to noise and struggle to generalize across odometry sources. To address the robustness problem in map utilization, we propose a data-driven…

Cited by 2SourceScholar
2022

Modality-Agnostic Learning for Radar-Lidar Fusion in Vehicle Detection

CVPR 2022poster

Fusion of multiple sensor modalities such as camera, Lidar, and Radar, which are commonly found on autonomous vehicles, not only allows for accurate detection but also robustifies perception against adverse weather conditions and individual sensor failures. Due to inherent sensor characteristics, Ra…

Cited by 47PDFScholar
2022

REvolveR: Continuous Evolutionary Models for Robot-to-robot Policy Transfer

ICML 2022oral

A popular paradigm in robotic learning is to train a policy from scratch for every new robot. This is not only inefficient but also often impractical for complex robots. In this work, we consider the problem of transferring a policy across two different robots with significantly different parameters…

2022

Whose Track Is It Anyway? Improving Robustness to Tracking Errors With Affinity-Based Trajectory Prediction

CVPR 2022poster

Multi-agent trajectory prediction is critical for planning and decision-making in human-interactive autonomous systems, such as self-driving cars. However, most prediction models are developed separately from their upstream perception (detection and tracking) modules, assuming ground truth past traj…

Cited by 27PDFScholar
2021

Inverse Reinforcement Learning with Explicit Policy Estimates

AAAI 2021technical

Various methods for solving the inverse reinforcement learning (IRL) problem have been developed independently in machine learning and economics. In particular, the method of Maximum Causal Entropy IRL is based on the perspective of entropy maximization, while related advances in the field of econom…

Cited by 6SourcePDFScholar
2021

Joint Object Detection and Multi-Object Tracking with Graph Neural Networks

ICRA 2021poster

Object detection and data association are critical components in multi-object tracking (MOT) systems. Despite the fact that the two components are dependent on each other, prior works often design detection and data association modules separately which are trained with separate objectives. As a resu…

Cited by 382SourcecodeScholar
2021

SimPoE: Simulated Character Control for 3D Human Pose Estimation

CVPR 2021poster

Accurate estimation of 3D human motion from monocular video requires modeling both kinematics (body motion without physical forces) and dynamics (motion with physical forces). To demonstrate this, we present SimPoE, a Simulation-based approach for 3D human Pose Estimation, which integrates image-bas…

Cited by 163PDFScholar
2020

3D Multi-Object Tracking: A Baseline and New Evaluation Metrics

IROS 2020poster

3D multi-object tracking (MOT) is an essential component for many applications such as autonomous driving and assistive robotics. Recent work on 3D MOT focuses on developing accurate systems giving less attention to practical considerations such as computational cost and system complexity. In contra…

Cited by 543SourcecodeScholar
2020

Efficient Non-Line-of-Sight Imaging from Transient Sinograms

ECCV 2020poster

Non-line-of-sight (NLOS) imaging techniques use light that diffusely reflects off of visible surfaces (e.g., walls) to see around corners. One approach involves using pulsed lasers and ultrafast sensors to measure the travel time of multiply scattered light. Unlike existing NLOS techniques that gene…

Cited by 48SourcePDFScholar
2020

Inverting the Pose Forecasting Pipeline with SPF2: Sequential Pointcloud Forecasting for Sequential Pose Forecasting

CoRL 2020

Many autonomous systems forecast aspects of the future in order to aid decision-making. For example, self-driving vehicles and robotic manipulation systems often forecast future object poses by first detecting and tracking objects. However, this detect-then-forecast pipeline is expensive to scale, a

Cited by 0SourcePDFScholar
2020

Residual Force Control for Agile Human Behavior Imitation and Extended Motion Synthesis

NeurIPS 2020poster

Reinforcement learning has shown great promise for synthesizing realistic human behaviors by learning humanoid control policies from motion capture data. However, it is still very challenging to reproduce sophisticated human skills like ballet dance, or to stably imitate long-term human behaviors wi…

2019

A-EXP4: Online Social Policy Learning for Adaptive Robot-Pedestrian Interaction

IROS 2019poster

We study self-supervised adaptation of a robot's policy for social interaction, i.e., a policy for active communication with surrounding pedestrians through audio or visual signals. Inspired by the observation that humans continually adapt their behavior when interacting under varying social context…

Cited by 3SourceScholar
2019

Incremental Class Discovery for Semantic Segmentation With RGBD Sensing

ICCV 2019poster

This work addresses the task of open world semantic segmentation using RGBD sensing to discover new semantic classes over time. Although there are many types of objects in the real-word, current semantic segmentation methods make a closed world assumption and are trained only to segment a limited nu…

Cited by 23PDFScholar
2019

PRECOG: PREdiction Conditioned on Goals in Visual Multi-Agent Settings

ICCV 2019poster

For autonomous vehicles (AVs) to behave appropriately on roads populated by human-driven vehicles, they must be able to reason about the uncertain intentions and decisions of other drivers from rich perceptual information. Towards these capabilities, we present a probabilistic forecasting model of f…

Cited by 467PDFcodeScholar
2017

Predictive-State Decoders: Encoding the Future into Recurrent Networks

NeurIPS 2017poster

Recurrent neural networks (RNNs) are a vital modeling technique that rely on internal states learned indirectly by optimization of a supervised, unsupervised, or reinforcement training loss. RNNs are used to model dynamic processes that are characterized by underlying latent states whose form is oft…

Cited by 46SourcePDFScholar