← Search

Hang Zhao

114 accepted papers

2026

Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots

RSS 2026poster

Animatronic faces are a central component of socially interactive robots, enabling rich nonverbal communication through facial articulation. However, state-of-the-art animatronic faces are typically tailored systems: each new facial geometry requires extensive manual mechanical redesign, making larg…

Cited by 0SourceScholar
2026

DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning

ICRA 2026poster

Vision-Language-Action (VLA) models have recently shown impressive generalization and language-guided manipulation capabilities. However, their performance degrades on tasks requiring precise spatial reasoning due to limited spatial reasoning inherited from Vision-Language Models (VLMs). Existing VL…

2026

DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers

CVPR 2026

Diffusion models have recently motivated great success in many generation tasks like object removal. Nevertheless, existing image decomposition methods struggle to disentangle semi-transparent or transparent layer occlusions due to mask prior dependencies, static object assumptions, and the lack of

Cited by 0SourcecodeScholar
2026

Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving

ICLR 2026poster

End-to-End (E2E) solutions have emerged as a mainstream approach for autonomous driving systems, with Vision-Language-Action (VLA) models representing a new paradigm that leverages pre-trained multimodal knowledge from Vision-Language Models (VLMs) to interpret and interact with complex real-world e…

Cited by 0SourcecodeScholar
2026

DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking

ICLR 2026poster

The advent of Vision-Language Models (VLMs) has significantly advanced end-to-end autonomous driving, demonstrating powerful reasoning abilities for high-level behavior planning tasks. However, existing methods are often constrained by a passive perception paradigm, relying solely on text-based reas…

Cited by 0SourcecodeScholar
2026

FASTer: Toward Powerful and Efficient Autoregressive Vision–Language–Action Models with Learnable Action Tokenizer and Block-wise Decoding

ICLR 2026poster

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce \textbf{FASTer}, a unified f…

Cited by 0SourceScholar
2026

Intrinsic Entropy of Context Length Scaling in LLMs

ICLR 2026oral

There has been work discussing the impact of long context on Language Model performance: some find that long irrelevant context could harm performance, while some experimentally summarize loss reduction by relevant long context as Scaling Laws. This calls for a more thorough understanding on how lon…

Cited by 0SourcecodeScholar
2026

Multimodal Dataset Distillation via Phased Teacher Models

ICLR 2026poster

Multimodal dataset distillation aims to construct compact synthetic datasets that enable efficient compression and knowledge transfer from large-scale image-text data. However, existing approaches often fail to capture the complex, dynamically evolving knowledge embedded in the later training stages…

Cited by 0SourcecodeScholar
2026

ORV: 4D Occupancy-centric Robot Video Generation

CVPR 2026

Recent embodied intelligence suffers from data scarcity, while conventional simulators lack visual realism. Controllable video generation is emerging as a promising data engine, yet current action-conditioned methods still fall short: generated videos are limited in fidelity and temporal consistency

Cited by 0SourcecodeScholar
2026

Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale

ICLR 2026poster

The rapid advancement of video generation models has enabled the creation of highly realistic synthetic media, raising significant societal concerns regarding the spread of misinformation. However, current detection methods suffer from critical limitations. They often rely on preprocessing operation…

Cited by 0SourceScholar
2026

TOWARDS RELIABLE TIME SERIES FORECASTING UNDER FUTURE UNCERTAINTY: AMBIGUITY AND NOVELTY REJECTION MECHANISMS

ICASSP 2026poster

In real-world time series forecasting, uncertainty and lack of reliable evaluation pose significant challenges. Notably, forecasting errors often arise from underfitting in-distribution data and failing to handle out-of-distribution inputs. To enhance model reliability, we introduce a dual rejection…

Cited by 0SourcePDFScholar
2026

UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos

CVPR 2026

Dexterous manipulation remains challenging due to the cost of collecting real-robot teleoperation data, the heterogeneity of hand embodiments, and the high dimensionality of control. We present UniDex, a robot foundation suite that couples a large-scale robot-centric dataset with a unified vision-la

Cited by 0SourcecodeScholar
2025

CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting

IROS 2025

Vehicle-to-everything (V2X) communication plays a crucial role in autonomous driving, enabling cooperation between vehicles and infrastructure. While simulation has significantly contributed to various autonomous driving tasks, its potential for data generation and augmentation in V2X scenarios rema

Cited by 4SourcecodeScholar
2025

Chameleon: Fast-Slow Neuro-Symbolic Lane Topology Extraction

ICRA 2025

Lane topology extraction involves detecting lanes and traffic elements and determining their relationships, a key perception task for mapless autonomous driving. This task requires complex reasoning, such as determining whether it is possible to turn left into a specific lane. To address this challe

Cited by 12SourcecodeScholar
2025

Conditioning Matters: Training Diffusion Policies is Faster Than You Think

NeurIPS 2025poster

Diffusion policies have emerged as a mainstream paradigm for building vision-language-action (VLA) models. Although they demonstrate strong robot control capabilities, their training efficiency remains suboptimal. In this work, we identify a fundamental challenge in conditional diffusion policy trai…

Cited by 0SourceScholar
2025

Delving into Mapping Uncertainty for Mapless Trajectory Prediction

IROS 2025

Recent advances in autonomous driving are moving towards mapless approaches, where High-Definition (HD) maps are generated online directly from sensor data, reducing the need for expensive labeling and maintenance. However, the reliability of these online-generated maps remains uncertain. While inco

Cited by 4SourcecodeScholar
2025

Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving

ICRA 2025

Accurately predicting 3D occupancy grids from visual inputs is critical for autonomous driving, but current discriminative methods struggle with noisy data, incomplete observations, and the complex structures inherent in 3D scenes. In this work, we reframe 3D occupancy prediction as a generative mod

Cited by 5SourceScholar
2025

GLIC-Calib: Targetless and One-Shot Spatial-Temporal Calibration of LiDAR-IMU-Camera for Ground Vehicles

IROS 2025

Accurate spatial-temporal parameters of LiDAR, IMU and camera, including extrinsic parameter and time offset, is the key to ensure multi-sensor fusion performance for ground vehicles. Compared with target-based calibration method, targetless method does not need artificial targets which is more conv

Cited by 1SourceScholar
2025

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

ICCV 2025poster

Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy annotations, which limits scalability and prevents leveraging vast amounts of potential crowdsourced data for auto-labe…

2025

Generalizing Motion Planners with Mixture of Experts for Autonomous Driving

ICRA 2025

Large real-world driving datasets have sparked significant research into various aspects of learning-based motion planners for autonomous driving. These include data augmentation, model architecture, reward design, training strategies, and planner pipelines. In this paper, we review and benchmark pr

Cited by 23SourcecodeScholar
2025

Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

NeurIPS 2025poster

Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80,000 meticulously curate…

Cited by 0SourcecodeScholar
2025

LONG3R: Long Sequence Streaming 3D Reconstruction

ICCV 2025poster

Recent advancements in multi-view scene reconstruction have been significant, yet existing methods face limitations when processing streams of input images. These methods either rely on time-consuming offline optimization or are restricted to shorter sequences, hindering their applicability in real-…

2025

Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Control

RSS 2025poster

Previous animatronic faces struggle to effectively express emotions due to both hardware and software limitations. On the hardware side, earlier approaches either used rigid-driven mechanisms, which provide precise control but are difficult to design within constrained spaces, or tendon-driven mecha…

Cited by 0PDFScholar
2025

PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation

RSS 2025poster

Non-prehensile manipulation, such as pushing and poking, involves moving objects without grasping, offering cost-effective solutions in constrained environments. However, it presents challenges due to sensitivity to complex physics like friction and restitution. Existing approaches either rely on ex…

Cited by 1PDFScholar
2025

PINN-Based Predictive Control Combined With Unknown Payload Identification for Robots With Prismatic Quasi-Direct-Drives

RA-L 2025

This study introduces a unified control framework that addresses the challenge of precise robots with Quasi-Direct-Drives under unknown payloads, named as online payload identification-based physics-informed neural network predictive control (OPI-PINNPC). By integrating online payload identification

Cited by 3SourceScholar
2025

Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset Distillation

NeurIPS 2025poster

Dataset distillation compresses large-scale datasets into compact, highly informative synthetic data, significantly reducing storage and training costs. However, existing research primarily focuses on balanced datasets and struggles to perform under real-world long-tailed distributions. In this work…

Cited by 0SourceScholar
2025

Reusing Attention for One-stage Lane Topology Understanding

IROS 2025

Understanding lane topology relationships accurately is critical for safe autonomous driving. However, existing two-stage methods suffer from inefficiencies due to error propagations and increased computational overheads. To address these challenges, we propose a one-stage architecture that simultan

Cited by 6SourcecodeScholar
2025

RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation

IROS 2025

Visual augmentation has become a crucial technique for enhancing the visual robustness of imitation learning. However, existing methods are often limited by prerequisites such as camera calibration or the need for controlled environments (e.g., green screen setups). In this work, we introduce RoboEn

Cited by 36SourcecodeScholar
2025

SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model

ICRA 2025

The application of vision-language models (VLMs) has achieved impressive success in various robotics tasks. However, there are few explorations for foundation models used in quadruped robot navigation through terrains in 3D environments. We introduce SARO (Space-Aware Robot System for Terrain Crossi

Cited by 5SourcecodeScholar
2025

TrackOcc: Camera-Based 4D Panoptic Occupancy Tracking

ICRA 2025

Comprehensive and consistent dynamic scene understanding from camera input is essential for advanced autonomous systems. Traditional camera-based perception tasks like 3D object tracking and semantic occupancy prediction lack either spatial comprehensiveness or temporal consistency. In this work, we

Cited by 4SourcecodeScholar
2025

VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion

RA-L 2025

Recent success in legged robot locomotion is attributed to the integration of reinforcement learning and physical simulators. However, these policies often encounter challenges when deployed in real-world environments due to sim-to-real gaps, as simulators typically fail to replicate visual realism

Cited by 26SourceScholar
2024

Boosting Offline Reinforcement Learning for Autonomous Driving with Hierarchical Latent Skills

ICRA 2024poster

Learning-based vehicle planning is receiving increasing attention with the emergence of diverse driving simulators and large-scale driving datasets. While offline reinforcement learning (RL) is well suited for these safety-critical tasks, it still struggles to plan over extended periods. In this wor…

Cited by 11SourceScholar
2024

Deep Demonstration Tracing: Learning Generalizable Imitator Policy for Runtime Imitation from a Single Demonstration

ICML 2024poster

One-shot imitation learning (OSIL) is to learn an imitator agent that can execute multiple tasks with only a single demonstration. In real-world scenario, the environment is dynamic, e.g., unexpected changes can occur after demonstration. Thus, achieving generalization of the imitator agent is cruci…

2024

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

CoRL 2024poster

A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors. We introduce DriveVLM, an autonomous driving system leveraging Vision-Language Models (VLMs) for enhanced scene understandi…

Cited by 190SourceScholar
2024

Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation

ECCV 2024poster

"CLIP, as a vision-language model, has significantly advanced Open-Vocabulary Semantic Segmentation (OVSS) with its zero-shot capabilities. Despite its success, its application to OVSS faces challenges due to its initial image-level alignment training, which affects its performance in tasks requirin…

2024

LiDAR-based 4D Occupancy Completion and Forecasting

IROS 2024poster

Scene completion and forecasting are two popular perception problems in research for mobile agents like autonomous vehicles. Existing approaches treat the two problems in isolation, resulting in a separate perception of the two aspects. In this paper, we introduce a novel LiDAR perception task of Oc…

Cited by 18SourcecodeScholar
2024

Multi-Task Self-Supervised Learning for Medical Image Segmentation

ICASSP 2024accepted

Although medical image segmentation has achieved remarkable results with supervised learning, obtaining labeled data remains challenging and costly. To counteract this, we present the MTSPSeg, a multi-task self-supervised learning framework. We establish the dynamic gradient learning rate (DGLR) str…

Cited by 0SourceScholar
2024

P-MapNet: Far-Seeing Map Generator Enhanced by Both SDMap and HDMap Priors

RA-L 2024

Autonomous vehicles are gradually entering city roads today, with the help of high-definition maps (HDMaps). However, the reliance on HDMaps prevents autonomous vehicles from stepping into regions without this expensive digital infrastructure. This fact drives many researchers to study online HDMap

Cited by 62SourceScholar
2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

IROS 2024

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details

Cited by 90SourcecodeScholar
2024

ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information

EMNLP 2024finding

In different NLP tasks, detecting harmful content is crucial for online environments, especially with the growing influence of social media. However, previous research has two main issues: 1) a lack of data in low-resource settings, and 2) inconsistent definitions and criteria for judging harmful co…

Cited by 3SourcePDFScholar
2024

Uncertainty-Aware Decision Transformer for Stochastic Driving Environments

CoRL 2024poster

Offline Reinforcement Learning (RL) enables policy learning without active interactions, making it especially appealing for self-driving tasks. Recent successes of Transformers inspire casting offline RL as sequence modeling, which, however, fails in stochastic environments with incorrect assumption…

Cited by 5SourceScholar
2023

A Universal Semantic-Geometric Representation for Robotic Manipulation

CoRL 2023poster

Robots rely heavily on sensors, especially RGB and depth cameras, to perceive and interact with the world. RGB cameras record 2D images with rich semantic information while missing precise spatial information. On the other side, depth cameras offer critical 3D geometry data but capture limited seman…

Cited by 22SourcecodeScholar
2023

Cross-Dataset Sensor Alignment: Making Visual 3D Object Detector Generalizable

CoRL 2023poster

While camera-based 3D object detection has evolved rapidly, these models are susceptible to overfitting to specific sensor setups. For example, in autonomous driving, most datasets are collected using a single sensor configuration. This paper evaluates the generalization capability of camera-based 3…

Cited by 3SourceScholar
2023

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

NeurIPS 2023poster

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation quality in terms of temporal synchronization and audio-visual re…

2023

GeoMAE: Masked Geometric Target Prediction for Self-Supervised Point Cloud Pre-Training

CVPR 2023poster

This paper tries to address a fundamental question in point cloud self-supervised learning: what is a good signal we should leverage to learn features from point clouds without annotations? To answer that, we introduce a point cloud representation learning framework, based on geometric feature recon…

2023

INT2: Interactive Trajectory Prediction at Intersections

ICCV 2023poster

Motion forecasting is an important component in autonomous driving systems. One of the most challenging problems in motion forecasting is interactive trajectory prediction, whose goal is to jointly forecasts the future trajectories of interacting agents. To this end, we present a large-scale interac…

Cited by 10PDFcodeScholar
2023

Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving

NeurIPS 2023poster

Robotic perception requires the modeling of both 3D geometry and semantics. Existing methods typically focus on estimating 3D bounding boxes, neglecting finer geometric details and struggling to handle general, out-of-vocabulary objects. 3D occupancy prediction, which estimates the detailed occupanc…

2023

On Uni-Modal Feature Learning in Supervised Multi-Modal Learning

ICML 2023poster

We abstract the features (i.e. learned representations) of multi-modal data into 1) uni-modal features, which can be learned from uni-modal training, and 2) paired features, which can only be learned from cross-modal interactions. Multi-modal models are expected to benefit from cross-modal interacti…

2023

P4P: Conflict-Aware Motion Prediction for Planning in Autonomous Driving

IROS 2023poster

Motion prediction is crucial in enabling safe motion planning for autonomous vehicles in interactive scenarios. It allows the planner to identify potential conflicts with other traffic agents and generate safe plans. Existing motion predictors often focus on reducing prediction errors, yet it remain…

Cited by 4SourceScholar
2023

PVT++: A Simple End-to-End Latency-Aware Visual Tracking Framework

ICCV 2023poster

Visual object tracking is essential to intelligent robots. Most existing approaches have ignored the online latency that can cause severe performance degradation during real-world processing. Especially for unmanned aerial vehicles (UAVs), where robust tracking is more challenging and onboard comput…

Cited by 11PDFcodeScholar
2023

Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

ICLR 2023top-25%

Robots operating in the real world require both rich manipulation skills as well as the ability to semantically reason about when to apply those skills. Towards this goal, recent works have integrated semantic representations from large-scale pretrained vision-language (VL) models into manipulation…

2023

Robot Parkour Learning

CoRL 2023oral

Parkour is a grand challenge for legged locomotion that requires robots to overcome various obstacles rapidly in complex environments. Existing methods can generate either diverse but blind locomotion skills or vision-based but specialized skills by using reference animal data or complex rewards. Ho…

Cited by 195SourcecodeScholar
2023

Self-supervision through Random Segments with Autoregressive Coding (RandSAC)

ICLR 2023poster

Inspired by the success of self-supervised autoregressive representation learning in natural language (GPT and its variants), and advances in recent visual architecture design with Vision Transformers (ViTs), in this paper, we explore the effects various design choices have on the success of applyin…

Cited by 15SourcePDFScholar
2023

SparseViT: Revisiting Activation Sparsity for Efficient High-Resolution Vision Transformer

CVPR 2023poster

High-resolution images enable neural networks to learn richer visual representations. However, this improved performance comes at the cost of growing computational complexity, hindering their usage in latency-sensitive applications. As not all pixels are equal, skipping computations for less-importa…

Cited by 55SourcePDFScholar
2023

The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation

ICLR 2023top-5%

Crossmodal knowledge distillation (KD) extends traditional knowledge distillation to the area of multimodal learning and demonstrates great success in various applications. To achieve knowledge transfer across modalities, a pretrained network from one modality is adopted as the teacher to provide su…

2023

VectorMapNet: End-to-end Vectorized HD Map Learning

ICML 2023poster

Autonomous driving systems require High-Definition (HD) semantic maps to navigate around urban roads. Existing solutions approach the semantic mapping problem by offline manual annotation, which suffers from serious scalability issues. Recent learning-based methods produce dense rasterized segmentat…

2023

ViP3D: End-to-End Visual Trajectory Prediction via 3D Agent Queries

CVPR 2023poster

Perception and prediction are two separate modules in the existing autonomous driving systems. They interact with each other via hand-picked features such as agent bounding boxes and trajectories. Due to this separation, prediction, as a downstream module, only receives limited information from the…

2023

What Happened 3 Seconds Ago? Inferring the Past With Thermal Imaging

CVPR 2023poster

Inferring past human motion from RGB images is challenging due to the inherent uncertainty of the prediction problem. Thermal images, on the other hand, encode traces of past human-object interactions left in the environment via thermal radiation measurement. Based on this observation, we collect th…

2022

AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection

IJCAI 2022poster

Object detection through either RGB images or the LiDAR point clouds has been extensively explored in autonomous driving. However, it remains challenging to make these two data sources complementary and beneficial to each other. In this paper, we propose AutoAlign, an automatic feature fusion strat…

Cited by 140SourcePDFScholar
2022

CYBORGS: Contrastively Bootstrapping Object Representations by Grounding in Segmentation

ECCV 2022poster

"Many recent approaches in contrastive learning have worked to close the gap between pretraining on iconic images like ImageNet and pretraining on complex scenes like COCO. This gap exists largely because commonly used random crop augmentations obtain semantically inconsistent content in crowded sce…

2022

Co-Advise: Cross Inductive Bias Distillation

CVPR 2022poster

The inductive bias of vision transformers is more relaxed that cannot work well with insufficient data. Knowledge distillation is thus introduced to assist the training of transformers. Unlike previous works, where merely heavy convolution-based teachers are provided, in this paper, we delve into th…

Cited by 82PDFcodeScholar
2022

Depth Estimation Matters Most: Improving Per-Object Depth Estimation for Monocular 3D Detection and Tracking

ICRA 2022poster

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior performance when compared to LiDAR-based techniques. Through s…

Cited by 24SourceScholar
2022

Egocentric Prediction of Action Target in 3D

CVPR 2022poster

We are interested in anticipating as early as possible the target location of a person's object manipulation action in a 3D workspace from egocentric vision. It is important in fields like human-robot collaboration, but has not yet received enough attention from vision and learning communities. To s…

Cited by 17PDFScholar
2022

Embracing Single Stride 3D Object Detector With Sparse Transformer

CVPR 2022poster

In LiDAR-based 3D object detection for autonomous driving, the ratio of the object size to input scene size is significantly smaller compared to 2D detection cases. Overlooking this difference, many 3D detectors directly follow the common practice of 2D detectors, which downsample the feature maps e…

Cited by 305PDFcodeScholar
2022

IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor Scenes

ICLR 2022poster

Building embodied intelligent agents that can interact with 3D indoor environments has received increasing research attention in recent years. While most works focus on single-object or agent-object visual functionality and affordances, our work proposes to study a novel, underexplored, kind of visu…

Cited by 7SourcePDFScholar
2022

InterSim: Interactive Traffic Simulation via Explicit Relation Modeling

IROS 2022poster

Interactive traffic simulation is crucial to autonomous driving systems by enabling testing for planners in a more scalable and safe way compared to real-world road testing. Existing approaches learn an agent model from large-scale driving data to simulate realistic traffic scenarios, yet it remains…

Cited by 35SourcecodeScholar
2022

Intrinsically Motivated Self-supervised Learning in Reinforcement Learning

ICRA 2022poster

In vision-based reinforcement learning (RL) tasks, it is prevalent to assign auxiliary tasks with a surrogate self-supervised loss so as to obtain more semantic representations and improve sample efficiency. However, abundant information in self-supervised auxiliary tasks has been disregarded, since…

Cited by 5SourceScholar
2022

M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction

CVPR 2022poster

Predicting future motions of road participants is an important task for driving autonomously in urban scenes. Existing models excel at predicting marginal trajectories for single agents, yet it remains an open question to jointly predict scene compliant trajectories over multiple agents. The challen…

Cited by 122PDFScholar
2022

R4D: Utilizing Reference Objects for Long-Range Distance Estimation

ICLR 2022poster

Estimating the distance of objects is a safety-critical task for autonomous driving. Focusing on short-range objects, existing methods and datasets neglect the equally important long-range objects. In this paper, we introduce a challenging and under-explored task, which we refer to as Long-Range Dis…

Cited by 7SourcePDFScholar
2022

S3T: Self-Supervised Pre-Training with Swin Transformer For Music Classification

ICASSP 2022accepted

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S3T introduces a momentum-based paradigm, MoCo, with Swin Transformer as its feat…

Cited by 0SourceScholar
2022

SimIPU: Simple 2D Image and 3D Point Cloud Unsupervised Pre-training for Spatial-Aware Visual Representations

AAAI 2022technical

Pre-training has become a standard paradigm in many computer vision tasks. However, most of the methods are generally designed on the RGB image domain. Due to the discrepancy between the two-dimensional image plane and the three-dimensional space, such pre-trained models fail to perceive spatial inf…

2022

Sound2Synth: Interpreting Sound via FM Synthesizer Parameters Estimation

IJCAI 2022poster

Synthesizer is a type of electronic musical instrument that is now widely used in modern music production and sound design. Each parameters configuration of a synthesizer produces a unique timbre and can be viewed as a unique instrument. The problem of estimating a set of parameters configuration th…

2021

DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries

CoRL 2021poster

We introduce a framework for multi-camera 3D object detection. In contrast to existing works, which estimate 3D bounding boxes directly from monocular images or use depth prediction networks to generate input for 3D object detection from 2D information, our method manipulates predictions directly in…

Cited by 899SourcecodeScholar
2021

HDMapGen: A Hierarchical Graph Generative Model of High Definition Maps

CVPR 2021poster

High Definition (HD) maps are maps with precise definitions of road lanes with rich semantics of the traffic rules. They are critical for several key stages in an autonomous driving system, including motion forecasting and planning. However, there are only a small amount of real-world road topologie…

Cited by 71PDFScholar
2021

Large Scale Interactive Motion Forecasting for Autonomous Driving: The Waymo Open Motion Dataset

ICCV 2021poster

As autonomous driving systems mature, motion forecasting has received increasing attention as a critical requirement for planning. Of particular importance are interactive situations such as merges, unprotected turns, etc., where predicting individual object motion is not sufficient. Joint predictio…

Cited by 624PDFScholar
2021

Multi-Agent Trajectory Prediction by Combining Egocentric and Allocentric Views

CoRL 2021poster

Trajectory prediction of road participants such as vehicles and pedestrians is crucial for autonomous driving. Recently, graph neural network (GNN) is widely adopted to capture the social interactions among the agents. Many GNN-based models formulate the prediction task as a single-agent prediction…

Cited by 55SourceScholar
2021

Neural Dubber: Dubbing for Videos According to Scripts

NeurIPS 2021poster

Dubbing is a post-production process of re-recording actors’ dialogues, which is extensively used in filmmaking and video production. It is usually performed manually by professional voice actors who read lines with proper prosody, and in synchronization with the pre-recorded videos. In this work, w…

Cited by 44SourcePDFScholar
2021

On Feature Decorrelation in Self-Supervised Learning

ICCV 2021poster

In self-supervised representation learning, a common idea behind most of the state-of-the-art approaches is to enforce the robustness of the representations to predefined augmentations. A potential issue of this idea is the existence of completely collapsed solutions (i.e., constant features), which…

Cited by 236PDFcodeScholar
2021

Online 3D Bin Packing with Constrained Deep Reinforcement Learning

AAAI 2021technical

We solve a challenging yet practically useful variant of 3D Bin Packing Problem (3D-BPP). In our problem, the agent has limited information about the items to be packed into a single bin, and an item must be packed immediately after its arrival without buffering or readjusting. The item's placement…

2021

What Makes Multi-Modal Learning Better than Single (Provably)

NeurIPS 2021poster

The world provides us with data of multiple modalities. Intuitively, models fusing data from different modalities outperform their uni-modal counterparts, since more information is aggregated. Recently, joining the success of deep learning, there is an influential line of work on deep multi-modal le…

Cited by 340SourcePDFScholar
2020

Scalability in Perception for Autonomous Driving: Waymo Open Dataset

CVPR 2020poster

The research community has increasing interest in autonomous driving research, despite the resource intensity of obtaining representative real world data. Existing self-driving datasets are limited in the scale and variation of the environments they capture, even though generalization within and bet…

Cited by 3735PDFScholar
2020

UnModNet: Learning to Unwrap a Modulo Image for High Dynamic Range Imaging

NeurIPS 2020poster

A conventional camera often suffers from over- or under-exposure when recording a real-world scene with a very high dynamic range (HDR). In contrast, a modulo camera with a Markov random field (MRF) based unwrapping algorithm can theoretically accomplish unbounded dynamic range but shows degenerate…

Cited by 18SourcePDFScholar
2020

Unsupervised Monocular Depth Learning in Dynamic Scenes

CoRL 2020

We present a method for jointly training the estimation of depth, ego-motion, and a dense 3D translation field of objects relative to the scene, with monocular photometric consistency being the sole source of supervision. We show that this apparently heavily underdetermined problem can be regularize

2020

VectorNet: Encoding HD Maps and Agent Dynamics From Vectorized Representation

CVPR 2020poster

Behavior prediction in dynamic, multi-agent systems is an important problem in the context of self-driving cars, due to the complex representations and interactions of road components, including moving agents (e.g. pedestrians and vehicles) and road context information (e.g. lanes, traffic lights).…

Cited by 1022PDFScholar
2019

HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization

ICCV 2019poster

This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage consensus and disagreement among visual classifiers to automatically mine candidate short clips fr…

Cited by 347PDFScholar
2019

Self-Supervised Moving Vehicle Tracking With Stereo Sound

ICCV 2019poster

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audiovisual data to learn to localize objects (moving vehicles) in a visual re…

Cited by 174PDFScholar
2019

Self-supervised Audio-visual Co-segmentation

ICASSP 2019accepted

Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this paper we develop a neural network model for visual object segmentation and sound source separation that learns from natural…

Cited by 0SourceScholar
2019

Through-Wall Human Mesh Recovery Using Radio Signals

ICCV 2019poster

This paper presents RF-Avatar, a neural network model that can estimate 3D meshes of the human body in the presence of occlusions, baggy clothes, and bad lighting conditions. We leverage that radio frequency (RF) signals in the WiFi range traverse clothes and occlusions and bounce off the human body…

Cited by 123PDFScholar
2018

The Sound of Pixels

ECCV 2018poster

We introduce PixelPlayer, a system that, by leveraging large amounts of unlabeled videos, learns to locate image regions which produce sounds and separate the input sounds into a set of components that represents the sound from each pixel. Our approach capitalizes on the natural synchronization of t…

Cited by 638SourcePDFScholar
2018

Through-Wall Human Pose Estimation Using Radio Signals

CVPR 2018poster

This paper demonstrates accurate human pose estimation through walls and occlusions. We leverage the fact that wireless signals in the WiFi frequencies traverse walls and reflect off the human body. We introduce a deep neural network approach that parses such radio signals to estimate 2D poses. Sinc…

Cited by 731SourcePDFScholar
2017

Duckietown: An open, inexpensive and flexible platform for autonomy education and research

ICRA 2017poster

Duckietown is an open, inexpensive and flexible platform for autonomy education and research. The platform comprises small autonomous vehicles (“Duckiebots”) built from off-the-shelf components, and cities (“Duckietowns”) complete with roads, signage, traffic lights, obstacles, and citizens (duckies…

Cited by 281SourceScholar