← Search

Jianfei Yang

44 accepted papers

2026

DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Driving

CVPR 2026

Autonomous driving has made substantial progress recently, achieving reliable performance in most real-world environments. However, existing algorithms still depend heavily on high-definition maps, making them ineffective in mapless scenarios such as indoor parking lots. These limitations hinder sea

Cited by 0SourceScholar
2026

HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning

CVPR 2026

Recent advances in diffusion models have demonstrated impressive capability in generating high-quality images for simple prompts. However, when confronted with complex prompts involving multiple objects and hierarchical structures, existing models struggle to accurately follow instructions, leading

Cited by 0SourcecodeScholar
2026

M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction

CVPR 2026

Human mesh reconstruction (HMR) provides direct insights into body-environment interaction, enabling various immersive applications. However, existing large-scale HMR benchmarks largely rely on line-of-sight RGB sensing, causing HMR systems to inherit the limitations of vision-based systems, includi

Cited by 0SourcecodeScholar
2026

Mask2IV: Interaction-Centric Video Generation via Mask Trajectories

AAAI 2026technical

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training, and affordance reasoning. However, existing methods often s

Cited by 0SourcePDFScholar
2026

REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?

ICLR 2026poster

Robot task planning decomposes human instructions into executable action sequences that enable robots to complete a series of complex tasks. Although recent large language model (LLM)-based task planners achieve amazing performance, they assume that human instructions are clear and straightforward.…

Cited by 0SourceScholar
2026

RF-MatID: Dataset and Benchmark for Radio Frequency Material Identification

ICLR 2026poster

Accurate material identification plays a crucial role in embodied AI systems, enabling a wide range of applications. However, current vision-based solutions are limited by the inherent constraints of optical sensors, while radio-frequency (RF) approaches, which can reveal intrinsic material properti…

Cited by 0SourcecodeScholar
2026

RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation

ICRA 2026poster

Precise robot manipulation is critical for fine-grained applications such as chemical and biological experiments, where even small errors (e.g., reagent spillage) can invalidate an entire task. Existing approaches often rely on pre-collected expert demonstrations and train policies via imitation lea…

2026

When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answering

CVPR 2026

Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied agents should know when they do not have sufficient information t

Cited by 0SourceScholar
2026

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

ICML 2026poster

Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hindered by the limited availability of sensor data and resource constraints of edge systems. While transferring pre-trained models to different sensing ap…

Cited by 0SourceScholar
2026

Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training

AAAI 2026technical

Understanding complex human activities demands the ability to decompose motion into fine-grained, semantic-aligned sub-actions. This motion grounding process is crucial for behavior analysis, embodied AI and virtual reality. Yet, most existing methods rely on dense supervision with predefined action

Cited by 0SourcePDFScholar
2026

mmPred: Radar-based Human Motion Prediction in the Dark

AAAI 2026technical

Existing Human Motion Prediction (HMP) methods based on RGB(D) cameras are sensitive to lighting conditions and raise privacy concerns, limiting their real-world applications such as firefighting and elderly care. Motivated by the robustness and privacy-preserving nature of millimeter-wave (mmWave)

Cited by 0SourcePDFScholar
2025

Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling

ICASSP 2025accepted

As small unmanned aerial vehicles (UAVs) become increasingly prevalent, there is growing concern regarding their impact on public safety and privacy, highlighting the need for advanced tracking and trajectory estimation solutions. In response, this paper introduces a novel framework that utilizes au…

Cited by 0SourceScholar
2025

CGS-SLAM: Compact 3D Gaussian Splatting for Dense Visual SLAM

IROS 2025

Recent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes. However, these approaches are built on a tremendous number of redundant 3D Gaussian ellipsoids, leading to high memory and storage costs and slow traini

Cited by 61SourceScholar
2025

Feedback Favors the Generalization of Neural ODEs

ICLR 2025oral

The well-known generalization problem hinders the application of artificial neural networks in continuous-time prediction tasks with varying latent dynamics. In sharp contrast, biological systems can neatly adapt to evolving environments benefiting from real-time feedback mechanisms. Inspired by the…

Cited by 1SourcePDFScholar
2025

GERA: Geometric Embedding for Efficient Point Registration Analysis

ICRA 2025

Point cloud registration aims to provide estimated transformations to align point clouds, which plays a crucial role in pose estimation of various navigation systems, such as surgical guidance systems and autonomous vehicles. Despite the impressive performance of recent models on benchmark datasets,

Cited by 3SourceScholar
2025

Identifying and Mitigating Social Bias Knowledge in Language Models

NAACL 2025findings

Generating fair and accurate predictions plays a pivotal role in deploying pre-trained language models (PLMs) in the real world. However, existing debiasing methods may inevitably generate incorrect or nonsensical predictions as they are designed and evaluated to achieve parity across different soci…

Cited by 1SourcePDFScholar
2025

Intend to Move: A Multimodal Dataset for Intention-Aware Human Motion Understanding

NeurIPS 2025poster

Human motion is inherently intentional, yet most motion modeling paradigms focus on low-level kinematics, overlooking the semantic and causal factors that drive behavior. Existing datasets further limit progress: they capture short, decontextualized actions in static scenes, providing little groundi…

Cited by 0SourceScholar
2025

Unsupervised UAV 3D Trajectories Estimation with Sparse Point Clouds

ICASSP 2025accepted

Compact UAV systems, while advancing delivery and surveillance, pose significant security challenges due to their small size, which hinders detection by traditional methods. This paper presents a cost-effective, unsupervised UAV detection method using spatial-temporal sequence processing to fuse mul…

Cited by 0SourceScholar
2024

Can We Evaluate Domain Adaptation Models Without Target-Domain Labels?

ICLR 2024poster

Unsupervised domain adaptation (UDA) involves adapting a model trained on a label-rich source domain to an unlabeled target domain. However, in real-world scenarios, the absence of target-domain labels makes it challenging to evaluate the performance of UDA models. Furthermore, prevailing UDA method…

Cited by 13SourcePDFScholar
2024

Fully-Connected Spatial-Temporal Graph for Multivariate Time-Series Data

AAAI 2024technical

Multivariate Time-Series (MTS) data is crucial in various application fields. With its sequential and multi-source (multiple sensors) properties, MTS data inherently exhibits Spatial-Temporal (ST) dependencies, involving temporal correlations between timestamps and spatial correlations between senso…

2024

Graph-Aware Contrasting for Multivariate Time-Series Classification

AAAI 2024technical

Contrastive learning, as a self-supervised learning paradigm, becomes popular for Multivariate Time-Series (MTS) classification. It ensures the consistency across different views of unlabeled samples and then learns effective representations for these samples. Existing contrastive learning methods m…

2024

MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone Threats

ICRA 2024poster

In response to the evolving challenges posed by small unmanned aerial vehicles (UAVs), which possess the potential to transport harmful payloads or independently cause damage, we introduce MMAUD: a comprehensive Multi-Modal Anti-UAV Dataset. MMAUD addresses a critical gap in contemporary threat dete…

Cited by 22SourcecodeScholar
2024

MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation

ICRA 2024poster

Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-i…

Cited by 19SourcecodeScholar
2024

Salient Sparse Visual Odometry With Pose-Only Supervision

RA-L 2024

Visual Odometry (VO) is vital for the navigation of autonomous systems, providing accurate position and orientation estimates at reasonable costs. While traditional VO methods excel in some conditions, they struggle with challenges like variable lighting and motion blur. Deep learning-based VO, thou

Cited by 14SourceScholar
2023

AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian Awareness

IROS 2023poster

In this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications. However, traditional approaches that rely on cameras and LID…

Cited by 9SourcecodeScholar
2023

Augmenting and Aligning Snippets for Few-Shot Video Domain Adaptation

ICCV 2023poster

For video models to be transferred and applied seamlessly across video tasks in varied environments, Video Unsupervised Domain Adaptation (VUDA) has been introduced to improve the robustness and transferability of video models. However, current VUDA methods rely on a vast amount of high-quality unla…

Cited by 7PDFcodeScholar
2023

Divide to Adapt: Mitigating Confirmation Bias for Domain Adaptation of Black-Box Predictors

ICLR 2023top-25%

Domain Adaptation of Black-box Predictors (DABP) aims to learn a model on an unlabeled target domain supervised by a black-box predictor trained on a source domain. It does not require access to both the source-domain data and the predictor parameters, thus addressing the data privacy and portabilit…

2023

Fast Model DeBias with Machine Unlearning

NeurIPS 2023poster

Recent discoveries have revealed that deep neural networks might behave in a biased manner in many real-world scenarios. For instance, deep networks trained on a large-scale face recognition dataset CelebA tend to predict blonde hair for females and black hair for males. Such biases not only jeopard…

Cited by 60SourcePDFScholar
2023

MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensing

NeurIPS 2023poster

4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensi…

2023

Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation

ICCV 2023poster

Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmenta…

Cited by 20PDFScholar
2023

SEnsor Alignment for Multivariate Time-Series Unsupervised Domain Adaptation

AAAI 2023technical

Unsupervised Domain Adaptation (UDA) methods can reduce label dependency by mitigating the feature discrepancy between labeled samples in a source domain and unlabeled samples in a similar yet shifted target domain. Though achieving good performance, these methods are inapplicable for Multivariate T…

2022

Global Context With Discrete Diffusion in Vector Quantised Modelling for Image Generation

CVPR 2022poster

The integration of Vector Quantised Variational AutoEncoder (VQ-VAE) with autoregressive models as generation part has yielded high-quality results on image generation. However, the autoregressive models will strictly follow the progressive scanning order during the sampling phase. This leads the ex…

Cited by 43PDFScholar
2022

Source-Free Video Domain Adaptation by Learning Temporal Consistency for Action Recognition

ECCV 2022poster

"Video-based Unsupervised Domain Adaptation (VUDA) methods improve the robustness of video models, enabling them to be applied to action recognition tasks across different environments. However, these methods require constant access to source data during the adaptation process. Yet in many real-worl…

2021

Deep Reinforcement Learning Boosted Partial Domain Adaptation

IJCAI 2021poster

Domain adaptation is critical for learning transferable features that effectively reduce the distribution difference among domains. In the era of big data, the availability of large-scale labeled datasets motivates partial domain adaptation (PDA) which deals with adaptation from large source domains…

Cited by 7SourcePDFScholar
2021

Partial Video Domain Adaptation With Partial Adversarial Temporal Attentive Network

ICCV 2021poster

Partial Domain Adaptation (PDA) is a practical and general domain adaptation scenario, which relaxes the fully shared label space assumption such that the source label space subsumes the target one. The key challenge of PDA is the issue of negative transfer caused by source-only classes. For videos,…

Cited by 36PDFcodeScholar
2020

Mind the Discriminability: Asymmetric Adversarial Domain Adaptation

ECCV 2020poster

Adversarial domain adaptation has made tremendous success by learning domain-invariant feature representations. However, conventional adversarial training pushes two domains together and brings uncertainty to feature learning, which deteriorates the discriminability in the target domain. In this pap…

Cited by 59SourcePDFScholar
2020

Suppressing Mislabeled Data via Grouping and Self-Attention

ECCV 2020poster

Deep networks achieve excellent results on large-scale clean data but degrade significantly when learning from noisy labels. To suppressing the impact of mislabeled data, this paper proposes a conceptually simple yet efficient training block, termed as Attentive Feature Mixup (AFM), which allows pay…

2020

Suppressing Uncertainties for Large-Scale Facial Expression Recognition

CVPR 2020poster

Annotating a qualitative large-scale facial expression dataset is extremely difficult due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. These uncertainties suspend the progress of large-scale Facial Expression Recognitio…

Cited by 783PDFcodeScholar