← Search

Yuchen Wang

28 accepted papers

2026

A Generalizable Physics-Guided Causal Model for Trajectory Prediction in Autonomous Driving

ICRA 2026poster

Trajectory prediction for traffic agents is critical for safe autonomous driving. However, achieving effective zero-shot generalization in previously unseen domains remains a significant challenge. Motivated by the consistent nature of kinematics across diverse domains, we aim to incorporate domain-…

2026

Beyond Single Transactions: D-EMAML---Dual-Edge Motif Neural Networks for Enhanced Anti-Money Laundering Detection

AAAI 2026technical

Anti-money laundering (AML) detection is of vital importance in financial risk control. Although Graph Neural Networks (GNN) have yielded promising results, existing motif-based approaches primarily focus on node anomaly detection on simple graphs, which hinders the direct identification of anomalou

Cited by 0SourcePDFScholar
2026

Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language Models

AAAI 2026technical

Dropout is a widely used regularization technique which improves the generalization ability of a model by randomly dropping neurons. In light of this, we propose Dropout Prompt Learning, which aims for applying dropout to improve the robustness of the vision-language models. Different from the vanil

Cited by 0SourcePDFScholar
2026

Exploring Data-Free LoRA Transferability for Video Diffusion Models

ICML 2026poster

Video diffusion models leveraging step distillation or causal distillation have achieved remarkable performance. However, adapting existing LoRAs to these variants remains a critical challenge due to weight space mismatches. We observe that direct application leads to style degradation and structura…

Cited by 0SourceScholar
2026

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

CVPR 2026

In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time or resource-constrained applications.Visual token pruning is a promising strategy for reducing the cost of MLLM inferen

Cited by 0SourcecodeScholar
2026

Modeling Trend Dynamics with Variational Neural ODEs for Information Popularity Prediction

AAAI 2026technical

Predicting the future popularity of information in online social networks is a crucial yet challenging task, due to the complex spatiotemporal dynamics underlying information diffusion. Existing methods typically use structural or sequential patterns within the observation window as direct inputs fo

Cited by 0SourcePDFScholar
2026

Robust Decentralized Multi-armed Bandits: From Corruption-Resilience to Byzantine-Resilience

AAAI 2026technical

Decentralized cooperative multi-agent multi-armed bandits (DeCMA2B) considers how multiple agents collaborate in a decentralized multi-armed bandit setting. Though this problem has been extensively studied in previous work, most existing methods remain susceptible to various adversarial attacks. In

Cited by 0SourcePDFScholar
2026

Towards Real-Time Neutral Atom Array Assembly via Unsupervised Hologram Generation and Path Optimization

AAAI 2026technical

The rapid and reliable assembly of defect-free atom arrays poses a fundamental challenge for neutral atom quantum computing. While parallel rearrangement methods using spatial light modulators show promise, they suffer from significant overhead in two sub-tasks: atom-site matching and hologram gener

Cited by 0SourcePDFScholar
2026

WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems

ICML 2026spotlight

Trajectory world models play a crucial role in robotic dynamics learning, planning, and control. While recent works have explored trajectory world models for diverse robotic systems, they struggle to scale to a large number of distinct system dynamics and overlook domain knowledge of physical struct…

Cited by 0SourcecodeScholar
2025

A Generalizable Physics-Enhanced State Space Model for Long-Term Dynamics Forecasting in Complex Environments

ICML 2025poster

This work aims to address the problem of long-term dynamic forecasting in complex environments where data are noisy and irregularly sampled. While recent studies have introduced some methods to improve prediction performance, these approaches still face a significant challenge in handling long-term…

Cited by 0SourcePDFScholar
2025

A Generalized Diffusion Framework with Learnable Propagation Dynamics for Source Localization

IJCAI 2025

Source localization has been widely studied in recent years due to its crucial role in controlling the spread of harmful information. Existing methods only achieve satisfactory performance within a specific propagation model, which restricts their applicability and generalizability across different

2025

Accelerating Neural ODEs: A Variational Formulation-based Approach

ICLR 2025poster

Neural Ordinary Differential Equations (Neural ODEs or NODEs) excel at modeling continuous dynamical systems from observational data, especially when the data is irregularly sampled. However, existing training methods predominantly rely on numerical ODE solvers, which are time-consuming and prone to…

2025

Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image Denoising

CVPR 2025poster

Existing single-image denoising algorithms often struggle to restore details when dealing with complex noisy images. The introduction of near-infrared (NIR) images offers new possibilities for RGB image denoising. However, due to the inconsistency between NIR and RGB images, the existing works still…

2025

Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin

ICML 2025poster

Adapting vision-language models (VLMs) to downstream tasks with pseudolabels has gained increasing attention. A major obstacle is that the pseudolabels generated by VLMs tend to be imbalanced, leading to inferior performance. While existing methods have explored various strategies to address this,…

2025

Learning Neural Jump Stochastic Differential Equations with Latent Graph for Multivariate Temporal Point Processes

IJCAI 2025

Multivariate Temporal Point Processes (MTPPs) play an important role in diverse domains such as social networks and finance for predicting event sequence data. In recent years, MTPPs based on Ordinary Differential Equations (ODEs) and Stochastic Differential Equations (SDEs) have demonstrated their

2025

Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual Processing

AAAI 2025technical

Event cameras encode visual information by generating asynchronous and sparse event streams, which hold great potential for low latency and low power consumption. Despite many successful implementations of event camera-based applications, most of them accumulate the events into frames and then utili…

Cited by 0SourcePDFScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Fusing Personal and Environmental Cues for Identification and Segmentation of First-Person Camera Wearers in Third-Person Views

CVPR 2024poster

As wearable cameras become more popular an important question emerges: how to identify camera wearers within the perspective of conventional static cameras. The drastic difference between first-person (egocentric) and third-person (exocentric) camera views makes this a challenging task. We present P…

2023

Audio-Driven High Definetion and Lip-Synchronized Talking Face Generation Based on Face Reenactment

ICASSP 2023accepted

Generating audio-driven photo-realistic talking face has received intensive attention due to its ability to bring more new human-computer interaction experiences. However, previous works struggled to balance high definition, lip synchronization, and low customization costs, which would degrade the u…

Cited by 0SourceScholar
2023

Spatial-Temporal Self-Attention for Asynchronous Spiking Neural Networks

IJCAI 2023poster

The brain-inspired spiking neural networks (SNNs) are receiving increasing attention due to their asynchronous event-driven characteristics and low power consumption. As attention mechanisms recently become an indispensable part of sequence dependence modeling, the combination of SNNs and attention…

2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Prediction of Whole-Body Velocity and Direction From Local Leg Joint Movements in Insect Walking via LSTM Neural Networks

RA-L 2022

Extracting motion information from videos is important for quantifying data from behavioral experiments to deepen the understanding of generation mechanisms of animal behavior. For insect walking, inter-leg coordination plays a crucial role, and the thorax-coxa (ThC) and femur-tibia (FTi) joint moti

Cited by 6SourceScholar
2022

Signed Neuron with Memory: Towards Simple, Accurate and High-Efficient ANN-SNN Conversion

IJCAI 2022poster

Spiking Neural Networks (SNNs) are receiving increasing attention due to their biological plausibility and the potential for ultra-low-power event-driven neuromorphic hardware implementation. Due to the complex temporal dynamics and discontinuity of spikes, training SNNs directly usually suffers fro…

2021

Deep Spiking Neural Network with Neural Oscillation and Spike-Phase Information

AAAI 2021technical

Deep spiking neural network (DSNN) is a promising computational model towards artificial intelligence. It benefits from both the DNNs and SNNs through a hierarchy structure to extract multiple levels of abstraction and the event-driven computational manner to provide ultra-low-power neuromorphic imp…

Cited by 17SourcePDFScholar
2019

Unsupervised Traffic Accident Detection in First-Person Videos

IROS 2019poster

Recognizing abnormal events such as traffic violations and accidents in natural driving scenes is essential for successful autonomous driving and advanced driver assistance systems. However, most work on video anomaly detection suffers from two crucial drawbacks. First, they assume cameras are fixed…

Cited by 209SourcecodeScholar
2018

Design, Modeling and Control of a Solar-Powered Quadcopter

ICRA 2018poster

This paper presents the design, modeling, control, and experimental test of a solar-powered quadcopter to allow for long-endurance missions. We first present the design of a large-scale quadcopter that incorporates solar energy harvesting capabilities. Based on the design results, we built the dynam…

Cited by 43SourceScholar
2018

Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos

ECCV 2018poster

In a world of pervasive cameras, public spaces are often captured from multiple perspectives by cameras of different types, both fixed and mobile. An important problem is to organize these heterogeneous collections of videos by finding connections between them, such as identifying correspondences be…

Cited by 46SourcePDFScholar