← Search

Le Wang

58 accepted papers

2026

A Diagnostic Study of Multi-Agent LLMs for Real-World Debates

ICML 2026poster

Multi-agent LLM debates are increasingly deployed in domains such as policy analysis and city planning, where no objective ground truth exists. Despite this, debate quality is typically evaluated using outcome-based proxies such as LLM-as-judge scores that provide little insight into whether meaning…

Cited by 0SourceScholar
2026

AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions

CVPR 2026

The integration of vision-language models (VLMs) is driving a new generation of embodied agents capable of operating in human-centered environments. However, as deployment expands, these systems face growing safety risks, particularly when executing hazardous instructions. Current safety evaluation

Cited by 29SourceScholar
2026

AUDIOGEN-OMNI: A UNIFIED MULTIMODAL DIFFUSION TRANSFORMER FOR VIDEO-SYNCHRONIZED AUDIO, SPEECH, AND SONG GENERATION

ICASSP 2026poster

We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and song coherently synchronized with the input video. AudioGen-Omni introduces a novel joint training paradigm that seamlessly integrates large-scale…

Cited by 0SourcePDFScholar
2026

Active Perception Driven Tactile Modeling of Deformable Objects for Robots

RA-L 2026

Despite extensive research on tactile sensing, exploiting it to model and interpret the external world remains largely underexplored. We propose a unified framework that leverages vision-based tactile sensors to construct fine-grained object models, enhancing situational understanding. This frame wo

Cited by 0SourceScholar
2026

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

ICRA 2026poster

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT ) tend to treat multi-view features equally and directly concatenate them for policy learning. How ever, it will introduce redundant visual information and bring hig…

2026

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs

AAAI 2026technical

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of e

Cited by 0SourcePDFScholar
2026

Spatial Matters: Position-Guided 3D Referring Expression Segmentation

CVPR 2026

3D Referring Expression segmentation (3D-RES) is an emerging field that segments 3D objects in point cloud scenes based on given referring expressions. Although existing methods have achieved substantial progress, they primarily focus on semantic cues and often overlook spatial relations, which are

Cited by 0SourcecodeScholar
2025

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

RA-L 2025

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT [1]) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it will introduce redundant visual information and bring h

Cited by 8SourceScholar
2025

BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration

EMNLP 2025

Mixture-of-Experts (MoE) models have become increasingly powerful in multimodal learning by enabling modular specialization across modalities. However, their effectiveness remains unclear when additional modalities introduce more noise than complementary information. Existing approaches, such as the

2025

Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding

AAAI 2025technical

Temporal sentence grounding is a challenging task that aims to localize the moment spans relevant to a language description. Although recent DETR-based models have achieved notable progress by leveraging multiple learnable moment queries, they suffer from overlapped and redundant proposals, leading…

2025

DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation

NeurIPS 2025poster

Learning generalizable robotic manipulation policies remains a key challenge due to the scarcity of diverse real-world training data. While recent approaches have attempted to mitigate this through self-supervised representation learning, most either rely on 2D vision pretraining paradigms such as m…

Cited by 0SourceScholar
2025

Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction

ICCV 2025poster

Sufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing existing methods to face the trade-off between interaction completeness and computational efficiency. To address this limit…

2025

FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation

ACL 2025long

Large language models (LLMs) augmented with retrieval systems have demonstrated significant potential in handling knowledge-intensive tasks. However, these models often struggle with unfaithfulness issues, generating outputs that either ignore the retrieved context or inconsistently blend it with th…

2025

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

CVPR 2025poster

Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during in…

Cited by 0SourcePDFScholar
2025

Moment Quantization for Video Temporal Grounding

ICCV 2025poster

Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation…

2025

PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation

CVPR 2025poster

Robotic manipulation based on visual observations and natural language instructions is a long-standing challenge in robotics. Yet prevailing approaches model action distribution by adopting explicit or implicit representations, which often struggle to achieve a trade-off between accuracy and efficie…

Cited by 0SourcePDFScholar
2025

RefDetector: A Simple Yet Effective Matching-based Method for Referring Expression Comprehension

AAAI 2025technical

Despite the rapid and substantial advancements in object detection, it continues to face limitations imposed by pre-defined category sets. Current methods for visual grounding primarily focus on how to better leverage the visual backbone to generate text-tailored visual features, which may require a…

Cited by 0SourcePDFScholar
2025

SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World Models

NeurIPS 2025poster

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and…

Cited by 0SourceScholar
2025

Semantic Graph Embedded Energy Minimization Learning for Scene Graph Generation

ICASSP 2025accepted

The performance of current scene graph generation models is affected by training with cross-entropy loss, exacerbating the problem of prediction bias stemming from biased training data. Energy-based model adopts a learning method for joint image and scene graph to alleviate this challenge. However,…

Cited by 0SourceScholar
2025

Towards Precise Embodied Dialogue Localization via Causality Guided Diffusion

CVPR 2025poster

Embodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently ex…

Cited by 0SourcePDFScholar
2025

VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible Regions

ICCV 2025poster

This paper introduces VisHall3D, a novel two-stage framework for monocular semantic scene completion that aims to address the issues of feature entanglement and geometric inconsistency prevalent in existing methods. VisHall3D decomposes the scene completion task into two stages: reconstructing the v…

2024

Analysis-by-Synthesis Transformer for Single-View 3D Reconstruction

ECCV 2024poster

"Deep learning approaches have made significant success in single-view 3D reconstruction, but they often rely on expensive 3D annotations for training. Recent efforts tackle this challenge by adopting an analysis-by-synthesis paradigm to learn 3D reconstruction with only 2D annotations. However, exi…

2024

PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation

ECCV 2024poster

"Semi-supervised learning has emerged as a widely adopted technique in the field of medical image segmentation. The existing works either focuses on the construction of consistency constraints or the generation of pseudo labels to provide high-quality supervisory signals, whose main challenge mainly…

2024

R3D-AD: Reconstruction via Diffusion for 3D Anomaly Detection

ECCV 2024poster

"3D anomaly detection plays a crucial role in monitoring parts for localized inherent defects in precision manufacturing. Embedding-based and reconstruction-based approaches are among the most popular and successful methods. However, there are two major challenges to the practical application of the…

Cited by 12SourcePDFScholar
2024

Referencing Where to Focus: Improving Visual Grounding with Referential Query

NeurIPS 2024poster

Visual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional effort…

Cited by 1SourcePDFScholar
2024

Temporal Correlation Vision Transformer for Video Person Re-Identification

AAAI 2024technical

Video Person Re-Identification (Re-ID) is a task of retrieving persons from multi-camera surveillance systems. Despite the progress made in leveraging spatio-temporal information in videos, occlusion in dense crowds still hinders further progress. To address this issue, we propose a Temporal Correla…

Cited by 5SourcePDFScholar
2024

Towards Generalizable Multi-Object Tracking

CVPR 2024poster

Multi-Object Tracking (MOT) encompasses various tracking scenarios each characterized by unique traits. Effective trackers should demonstrate a high degree of generalizability across diverse scenarios. However existing trackers struggle to accommodate all aspects or necessitate hypothesis and experi…

2023

Hierarchical Motion Planning for Autonomous Vehicles in Unstructured Dynamic Environments

RA-L 2023

This letter presents a hierarchical motion planner for generating smooth and feasible trajectories for autonomous vehicles in unstructured environments with static and moving obstacles. The framework enables real-time computation by progressively shrinking the solution space. First, a graph searcher

Cited by 21SourceScholar
2023

Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action Localization

ICCV 2023poster

Semi-Supervised Temporal Action Localization (SS-TAL) aims to improve the generalization ability of action detectors with large-scale unlabeled videos. Albeit the recent advancement, one of the major challenges still remains: noisy pseudo labels hinder efficient learning on abundant unlabeled videos…

Cited by 10PDFcodeScholar
2023

MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object Tracking

CVPR 2023poster

The main challenge of Multi-Object Tracking (MOT) lies in maintaining a continuous trajectory for each target. Existing methods often learn reliable motion patterns to match the same target between adjacent frames and discriminative appearance features to re-identify the lost targets after a long pe…

2023

Multi-Stream Representation Learning for Pedestrian Trajectory Prediction

AAAI 2023technical

Forecasting the future trajectory of pedestrians is an important task in computer vision with a range of applications, from security cameras to autonomous driving. It is very challenging because pedestrians not only move individually across time but also interact spatially, and the spatial and tempo…

2023

Parallel Attention Interaction Network for Few-Shot Skeleton-Based Action Recognition

ICCV 2023poster

Learning discriminative features from very few labeled samples to identify novel classes has received increasing attention in skeleton-based action recognition. Existing works aim to learn action-specific embeddings by exploiting either intra-skeleton or inter-skeleton spatial associations, which ma…

Cited by 11PDFScholar
2023

Progressive Backdoor Erasing via Connecting Backdoor and Adversarial Attacks

CVPR 2023poster

Deep neural networks (DNNs) are known to be vulnerable to both backdoor attacks as well as adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct problems and solved separately, since they belong to training-time and inference-time attacks respectively. H…

Cited by 30SourcePDFScholar
2022

Complementary Attention Gated Network for Pedestrian Trajectory Prediction

AAAI 2022technical

Pedestrian trajectory prediction is crucial in many practical applications due to the diversity of pedestrian movements, such as social interactions and individual motion behaviors. With similar observable trajectories and social environments, different pedestrians may make completely different futu…

2022

Improving Robustness of Language Models from a Geometry-aware Perspective

ACL 2022findings

Recent studies have found that removing the norm-bounded projection and increasing search steps in adversarial training can significantly improve robustness. However, we observe that a too large number of search steps can hurt accuracy. We aim to obtain strong robustness efficiently using fewer step…

2022

Learning Disentangled Classification and Localization Representations for Temporal Action Localization

AAAI 2022technical

A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that thi…

Cited by 20SourcePDFScholar
2022

Learning To Refactor Action and Co-Occurrence Features for Temporal Action Localization

CVPR 2022poster

The main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer f…

Cited by 57PDFScholar
2022

Social Interpretable Tree for Pedestrian Trajectory Prediction

AAAI 2022technical

Understanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on th…

2021

ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization

AAAI 2021technical

The object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foregroun…

Cited by 86SourcePDFScholar
2021

Enriching Local and Global Contexts for Temporal Action Localization

ICCV 2021poster

Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by…

Cited by 148PDFcodeScholar
2021

Meta Pairwise Relationship Distillation for Unsupervised Person Re-Identification

ICCV 2021poster

Unsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccur…

Cited by 51PDFcodeScholar
2021

Practical Relative Order Attack in Deep Ranking

ICCV 2021poster

Recent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains unde…

Cited by 21PDFcodeScholar
2021

SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory Prediction

CVPR 2021poster

Pedestrian trajectory prediction is a key technology in autopilot, which remains to be very challenging due to complex interactions between pedestrians. However, previous works based on dense undirected interaction suffer from modeling superfluous interactions and neglect of trajectory motion tenden…

Cited by 327PDFcodeScholar
2021

Unlimited Neighborhood Interaction for Heterogeneous Trajectory Prediction

ICCV 2021poster

Understanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-loca…

Cited by 35PDFcodeScholar
2021

Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and Context

AAAI 2021technical

Weakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classificat…

Cited by 32SourcePDFScholar
2020

A Comprehensive Study of Weight Sharing in Graph Networks for 3D Human Pose Estimation

ECCV 2020poster

Graph convolutional networks (GCNs) have been applied to 3D human pose estimation (HPE) from 2D body joint detections and have shown encouraging performance. One limitation of the vanilla graph convolution is that it models the relationships between neighboring nodes via a shared weight matrix. This…

Cited by 180SourcePDFScholar
2020

Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization

ECCV 2020poster

Weakly-supervised Temporal Action Localization (W-TAL) aims to classify and localize all action instances in an untrimmed video under only video-level supervision. However, without frame-level annotations, it is challenging for W-TAL methods to identify false positive action proposals and generate a…

2019

Compressing Unknown Images With Product Quantizer for Efficient Zero-Shot Classification

CVPR 2019poster

For Zero-Shot Learning (ZSL), the Nearest Neighbor (NN) search is generally conducted for classification, which may cause unacceptable computational complexity for large-scale datasets. To compress zero-shot classes by the trained quantizer for efficient search, it tends to induce large quantization…

Cited by 48PDFScholar
2019

Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks

ICCV 2019poster

Weakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video…

Cited by 135PDFcodeScholar
2017

ER3: A Unified Framework for Event Retrieval, Recognition and Recounting

CVPR 2017poster

We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames an…

Cited by 28PDFScholar