← Search

You He

22 accepted papers

2026

From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection

CVPR 2026

Zero-shot anomaly detection (ZSAD) aims to identify unseen anomalies without abnormal supervision, which is essential for open-world scenarios. Recent vision-language models such as CLIP enable anomaly reasoning through shared visual-textual embeddings, but existing methods often rely on coarse prom

Cited by 0SourceScholar
2026

From Observations to Events: Event-Aware World Models for Reinforcement Learning

ICLR 2026poster

While model-based reinforcement learning (MBRL) improves sample efficiency by learning world models from raw observations, existing methods struggle to generalize across structurally similar scenes and remain vulnerable to spurious variations such as textures or color shifts. From a cognitive scienc…

Cited by 0SourcecodeScholar
2026

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

AAAI 2026technical

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs)

Cited by 0SourcePDFScholar
2025

ArenaSim: A High-Performance Simulation Platform for Multi-Robot Self-Play Learning

RA-L 2025

In this letter, we introduce <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ArenaSim</i>, a novel simulation platform designed for realistic and efficient self-play learning in multi-robot cooperative-competitive games. Compared to previous simulati

Cited by 0SourceScholar
2025

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

NeurIPS 2025poster

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particul…

Cited by 0SourceScholar
2025

ReNeg: Learning Negative Embedding with Reward Guidance

CVPR 2025highlight

In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In…

2025

Towards Survivability in Complex Motion Scenarios: RGB-Event Object Tracking via Historical Trajectory Prompting

ICRA 2025

Event data has recently emerged as a valuable complement to object tracking, offering dense temporal resolution and a high dynamic range. However, existing RGB-Event trackers struggle with targets exhibiting complex motion trajectories, where RGB features alone fail to provide sufficient discriminat

Cited by 4SourcecodeScholar
2025

TrackFusion: Enhancing Multi-Object Tracking With Temporal Trajectory Modeling and Frame-Integrated Detection

ICASSP 2025accepted

Although MOTIP is the SOTA multi-object tracking method, there are still some issues that limit its performance. First, MOTIP still has defects in temporal information modeling, which leads to the failure to fully utilize the historical information of the tracked target and affects the correlation p…

Cited by 0SourceScholar
2024

Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

CVPR 2024poster

Continual learning can empower vision-language models to continuously acquire new knowledge without the need for access to the entire historical dataset. However mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout lifelong learning and (…

2024

LLMs Can Evolve Continually on Modality for $\mathbb{X}$-Modal Reasoning

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily on extensive modal-specific pretraining and joint-modal tuning, leading to significant computational burdens when expand…

2024

Multi-agent Collaborative Perception via Motion-aware Robust Communication Network

CVPR 2024poster

Collaborative perception allows for information sharing between multiple agents such as vehicles and infrastructure to obtain a comprehensive view of the environment through communication and fusion. Current research on multi-agent collaborative perception systems often assumes ideal communication a…

Cited by 4SourcePDFScholar
2024

ReDiffuser: Reliable Decision-Making Using a Diffuser with Confidence Estimation

ICML 2024poster

The diffusion model has demonstrated impressive performance in offline reinforcement learning. However, non-deterministic sampling in diffusion models can lead to unstable performance. Furthermore, the lack of confidence measurements makes it difficult to evaluate the reliability and trustworthiness…

Cited by 4SourcePDFScholar
2023

Style-Content Metric Learning for Multidomain Remote Sensing Object Recognition

AAAI 2023technical

Previous remote sensing recognition approaches predominantly perform well on the training-testing dataset. However, due to large style discrepancies not only among multidomain datasets but also within a single domain, they suffer from obvious performance degradation when applied to unseen domains. I…

2023

Video Diffusion Models with Local-Global Context Guidance

IJCAI 2023poster

Diffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget, existing methods usually implement conditional diffusion models with an autoregressive inference pipeline, in which th…

2022

United Defocus Blur Detection and Deblurring via Adversarial Promoting Learning

ECCV 2022poster

"Understanding blur from a single defocused image contains two tasks of defocus detection and deblurring. This paper makes the earliest effort to jointly learn both defocus detection and deblurring without using pixel-level defocus detection annotation and paired defocus deblurring ground truth. We…

2020

Unsupervised Video Object Segmentation with Joint Hotspot Tracking

ECCV 2020poster

Object tracking is a well-studied problem in computer vision while identifying salient spots of objects in a video is a less explored direction in the literature. Video eye gaze estimation methods aim to tackle a related task but salient spots in those methods are not bounded by objects and tend to…

2019

CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection

CVPR 2019poster

Detecting salient objects in cluttered scenes is a big challenge. To address this problem, we argue that the model needs to learn discriminative semantic features for salient objects. To this end, we propose to leverage captioning as an auxiliary semantic task to boost salient object detection in c…

Cited by 136PDFScholar
2018

A Bi-Directional Message Passing Model for Salient Object Detection

CVPR 2018poster

Recent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detect…

Cited by 579SourcePDFScholar