← Search

Yadong Li

17 accepted papers

2026

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

ICLR 2026poster

The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models. A vision tokenizer trained for reconstruction excels at capturing low-level visual appearance, making it well-suited for…

Cited by 0SourcecodeScholar
2026

MAS-Architect: Declarative Multi-Agent System Design via Separation of Concerns

ICML 2026poster

The Automated Design of Multi-Agent Systems (Auto-MAS) has emerged as a promising framework for addressing complex reasoning tasks. However, existing approaches often suffer from structural rigidity and entangle the design of system topology with the implementation of individual agents. To overcome …

Cited by 0SourceScholar
2026

META: Meta Evolution of Tool Trajectory Adaptation for Long-Video Understanding

CVPR 2026

Long-video understanding remains challenging due to extreme temporal redundancy, sparse yet decisive events, and the instability of long-horizon reasoning in visual-language models (VLMs). Existing agent-based methods invoke external micro-tools but remain static, repeatedly rebuilding long chains o

Cited by 0SourceScholar
2026

UCPO: Uncertainty-Aware Policy Optimization

ICML 2026poster

The key to building trustworthy Large Language Models (LLMs) lies in endowing them with inherent uncertainty expression capabilities to mitigate the hallucinations that restrict their high-stakes applications. However, existing RL paradigms such as GRPO often suffer from Advantage Bias due to binary…

Cited by 0SourceScholar
2026

VRCLIP: Multimodal Canonical Correlation Alignment for CLIP-Driven Vision-Radio Person Re-Identification

CVPR 2026

Multimodal person Re-IDentification (ReID) aims to reliably associate specific individuals by utilizing complementary information from heterogeneous modalities. In contrast, low-frequency radio frequency (RF) signals, with their superior penetration capability and illumination invariance, provide id

Cited by 0SourceScholar
2025

MHBench: Demystifying Motion Hallucination in VideoLLMs

AAAI 2025technical

Similar to Language or Image LLMs, VideoLLMs are also plagued by hallucination issues. Hallucinations in videos not only manifest in the spatial dimension regarding the perception of the existence of visual objects (static) but also the temporal dimension influencing the perception of actions and ev…

2025

MoLE:Decoding by Mixture of Layer Experts Alleviates Hallucination in Large Vision-Language Models

AAAI 2025technical

Recent advancements in Large Vision-Language Models (LVLMs) highlight their ability to integrate and process multi-modal information. However, hallucinations—where generated content is inconsistent with input vision and instructions—remain a challenge. In this paper, we analyze LVLMs' layer-wise dec…

2025

RFMamba: Frequency-Aware State Space Model for RF-Based Human-Centric Perception

ICLR 2025poster

Human-centric perception with radio frequency (RF) signals has recently entered a new era of end-to-end processing with Transformers. Considering the long-sequence nature of RF signals, the State Space Model (SSM) has emerged as a superior alternative due to its effective long-sequence modeling and…

Cited by 1SourcePDFScholar
2024

Diffradar: High-Quality Mmwave Radar Perception With Diffusion Probabilistic Model

ICASSP 2024accepted

Millimeter-wave (mmWave) radar has gained increasing attention in environmental perception due to its robustness under low-light conditions. However, existing methods fail to address the challenges of multipath interference and low angle resolution. In this paper, we introduce DiffRadar which levera…

Cited by 0SourceScholar
2024

Enabling Orientation-Free Mmwave-Based Vital Sign Sensing with Multi-Domain Signal Analysis

ICASSP 2024accepted

Contactless vital signs estimation using mmWave radar has gained significant attention. However, existing studies are built upon the radar being directed facing the thorax to capture fine-grained vital signs, ignoring the angle variation between the radar and thorax in practical deployment. In this…

Cited by 0SourceScholar
2024

IFNet: Imaging and Focusing Network for handheld mmWave Devices

ICASSP 2024accepted

Recent advancements have showcased the potential of hand-held millimeter-wave (mmWave) imaging, which applies synthetic aperture radar (SAR) principles in portable settings. However, existing studies addressing handheld motion errors either rely on costly tracking devices or employ simplified imagin…

Cited by 0SourceScholar
2024

Learning-Based Tracking-before-Detect for RF-Based Unconstrained Indoor Human Tracking

IJCAI 2024poster

Existing efforts on human tracking using wireless signal are primarily focused on constrained scenarios with only a few individuals in empty spaces. However, in practical unconstrained scenarios with severe interference and attenuation, accurate multi-person tracking has been intractable. In this pa…

Cited by 0SourcePDFScholar
2024

SIMFALL: A Data Generator for RF-Based Fall Detection

ICASSP 2024accepted

Fall detection using Radio Frequency (RF) signals with deep learning has exhibited significant promise in recent years. However, the costly collection of RF data with falls has hampered the performance of existing methods. While there has been approaches which can generate RF signals using various s…

Cited by 0SourceScholar
2023

Two-Stream Networks for Weakly-Supervised Temporal Action Localization With Semantic-Aware Mechanisms

CVPR 2023poster

Weakly-supervised temporal action localization aims to detect action boundaries in untrimmed videos with only video-level annotations. Most existing schemes detect temporal regions that are most responsive to video-level classification, but they overlook the semantic consistency between frames. In t…

Cited by 29SourcePDFScholar
2022

Real-Time Fall Detection Using Mmwave Radar

ICASSP 2022accepted

Fall is a severe health threat for elders’ health care. While existing systems could achieve promising performance under specific scenarios, the required computing resources are usually not affordable, which is not applicable for real-time detection. In this paper, we propose mmFall, a real time fal…

Cited by 0SourceScholar
2020

Parsing-Based View-Aware Embedding Network for Vehicle Re-Identification

CVPR 2020poster

Vehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance distance caused by different views and the subtle inter-instance discrepancy caused by similar vehicles. In this paper, we pr…

Cited by 255PDFcodeScholar
2019

Vehicle Re-Identification With Viewpoint-Aware Metric Learning

ICCV 2019poster

This paper considers vehicle re-identification (re-ID) problem. The extreme viewpoint variation (up to 180 degrees) poses great challenges for existing approaches. Inspired by the behavior in human's recognition process, we propose a novel viewpoint-aware metric learning approach. It learns two metr…

Cited by 253PDFcodeScholar