← Search

Yubo Dong

8 accepted papers

2026

Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token Selection

CVPR 2026

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet processing long visual token sequences remains computationally expensive. Existing approaches mitigate this cost by reducing image tokens, either by discarding them after the visual encod

Cited by 0SourcecodeScholar
2026

Structured Reasoning for LLMs: A Unified Framework for Efficiency and Explainability

ICLR 2026poster

Recent Large Language Models (LLMs) have made remarkable progress, but they still struggle with complex reasoning tasks such as logical deduction and planning. This is partly because they rely primarily on token-level probability relationships, which limits their ability to reason effectively. In t…

Cited by 0SourcecodeScholar
2025

CLAP: A Closed-Loop Diffusion Transformer Action Foundation Model for Robotic Manipulation

IROS 2025

The development of large Vision-Language-Action (VLA) models has enhanced the robot’s ability to manipulate objects in unseen scenarios based on language instructions. While existing VLAs have demonstrated promise in various scenarios, they still struggle with effective multi-modal data feature extr

Cited by 1SourceScholar
2025

Multi Queue for Unsupervised Person Re-identification

ICASSP 2025accepted

Recently, cluster-based methods have achieved significant success in unsupervised re-ID tasks. The hierarchical clustering algorithm, exemplified by SpCL, has been widely adopted in unsupervised cross-domain adaptation and unsupervised learning. The momentum-based feature update mechanism in SpCL ha…

Cited by 0SourceScholar
2024

VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in Minecraft

ACL 2024findings

In this paper, we aim to evaluate multi-agent systems against complex dependencies, including spatial, causal, and temporal constraints. First, we construct a new benchmark, named VillagerBench, within the Minecraft environment. VillagerBench comprises diverse tasks crafted to test various aspects o…

2023

Residual Degradation Learning Unfolding Framework With Mixing Priors Across Spectral and Spatial for Compressive Spectral Imaging

CVPR 2023poster

To acquire a snapshot spectral image, coded aperture snapshot spectral imaging (CASSI) is proposed. A core problem of the CASSI system is to recover the reliable and fine underlying 3D spectral cube from the 2D measurement. By alternately solving a data subproblem and a prior subproblem, deep unfold…