← Search

Yuhao Chen

15 accepted papers

2026

IVISION-2DCD: A Long-Term Change Detection Dataset for Large-Scale Outdoor Construction Monitoring

ICRA 2026poster

Automation in construction is essential for reducing costs and human errors in large-scale projects. We approach the construction progress monitoring from the aspect of detecting changes in construction sites. As construction buildings continue to evolve in geometry and appearance over time, change …

Cited by 0Scholar
2026

Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning

AAAI 2026technical

Aiming to identify precise evidence sources from visual documents, visual evidence attribution for visual document retrieval–augmented generation (VD-RAG) ensures reliable and verifiable predictions from vision-language models (VLMs) in multimodal question answering. Most existing methods adopt end-

Cited by 0SourcePDFScholar
2026

MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding

AAAI 2026technical

Despite the rapid progress of Vision-Language Models (VLMs), their capabilities are inadequately assessed by existing benchmarks, which are predominantly English-centric, feature simplistic layouts, and support limited tasks. Consequently, they fail to evaluate model performance for Visually Rich Do

Cited by 0SourcePDFScholar
2025

Anima2: Cross-Species Animal Animation through Image-to-Video Synthesis with Subject Alignment

ICASSP 2025accepted

Recent video editing advancements rely on accurate pose sequences to animate human actors. However, these efforts are not suitable for cross-species animation due to pose misalignment between species (for example, the poses of a cat differ greatly from that of a pig due to their distinct body struct…

Cited by 0SourceScholar
2025

Asymmetric Decision-Making in Online Knowledge Distillation: Unifying Consensus and Divergence

ICML 2025poster

Online Knowledge Distillation (OKD) methods represent a streamlined, one-stage distillation training process that obviates the necessity of transferring knowledge from a pretrained teacher network to a more compact student network. In contrast to existing logits-based OKD methods, this paper present…

Cited by 0SourcePDFScholar
2025

LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding

NeurIPS 2025poster

Diffusion transformers (DiTs) struggle to generate images at resolutions higher than their training resolutions. The primary obstacle is that the explicit positional encodings (PE), such as RoPE, need extrapolating to unseen positions which degrades performance when the inference resolution differs…

Cited by 0SourcecodeScholar
2025

MGSO: Monocular Real-Time Photometric SLAM with Efficient 3D Gaussian Splatting

ICRA 2025

Real-time SLAM with dense 3D mapping is computationally challenging, especially on resource-limited devices. The recent development of 3D Gaussian Splatting (3DGS) offers a promising approach for real-time dense 3D reconstruction. However, existing 3DGS-based SLAM systems struggle to balance hardwar

Cited by 7SourceScholar
2025

Streamlining the Collaborative Chain of Models into A Single Forward Pass in Generation-Based Tasks

ACL 2025finding

In Retrieval-Augmented Generation (RAG) and agent-based frameworks, the “Chain of Models” approach is widely used, where multiple specialized models work sequentially on distinct sub-tasks. This approach is effective but increases resource demands as each model must be deployed separately. Recent ad…

2025

Think Wider, Detect Sharper: Reinforced Reference Coverage for Document-Level Self-Contradiction Detection

EMNLP 2025

Detecting self-contradictions within documents is a challenging task for ensuring textual coherence and reliability. While large language models (LLMs) have advanced in many natural language understanding tasks, document-level self-contradiction detection (DSCD) remains insufficiently studied. Recen

2024

HiDiffusion: Unlocking Higher-Resolution Creativity and Efficiency in Pretrained Diffusion Models

ECCV 2024poster

"Diffusion models have become a mainstream approach for high-resolution image synthesis. However, directly generating higher-resolution images from pretrained diffusion models will encounter unreasonable object duplication and exponentially increase the generation time. In this paper, we discover th…

Cited by 5SourcePDFScholar
2024

Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint Prediction

NeurIPS 2024poster

Accurate estimation of human pose and the pose of interacting objects, like a hockey stick, is crucial for action recognition and performance analysis, particularly in sports. Existing methods capture the object along with the human in the bounding boxes, assuming all keypoints are visible within th…

Cited by 0SourcePDFScholar
2023

Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data

CVPR 2023poster

Semi-supervised learning (SSL) has attracted enormous attention due to its vast potential of mitigating the dependence on large labeled datasets. The latest methods (e.g., FixMatch) use a combination of consistency regularization and pseudo-labeling to achieve remarkable successes. However, these me…

2021

Low Resolution Information Also Matters: Learning Multi-Resolution Representations for Person Re-Identification

IJCAI 2021poster

As a prevailing task in video surveillance and forensics field, person re-identification (re-ID) aims to match person images captured from non-overlapped cameras. In unconstrained scenarios, person images often suffer from the resolution mismatch problem, i.e., Cross-Resolution Person Re-ID. To over…

Cited by 30SourcePDFScholar