← Search

Yong Du

15 accepted papers

2026

Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual Refinement

AAAI 2026technical

Chain-of-Thought prompting has remarkably advanced LLM reasoning by generating explicit step-by-step tokens, yet its discrete nature inherently limits expressiveness and efficiency, struggling with abstract, ambiguous, or semantically divergent cognition beyond linguistic tokens. Latent reasoning of

Cited by 0SourcePDFScholar
2026

MetaphorVU: Towards Metaphorical Video Understanding

ICML 2026spotlight

Metaphorical videos are prevalent across various real-world scenarios to convey complex ideas, and understanding them typically requires high-order cognitive capabilities. The lack of systematic studies on metaphorical video understanding not only constrains the real-world applicability of MLLMs but…

Cited by 0SourceScholar
2026

NimbusGS: Unified 3D Scene Reconstruction under Hybrid Weather

CVPR 2026

We present NimbusGS, a unified framework for reconstructing high-quality 3D scenes from degraded multi-view inputs captured under diverse and mixed adverse weather conditions. Unlike existing methods that target specific weather types, NimbusGS addresses the broader challenge of generalization by mo

Cited by 0SourcecodeScholar
2026

Test-Time Reinforcement Learning for GUI Grounding via Region Consistency

AAAI 2026technical

Graphical User Interface (GUI) grounding, the task of mapping natural language instructions to precise screen coordinates, is fundamental to autonomous GUI agents. While existing methods achieve strong performance through extensive supervised training or reinforcement learning with labeled rewards,

Cited by 0SourcePDFScholar
2025

Cross-Subject Mind Decoding from Inaccurate Representations

ICCV 2025poster

Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings genera…

Cited by 0SourcePDFScholar
2025

NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian Splatting

CVPR 2025highlight

Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-…

2025

OmniVTON: Training-Free Universal Virtual Try-On

ICCV 2025poster

Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unifi…

2025

PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem Equilibrium

AAAI 2025technical

Personalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances…

2024

D3still: Decoupled Differential Distillation for Asymmetric Image Retrieval

CVPR 2024poster

Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However these one-to-one constraint approaches often fail to maintain retrieval order consistency especially when the query network has limited repr…

2023

Curricular Contrastive Regularization for Physics-Aware Single Image Dehazing

CVPR 2023poster

Considering the ill-posed nature, contrastive regularization has been developed for single image dehazing, introducing the information from negative images as a lower bound. However, the contrastive samples are nonconsensual, as the negatives are usually represented distantly from the clear (i.e., p…

2022

Editing Out-of-Domain GAN Inversion via Differential Activations

ECCV 2022poster

"Despite the demonstrated editing capacity in the latent space of a pretrained GAN model, inverting real-world images is stuck in a dilemma that the reconstruction cannot be faithful to the original input. The main reason for this is that the distributions between training and real-world data are mi…

2021

From Continuity to Editability: Inverting GANs With Consecutive Images

ICCV 2021poster

Existing GAN inversion methods are stuck in a paradox that the inverted codes can either achieve high-fidelity reconstruction, or retain the editing capability. Having only one of them clearly cannot realize real image editing. In this paper, we resolve this paradox by introducing consecutive images…

Cited by 46PDFcodeScholar
2021

Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip Reading

CVPR 2021poster

Lip reading aims to predict the spoken sentences from silent lip videos. Due to the fact that such a vision task usually performs worse than its counterpart speech recognition, one potential scheme is to distill knowledge from a teacher pretrained by audio signals. However, the latent domain gap bet…

Cited by 86PDFScholar