← Search

Yingcong Chen

15 accepted papers

2026

ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated exceptional proficiency in standard text extraction, but they encounter significant challenges when confronting real-world implicit text. Such content typically contains malicious information, intentionally concealed through physical deforma…

Cited by 0SourceScholar
2026

ScalingAR: Scaling Confidence for Autoregressive Image Generation

ICML 2026poster

Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregressive (AR) image generation remains largely underexplored. Existing test-time scaling (TTS) methods for visual autoregressive models (VAR) rely on freq…

Cited by 0SourceScholar
2026

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

ICML 2026poster

Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation--a hallmark of human creativity. Current reasoning-augmented paradigms mostly rely on explicit thought processes, where intermediate reasoni…

Cited by 0SourceScholar
2025

Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models

ICRA 2025

Large Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate LLMs with an important representation. To effectively encode

Cited by 23SourceScholar
2025

RhythmGuassian: Repurposing Generalizable Gaussian Model For Remote Physiological Measurement

ICCV 2025poster

Remote Photoplethysmography (rPPG) enables non-contact extraction of physiological signals, providing significant advantages in medical monitoring, emotion recognition, and face anti-spoofing. However, the extraction of reliable rPPG signals is hindered by motion variations in real-world environment…

2024

From Bird’s-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model

ICRA 2024poster

We explore Bird’s-Eye View (BEV) generation, converting a BEV map into its corresponding multi-view street images. Valued for its unified spatial representation aiding multi-sensor fusion, BEV is pivotal for various autonomous driving applications. Creating accurate street-view images from BEV maps…

Cited by 1SourceScholar
2024

LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching

CVPR 2024highlight

The recent advancements in text-to-3D generation mark a significant milestone in generative models unlocking new possibilities for creating imaginative 3D assets across various real-world scenarios. While recent advancements in text-to-3D generation have shown promise they often fall short in render…

2024

MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders

ECCV 2024poster

"Multi-task dense scene understanding, which learns a model for multiple dense prediction tasks, has a wide range of application scenarios. Modeling long-range dependency and enhancing cross-task interactions are crucial to multi-task dense prediction. In this paper, we propose MTMamba, a novel Mamb…

2023

Out-of-Domain GAN Inversion via Invertibility Decomposition for Photo-Realistic Human Face Manipulation

ICCV 2023poster

The fidelity of Generative Adversarial Networks (GAN) inversion is impeded by Out-Of-Domain (OOD) areas (e.g., background, accessories) in the image. Detecting the OOD areas beyond the generation ability of the pre-trained model and blending these regions with the input image can enhance fidelity.…

Cited by 4PDFcodeScholar
2022

DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation

ECCV 2022poster

"Unsupervised domain adaptation in semantic segmentation alleviates the reliance on expensive pixel-wise annotation. It uses a labeled source domain dataset as well as unlabeled target domain images to learn a segmentation network. In this paper, we observe two main issues of existing domain-invaria…

2022

RC-MVSNet: Unsupervised Multi-View Stereo with Neural Rendering

ECCV 2022poster

"Finding accurate correspondences among different views is the Achilles’ heel of unsupervised Multi-View Stereo (MVS). Existing methods are built upon the assumption that corresponding pixels share similar photometric features. However, multi-view images in real scenarios observe non-Lambertian surf…

2022

Semi-Supervised Monocular 3D Object Detection by Multi-View Consistency

ECCV 2022poster

"The success of monocular 3D object detection highly relies on considerable labeled data, which is costly to obtain. To alleviate the annotation effort, we propose MVC-MonoDet, the first semi-supervised training framework that improves Monocular 3D object detection by enforcing multi-view consistenc…

Cited by 15SourcePDFScholar
2020

Particularity beyond Commonality: Unpaired Identity Transfer with Multiple References

ECCV 2020poster

Unpaired image-to-image translation aims to translate images from the source class to target one by providing sufficient data for these classes. Current few-shot translation methods use multiple reference images to describe the target domain through extracting common features. In this paper, we focu…

Cited by 0SourcePDFScholar