← Search

Yuhe Liu

6 accepted papers

2026

Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers

CVPR 2026

Recent advances in diffusion-based controllable visual generation have led to remarkable improvements in image quality. However, these powerful models are typically deployed on cloud servers due to their large computational demands, raising serious concerns about user data privacy. To enable secure

Cited by 0SourceScholar
2026

MOGS: Monocular Object-Guided Gaussian Splatting in Large Scenes

ICRA 2026poster

Recent advances in 3D Gaussian Splatting (3DGS) deliver striking photorealism, and extending it to large scenes opens new opportunities for semantic reasoning and prediction in applications such as autonomous driving. Today's state-of-the-art systems for large scenes primarily originate from LiDAR-b…

2025

Enhancing Image Editing with Chain-of-Thought Reasoning and Multimodal Large Language Models

ICASSP 2025accepted

Image editing in our daily lives often requires models to first understand user’s intention and then proceed with the editing. Despite significant advancements in image editing technology, understanding and executing complex instructions remains a substantial challenge. Existing image editing models…

Cited by 0SourceScholar
2025

LeFusion: Controllable Pathology Synthesis via Lesion-Focused Diffusion Models

ICLR 2025spotlight

Patient data from real-world clinical practice often suffers from data scarcity and long-tail imbalances, leading to biased outcomes or algorithmic unfairness. This study addresses these challenges by generating lesion-containing image-segmentation pairs from lesion-free images. Previous efforts in…

2023

Boosting Semantic Segmentation from the Perspective of Explicit Class Embeddings

ICCV 2023poster

Semantic segmentation is a computer vision task that associates a label with each pixel in an image. Modern approaches tend to introduce class embeddings into semantic segmentation for deeply utilizing category semantics, and regard supervised class masks as final predictions. In this paper, we expl…

Cited by 11PDFcodeScholar