← Search

Hanyuan Liu

7 accepted papers

2026

Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering

ICML 2026poster

Large Multimodal Models (LMMs) have recently demonstrated their proficiency in holistic visual comprehension. However, most of them struggle to tackle region-level perception guided by visual prompts, especially for cases where multiple regions are referred simultaneously, or scenarios where global …

Cited by 0SourceScholar
2026

RegionCache: Semantic-Aware Region Reuse for Efficient Multi-Turn Image Generation

IJCAI 2026

Real-world image generation generally requires multi-turn editing, where users iteratively refine a small region while the majority of the image remains stable across turns. Despite this strong region-level stability, existing diffusion transformer (DiT)–based editing pipelines recompute the entire

Cited by 0Scholar
2025

Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 Dataset

CVPR 2025poster

Manga, a popular form of multimodal artwork, has traditionally been overlooked in deep learning advancements due to the absence of a robust dataset and comprehensive annotation. Manga segmentation is the key to the digital migration of manga. There exists a significant domain gap between the manga a…

Cited by 0SourcePDFScholar
2025

BlueNeg: A 35mm Negative Film Dataset for Restoring Channel-Heterogeneous Deterioration

ICCV 2025poster

While digitally acquired photographs have been dominating since around 2000, there remains a huge amount of legacy photographs being acquired by optical cameras and are stored in the form of film negatives. In this paper, we address the unique challenge of channel-heterogeneous deterioration in film…

2024

DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

ECCV 2024oral

"Animating a still image offers an engaging visual experience. Traditional image animation techniques mainly focus on animating natural scenes with stochastic dynamics (e.g. clouds and fluid) or domain-specific motions (e.g. human hair or body motions), and thus limits their applicability to more ge…