← Search

Ruihuang Li

20 accepted papers

2026

EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation

CVPR 2026

Visual effects (VFX) are essential for enhancing the expressiveness and creativity of video content, yet producing high-quality effects typically requires expert knowledge and costly production pipelines. Existing AIGC systems face significant challenges in VFX generation due to the scarcity of effe

Cited by 0SourcecodeScholar
2026

Fast Multi-view Consistent 3D Editing with Video Priors

AAAI 2026technical

Text-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employ 2D generation or editing models to process per-view individually, followed by iterative 2D-3D-2D updating. How

Cited by 0SourcePDFScholar
2026

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

ICML 2026poster

Humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings. However, existing omnimodal models still exhibit substantial performance degradation on visual tasks when the audio modality is incorporated. We identify this …

Cited by 0SourceScholar
2026

PromptEnhancer: Taming Your Rewriter for Text-to-Image Generation via Fine-Grained Reward

CVPR 2026

Recent text-to-image (T2I) diffusion models have achieved impressive progress in generating high-fidelity images, yet they often fail to faithfully follow complex user prompts, especially in attribute binding, negation, and compositional reasoning. To address this limitation, we propose PromptEnhanc

Cited by 0SourcecodeScholar
2025

FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded Sampling

ICLR 2025poster

While image generation with diffusion models has achieved a great success, generating images of higher resolution than the training size remains a challenging task due to the high computational cost. Current methods typically perform the entire sampling process at full resolution and process all fre…

2025

SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D Editing

AAAI 2025technical

Text-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achieve consistent edits across multiple viewpoints remains a challenge. While the it…

Cited by 0SourcePDFScholar
2024

ScatterFormer: Efficient Voxel Transformer with Scattered Linear Attention

ECCV 2024poster

"Window-based transformers excel in large-scale point cloud understanding by capturing context-aware representations with affordable attention computation in a more localized manner. However, the sparse nature of point clouds leads to a significant variance in the number of voxels per window. Existi…

2024

Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models

ECCV 2024poster

"Text-driven diffusion models have significantly advanced the image editing performance by using text prompts as inputs. One crucial step in text-driven image editing is to invert the original image into a latent noise code conditioned on the source prompt. While previous methods have achieved promi…

2023

DynaMask: Dynamic Mask Selection for Instance Segmentation

CVPR 2023poster

The representative instance segmentation methods mostly segment different object instances with a mask of the fixed resolution, e.g., 28x 28 grid. However, a low-resolution mask loses rich details, while a high-resolution mask incurs quadratic computation overhead. It is a challenging task to predic…

2023

FPR: False Positive Rectification for Weakly Supervised Semantic Segmentation

ICCV 2023poster

Many weakly supervised semantic segmentation (WSSS) methods employ the class activation map (CAM) to generate the initial segmentation results. However, CAM often fails to distinguish the foreground from its co-occurred background (e.g., train and railroad), resulting in inaccurate activation from t…

Cited by 46PDFcodeScholar
2023

MSF: Motion-Guided Sequential Fusion for Efficient 3D Object Detection From Point Cloud Sequences

CVPR 2023poster

Point cloud sequences are commonly used to accurately detect 3D objects in applications such as autonomous driving. Current top-performing multi-frame detectors mostly follow a Detect-and-Fuse framework, which extracts features from each frame of the sequence and fuses them to detect the objects in…

2023

One-to-Few Label Assignment for End-to-End Dense Detection

CVPR 2023poster

One-to-one (o2o) label assignment plays a key role for transformer based end-to-end detection, and it has been recently introduced in fully convolutional detectors for lightweight end-to-end dense detection. However, o2o can largely degrade the feature learning performance due to the limited number…

2023

SIM: Semantic-Aware Instance Mask Generation for Box-Supervised Instance Segmentation

CVPR 2023poster

Weakly supervised instance segmentation using only bounding box annotations has recently attracted much research attention. Most of the current efforts leverage low-level image features as extra supervision without explicitly exploiting the high-level semantic information of the objects, which will…

2022

Class-Balanced Pixel-Level Self-Labeling for Domain Adaptive Semantic Segmentation

CVPR 2022poster

Domain adaptive semantic segmentation aims to learn a model with the supervision of source domain data, and produce satisfactory dense predictions on unlabeled target domain. One popular solution to this challenging task is self-training, which selects high-scoring predictions on target samples as p…

Cited by 111PDFcodeScholar
2022

Exact Feature Distribution Matching for Arbitrary Style Transfer and Domain Generalization

CVPR 2022oral

Arbitrary style transfer (AST) and domain generalization (DG) are important yet challenging visual learning tasks, which can be cast as a feature distribution matching problem. With the assumption of Gaussian feature distribution, conventional feature distribution matching methods usually match the…

Cited by 244PDFcodeScholar
2022

Look Back and Forth: Video Super-Resolution With Explicit Temporal Difference Modeling

CVPR 2022poster

Temporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or com…

Cited by 60PDFcodeScholar
2022

Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection From Point Clouds

CVPR 2022poster

Transformer has demonstrated promising performance in many 2D vision tasks. However, it is cumbersome to apply the self-attention underlying transformer on large-scale point cloud data because point cloud is a long sequence and unevenly distributed in 3D space. To solve this issue, existing methods…

Cited by 211PDFcodeScholar
2021

T-SVDNet: Exploring High-Order Prototypical Correlations for Multi-Source Domain Adaptation

ICCV 2021poster

Most existing domain adaptation methods focus on adaptation from only one source domain, however, in practice there are a number of relevant sources that could be leveraged to help improve performance on target domain. We propose a novel approach named T-SVDNet to address the task of Multi-source Do…

Cited by 57PDFcodeScholar
2019

Reciprocal Multi-Layer Subspace Learning for Multi-View Clustering

ICCV 2019poster

Multi-view clustering is a long-standing important research topic, however, remains challenging when handling high-dimensional data and simultaneously exploring the consistency and complementarity of different views. In this work, we present a novel Reciprocal Multi-layer Subspace Learning (RMSL) al…

Cited by 158PDFScholar