← Search

Yukai Shi

14 accepted papers

2026

Contractive Anchor Resolvent Diffusion for Incomplete Multi-View Clustering

ICML 2026poster

Incomplete Multi-View Clustering (IMVC) is fundamentally challenged by structural degradation induced by missing views, rather than the absence of feature values. Existing graph-based approaches either rely on costly data imputation or adopt first-order linear fusion, which acts as a weak low-pass f…

Cited by 0SourceScholar
2026

SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model

CVPR 2026

We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing methods struggle to simultaneously produce high-quality geometry and accurate poses under severe occlusion and open-set settin

Cited by 0SourcecodeScholar
2026

SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features

AAAI 2026technical

In this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training images, SegDINO3D is designed to fully leverage 2D representation from a pre-trained 2D detection model, including both

Cited by 0SourcePDFScholar
2025

CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

AAAI 2025technical

Video inpainting is a crucial task with diverse applications, including fine-grained video editing, video recovery, and video dewatermarking. However, most existing video inpainting methods primarily focus on visual content completion while neglecting text information. There are only a limited numbe…

2025

Imbalance in Balance: Online Concept Balancing in Generation Models

ICCV 2025accepted

In visual generation tasks, the responses and combinations of complex concepts often lack stability and are error-prone, which remains an under-explored area. In this paper, we attempt to explore the causal factors for poor concept responses through elaborately designed experiments. We also design a…

Cited by 0SourcePDFScholar
2025

Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

CVPR 2025poster

With the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the performance of video generation models. We assert that temporal splitting, detailed captions, and video quality filtering are…

2024

DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation

ICLR 2024poster

Text-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization process suffers slow convergence and the resultant 3D models o…

Cited by 22SourcePDFScholar
2024

NegVSR: Augmenting Negatives for Generalized Noise Modeling in Real-world Video Super-Resolution

AAAI 2024technical

The capability of video super-resolution (VSR) to synthesize high-resolution (HR) video from ideal datasets has been demonstrated in many works. However, applying the VSR model to real-world video with unknown and complex degradation remains a challenging task. First, existing degradation metrics in…

2024

Structural Information Guided Multimodal Pre-training for Vehicle-Centric Perception

AAAI 2024technical

Understanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neg…

2024

TOSS: High-quality Text-guided Novel View Synthesis from a Single Image

ICLR 2024poster

In this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image. While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the cha…

Cited by 18SourcePDFScholar
2023

DreamWaltz: Make a Scene with Complex 3D Animatable Avatars

NeurIPS 2023poster

We present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains chall…

2023

LipsFormer: Introducing Lipschitz Continuity to Vision Transformers

ICLR 2023poster

We present a Lipschitz continuous Transformer, called LipsFormer, to pursue training stability both theoretically and empirically for Transformer-based models. In contrast to previous practical tricks that address training instability by learning rate warmup, layer normalization, attention formulati…

2023

Scale-Aware Squeeze-and-Excitation for Lightweight Object Detection

RA-L 2023

Lightweight object detection can promote intelligent robotics to recognize surroundings objects with limited computational resources, and thus receives increasing attention in robotics communities. Recently, high-resolution networks (HRNets) can learn high-resolution representation and it obtains ex

Cited by 13SourceScholar
2017

Attention-Aware Face Hallucination via Deep Reinforcement Learning

CVPR 2017poster

Face hallucination is a domain-specific super-resolution problem with the goal to generate high-resolution (HR) faces from low-resolution (LR) input images. In contrast to existing methods that often learn a single patch-to-patch mapping from LR to HR images and are regardless of the contextual inte…

Cited by 249PDFScholar