← Search

Anyi Rao

19 accepted papers

2026

Composing Concepts from Images and Videos via Concept-prompt Binding

CVPR 2026

Visual concept composition, which aims to integrate different elements from images and videos into a single, coherent visual output, still falls short in accurately extracting complex concepts from visual inputs and flexibly combining concepts from both images and videos. We introduce Bind & Compose

Cited by 0SourcecodeScholar
2026

Light of Normals: Unified Feature Representation for Universal Photometric Stereo

ICLR 2026poster

Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination models. Despite progress (e.g., SDM UniPS), two challenges remain. First, current encoders cannot guarantee that illumination…

Cited by 0SourcecodeScholar
2026

SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment

ICLR 2026poster

Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in textures and environments, and fail to include crucial compon…

Cited by 0SourceScholar
2025

Keyframe-Guided Creative Video Inpainting

CVPR 2025poster

Video inpainting, which aims to fill missing regions with visually coherent content, has emerged as a crucial technique for creative applications such as editing. While existing approaches achieve visual consistency or text-guided generation, they often struggle to balance coherence and creative div…

Cited by 0SourcePDFScholar
2025

Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport

ICLR 2025oral

Diffusion-based image generators are becoming unique methods for illumination harmonization and editing. The current bottleneck in scaling up the training of diffusion-based illumination editing models is mainly in the difficulty of preserving the underlying image details and maintaining intrinsic p…

Cited by 8SourcePDFScholar
2024

AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

ICLR 2024spotlight

With the advance of text-to-image (T2I) diffusion models (e.g., Stable Diffusion) and corresponding personalization techniques such as DreamBooth and LoRA, everyone can manifest their imagination into high-quality images at an affordable cost. However, adding motion dynamics to existing high-quality…

2024

Cinematic Behavior Transfer via NeRF-based Differentiable Filming

CVPR 2024poster

In the evolving landscape of digital media and video production the precise manipulation and reproduction of visual elements like camera movements and character actions are highly desired. Existing SLAM methods face limitations in dynamic scenes and human pose estimation often focuses on 2D projecti…

Cited by 6SourcePDFScholar
2024

CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers

ICML 2024poster

Recent vision-language models have achieved tremendous advances. However, their computational costs are also escalating dramatically, making model acceleration exceedingly critical. To pursue more efficient vision-language Transformers, this paper introduces Cross-Guided Ensemble of Tokens (CrossGET…

2024

SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models

ECCV 2024poster

"The development of text-to-video (T2V), i.e., generating videos with a given text prompt, has been significantly advanced in recent years. However, relying solely on text prompts often results in ambiguous frame composition due to spatial uncertainty. The research community thus leverages the dense…

2023

HireVAE: An Online and Adaptive Factor Model Based on Hierarchical and Regime-Switch VAE

IJCAI 2023poster

Factor model is a fundamental investment tool in quantitative investment, which can be empowered by deep learning to become more flexible and efficient in practical complicated investing situations. However, it is still an open question to build a factor model that can conduct stock prediction in an…

Cited by 4SourcePDFScholar
2023

Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences

AAAI 2023technical

Self-supervised learning has demonstrated remarkable capability in representation learning for skeleton-based action recognition. Existing methods mainly focus on applying global data augmentation to generate different views of the skeleton sequence for contrastive learning. However, due to the rich…

2022

AutoGPart: Intermediate Supervision Search for Generalizable 3D Part Segmentation

CVPR 2022poster

Training a generalizable 3D part segmentation network is quite challenging but of great importance in real-world applications. To tackle this problem, some works design task-specific solutions by translating human understanding of the task to machine's learning process, which faces the risk of missi…

Cited by 15PDFcodeScholar
2022

BungeeNeRF: Progressive Neural Radiance Field for Extreme Multi-Scale Scene Rendering

ECCV 2022poster

"Neural Radiance Field (NeRF) has achieved outstanding performance in modeling 3D objects and controlled scenes, usually under a single scale. In this work, we focus on multi-scale cases where large changes in imagery are observed at drastically different scales. This scenario vastly exists in the r…

Cited by 267SourcePDFScholar
2021

BlockPlanner: City Block Generation With Vectorized Graph Representation

ICCV 2021poster

City modeling is the foundation for computational urban planning, navigation, and entertainment. In this work, we present the first generative model of city blocks named BlockPlanner, and showcase its ability to synthesize valid city blocks with varying land lots configurations. We propose a novel v…

Cited by 22PDFScholar
2020

A Local-to-Global Approach to Multi-Modal Movie Scene Segmentation

CVPR 2020poster

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of movies. This is very challenging - compared to the videos st…

Cited by 155PDFcodeScholar
2020

A Unified Framework for Shot Type Classification Based on Subject Centric Lens

ECCV 2020poster

In film making, shot has a profound influence on how the story is delivered and how the audiences are echoed. As different scale and movement types of shots can express different emotions and contents, recognizing shots and their attributes is important to the understanding of movies as well as gene…

Cited by 85SourcePDFScholar
2020

MovieNet: A Holistic Dataset for Movie Understanding

ECCV 2020poster

Recent years have seen remarkable advances in visual understanding. However, how to understand a story-based long video with artistic styles, e.g. movie, remains challenging. In this paper, we introduce MovieNet -- a holistic dataset for movie understanding. MovieNet contains 1,100 movies with a lar…

2020

Online Multi-modal Person Search in Videos

ECCV 2020poster

The task of searching certain people in videos has seen increasing potential in real-world applications, such as video organization and editing. Most existing approaches are devised to work in an offline manner, where identifies can only be inferred after an entire video is examined. This working ma…

Cited by 35SourcePDFScholar