← Search

Qifeng Chen

157 accepted papers

2026

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

ICLR 2026poster

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data (e.g., conflating acoustically distinct sounds like different dog barks under co…

Cited by 0SourcecodeScholar
2026

Active Intelligence in Video Avatars via Closed-loop World Modeling

CVPR 2026

Current video avatar generation methods excel at identity preservation and motion alignment but lack genuine agency--they cannot autonomously pursue long-term goals through adaptive environmental interaction. We address this by introducing L-IVA (Long-horizon Interactive Visual Avatar), a task and b

Cited by 0SourceScholar
2026

AudioX: A Unified Framework for Anything-to-Audio Generation

ICLR 2026poster

Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework, and 2) large-scale, high-quality training data. As such, we propose AudioX, a unified framework for anything-to-audio g…

Cited by 0SourcecodeScholar
2026

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization

CVPR 2026

We introduce AvatarPointillist, a novel framework for generating dynamic 4D Gaussian avatars from a single portrait image. At the core of our method is a decoder-only Transformer that autoregressively generates a point cloud for 3D Gaussian Splatting. This sequential approach allows for precise, ada

Cited by 0SourcecodeScholar
2026

EasyCreator: Empowering 4D Creation through Video Inpainting

ICLR 2026poster

We introduce EasyCreator, a novel 4D video creation framework capable of both generating and editing 4D content from a single monocular video input. By leveraging a powerful video inpainting foundation model as a generative prior, we reformulate 4D video creation as a video inpainting task, enabling…

Cited by 0SourceScholar
2026

EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

ICLR 2026poster

Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly use two-stage Low-Rank Adaptations (LoRAs) finetuning to obtain better performance. However, existing adaptation-based…

Cited by 0SourceScholar
2026

FastVMT: Eliminating Redundancy in Video Motion Transfer

ICLR 2026poster

Video motion transfer aims to synthesize videos by generating visual content according to a text prompt while transferring the motion pattern observed in a reference video. Recent methods predominantly use the Diffusion Transformer (DiT) architecture. To achieve satisfactory runtime, several methods…

Cited by 0SourceScholar
2026

FlyAware: Inertia-Aware Aerial Manipulation via Vision-Based Estimation and Post-Grasp Adaptation

RA-L 2026

Aerial manipulators (AMs) are gaining increasing attention in automated transportation and emergency services due to their superior dexterity compared to conventional multirotor drones. However, their practical deployment is challenged by the complexity of time-varying inertial parameters, which are

Cited by 1SourceScholar
2026

Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control

ICLR 2026poster

While recent flow-based image editing models demonstrate general-purpose capabilities across diverse tasks, they often struggle to specialize in challenging scenarios---particularly those involving large-scale shape transformations. When performing such structural edits, these methods either fail t…

Cited by 0SourcecodeScholar
2026

Group Editing: Edit Multiple Images in One Go

CVPR 2026

In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images may vary significantly in pose, viewpoint, and spatial layout. Achieving coherent edits requires establishing reliable co

Cited by 0SourcecodeScholar
2026

LiDAR-GS++: Improving LiDAR Gaussian Reconstruction via Diffusion Priors

AAAI 2026technical

Recent GS-based rendering has made significant progress for LiDAR, surpassing Neural Radiance Fields (NeRF) in both quality and speed. However, these methods exhibit artifacts in extrapolated novel view synthesis due to the incomplete reconstruction from single traversal scans. To address this limit

Cited by 0SourcePDFScholar
2026

MagicQuill V2: Precise and Interactive Image Editing with Layered Visual Cues

CVPR 2026

We propose MagicQuill V2, a novel framework that introduces a layered composition paradigm to generative image editing, bridging the gap between the semantic power of modern diffusion models and the granular control of traditional graphics software. While state-of-the-art diffusion transformers exce

Cited by 0SourcecodeScholar
2026

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

AAAI 2026technical

Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering f

Cited by 0SourcePDFScholar
2026

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real-world visual corruptions. While existing robustness enhancement approaches exist, they are limited: black-box feature alignment lacks interpr…

Cited by 0SourceScholar
2026

ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data

ICLR 2026oral

Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously, showing great potential, yet progress is limited by the lack of large-scale, open-source computer use data and foundation models. In this work, we introduce ScaleCUA, a step toward scaling open-sour…

Cited by 0SourcecodeScholar
2026

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

CVPR 2026

Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel dat

Cited by 0SourcecodeScholar
2026

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

ICML 2026poster

Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation--a hallmark of human creativity. Current reasoning-augmented paradigms mostly rely on explicit thought processes, where intermediate reasoni…

Cited by 0SourceScholar
2026

Switch: Learning Agile Skills Switching for Humanoid Robots

ICRA 2026poster

Recent advancements in whole-body control through deep reinforcement learning have enabled humanoid robots to achieve remarkable progress in real-world challenging locomotion skills. However, existing approaches often struggle with flexible transitions between distinct skills, creating safety concer…

2026

TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment

ICLR 2026poster

While visual-language models have profoundly linked features between texts and images, the incorporation of 3D modality data, such as point clouds and 3D Gaussians, further enables pretraining for 3D-related tasks, e.g., cross-modal retrieval, zero-shot classification, and scene recognition. As chal…

Cited by 0SourcecodeScholar
2026

Threat-Aware UAV Dodging of Human-Thrown Projectiles With an RGB-D Camera

RA-L 2026

Uncrewed aerial vehicles (UAVs) performing tasks such as transportation and aerial photography are vulnerable to intentional projectile attacks from humans. Dodging such a sudden and fast projectile poses a significant challenge for UAVs, requiring ultra-low latency responses and agile maneuvers. Dr

Cited by 2SourceScholar
2026

Threat-Aware UAV Dodging of Human-Thrown Projectiles with an RGB-D Camera

ICRA 2026poster

Unmanned aerial vehicles (UAVs) performing tasks such as transportation and aerial photography are vulnerable to intentional projectile attacks from humans. Dodging such a sudden and fast projectile poses a significant challenge for UAVs, requiring ultra-low latency responses and agile maneuvers. Dr…

2025

AlignGuard: Scalable Safety Alignment for Text-to-Image Generation

ICCV 2025poster

Text-to-image (T2I) models have become widespread, but their limited safety guardrails expose end users to harmful content and potentially allow for model misuse. Current safety measures are typically limited to text-based filtering or concept removal strategies, able to remove just a few concepts f…

Cited by 0SourcePDFScholar
2025

DiT4Edit: Diffusion Transformer for Image Editing

AAAI 2025technical

Despite recent advances in UNet-based image editing, methods for shape-aware object editing in high-resolution images are still lacking. Compared to UNet, Diffusion Transformers (DiT) demonstrate superior capabilities to effectively capture the long-range dependencies among patches, leading to highe…

2025

EEdit : Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing

ICCV 2025poster

Inversion-based image editing is rapidly gaining momentum while suffering from significant computation overhead, hindering its application in real-time interactive scenarios. In this paper, we rethink that the redundancy in inversion-based image editing exists in both the spatial and temporal dimens…

2025

Edicho: Consistent Image Editing in the Wild

ICCV 2025poster

As a verified need, consistent editing across in-the-wild images remains a technical challenge arising from various unmanageable factors, like object poses, lighting conditions, and photography environments. Edicho steps in with a training-free solution based on diffusion models, featuring a fundame…

2025

Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis

CVPR 2025poster

Due to the difficulty in scaling up, generative adversarial networks (GANs) seem to be falling out of grace with the task of text-conditioned image synthesis. Sparsely activated mixture-of-experts (MoE) has recently been demonstrated as a valid solution to training large-scale models with limited re…

2025

Follow-Your-Click: Open-domain Regional Image Animation via Motion Prompts

AAAI 2025technical

Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. A…

Cited by 52SourcePDFScholar
2025

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

NeurIPS 2025poster

Recent advancements in video generation have enabled the creation of high-quality, visually compelling videos. However, generating videos that adhere to the laws of physics remains a critical challenge for applications requiring realism and accuracy. In this work, we propose **PhysHPO**, a novel fra…

Cited by 0SourceScholar
2025

Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving

ICCV 2025poster

In light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hin…

Cited by 0SourcePDFScholar
2025

Industrial-Grade Sensor Simulation via Gaussian Splatting: A Modular Framework for Scalable Editing and Full-Stack Validation

IROS 2025

Sensor simulation is pivotal for scalable validation of autonomous driving systems, yet existing Neural Radiance Fields (NeRF) based methods face applicability and efficiency challenges in industrial workflows. This paper introduces a Gaussian Splatting (GS) based system to address these challenges:

Cited by 3SourceScholar
2025

Infinite-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation

AAAI 2025technical

This paper explores higher-resolution video outpainting with extensive content generation. We point out common issues faced by existing methods when attempting to largely outpaint videos: the generation of low-quality content and limitations imposed by GPU memory. To address these challenges, we pro…

2025

InstantSwap: Fast Customized Concept Swapping across Sharp Shape Differences

ICLR 2025poster

Recent advances in Customized Concept Swapping (CCS) enable a text-to-image model to swap a concept in the source image with a customized target concept. However, the existing methods still face the challenges of $\textit{\textbf{inconsistency}}$ and $\textit{\textbf{inefficiency}}$. They struggle t…

2025

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

CVPR 2025highlight

The intuitive nature of drag-based interaction has led to its growing adoption for controlling object trajectories in image-to-video synthesis. Still, existing methods that perform dragging in the 2D space usually face ambiguity when handling out-of-plane movements. In this work, we augment the inte…

2025

MagicQuill: An Intelligent Interactive Image Editing System

CVPR 2025poster

As a highly practical application, image editing encounters a variety of user demands and thus prioritizes excellent ease of use. In this paper, we unveil MagicQuill, an integrated image editing system designed to support users in swiftly actualizing their creativity. Our system starts with a stream…

2025

MangaNinja: Line Art Colorization with Precise Reference Following

CVPR 2025highlight

Derived from diffusion models, MangaNinja specializes in the task of reference-guided line art colorization. We incorporate two thoughtful designs to ensure precise character detail transcription, including a patch shuffling module to facilitate correspondence learning between the reference color im…

Cited by 3SourcePDFScholar
2025

Master Rules from Chaos: Learning to Reason, Plan, and Interact from Chaos for Tangram Assembly

ICRA 2025

Tangram assembly, the art of human intelligence and manipulation dexterity, is a new challenge for robotics and reveals the limitations of state-of-the-arts. Here, we describe our initial exploration and highlight key problems in reasoning, planning, and manipulation for robotic tangram assembly. We

Cited by 1SourcecodeScholar
2025

RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion Priors

ICCV 2025poster

A single-pass driving clip frequently results in incomplete scanning of the road structure, making reconstructed scene expanding a critical requirement for sensor simulators to effectively regress driving actions. Although contemporary 3D Gaussian Splatting (3DGS) techniques achieve remarkable recon…

2025

Rethinking Layered Graphic Design Generation with a Top-Down Approach

ICCV 2025poster

Graphic design is crucial for conveying ideas and messages. Designers usually organize their work into objects, backgrounds, and vectorized text layers to simplify editing. However, this workflow demands considerable expertise. With the rise of GenAI methods, an endless supply of high-quality graphi…

Cited by 0SourcePDFScholar
2025

SkillMimic: Learning Basketball Interaction Skills from Demonstrations

CVPR 2025highlight

Traditional reinforcement learning methods for human-object interaction (HOI) rely on labor-intensive, manually designed skill rewards that do not generalize well across different interactions. We introduce SkillMimic, a unified data-driven framework that fundamentally changes how agents learn inter…

2025

VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

CVPR 2025poster

In this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset by collecting 360K video-music pairs, including various genres such as movie trailers, advertisements, and documentaries. Furthermore, we propose VidMuse, a simple framewor…

2025

VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

CVPR 2025poster

Recent progress in generative diffusion models has greatly advanced text-to-video generation. While text-to-video models trained on large-scale, diverse datasets can produce varied outputs, these generations often deviate from user preferences, highlighting the need for preference alignment on pre-t…

2025

VideoVAE+: Large Motion Video Autoencoding with Cross-modal Video VAE

ICCV 2025poster

Learning a robust video Variational Autoencoder (VAE) is essential for reducing video redundancy and facilitating efficient video generation. Directly applying image VAEs to individual frames in isolation results in temporal inconsistencies and fails to compress temporal redundancy effectively. Exis…

Cited by 0SourcePDFScholar
2024

A Diffusion Model with State Estimation for Degradation-Blind Inverse Imaging

AAAI 2024technical

Solving the task of inverse imaging problems can restore unknown clean images from input measurements that have incomplete information. Utilizing powerful generative models, such as denoising diffusion models, could better tackle the ill-posed issues of inverse problems with the distribution prior o…

Cited by 2SourcePDFScholar
2024

Adaptive Domain Learning for Cross-domain Image Denoising

NeurIPS 2024poster

Different camera sensors have different noise patterns, and thus an image denoising model trained on one sensor often does not generalize well to a different sensor. One plausible solution is to collect a large dataset for each sensor for training or fine-tuning, which is inevitably time-consuming.…

Cited by 0SourcePDFScholar
2024

CMDFusion: Bidirectional Fusion Network With Cross-Modality Knowledge Distillation for LiDAR Semantic Segmentation

RA-L 2024

2D RGB images and 3D LIDAR point clouds provide complementary knowledge for the perception system of autonomous vehicles. Several 2D and 3D fusion methods have been explored for the LIDAR semantic segmentation task, but they suffer from different problems. 2D-to-3D fusion methods require strictly pa

Cited by 21SourcecodeScholar
2024

CoDeF: Content Deformation Fields for Temporally Consistent Video Processing

CVPR 2024highlight

We present the content deformation field (CoDeF) as a new type of video representation which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e. rendered from the canon…

2024

Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving

ICRA 2024poster

We present a new 3D point-based detector model, named Shift-SSD, for precise 3D object detection in autonomous driving. Traditional point-based 3D object detectors often employ architectures that rely on a progressive downsampling of points. While this method effectively reduces computational demand…

Cited by 1SourceScholar
2024

DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation

CVPR 2024poster

We propose DiffSHEG a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation. While previous works focused on co-speech gesture or expression generation individually the joint generation of synchronized expressions and gestures remains barely explored. To address th…

Cited by 37SourcePDFScholar
2024

Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free Videos

AAAI 2024technical

Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose captions and the generative prior models for videos. In this w…

2024

GIC: Gaussian-Informed Continuum for Physical Property Identification and Simulation

NeurIPS 2024oral

This paper studies the problem of estimating physical properties (system identification) through visual observations. To facilitate geometry-aware guidance in physical property estimation, we introduce a novel hybrid framework that leverages 3D Gaussian representation to not only capture explicit sh…

2024

Gaussian Shell Maps for Efficient 3D Human Generation

CVPR 2024poster

Efficient generation of 3D digital humans is important in several industries including virtual reality social media and cinematic production. 3D generative adversarial networks (GANs) have demonstrated state-of-the-art (SOTA) quality and diversity for generated assets. Current 3D GAN architectures h…

2024

HAWK: Learning to Understand Open-World Video Anomalies

NeurIPS 2024poster

Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prev…

2024

Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation

ECCV 2024poster

"Diffusion models have proven to be highly effective in image and video generation; however, they encounter challenges in the correct composition of objects when generating images of varying sizes due to single-scale training data. Adapting large pre-trained diffusion models to higher resolution dem…

2024

Multitarget Device-Free Localization via Cross-Domain Wi-Fi RSS Training Data and Attentional Prior Fusion

AAAI 2024technical

Device-free localization (DFL) using easily-obtained Wi-Fi received signal strength (RSS) has wide real-world applications for not requiring people to carry trackable devices. However, accurate multitarget DFL remains challenging due to the unknown number of targets, multipath interference (MPI), es…

Cited by 0SourcePDFScholar
2024

Open-Vocabulary Category-Level Object Pose and Size Estimation

RA-L 2024

This letter studies a new open-set problem, the open-vocabulary category-level object pose and size estimation. Given human text descriptions of arbitrary novel object categories, the robot agent seeks to predict the position, orientation, and size of the target object in the observed scene image. T

Cited by 11SourceScholar
2024

Robust Depth Enhancement via Polarization Prompt Fusion Tuning

CVPR 2024poster

Existing depth sensors are imperfect and may provide inaccurate depth values in challenging scenarios such as in the presence of transparent or reflective objects. In this work we present a general framework that leverages polarization imaging to improve inaccurate depth measurements from various de…

2024

SPIRE: Semantic Prompt-Driven Image Restoration

ECCV 2024poster

"Text-driven diffusion models have become increasingly popular for various image editing tasks, including inpainting, stylization, and object replacement. However, it still remains an open research problem to adopt this language-vision paradigm for more fine-level image processing tasks, such as den…

Cited by 4SourcePDFScholar
2024

ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion Models

ICLR 2024spotlight

In this work, we investigate the capability of generating images from pre-trained diffusion models at much higher resolutions than the training image sizes. In addition, the generated images should have arbitrary image aspect ratios. When generating images directly at a higher resolution, 1024 x 102…

2024

Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

CVPR 2024poster

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently existing diffusion-based methods tackle video and audio generation separately which hinders the technique transfer from academia to industry. In this work we aim at filling the gap w…

2024

TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering

ECCV 2024oral

"The diffusion model has been proven a powerful generative model in recent years, yet it remains a challenge in generating visual text. Although existing work has endeavored to enhance the accuracy of text rendering, these methods still suffer from several drawbacks, such as (1) limited flexibility…

2024

Using Left and Right Brains Together: Towards Vision and Language Planning

ICML 2024poster

Large Language Models (LLMs) and Large Multi-modality Models (LMMs) have demonstrated remarkable decision masking capabilities on a variety of tasks. However, they inherently operate planning within the language space, lacking the vision and spatial imagination ability. In contrast, humans utilize b…

Cited by 5SourcePDFScholar
2023

4D Panoptic Scene Graph Generation

NeurIPS 2023spotlight

We are living in a three-dimensional space while moving forward through a fourth dimension: time. To allow artificial intelligence to develop a comprehensive understanding of such a 4D environment, we introduce **4D Panoptic Scene Graph (PSG-4D)**, a new representation that bridges the raw visual da…

Cited by 16SourcePDFScholar
2023

Blind Video Deflickering by Neural Filtering With a Flawed Atlas

CVPR 2023poster

Many videos contain flickering artifacts; common causes of flicker include video processing algorithms, video generation algorithms, and capturing videos under specific situations. Prior work usually requires specific guidance such as the flickering frequency, manual annotations, or extra consistent…

2023

Bootstrap Motion Forecasting With Self-Consistent Constraints

ICCV 2023poster

We present a novel framework to bootstrap Motion forecasting with Self-consistent Constraints (MISC). The motion forecasting task aims at predicting future trajectories of vehicles by incorporating spatial and temporal information from the past. A key design of MISC is the proposed Dual Consistency…

Cited by 19PDFScholar
2023

Delving StyleGAN Inversion for Image Editing: A Foundation Latent Space Viewpoint

CVPR 2023poster

GAN inversion and editing via StyleGAN maps an input image into the embedding spaces (W, W^+, and F) to simultaneously maintain image fidelity and meaningful manipulation. From latent space W to extended latent space W^+ to feature space F in StyleGAN, the editability of GAN inversion decreases whil…

2023

DynaFed: Tackling Client Data Heterogeneity With Global Dynamics

CVPR 2023poster

The Federated Learning (FL) paradigm is known to face challenges under heterogeneous client data. Local training on non-iid distributed data results in deflected local optimum, which causes the client models drift further away from each other and degrades the aggregated global model's performance. A…

Cited by 30SourcePDFScholar
2023

ERRA: An Embodied Representation and Reasoning Architecture for Long-Horizon Language-Conditioned Manipulation Tasks

RA-L 2023

This letter introduces ERRA, an embodied learning architecture that enables robots to jointly obtain three fundamental capabilities (reasoning, planning, and interaction) for solving long-horizon language-conditioned manipulation tasks. ERRA is based on tightly-coupled probabilistic inferences at tw

Cited by 17SourceScholar
2023

Enlarging Instance-Specific and Class-Specific Information for Open-Set Action Recognition

CVPR 2023poster

Open-set action recognition is to reject unknown human action cases which are out of the distribution of the training set. Existing methods mainly focus on learning better uncertainty scores but dismiss the importance of feature representations. We find that features with richer semantic diversity c…

2023

FateZero: Fusing Attentions for Zero-shot Text-based Video Editing

ICCV 2023oral

The diffusion-based generative models have achieved remarkable success in text-based image generation. However, since it contains enormous randomness in generation progress, it is still challenging to apply such models for real-world visual content editing, especially in videos. In this paper, we pr…

Cited by 338PDFcodeScholar
2023

Flipbot: Learning Continuous Paper Flipping via Coarse-to-Fine Exteroceptive-Proprioceptive Exploration

ICRA 2023poster

This paper tackles the task of singulating and grasping paper-like deformable objects. We refer to such tasks as paper-flipping. In contrast to manipulating deformable objects that lack compression strength (such as shirts and ropes), minor variations in the physical properties of the paper-like def…

Cited by 4SourcecodeScholar
2023

High-Fidelity 3D GAN Inversion by Pseudo-Multi-View Optimization

CVPR 2023poster

We present a high-fidelity 3D generative adversarial network (GAN) inversion framework that can synthesize photo-realistic novel views while preserving specific details of the input image. High-fidelity 3D GAN inversion is inherently challenging due to the geometry-texture trade-off, where overfitti…

2023

Human MotionFormer: Transferring Human Motions with Vision Transformers

ICLR 2023poster

Human motion transfer aims to transfer motions from a target dynamic person to a source static one for motion synthesis. An accurate matching between the source person and the target motion in both large and subtle motion changes is vital for improving the transferred motion quality. In this paper,…

2023

Learn to Grasp Via Intention Discovery and Its Application to Challenging Clutter

RA-L 2023

Humans excel in grasping objects through diverse and robust policies, many of which are so probabilistically rare that exploration-based learning methods hardly observe and learn. Inspired by the human learning process, we propose a method to extract and exploit latent intents from demonstrations, a

Cited by 1SourceScholar
2023

Learning 3D-Aware Image Synthesis With Unknown Pose Distribution

CVPR 2023poster

Existing methods for 3D-aware image synthesis largely depend on the 3D pose distribution pre-estimated on the training set. An inaccurate estimation may mislead the model into learning faulty geometry. This work proposes PoF3D that frees generative radiance fields from the requirements of 3D pose pr…

2023

LinkGAN: Linking GAN Latents to Pixels for Controllable Image Synthesis

ICCV 2023poster

This work presents an easy-to-use regularizer for GAN training, which helps explicitly link some axes of the latent space to a set of pixels in the synthesized image. Establishing such a connection facilitates a more convenient local control of GAN generation, where users can alter the image content…

Cited by 30PDFScholar
2023

MetaPortrait: Identity-Preserving Talking Head Generation With Fast Personalized Adaptation

CVPR 2023poster

In this work, we propose an ID-preserving talking head generation framework, which advances previous methods in two aspects. First, as opposed to interpolating from sparse flow, we claim that dense landmarks are crucial to achieving accurate geometry-aware flow fields. Second, inspired by face-swapp…

2023

RODIN: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion

CVPR 2023highlight

This paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle thi…

Cited by 369SourcePDFScholar
2023

Randomized Quantization: A Generic Augmentation for Data Agnostic Self-supervised Learning

ICCV 2023poster

Self-supervised representation learning follows a paradigm of withholding some part of the data and tasking the network to predict it from the remaining part. Among many techniques, data augmentation lies at the core for creating the information gap. Towards this end, masking has emerged as a generi…

Cited by 11PDFcodeScholar
2023

Real-Time 6K Image Rescaling With Rate-Distortion Optimization

CVPR 2023poster

The task of image rescaling aims at embedding an high-resolution (HR) image into a low-resolution (LR) one that can contain embedded information for HR image reconstruction. Existing image rescaling methods do not optimize the LR image file size and recent flow-based rescaling methods are not real-t…

2023

Rotating without Seeing: Towards In-hand Dexterity through Touch

RSS 2023

Tactile information plays a critical role in human dexterity. It reveals useful contact information that may not be inferred directly from vision. In fact, humans can even perform in-hand dexterous manipulation without using vision. Can we enable the same ability for the multi-finger robot hand? In

Cited by 70SourceScholar
2023

Scene-level Point Cloud Colorization with Semantics-and-geometry-aware Networks

ICRA 2023poster

In robotic applications, we often obtain tons of 3D point cloud data without color information, and it is difficult to visualize point clouds in a meaningful and colorful way. Can we colorize 3D point clouds for better visualization? Existing deep learning-based colorization methods usually only tak…

Cited by 3SourceScholar
2023

TextDiffuser: Diffusion Models as Text Painters

NeurIPS 2023poster

Diffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue, we introduce TextDiffuser, focusing on generating images with visually appealing text that is coherent with backgrounds…

Cited by 128SourcePDFScholar
2023

The Devil is in the Wrongly-classified Samples: Towards Unified Open-set Recognition

ICLR 2023poster

Open-set Recognition (OSR) aims to identify test samples whose classes are not seen during the training process. Recently, Unified Open-set Recognition (UOSR) has been proposed to reject not only unknown samples but also known but wrongly classified samples, which tends to be more practical in real-…

2022

3D-Aware Indoor Scene Synthesis with Depth Priors

ECCV 2022poster

"Despite the recent advancement of Generative Adversarial Networks (GANs) in learning 3D-aware image synthesis from 2D data, existing methods fail to model indoor scenes due to the large diversity of room layouts and the objects inside. We argue that indoor scenes do not have a shared intrinsic stru…

Cited by 37SourcePDFScholar
2022

A Portable Multiscopic Camera for Novel View and Time Synthesis in Dynamic Scenes

IROS 2022poster

We present a portable multiscopic camera system with a dedicated model for novel view and time synthesis in dynamic scenes. Our goal is to render high-quality images for a dynamic scene from any viewpoint at any time using our portable multiscopic camera. To achieve such novel view and time synthesi…

Cited by 4SourcecodeScholar
2022

AniFaceGAN: Animatable 3D-Aware Face Image Generation for Video Avatars

NeurIPS 2022accept

Although 2D generative models have made great progress in face image generation and animation, they often suffer from undesirable artifacts such as 3D inconsistency when rendering images from different camera viewpoints. This prevents them from synthesizing video animations indistinguishable from re…

2022

Efficient Point Cloud Segmentation with Geometry-Aware Sparse Networks

ECCV 2022poster

"In point cloud learning, sparsity and geometry are two core properties. Recently, many approaches have been proposed through single or multiple representations to improve the performance of point cloud semantic segmentation. However, these works fail to maintain the balance among performance, effic…

Cited by 26SourcePDFScholar
2022

Harvesting Partially-Disjoint Time-Frequency Information for Improving Degenerate Unmixing Estimation Technique

ICASSP 2022accepted

The degenerate unmixing estimation technique (DUET) is one of the most efficient blind source separation algorithms tackling the challenging situation when the number of sources exceeds the number of microphones. However, as a time-frequency mask-based method, DUET erroneously results in interferenc…

Cited by 0SourceScholar
2022

High-Fidelity GAN Inversion for Image Attribute Editing

CVPR 2022poster

We present a novel high-fidelity generative adversarial network (GAN) inversion framework that enables attribute editing with image-specific details well-preserved (e.g., background, appearance, and illumination). We first analyze the challenges of high-fidelity GAN inversion from the perspective of…

Cited by 313PDFcodeScholar
2022

Improving 3D-aware Image Synthesis with A Geometry-aware Discriminator

NeurIPS 2022accept

3D-aware image synthesis aims at learning a generative model that can render photo-realistic 2D images while capturing decent underlying 3D shapes. A popular solution is to adopt the generative adversarial network (GAN) and replace the generator with a 3D renderer, where volume rendering with neural…

2022

One Model to Edit Them All: Free-Form Text-Driven Image Manipulation with Semantic Modulations

NeurIPS 2022accept

Free-form text prompts allow users to describe their intentions during image manipulation conveniently. Based on the visual latent space of StyleGAN[21] and text embedding space of CLIP[34], studies focus on how to map these two latent spaces for text-driven attribute manipulations. Currently, the l…

2022

Point Cloud Compression with Sibling Context and Surface Priors

ECCV 2022poster

"We present a novel octree-based multi-level framework for large-scale point cloud compression, which can organize sparse and unstructured point clouds in a memory-efficient way. In this framework, we propose a new entropy model that explores the hierarchical dependency in an octree using the contex…

2022

Real-Time Neural Character Rendering with Pose-Guided Multiplane Images

ECCV 2022poster

"We propose pose-guided multiplane image (MPI) synthesis which can render an animatable character in real scenes with photorealistic quality. We use a portable camera rig to capture the multi-view images along with the driving signal for the moving subject. Our method generalizes the image-to-image…

2022

Region-Based Semantic Factorization in GANs

ICML 2022spotlight

Despite the rapid advancement of semantic discovery in the latent space of Generative Adversarial Networks (GANs), existing approaches either are limited to finding global attributes or rely on a number of segmentation masks to identify local attributes. In this work, we present a highly efficient a…

2022

Shape From Polarization for Complex Scenes in the Wild

CVPR 2022poster

We present a new data-driven approach with physics-based priors to scene-level normal estimation from a single polarization image. Existing shape from polarization (SfP) works mainly focus on estimating the normal of a single object rather than complex scenes in the wild. A key barrier to high-quali…

Cited by 69PDFcodeScholar
2022

Volumetric-based Contact Point Detection for 7-DoF Grasping

CoRL 2022poster

In this paper, we propose a novel grasp pipeline based on contact point detection on the truncated signed distance function (TSDF) volume to achieve closed-loop 7-degree-of-freedom (7-DoF) grasping on cluttered environments. The key aspects of our method are that 1) the proposed pipeline exploits th…

Cited by 11SourcecodeScholar
2021

DRINet: A Dual-Representation Iterative Learning Network for Point Cloud Segmentation

ICCV 2021poster

We present a novel and flexible architecture for point cloud segmentation with dual-representation iterative learning. In point cloud processing, different representations have their own pros and cons. Thus, finding suitable ways to represent point cloud data structure while keeping its own internal…

Cited by 50PDFScholar
2021

Dual-Camera Super-Resolution With Aligned Attention Modules

ICCV 2021poster

We present a novel approach to reference-based super-resolution (RefSR) with the focus on dual-camera super-resolution (DCSR), which utilizes reference images for high-quality and high-fidelity results. Our proposed method generalizes the standard patch-based feature matching with spatial alignment…

Cited by 54PDFcodeScholar
2021

FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose Estimation

CVPR 2021poster

In this work, we present FFB6D, a full flow bidirectional fusion network designed for 6D pose estimation from a single RGBD image. Our key insight is that appearance information in the RGB image and geometry information from the depth image are two complementary data sources, and it still remains un…

Cited by 364PDFcodeScholar
2021

Involution: Inverting the Inherence of Convolution for Visual Recognition

CVPR 2021poster

Convolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatial-agnostic and channel-specific. Instead, we present a novel atomic opera…

Cited by 468PDFcodeScholar
2021

Joint Depth and Normal Estimation from Real-world Time-of-flight Raw Data

IROS 2021poster

We present a novel approach to joint depth and normal estimation for time-of-flight (ToF) sensors. Our model learns to predict the high-quality depth and normal maps jointly from ToF raw sensor data. To achieve this, we meticulously constructed the first large-scale dataset (named ToF-100) with pair…

Cited by 4SourceScholar
2021

Learning to Predict Vehicle Trajectories with Model-based Planning

CoRL 2021poster

Predicting the future trajectories of on-road vehicles is critical for autonomous driving. In this paper, we introduce a novel prediction framework called PRIME, which stands for Prediction with Model-based Planning. Unlike recent prediction works that utilize neural networks to model scene context…

Cited by 160SourceScholar
2021

Low-Rank Subspaces in GANs

NeurIPS 2021poster

The latent space of a Generative Adversarial Network (GAN) has been shown to encode rich semantics within some subspaces. To identify these subspaces, researchers typically analyze the statistical information from a collection of synthesized data, and the identified subspaces tend to control image a…

2021

Normalized Human Pose Features for Human Action Video Alignment

ICCV 2021poster

We present a novel approach for extracting human pose features from human action videos. The goal is to let the pose features capture only the poses of the action while being invariant to other factors, including video backgrounds, the video subject's anthropometric characteristics and viewpoints. S…

Cited by 17PDFScholar
2021

Safety-Aware Motion Prediction With Unseen Vehicles for Autonomous Driving

ICCV 2021poster

Motion prediction of vehicles is critical but challenging due to the uncertainties in complex environments and the limited visibility caused by occlusions and limited sensor ranges. In this paper, we study a new task, safety-aware motion prediction with unseen vehicles for autonomous driving. Unlike…

Cited by 32PDFcodeScholar
2021

SinIR: Efficient General Image Manipulation with Single Image Reconstruction

ICML 2021spotlight

We propose SinIR, an efficient reconstruction-based framework trained on a single natural image for general image manipulation, including super-resolution, editing, harmonization, paint-to-image, photo-realistic style transfer, and artistic style transfer. We train our model on a single image with c…

2021

Stereo Matching by Self-supervision of Multiscopic Vision

IROS 2021poster

Self-supervised learning for depth estimation possesses several advantages over supervised learning. The benefits of no need for ground-truth depth, online fine-tuning, and better generalization with unlimited data attract researchers to seek self-supervised solutions. In this work, we propose a new…

Cited by 18SourceScholar
2020

Fully Convolutional Networks for Continuous Sign Language Recognition

ECCV 2020poster

Continuous sign language recognition (SLR) is a challenging task that requires learning on both spatial and temporal dimensions of signing frame sequences. Most recent work accomplishes this by using CNN and RNN hybrid networks. However, training these networks is generally non-trivial, and most of…

Cited by 200SourcePDFScholar
2020

Learning to Learn Parameterized Classification Networks for Scalable Input Images

ECCV 2020poster

Convolutional Neural Networks (CNNs) do not have a predictable recognition behavior with respect to the input resolution change. This prevents the feasibility of deployment on different input image resolutions for a specific model. To achieve efficient and flexible image classification at runtime, w…

2020

PSConv: Squeezing Feature Pyramid into One Compact Poly-Scale Convolutional Layer

ECCV 2020poster

Despite their strong modeling capacities, Convolutional Neural Networks (CNNs) are often scale-sensitive. For enhancing the robustness of CNNs to scale variance, multi-scale feature fusion from different layers or filters attracts great attention among existing solutions, while the more granular ker…

2020

PiP: Planning-informed Trajectory Prediction for Autonomous Driving

ECCV 2020poster

It is critical to predict the motion of surrounding vehicles for self-driving planning, especially in a socially compliant and flexible way. However, future prediction is challenging due to the interaction and uncertainty in driving behaviors. We propose planning-informed trajectory prediction (PiP)…

2020

Polarized Reflection Removal With Perfect Alignment in the Wild

CVPR 2020poster

We present a novel formulation to removing reflection from polarized images in the wild. We first identify the misalignment issues of existing reflection removal datasets where the collected reflection-free images are not perfectly aligned with input mixed images due to glass refraction. Then we bui…

Cited by 132PDFcodeScholar
2020

Self-supervised Object Tracking with Cycle-consistent Siamese Networks

IROS 2020poster

Self-supervised learning for visual object tracking possesses valuable advantages compared to supervised learning, such as the non-necessity of laborious human annotations and online training. In this work, we exploit an end-to-end Siamese network in a cycle-consistent self-supervised framework for…

Cited by 21SourcecodeScholar
2018

Combinatorial Optimization with Graph Convolutional Networks and Guided Tree Search

NeurIPS 2018poster

We present a learning-based approach to computing solutions for certain NP-hard problems. Our approach combines deep learning techniques with useful algorithmic elements from classic heuristics. The central component is a graph convolutional network that is trained to estimate the likelihood, for ea…

Cited by 637SourcePDFScholar