← Search

Peng Dai

36 accepted papers

2026

ObjectMorpher: 3D-Aware Image Editing via Deformable 3DGS

CVPR 2026

Achieving precise, object-level control in image editing remains challenging: 2D methods lack 3D awareness and often yield ambiguous or implausible results, while existing 3D-aware approaches rely on heavy optimization or incomplete monocular reconstructions. We present ObjectMorpher, a unified, int

Cited by 0SourceScholar
2026

SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation

AAAI 2026technical

Egocentric human pose estimation (HPE) plays a crucial role in immersive applications such as virtual and augmented reality. However, existing methods relying on either visual or sparse inertial data alone often suffer from occlusion or ill-posed problems. In this work, we propose SAME, a novel spat

Cited by 0SourcePDFScholar
2025

EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs

AAAI 2025technical

Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to self-occlusion in images or the sparseness and drift of inert…

Cited by 2SourcePDFScholar
2025

SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix

ICLR 2025poster

Video generation models have demonstrated great capability of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free approach for generating 3D stereoscopic videos using an off-the-shelf monocular video…

2024

EIVEN: Efficient Implicit Attribute Value Extraction using Multimodal LLM

NAACL 2024industry

In e-commerce, accurately extracting product attribute values from multimodal data is crucial for improving user experience and operational efficiency of retailers. However, previous approaches to multimodal attribute value extraction often struggle with implicit attribute values embedded in images…

2024

GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction

ECCV 2024poster

"We present GSD, a diffusion model approach based on Gaussian Splatting (GS) representation for 3D object reconstruction from a single view. Prior works suffer from inconsistent 3D geometry or mediocre rendering quality due to improper representations. We take a step towards resolving these shortcom…

Cited by 7SourcePDFScholar
2024

Generative Human Motion Stylization in Latent Space

ICLR 2024poster

Human motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction a…

Cited by 13SourcePDFScholar
2024

HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations

CVPR 2024poster

It is especially challenging to achieve real-time human motion tracking on a standalone VR Head-Mounted Display (HMD) such as Meta Quest and PICO. In this paper we propose HMD-Poser the first unified approach to recover full-body motions using scalable sparse observations from HMD and body-worn IMUs…

2024

Sequential LLM Framework for Fashion Recommendation

EMNLP 2024industry

The fashion industry is one of the leading domains in the global e-commerce sector, prompting major online retailers to employ recommendation systems for product suggestions and customer convenience. While recommendation systems have been widely studied, most are designed for general e-commerce prob…

2024

TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling

ECCV 2024poster

"Given a 3D mesh, we aim to synthesize 3D textures that correspond to arbitrary textual descriptions. Current methods for generating and assembling textures from sampled views often result in prominent seams or excessive smoothing. To tackle these issues, we present TexGen, a novel multi-view sampli…

2024

Total-Decom: Decomposed 3D Scene Reconstruction with Minimal Interaction

CVPR 2024highlight

Scene reconstruction from multi-view images is a fundamental problem in computer vision and graphics. Recent neural implicit surface reconstruction methods have achieved high-quality results; however editing and manipulating the 3D geometry of reconstructed scenes remains challenging due to the abse…

2023

CL-NeRF: Continual Learning of Neural Radiance Fields for Evolving Scene Representation

NeurIPS 2023poster

Existing methods for adapting Neural Radiance Fields (NeRFs) to scene changes require extensive data capture and model retraining, which is both time-consuming and labor-intensive. In this paper, we tackle the challenge of efficiently adapting NeRFs to real-world scene changes over time using a few…

Cited by 10SourcePDFScholar
2023

CLIPPING: Distilling CLIP-Based Models With a Student Base for Video-Language Retrieval

CVPR 2023poster

Pre-training a vison-language model and then fine-tuning it on downstream tasks have become a popular paradigm. However, pre-trained vison-language models with the Transformer architecture usually take long inference time. Knowledge distillation has been an efficient technique to transfer the capabi…

Cited by 47SourcePDFScholar
2023

Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the Wild

NeurIPS 2023poster

This paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependen…

2023

HiVLP: Hierarchical Interactive Video-Language Pre-Training

ICCV 2023poster

Video-Language Pre-training (VLP) has become one of the most popular research topics in deep learning. However, compared to image-language pre-training, VLP has lagged far behind due to the lack of large amounts of video-text pairs. In this work, we train a VLP model with a hybrid of image-text and…

Cited by 6PDFScholar
2023

Hybrid Neural Rendering for Large-Scale Scenes With Motion Blur

CVPR 2023poster

Rendering novel view images is highly desirable for many applications. Despite recent progress, it remains challenging to render high-fidelity and view-consistent novel views of large-scale scenes from in-the-wild images with inevitable artifacts (e.g., motion blur). To this end, we develop a hybrid…

2023

ISS: Image as Stepping Stone for Text-Guided 3D Shape Generation

ICLR 2023top-25%

Text-guided 3D shape generation remains challenging due to the absence of large paired text-shape dataset, the substantial semantic gap between these two modalities, and the structural complexity of 3D shapes. This paper presents a new framework called Image as Stepping Stone (ISS) for the task by i…

2023

Learning a Room with the Occ-SDF Hybrid: Signed Distance Function Mingled with Occupancy Aids Scene Representation

ICCV 2023poster

Implicit neural rendering, using signed distance function (SDF) representation with geometric priors like depth or surface normal, has made impressive strides in the surface reconstruction of large-scale scenes. However, applying this method to reconstruct a room-level scene from images may miss str…

Cited by 13PDFcodeScholar
2023

Texture Generation on 3D Meshes with Point-UV Diffusion

ICCV 2023oral

In this work, we focus on synthesizing high-quality textures on 3D meshes. We present Point-UV diffusion, a coarse-to-fine pipeline that marries the denoising diffusion model with UV mapping to generate 3D consistent and high-quality texture images in UV space. We start with introducing a point diff…

Cited by 45PDFcodeScholar
2022

Decompose the Sounds and Pixels, Recompose the Events

AAAI 2022technical

In this paper, we propose a framework centering around a novel architecture called the Event Decomposition Recomposition Network (EDRNet) to tackle the Audio-Visual Event (AVE) localization problem in the supervised and weakly supervised settings. AVEs in the real world exhibit common unraveling pat…

Cited by 3SourcePDFScholar
2022

Dual Perspective Network for Audio-Visual Event Localization

ECCV 2022poster

"The Audio-Visual Event Localization (AVEL) problem involves tackling three core sub-tasks: the creation of efficient audio-visual representations using cross-modal guidance, the formation of short-term temporal feature aggregations, and its accumulation to achieve long-term dependency resolution. T…

Cited by 21SourcePDFScholar
2022

Self-Supervised Spatiotemporal Representation Learning by Exploiting Video Continuity

AAAI 2022technical

Recent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored property of videos, the textit{video continuity}, to obtain supervision signals…

Cited by 33SourcePDFScholar
2022

Towards Efficient and Scale-Robust Ultra-High-Definition Image Demoiréing

ECCV 2022poster

"With the rapid development of mobile devices, modern widely-used mobile phones typically allow users to capture 4K resolution (i.e., ultra-high-definition) images. However, for image demoiréing, a challenging task in low-level vision, existing works are generally carried out on low-resolution or sy…

2022

Video Demoireing With Relation-Based Temporal Consistency

CVPR 2022poster

Moire patterns, appearing as color distortions, severely degrade the image and video qualities when filming a screen with digital cameras. Considering the increasing demands for capturing videos, we study how to remove such undesirable moire patterns in videos, namely video demoireing. To this end,…

Cited by 27PDFcodeScholar
2021

Boosting the Generalization Capability in Cross-Domain Few-Shot Learning via Noise-Enhanced Supervised Autoencoder

ICCV 2021poster

State of the art (SOTA) few-shot learning (FSL) methods suffer significant performance drop in the presence of domain differences between source and target datasets. The strong discrimination ability on the source dataset does not necessarily translate to high classification accuracy on the target d…

Cited by 79PDFScholar
2021

Class Semantics-Based Attention for Action Detection

ICCV 2021poster

Action localization networks are often structured as a feature encoder sub-network and a localization sub-network, where the feature encoder learns to transform an input video to features that are useful for the localization sub-network to generate reliable action proposals. While some of the encode…

Cited by 83PDFScholar
2021

Graph-Enhanced Multi-Task Learning of Multi-Level Transition Dynamics for Session-based Recommendation

AAAI 2021technical

Session-based recommendation plays a central role in a wide spectrum of online applications, ranging from e-commerce to online advertising services. However, the majority of existing session-based recommendation techniques (e.g., attention-based recurrent network or graph neural network) are not wel…

2021

Knowledge-Enhanced Hierarchical Graph Transformer Network for Multi-Behavior Recommendation

AAAI 2021technical

Accurate user and item embedding learning is crucial for modern recommender systems. However, most existing recommendation techniques have thus far focused on modeling users' preferences over singular type of user-item interactions. Many practical recommendation scenarios involve multi-typed user in…

2021

Knowledge-aware Coupled Graph Neural Network for Social Recommendation

AAAI 2021technical

Social recommendation task aims to predict users' preferences over items with the incorporation of social connections among users, so as to alleviate the sparse issue of collaborative filtering. While many recent efforts show the effectiveness of neural network-based social recommender systems, seve…

2021

Learning a Proposal Classifier for Multiple Object Tracking

CVPR 2021poster

The recent trend in multiple object tracking (MOT) is heading towards leveraging deep learning to boost the tracking performance. However, it is not trivial to solve the data-association problem in an end-to-end fashion. In this paper, we propose a novel proposal-based learnable framework, which mod…

Cited by 140PDFcodeScholar
2021

Spatial-Temporal Sequential Hypergraph Network for Crime Prediction with Dynamic Multiplex Relation Learning

IJCAI 2021poster

Crime prediction is crucial for public safety and resource optimization, yet is very challenging due to two aspects: i) the dynamics of criminal patterns across time and space, crime events are distributed unevenly on both spatial and temporal domains; ii) time-evolving dependencies between differen…

2020

Cross-Interaction Hierarchical Attention Networks for Urban Anomaly Prediction

IJCAI 2020poster

Predicting anomalies (e.g., blocked driveway and vehicle collisions) in urban space plays an important role in assisting governments and communities for building smart city applications, ranging from intelligent transportation to public safety. However, predicting urban anomalies is not trivial due…

Cited by 0SourcePDFScholar
2020

Towards Efficient Coarse-to-Fine Networks for Action and Gesture Recognition

ECCV 2020poster

State-of-the-art approaches to video-based action and gesture recognition often employ two key concepts: First, they employ multistream processing; second, they use an ensemble of convolutional networks. We improve and extend both aspects. First, we systematically yield enhanced receptive fields for…

Cited by 16SourcePDFScholar
2020

Weight Excitation: Built-in Attention Mechanisms in Convolutional Neural Networks

ECCV 2020poster

We propose novel approaches for simultaneously identifying important weights of a convolutional neural network (ConvNet) and providing more attention to the important weights during training. More formally, we identify two characteristics of a weight, its magnitude and its location, which can be lin…