← Search

Baoyuan Wang

35 accepted papers

2025

Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations

COLING 2025main

Existing retrieval-based methods have made significant strides in maintaining long-term conversations. However, these approaches face challenges in memory database management and accurate memory retrieval, hindering their efficacy in dynamic, real-world interactions. This study introduces a novel fr…

2025

LLM-driven Multimodal and Multi-Identity Listening Head Generation

CVPR 2025poster

Generating natural listener responses in conversational scenarios is crucial for creating engaging digital humans and avatars. Recent work has shown that large language models (LLMs) can be effectively leveraged for this task, demonstrating remarkable capabilities in generating contextually appropri…

Cited by 0SourcePDFScholar
2025

Subobject-level Image Tokenization

ICML 2025poster

Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we introduce subobject-level adaptive token segmentation and explore several approaches, including superpixel, SAM, and a pro…

2024

AvatarGPT: All-in-One Framework for Motion Understanding Planning Generation and Beyond

CVPR 2024poster

Large Language Models(LLMs) have shown remarkable emergent abilities in unifying almost all (if not every) NLP tasks. In the human motion-related realm however researchers still develop siloed models for each task. Inspired by InstuctGPT[??] and the generalist concept behind Gato [??] we introduce A…

Cited by 32SourcePDFScholar
2024

PICTURE: PhotorealistIC virtual Try-on from UnconstRained dEsigns

CVPR 2024poster

In this paper we propose a novel virtual try-on from unconstrained designs (ucVTON) task to enable photorealistic synthesis of personalized composite clothing on input human images. Unlike prior arts constrained by specific input types our method allows flexible specification of style (text or image…

Cited by 8SourcePDFScholar
2024

Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer

ECCV 2024poster

"In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseudo multi-view videos to learn a 4D head synthesizer in a data-driven manner, av…

2024

Portrait4D: Learning One-Shot 4D Head Avatar Synthesis using Synthetic Data

CVPR 2024poster

Existing one-shot 4D head synthesis methods usually learn from monocular videos with the aid of 3DMM reconstruction yet the latter is evenly challenging which restricts them from reasonable 4D head synthesis. We present a method to learn one-shot 4D head synthesis via large-scale synthetic data. The…

Cited by 19SourcePDFScholar
2024

Towards Objectively Benchmarking Social Intelligence of Language Agents at the Action Level

ACL 2024findings

Prominent large language models have exhibited human-level performance in many domains, even enabling the derived agents to simulate human and social interactions. While practical works have substantiated the practicability of grounding language agents in sandbox simulation or embodied simulators, c…

2023

DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language Models

EMNLP 2023long main

Chain-of-Thought (CoT) prompting has successfully enhanced the reasoning capabilities of Large Language Models~(LLMs) with at least 100 billion parameters. However, it is ineffective, or even detrimental, to the performance on reasoning tasks in Smaller Language Models (SLMs) with less than 10 billi…

Cited by 0SourcecodeScholar
2023

Hand Avatar: Free-Pose Hand Animation and Rendering From Monocular Video

CVPR 2023poster

We present HandAvatar, a novel representation for hand animation and rendering, which can generate smoothly compositional geometry and self-occlusion-aware texture. Specifically, we first develop a MANO-HD model as a high-resolution mesh topology to fit personalized hand shapes. Sequentially, we dec…

Cited by 47SourcePDFScholar
2023

Hierarchical Verbalizer for Few-Shot Hierarchical Text Classification

ACL 2023long

Due to the complex label hierarchy and intensive labeling cost in practice, the hierarchical text classification (HTC) suffers a poor performance especially when low-resource or few-shot settings are considered. Recently, there is a growing trend of applying prompts on pre-trained language models (P…

2023

Learning Detailed Radiance Manifolds for High-Fidelity and 3D-Consistent Portrait Synthesis From Monocular Image

CVPR 2023poster

A key challenge for novel view synthesis of monocular portrait images is 3D consistency under continuous pose variations. Most existing methods rely on 2D generative models which often leads to obvious 3D inconsistency artifacts. We present a 3D-consistent novel view synthesis approach for monocular…

2023

LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming

ACL 2023long

Open-domain dialogue systems have made promising progress in recent years. While the state-of-the-art dialogue agents are built upon large-scale social media data and large pre-trained models, there is no guarantee these agents could also perform well in fast-growing scenarios, such as live streamin…

2023

Natural Response Generation for Chinese Reading Comprehension

EMNLP 2023long findings

Machine reading comprehension (MRC) is an important area of conversation agents and draws a lot of attention. However, there is a notable limitation to current MRC benchmarks: The labeled answers are mostly either spans extracted from the target corpus or the choices of the given candidates, ignorin…

Cited by 0SourcecodeScholar
2023

Orca: A Few-shot Benchmark for Chinese Conversational Machine Reading Comprehension

EMNLP 2023long findings

The conversational machine reading comprehension (CMRC) task aims to answer questions in conversations, which has been a hot research topic in recent years because of its wide applications. However, existing CMRC benchmarks in which each conversation is assigned a static passage are inconsistent wit…

Cited by 0SourcecodeScholar
2023

Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head Synthesis

CVPR 2023poster

We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disentangled latent representations and leverage an image generator to synthesize tal…

2023

Reinforced Disentanglement for Face Swapping without Skip Connection

ICCV 2023poster

The SOTA face swap models still suffer the problem of either target identity (i.e., shape) being leaked or the target non-identity attributes (i.e., background, hair) failing to be fully preserved in the final results. We show that this insufficient disentanglement is caused by two flawed designs t…

Cited by 14PDFcodeScholar
2023

Talking Head Generation with Probabilistic Audio-to-Visual Diffusion Priors

ICCV 2023poster

We introduce a novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead sample all holistic lip-irrelevant facial motions (i.e. pose, expression, blink, gaze, etc.) to…

Cited by 41PDFScholar
2022

Beyond 3D Siamese Tracking: A Motion-Centric Paradigm for 3D Single Object Tracking in Point Clouds

CVPR 2022oral

3D single object tracking (3D SOT) in LiDAR point clouds plays a crucial role in autonomous driving. Current approaches all follow the Siamese paradigm based on appearance matching. However, LiDAR point clouds are usually textureless and incomplete, which hinders effective appearance matching. Besid…

Cited by 109PDFcodeScholar
2022

Local-Adaptive Face Recognition via Graph-Based Meta-Clustering and Regularized Adaptation

CVPR 2022poster

Due to the rising concern of data privacy, it's reasonable to assume the local client data can't be transferred to a centralized server, nor their associated identity label is provided. To support continuous learning and fill the last-mile quality gap, we introduce a new problem setup called Local-A…

Cited by 14PDFScholar
2022

Privacy-Preserving Online AutoML for Domain-Specific Face Detection

CVPR 2022poster

Despite the impressive progress of general face detection, the tuning of hyper-parameters and architectures is still critical for the performance of a domain-specific face detector. Though existing AutoML works can speedup such process, they either require tuning from scratch for a new scenario or d…

Cited by 20PDFcodeScholar
2020

JNR: Joint-based Neural Rig Representation for Compact 3D Face Modeling

ECCV 2020poster

In this paper, we introduce a novel approach to learn a 3D face model using a joint-based face rig and a neural skinning network. Thanks to the joint-based representation, our model enjoys some significant advantages over prior blendshape-based models. First, it is very compact such that we are orde…

Cited by 7SourcePDFScholar
2020

Personalized Face Modeling for Improved Face Reconstruction and Motion Retargeting

ECCV 2020poster

Traditional methods for image-based 3D face reconstruction and facial motion retargeting fit a 3D morphable model (3DMM) to the face, which has limited modeling capacity and fail to generalize well to in-the-wild data. Use of deformation transfer or multilinear tensor as a personalized 3DMM for blen…

Cited by 74SourcePDFScholar
2020

ReDA:Reinforced Differentiable Attribute for 3D Face Reconstruction

CVPR 2020oral

The key challenge for 3D face shape reconstruction is to build the correct dense face correspondence between the deformable mesh and the single input image. Given the ill-posed nature, previous works heavily rely on prior knowledge (such as 3DMM [2]) to reduce depth ambiguity. Although impressive re…

Cited by 47PDFScholar
2017

Personalized Cinemagraphs Using Semantic Understanding and Collaborative Learning

ICCV 2017poster

Cinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically pleasing cinemagraph requires isolating objects in a semantically meaningful way an…

Cited by 19PDFScholar
2015

Harvesting Discriminative Meta Objects With Deep CNN Features for Scene Classification

ICCV 2015poster

Recent work on scene classification still makes use of generic CNN features in a rudimentary manner. In this paper, we present a novel pipeline built upon deep CNN features to harvest discriminative visual objects and parts for scene classification. We first use a region proposal technique to genera…

Cited by 153PDFScholar
2015

Unsupervised Extraction of Video Highlights Via Robust Recurrent Auto-Encoders

ICCV 2015poster

With the growing popularity of short-form video sharing platforms such as Instagram and Vine, there has been an increasing need for techniques that automatically extract highlights from video. Whereas prior works have approached this problem with heuristic rules or supervised learning, we present an…

Cited by 221PDFScholar