← Search

Bang Zhang

21 accepted papers

2026

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

CVPR 2026

Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment models often sacrifice granular control for expressiveness, while methods designed for control may struggle with fidelity and robust disentanglement. Ins

Cited by 0SourceScholar
2026

RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents

ICLR 2026poster

Large language models (LLMs) excel at logical and algorithmic reasoning, yet their emotional intelligence (EQ) still lags far behind their cognitive prowess. While reinforcement learning from verifiable rewards (RLVR) has advanced in other domains, its application to dialogue—especially for emotion…

Cited by 0SourcecodeScholar
2025

Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

ICCV 2025poster

Recent character image animation methods based on diffusion models, such as Animate Anyone, have made significant progress in generating consistent and generalizable character animations. However, these approaches fail to produce reasonable associations between characters and their environments. To…

Cited by 0SourcePDFScholar
2025

Controllable and Expressive One-Shot Video Head Swapping

ICCV 2025poster

In this paper, we propose a novel diffusion-based multi-condition controllable framework for video head swapping, which seamlessly transplant a human head from a static image into a dynamic video, while preserving the original body and background of target video, and further allowing to tweak head e…

Cited by 0SourcePDFScholar
2025

EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers

ICML 2025poster

Removing unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable Diffusion (SD) v3 and Flux, which incorporate flow matching an…

2025

Exploring Timeline Control for Facial Motion Generation

CVPR 2025poster

This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arran…

Cited by 0SourcePDFScholar
2025

ExtPose: Robust and Coherent Pose Estimation by Extending ViTs

ICML 2025poster

Vision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel…

Cited by 0SourcePDFScholar
2025

OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking

NeurIPS 2025poster

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video content from input text while emulating the target identity'…

Cited by 0SourceScholar
2025

S2R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

ACL 2025long

Recent studies have demonstrated the effectiveness of LLM test-time scaling. However, existing approaches to incentivize LLMs’ deep thinking abilities generally require large-scale data or significant training efforts. Meanwhile, it remains unclear how to improve the thinking abilities of less power…

2025

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

NeurIPS 2025poster

Evaluating the step-by-step reliability of large language model (LLM) reasoning, such as Chain-of-Thought, remains challenging due to the difficulty and cost of obtaining high-quality step-level supervision. In this paper, we introduce Self-Play Critic (SPC), a novel approach where a critic model ev…

Cited by 0SourceScholar
2024

Cross-Sentence Gloss Consistency for Continuous Sign Language Recognition

AAAI 2024technical

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from continuous sign videos. Recent works enhance the gloss representation consistency by mining correlations between visual and contextual modules within individual sentences. However, there still remain much richer corre…

Cited by 2SourcePDFScholar
2023

Gloss-Free End-to-End Sign Language Translation

ACL 2023long

In this paper, we tackle the problem of sign language translation (SLT) without gloss annotations. Although intermediate representation like gloss has been proven effective, gloss annotations are hard to acquire, especially in large quantities. This limits the domain coverage of translation datasets…

2023

One-Shot High-Fidelity Talking-Head Synthesis With Deformable Neural Radiance Field

CVPR 2023poster

Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encou…

Cited by 54SourcePDFScholar
2023

RenderIH: A Large-Scale Synthetic Dataset for 3D Interacting Hand Pose Estimation

ICCV 2023poster

The current interacting hand (IH) datasets are relatively simplistic in terms of background and texture, with hand joints being annotated by a machine annotator, which may result in inaccuracies, and the diversity of pose distribution is limited. However, the variability of background, pose distribu…

Cited by 19PDFcodeScholar
2023

Towards Stable Human Pose Estimation via Cross-View Fusion and Foot Stabilization

CVPR 2023poster

Towards stable human pose estimation from monocular images, there remain two main dilemmas. On the one hand, the different perspectives, i.e., front view, side view, and top view, appear the inconsistent performances due to the depth ambiguity. On the other hand, foot posture plays a significant rol…

Cited by 5SourcePDFScholar
2022

A Speech-driven Sign Language Avatar Animation System for Hearing Impaired Applications

IJCAI 2022poster

Sign language is the communication language used in hearing impaired community. Recently, the research of sign language production has made great progress but still need to cope with some critical challenges. In this paper, we propose a system-level scheme and push forward the implementation of sign…

Cited by 6SourcePDFScholar
2022

DART: Articulated Hand Model with Diverse Accessories and Rich Textures

NeurIPS 2022accept

Hand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits it…

2022

Multi-View Consistent Generative Adversarial Networks for 3D-Aware Image Synthesis

CVPR 2022poster

3D-aware image synthesis aims to generate images of objects from multiple views by learning a 3D representation. However, one key challenge remains: existing approaches lack geometry constraints, hence usually fail to generate multi-view consistent images. To address this challenge, we propose Multi…

Cited by 55PDFcodeScholar
2022

Recurrent Dynamic Embedding for Video Object Segmentation

CVPR 2022poster

Space-time memory (STM) based video object segmentation (VOS) networks usually keep increasing memory bank every several frames, which shows excellent performance. However, 1) the hardware cannot withstand the ever-increasing memory requirements as the video length increases. 2) Storing lots of info…

Cited by 95PDFcodeScholar
2021

Learning Position and Target Consistency for Memory-Based Video Object Segmentation

CVPR 2021poster

This paper studies the problem of semi-supervised video object segmentation(VOS). Multiple works have shown that memory-based approaches can be effective for video object segmentation. They are mostly based on pixel-level matching, both spatially and temporally. The main shortcoming of memory-based…

Cited by 134PDFScholar