← Search

Fa-ting Hong

14 accepted papers

2026

PhysiGen: Integrating Collision-Aware Physical Constraints for High-Fidelity Human-Human Interaction Generation

ICASSP 2026poster

Despite substantial progress in text-driven 3D human motion synthesis, generating realistic multi-person interaction sequences remains challenging. Notably, body inter-penetration is a pervasive issue from both data acquisition to the generated results, which significantly undermines the realism and…

Cited by 0SourcePDFScholar
2025

AnyTalk: Multi-modal Driven Multi-domain Talking Head Generation

AAAI 2025technical

Cross-domain talking head generation, such as animating a static cartoon animal photo with real human video, is crucial for personalized content creation. However, prior works typically rely on domain-specific frameworks and paired videos, limiting its utility and complicating its architecture with…

Cited by 0SourcePDFScholar
2025

Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

ICCV 2025poster

Talking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end, we introduce ACTalker, an end-to-end video diffusion framewo…

2025

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

CVPR 2025poster

Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of visual language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-gra…

Cited by 2SourcePDFScholar
2025

Free-viewpoint Human Animation with Pose-correlated Reference Selection

CVPR 2025highlight

Diffusion-based human animation aims to animate a human character based on a source human image as well as driving signals such as a sequence of poses. Leveraging the generative capacity of diffusion model, existing approaches are able to generate high-fidelity poses, but struggle with significant v…

Cited by 1SourcePDFScholar
2025

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

CVPR 2025poster

We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the chara…

2025

Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation

CVPR 2025poster

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a challenging and critical problem to generate videos with accurate…

2023

Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head video Generation

ICCV 2023poster

Talking head video generation aims to animate a human face in a still image with dynamic poses and expressions using motion information derived from a target-driving video, while maintaining the person's identity in the source image. However, dramatic and complex motions in the driving video cause a…

Cited by 45PDFcodeScholar
2022

Depth-Aware Generative Adversarial Network for Talking Head Video Generation

CVPR 2022poster

Talking head video generation aims to produce a synthetic human face video that contains the identity and pose information respectively from a given source image and a driving video. Existing works for this task heavily rely on 2D representations (e.g. appearance and motion) learned from the input i…

Cited by 203PDFcodeScholar
2021

MIST: Multiple Instance Self-Training Framework for Video Anomaly Detection

CVPR 2021poster

Weakly supervised video anomaly detection (WS-VAD) is to distinguish anomalies from normal events based on discriminative representations. Most existing works are limited in insufficient video representations. In this work, we develop a multiple instance self-training framework (MIST) to efficiently…

Cited by 342PDFcodeScholar
2020

Learning to Detect Important People in Unlabelled Images for Semi-Supervised Important People Detection

CVPR 2020poster

Important people detection is to automatically detect the individuals who play the most important roles in a social event image, which requires the designed model to understand a high-level pattern. However, existing methods rely heavily on supervised learning using large quantities of annotated ima…

Cited by 21PDFScholar
2020

MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection

ECCV 2020poster

We address the weakly supervised video highlight detection problem for learning to detect segments that are more attractive in training videos given their video event label but without expensive supervision of manually annotating highlight segments. While manually averting localizing highlight segme…

Cited by 87SourcePDFScholar