← Search

Jiazheng Xing

11 accepted papers

2026

CyC3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization

AAAI 2026technical

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remains a significant challenge. Existing methods often struggle to maintain accurate alignment, leading to noticeable discre

Cited by 0SourcePDFScholar
2026

Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective

ICLR 2026poster

Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video generators either diverge from standard LLM architectures, depend on bulky external text encoders, or incur prohibitive la…

Cited by 0SourcecodeScholar
2026

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

ICLR 2026poster

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face–attribute alignment across subjects remains challenging, as existing methods…

Cited by 0SourcecodeScholar
2026

OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization

AAAI 2026technical

Watermarking diffusion-generated images is crucial for copyright protection and user tracking. However, current diffusion watermarking methods face significant limitations: zero-bit watermarking systems lack the capacity for large-scale user tracking, while multi-bit methods are highly sensitive to

Cited by 0SourcePDFScholar
2025

CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion

ICASSP 2025accepted

Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today’s video proliferation era. Multi-modal video summarization that accomodates user input has become a research hot…

Cited by 0SourceScholar
2025

UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback

NeurIPS 2025poster

Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown strong potential by enabling rich and controllable lighting effects. However, as they are typically optimized in semantic latent space, where proximity does not guarantee physical corre…

Cited by 0SourcecodeScholar
2024

A Multimodal, Multi-Task Adapting Framework for Video Action Recognition

AAAI 2024technical

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance a…

Cited by 17SourcePDFScholar
2024

FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio

CVPR 2024poster

In this paper we abstract the process of people hearing speech extracting meaningful cues and creating various dynamically audio-consistent talking faces termed Listening and Imagining into the task of high-fidelity diverse talking faces generation from a single audio. Specifically it involves two c…

2024

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

CVPR 2024poster

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore recent studies have ut…

2023

Boosting Few-shot Action Recognition with Graph-guided Hybrid Matching

ICCV 2023poster

Class prototype construction and matching are core aspects of few-shot action recognition. Previous methods mainly focus on designing spatiotemporal relation modeling modules or complex temporal alignment algorithms. Despite the promising results, they ignored the value of class prototype constructi…

Cited by 37PDFcodeScholar
2023

Revisiting the Spatial and Temporal Modeling for Few-Shot Action Recognition

AAAI 2023technical

Spatial and temporal modeling is one of the most core aspects of few-shot action recognition. Most previous works mainly focus on long-term temporal relation modeling based on high-level spatial representations, without considering the crucial low-level spatial features and short-term temporal relat…

Cited by 45SourcePDFScholar