← Search

Zhan Qu

7 accepted papers

2026

PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs

AAAI 2026technical

This paper presents a systematic investigation into the constrained generation capabilities of large language models (LLMs) in producing Songci, a classical Chinese poetry form characterized by strict structural, tonal, and rhyme constraints defined by Cipai templates. We first develop a comprehensi

Cited by 0SourcePDFScholar
2026

WPT: World-to-Policy Transfer via Online World Model Distillation

CVPR 2026

Recent years have witnessed remarkable progress in world models, which primarily aim to capture the spatiotemporal correlations between an agent's actions and the evolving environment. However, existing approaches often suffer from tight runtime coupling or depend on offline reward signals, resultin

Cited by 0SourceScholar
2025

ExpTalk: Diverse Emotional Expression via Adaptive Disentanglement and Refined Alignment for Speech-Driven 3D Facial Animation

IJCAI 2025

Speech-driven 3D facial animation aims to create lifelike facial expressions that synchronize accurately with speech. Despite significant progress, many existing methods may focus on generating facial animation with a fixed emotional state, neglecting the diverse transformations of facial emotions u

Cited by 0SourcePDFScholar
2025

Prioritizing Perception-Guided Self-Supervision: A New Paradigm for Causal Modeling in End-to-End Autonomous Driving

NeurIPS 2025poster

End-to-end autonomous driving systems, predominantly trained through imitation learning, have demonstrated considerable effectiveness in leveraging large-scale expert driving data. Despite their success in open-loop evaluations, these systems often exhibit significant performance degradation in clos…

Cited by 0SourceScholar
2022

Diversity Matters: Fully Exploiting Depth Clues for Reliable Monocular 3D Object Detection

CVPR 2022oral

As an inherently ill-posed problem, depth estimation from single images is the most challenging part of monocular 3D object detection (M3OD). Many existing methods rely on preconceived assumptions to bridge the missing spatial information in monocular images, and predict a sole depth value for every…

Cited by 78PDFScholar