← Search

Lingyun Sun

20 accepted papers

2026

Circular-DPO: Aligning Multi-Stage 3D Generative Models via Preference Feedback Loop

CVPR 2026

Multi-stage generative models have shown great promise in 3D content creation due to focused generation of structure or texture in different stages, but their outputs often fail to align with human preferences. The key bottleneck to apply alignment methods is the presence of non-differentiable opera

Cited by 0SourceScholar
2026

Diffusion Distillation with Direct Preference Optimization for Efficient 3D LiDAR Scene Completion

AAAI 2026technical

The slow sampling speed of diffusion models hinders their application in 3D LiDAR scene completion. To address this, we propose Distillation-DPO, a novel framework that accelerates sampling through score distillation while simultaneously enhancing generation quality via preference alignment. Disti

Cited by 0SourcePDFScholar
2026

IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by $\textit{IChing}$

ICML 2026poster

Vector Quantization (VQ) has been widely used in visual and audio representation due to its effectiveness in compressing high-dimensional signals. However, existing VQ methods often rely on large and unstructured codebooks, which leads to inefficient code utilization and frequent codebook collapse. …

Cited by 0SourceScholar
2026

Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models

ICML 2026poster

Flow Matching models have demonstrated strong performance across a wide range of generative tasks. However, their reliance on ODE-based iterative sampling incurs substantial computational overhead, which limits their applicability in real-time scenes. While distillation is a promising solution, exis…

Cited by 0SourceScholar
2026

When Diffusion Language Models Hesitate: Detecting and Correcting Visual Hallucinations via Confidence Fluctuation

ICML 2026poster

Multi-modal Diffusion Language Models (MDLMs) have emerged as a powerful alternative to autoregressive models in vision-language understanding, offering advantages in bidirectional context modeling and parallel decoding. However, existing MDLMs suffer from severe visual hallucinations due to the sta…

Cited by 0SourceScholar
2025

Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion

ICCV 2025poster

Diffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the practical application of diffusion-based scene completion models since autonomous vehicles require an efficient perception o…

2025

Distribution Backtracking Builds A Faster Convergence Trajectory for Diffusion Distillation

ICLR 2025poster

Accelerating the sampling speed of diffusion models remains a significant challenge. Recent score distillation methods distill a heavy teacher model into a student generator to achieve one-step generation, which is optimized by calculating the difference between two score functions on the samples ge…

2025

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

AAAI 2025technical

Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative diversity, primarily due to inadequate feature alignment methods…

Cited by 3SourcePDFScholar
2025

Hand by Hand: LLM Driving EMS Assistant for Operational Skill Learning

IJCAI 2025

Operational skill learning, inherently physical and reliant on hands-on practice and kinesthetic feedback, has yet to be effectively replicated in large language model (LLM)-supported training. Current LLM training assistants primarily generate customized textual feedback, neglecting the crucial kin

2025

Integrating Sequence and Image Modeling in Irregular Medical Time Series Through Self-Supervised Learning

AAAI 2025technical

Medical time series are often irregular and face significant missingness, posing challenges for data analysis and clinical decision-making. Existing methods typically adopt a single modeling perspective, either treating series data as sequences or transforming them into image representations for fur…

2025

Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning

AAAI 2025technical

Dynamic Music Emotion Recognition (DMER) aims to predict the emotion of different moments in music, playing a crucial role in music information retrieval. The existing DMER methods struggle to capture long-term dependencies when dealing with sequence data, which limits their performance. Furthermore…

2024

Rapid 3D Model Generation with Intuitive 3D Input

CVPR 2024highlight

With the emergence of AR/VR 3D models are in tremendous demand. However conventional 3D modeling with Computer-Aided Design software requires much expertise and is difficult for novice users. We find that AR/VR devices in addition to serving as effective display mediums can offer a promising potenti…

Cited by 5SourcePDFScholar
2024

Reducing Spatial Fitting Error in Distillation of Denoising Diffusion Models

AAAI 2024technical

Denoising Diffusion models have exhibited remarkable capabilities in image generation. However, generating high-quality samples requires a large number of iterations. Knowledge distillation for diffusion models is an effective method to address this limitation with a shortened sampling process but c…

2023

Deep3DSketch: 3D Modeling from Free-Hand Sketches with View- and Structural-Aware Adversarial Training

ICASSP 2023accepted

This work aims to investigate the problem of 3D modeling using single free-hand sketches, which is one of the most natural ways we humans express ideas. Although sketch-based 3D modeling can drastically make the 3D modeling process more accessible, the sparsity and ambiguity of sketches bring signif…

Cited by 0SourceScholar
2023

Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation

ICCV 2023poster

Audio-driven talking-head synthesis is a popular research topic for virtual human-related applications. However, the inflexibility and inefficiency of existing methods, which necessitate expensive end-to-end training to transfer emotions from guidance videos to talking-head predictions, are signific…

Cited by 122PDFcodeScholar
2023

Learning Object Consistency and Interaction in Image Generation from Scene Graphs

IJCAI 2023poster

This paper is concerned with synthesizing images conditioned on a scene graph (SG), a set of object nodes and their edges of interactive relations. We divide existing works into image-oriented and code-oriented methods. In our analysis, the image-oriented methods do not consider object interaction i…

2023

Preserving Structural Consistency in Arbitrary Artist and Artwork Style Transfer

AAAI 2023technical

Deep generative models are effective in style transfer. Previous methods learn one or several specific artist-style from a collection of artworks. These methods not only homogenize the artist-style of different artworks of the same artist but also lack generalization for the unseen artists. To solv…

Cited by 5SourcePDFScholar
2021

Image Synthesis From Layout With Locality-Aware Mask Adaption

ICCV 2021poster

This paper is concerned with synthesizing images conditioned on a layout (a set of bounding boxes with object categories). Existing works construct a layout-mask-image pipeline. Object masks are generated separately and mapped to bounding boxes to form a whole semantic segmentation mask (layout-to-m…

Cited by 80PDFcodeScholar