← Search

Dandan Zheng

16 accepted papers

2026

Perceptual Flow Network for Visually Grounded Reasoning

ICML 2026poster

Despite the success of LVLMs, general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories, leading to language bias and hallucination. To mitigate this, current methods introduce geometric priors from visual experts as additional supervision. However, we observe that s…

Cited by 0SourceScholar
2026

TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion Generation

ICLR 2026poster

Text-to-motion generation, a rapidly evolving field in computer vision, aims to produce realistic and text-aligned motion sequences. Current methods primarily focus on spatial-temporal modeling or independent frequency domain analysis, lacking a unified framework for joint optimization across spatia…

Cited by 0SourcecodeScholar
2026

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

AAAI 2026technical

The remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual a

Cited by 0SourcePDFScholar
2026

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

CVPR 2026

The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a sig

Cited by 0SourceScholar
2026

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

IJCAI 2026

The rapid advancement of AIGC video generation calls for evaluation frameworks that move beyond technical fidelity and incorporate human-centered aesthetic assessment. Existing benchmarks often overlook fine-grained perceptual qualities such as visual aesthetics, artistic style, and human preference

Cited by 0Scholar
2025

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

NeurIPS 2025poster

We propose a novel AutoRegressive Generation-based paradigm for image Segmentation (ARGenSeg), achieving multimodal understanding and pixel-level perception within a unified framework. Prior works integrating image segmentation into multimodal large language models (MLLMs) typically employ either b…

Cited by 0SourceScholar
2025

Animate-X: Universal Character Image Animation with Enhanced Motion Representation

ICLR 2025poster

Character image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not generalize well on anthropomorphic characters commonly used…

Cited by 14SourcePDFScholar
2025

Mimir: Improving Video Diffusion Models for Precise Text Understanding

CVPR 2025poster

Text serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features from text encoders yet struggle with limited text comprehension. The recent success of large language models (LLMs) show…

Cited by 4SourcePDFScholar
2025

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation

CVPR 2025poster

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such…

Cited by 4SourcePDFScholar
2025

Reversing Flow for Image Restoration

CVPR 2025poster

Image restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, wh…

Cited by 0SourcePDFScholar
2025

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

NeurIPS 2025poster

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-bas…

Cited by 0SourcecodeScholar
2025

VideoMAR: Autoregressive Video Generation with Continuous Tokens

NeurIPS 2025poster

Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored. Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. Howev…

Cited by 0SourceScholar
2025

Visual-Instructed Degradation Diffusion for All-in-One Image Restoration

CVPR 2025poster

Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that…

2024

SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models

ECCV 2024poster

"Text-to-image diffusion models (SD) exhibit significant advancements while requiring extensive computational resources. Existing acceleration methods usually require extensive training and are not universally applicable. LCM-LoRA, trainable once for diverse models, offers universality but rarely co…

2022

Leveraging Inter-Layer Dependency for Post -Training Quantization

NeurIPS 2022accept

Prior works on Post-training Quantization (PTQ) typically separate a neural network into sub-nets and quantize them sequentially. This process pays little attention to the dependency across the sub-nets, hence is less optimal. In this paper, we propose a novel Network-Wise Quantization (NWQ) approac…

Cited by 23SourcePDFScholar