← Search

Yingda Chen

5 accepted papers

2026

Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation

ICML 2026poster

Inference-time scaling offers a versatile paradigm for aligning visual generative models with downstream objectives without parameter updates. However, existing approaches that optimize the high-dimensional initial noise suffer from severe inefficiency, as many search directions exert negligible inf…

Cited by 0SourceScholar
2025

Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion models

NeurIPS 2025poster

Text-to-image diffusion models have gained widespread application across various domains, demonstrating remarkable creative potential. However, the strong generalization capabilities of diffusion models can inadvertently lead to the generation of not-safe-for-work (NSFW) content, posing significant…

Cited by 0SourceScholar
2025

ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning

IJCAI 2025

Recently, advancements in video synthesis have attracted significant attention. Video synthesis models have demonstrated the practical applicability of diffusion models in creating dynamic visual content. Despite these advancements, the extension of video lengths remains constrained by computational

2025

SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning

AAAI 2025technical

Recent development in Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) have achieved superior performance and generalization capabilities, covered extensive areas of traditional tasks. However, existing large model training frameworks support only a limited number of models…

2025

Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video Understanding

ACL 2025long

Despite existing multimodal language models showing impressive performance on the video understanding task, extremely long videos still pose significant challenges to language model’s context length, memory consumption, and computational complexity. To address these issues, we propose a vision-langu…