← Search

Haoyuan Sun

16 accepted papers

2026

CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation

CVPR 2026

Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for semantic synthesis, while others rely on localization cues for spatial precision. Forcing these heterogeneous tasks to

Cited by 3SourcecodeScholar
2026

Principled RL for Flow Matching Emerges From the Chunk-level Policy Optimization

ICML 2026poster

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive …

Cited by 0SourceScholar
2026

The Secret Engine Behind RLHF: It's Contarstive Learning All Along

ICML 2026poster

Alignment of large language models (LLMs) with human values has recently garnered significant attention, with prominent examples including the canonical yet costly Reinforcement Learning from Human Feedback (RLHF) and the simple Direct Preference Optimization (DPO). In this work, we demonstrate that…

Cited by 0SourceScholar
2026

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

ICLR 2026poster

Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e., visual tokens) versus the limited context length of language models. Traditional uniform sampling often leads to sele…

Cited by 0SourcecodeScholar
2025

Entropy-based Activation Function Optimization: A Method on Searching Better Activation Functions

ICLR 2025poster

The success of artificial neural networks (ANNs) hinges greatly on the judicious selection of an activation function, introducing non-linearity into network and enabling them to model sophisticated relationships in data. However, the search of activation functions has largely relied on empirical kno…

Cited by 0SourcePDFScholar
2025

FloorPlan-LLaMa: Aligning Architects’ Feedback and Domain Knowledge in Architectural Floor Plan Generation

ACL 2025long

Floor plans serve as a graphical language through which architects sketch and communicate their design ideas. Actually, in the Architecture, Engineering, and Construction (AEC) design stages, generating floor plans is a complex task requiring domain expertise and alignment with user requirements. Ho…

Cited by 0SourcePDFScholar
2025

Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence Minimization

AAAI 2025technical

Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches re…

Cited by 4SourcePDFScholar
2025

Identical Human Preference Alignment Paradigm for Text-to-Image Models

ICASSP 2025accepted

Implicit reward mechanism of Direct Preference Optimization (DPO) has facilitated its recent applications beyond large language models (LLMs), notably in aligning text-to-image models with human preferences. While promising results have been achieved with algorithms such as Diffusion-DPO, their reli…

Cited by 0SourceScholar
2025

Learning Statistical and Physical Modeling for Consistency Human Motion Prediction

ICASSP 2025accepted

Diffusion denoising models have great potential in generating diverse and realistic human motions. However, despite the impressive performance of existing methods, they still face some issues. The diffusion process often significantly overlooks physical laws, leading to physically implausible motion…

Cited by 0SourceScholar
2025

Positive Enhanced Preference Alignment for Text-to-Image Models

ICASSP 2025accepted

Direct Preference Optimization (DPO) has recently expanded its successful application beyond aligning large language models (LLMs), further targeting the alignment of text-to-image models with human preferences. However, traditional DPO approach would inadvertently result in a simultaneous reduction…

Cited by 0SourceScholar
2025

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

NeurIPS 2025poster

Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO,…

Cited by 0SourceScholar
2022

Mirror Descent Maximizes Generalized Margin and Can Be Implemented Efficiently

NeurIPS 2022accept

Driven by the empirical success and wide use of deep neural networks, understanding the generalization performance of overparameterized models has become an increasingly popular question. To this end, there has been substantial effort to characterize the implicit bias of the optimization algorithms…

Cited by 27SourcePDFScholar
2021

Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying Systems

NeurIPS 2021spotlight

We study predictive control in a setting where the dynamics are time-varying and linear, and the costs are time-varying and well-conditioned. At each time step, the controller receives the exact predictions of costs, dynamics, and disturbances for the future $k$ time steps. We show that when the pre…

Cited by 46SourcePDFScholar
2019

Beyond Online Balanced Descent: An Optimal Algorithm for Smoothed Online Optimization

NeurIPS 2019spotlight

We study online convex optimization in a setting where the learner seeks to minimize the sum of a per-round hitting cost and a movement cost which is incurred when changing decisions between rounds. We prove a new lower bound on the competitive ratio of any online algorithm in the setting where the…

Cited by 79SourcePDFScholar