← Search

Yutong He

17 accepted papers

2026

Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation

ICLR 2026poster

Although diffusion models have achieved strong results in decision-making tasks, their slow inference speed remains a key limitation. While consistency models offer a potential solution, existing applications to decision-making either struggle with suboptimal demonstrations under behavior cloning or…

Cited by 0SourceScholar
2026

Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps

ICML 2026poster

Flow and diffusion models produce high-quality samples, but adapting them to user preferences or constraints post-training remains costly and brittle, a challenge commonly called reward alignment. We argue that efficient reward alignment should be a property of the generative model itself, not an af…

Cited by 0SourceScholar
2026

Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models

ICLR 2026poster

Log-likelihood evaluation enables important capabilities in generative models, including model comparison, certain fine-tuning objectives, and many downstream applications. Yet paradoxically, some of today's best generative models -- diffusion and flow-based models -- still require hundreds to thous…

Cited by 0SourceScholar
2025

MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling

NeurIPS 2025poster

The substantial memory demands of pre-training and fine-tuning large language models (LLMs) require memory-efficient optimization algorithms. One promising approach is layer-wise optimization, which treats each transformer block as a single layer and optimizes it sequentially, while freezing the oth…

Cited by 0SourceScholar
2025

MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization

NeurIPS 2025poster

As distributed optimization scales to meet the demands of Large Language Model (LLM) training, hardware failures become increasingly non-negligible. Existing fault-tolerant training methods often introduce significant computational or memory overhead, demanding additional resources. To address this…

Cited by 0SourceScholar
2025

Subspace Optimization for Large Language Models with Convergence Guarantees

ICML 2025poster

Subspace optimization algorithms, such as GaLore (Zhao et al., 2024), have gained attention for pre-training and fine-tuning large language models (LLMs) due to their memory efficiency. However, their convergence guarantees remain unclear, particularly in stochastic settings. In this paper, we revea…

2024

Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion

ICLR 2024poster

Consistency Models (CM) (Song et al., 2023) accelerate score-based diffusion model sampling at the cost of sample quality but lack a natural way to trade-off quality for speed. To address this limitation, we propose Consistency Trajectory Model (CTM), a generalization encompassing CM and score-based…

2024

Distributed Bilevel Optimization with Communication Compression

ICML 2024poster

Stochastic bilevel optimization tackles challenges involving nested optimization structures. Its fast-growing scale nowadays necessitates efficient distributed algorithms. In conventional distributed bilevel methods, each worker must transmit full-dimensional stochastic gradients to the server every…

Cited by 2SourcePDFScholar
2024

Manifold Preserving Guided Diffusion

ICLR 2024poster

Despite the recent advancements, conditional image generation still faces challenges of cost, generalizability, and the need for task-specific training. In this paper, we propose Manifold Preserving Guided Diffusion (MPGD), a training-free conditional generation framework that leverages pretrained d…

Cited by 50SourcePDFScholar
2023

CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual Representations

ICML 2023poster

Geo-tagged images are publicly available in large quantities, whereas labels such as object classes are rather scarce and expensive to collect. Meanwhile, contrastive learning has achieved tremendous success in various natural image and language tasks with limited labeled data. However, existing met…

Cited by 72SourcePDFScholar
2023

Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?

NeurIPS 2023poster

Communication compression is a common technique in distributed optimization that can alleviate communication overhead by transmitting compressed gradients and model parameters. However, compression can introduce information distortion, which slows down convergence and incurs more communication round…

Cited by 9SourcePDFScholar
2022

Comparing Distributions by Measuring Differences that Affect Decision Making

ICLR 2022oral

Measuring the discrepancy between two probability distributions is a fundamental problem in machine learning and statistics. We propose a new class of discrepancies based on the optimal loss for a decision task -- two distributions are different if the optimal decision loss is higher on their mixtur…

Cited by 34SourcePDFScholar
2022

SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

ICLR 2022poster

Guided image synthesis enables everyday users to create and edit photo-realistic images with minimum effort. The key challenge is balancing faithfulness to the user inputs (e.g., hand-drawn colored strokes) and realism of the synthesized images. Existing GAN-based methods attempt to achieve such bal…

2022

SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery

NeurIPS 2022accept

Unsupervised pre-training methods for large vision models have shown to enhance performance on downstream supervised tasks. Developing similar techniques for satellite imagery presents significant opportunities as unlabelled data is plentiful and the inherent temporal and multi-spectral structure pr…

2021

Spatial-Temporal Super-Resolution of Satellite Imagery via Conditional Pixel Synthesis

NeurIPS 2021poster

High-resolution satellite imagery has proven useful for a broad range of tasks, including measurement of global human population, local economic livelihoods, and biodiversity, among many others. Unfortunately, high-resolution imagery is both infrequently collected and expensive to purchase, making i…

2020

Fine-Grained Image-to-Image Transformation Towards Visual Recognition

CVPR 2020poster

Existing image-to-image transformation approaches primarily focus on synthesizing visually pleasing data. Generating images with correct identity labels is challenging yet much less explored. It is even more challenging to deal with image transformation tasks with large deformation in poses, viewpoi…

Cited by 35PDFScholar