← Search

Yinxiao Li

15 accepted papers

2025

A Simple Approach to Unifying Diffusion-based Conditional Generation

ICLR 2025poster

Recent progress in image generation has sparked research into controlling these models through condition signals, with various methods addressing specific challenges in conditional generation. Instead of proposing another specialized technique, we introduce a simple, unified framework to handle dive…

Cited by 1SourcePDFScholar
2025

Calibrated Multi-Preference Optimization for Aligning Diffusion Models

CVPR 2025poster

Aligning text-to-image (T2I) diffusion models with prefer-ence optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternative, however, current preference optimization methods fall short in exploiting…

Cited by 5SourcePDFScholar
2025

Cropper: Vision-Language Model for Image Cropping through In-Context Learning

CVPR 2025poster

The goal of image cropping is to identify visually appealing crops in an image. Conventional methods are trained on specific datasets and fail to adapt to new requirements. Recent breakthroughs in large vision-language models (VLMs) enable visual in-context learning without explicit training. Howeve…

Cited by 2SourcePDFScholar
2025

DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes

CVPR 2025poster

The increasing demand for immersive AR/VR applications and spatial intelligence has heightened the need to generate high-quality scene-level and 360deg panoramic video. However, most video diffusion models are constrained by limited resolution and aspect ratio, which restricts their applicability to…

Cited by 5SourcePDFScholar
2025

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

CVPR 2025highlight

Text-to-image (T2I) generation has made significant advances in recent years, but challenges still remain in the generation of perceptual artifacts, misalignment with complex prompts, and safety. The prevailing approach to address these issues involves collecting human feedback on generated images,…

Cited by 2SourcePDFScholar
2025

HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis

NeurIPS 2025poster

We propose HoliGS, a novel deformable Gaussian splatting framework that addresses embodied view synthesis from long monocular RGB videos. Unlike prior 4D Gaussian splatting and dynamic NeRF pipelines, which struggle with training overhead in minute-long captures, our method leverages invertible Gaus…

Cited by 0SourceScholar
2025

Toward Material-Agnostic System Identification from Videos

ICCV 2025poster

System identification from videos aims to recover object geometry and governing physical laws. Existing methods integrate differentiable rendering with simulation but rely on predefined material priors, limiting their ability to handle unknown ones. We introduce MASIV, the first vision-based framewo…

2024

Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation

ECCV 2024oral

"Recent works have demonstrated that using reinforcement learning (RL) with multiple quality rewards can improve the quality of generated images in text-to-image (T2I) generation. However, manually adjusting reward weights poses challenges and may cause over-optimization in certain metrics. To solve…

Cited by 22SourcePDFScholar
2023

SVDiff: Compact Parameter Space for Diffusion Fine-Tuning

ICCV 2023poster

Recently, diffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts and various conditions. However, existing methods for customizing these models are limited by handling multiple personalized subjects and the risk o…

Cited by 271PDFScholar
2022

MAXIM: Multi-Axis MLP for Image Processing

CVPR 2022oral

Recent progress on Transformers and multi-layer perceptron (MLP) models provide new network architectural designs for computer vision tasks. Although these models proved to be effective in many vision tasks such as image recognition, there remain challenges in adapting them for low-level vision. The…

Cited by 624PDFcodeScholar
2022

MaxViT: Multi-axis Vision Transformer

ECCV 2022poster

"Transformers have recently gained significant attention in the computer vision community. However, the lack of scalability of self-attention mechanisms with respect to image size has limited their wide adoption in state-of-the-art vision backbones. In this paper we introduce an efficient and scalab…

2021

COMISR: Compression-Informed Video Super-Resolution

ICCV 2021poster

Most video super-resolution methods focus on restoring high-resolution video frames from low-resolution videos without taking into account compression. However, most videos on the web or mobile devices are compressed, and the compression can be severe when the bandwidth is limited. In this paper, we…

Cited by 48PDFcodeScholar
2015

Folding deformable objects using predictive simulation and trajectory optimization

IROS 2015poster

Robotic manipulation of deformable objects remains a challenging task. One such task is folding a garment autonomously. Given start and end folding positions, what is an optimal trajectory to move the robotic arm to fold a garment? Certain trajectories will cause the garment to move, creating wrinkl…

Cited by 158SourceScholar
2015

Regrasping and unfolding of garments using predictive thin shell modeling

ICRA 2015poster

Deformable objects such as garments are highly unstructured, making them difficult to recognize and manipulate. In this paper, we propose a novel method to teach a two-arm robot to efficiently track the states of a garment from an unknown state to a known state by iterative regrasping. The problem i…

Cited by 95SourceScholar