← Search

Ruibin Li

10 accepted papers

2026

CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning

ICML 2026poster

Image editing has achieved impressive results with the development of large-scale generative models. However, existing models mainly focus on the editing effects of intended objects and regions, often leading to unwanted changes in unintended regions. We present a post-training framework for \textbf…

Cited by 0SourceScholar
2026

Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis

ICML 2026poster

Distribution matching distillation (DMD) aligns a multi-step generator with its few-step counterpart to enable high-quality generation under low inference cost. However, DMD tends to suffer from mode collapse, as its reverse-KL formulation inherently encourages mode-seeking behavior, for which exist…

Cited by 0SourceScholar
2026

Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks

ICLR 2026poster

Diffusion models have shown impressive performance in many visual generation and manipulation tasks. Many existing methods focus on training a model for a specific task, especially, text-to-video (T2V) generation, while many other works focus on finetuning the pretrained T2V model for image-to-video…

Cited by 0SourcecodeScholar
2025

DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement

AAAI 2025technical

Recent years have witnessed the rise of Neural-enhanced Video Streaming (NeVS), which integrates neural restoration models into video codecs for higher compression-restoration performance. Despite its benefit, existing work has not well explored the full potential of NeVS paradigm, due to: (1) post-…

2025

InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction

ICCV 2025poster

Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video, instruction) is a challenging task. Existing datasets mostly co…

2025

Mjölnir: Breaking the Shield of Perturbation-Protected Gradients via Adaptive Diffusion

AAAI 2025technical

Perturbation-based mechanisms, such as differential privacy, mitigate gradient leakage attacks by introducing noise into the gradients, thereby preventing attackers from reconstructing clients' private data from the leaked gradients. However, can gradient perturbation protection mechanisms truly def…

Cited by 0SourcePDFScholar
2025

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

IJCAI 2025

Vision-Language Models (VLMs) are essential for multimodal tasks, especially compositional reasoning (CR) tasks, which require distinguishing fine-grained semantic differences between visual and textual embeddings. However, existing methods primarily fine-tune the model by generating text-based hard

2024

On the Robustness of Neural-Enhanced Video Streaming against Adversarial Attacks

AAAI 2024technical

The explosive growth of video traffic on today's Internet promotes the rise of Neural-enhanced Video Streaming (NeVS), which effectively improves the rate-distortion trade-off by employing a cheap neural super-resolution model for quality enhancement on the receiver side. Missing by existing work, w…

Cited by 10SourcePDFScholar
2024

ParsNets: A Parsimonious Composition of Orthogonal and Low-Rank Linear Networks for Zero-Shot Learning

IJCAI 2024poster

This paper provides a novel parsimonious yet efficient design for zero-shot learning (ZSL), dubbed ParsNets, in which we are interested in learning a composition of on-device friendly linear networks, each with orthogonality and low-rankness properties, to achieve equivalent or better performance ag…

Cited by 10SourcePDFScholar