← Search

Sucheng Ren

26 accepted papers

2026

Frequency-Aware Flow Matching for High-Quality Image Generation

CVPR 2026

Flow matching models have emerged as a powerful framework for realistic image generation by learning to reverse a corruption process that progressively adds Gaussian noise. However, because noise is injected in the latent domain, its impact on different frequency components is non-uniform. As a resu

Cited by 0SourcecodeScholar
2026

Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane

ICML 2026poster

Rotary Position Embedding (RoPE) is the de facto positional encoding in large language models due to its ability to encode relative positions and support length extrapolation. When adapted to vision transformers, the standard axial formulation decomposes two-dimensional spatial positions into horizo…

Cited by 0SourceScholar
2026

WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark

ICLR 2026poster

Recent advances in image editing models have demonstrated remarkable capabilities in executing explicit instructions, such as attribute manipulation, style transfer, and pose synthesis. However, these models often face challenges when dealing with implicit editing instructions, which describe the…

Cited by 0SourceScholar
2026

iGRPO: Fast Online RL for Flow Matching Model with Dense Reward

ICML 2026poster

Conventional practice assumes that online reinforcement learning for flow-matching models requires sampling full denoising trajectories to compute rewards. This assumption underlies methods such as Group Relative Policy Optimization (GRPO), where the policy must traverse the entire reverse process b…

Cited by 0SourceScholar
2025

Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency

CVPR 2025poster

In this work, we introduce the Adventurer series models where we treat images as sequences of patch tokens and employ uni-directional language models to learn visual representations. This modeling paradigm allows us to process images in a recurrent formulation with linear complexity relative to the…

Cited by 0SourcePDFScholar
2025

Autoregressive Pretraining with Mamba in Vision

ICLR 2025poster

The vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that Mamba's visual capability can be significantly enhanced through autoregressive pretraining, a direction not previously explored. Efficiency-…

2025

Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

ICCV 2025poster

Autoregressive (AR) modeling, known for its next-token prediction paradigm, underpins state-of-the-art language and visual generative models. Traditionally, a "token" is treated as the smallest prediction unit, often a discrete symbol in language or a quantized patch in vision. However, the optimal…

2025

FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching

ICML 2025poster

Autoregressive (AR) modeling has achieved remarkable success in natural language processing by enabling models to generate text with coherence and contextual understanding through next token prediction. Recently, in image generation, VAR proposes scale-wise autoregressive modeling, which extends the…

2025

Mamba-Reg: Vision Mamba Also Needs Registers

CVPR 2025poster

Similar to Vision Transformers, this paper identifies artifacts also present within the feature maps of Vision Mamba. These artifacts, corresponding to high-norm tokens emerging in low-information background areas of images, appear much more severe in Vision Mamba---they exist prevalently even with…

2025

What If We Recaption Billions of Web Images with LLaMA-3?

ICML 2025poster

Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investi…

Cited by 38SourcePDFScholar
2024

Rejuvenating image-GPT as Strong Visual Representation Learners

ICML 2024oral

This paper enhances image-GPT (iGPT), one of the pioneering works that introduce autoregressive pretraining to predict the next pixels for visual representation learning. Two simple yet essential changes are made. First, we shift the prediction target from raw pixels to semantic tokens, enabling a h…

2023

SG-Former: Self-guided Transformer with Evolving Token Reallocation

ICCV 2023poster

Vision Transformer has demonstrated impressive success across various vision tasks. However, its heavy computation cost, which grows quadratically with respect to the token sequence length, largely limits its power in handling large feature maps. To alleviate the computation cost, previous works rel…

Cited by 74PDFcodeScholar
2023

Self-supervision through Random Segments with Autoregressive Coding (RandSAC)

ICLR 2023poster

Inspired by the success of self-supervised autoregressive representation learning in natural language (GPT and its variants), and advances in recent visual architecture design with Vision Transformers (ViTs), in this paper, we explore the effects various design choices have on the success of applyin…

Cited by 15SourcePDFScholar
2023

The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation

ICLR 2023top-5%

Crossmodal knowledge distillation (KD) extends traditional knowledge distillation to the area of multimodal learning and demonstrates great success in various applications. To achieve knowledge transfer across modalities, a pretrained network from one modality is adopted as the teacher to provide su…

2023

TinyMIM: An Empirical Study of Distilling MIM Pre-Trained Models

CVPR 2023poster

Masked image modeling (MIM) performs strongly in pre-training large vision Transformers (ViTs). However, small models that are critical for real-world applications cannot or only marginally benefit from this pre-training approach. In this paper, we explore distillation techniques to transfer the suc…

2022

A Simple Data Mixing Prior for Improving Self-Supervised Learning

CVPR 2022poster

Data mixing (e.g., Mixup, Cutmix, ResizeMix) is an essential component for advancing recognition models. In this paper, we focus on studying its effectiveness in the self-supervised setting. By noticing the mixed images that share the same source images are intrinsically related to each other, we he…

Cited by 48PDFcodeScholar
2022

Co-Advise: Cross Inductive Bias Distillation

CVPR 2022poster

The inductive bias of vision transformers is more relaxed that cannot work well with insufficient data. Knowledge distillation is thus introduced to assist the training of transformers. Unlike previous works, where merely heavy convolution-based teachers are provided, in this paper, we delve into th…

Cited by 82PDFcodeScholar
2022

DynaST: Dynamic Sparse Transformer for Exemplar-Guided Image Generation

ECCV 2022poster

"One key challenge of exemplar-guided image generation lies in establishing fine-grained correspondences between input and guided images. Prior approaches, despite the promising results, have relied on either estimating dense attention to compute per-point matching, which is limited to only coarse s…

2022

Learning from Multiple Annotator Noisy Labels via Sample-Wise Label Fusion

ECCV 2022poster

"Data lies at the core of modern deep learning. The impressive performance of supervised learning is built upon a base of massive accurately labeled data. However, in some real-world applications, accurate labeling might not be viable; instead, multiple noisy labels (instead of one accurate label) a…

2022

Shunted Self-Attention via Multi-Scale Token Aggregation

CVPR 2022oral

Recent Vision Transformer (ViT) models have demonstrated encouraging results across various computer vision tasks, thanks to its competence in modeling long-range dependencies of image patches or tokens via self-attention. These models, however, usually designate the similar receptive fields of each…

Cited by 335PDFcodeScholar
2021

Delving Deep Into Many-to-Many Attention for Few-Shot Video Object Segmentation

CVPR 2021poster

This paper tackles the task of Few-Shot Video Object Segmentation (FSVOS), i.e., segmenting objects in the query videos with certain class specified in a few labeled support images. The key is to model the relationship between the query videos and the support images for propagating the object inform…

Cited by 25PDFcodeScholar
2021

Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip Reading

CVPR 2021poster

Lip reading aims to predict the spoken sentences from silent lip videos. Due to the fact that such a vision task usually performs worse than its counterpart speech recognition, one potential scheme is to distill knowledge from a teacher pretrained by audio signals. However, the latent domain gap bet…

Cited by 86PDFScholar
2021

On Feature Decorrelation in Self-Supervised Learning

ICCV 2021poster

In self-supervised representation learning, a common idea behind most of the state-of-the-art approaches is to enforce the robustness of the representations to predefined augmentations. A potential issue of this idea is the existence of completely collapsed solutions (i.e., constant features), which…

Cited by 236PDFcodeScholar
2021

Reciprocal Transformations for Unsupervised Video Object Segmentation

CVPR 2021poster

Unsupervised video object segmentation (UVOS) aims at segmenting the primary objects in videos without any human intervention. Due to the lack of prior knowledge about the primary objects, identifying them from videos is the major challenge of UVOS. Previous methods often regard the moving objects a…

Cited by 107PDFcodeScholar
2020

TENet: Triple Excitation Network for Video Salient Object Detection

ECCV 2020poster

In this paper, we propose a simple yet effective approach, named Triple Excitation Network, to reinforce the training of video salient object detection (VSOD) from three aspects, spatial, temporal, and online excitations. These excitation mechanisms are designed following the spirit of curriculum le…

Cited by 75SourcePDFScholar