← Search

Shuai Yang

66 accepted papers

2026

FlowPortal: Residual-Corrected Flow for Training-Free Video Relighting and Background Replacement

CVPR 2026

Video relighting with background replacement is a challenging task critical for applications in film production and creative media. Existing methods struggle to balance temporal consistency, spatial fidelity, and illumination naturalness. To address these issues, we introduce FlowPortal, a novel tra

Cited by 0SourcecodeScholar
2026

LongLive: Real-time Interactive Long Video Generation

ICLR 2026poster

We present LongLive, a frame-level autoregressive (AR) framework for real-time and interactive long video generation. Long video generation presents challenges in both efficiency and quality. Diffusion and Diffusion-Forcing models can produce high-quality videos but suffer from low efficiency due to…

Cited by 188SourcecodeScholar
2026

On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional Study

ICLR 2026poster

Counterfactual reasoning has emerged as a crucial technique for generalizing the reasoning capabilities of large language models (LLMs). By generating and analyzing counterfactual scenarios, researchers can assess the adaptability and reliability of model decision-making. Although prior work has sho…

Cited by 0SourceScholar
2026

QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMs

ICLR 2026poster

We propose QeRL, a Quantization-enhanced Reinforcement Learning framework for large language models (LLMs). While RL is essential for LLMs' reasoning capabilities, it is resource-intensive, requiring substantial GPU memory and long rollout duration. QeRL addresses these issues by combining NVFP4 qua…

Cited by 0SourcecodeScholar
2026

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

ICLR 2026poster

Large language and vision-language models have inspired end-to-end vision-language-action (VLA) systems in robotics, yet existing robot datasets remain costly, embodiment-specific, and insufficient, limiting robustness and generalization. Recent approaches address this by adopting a plan-then-execut…

Cited by 0SourcecodeScholar
2026

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

ICLR 2026oral

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720×1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Tw…

Cited by 0SourcecodeScholar
2026

STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

ICLR 2026poster

We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global optimization or rely on simplistic memory mechanisms that scale po…

Cited by 0SourcecodeScholar
2026

Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling

AAAI 2026technical

Recent vision-language-action (VLA) models built on pretrained vision-language models (VLMs) have demonstrated strong performance in robotic manipulation. However, these models remain constrained by the single-frame image paradigm and fail to fully leverage the temporal information offered by multi-

Cited by 0SourcePDFScholar
2026

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers

CVPR 2026

Diffusion Transformers (DiTs) set the state of the art in visual generation, yet their quadratic self-attention cost fundamentally limits scaling to long token sequences. Recent Top-K sparse attention approaches reduce the computation of DiTs by compressing tokens into block-wise representation and

Cited by 4SourceScholar
2026

Vision-Language-Action Instruction Tuning: From Understanding to Manipulation

ICLR 2026poster

To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often sacrifice one for the other, narrow their abilities to task-specific manipulation data, and suffer catastrophic forgettin…

Cited by 0SourcecodeScholar
2025

Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space

CVPR 2025poster

Latent Diffusion Models (LDMs) are known to have an unstable generation process, where even small perturbations or shifts in the input noise can lead to significantly different outputs. This hinders their applicability in applications requiring consistent results. In this work, we redesign LDMs to e…

2025

Balanced Image Stylization with Style Matching Score

ICCV 2025poster

We present Style Matching Score (SMS), a novel optimization method for image stylization with diffusion models. Balancing effective style transfer with content preservation is a long-standing challenge. Unlike existing efforts, our method reframes image stylization as a style distribution matching p…

Cited by 0SourcePDFScholar
2025

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation

CVPR 2025poster

Robotic manipulation in real-world settings remains challenging, especially regarding robust generalization. Existing simulation platforms lack sufficient support for exploring how policies adapt to varied instructions and scenarios. Thus, they lag behind the growing interest in instruction-followin…

Cited by 0SourcePDFScholar
2025

GaussianAnything: Interactive Point Cloud Flow Matching for 3D Generation

ICLR 2025poster

Recent advancements in diffusion models and large-scale datasets have revolutionized image and video generation, with increasing focus on 3D content generation. While existing methods show promise, they face challenges in input formats, latent space structures, and output representations. This paper…

Cited by 0SourcePDFScholar
2025

Imagine360: Immersive 360 Video Generation from Perspective Anchor

NeurIPS 2025poster

$360^\circ$ videos offer a hyper-immersive experience that allows the viewers to explore a dynamic scene from full 360 degrees. To achieve more accessible and personalized content creation in $360^\circ$ video format, we seek to lift standard perspective videos into $360^\circ$ equirectangular vide…

Cited by 0SourceScholar
2025

MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention

CVPR 2025poster

Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we…

2025

Omni-Chart-600K: A Comprehensive Dataset of Chart Types for Chart Understanding

NAACL 2025findings

To address the deficiencies in chart types and the limited scope of chart tasks in existing datasets, we conducted a comprehensive review of current data collection methodologies. By integrating manual annotation with data generation leveraging GPT-4, we developed a dataset that includes 21 diverse…

Cited by 0SourcePDFScholar
2025

OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing Images

ICCV 2025poster

Remote sensing object detection has made significant progress, but most studies still focus on closed-set detection, limiting generalization across diverse datasets. Open-vocabulary object detection (OVD) provides a solution by leveraging multimodal associations between text prompts and visual featu…

2025

PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion Model

CVPR 2025poster

Optical illusion hidden picture is an interesting visual perceptual phenomenon where an image is cleverly integrated into another picture. Established on the off-the-shelf text-to-image (T2I) diffusion model, we propose a novel text-guided image-to-image (I2I) translation framework dubbed as Phase-T…

2025

REAR: Reinforced Reasoning Optimization for Event Argument Extraction with Relation-Aware Support

EMNLP 2025

Event argument extraction aims to identify event arguments and classify their roles within events, whereas relation extraction classifies semantic relationships between entities. Existing methods typically design task-specific models for EAE, which restricts the integration of relation-level semanti

Cited by 0SourcePDFScholar
2025

SGAR: Structural Generative Augmentation for 3D Human Motion Retrieval

NeurIPS 2025poster

3D human motion-text retrieval is essential for accurate motion understanding, targeted at cross-modal alignment learning. Existing methods typically align the global motion-text concepts directly, suffering from sub-optimal generalization due to the uncertainty of correspondence learning between mu…

Cited by 0SourceScholar
2025

Split-and-Combine: Enhancing Style Augmentation for Single Domain Generalization

ICCV 2025poster

Single domain generalization aims to learn a model with good generalization ability from a single source domain. Recent advances in this field have focused on increasing the diversity of the training data through style (e.g., color and texture) augmentation. However, most existing methods apply unif…

Cited by 0SourcePDFScholar
2025

State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator

IJCAI 2025

In reinforcement learning (RL) based robot skill acquisition, a high-fidelity simulator is usually indispensable but unattainable since the real environment dynamics are difficult to model, which leads to severe sim-to-real gaps. Existing methods solve this problem by combining offline and online RL

Cited by 0SourcePDFScholar
2025

TokensGen: Harnessing Condensed Tokens for Long Video Generation

ICCV 2025poster

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In this paper, we propose TokensGen, a novel two-stage framewo…

Cited by 0SourcePDFScholar
2025

Trajectory attention for fine-grained video motion control

ICLR 2025poster

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixe…

Cited by 0SourcePDFScholar
2025

WorldMem: Long-term Consistent World Simulation with Memory

NeurIPS 2025poster

World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term consistency, particularly in preserving 3D spatial consistency. In…

Cited by 0SourceScholar
2024

AutoCali: Enhancing AoA-based Indoor Localization through Automatic Phase Calibration

ICASSP 2024accepted

Recent advancements in WiFi indoor localization have demonstrated the potential for achieving decimeter-level accuracy based on Angle of Arrival (AoA). However, existing commercial WiFi Access Points (APs) suffer from phase offset across different antennas, which significantly degrade the performanc…

Cited by 0SourceScholar
2024

Automotive Radar Interference Mitigation Via SINR Maximization

ICASSP 2024accepted

The mutual interference mitigation between identical or similar radar systems in autonomous driving has gained wide spread attention from both academia and industry. The resulted ghost target interference will reduce the sensitivity of the radar sensor and increase the false alarm rate. To tackle th…

Cited by 0SourceScholar
2024

Defect Spectrum: A Granular Look of Large-scale Defect Datasets with Rich Semantics

ECCV 2024poster

"Defect inspection is paramount within the closed-loop manufacturing system. However, existing datasets for defect inspection often lack the precision and semantic granularity required for practical applications. In this paper, we introduce the Defect Spectrum, a comprehensive benchmark that offers…

2024

Demonstration Retrieval-Augmented Generative Event Argument Extraction

COLING 2024main

We tackle Event Argument Extraction (EAE) in the manner of template-based generation. Based on our exploration of generative EAE, it suffers from several issues, such as multiple arguments of one role, generating words out of context and inconsistency with prescribed format. We attribute it to the w…

Cited by 2SourcePDFScholar
2024

FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation

CVPR 2024poster

The remarkable efficacy of text-to-image diffusion models has motivated extensive exploration of their potential application in video domains. Zero-shot methods seek to extend image diffusion models to videos without necessitating model training. Recent methods mainly focus on incorporating inter-fr…

2024

Forward Learning of Graph Neural Networks

ICLR 2024poster

Graph neural networks (GNNs) have achieved remarkable success across a wide range of applications, such as recommendation, drug discovery, and question answering. Behind the success of GNNs lies the backpropagation (BP) algorithm, which is the de facto standard for training deep neural networks (NNs…

2024

GroupDiff: Diffusion-based Group Portrait Editing

ECCV 2024poster

"Group portrait editing is highly desirable since users constantly want to add a person, delete a person, or manipulate existing persons. It is also challenging due to the intricate dynamics of human interactions and the diverse gestures. In this work, we present GroupDiff, a pioneering effort to ta…

2024

KnowVrDU: A Unified Knowledge-aware Prompt-Tuning Framework for Visually-rich Document Understanding

COLING 2024main

In Visually-rich Document Understanding (VrDU), recent advances of incorporating layout and image features into the pre-training language models have achieved significant progress. Existing methods usually developed complicated dedicated architectures based on pre-trained models and fine-tuned them…

2024

LN3Diff: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation

ECCV 2024poster

"The field of neural rendering has witnessed significant progress with advancements in generative models and differentiable rendering techniques. Though 2D diffusion has achieved success, a unified 3D diffusion pipeline remains unsettled. This paper introduces a novel framework called to address thi…

2024

Low-Rank Approximation for Sparse Attention in Multi-Modal LLMs

CVPR 2024poster

This paper focuses on the high computational complexity in Large Language Models (LLMs) a significant challenge in both natural language processing (NLP) and multi-modal tasks. We propose Low-Rank Approximation for Sparse At- tention (LoRA-Sparse) an innovative approach that strate- gically reduces…

Cited by 11SourcePDFScholar
2024

MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations

NeurIPS 2024poster

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets, previous works mainly focus on understanding object properties…

2024

Unified Generative and Discriminative Training for Multi-modal Large Language Models

NeurIPS 2024poster

In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle various complex tasks, yet issues such as hallucinations and weak object discrimination persist. Discriminative trainin…

Cited by 3SourcePDFScholar
2024

Video Diffusion Models are Training-free Motion Interpreter and Controller

NeurIPS 2024poster

Video generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with training-based paradigms, which, however, demands substantial train…

Cited by 15SourcePDFScholar
2024

VideoBooth: Diffusion-based Video Generation with Image Prompts

CVPR 2024poster

Text-driven video generation witnesses rapid progress. However merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents especially for customized content creation. In this paper we study the task of video generation with image prompt…

Cited by 67SourcePDFScholar
2024

Word-level Commonsense Knowledge Selection for Event Detection

COLING 2024main

Event Detection (ED) is a task of automatically extracting multi-class trigger words. The understanding of word sense is crucial for ED. In this paper, we utilize context-specific commonsense knowledge to strengthen word sense modeling. Specifically, we leverage a Context-specific Knowledge Selector…

2023

A Multi-Configuration Track-Legged Humanoid Robot for Dexterous Manipulation and High Mobility: Design and Development

RA-L 2023

Various applications (e.g., disaster response) put forward higher requirements for the capabilities of robots. However, to date, few platforms can be compatible with the required mobility, stability and dexterity. A major challenge is mobile stability may severely limit the variety of manipulation s

Cited by 7SourceScholar
2023

DeformToon3D: Deformable Neural Radiance Fields for 3D Toonification

ICCV 2023poster

In this paper, we address the challenging problem of 3D toonification, which involves transferring the style of an artistic domain onto a target 3D face with stylized geometry and texture. Although fine-tuning a pre-trained 3D GAN on the artistic domain can produce reasonable performance, this strat…

Cited by 14PDFScholar
2023

Not All Steps are Created Equal: Selective Diffusion Distillation for Image Manipulation

ICCV 2023poster

Conditional diffusion models have demonstrated impressive performance in image manipulation tasks. The general pipeline involves adding noise to the image and then denoising it. However, this method faces a trade-off problem: adding too much noise affects the fidelity of the image while adding too l…

Cited by 12PDFcodeScholar
2023

Scenimefy: Learning to Craft Anime Scene via Semi-Supervised Image-to-Image Translation

ICCV 2023poster

Automatic high-quality rendering of anime scenes from complex real-world images is of significant practical value. The challenges of this task lie in the complexity of the scenes, the unique features of anime style, and the lack of high-quality datasets to bridge the domain gap. Despite promising at…

Cited by 21PDFcodeScholar
2023

Self-Supervised Geometry-Aware Encoder for Style-Based 3D GAN Inversion

CVPR 2023poster

StyleGAN has achieved great progress in 2D face reconstruction and semantic editing via image inversion and latent editing. While studies over extending 2D StyleGAN to 3D faces have emerged, a corresponding generic 3D GAN inversion framework is still missing, limiting the applications of 3D face rec…

Cited by 39SourcePDFScholar
2023

StyleGANEX: StyleGAN-Based Manipulation Beyond Cropped Aligned Faces

ICCV 2023poster

Recent advances in face manipulation using StyleGAN have produced impressive results. However, StyleGAN is inherently limited to cropped aligned faces at a fixed image resolution it is pre-trained on. In this paper, we propose a simple and effective solution to this limitation by using dilated convo…

Cited by 26PDFcodeScholar
2023

Text2Performer: Text-Driven Human Video Generation

ICCV 2023poster

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the appearance and motions of a target performer. Compared to general te…

Cited by 57PDFcodeScholar
2022

Pastiche Master: Exemplar-Based High-Resolution Portrait Style Transfer

CVPR 2022poster

Recent studies on StyleGAN show high performance on artistic portrait generation by transfer learning with limited data. In this paper, we explore more challenging exemplar-based high-resolution portrait style transfer by introducing a novel DualStyleGAN with flexible control of dual styles of the o…

Cited by 145PDFcodeScholar
2021

Instance-Aware Coherent Video Style Transfer for Chinese Ink Wash Painting

IJCAI 2021poster

Recent researches have made remarkable achievements in fast video style transfer based on western paintings. However, due to the inherent different drawing techniques and aesthetic expressions of Chinese ink wash painting, existing methods either achieve poor temporal consistency or fail to transfer…

2020

Deep Plastic Surgery: Robust and Controllable Image Editing with Human-Drawn Sketches

ECCV 2020poster

Sketch-based image editing aims to synthesize and modify photos based on the structural information provided by the human-drawn sketches. Since sketches are difficult to collect, previous methods mainly use edge maps instead of sketches to train models (referred to as edge-based models). However, hu…

2019

Controllable Artistic Text Style Transfer via Shape-Matching GAN

ICCV 2019oral

Artistic text style transfer is the task of migrating the style from a source image to the target text to create artistic typography. Recent style transfer methods have considered texture control to enhance usability. However, controlling the stylistic degree in terms of shape deformation remains an…

Cited by 129PDFcodeScholar
2018

Erase or Fill? Deep Joint Recurrent Rain Removal and Reconstruction in Videos

CVPR 2018poster

In this paper, we address the problem of video rain removal by constructing deep recurrent convolutional networks. We visit the rain removal case by considering rain occlusion regions, i.e. light transmittance of rain streaks is low. Different from additive rain streaks, in such rain occlusion regio…

Cited by 224SourcePDFScholar
2018

Soft Decoding of Light Field Images Using Pocs and Fast Graph Spectrayl Filters

ICASSP 2018accepted

Light field data captured by a lenslet-based image sensor is typically demosaicked, aligned and rearranged into a series of sub-aperture (viewpoint) images, before a disparity-compensated coding scheme is employed for compression. In this paper, we focus on the problem of soft decoding of block-base…

Cited by 0SourceScholar
2017

Geometric Map-Assisted Localization for Mobile Robots Based on Uniform-Gaussian Distribution

RA-L 2017

Drift and scale ambiguity are two main issues which reduce localization accuracy in monocular visual odometry (MVO). It is necessary to propose a unified model to represent these measurement uncertainties. In this paper, we present a geometric map-assisted localization approach for mobile robots equ

Cited by 18SourceScholar
2016

Structure-guided image completion via regularity statistics

ICASSP 2016accepted

In this paper, we propose a novel hierarchical image completion approach using regularity statistics, considering structure features. Guided by dominant structures, the target image is used to generate reference images in a self-reproductive way by image data enhancement. The structure-guided image…

Cited by 0SourceScholar
2015

Novel autoregressive model based on adaptive window-extension and patch-geodesic distance for image interpolation

ICASSP 2015accepted

In this paper, we propose a novel autoregressive (AR) model based on the adaptive window and the patch-geodesic distance for the image interpolation. The model combines the information of inner/inter-patch correlation. To model the inner-patch correlation, we introduce a patch-geodesic distance simi…

Cited by 0SourceScholar