← Search

Yong Guo

42 accepted papers

2026

Condition Number Based Low-Bit Quantization for Image Super-Resolution

ICML 2026poster

Low-bit model quantization for image super-resolution (SR) is a longstanding task that is renowned for its surprising compression and acceleration ability. However, accuracy degradation is inevitable when compressing the full-precision (FP) model to ultra-low bit widths ($2\sim4$ bits). Experimental…

Cited by 0SourceScholar
2026

FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training

AAAI 2026technical

Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, their massive parameter sizes lead to slow inference, high memory usage, and poor deployability. Existing acceleration meth

Cited by 0SourcePDFScholar
2026

Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment

ICLR 2026poster

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on general contextual descriptions, sometimes limitin…

Cited by 0SourcecodeScholar
2026

LSGQuant: Layer-Sensitivity Guided Quantization for One-Step Diffusion Real-World Video Super-Resolution

ICML 2026poster

One-Step Diffusion Models have demonstrated promising capability and fast inference in real-world Video Super-Resolution (VSR). However, the substantial model size and high computational cost of Diffusion Transformers (DiTs) hinder their practical deployment. While low-bit quantization is a common a…

Cited by 0SourceScholar
2026

Learning from Disagreement: A Group Decision Simulation Framework for Robust Medical Image Segmentation

ICASSP 2026poster

Medical image segmentation annotation suffers from inter-rater variability (IRV) due to differences in annotators' expertise and the inherent blurriness of medical images. Standard approaches that simply average expert labels are flawed, as they discard the valuable clinical uncertainty revealed in…

Cited by 0SourcePDFScholar
2026

Q-DiT4SR: Exploration of Detail-Preserving Diffusion Transformer Quantization for Real-World Image Super-Resolution

ICML 2026poster

Recently, Diffusion Transformers (DiTs) have emerged in Real-World Image Super-Resolution (Real-ISR) to generate high-quality textures, yet their heavy inference burden hinders real-world deployment. While Post-Training Quantization (PTQ) is a promising solution for acceleration, existing methods in…

Cited by 0SourceScholar
2026

Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models

ICLR 2026poster

Diffusion large language models (dLLMs), which offer bidirectional context and flexible masked-denoising generation, are emerging as a compelling alternative to autoregressive (AR) LLMs. However, like AR LLMs, their model sizes continue to grow, motivating weight compression for deployment. Although…

Cited by 0SourceScholar
2026

QuantVSR: Low-Bit Post-Training Quantization for Real-World Video Super-Resolution

AAAI 2026technical

Diffusion models have shown superior performance in real-world video super-resolution (VSR). However, the slow processing speeds and heavy resource consumption of diffusion models hinder their practical application and deployment. Quantization offers a potential solution for compressing the VSR mode

Cited by 0SourcePDFScholar
2026

SODiff:Semantic-Oriented Diffusion Model for JPEG Compression Artifacts Removal

AAAI 2026technical

JPEG, as a widely used image compression standard, often introduces severe visual artifacts when achieving high compression ratios. Although existing deep learning-based restoration methods have made considerable progress, they often struggle to recover complex texture details, resulting in over-smo

Cited by 0SourcePDFScholar
2025

BiMaCoSR: Binary One-Step Diffusion Model Leveraging Flexible Matrix Compression for Real Super-Resolution

ICML 2025poster

While super-resolution (SR) methods based on diffusion models (DM) have demonstrated inspiring performance, their deployment is impeded due to the heavy request of memory and computation. Recent researchers apply two kinds of methods to compress or fasten the DM. One is to compress the DM into 1-bit…

2025

Compression-Aware One-Step Diffusion Model for JPEG Artifact Removal

ICCV 2025poster

Diffusion models have demonstrated remarkable success in image restoration tasks. However, their multi-step denoising process introduces significant computational overhead, limiting their practical deployment. Furthermore, existing methods struggle to effectively remove severe JPEG artifact, especia…

2025

DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution

NeurIPS 2025poster

Diffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one…

Cited by 0SourcecodeScholar
2025

Effective Diffusion Transformer Architecture for Image Super-Resolution

AAAI 2025technical

Recent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image…

2025

Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic Teacher

ICLR 2025poster

Knowledge distillation aims to transfer knowledge from a large teacher model to a compact student counterpart, often coming with a significant performance gap between them. Interestingly, we find that a too-large performance gap can hamper the training process. To alleviate this, we propose a **Gap…

Cited by 0SourcePDFScholar
2025

MambaIRv2: Attentive State Space Restoration

CVPR 2025poster

The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts t…

2025

OSCAR: One-Step Diffusion Codec Across Multiple Bit-rates

NeurIPS 2025poster

Pretrained latent diffusion models have shown strong potential for lossy image compression, owing to their powerful generative priors. Most existing diffusion-based methods reconstruct images by iteratively denoising from random noise, guided by compressed latent representations. While these approac…

Cited by 0SourcecodeScholar
2025

One Diffusion Step to Real-World Super-Resolution via Flow Trajectory Distillation

ICML 2025poster

Diffusion models (DMs) have significantly advanced the development of real-world image super-resolution (Real-ISR), but the computational cost of multi-step diffusion models limits their application. One-step diffusion models generate high-quality images in a one sampling step, greatly reducing comp…

2025

PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement

NeurIPS 2025poster

Multi-frame video enhancement tasks aim to improve the spatial and temporal resolution and quality of video sequences by leveraging temporal information from multiple frames, which are widely used in streaming video processing, surveillance, and generation. Although numerous Transformer-based enhanc…

Cited by 0SourcecodeScholar
2025

PassionSR: Post-Training Quantization with Adaptive Scale in One-Step Diffusion based Image Super-Resolution

CVPR 2025poster

Diffusion-based image super-resolution (SR) models have shown superior performance at the cost of multiple denoising steps. However, even though the denoising step has been reduced to one, they require high computational costs and storage requirements, making it difficult for deployment on hardware…

2025

PocketSR: The Super-Resolution Expert in Your Pocket Mobiles

NeurIPS 2025poster

Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for ed…

Cited by 0SourceScholar
2025

Restabilizing Diffusion Models with Predictive Noise Fusion Strategy for Image Super-Resolution

AAAI 2025technical

Diffusion models are prominent in image generation for producing detailed and realistic images from Gaussian noises. However, they often encounter instability issues in image restoration tasks, e.g., super-resolution. Existing methods typically rely on multiple runs to find an initial noise that pro…

2025

Segment Any-Quality Images with Generative Latent Space Enhancement

CVPR 2025poster

Despite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting their effectiveness in real-world scenarios. To address this, we propose GleSAM, which utilizes Generative Latent space Enhancement to boost robustness on…

Cited by 0SourcePDFScholar
2025

Temporal Action Detection Model Compression by Progressive Block Drop

CVPR 2025poster

Temporal action detection (TAD) aims to identify and localize action instances in untrimmed videos, which is essential for various video understanding tasks. However, recent improvements in model performance, driven by larger feature extractors and datasets, have led to increased computational deman…

Cited by 0SourcePDFScholar
2025

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis

ICCV 2025poster

Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals.While diffusion transformers (DiTs) have demonstrated remarkable capabilities in high-quality video generation, scaling them to 2K resolution remains computationally prohibitive due to quadratic gro…

Cited by 0SourcePDFScholar
2025

Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator

NeurIPS 2025poster

Diffusion models have demonstrated excellent performance for real-world image super-resolution (Real-ISR), albeit at high computational costs. Most existing methods are trying to derive one-step diffusion models from multi-step counterparts through knowledge distillation (KD) or variational score di…

Cited by 0SourceScholar
2024

2DQuant: Low-bit Post-Training Quantization for Image Super-Resolution

NeurIPS 2024poster

Low-bit quantization has become widespread for compressing image super-resolution (SR) models for edge deployment, which allows advanced SR models to enjoy compact low-bit parameters and efficient integer/bitwise constructions for storage compression and inference acceleration, respectively. However…

2024

Benchmarking the Robustness of Temporal Action Detection Models Against Temporal Corruptions

CVPR 2024poster

Temporal action detection (TAD) aims to locate action positions and recognize action categories in long-term untrimmed videos. Although many methods have achieved promising results their robustness has not been thoroughly studied. In practice we observe that temporal information in videos can be occ…

2024

Binarized Diffusion Model for Image Super-Resolution

NeurIPS 2024poster

Advanced diffusion models (DMs) perform impressively in image super-resolution (SR), but the high memory and computational costs hinder their deployment. Binarization, an ultra-compression algorithm, offers the potential for effectively accelerating DMs. Nonetheless, due to the model structure and t…

2024

DriveGPT4: Interpretable End-to-End Autonomous Driving Via Large Language Model

RA-L 2024

Multimodallarge language models (MLLMs) have emerged as a prominent area of interest within the research community, given their proficiency in handling and reasoning with non-textual data, including images and videos. This study seeks to extend the application of MLLMs to the realm of autonomous dri

Cited by 603SourceScholar
2024

UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New Peaks

NeurIPS 2024poster

Ultra-high-resolution image generation poses great challenges, such as increased semantic planning complexity and detail synthesis difficulties, alongside substantial training resource demands. We present UltraPixel, a novel architecture utilizing cascade diffusion models to generate high-quality im…

Cited by 17SourcePDFScholar
2023

Downscaled Representation Matters: Improving Image Rescaling with Collaborative Downscaled Images

ICCV 2023poster

Deep networks have achieved great success in image rescaling (IR) task that seeks to learn the optimal downscaled representations, i.e., low-resolution (LR) images, to reconstruct the original high-resolution (HR) images. Compared with super-resolution methods that consider a fixed downscaling schem…

Cited by 13PDFcodeScholar
2023

Improving Robustness of Vision Transformers by Reducing Sensitivity To Patch Corruptions

CVPR 2023poster

Despite their success, vision transformers still remain vulnerable to image corruptions, such as noise or blur. Indeed, we find that the vulnerability mainly stems from the unstable self-attention mechanism, which is inherently built upon patch-based inputs and often becomes overly sensitive to the…

2021

AdaXpert: Adapting Neural Architecture for Growing Data

ICML 2021spotlight

In real-world applications, data often come in a growing manner, where the data volume and the number of classes may increase dynamically. This will bring a critical challenge for learning: given the increasing data volume or the number of classes, one has to instantaneously adjust the neural model…

2021

Contrastive Neural Architecture Search With Neural Architecture Comparators

CVPR 2021poster

One of the key steps in Neural Architecture Search (NAS) is to estimate the performance of candidate architectures. Existing methods either directly use the validation performance or learn a predictor to estimate the performance. However, these methods can be either computationally expensive or very…

Cited by 87PDFcodeScholar
2020

Breaking the Curse of Space Explosion: Towards Efficient NAS with Curriculum Search

ICML 2020poster

Neural architecture search (NAS) has become an important approach to automatically find effective architectures. To cover all possible good architectures, we need to search in an extremely large search space with billions of candidate architectures. More critically, given a large search space, we ma…

2020

Closed-Loop Matters: Dual Regression Networks for Single Image Super-Resolution

CVPR 2020poster

Deep neural networks have exhibited promising performance in image super-resolution (SR) by learning a nonlinear mapping function from low-resolution (LR) images to high-resolution (HR) images. However, there are two underlying limitations to existing SR methods. First, learning the mapping function…

Cited by 425PDFcodeScholar
2019

NAT: Neural Architecture Transformer for Accurate and Compact Architectures

NeurIPS 2019poster

Designing effective architectures is one of the key factors behind the success of deep neural networks. Existing deep architectures are either manually designed or automatically searched by some Neural Architecture Search (NAS) methods. However, even a well-searched architecture may still contain ma…

2018

Adversarial Learning with Local Coordinate Coding

ICML 2018oral

Generative adversarial networks (GANs) aim to generate realistic data from some prior distribution (e.g., Gaussian noises). However, such prior distribution is often independent of real data and thus may lose semantic information (e.g., geometric structure or content in images) of data. In practice,…

Cited by 44SourcePDFScholar
2018

Discrimination-aware Channel Pruning for Deep Neural Networks

NeurIPS 2018poster

Channel pruning is one of the predominant approaches for deep model compression. Existing pruning methods either train from scratch with sparsity constraints on channels, or minimize the reconstruction error between the pre-trained feature maps and the compressed ones. Both strategies suffer from s…