← Search

Mingrui Zhu

11 accepted papers

2026

Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution

AAAI 2026technical

The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image super-resolution (Real-ISR), existing approaches mainly rely on fine-t

Cited by 0SourcePDFScholar
2026

PSMix: Robust Point Cloud Recognition through Spectral Domain Mixing

ICML 2026poster

While data augmentation is essential for robust point cloud recognition, conventional spatial mixup strategies often compromise geometric integrity by generating physically unrealistic samples. To overcome this limitation, we propose PSMix, which shifts the mixing paradigm to the spectral domain via…

Cited by 0SourceScholar
2026

Robust Cross-Modal Retrieval via Generative Semantic Refinement and Exclusion-Guided Adaptation

ICML 2026poster

Vision-Language Pre-trained (VLP) models are vulnerable to real-world query noise. Current cross-modal Test-Time Adaptation (TTA) methods often rely on high-confidence predictions, which induces confirmation bias and neglects the informative signals in ambiguous Low-Confidence Queries. To address th…

Cited by 0SourceScholar
2026

When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity Correspondence

CVPR 2026

Establishing accurate correspondence between sparse line representations and rich textured imagery remains a formidable challenge. While diffusion features excel in semantic correspondence, they struggle to bridge the fundamental gap between abstract sketches and texture-rich photographs. We identif

Cited by 0SourcecodeScholar
2025

3D Test-time Adaptation via Graph Spectral Driven Point Shift

ICCV 2025poster

While test-time adaptation (TTA) methods effectively address domain shifts by dynamically adapting pre-trained models to target domain data during online inference, their application to 3D point clouds is hindered by their irregular and unordered structure. Current 3D TTA methods often rely on compu…

Cited by 0SourcePDFScholar
2025

Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts

ICML 2025poster

Diffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformer…

Cited by 0SourcePDFScholar
2025

Effective Diffusion Transformer Architecture for Image Super-Resolution

AAAI 2025technical

Recent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image…

2024

Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors

IJCAI 2024poster

The remarkable prowess of diffusion models in image generation has spurred efforts to extend their application beyond generative tasks. However, a persistent challenge exists in lacking a unified approach to apply diffusion models to visual perception tasks with diverse semantic granularity requirem…

Cited by 3SourcePDFScholar
2024

On the Analysis of GAN-based Image-to-Image Translation with Gaussian Noise Injection

ICLR 2024poster

Image-to-image (I2I) translation is vital in computer vision tasks like style transfer and domain adaptation. While recent advances in GAN have enabled high-quality sample generation, real-world challenges such as noise and distortion remain significant obstacles. Although Gaussian noise injection d…

Cited by 2SourcePDFScholar
2021

A Sketch-Transformer Network for Face Photo-Sketch Synthesis

IJCAI 2021poster

We present a face photo-sketch synthesis model, which converts a face photo into an artistic face sketch or recover a photo-realistic facial image from a sketch portrait. Recent progress has been made by convolutional neural networks (CNNs) and generative adversarial networks (GANs), so that promisi…

Cited by 32SourcePDFScholar