← Search

Weiqi Li

14 accepted papers

2026

360Explorer: Exploring 4D Controllable World in Panoramic Videos

AAAI 2026technical

We present 360Explorer, a novel approach for generating 4D controllable panoramic videos conditioned on user-provided 3D instructions for exploring and manipulating dynamic worlds. Compared to existing perspective-based methods struggle to address spatial consistency during camera rotation in place,

Cited by 1SourcePDFScholar
2026

Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution

ICLR 2026poster

While many diffusion models have achieved impressive results in real-world video super-resolution (Real-VSR) by generating rich and realistic details, their reliance on multi-step sampling leads to slow inference. One-step networks like SeedVR2, DOVE, and DLoRAL alleviate this through condensing gen…

Cited by 0SourceScholar
2026

Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment

ICLR 2026oral

Reasoning-based image quality assessment (IQA) models trained through reinforcement learning (RL) exhibit exceptional generalization, yet the underlying mechanisms and critical factors driving this capability remain underexplored in current research. Moreover, despite their superior performance, the…

Cited by 0SourcecodeScholar
2026

UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement

CVPR 2026

Image quality assessment (IQA) and image restoration are fundamental problems in low-level vision. Although IQA and restoration are closely connected conceptually, most existing work treats them in isolation. Recent advances in unified multimodal understanding-generation models demonstrate promising

Cited by 0SourcecodeScholar
2026

VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling

CVPR 2026

Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose

Cited by 0SourceScholar
2026

VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning

AAAI 2026technical

Recent advances in AI-generated content (AIGC) have led to the emergence of powerful text-to-video generation models. Despite these successes, evaluating the quality of AIGC-generated videos remains challenging due to limited generalization, lack of temporal awareness, heavy reliance on large-scale

Cited by 0SourcePDFScholar
2025

AlignedGen: Aligning Style Across Generated Images

NeurIPS 2025poster

Diffusion-based generative models struggle to maintain high style consistency across generated images via text description. Although several style-aligned image generation methods have been proposed to address this issue, they exhibit suboptimal performance and are primarily built upon the U-Net arc…

Cited by 0SourcecodeScholar
2025

Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

NeurIPS 2025spotlight

Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language models (MLLMs) has significantly broadened the scope of IQA, mo…

Cited by 0SourcecodeScholar
2024

360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model

CVPR 2024poster

Panorama video recently attracts more interest in both study and application courtesy of its immersive experience. Due to the expensive cost of capturing 360-degree panoramic videos generating desirable panorama videos by prompts is urgently required. Lately the emerging text-to-video (T2V) diffusio…

2024

Automatic Field of View Adjustment of an RCM Constraint-Free Continuum Laparoscopic Robot

IROS 2024poster

Automatic laparoscopic field of view (FOV) adjustment can effectively assist surgeons in minimally invasive surgery (MIS). However, existing work based on rod-shaped laparoscopes is inevitably constrained by the remote center of motion (RCM) during the process of FOV adjustment. The RCM limits lapar…

Cited by 2SourceScholar
2024

EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection

CVPR 2024poster

In the era of AI-generated content (AIGC) malicious tampering poses imminent threats to copyright integrity and information security. Current deep image watermarking while widely accepted for safeguarding visual content can only protect copyright and ensure traceability. They fall short in localizin…

Cited by 54SourcePDFScholar
2024

Long-Tailed Learning as Multi-Objective Optimization

AAAI 2024technical

Real-world data is extremely imbalanced and presents a long-tailed distribution, resulting in models biased towards classes with sufficient samples and performing poorly on rare classes. Recent methods propose to rebalance classes but they undertake the seesaw dilemma (what is increasing performance…

2024

OmniSSR: Zero-shot Omnidirectional Image Super-Resolution using Stable Diffusion Model

ECCV 2024oral

"Omnidirectional images (ODIs) are commonly used in real-world visual tasks, and high-resolution ODIs help improve the performance of related visual tasks. Most existing super-resolution methods for ODIs use end-to-end learning strategies, resulting in inferior realness of generated images and a lac…