← Search

Qiang Zhu

17 accepted papers

2026

Linking Perception, Confidence and Accuracy in MLLMs

CVPR 2026

Recent advances in Multi-modal Large Language Models (MLLMs) have predominantly focused on enhancing visual \perception to improve \accuracy. However, a critical question remains unexplored: Do models know when they do not know? Through a probing experiment, we reveal a severe \confidence miscalibra

Cited by 0SourcecodeScholar
2026

MAS-Architect: Declarative Multi-Agent System Design via Separation of Concerns

ICML 2026poster

The Automated Design of Multi-Agent Systems (Auto-MAS) has emerged as a promising framework for addressing complex reasoning tasks. However, existing approaches often suffer from structural rigidity and entangle the design of system topology with the implementation of individual agents. To overcome …

Cited by 0SourceScholar
2026

META: Meta Evolution of Tool Trajectory Adaptation for Long-Video Understanding

CVPR 2026

Long-video understanding remains challenging due to extreme temporal redundancy, sparse yet decisive events, and the instability of long-horizon reasoning in visual-language models (VLMs). Existing agent-based methods invoke external micro-tools but remain static, repeatedly rebuilding long chains o

Cited by 0SourceScholar
2026

NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting

ICLR 2026poster

With the rapid development and adoption of 3D Gaussian Splatting (3DGS), the need for effective copyright protection has become increasingly critical. Existing watermarking techniques for 3DGS mainly focus on protecting rendered images via pre-trained decoders, leaving the underlying 3D Gaussian pri…

Cited by 0SourceScholar
2026

Trajectory-aware Shifted State Space Models for Online Video Super-Resolution

ICLR 2026poster

Online video super-resolution (VSR) is an important technique for many real-world video processing applications, which aims to restore the current high-resolution video frame based on temporally previous frames. Most of the existing online VSR methods solely employ one neighboring previous frame to…

Cited by 0SourcecodeScholar
2026

UCPO: Uncertainty-Aware Policy Optimization

ICML 2026poster

The key to building trustworthy Large Language Models (LLMs) lies in endowing them with inherent uncertainty expression capabilities to mitigate the hallucinations that restrict their high-stakes applications. However, existing RL paradigms such as GRPO often suffer from Advantage Bias due to binary…

Cited by 0SourceScholar
2025

Blind Video Super-Resolution based on Implicit Kernels

ICCV 2025poster

Blind video super-resolution (BVSR) is a low-level vision task which aims to generate high-resolution videos from low-resolution counterparts in unknown degradation scenarios. Existing approaches typically predict blur kernels that are spatially invariant in each video frame or even the entire video…

2025

MHBench: Demystifying Motion Hallucination in VideoLLMs

AAAI 2025technical

Similar to Language or Image LLMs, VideoLLMs are also plagued by hallucination issues. Hallucinations in videos not only manifest in the spatial dimension regarding the perception of the existence of visual objects (static) but also the temporal dimension influencing the perception of actions and ev…

2025

MoLE:Decoding by Mixture of Layer Experts Alleviates Hallucination in Large Vision-Language Models

AAAI 2025technical

Recent advancements in Large Vision-Language Models (LVLMs) highlight their ability to integrate and process multi-modal information. However, hallucinations—where generated content is inconsistent with input vision and instructions—remain a challenge. In this paper, we analyze LVLMs' layer-wise dec…

2025

SymmCD: Symmetry-Preserving Crystal Generation with Diffusion Models

ICLR 2025poster

Generating novel crystalline materials has potential to lead to advancements in fields such as electronics, energy storage, and catalysis. The defining characteristic of crystals is their symmetry, which plays a central role in determining their physical properties. However, existing crystal generat…

2024

CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality Enhancement

CVPR 2024poster

Recently numerous approaches have achieved notable success in compressed video quality enhancement (VQE). However these methods usually ignore the utilization of valuable coding priors inherently embedded in compressed videos such as motion vectors and residual frames which carry abundant temporal a…

2024

EiffHDR: An Efficient Network for Multi-Exposure High Dynamic Range Imaging

ICASSP 2024accepted

While recent progress in Multi-exposure HDR imaging is promising, the growing complexity of state-of-the-art (SOTA) methods poses challenges for their analysis and comparison. In this paper, we analyze the motivations and approaches behind previous SOTA works and introduce EiffHDR, an efficient Mult…

Cited by 0SourceScholar
2024

OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal

ECCV 2024poster

"Deep learning-based methods have shown remarkable performance in single JPEG artifacts removal task. However, existing methods tend to degrade on double JPEG images, which are prevalent in real-world scenarios. To address this issue, we propose Offset-Aware Partition Transformer for double JPEG art…

2024

Querying as Prompt: Parameter-Efficient Learning for Multimodal Language Model

CVPR 2024poster

Recent advancements in language models pre-trained on large-scale corpora have significantly propelled developments in the NLP domain and advanced progress in multimodal tasks. In this paper we propose a Parameter-Efficient multimodal language model learning strategy named QaP (Querying as Prompt).…

Cited by 5SourcePDFScholar