← Search

Ming Kong

9 accepted papers

2026

Linking Perception, Confidence and Accuracy in MLLMs

CVPR 2026

Recent advances in Multi-modal Large Language Models (MLLMs) have predominantly focused on enhancing visual \perception to improve \accuracy. However, a critical question remains unexplored: Do models know when they do not know? Through a probing experiment, we reveal a severe \confidence miscalibra

Cited by 0SourcecodeScholar
2026

MAS-Architect: Declarative Multi-Agent System Design via Separation of Concerns

ICML 2026poster

The Automated Design of Multi-Agent Systems (Auto-MAS) has emerged as a promising framework for addressing complex reasoning tasks. However, existing approaches often suffer from structural rigidity and entangle the design of system topology with the implementation of individual agents. To overcome …

Cited by 0SourceScholar
2026

META: Meta Evolution of Tool Trajectory Adaptation for Long-Video Understanding

CVPR 2026

Long-video understanding remains challenging due to extreme temporal redundancy, sparse yet decisive events, and the instability of long-horizon reasoning in visual-language models (VLMs). Existing agent-based methods invoke external micro-tools but remain static, repeatedly rebuilding long chains o

Cited by 0SourceScholar
2026

NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting

ICLR 2026poster

With the rapid development and adoption of 3D Gaussian Splatting (3DGS), the need for effective copyright protection has become increasingly critical. Existing watermarking techniques for 3DGS mainly focus on protecting rendered images via pre-trained decoders, leaving the underlying 3D Gaussian pri…

Cited by 0SourceScholar
2025

MHBench: Demystifying Motion Hallucination in VideoLLMs

AAAI 2025technical

Similar to Language or Image LLMs, VideoLLMs are also plagued by hallucination issues. Hallucinations in videos not only manifest in the spatial dimension regarding the perception of the existence of visual objects (static) but also the temporal dimension influencing the perception of actions and ev…

2025

MoLE:Decoding by Mixture of Layer Experts Alleviates Hallucination in Large Vision-Language Models

AAAI 2025technical

Recent advancements in Large Vision-Language Models (LVLMs) highlight their ability to integrate and process multi-modal information. However, hallucinations—where generated content is inconsistent with input vision and instructions—remain a challenge. In this paper, we analyze LVLMs' layer-wise dec…

2024

Querying as Prompt: Parameter-Efficient Learning for Multimodal Language Model

CVPR 2024poster

Recent advancements in language models pre-trained on large-scale corpora have significantly propelled developments in the NLP domain and advanced progress in multimodal tasks. In this paper we propose a Parameter-Efficient multimodal language model learning strategy named QaP (Querying as Prompt).…

Cited by 5SourcePDFScholar