← Search

Guoqing Wang

33 accepted papers

2026

Careful Queries, Credible Results: Teaching RAG Models Advanced Web Search Tools with Reinforcement Learning

AAAI 2026technical

Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating up-to-date external knowledge, yet real-world web environments present unique challenges. These limitations manifest as two key challenges: pervasive misinformation in the web environment, which introduces unre

Cited by 0SourcePDFScholar
2026

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR

CVPR 2026

Optical Character Recognition (OCR) is fundamental to Vision-Language Models (VLMs) and high-quality data generation for LLM training. Yet, despite progress in average OCR accuracy, state-of-the-art VLMs still struggle with detecting sample-level errors and lack effective unsupervised quality contro

Cited by 0SourcecodeScholar
2026

Experience Transfer for Multimodal LLM Agents in Minecraft Game

CVPR 2026

Multimodal LLM agents operating in complex game environments must continually reuse past experience to solve new tasks efficiently. In this work, we propose Echo, a transfer-oriented memory framework that enables agents to derive actionable knowledge from prior interactions rather than treating memo

Cited by 0SourceScholar
2026

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

ICLR 2026poster

Recent attempts to transfer features from 2D Vision–Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields noisy and fragmented predictions, whereas enforcing geometric coherence necessitates costly training pipelines and larg…

Cited by 0SourcecodeScholar
2026

Grounding Everything in Tokens for Multimodal Large Language Models

CVPR 2026

Multimodal large language models (MLLMs) have made significant advancements in vision understanding and reasoning. However, the autoregressive Transformer architecture used by MLLMs requires tokenization on input images, which limits their ability to accurately ground objects within the 2D image spa

Cited by 0SourceScholar
2026

Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents

ICLR 2026poster

Large language model (LLM)–based agents are increasingly trained with reinforcement learning (RL) to enhance their ability to interact with external environments through tool use, particularly in search-based settings that require multi-turn reasoning and knowledge acquisition. However, existing app…

Cited by 0SourcecodeScholar
2026

LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models

ICLR 2026poster

Large multimodal models (LMMs) have achieved impressive performance on various vision-language tasks, but their substantial computational and memory costs hinder their practical deployment. Existing compression methods often decouple low-rank decomposition and quantization, leading to compounded rec…

Cited by 0SourceScholar
2026

MetaEval: Measuring the Discrimination of Benchmarks for Efficient LLM Evaluation

AAAI 2026technical

Benchmarks serve as standardized test systems to distinguish capabilities among large language models (LLMs). Discriminative items enable high-ability LLMs to favor correct answers, while causing low-ability models to assign lower plausibility to these answers and tend toward incorrect answers. Curr

Cited by 0SourcePDFScholar
2026

Seeing Beyond Illusion: Generalized and Efficient Mirror Detection

AAAI 2026technical

Reflective imaging enables the mirror imagings and physical entities to possess identical attributes, e.g., color and shape. Current mirror detection (MD) methods primarily rely on designing functional components to establish the correlation and disparities between the imagings and entities, thereby

Cited by 0SourcePDFScholar
2026

Text summarization via global structure awareness

ICLR 2026poster

Text summarization is a core task in natural language processing (NLP). With the rapid growth of information, handling long documents has become increasingly demanding, making summarization essential. Existing research mainly focuses on model improvements and sentence-level pruning, but often overlo…

Cited by 0SourceScholar
2025

Decoupling Metacognition from Cognition: A Framework for Quantifying Metacognitive Ability in LLMs

AAAI 2025technical

Large Language Models (LLMs) are known to hallucinate facts and make non-factual statements which can undermine trust in their output. The essence of hallucination lies in the absence of metacognition in LLMs, namely the understanding of their own cognitive processes. However, there has been limited…

2025

Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

NeurIPS 2025poster

The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiote…

Cited by 0SourceScholar
2025

Disentangled Modeling of Preferences and Social Influence for Group Recommendation

AAAI 2025technical

The group recommendation (GR) aims to suggest items for a group of users in social networks. Existing work typically considers individual preferences as the sole factor in aggregating group preferences. Actually, social influence is also an important factor in modeling users' contributions to the fi…

2025

Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy

ICCV 2025poster

A prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived thro…

2025

Hybrid Boundary Physics-Informed Neural Networks for Solving Navier-Stokes Equations with Complex Boundary

NeurIPS 2025poster

Physics-informed neural networks (PINN) have achieved notable success in solving partial differential equations (PDE), yet solving the Navier-Stokes equations (NSE) with complex boundary conditions remains a challenging task. In this paper, we introduce a novel Hybrid Boundary PINN (HB-PINN) method…

Cited by 0SourceScholar
2025

Implicit Counterfactual Learning for Audio-Visual Segmentation

ICCV 2025poster

Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfac…

Cited by 0SourcePDFScholar
2025

Iterative Predictor-Critic Code Decoding for Real-World Image Dehazing

CVPR 2025poster

We propose a novel Iterative Predictor-Critic Code Decoding framework for real-world image dehazing, abbreviated as IPC-Dehaze, which leverages the high-quality codebook prior encapsulated in a pre-trained VQGAN. Apart from previous codebook-based methods that rely on one-shot decoding, our method u…

2025

Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual Processing

AAAI 2025technical

Event cameras encode visual information by generating asynchronous and sparse event streams, which hold great potential for low latency and low power consumption. Despite many successful implementations of event camera-based applications, most of them accumulate the events into frames and then utili…

Cited by 0SourcePDFScholar
2025

S$^2$NN: Sub-bit Spiking Neural Networks

NeurIPS 2025poster

Spiking Neural Networks (SNNs) offer an energy-efficient paradigm for machine intelligence, but their continued scaling poses challenges for resource-limited deployment. Despite recent advances in binary SNNs, the storage and computational demands remain substantial for large-scale networks. To furt…

Cited by 0SourceScholar
2025

Towards Generalizable Multi-Camera 3D Object Detection via Perspective Rendering

AAAI 2025technical

Detecting and localizing objects in 3D space using multiple cameras, known as Multi-Camera 3D Object Detection (MC3D-Det), has gained prominence with the advent of bird's-eye view (BEV) approaches. However, these methods often struggle with the serious domain gaps caused by various viewpoints and en…

2025

VSS-SLAM: Voxelized Surfel Splatting for Geometally Accurate SLAM

ICRA 2025

[1] Visual Simultaneous Localization and Mapping (SLAM) helps robots estimate their poses and perceive the environment in unknown settings. Recent work has demonstrated that implicit neural radiance fields and 3D Gaussian Splatting (3DGS) offer higher fidelity scene representation than traditional m

Cited by 1SourceScholar
2024

Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance

CoRL 2024poster

In this work, we address the challenging problem of long-horizon goal-reaching policy learning from non-expert, action-free observation data. Unlike fully labeled expert data, our data is more accessible and avoids the costly process of action labeling. Additionally, compared to online learning, whi…

Cited by 1SourcecodeScholar
2024

OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving

ECCV 2024poster

"Existing 3D semantic occupancy prediction methods typically treat the task as a one-shot 3D voxel-wise segmentation problem, focusing on a single-step mapping between the inputs and occupancy maps, which limits their ability to refine and complete local regions gradually. In this paper, we introduc…

Cited by 21SourcePDFScholar
2024

Open-Vocabulary Calibration for Fine-tuned CLIP

ICML 2024poster

Vision-language models (VLMs) have emerged as formidable tools, showing their strong capability in handling various open-vocabulary tasks in image recognition, text-driven visual content generation, and visual chatbots, to name a few. In recent years, considerable efforts and resources have been dev…

2024

Physics-Constrained Comprehensive Optical Neural Networks

NeurIPS 2024poster

With the advantages of low latency, low power consumption, and high parallelism, optical neural networks (ONN) offer a promising solution for time-sensitive and resource-limited artificial intelligence applications. However, the performance of the ONN model is often diminished by the gap between the…

Cited by 1SourcePDFScholar
2024

ScanERU: Interactive 3D Visual Grounding Based on Embodied Reference Understanding

AAAI 2024technical

Aiming to link natural language descriptions to specific regions in a 3D scene represented as 3D point clouds, 3D visual grounding is a very fundamental task for human-robot interaction. The recognition errors can significantly impact the overall accuracy and then degrade the operation of AI systems…

2024

SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction

CVPR 2024poster

Vision-based perception for autonomous driving requires an explicit modeling of a 3D space where 2D latent representations are mapped and subsequent 3D operators are applied. However operating on dense latent spaces introduces a cubic time and space complexity which limits scalability in terms of pe…

Cited by 42SourcePDFScholar
2024

VEON: Vocabulary-Enhanced Occupancy Prediction

ECCV 2024poster

"Perceiving the world as 3D occupancy supports embodied agents to avoid collision with any types of obstacle. While open-vocabulary image understanding has prospered recently, how to bind the predicted 3D occupancy grids with open-world semantics still remains under-explored due to limited open-worl…

2024

Weakly-Supervised Mirror Detection via Scribble Annotations

AAAI 2024technical

Mirror detection is of great significance for avoiding false recognition of reflected objects in computer vision tasks. Existing mirror detection frameworks usually follow a supervised setting, which relies heavily on high quality labels and suffers from poor generalization. To resolve this, we inst…

2023

Cross-Subject Mental Fatigue Detection based on Separable Spatio-Temporal Feature Aggregation

ICASSP 2023accepted

Cross-subject mental fatigue detection via Electroencephalography (EEG) is challenging because EEG from different individuals varies greatly. Existing works have exploited domain adaption to alleviate the individual discrepancy due to personality, gender and so on. However, the distributions of data…

Cited by 0SourceScholar
2023

Learning Semantic-Aware Knowledge Guidance for Low-Light Image Enhancement

CVPR 2023poster

Low-light image enhancement (LLIE) investigates how to improve illumination and produce normal-light images. The majority of existing methods improve low-light images via a global and uniform manner, without taking into account the semantic information of different regions. Without semantic priors,…

2020

Cross-Domain Face Presentation Attack Detection via Multi-Domain Disentangled Representation Learning

CVPR 2020poster

Face presentation attack detection (PAD) has been an urgent problem to be solved in the face recognition systems. Conventional approaches usually assume the testing and training are within the same domain; as a result, they may not generalize well into unseen scenarios because the representations le…

Cited by 234PDFcodeScholar