← Search

Shuai Chen

22 accepted papers

2026

Backjump-on-Graph: Empowering LLMs with Reinforced Retrospective Exploration for Agentic KG Reasoning

ICML 2026poster

Grounding Large Language Models (LLMs) in Knowledge Graphs (KGs) has shown significant promise for complex Question Answering (QA) tasks. Since LLMs' limited context window cannot accommodate the sheer volume of large-scale KGs, existing work usually utilizes agents to reason on real-world KGs, whic…

Cited by 0SourceScholar
2026

CPiRi: Channel Permutation-Invariant Relational Interaction for Multivariate Time Series Forecasting

ICLR 2026poster

Current methods for multivariate time series forecasting can be classified into channel-dependent and channel-independent models. Channel-dependent models learn cross-channel features but often overfit the channel ordering, which hampers adaptation when channels are added or reordered. Channel-indep…

Cited by 0SourcecodeScholar
2026

Do 3D Large Language Models Really Understand 3D Spatial Relationships?

ICLR 2026poster

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on text-only question-answer pairs can perform comparably or even surpass these methods on the SQA3D benchmark without using…

Cited by 0SourceScholar
2026

EventGait: Towards Robust Gait Recognition with Event Streams

CVPR 2026

Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensitivity in conventional cameras. In this work, we explore gait recognition using event cameras, which offer microsecond temporal resolution and high d

Cited by 0SourcecodeScholar
2026

Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling

ICLR 2026poster

The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answers in verifiable knowledge sources, such as Knowledge Graphs (KGs). Prevailing KG-enhanced methods typically constrain…

Cited by 0SourceScholar
2025

Aesthetic Perception Prompting for Interpretable Image Aesthetics Assessment with MLLMs

ICASSP 2025accepted

Image Aesthetic Assessment (IAA) aims to rate the aesthetic quality of images and has many practical applications. However, existing methods typically rely on limited annotated data for training, leading to two key issues: 1) score-only predictions lack interpretability, making it hard for users to…

Cited by 0SourceScholar
2025

Compress Large Language Models via Collaboration Between Learning and Matrix Approximation

NeurIPS 2025poster

Sparse and low-rank matrix composite approximation has emerged as a promising paradigm for compressing large language models (LLMs), offering a more flexible pruning structure than conventional methods based solely on sparse matrices. The significant variation in weight redundancy across layers, alo…

Cited by 0SourceScholar
2025

CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection

IJCAI 2025

With the swift progression of image generation technology, the widespread emergence of facial deepfakes poses significant challenges to the field of security, thus amplifying the urgent need for effective deepfake detection. Existing techniques for face forgery detection can broadly be categorized i

Cited by 0SourcePDFScholar
2025

Efficient Representativeness-Aware Coreset Selection

NeurIPS 2025poster

Dynamic coreset selection is a promising approach for improving the training efficiency of deep neural networks by periodically selecting a small subset of the most representative or informative samples, thereby avoiding the need to train on the entire dataset. However, it remains inherently challen…

Cited by 0SourceScholar
2025

GS-CPR: Efficient Camera Pose Refinement via 3D Gaussian Splatting

ICLR 2025poster

We leverage 3D Gaussian Splatting (3DGS) as a scene representation and propose a novel test-time camera pose refinement (CPR) framework, GS-CPR. This framework enhances the localization accuracy of state-of-the-art absolute pose regression and scene coordinate regression methods. The 3DGS model rend…

Cited by 7SourcePDFScholar
2025

RETAIL: Towards Real-world Travel Planning for Large Language Models

EMNLP 2025

Although large language models have enhanced automated travel planning abilities, current systems remain misaligned with real-world scenarios. First, they assume users provide explicit queries, while in reality requirements are often implicit. Second, existing solutions ignore diverse environmental

Cited by 0SourcePDFScholar
2024

Fast Cross-Modality Knowledge Transfer via a Contextual Autoencoder Transformation

ICASSP 2024accepted

Cross-modality knowledge transfer aims to apply knowledge learned in the source modality to the target modality. It is more challenging than the general knowledge transfer task because of the aggravated modality shift problem due to introducing heterogeneous data. This paper proposes a novel fast cr…

Cited by 0SourceScholar
2024

HR-APR: APR-agnostic Framework with Uncertainty Estimation and Hierarchical Refinement for Camera Relocalisation

ICRA 2024poster

Absolute Pose Regressors (APRs) directly estimate camera poses from monocular images, but their accuracy is unstable for different queries. Uncertainty-aware APRs provide uncertainty information on the estimated pose, alleviating the impact of these unreliable predictions. However, existing uncertai…

Cited by 6SourcecodeScholar
2024

Information Bottleneck Based Data Correction in Continual Learning

ECCV 2024poster

"Continual Learning (CL) requires model to retain previously learned knowledge while learning new tasks. Recently, experience replay-based methods have made significant progress in addressing this challenge. These methods primarily select data from old tasks and store them in a buffer. When learning…

Cited by 0SourcePDFScholar
2024

KnowVrDU: A Unified Knowledge-aware Prompt-Tuning Framework for Visually-rich Document Understanding

COLING 2024main

In Visually-rich Document Understanding (VrDU), recent advances of incorporating layout and image features into the pre-training language models have achieved significant progress. Existing methods usually developed complicated dedicated architectures based on pre-trained models and fine-tuned them…

2024

Map-Relative Pose Regression for Visual Re-Localization

CVPR 2024highlight

Pose regression networks predict the camera pose of a query image relative to a known environment. Within this family of methods absolute pose regression (APR) has recently shown promising accuracy in the range of a few centimeters in position error. APR networks encode the scene geometry implicitly…

2024

Neural Refinement for Absolute Pose Regression with Feature Synthesis

CVPR 2024poster

Absolute Pose Regression (APR) methods use deep neural networks to directly regress camera poses from RGB images. However the predominant APR architectures only rely on 2D operations during inference resulting in limited accuracy of pose estimation due to the lack of 3D geometry constraints or prior…

2024

Scene Coordinate Reconstruction: Posing of Image Collections via Incremental Learning of a Relocalizer

ECCV 2024oral

"We address the task of estimating camera parameters from a set of images depicting a scene. Popular feature-based structure-from-motion (SfM) tools solve this task by incremental reconstruction: they repeat triangulation of sparse 3D points and registration of more camera views to the sparse point…

2024

Task-Wise Prompt Query Function for Rehearsal-Free Continual Learning

ICASSP 2024accepted

Continual learning (CL) aims to enable a model to retain knowledge of old tasks while learning new ones. One effective approach to CL is based on data rehearsal method. However, this approach increases the cost of storing data and cannot be used when data from old tasks is unavailable for some reaso…

Cited by 0SourceScholar
2023

Dialogue Rewriting via Skeleton-Guided Generation

AAAI 2023technical

Dialogue rewriting aims to transform multi-turn, context-dependent dialogues into well-formed, context-independent text for most NLP systems. Previous dialogue rewriting benchmarks and systems assume a fluent and informative utterance to rewrite. Unfortunately, dialogue utterances from real-world sy…

2023

Pay More Attention to Relation Exploration for Knowledge Base Question Answering

ACL 2023findings

Knowledge base question answering (KBQA) is a challenging task that aims to retrieve correct answers from large-scale knowledge bases. Existing attempts primarily focus on entity representation and final answer reasoning, which results in limited supervision for this task. Moreover, the relations, w…

2022

DFNet: Enhance Absolute Pose Regression with Direct Feature Matching

ECCV 2022poster

"We introduce a camera relocalization pipeline that combines absolute pose regression (APR) and direct feature matching. By incorporating exposure-adaptive novel view synthesis, our method successfully addresses photometric distortions in outdoor environments that existing photometric-based methods…