← Search

Qichuan Geng

8 accepted papers

2026

GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian Updating

AAAI 2026technical

Image geo-localization aims to determine the geographic location of a query image. While Multimodal Large Language Models (MLLMs) show potential for this task due to their rich world knowledge and explainable abilities, they often struggle with confirmation bias, i.e., committing to early, potential

Cited by 0SourcePDFScholar
2026

VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation

CVPR 2026

The rapid evolution of generative AI, including such models as Sora, has intensified the threat of video misinformation. A critical challenge in detecting these AI-generated video misinformation lies in a fundamental disconnect between existing datasets and practical deception tactics. Current datas

Cited by 0SourceScholar
2025

3D Lane Detection Based on Projection-Consistent Reference Points and Intra- & Inter-lane Context

ICRA 2025

3D lane detection aims to identify lane categories and trends in 3D space, which is a vital and challenging task in autonomous driving. Existing methods introduce various priors to guide 3D lane prediction, which generally consist of a series of reference points for context aggregation. However, due

Cited by 1SourceScholar
2025

Boosting Few-Shot Open-Set Object Detection via Prompt Learning and Robust Decision Boundary

IJCAI 2025

Few-shot Open-set Object Detection (FOOD) poses a challenge in many open-world scenarios. It aims to train an open-set detector to detect known objects while rejecting unknowns with scarce training samples. Existing FOOD methods are subject to limited visual information, and often exhibit an ambiguo

2025

Incremental Few-Shot Semantic Segmentation via Multi-Level Switchable Visual Prompts

ICCV 2025poster

Existing incremental few-shot semantic segmentation (IFSS) methods often learn novel classes by fine-tuning parameters from previous stages. This inevitably reduces the distinguishability of old class features, leading to catastrophic forgetting and overfitting to limited new samples. In this paper,…

2025

Understanding Matters: Semantic-Structural Determined Visual Relocalization for Large Scenes

IJCAI 2025

Scene Coordinate Regression (SCR) estimates 3D scene coordinates from 2D images, and has become an important approach in visual relocalization. Existing methods exhibit high localization accuracy in small scenes, but still face substantial challenges in large-scale scenes, which usually have signifi

Cited by 0SourcePDFScholar
2023

Exploiting 3D Human Recovery for Action Recognition with Spatio-Temporal Bifurcation Fusion

ICASSP 2023accepted

Action recognition utilizes information in images or videos to analyze and classify human behaviors. The existing methods usually exploit 2D pose to improve classification features. Due to the lack of 3D cues, some approximate behaviors in 2D perspective cannot be recognized. In this paper, we propo…

Cited by 0SourceScholar
2021

Frequency Domain Image Translation: More Photo-Realistic, Better Identity-Preserving

ICCV 2021poster

Image-to-image translation has been revolutionized with GAN-based methods. However, existing methods lack the ability to preserve the identity of the source domain. As a result, synthesized images can often over-adapt to the reference domain, losing important structural characteristics and suffering…

Cited by 106PDFcodeScholar