← Search

Chengxin Liu

7 accepted papers

2026

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow

CVPR 2026

Vision-Language Models (VLMs) have demonstrated strong capability in a wide range of tasks such as visual recognition, document parsing, and visual grounding. Nevertheless, recent work shows that while VLMs often manage to capture the correct image region corresponding to the question, they do not n

Cited by 0SourcecodeScholar
2026

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

CVPR 2026

Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate event occurrence. In contrast, human speech carries fundamentally different, ri

Cited by 0SourcecodeScholar
2024

LPS-Net: Lightweight Parameter-shared Network for Point Cloud-based Place Recognition

ICRA 2024poster

With innovation in fields such as autonomous driving and augmented reality, point cloud-based place recognition has gained significant attention. Many methods try to address this problem by extracting and matching global descriptors in a database, but they often must balance the extraction of compre…

Cited by 5SourcecodeScholar
2022

Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic Counting

CVPR 2022poster

Class-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with query images to infer object counts. Two essential components in this pipeline are feature representation and similarit…

Cited by 109PDFcodeScholar
2022

Robust Object Detection with Inaccurate Bounding Boxes

ECCV 2022poster

"Learning accurate object detectors often requires large-scale training data with precise object bounding boxes. However, labeling such data is expensive and time-consuming. As the crowd-sourcing labeling process and the ambiguities of the objects may raise noisy bounding box annotations, the object…

2019

From Open Set to Closed Set: Counting Objects by Spatial Divide-and-Conquer

ICCV 2019poster

Visual counting, a task that predicts the number of objects from an image/video, is an open-set problem by nature, i.e., the number of population can vary in [0,+[?]) in theory. However, the collected images and labeled count values are limited in reality, which means only a small closed set is obse…

Cited by 214PDFcodeScholar