← Search

Zheng Qin

25 accepted papers

2026

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs

AAAI 2026technical

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of e

Cited by 0SourcePDFScholar
2026

PENTESTLLMAGENT: A Task Dependency Graph Planning-Based Multi-Agent Framework for Automated Penetration Testing

IJCAI 2026

Fully autonomous IP-to-Root penetration testing remains challenging for LLM agents. We conduct an exploratory study on 10 LLMs and introduce AutoPentest-Bench, an end-to-end benchmark with 13 VulnHub targets and 93 sub-tasks. From 130 interaction logs, we identify three challenges: Rigid Strategy, C

Cited by 0Scholar
2026

Spatial Matters: Position-Guided 3D Referring Expression Segmentation

CVPR 2026

3D Referring Expression segmentation (3D-RES) is an emerging field that segments 3D objects in point cloud scenes based on given referring expressions. Although existing methods have achieved substantial progress, they primarily focus on semantic cues and often overlook spatial relations, which are

Cited by 0SourcecodeScholar
2025

CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs

ICCV 2025poster

Object goal navigation (ObjectNav) is a fundamental task in embodied AI, requiring an agent to locate a target object in previously unseen environments. This task is particularly challenging because it requires both perceptual and cognitive processes, including object recognition and decision-making…

Cited by 0SourcePDFScholar
2025

RefDetector: A Simple Yet Effective Matching-based Method for Referring Expression Comprehension

AAAI 2025technical

Despite the rapid and substantial advancements in object detection, it continues to face limitations imposed by pre-defined category sets. Current methods for visual grounding primarily focus on how to better leverage the visual backbone to generate text-tailored visual features, which may require a…

Cited by 0SourcePDFScholar
2025

Towards Precise Embodied Dialogue Localization via Causality Guided Diffusion

CVPR 2025poster

Embodied localization based on vision and natural language dialogues presents a persistent challenge in embodied intelligence. Existing methods often approach this task as an image translation problem, leveraging encoder-decoder architectures to predict heatmaps. However, these methods frequently ex…

Cited by 0SourcePDFScholar
2024

Are Watermarks Bugs for Deepfake Detectors? Rethinking Proactive Forensics

IJCAI 2024poster

AI-generated content has accelerated the topic of media synthesis, particularly Deepfake, which can manipulate our portraits for positive or malicious purposes. Before releasing these threatening face images, one promising forensics solution is the injection of robust watermarks to track their own p…

2024

Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces

IJCAI 2024poster

Deepfake videos are becoming increasingly realistic, showing few tampering traces on facial areas that vary between frames. Consequently, existing Deepfake detection methods struggle to detect unknown domain Deepfake videos while accurately locating the tampered region. To address this limitation,…

2024

Fuel-Saving Route Planning with Data-Driven and Learning-Based Approaches – A Systematic Solution for Harbor Tugs

IJCAI 2024poster

In recent years, there are trends toward cleaner port environments through enforcement by imposed legislation. Transit optimisation of fuel-based port service boats like harbour tugs has emerged as a critical task to reduce fuel consumption and carbon emission. In this paper, an innovative learning-…

Cited by 3SourcePDFScholar
2024

Learning Instance-Aware Correspondences for Robust Multi-Instance Point Cloud Registration in Cluttered Scenes

CVPR 2024poster

Multi-instance point cloud registration estimates the poses of multiple instances of a model point cloud in a scene point cloud. Extracting accurate point correspondences is to the center of the problem. Existing approaches usually treat the scene point cloud as a whole overlooking the separation of…

2024

Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identification

AAAI 2024technical

The fine-grained attribute descriptions can significantly supplement the valuable semantic information for person image, which is vital to the success of person re-identification (ReID) task. However, current ReID algorithms typically failed to effectively leverage the rich contextual information av…

2024

Referencing Where to Focus: Improving Visual Grounding with Referential Query

NeurIPS 2024poster

Visual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional effort…

Cited by 1SourcePDFScholar
2024

Towards Generalizable Multi-Object Tracking

CVPR 2024poster

Multi-Object Tracking (MOT) encompasses various tracking scenarios each characterized by unique traits. Effective trackers should demonstrate a high degree of generalizability across diverse scenarios. However existing trackers struggle to accommodate all aspects or necessitate hypothesis and experi…

2023

2D3D-MATR: 2D-3D Matching Transformer for Detection-Free Registration Between Images and Point Clouds

ICCV 2023poster

The commonly adopted detect-then-match approach to registration finds difficulties in the cross-modality cases due to the incompatible keypoint detection and inconsistent feature description. We propose, 2D3D-MATR, a detection-free method for accurate and robust registration between images and point…

Cited by 18PDFcodeScholar
2023

Deep Graph-Based Spatial Consistency for Robust Non-Rigid Point Cloud Registration

CVPR 2023poster

We study the problem of outlier correspondence pruning for non-rigid point cloud registration. In rigid registration, spatial consistency has been a commonly used criterion to discriminate outliers from inliers. It measures the compatibility of two correspondences by the discrepancy between the resp…

2023

Instance-Aware Hierarchical Structured Policy for Prompt Learning in Vision-Language Models

ICASSP 2023accepted

In recent years, learnable prompts have emerged as a major prompt learning paradigm, enhancing the performance of large-scale vision-language pre-trained models in few-shot image classification. However, enhancing methods are often time-consuming and inflexible because 1) class-specific prompts are…

Cited by 0SourceScholar
2023

MotionTrack: Learning Robust Short-Term and Long-Term Motions for Multi-Object Tracking

CVPR 2023poster

The main challenge of Multi-Object Tracking (MOT) lies in maintaining a continuous trajectory for each target. Existing methods often learn reliable motion patterns to match the same target between adjacent frames and discriminative appearance features to re-identify the lost targets after a long pe…

2023

Rotation-Invariant Transformer for Point Cloud Matching

CVPR 2023poster

The intrinsic rotation invariance lies at the core of matching point clouds with handcrafted descriptors. However, it is widely despised by recent deep matchers that obtain the rotation invariance extrinsically via data augmentation. As the finite number of augmented rotations can never span the con…

2022

FInfer: Frame Inference-Based Deepfake Detection for High-Visual-Quality Videos

AAAI 2022technical

Deepfake has ignited hot research interests in both academia and industry due to its potential security threats. Many countermeasures have been proposed to mitigate such risks. Current Deepfake detection methods achieve superior performances in dealing with low-visual-quality Deepfake media which ca…

2022

Geometric Transformer for Fast and Robust Point Cloud Registration

CVPR 2022oral

We study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods bypass the detection of repeatable keypoints which is difficult in low-overlap scenarios, showing great potential in registration. They seek correspondences over downsampled superpo…

Cited by 454PDFcodeScholar
2022

Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast

IJCAI 2022poster

We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice and face. These methods neglect the semantic content of different videos, introd…

2021

A Features Decoupling Method for Multiple Manipulations Identification in Image Operation Chains

ICASSP 2021accepted

Recently, many forensic techniques have been developed to detect the use of a certain processing operation. When utilizing several manipulations to alter an image, artifacts left by manipulations that have been applied later can potentially disguise traces left by manipulations that were applied ear…

Cited by 0SourceScholar
2021

Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval

CVPR 2021poster

Given an untrimmed video and a query sentence, cross-modal video moment retrieval aims to rank a video moment from pre-segmented video moment candidates that best matches the query sentence. Pioneering work typically learns the representations of the textual and visual content separately and then ob…

Cited by 86PDFcodeScholar
2019

ThunderNet: Towards Real-Time Generic Object Detection on Mobile Devices

ICCV 2019poster

Real-time generic object detection on mobile platforms is a crucial but challenging computer vision task. Prior lightweight CNN-based detectors are inclined to use one-stage pipeline. In this paper, we investigate the effectiveness of two-stage detectors in real-time generic detection and propose a…

Cited by 287PDFScholar