← Search

Qingchao Chen

24 accepted papers

2026

Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

ICML 2026poster

Current image editing methods excels at static attributes but fails at complex Human-Object Interactions (HOI), a critical challenge unaddressed by existing benchmarks that conflate HOI with static attributes, relying on global metrics incapable of simultaneously assessing dynamic interaction validi…

Cited by 0SourceScholar
2025

CubeDN: Real-Time Drone Detection in 3D Space from Dual mmWave Radar Cubes

ICRA 2025

As drone use has become more widespread, there is a critical need to ensure safety and security. A key element of this is robust and accurate drone detection and localization. While cameras and other optical sensors like LiDAR are commonly used for object detection, their performance degrades under

Cited by 1SourceScholar
2025

Joint Class-level and Instance-level Relationship Modeling for Novel Class Discovery

AAAI 2025technical

Novel class discovery(NCD) aims to cluster the unlabeled data with the help of a labeled set containing different but related classes. The key to solving NCD is the knowledge transfer between labeled and unlabeled sets.Since NCD requires that known classes and unknown classes are related, it is sign…

Cited by 0SourcePDFScholar
2025

Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration

ICCV 2025poster

Open Vocabulary Human-Object Interaction (HOI) detection aims to detect interactions between humans and objects while generalizing to novel interaction classes beyond the training set. Current methods often rely on Vision and Language Models (VLMs) but face challenges due to suboptimal image encoder…

2025

Radar2ECG: Multi-Scale Bottleneck Fusion and Cross-modal Semantic Distillation for Conditional Electrocardiogram Generation from Radar Heart Sound

ICASSP 2025accepted

The field of conditional Electrocardiogram(ECG) generation focuses on generating specified ECGs under given conditions for medical purposes. Existing methods are typically based on conditions of simple inputs like text or lead types. However, they struggle to handle the complexity of radar heart sou…

Cited by 0SourceScholar
2025

TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge Transferring

ICCV 2025poster

Dynamic Scene Graph Generation (DSGG) aims to create a scene graph for each video frame by detecting objects and predicting their relationships. Weakly Supervised DSGG (WS-DSGG) reduces annotation workload by using an un- localized scene graph from a single frame per video for training. Existing WS-…

2024

3D Vision and Language Pretraining with Large-Scale Synthetic Data

IJCAI 2024poster

3D Vision-Language Pre-training (3D-VLP) aims to provide a pre-train model which can bridge 3D scenes with natural language, which is an important technique for embodied intelligence. However, current 3D-VLP datasets are hindered by limited scene-level diversity and insufficient fine-grained annot…

2024

Novel Class Discovery in Chest X-rays via Paired Images and Text

AAAI 2024technical

Novel class discover(NCD) aims to identify new classes undefined during model training phase with the help of knowledge of known classes. Many methods have been proposed and notably boosted performance of NCD in natural images. However, there has been no work done in discovering new classes based on…

2024

OED: Towards One-stage End-to-End Dynamic Scene Graph Generation

CVPR 2024poster

Dynamic Scene Graph Generation (DSGG) focuses on identifying visual relationships within the spatial-temporal domain of videos. Conventional approaches often employ multi-stage pipelines which typically consist of object detection temporal association and multi-relation classification. However these…

2024

Training-free Video Temporal Grounding using Large-scale Pre-trained Models

ECCV 2024poster

"Video temporal grounding aims to identify video segments within untrimmed videos that are most relevant to a given natural language query. Existing video temporal localization models rely on specific datasets for training, with high data collection costs, but exhibit poor generalization capability…

2023

Confidence-aware Pseudo-label Learning for Weakly Supervised Visual Grounding

ICCV 2023poster

Visual grounding aims at localizing the target object in image which is most related to the given free-form natural language query. As labeling the position of target object is labor-intensive, the weakly supervised methods, where only image-sentence annotations are required during model training ha…

Cited by 12PDFcodeScholar
2023

Efficient Adaptive Human-Object Interaction Detection with Concept-guided Memory

ICCV 2023poster

Human Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare classes and the high computational cost and time required to…

Cited by 24PDFcodeScholar
2023

Masked Retraining Teacher-Student Framework for Domain Adaptive Object Detection

ICCV 2023poster

Domain adaptive Object Detection (DAOD) leverages a labeled domain (source) to learn an object detector generalizing to a novel domain without annotation (target). Recent advances use a teacher-student framework, i.e., a student model is supervised by the pseudo labels from a teacher model. Though g…

Cited by 31PDFcodeScholar
2023

Phrase-Level Temporal Relationship Mining for Temporal Sentence Localization

AAAI 2023technical

In this paper, we address the problem of video temporal sentence localization, which aims to localize a target moment from videos according to a given language query. We observe that existing models suffer from a sheer performance drop when dealing with simple phrases contained in the sentence. It r…

2022

Mixed In Time And Modality: Curse Or Blessingƒ Cross-Instance Data Augmentation for Weakly Supervised Multimodal Temporal Fusion

ICASSP 2022accepted

In multimodal video event localization, we usually leverage feature fusion across different axes, such as the modality and temporal axes, for better context. To reduce the costs of detailed annotations, recent solutions explore weakly supervised settings. However, we observe that when feature fusion…

Cited by 0SourceScholar
2022

Weakly Supervised Temporal Sentence Grounding With Gaussian-Based Contrastive Proposal Learning

CVPR 2022poster

Temporal sentence grounding aims to detect the most salient moment corresponding to the natural language query from untrimmed videos. As labeling the temporal boundaries is labor-intensive and subjective, the weakly-supervised methods have recently received increasing attention. Most of the existing…

Cited by 109PDFcodeScholar
2022

Weakly Supervised Video Moment Localization with Contrastive Negative Sample Mining

AAAI 2022technical

Video moment localization aims at localizing the video segments which are most related to the given free-form natural language query. The weakly supervised setting, where only video level description is available during training, is getting more and more attention due to its lower annotation cost. P…

2018

Re-Weighted Adversarial Adaptation Network for Unsupervised Domain Adaptation

CVPR 2018poster

Unsupervised Domain Adaptation (UDA) aims to transfer domain knowledge from existing well-defined tasks to new ones where labels are unavailable. In the real-world applications, as the domain (task) discrepancies are usually uncontrollable, it is significantly motivated to match the feature distribu…

Cited by 174SourcePDFScholar
2015

Indoor target tracking using high doppler resolution passive Wi-Fi radar

ICASSP 2015accepted

This paper describes two Doppler only indoor passive Wi-Fi tracking methods based on high Doppler resolution passive radar. Two filters are investigated in this paper, the extended Kalman filter and the sequential importance resampling (SIR) particle filter. Experimental results for these two tracki…

Cited by 0SourceScholar