← Search

Qingjie Liu

29 accepted papers

2026

Diffusion Trajectory-Guided Policy for Long-Horizon Robot Manipulation

ICRA 2026poster

Recently, Vision-Language-Action Models (VLA) have advanced robot imitation learning, but high data collection costs and limited demonstrations hinder generalization and current imitation learning methods struggle in out-of-distribution scenarios, especially for long-horizon tasks. A key challenge i…

2026

LIBERO-X: Robustness Litmus for Vision-Language-Action Models

RSS 2026poster

Reliable benchmarking is critical for advancing Vision–Language–Action (VLA) models, as it reveals their generalization, robustness, and alignment of perception with language-driven manipulation tasks. However, existing benchmarks often provide limited or misleading assessments due to insufficient e…

Cited by 0SourceScholar
2026

Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision

CVPR 2026

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to image-level anomaly detection and textual reasoning, while pixel-level localization still relies on external vision mod

Cited by 0SourcecodeScholar
2026

ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving

ICLR 2026poster

The comprehensive understanding capabilities of world models for driving scenarios have significantly improved the planning accuracy of end-to-end autonomous driving frameworks. However, the redundant modeling of static regions and the lack of deep interaction with trajectories hinder world models f…

Cited by 0SourcecodeScholar
2026

Semantic-Aware Motion Encoding for Topology-Agnostic Character Animation

ICML 2026poster

Generalizing motion representation across diverse characters remains challenging due to significant topological variations in skeletal structures across datasets and species, which hinders the development of scalable generative models. To bridge this gap, we propose a Semantic-Aware Topology-Agnosti…

Cited by 0SourceScholar
2026

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

ICML 2026oral

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (\textit{i}) producing precise low-level actions from high-dimensional observations, (\t…

Cited by 0SourcecodeScholar
2025

Diffusion Trajectory-Guided Policy for Long-Horizon Robot Manipulation

RA-L 2025

Recently, Vision-Language-Action models (VLA) have advanced robot imitation learning, but high data collection costs and limited demonstrations hinder generalization and current imitation learning methods struggle in out-of-distribution scenarios, especially for long-horizon tasks. A key challenge i

Cited by 14SourcecodeScholar
2025

GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art

ACL 2025long

***Video Comment Art*** enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural and contextual subtleties. Although Multimodal Large Language Models (MLLMs) and Chain-of-Thought (CoT) have demo…

2025

GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection

AAAI 2025technical

Bird's-Eye-View (BEV) representation has emerged as a mainstream paradigm for multi-view 3D object detection, demonstrating impressive perceptual capabilities. However, existing methods overlook the geometric quality of BEV representation, leaving it in a low-resolution state and failing to restore…

2025

KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus

NAACL 2025findings

Video-based dialogue systems have compelling application value, such as education assistants, thereby garnering growing interest. However, the current video-based dialogue systems are limited by their reliance on a single dialogue type, which hinders their versatility in practical applications acros…

2025

OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing Images

ICCV 2025poster

Remote sensing object detection has made significant progress, but most studies still focus on closed-set detection, limiting generalization across diverse datasets. Open-vocabulary object detection (OVD) provides a solution by leveraging multimodal associations between text prompts and visual featu…

2025

SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking

CVPR 2025poster

Most state-of-the-art trackers adopt one-stream paradigm, using a single Vision Transformer for joint feature extraction and relation modeling of template and search region images. However, relation modeling between different image patches exhibits significant variations. For instance, background re…

2025

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding

CVPR 2025poster

With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmarks focus on standalone videos and only assess "visual elements" like human action…

2025

SkeletonMix: A Mixup-Based Data Augmentation Framework for Skeleton-Based Action Recognition

ICASSP 2025accepted

Skeleton-based human action recognition has received widespread attention for its robustness to changes in the background and appearance of actors compared to the RGB modality. Data augmentation is widely used to explicitly regularize the model to prevent overfitting, especially when the number of l…

Cited by 0SourceScholar
2025

Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic Segmentation

AAAI 2025technical

Training-free open-vocabulary semantic segmentation aims to explore the potential of frozen vision-language models (VLM) for segmentation tasks. Recent works reform the inference process of CLIP and utilize the features from the final layer to reconstruct dense representations for segmentation, dem…

Cited by 0SourcePDFScholar
2024

ActiveDC: Distribution Calibration for Active Finetuning

CVPR 2024poster

The pretraining-finetuning paradigm has gained popularity in various computer vision tasks. In this paradigm the emergence of active finetuning arises due to the abundance of large-scale data and costly annotation requirements. Active finetuning involves selecting a subset of data from an unlabeled…

Cited by 2SourcePDFScholar
2024

DSD-DA: Distillation-based Source Debiasing for Domain Adaptive Object Detection

ICML 2024poster

Though feature-alignment based Domain Adaptive Object Detection (DAOD) methods have achieved remarkable progress, they ignore the source bias issue, i.e., the detector tends to acquire more source-specific knowledge, impeding its generalization capabilities in the target domain. Furthermore, these m…

Cited by 2SourcePDFScholar
2024

Read, Spell and Repeat: Scene Text Recognition with Vision-Language Circular Refinement

ICASSP 2024accepted

Scene Text Recognition (STR) has long been considered an important yet challenging task in the field of computer vision. Recent works have demonstrated that utilizing language information is effective for the visually difficult images, like ones with occultation or blurring. However, the use of lang…

Cited by 0SourceScholar
2023

BISVP: Building Footprint Extraction Via Bidirectional Serialized Vertex Prediction

ICASSP 2023accepted

Extracting building footprints from remote sensing images has been attracting extensive attention recently. Dominant approaches address this challenging problem by generating vectorized building masks with cumbersome refinement stages, which limits the application of such methods. In this paper, we…

Cited by 0SourceScholar
2023

Learning Discriminative Representations for Skeleton Based Action Recognition

CVPR 2023poster

Human action recognition aims at classifying the category of human action from a segment of a video. Recently, people have dived into designing GCN-based models to extract features from skeletons for performing this task, because skeleton representations are much more efficient and robust than other…

2023

SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object Detection

ICCV 2023poster

Recently, the pure camera-based Bird's-Eye-View (BEV) perception provides a feasible solution for economical autonomous driving. However, the existing BEV-based multi-view 3D detectors generally transform all image features into BEV features, without considering the problem that the large proportion…

Cited by 34PDFcodeScholar
2022

Learning from Future: A Novel Self-Training Framework for Semantic Segmentation

NeurIPS 2022accept

Self-training has shown great potential in semi-supervised learning. Its core idea is to use the model learned on labeled data to generate pseudo-labels for unlabeled samples, and in turn teach itself. To obtain valid supervision, active attempts typically employ a momentum teacher for pseudo-label…

2022

SparseTT: Visual Tracking with Sparse Transformers

IJCAI 2022poster

Transformers have been successfully applied to the visual tracking task and significantly promote tracking performance. The self-attention mechanism designed to model long-range dependencies is the key to the success of Transformers. However, self-attention lacks focusing on the most relevant inform…

2021

STMTrack: Template-Free Visual Tracking With Space-Time Memory Networks

CVPR 2021poster

Boosting performance of the offline trained siamese trackers is getting harder nowadays since the fixed information of the template cropped from the first frame has been almost thoroughly mined, but they are poorly capable of resisting target appearance changes. Existing trackers with template updat…

Cited by 364PDFcodeScholar
2018

Hough Transform Guided Deep Feature Extraction for Dense Building Detection in Remote Sensing Images

ICASSP 2018accepted

Detecting dense buildings without elevation information is an important and challenging task in remote sensing applications. In this paper, we present a novel cascaded deep neural network architecture, incorporating multi -stage region proposal detection and Hough transform to obtain better mid-leve…

Cited by 0SourceScholar