← Search

Jie Qin

65 accepted papers

2026

DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenges: constructing an effective long-term memory bank and overcoming the compounding errors problem. To address these issue

Cited by 0SourceScholar
2026

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos

ICML 2026poster

The rapid advancement of AI-driven video generation has transformed content creation, while simultaneously increasing the risk of misinformation through localized manipulations in long-form videos. Existing video forensic methods predominantly operate on short, independent clips, and thus fail to ca…

Cited by 0SourceScholar
2026

GarmentGPT: Compositional Garment Pattern Generation via Discrete Latent Tokenization

ICLR 2026poster

Apparel is a fundamental component of human appearance, making garment digitalization critical for digital human creation. However, sewing pattern creation traditionally relies on the intuition and extensive experience of skilled artisans. This manual bottleneck significantly hinders the scalability…

Cited by 0SourcecodeScholar
2026

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation

CVPR 2026

Inspired by the general Vision-and-Language Navigation (VLN) task, aerial VLN has attracted widespread attention, owing to its significant practical value in applications such as logistics delivery and urban inspection. However, existing methods face several challenges in complex urban environments,

Cited by 0SourceScholar
2026

History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation

AAAI 2026technical

Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental reasoning and local scene comprehension, existing UAV agent

Cited by 0SourcePDFScholar
2026

Instruction Decomposition and Action Alignment for Vision-Language Navigation

ICML 2026poster

Vision-and-Language Navigation (VLN) empowered by Multimodal Large Language Models (MLLMs) is promise, yet remains challenged by long-horizon tasks with complex user instructions. Existing approaches that continuously condition on full instructions incur high latency due to abundant visual tokens an…

Cited by 0SourceScholar
2026

LoFA: Learning to Predict Personalized Prior for Fast Adaptation of Visual Generative Models

CVPR 2026

Personalizing visual generative models to meet specific user needs has gained increasing attention, yet current methods like Low-Rank Adaptation (LoRA) remain impractical due to their demand for task-specific data and lengthy optimization. While a few hypernetwork-based approaches attempt to predict

Cited by 0SourcecodeScholar
2026

RunawayEvil: Jailbreaking the Image-to-Video Generative Models

CVPR 2026

Image-to-Video (I2V) generation represents a frontier in content creation, where models synthesize dynamic visual sequences by jointly reasoning from both image and text prompts. This multimodal grounding enables diverse controllability over video attributes. However, it is precisely this capability

Cited by 0SourcecodeScholar
2026

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization

ICLR 2026poster

Vector quantization (VQ) is a key component in discrete tokenizers for image generation, but its training is often unstable due to straight-through estimation bias, one-step-behind updates, and sparse codebook gradients, which lead to suboptimal reconstruction performance and low codebook usage. In…

Cited by 0SourceScholar
2026

Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion

AAAI 2026technical

Point cloud completion is a fundamental task in 3D vision. A persistent challenge in this field is simultaneously preserving fine-grained details present in the input while ensuring the global structural integrity of the completed shape. While recent works leveraging local symmetry transformations v

Cited by 0SourcePDFScholar
2026

Text-guided Controllable Diffusion for Realistic Camouflage Images Generation

AAAI 2026technical

Camouflage Images Generation (CIG) is an emerging research area that focuses on synthesizing images in which objects are harmoniously blended and exhibit high visual consistency with their surroundings. Existing methods perform CIG by either fusing objects into specific backgrounds or outpainting th

Cited by 0SourcePDFScholar
2025

Doubly Contrastive Learning for Source-Free Domain Adaptive Person Search

AAAI 2025technical

Domain Adaptive Person Search (DAPS) aims to improve the generalization capability of person search models by training on both labeled source data and unlabeled target data, which is not that practical in real-world applications considering the storage/transmission costs and the privacy of source da…

Cited by 0SourcePDFScholar
2025

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

NeurIPS 2025poster

We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed videos containing relevant segments using multi-modal queries. It has the following features: **1) Practical retrieval par…

Cited by 0SourcecodeScholar
2025

MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection

NeurIPS 2025poster

Video Anomaly Detection (VAD) aims to locate unusual activities or behaviors within videos. Recently, offline VAD has garnered substantial research attention, which has been invigorated by the progress in large language models (LLMs) and vision-language models (VLMs), offering the potential for a mo…

Cited by 0SourcecodeScholar
2025

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

ICCV 2025poster

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-per…

2025

VITRIX-UniViTAR: Unified Vision Transformer with Native Resolution

NeurIPS 2025poster

Conventional Vision Transformer streamlines visual modeling by employing a uniform input resolution, which underestimates the inherent variability of natural visual data and incurs a cost in spatial-contextual fidelity. While preliminary explorations have superficially investigated native resolution…

Cited by 0SourceScholar
2024

CDFormer: When Degradation Prediction Embraces Diffusion Model for Blind Image Super-Resolution

CVPR 2024poster

Existing Blind image Super-Resolution (BSR) methods focus on estimating either kernel or degradation information but have long overlooked the essential content details. In this paper we propose a novel BSR approach Content-aware Degradation-driven Transformer (CDFormer) to capture both degradation a…

2024

Cross-Domain Few-Shot Semantic Segmentation via Doubly Matching Transformation

IJCAI 2024poster

Cross-Domain Few-shot Semantic Segmentation (CD-FSS) aims to train generalized models that can segment classes from different domains with a few labeled images. Previous works have proven the effectiveness of feature transformation in addressing CD-FSS. However, they completely rely on support image…

2024

DSMix: Distortion-Induced Saliency Map Based Pre-training for No-Reference Image Quality Assessment

ECCV 2024poster

"Image quality assessment (IQA) has long been a fundamental challenge in image understanding. In recent years, deep learning-based IQA methods have shown promising performance. However, the lack of large amounts of labeled data in the IQA field has hindered further advancements of these methods. Thi…

2024

Puff-Net: Efficient Style Transfer with Pure Content and Style Feature Fusion Network

CVPR 2024poster

Style transfer aims to render an image with the artistic features of a style image while maintaining the original structure. Various methods have been put forward for this task but some challenges still exist. For instance it is difficult for CNN-based methods to handle global information and long-r…

2024

Relevant Intrinsic Feature Enhancement Network for Few-Shot Semantic Segmentation

AAAI 2024technical

For few-shot semantic segmentation, the primary task is to extract class-specific intrinsic information from limited labeled data. However, the semantic ambiguity and inter-class similarity of previous methods limit the accuracy of pixel-level foreground-background classification. To alleviate these…

Cited by 16SourcePDFScholar
2024

Transformer-Based No-Reference Image Quality Assessment via Supervised Contrastive Learning

AAAI 2024technical

Image Quality Assessment (IQA) has long been a research hotspot in the field of image processing, especially No-Reference Image Quality Assessment (NR-IQA). Due to the powerful feature extraction ability, existing Convolution Neural Network (CNN) and Transformers based NR-IQA methods have achieved c…

2024

Unified Unsupervised Salient Object Detection via Knowledge Transfer

IJCAI 2024poster

Recently, unsupervised salient object detection (USOD) has gained increasing attention due to its annotation-free nature. However, current methods mainly focus on specific tasks such as RGB and RGB-D, neglecting the potential for task migration. In this paper, we propose a unified USOD framework for…

2024

WPS-SAM: Towards Weakly-Supervised Part Segmentation with Foundation Models

ECCV 2024oral

"Segmenting and recognizing diverse object parts is crucial in computer vision and robotics. Despite significant progress in object segmentation, part-level segmentation remains underexplored due to complex boundaries and scarce annotated data. To address this, we propose a novel Weakly-supervised P…

2023

AlignDet: Aligning Pre-training and Fine-tuning in Object Detection

ICCV 2023poster

The paradigm of large-scale pre-training followed by downstream fine-tuning has been widely employed in various object detection algorithms. In this paper, we reveal discrepancies in data, model, and task between the pre-training and fine-tuning procedure in existing practices, which implicitly limi…

Cited by 22PDFcodeScholar
2023

Feature Shrinkage Pyramid for Camouflaged Object Detection With Transformers

CVPR 2023poster

Vision transformers have recently shown strong global context modeling capabilities in camouflaged object detection. However, they suffer from two major limitations: less effective locality modeling and insufficient feature aggregation in decoders, which are not conducive to camouflaged object detec…

2023

FreeSeg: Unified, Universal and Open-Vocabulary Image Segmentation

CVPR 2023poster

Recently, open-vocabulary learning has emerged to accomplish segmentation for arbitrary categories of text-based descriptions, which popularizes the segmentation system to more general-purpose application scenarios. However, existing methods devote to designing specialized architectures or parameter…

2023

ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target Detection

ICASSP 2023accepted

Small targets are often submerged in cluttered backgrounds of infrared images. Conventional detectors tend to generate false alarms, while CNN-based detectors lose small targets in deep layers. To this end, we propose iSmallNet, a multi-stream densely nested network with label decoupling for infrare…

Cited by 0SourceScholar
2023

Memory-Aided Contrastive Consensus Learning for Co-salient Object Detection

AAAI 2023technical

Co-salient object detection (CoSOD) aims at detecting common salient objects within a group of relevant source images. Most of the latest works employ the attention mechanism for finding common objects. To achieve accurate CoSOD results with high-quality maps and high efficiency, we propose a novel…

2023

Movienet-PS: A Large-Scale Person Search Dataset in the Wild

ICASSP 2023accepted

Person search (PS) aims to jointly localize and identify a query person from natural, uncropped images. Existing works unintentionally adopt pedestrians (with similar poses and unchanging clothing) as the query and restrict the application scenarios in surveillance. This is due to that most PS datas…

Cited by 0SourceScholar
2023

Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation for Anomaly Detection

ICCV 2023poster

Anomaly detection (AD), aiming to find samples that deviate from the training distribution, is essential in safety-critical applications. Though recent self-supervised learning based attempts achieve promising results by creating virtual outliers, their training objectives are less faithful to AD wh…

Cited by 10PDFScholar
2022

ACGNet: Action Complement Graph Network for Weakly-Supervised Temporal Action Localization

AAAI 2022technical

Weakly-supervised temporal action localization (WTAL) in untrimmed videos has emerged as a practical but challenging task since only video-level labels are available. Existing approaches typically leverage off-the-shelf segment-level features, which suffer from spatial incompleteness and temporal in…

Cited by 64SourcePDFScholar
2022

Activation Modulation and Recalibration Scheme for Weakly Supervised Semantic Segmentation

AAAI 2022technical

Image-level weakly supervised semantic segmentation (WSSS) is a fundamental yet challenging computer vision task facilitating scene understanding and automatic driving. Most existing methods resort to classification-based Class Activation Maps (CAMs) to play as the initial pseudo labels, which tend…

2022

Exploring Visual Context for Weakly Supervised Person Search

AAAI 2022technical

Person search has recently emerged as a challenging task that jointly addresses pedestrian detection and person re-identification. Existing approaches follow a fully supervised setting where both bounding box and identity annotations are available. However, annotating identities is labor-intensive,…

2022

Motion Sensitive Contrastive Learning for Self-Supervised Video Representation

ECCV 2022poster

"Contrastive learning has shown great potential in video representation learning. However, existing approaches fail to sufficiently exploit short-term motion dynamics, which are crucial to various down-stream video understanding tasks. In this paper, we propose Motion Sensitive Contrastive Learning…

Cited by 20SourcePDFScholar
2022

Multi-Granularity Distillation Scheme towards Lightweight Semi-Supervised Semantic Segmentation

ECCV 2022poster

"Albeit with varying degrees of progress in the field of Semi-Supervised Semantic Segmentation, most of its recent successes are involved in unwieldy models and the lightweight solution is still not yet explored. We find that existing knowledge distillation techniques pay more attention to pixel-lev…

2022

PACE: Predictive and Contrastive Embedding for Unsupervised Action Segmentation

IJCAI 2022poster

Action segmentation, inferring temporal positions of human actions in an untrimmed video, is an important prerequisite for various video understanding tasks. Recently, unsupervised action segmentation (UAS) has emerged as a more challenging task due to the unavailability of frame-level annotations.…

Cited by 0SourcePDFScholar
2022

Video Anomaly Detection by Solving Decoupled Spatio-Temporal Jigsaw Puzzles

ECCV 2022poster

"Video Anomaly Detection (VAD) is an important topic in computer vision. Motivated by the recent advances in self-supervised learning, this paper addresses VAD by solving an intuitive yet challenging pretext task, i.e., spatio-temporal jigsaw puzzles, which is cast as a multi-label fine-grained clas…

2021

Contour Primitive of Interest Extraction Network Based on One-Shot Learning for Object-Agnostic Vision Measurement

ICRA 2021poster

Image contour based vision measurement is widely applied in robot manipulation and industrial automation. It is appealing to realize object-agnostic vision system, which can be conveniently reused for various types of objects. We propose the contour primitive of interest extraction network (CPieNet)…

Cited by 7SourceScholar
2021

P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching

ICCV 2021poster

Accurately describing and detecting 2D and 3D keypoints is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint dete…

Cited by 62PDFcodeScholar
2021

Partial Is Better Than All: Revisiting Fine-tuning Strategy for Few-shot Learning

AAAI 2021technical

The goal of few-shot learning is to learn a classifier that can recognize unseen classes from limited support data with labels. A common practice for this task is to train a model on the base set first and then transfer to novel classes through fine-tuning or meta-learning. However, as the base clas…

Cited by 193SourcePDFScholar
2021

S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution Calibration

CVPR 2021poster

Previous studies dominantly target at self-supervised learning on real-valued networks and have achieved many promising results. However, on the more challenging binary neural networks (BNNs), this task has not yet been fully explored in the community. In this paper, we focus on this more difficult…

Cited by 23PDFcodeScholar
2020

Layer-wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs

ECCV 2020poster

Conditioning analysis uncovers the landscape of an optimization objective by exploring the spectrum of its curvature matrix. This has been well explored theoretically for linear models. We extend this analysis to deep neural networks (DNNs) in order to investigate their learning dynamics. To this en…

Cited by 12SourcePDFScholar
2020

Learning Attentive and Hierarchical Representations for 3D Shape Recognition

ECCV 2020poster

This paper proposes a novel method for 3D shape representation learning, namely Hyperbolic Embedded Attentive Representation (HEAR). Different from existing multi-view based methods, HEAR develops a unified framework to address both multi-view redundancy and single-view incompleteness. Specifically,…

Cited by 36SourcePDFScholar
2020

Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification

CVPR 2020poster

Video-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), t…

Cited by 189PDFcodeScholar
2020

Region Graph Embedding Network for Zero-Shot Learning

ECCV 2020poster

Most of the existing Zero-Shot Learning (ZSL) approaches learn direct embeddings from global features or image parts (regions) to the semantic space, which, however, fail to capture the appearance relationships between different local regions within a single image. In this paper, to model the relati…

Cited by 195SourcePDFScholar
2019

Attentive Region Embedding Network for Zero-Shot Learning

CVPR 2019poster

Zero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of t…

Cited by 351PDFScholar
2019

Deep Sketch-Shape Hashing With Segmented 3D Stochastic Viewing

CVPR 2019poster

Sketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH),…

Cited by 48PDFScholar
2019

KE-GAN: Knowledge Embedded Generative Adversarial Networks for Semi-Supervised Scene Parsing

CVPR 2019poster

In recent years, scene parsing has captured increasing attention in computer vision. Previous works have demonstrated promising performance in this task. However, they mainly utilize holistic features, whilst neglecting the rich semantic knowledge and inter-object relationships in the scene. In addi…

Cited by 59PDFScholar
2018

Hierarchical Attention and Context Modeling for Group Activity Recognition

ICASSP 2018accepted

Group activity recognition in videos is a challenging task, with two major issues, i.e. attending to those persons and their body parts that contribute significantly to the activity, and modeling contextual person structures in the group. Most previous approaches fail to provide a practical solution…

Cited by 0SourceScholar
2018

Highly-Economized Multi-View Binary Compression for Scalable Image Clustering

ECCV 2018poster

How to economically cluster large-scale multi-view images is a long-standing problem in computer vision. To tackle this challenge, this paper introduces a novel approach named Highly-economized Scalable Image Clustering (HSIC) that radically surpasses conventional image clustering methods via binary…

Cited by 55SourcePDFScholar
2018

TBN: Convolutional Neural Network with Ternary Inputs and Binary Weights

ECCV 2018poster

Despite the remarkable success of Convolutional Neural Networks (CNNs) on generalized visual tasks, high computational and memory costs restrict their comprehensive applications on consumer electronics (e.g., portable or smart wearable devices). Recent advancements in binarized networks have demonst…

2018

stagNet: An Attentive Semantic RNN for Group Activity Recognition

ECCV 2018poster

Group activity recognition plays a fundamental role in a variety of applications, e.g. sports video analysis and intelligent surveillance. How to model the spatio-temporal contextual information in a scene still remains a crucial yet challenging issue. We propose a novel attentive semantic recurrent…

Cited by 180SourcePDFScholar
2017

Binary Coding for Partial Action Analysis With Limited Observation Ratios

CVPR 2017poster

Traditional action recognition methods aim to recognize actions with complete observations/executions. However, it is often difficult to capture fully executed actions due to occlusions, interruptions, etc. Meanwhile, action prediction/recognition in advance based on partial observations is essentia…

Cited by 34PDFScholar
2017

Fast Person Re-Identification via Cross-Camera Semantic Binary Transformation

CVPR 2017poster

Numerous methods have been proposed for person re-identification, most of which however neglect the matching efficiency. Recently, several hashing based approaches have been developed to make re-identification more scalable for large-scale gallery sets. Despite their efficiency, these works ignore c…

Cited by 94PDFScholar
2017

Zero-Shot Action Recognition With Error-Correcting Output Codes

CVPR 2017poster

Recently, zero-shot action recognition (ZSAR) has emerged with the explosive growth of action categories. In this paper, we explore ZSAR from a novel perspective by adopting the Error-Correcting Output Codes (dubbed ZSECOC). Our ZSECOC equips the conventional ECOC with the additional capability of Z…

Cited by 186PDFScholar