← Search

Yu Wu

84 accepted papers

2026

AdaNODEs: Test Time Adaptation for Time Series Forecasting Using Neural ODEs

ICASSP 2026poster

Test time adaptation (TTA) has emerged as a promising solution to adapt pre-trained models to new, unseen data distributions using unlabeled target domain data. However, most TTA methods are designed for independent data, often overlooking the time series data and rarely addressing forecasting tasks…

Cited by 0SourcePDFScholar
2026

Beyond Hearing: Learning Task-agnostic ExG Representations from Earphones via Physiology-informed Tokenization

ICLR 2026poster

Electrophysiological (ExG) signals offer valuable insights into human physiology, yet building foundation models that generalize across everyday tasks remains challenging due to two key limitations: (i) insufficient data diversity, as most ExG recordings are collected in controlled labs with bulky,…

Cited by 0SourceScholar
2026

Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning

ICLR 2026poster

Large Language Models (LLMs) are widely used as judges to evaluate response quality, providing a scalable alternative to human evaluation. However, most LLM judges operate solely on intrinsic text-based reasoning, limiting their ability to verify complex constraints or perform accurate computation.…

Cited by 0SourceScholar
2026

MIST: Moment-Aligned Invariant Stability Transform for Robust Flow Matching

ICML 2026poster

Classifier-Free Guidance (CFG) is a cornerstone of flow-matching models, significantly enhancing visual quality and prompt adherence. However, high guidance scales inherently violate the optimal transport dynamics, leading to visual artifacts and mode collapse. In this paper, we investigate the mech…

Cited by 0SourceScholar
2026

RETHINKING LARGE LANGUAGE MODELS FOR IRREGULAR TIME SERIES CLASSIFICATION IN CRITICAL CARE

ICASSP 2026oral

Time series data from the Intensive Care Unit (ICU) provides critical information for patient monitoring. While recent advancements in applying Large Language Models (LLMs) to time series modeling (TSM) have shown great promise, their effectiveness on the irregular ICU data, characterized by particu…

Cited by 0SourcePDFScholar
2026

Restoring Initial Noise Sensitivity in Text-to-Image Distillation through Geometric Alignment

ICML 2026poster

Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models while preserving perceptual quality. However, existing distillation methods prioritize efficiency and output fidelity, often overlooking the preservati…

Cited by 0SourceScholar
2026

RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident Anticipation

CVPR 2026

Accident anticipation aims to predict impending collisions from dashcam videos and trigger early alerts. Existing methods rely on binary supervision with manually annotated "anomaly onset" frames, which are subjective and inconsistent, leading to inaccurate risk estimation. In contrast, we propose R

Cited by 0SourcecodeScholar
2026

Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models

ICML 2026poster

We present a mechanism to steer the sampling diversity of denoising diffusion and flow matching models, allowing users to sample from a sharper or broader distribution than the training distribution. We build on the observation that these models leverage (learned) score functions of noisy data distr…

Cited by 0SourceScholar
2026

Towards One-step Causal Video Generation via Adversarial Self-Distillation

ICLR 2026poster

Recent hybrid video generation models combine autoregressive temporal dynamics with diffusion-based spatial denoising, but their sequential, iterative nature leads to error accumulation and long inference times. In this work, we propose a distillation-based framework for efficient causal video gener…

Cited by 0SourcecodeScholar
2026

Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection

CVPR 2026

Unmanned aerial vehicle (UAV) based object detection is a critical but challenging task, when applied in dynamically changing scenarios with limited annotated training data. Layout-to-image generation approaches have proved effective in promoting detection accuracy by synthesizing labeled images bas

Cited by 0SourcecodeScholar
2025

Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement

CVPR 2025poster

Generalized Category Discovery (GCD) aims to recognize unlabeled images from known and novel classes by distinguishing novel classes from known ones, while also transferring knowledge from another set of labeled images with known classes. Existing GCD methods rely on self-supervised vision transform…

Cited by 0SourcePDFScholar
2025

An Online Motion Planning Framework for Navigating Torpedo-shaped Autonomous Underwater Vehicles in Unknown Underwater Environments

IROS 2025

Navigating unknown underwater environments is a significant challenge for autonomous underwater vehicles (AUVs), especially those with torpedo-like shapes. Lacking a prior map, these vehicles rely on real-time sensor data for perception. Although online motion planning addresses this challenge, many

Cited by 0SourceScholar
2025

BNMusic: Blending Environmental Noises into Personalized Music

NeurIPS 2025poster

While being disturbed by environmental noises, the acoustic masking technique is a conventional way to reduce the annoyance in audio engineering that seeks to cover up the noises with other dominant yet less intrusive sounds. However, misalignment between the dominant sound and the noise—such as mis…

Cited by 0SourcecodeScholar
2025

CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

CVPR 2025highlight

Interleaved image-text generation has emerged as a crucial multimodal task, aiming at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models (MLLMs), generating integrated image-text sequences that exhibit n…

2025

CodeIO: Condensing Reasoning Patterns via Code Input-Output Prediction

ICML 2025oral

Reasoning is a fundamental capability of Large Language Models. While prior research predominantly focuses on enhancing narrow skills like math or code generation, improving performance on many other reasoning tasks remains challenging due to sparse and fragmented training data. To address this issu…

2025

DPSN: Dual Prior Knowledge Induced Tactile paving and Obstacle Joint Segmentation Network

IROS 2025

Accurate semantic segmentation of both tactile paving and the obstacle is crucial for the safe mobility of visually impaired individuals. However, existing methods face two major challenges: (i) discontinuous segmentation fragments; (ii) Inaccurate obstacle recognition. To address challenge (i), we

Cited by 0SourceScholar
2025

D^3: Scaling Up Deepfake Detection by Learning from Discrepancy

CVPR 2025poster

The boom of Generative AI brings opportunities entangled with risks and concerns. Existing literature emphasizes the generalization capability of deepfake detection on unseen generators, significantly promoting the detector's ability to identify more universal artifacts. This work seeks a step towar…

2025

Implicit Bias Injection Attacks against Text-to-Image Diffusion Models

CVPR 2025poster

The proliferation of text-to-image diffusion models (T2I DMs) has led to an increased presence of AI-generated images in daily life. However, biased T2I models can generate content with specific tendencies, potentially influencing people's perceptions. Intentional exploitation of these biases risks…

2025

RQTalker: Speech-driven 3D Facial Animation via Region-aware Vector Quantization

ICASSP 2025accepted

Speech-driven 3D facial animation has been a long-standing topic due to the complex geometry and motion modeling as well as difficulties in cross-modality learning. Current studies struggle to synthesize human-like lip motions, as they usually represent the movement of the entire face with a compres…

Cited by 0SourceScholar
2025

Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model

CVPR 2025poster

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g., objects) have not been well investigated. Despite human hand sy…

2025

Rethinking Query-based Transformer for Continual Image Segmentation

CVPR 2025poster

Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods of…

2025

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation

IJCAI 2025

Video-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language, a critical component of artistic expression in filmmaking. As a result, their performance deteriorates in scenarios wher

Cited by 0SourcePDFScholar
2025

The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation

ICCV 2025poster

In this work, we introduce NoiseQuery as a novel method for enhanced noise initialization in versatile goal-driven text-to-image (T2I) generation. Specifically, we propose to leverage an aligned Gaussian noise as implicit guidance to complement explicit user-defined inputs, such as text prompts, for…

2025

WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration

AAAI 2025technical

LLM-based autonomous agents often fail to execute complex web tasks that require dynamic interaction, largely due to the inherent uncertainty and complexity of these environments. Existing LLM-based web agents typically rely on rigid, expert-designed policies specific to certain states and actions,…

Cited by 21SourcePDFScholar
2024

Diffusion in Diffusion: Cyclic One-Way Diffusion for Text-Vision-Conditioned Generation

ICLR 2024poster

Originating from the diffusion phenomenon in physics that describes particle movement, the diffusion generative models inherit the characteristics of stochastic random walk in the data space along the denoising trajectory. However, the intrinsic mutual interference among image regions contradicts th…

2024

Improving Bird's Eye View Semantic Segmentation by Task Decomposition

CVPR 2024poster

Semantic segmentation in bird's eye view (BEV) plays a crucial role in autonomous driving. Previous methods usually follow an end-to-end pipeline directly predicting the BEV segmentation map from monocular RGB inputs. However the challenge arises when the RGB inputs and BEV targets from distinct per…

2024

Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning

CVPR 2024poster

Generative vision-language models (VLMs) have shown impressive performance in zero-shot vision-language tasks like image captioning and visual question answering.However improving their zero-shot reasoning typically requires second-stage instruction tuning which relies heavily on human-labeled or la…

2024

Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models

EMNLP 2024main

Parameter-efficient fine-tuning (PEFT) is crucial for customizing Large Language Models (LLMs) with constrained resource. Although there have been various PEFT methods for dense-architecture LLMs, PEFT for sparse-architecture LLMs is still underexplored. In this work, we study the PEFT method for LL…

2024

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

ACL 2024long

In this paper, we present an innovative process-oriented math process reward model called Math-shepherd, which assigns a reward score to each step of math problem solutions. The training of Math-shepherd is achieved using automatically constructed process-wise supervision data, breaking the bottlene…

Cited by 242SourcePDFScholar
2024

Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

ICML 2024poster

Large Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks. Many efforts are being made to transfer these attributes to video modality, which are termed Video-LLMs. However, existing Video-LLMs can only capture the coarse-grained semantics and are…

2024

ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial Optimization

NeurIPS 2024poster

Watermarking generative content serves as a vital tool for authentication, ownership protection, and mitigation of potential misuse. Existing watermarking methods face the challenge of balancing robustness and concealment. They empirically inject a watermark that is both invisible and robust and pas…

2024

Toward Real Ultra Image Segmentation: Leveraging Surrounding Context to Cultivate General Segmentation Model

NeurIPS 2024poster

Existing ultra image segmentation methods suffer from two major challenges, namely the scalability issue (i.e. they lack the stability and generality of standard segmentation models, as they are tailored to specific datasets), and the architectural issue (i.e. they are incompatible with real-world u…

Cited by 1SourcePDFScholar
2024

Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking

NeurIPS 2024poster

Respiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Genera…

2023

BEATs: Audio Pre-Training with Acoustic Tokenizers

ICML 2023oral

We introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is tra…

2023

Boundary Guided Learning-Free Semantic Control with Diffusion Models

NeurIPS 2023poster

Applying pre-trained generative denoising diffusion models (DDMs) for downstream tasks such as image semantic editing usually requires either fine-tuning DDMs or learning auxiliary editing networks in the existing literature. In this work, we present our BoundaryDiffusion method for efficient, effec…

2023

DVIS: Decoupled Video Instance Segmentation Framework

ICCV 2023poster

Video instance segmentation (VIS) is a critical task with diverse applications, including autonomous driving and video editing. Existing methods often underperform on complex and long videos in real world, primarily due to two factors. Firstly, offline methods are limited by the tightly-coupled mode…

Cited by 59PDFcodeScholar
2023

Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation

ICLR 2023poster

Diffusion probabilistic models (DPMs) have become a popular approach to conditional generation, due to their promising results and support for cross-modal synthesis. A key desideratum in conditional synthesis is to achieve high correspondence between the conditioning input and generated output. Most…

2023

Good Is Bad: Causality Inspired Cloth-Debiasing for Cloth-Changing Person Re-Identification

CVPR 2023poster

Entangled representation of clothing and identity (ID)-intrinsic clues are potentially concomitant in conventional person Re-IDentification (ReID). Nevertheless, eliminating the negative impact of clothing on ID remains challenging due to the lack of theory and the difficulty of isolating the exact…

2023

Grounded Image Text Matching with Mismatched Relation Reasoning

ICCV 2023poster

This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image,…

Cited by 8PDFcodeScholar
2023

Learning To Segment Every Referring Object Point by Point

CVPR 2023poster

Referring Expression Segmentation (RES) can facilitate pixel-level semantic alignment between vision and language. Most of the existing RES approaches require massive pixel-level annotations, which are expensive and exhaustive. In this paper, we propose a new partially supervised training paradigm f…

2023

Local-Global Progressive U-Transformers for Accurate Hepatic and Portal Veins Segmentation in Abdominal MR Images

ICASSP 2023accepted

Segmentation of hepatic and portal veins in abdominal magnetic resonance images plays an essential role in surgical planning of liver tumor ablation and resection. Accurately extracting these blood vessels is a challenging task due to the complex vessel structures with high noise and irregular vesse…

Cited by 0SourceScholar
2023

LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer

ICASSP 2023accepted

Traditional automatic speech recognition (ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is more practical in real scenarios. Simply attending longer transcription history for a vanilla neural transducer model shows…

Cited by 0SourceScholar
2023

Magneto: A Foundation Transformer

ICML 2023poster

A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name ''Transformers'', the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers.…

Cited by 12SourcePDFScholar
2023

RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments

NeurIPS 2023poster

Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctively seek out a suitable option such as a "bed" or a "sofa" that can fulfill our needs. Previous work in this area is limi…

Cited by 14SourcePDFScholar
2023

Real-Time Speech Interruption Analysis: from Cloud to Client Deployment

ICASSP 2023accepted

Meetings are an essential form of communication for all types of organizations, and remote collaboration systems have been much more widely used since the COVID-19 pandemic. One major issue with remote meetings is that it is challenging for remote participants to interrupt and speak. We have recentl…

Cited by 0SourceScholar
2023

Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language Perspective

NeurIPS 2023poster

We focus on the weakly-supervised audio-visual video parsing task (AVVP), which aims to identify and locate all the events in audio/visual modalities. Previous works only concentrate on video-level overall label denoising across modalities, but overlook the segment-level label noise, where adjacent…

2023

Speech Separation with Large-Scale Self-Supervised Learning

ICASSP 2023accepted

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the exploration of the SSL-based SS by massively scaling up both the pre-training data (more than 300K hours) and fine-tuning…

Cited by 0SourceScholar
2023

Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open World

ICCV 2023poster

Scene Graph Generation (SGG) aims to extract <subject, predicate, object> relationships in images for vision understanding. Although recent works have made steady progress on SGG, they still suffer long-tail distribution that tail-predicates are more costly to train and hard to distinguish due to a…

Cited by 35PDFcodeScholar
2022

Enabling Detailed Action Recognition Evaluation Through Video Dataset Augmentation

NeurIPS 2022accept

It is well-known in the video understanding community that human action recognition models suffer from background bias, i.e., over-relying on scene cues in making their predictions. However, it is difficult to quantify this effect using existing evaluation frameworks. We introduce the Human-centric…

Cited by 14SourcePDFScholar
2022

Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer Supervision

ICASSP 2022accepted

Recently, pioneer work finds that self-supervised pre-training methods can improve multiple downstream speech tasks, because the model utilizes bottom layers to learn speaker-related information and top layers to encode content-related information. Since the network capacity is limited, we believe t…

Cited by 0SourceScholar
2022

Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification

ICASSP 2022accepted

The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we explore the limits of speech representations learned by different…

Cited by 0SourceScholar
2022

Large-Scale Video Panoptic Segmentation in the Wild: A Benchmark

CVPR 2022poster

In this paper, we present a new large-scale dataset for the video panoptic segmentation task, which aims to assign semantic classes and track identities to all pixels in a video. As the ground truth for this task is difficult to annotate, previous datasets for video panoptic segmentation are limited…

Cited by 100PDFcodeScholar
2022

Learning To Learn by Jointly Optimizing Neural Architecture and Weights

CVPR 2022poster

Meta-learning enables models to adapt to new environments rapidly with a few training examples. Current gradient-based meta-learning methods concentrate on finding good initialization (meta-weights) for learners but ignore the impact of neural architectures. In this paper, we aim to obtain better me…

Cited by 13PDFScholar
2022

Quantized GAN for Complex Music Generation from Dance Videos

ECCV 2022poster

"We present Dance2Music-GAN (D2M-GAN), a novel adversarial multi-modal framework that generates complex musical samples conditioned on dance videos. Our proposed framework takes dance video frames and human body motions as input, and learns to generate music samples that plausibly accompany the corr…

2022

SiRi: A Simple Selective Retraining Mechanism for Transformer-Based Visual Grounding

ECCV 2022poster

"In this paper, we investigate how to achieve better referring visual grounding with modern vision-language transformers, and propose a simple yet powerful Selective Retraining (SiRi) mechanism. Particularly, SiRi conveys a significant principle to the research of visual grounding, i.e, a better ini…

2022

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

ACL 2022long

Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for self-supervised speech/text representation learning. The SpeechT5 framework consists…

2022

Two-Stream Network for Sign Language Recognition and Translation

NeurIPS 2022accept

Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos into hidden representations. RGB videos, however, are raw signals with substanti…

2022

Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training

ICASSP 2022accepted

Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attemp…

Cited by 0SourceScholar
2022

Wav2vec-Switch: Contrastive Learning from Original-Noisy Speech Pairs for Robust Speech Recognition

ICASSP 2022accepted

The goal of self-supervised learning (SSL) for automatic speech recognition (ASR) is to learn good speech representations from a large amount of unlabeled speech for the downstream ASR task. However, most SSL frameworks do not consider noise robustness which is crucial for real-world applications. I…

Cited by 66SourceScholar
2021

Detecting Speaker Personas from Conversational Texts

EMNLP 2021main

Personas are useful for dialogue response prediction. However, the personas used in current studies are pre-defined and hard to obtain before a conversation. To tackle this issue, we study a new task, named Speaker Persona Detection (SPD), which aims to detect speaker personas based on the plain con…

2021

Developing Real-Time Streaming Transformer Transducer for Speech Recognition on Large-Scale Dataset

ICASSP 2021accepted

Recently, Transformer based end-to-end models have achieved great success in many areas including speech recognition. However, compared to LSTM models, the heavy computational cost of the Transformer during inference is a key issue to prevent their applications. In this work, we explored the potenti…

Cited by 0SourceScholar
2021

Don't Shoot Butterfly with Rifles: Multi-Channel Continuous Speech Separation with Early Exit Transformer

ICASSP 2021accepted

With its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been successfully applied to speech separation recently. However, multi-channel speech separation sometimes does not necessa…

Cited by 0SourceScholar
2021

Knowledge Enhanced Fine-Tuning for Better Handling Unseen Entities in Dialogue Generation

EMNLP 2021main

Although pre-training models have achieved great success in dialogue generation, their performance drops dramatically when the input contains an entity that does not appear in pre-training and fine-tuning datasets (unseen entity). To address this issue, existing methods leverage an external knowledg…

2021

Learning Audio-Visual Correlations From Variational Cross-Modal Generation

ICASSP 2021accepted

People can easily imagine the potential sound while seeing an event. This natural synchronization between audio and visual signals reveals their intrinsic correlations. To this end, we propose to learn the audio-visual correlations from the perspective of cross-modal generation in a self-supervised…

Cited by 0SourceScholar
2021

Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020

ICASSP 2021accepted

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker reco…

Cited by 0SourceScholar
2021

UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data

ICML 2021spotlight

In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both labeled and unlabeled data, in which supervised phonetic CTC learning and phonetically-aware contrastive self-supervised learning are conducted in a multi-task learning manner. The re…

2021

VSPW: A Large-scale Dataset for Video Scene Parsing in the Wild

CVPR 2021poster

In this paper, we present a new dataset with the target of advancing the scene parsing task from images to videos. Our dataset aims to perform Video Scene Parsing in the Wild (VSPW), which covers a wide range of real-world scenarios and categories. To be specific, our VSPW is featured from the follo…

Cited by 136PDFScholar
2020

Unsupervised Person Re-Identification via Softened Similarity Learning

CVPR 2020poster

Person re-identification (re-ID) is an important topic in computer vision. This paper studies the unsupervised setting of re-ID, which does not require any labeled information and thus is freely deployed to new scenarios. There are very few studies under this setting, and one of the best approach ti…

Cited by 346PDFScholar
2019

Auto-ReID: Searching for a Part-Aware ConvNet for Person Re-Identification

ICCV 2019poster

Prevailing deep convolutional neural networks (CNNs) for person re-IDentification (reID) are usually built upon ResNet or VGG backbones, which were originally designed for classification. Because reID is different from classification, the architecture should be modified accordingly. We propose to au…

Cited by 315PDFScholar
2019

Pose-Guided Feature Alignment for Occluded Person Re-Identification

ICCV 2019poster

Persons are often occluded by various obstacles in person retrieval scenarios. Previous person re-identification (re-id) methods, either overlook this issue or resolve it based on an extreme assumption. To alleviate the occlusion problem, we propose to detect the occluded regions, and explicitly exc…

Cited by 696PDFcodeScholar
2018

Exploit the Unknown Gradually: One-Shot Video-Based Person Re-Identification by Stepwise Learning

CVPR 2018poster

We focus on the one-shot learning for video-based person re-Identification (re-ID). Unlabeled tracklets for the person re-ID tasks can be easily obtained by pre-processing, such as pedestrian detection and tracking. In this paper, we propose an approach to exploiting unlabeled tracklets by gradually…

Cited by 456SourcePDFScholar