← Search

jiawei liu

72 accepted papers

2026

ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking

ICML 2026poster

Reinforcement learning (RL) has advanced LLM agents on verifiable tasks but remains challenging for open-ended tasks with vast solution spaces (e.g., complex travel planning). Lacking objective ground truth, current RL algorithms rely on reward models assigning scalar scores to individual responses.…

Cited by 0SourceScholar
2026

Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

ICML 2026poster

Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as state-of-the-art generative models not only optimize for the lip synchronization but also si…

Cited by 0SourceScholar
2026

Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation

CVPR 2026

We propose Decoupled Residual Denoising Diffusion models (DRDD) for unified and data-efficient image-to-image (I2I) translation. While diffusion models have advanced I2I translation in terms of quality and diversity, we uncover a previously under-explored property in diffusion models. Crucially, bey

Cited by 0SourcecodeScholar
2026

Disturbance-Aware Sliding Mode Control for Automated Optical Tweezers Using Nonlinear Observation

RA-L 2026

This paper presents a control-oriented framework for automated optical tweezers (AOT), explicitly accounting for fluid dynamics and common real-world disturbances. AOT operation is strongly affected by unmeasurable flow-induced forces and stochastic perturbations under noisy vision feedback, which p

Cited by 0SourceScholar
2026

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

ICML 2026poster

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and audio-driven video animation (RA2V) as isolated objectives. F…

Cited by 0SourceScholar
2026

HUMOF: Human Motion Forecasting in Interactive Social Scenes

ICLR 2026poster

Complex dynamic scenes present significant challenges for predicting human behavior due to the abundance of interaction information, such as human-human and human-environment interactions. These factors complicate the analysis and understanding of human behavior, thereby increasing the uncertainty i…

Cited by 0SourceScholar
2026

Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

AAAI 2026technical

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty

Cited by 0SourcePDFScholar
2026

Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection

CVPR 2026

Open-Vocabulary Human-Object Interaction (OV-HOI) detection aims to recognize novel HOI categories beyond the training set. Existing OV-HOI detection approaches typically leverage CLIP to extract global visual representations and perform cross-attention between learnable queries and global features

Cited by 0SourceScholar
2026

Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics

ICML 2026poster

Agentic systems have recently become the dominant paradigm for formal theorem proving, achieving strong performance by coordinating multiple models and tools. However, existing approaches often rely on task-specific pipelines and trained formal provers, limiting their flexibility and reproducibility…

Cited by 0SourceScholar
2026

OCT-DeformNet: Optical Coherence Tomography-Guided Biological Tissue Shape Prediction for Robot Palpation in Microsurgery

ICRA 2026poster

In medical robotics, biological shape deformation resulting from arbitrary tool-tissue interaction commonly occurs and motivates the need in microsurgery to predict the new geometry of tissue structures. However, handling deformation is challenging due to the lack of a general prediction model for v…

Cited by 0Scholar
2026

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

CVPR 2026

Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals--without sacrificing interpretability and traceability--remai

Cited by 0SourceScholar
2026

PSP: Prompt-Guided Self-Training Sampling Policy for Active Prompt Learning

ICLR 2026poster

Active Prompt Learning (APL) using vision-language models (\textit{e.g.}, CLIP) has attracted considerable attention for mitigating the dependence on fully labeled dataset in downstream task adaptation. However, existing methods fail to explicitly leverage prompt to guide sample selection, resulting…

Cited by 0SourcecodeScholar
2026

R2G: A Multi-View Circuit Graph Benchmark Suite from RTL to GDSII

CVPR 2026

Graph neural networks (GNNs) are increasingly applied to physical design tasks such as congestion prediction and wirelength estimation, yet progress is hindered by inconsistent circuit representations and the absence of controlled evaluation protocols. We present R2G (RTL-to-GDSII), a multi-view cir

Cited by 0SourcecodeScholar
2026

Topology Matters in RTL Circuit Representation Learning

ICLR 2026poster

Representation learning for register transfer level (RTL) circuits is fundamental to enabling accurate performance, power, and area (PPA) prediction, efficient circuit generation, and retrieval in automated chip design. Unlike general programming languages, RTL is inherently a structured dataflow gr…

Cited by 0SourcecodeScholar
2026

Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation Alignment

AAAI 2026technical

We present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM

Cited by 0SourcePDFScholar
2025

AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion

CVPR 2025poster

The task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge. However, asynchronous auto-regressive models often suffer from inc…

2025

BearLLM: A Prior Knowledge-Enhanced Bearing Health Management Framework with Unified Vibration Signal Representation

AAAI 2025technical

We propose a bearing health management framework leveraging large language models (BearLLM), a novel multimodal model that unifies multiple bearing-related tasks by processing user prompts and vibration signals. Specifically, we introduce a prior knowledge-enhanced unified vibration signal represent…

2025

Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

CVPR 2025poster

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this p…

2025

BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

ICLR 2025oral

Task automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks range from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human develop…

Cited by 609SourcePDFScholar
2025

Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation

EMNLP 2025

Large Language Models (LLMs) often exhibit hallucinations, generating factually incorrect or semantically irrelevant content in response to prompts. Chain-of-Thought (CoT) prompting can mitigate hallucinations by encouraging step-by-step reasoning, but its impact on hallucination detection remains u

2025

Dual-Arm Teleoperated Robotic Microsurgery System with Live Volumetric OCT Image Feedback

IROS 2025

In microsurgery, surgeons frequently encounter challenges due to the need for exceptional precision and dexterity, the lack of depth perception for micro-scale surgical maneuvers, and the inevitable effects of fatigue and hand tremor. In surgical robotics, conventional intraoperative perception syst

Cited by 0SourceScholar
2025

Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

NeurIPS 2025poster

The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive conte…

Cited by 0SourcecodeScholar
2025

HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI Detection

AAAI 2025technical

Human-object interaction (HOI) detection aims to detect the spatial positions of human-object pairs and recognize their interactions. Existing single-branch, two-branch, and three-branch methods are challenging to make an appropriate trade-off on efficiency, multi-task decoupling, and collaborative…

Cited by 0SourcePDFScholar
2025

Harnessing and Evaluating the Intrinsic Extrapolation Ability of Large Language Models for Vehicle Trajectory Prediction

NAACL 2025long

Emergent abilities of large language models (LLMs) have significantly advanced their application in autonomous vehicle (AV) research. Safe integration of LLMs into vehicles, however, necessitates their thorough understanding of dynamic traffic environments. Towards this end, this study introduces a…

Cited by 0SourcePDFScholar
2025

Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptation

CVPR 2025poster

Test-time adaptation using vision-language models (such as CLIP) to quickly adjust to distributional shifts of downstream tasks has shown great potential. Despite significant progress, existing methods are still limited to single-task test-time adaptation scenarios and have not effectively explored…

Cited by 0SourcePDFScholar
2025

I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength

ICLR 2025poster

Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However, existing camera control methods still suffer from several limita…

2025

I2VControl: Disentangled and Unified Video Motion Synthesis Control

ICCV 2025poster

Motion controllability is crucial in video synthesis. However, most previous methods are limited to single control types, and combining them often results in logical conflicts. In this paper, we propose a disentangled and unified framework, namely I2VControl, to overcome the logical conflicts. We re…

2025

Interweaving Memories of a Siamese Large Language Model

AAAI 2025technical

Parameter-efficient fine-tuning (PEFT) methods optimize large language models (LLMs) by modifying or introducing a small number of parameters to enhance alignment with downstream tasks. However, they can result in catastrophic forgetting, where LLMs prioritize new knowledge at the expense of compreh…

2025

Learnable Frequency Decomposition for Image Forgery Detection and Localization

IJCAI 2025

Concern for image authenticity spurs research in image forgery detection and localization (IFDL). Most deep learning-based methods focus primarily on spatial domain modeling and have not fully explored frequency domain strategies. In this paper, we observe and analyze the frequency characteristic ch

Cited by 0SourcePDFScholar
2025

Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

CVPR 2025poster

Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask^2DiT,…

2025

Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment

ICCV 2025poster

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-con…

Cited by 0SourcePDFScholar
2025

PurpCode: Reasoning for Safer Code Generation

NeurIPS 2025poster

We introduce PurpCode, the first post-training recipe for training safe code reasoning models towards generating secure code and defending against malicious cyberactivities. PurpCode trains a reasoning model in two stages: (i) Rule Learning, which explicitly teaches the model to reference cybersafet…

Cited by 0SourceScholar
2025

Reducing AUV Energy Consumption Through Dynamic Sensor Directions Switching via Deep Reinforcement Learning

AAAI 2025technical

Autonomous underwater vehicle (AUV) is crucial for marine applications such as ocean data collection, pollution monitoring, and navigation. However, their limited energy resources constrain their operational duration, posing a significant challenge for long-term operations. Due to the complex and un…

Cited by 0SourcePDFScholar
2025

Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level Guidance

NeurIPS 2025poster

The dynamic nature of real-world information demands efficient knowledge editing in multimodal large language models (MLLMs) to ensure continuous knowledge updates. However, existing methods often struggle with precise matching in large-scale knowledge retrieval and lack multi-level guidance for coo…

Cited by 0SourceScholar
2024

Bevel-Tip Needle Deflection Modeling, Simulation, and Validation in Multi-Layer Tissues

ICRA 2024poster

Percutaneous needle insertions are commonly performed for diagnostic and therapeutic purposes as an effective alternative to more invasive surgical procedures. However, the outcome of needle-based approaches relies heavily on the accuracy of needle placement, which remains a challenge even with robo…

Cited by 3SourceScholar
2024

DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations

CVPR 2024highlight

The diffusion-based text-to-image model harbors immense potential in transferring reference style. However current encoder-based approaches significantly impair the text controllability of text-to-image models while transferring styles. In this paper we introduce DEADiff to address this issue using…

2024

Enhance Robustness of Language Models against Variation Attack through Graph Integration

COLING 2024main

The widespread use of pre-trained language models (PLMs) in natural language processing (NLP) has greatly improved performance outcomes. However, these models’ vulnerability to adversarial attacks (e.g., camouflaged hints from drug dealers), particularly in the Chinese language with its rich charact…

2024

Fall Prediction by a Spatio-Temporal Multi-Channel Causal Model from Wearable Sensors Data

ICASSP 2024accepted

Predicting human falls from wearable devices is a complex task due to the inherent diversity and causality of multivariate physical changes, where each instance exhibits a unique style of motion events and their spatio-temporal causal dependencies. Consequently, we propose a multichannel causal mode…

Cited by 0SourceScholar
2024

From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications

ACL 2024findings

Evaluating large language models (LLMs) is fundamental, particularly in the context of practical applications. Conventional evaluation methods, typically designed primarily for LLM development, yield numerical scores that ignore the user experience. Therefore, our study shifts the focus from model-c…

Cited by 0SourcePDFScholar
2024

Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes

NeurIPS 2024poster

In recent years, DeepFake technology has achieved unprecedented success in high-quality video synthesis, but these methods also pose potential and severe security threats to humanity. DeepFake can be bifurcated into entertainment applications like face swapping and illicit uses such as lip-syncing f…

2024

Magicoder: Empowering Code Generation with OSS-Instruct

ICML 2024poster

We introduce Magicoder, a series of fully open-source (code, weights, and data) Large Language Models (LLMs) for code that significantly closes the gap with top code models while having no more than 7B parameters. Magicoder models are trained on 75K synthetic instruction data using **OSS-Instruct**,…

2024

Natural Language-centered Inference Network for Multi-modal Fake News Detection

IJCAI 2024poster

The proliferation of fake news with image and text in the internet has triggered widespread concern. Existing research has made important contributions in cross-modal information interaction and fusion, but fails to fundamentally address the modality gap among news image, text, and news-related exte…

Cited by 4SourcePDFScholar
2024

Noise-assisted Prompt Learning for Image Forgery Detection and Localization

ECCV 2024poster

"We present CLIP-IFDL, a novel image forgery detection and localization (IFDL) model that harnesses the power of Contrastive Language Image Pre-Training (CLIP). However, directly incorporating CLIP in forgery detection poses challenges, given its lack of specific prompts and forgery consciousness. T…

Cited by 4SourcePDFScholar
2024

Predicting Fall Events by a Spatio-Temporal Topological Network with Multiple Wearable Sensors

ICASSP 2024accepted

A key challenge in sensor-based fall prediction is the fact that a fall event can often occur in various configurations of fall poses together with their own spatio-temporal dependencies. This leads us to define a spatio-temporal model to explicitly characterize these internal configurations of pose…

Cited by 0SourceScholar
2024

Residual Denoising Diffusion Models

CVPR 2024poster

We propose residual denoising diffusion models (RDDM) a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the denoising-based diffusion models initially uninterpretable for…

2024

Self-Calibrating Vicinal Risk Minimisation for Model Calibration

CVPR 2024poster

Model calibration measuring the alignment between the prediction accuracy and model confidence is an important metric reflecting model trustworthiness. Existing dense binary classification methods without proper regularisation of model confidence are prone to being over-confident. To calibrate Deep…

2024

SelfCodeAlign: Self-Alignment for Code Generation

NeurIPS 2024poster

Instruction tuning is a supervised fine-tuning approach that significantly improves the ability of large language models (LLMs) to follow human instructions. For programming tasks, most models are finetuned with costly human-annotated instruction-response pairs or those generated by large, proprieta…

2024

Tracking Tumors under Deformation from Partial Point Clouds using Occupancy Networks

IROS 2024poster

To track tumors during surgery, information from preoperative CT scans is used to determine their position. However, as the surgeon operates, the tumor may be deformed which presents a major hurdle for accurately resecting the tumor, and can lead to surgical inaccuracy, increased operation time, and…

Cited by 2SourceScholar
2024

View From Above: Orthogonal-View aware Cross-view Localization

CVPR 2024poster

This paper presents a novel aerial-to-ground feature aggregation strategy tailored for the task of cross-view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database often through establishing planar homog…

Cited by 5SourcePDFScholar
2024

Visual Text Generation in the Wild

ECCV 2024poster

"Recently, with the rapid advancements of generative models, the field of visual text generation has witnessed significant progress. However, it is still challenging to render high-quality text images in real-world scenarios, as three critical criteria should be satisfied: (1) Fidelity: the generate…

2024

XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts

ACL 2024long

We introduce XFT, a simple yet powerful training scheme, by simply merging upcycled Mixture-of-Experts (MoE) to unleash the performance limit of instruction-tuned code Large Language Models (LLMs). While vanilla sparse upcycling fails to improve instruction tuning, XFT introduces a shared expert mec…

2023

A Bio-Inspired Simultaneous Surface and Underwater Risk Assessment Method Based on Stereo Vision for USVs in Nearshore Clean Waters

RA-L 2023

Ensuring safety of unmanned surface vehicles (USVs) during nearshore field operations is challenging as they face unknown risks of collision with surface and underwater obstacles. To address this problem, we propose a bio-inspired collision risk-assessment method for ensuring safe USVs operation in

Cited by 18SourceScholar
2023

A Speaker Turn-Aware Multi-Task Adversarial Network for Joint User Satisfaction Estimation and Sentiment Analysis

AAAI 2023technical

User Satisfaction Estimation is an important task and increasingly being applied in goal-oriented dialogue systems to estimate whether the user is satisfied with the service. It is observed that whether the user’s needs are met often triggers various sentiments, which can be pertinent to the success…

Cited by 10SourcePDFScholar
2023

Development and Evaluation of a Single-arm Robotic System for Autonomous Suturing

IROS 2023poster

This article introduces a novel suture managing device (SMD) and new suture management controller to enable single-arm suture management during autonomous suturing with the Smart Tissue Autonomous Robot (STAR). The primary function of the SMD is to tension and manage the suture thread, a task that w…

Cited by 0SourceScholar
2023

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

ICCV 2023poster

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance a…

Cited by 5PDFcodeScholar
2023

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

NeurIPS 2023poster

Program synthesis has been long studied with recent approaches focused on directly using the power of Large Language Models (LLMs) to generate code. Programming benchmarks, with curated synthesis problems and test-cases, are used to measure the performance of various LLMs on code synthesis. However,…

2023

Model Calibration in Dense Classification with Adaptive Label Perturbation

ICCV 2023poster

For safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of correctness for subsequent decision-making. Existing dense binary classification models are prone to being over-confident. To…

Cited by 4PDFcodeScholar
2023

P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds

ICCV 2023poster

Point cloud completion aims to recover the complete shape based on a partial observation. Existing methods require either complete point clouds or multiple partial observations of the same object for learning. In contrast to previous approaches, we present Partial2Complete (P2C), the first self-supe…

Cited by 29PDFcodeScholar
2023

PD-Quant: Post-Training Quantization Based on Prediction Difference Metric

CVPR 2023poster

Post-training quantization (PTQ) is a neural network compression technique that converts a full-precision model into a quantized model using lower-precision data types. Although it can help reduce the size and computational cost of deep neural networks, it can also introduce quantization noise and r…

2023

Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-Trained Vision-Language Models

ICCV 2023poster

Prompt tuning and adapter tuning have shown great potential in transferring pre-trained vision-language models (VLMs) to various downstream tasks. In this work, we design a new type of tuning method, termed as regularized mask tuning, which masks the network parameters through a learnable selection.…

Cited by 12PDFScholar
2023

WL-MSR: Watch and Listen for Multimodal Subtitle Recognition

ICASSP 2023accepted

Video subtitles could be defined as the combination of visualized subtitles in frames and textual content recognized from speech, which play a significant role in video understanding for both humans and machines. In this paper, we propose a novel Watch and Listen for Multimodal Subtitle Recognition…

Cited by 0SourceScholar
2022

Debiased Batch Normalization via Gaussian Process for Generalizable Person Re-identification

AAAI 2022technical

Generalizable person re-identification aims to learn a model with only several labeled source domains that can perform well on unseen domains. Without access to the unseen domain, the feature statistics of the batch normalization (BN) layer learned from a limited number of source domains is doubtles…

Cited by 36SourcePDFScholar
2022

Modality-Adaptive Mixup and Invariant Decomposition for RGB-Infrared Person Re-identification

AAAI 2022technical

RGB-infrared person re-identification is an emerging cross-modality re-identification task, which is very challenging due to significant modality discrepancy between RGB and infrared images. In this work, we propose a novel modality-adaptive mixup and invariant decomposition (MID) approach for RGB-i…

Cited by 105SourcePDFScholar
2022

Temporal Complementarity-Guided Reinforcement Learning for Image-to-Video Person Re-Identification

CVPR 2022poster

Image-to-video person re-identification aims to retrieve the same pedestrian as the image-based query from a video-based gallery set. Existing methods treat it as a cross-modality retrieval task and learn the common latent embeddings from image and video modalities, which are both less effective and…

Cited by 17PDFScholar
2021

A Role-Selected Sharing Network for Joint Machine-Human Chatting Handoff and Service Satisfaction Analysis

EMNLP 2021main

Chatbot is increasingly thriving in different domains, however, because of unexpected discourse complexity and training data sparseness, its potential distrust hatches vital apprehension. Recently, Machine-Human Chatting Handoff (MHCH), predicting chatbot failure and enabling human-algorithm collabo…

2021

Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in Videos

CVPR 2021poster

Video-based person re-identification aims to match pedestrians from video sequences across non-overlapping camera views. The key factor for video person re-identification is to effectively exploit both spatial and temporal clues from video sequences. In this work, we propose a novel Spatial-Temporal…

Cited by 81PDFScholar
2021

Time to Transfer: Predicting and Evaluating Machine-Human Chatting Handoff

AAAI 2021technical

Is chatbot able to completely replace the human agent? The short answer could be – ``it depends...''. For some challenging cases, e.g., dialogue's topical spectrum spreads beyond the training corpus coverage, the chatbot may malfunction and return unsatisfied utterances. This problem can be addresse…

2020

Co-Saliency Spatio-Temporal Interaction Network for Person Re-Identification in Videos

IJCAI 2020poster

Person re-identification aims at identifying a certain pedestrian across non-overlapping camera networks. Video-based person re-identification approaches have gained significant attention recently, expanding image-based approaches by learning features from multiple frames. In this work, we propose a…

Cited by 0SourcePDFScholar
2020

Decorrelated Clustering with Data Selection Bias

IJCAI 2020poster

Most of existing clustering algorithms are proposed without considering the selection bias in data. In many real applications, however, one cannot guarantee the data is unbiased. Selection bias might bring the unexpected correlation between features and ignoring those unexpected correlations will hu…

2020

Multi-Scale Spatial-Temporal Integration Convolutional Tube for Human Action Recognition

IJCAI 2020poster

Applying multi-scale representations leads to consistent performance improvements on a wide range of image recognition tasks. However, with the addition of the temporal dimension in video domain, directly obtaining layer-wise multi-scale spatial-temporal features will add a lot extra computational c…

Cited by 0SourcePDFScholar
2019

Adaptive Transfer Network for Cross-Domain Person Re-Identification

CVPR 2019poster

Recent deep learning based person re-identification approaches have steadily improved the performance for benchmarks, however they often fail to generalize well from one domain to another. In this work, we propose a novel adaptive transfer network (ATNet) for effective cross-domain person re-identif…

Cited by 350PDFScholar