← Search

Li Liu

149 accepted papers

2026

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

ICML 2026poster

REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuri…

Cited by 0SourceScholar
2026

AUDIOGENIE-REASONER: A TRAINING-FREE MULTI-AGENT FRAMEWORK FOR COARSE-TO-FINE AUDIO DEEP REASONING

ICASSP 2026oral

Audio deep reasoning is a challenging task that requires expert-level perception, multi-step logical inference, and the integration of contextual knowledge. However, existing models suffer from a gap between audio perception and reasoning abilities due to the lack of training data with explicit reas…

Cited by 0SourcePDFScholar
2026

Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models

ICML 2026poster

The integration of audio modality into Large Audio Language Models (LALMs) significantly expands their attack surface. Existing jailbreak paradigms predominantly treat audio as a carrier for malicious payloads, relying on semantic optimization, acoustic parameter control, or additive perturbation to…

Cited by 1SourceScholar
2026

Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards

AAAI 2026technical

Recently, Automatic Speech Recognition (ASR) systems (e.g., Whisper) have achieved remarkable accuracy improvements but remain highly sensitive to real-world unseen data (data with large distribution shifts), including noisy environments and diverse accents. To address this issue, test-time adaptati

Cited by 0SourcePDFScholar
2026

Breaking the Latency Barrier: Synergistic Perception and Control for High-Frequency 3D Ultrasound Servoing

ICRA 2026poster

Tracking moving anatomical targets with robotic ultrasound is particularly challenging when the target motion is both fast and large in scale, as the end-to-end latency of existing systems prevents the perception–control loop from closing fast enough. In this paper, we argue that overcoming this lim…

2026

Cueing Without Gapping: Cuer-Independent Cued Speech Recognition Powered by Cross-Cuer Invariant Modeling

AAAI 2026technical

Automatic Cued Speech Recognition (ACSR) is a vital communication system designed to enhance spoken language accessibility for the hearing-impaired by combining lip movements and hand gestures to encode phonemes. Despite its effectiveness, current ACSR methods face significant challenges, including

Cited by 0SourcePDFScholar
2026

HSA-Net: Hierarchical and Structure-Aware Framework for Efficient and Scalable Molecular Language Modeling

AAAI 2026technical

Molecular representation learning, a cornerstone for downstream tasks like molecular captioning and molecular property prediction, heavily relies on Graph Neural Networks (GNN). However, GNN suffers from the over-smoothing problem, where node-level features collapse in deep GNN layers. While existin

Cited by 0SourcePDFScholar
2026

LEND A HAND: SEMI TRAINING-FREE CUED SPEECH RECOGNITION VIA MLLM-DRIVEN HAND MODELING FOR BARRIER-FREE COMMUNICATION

ICASSP 2026poster

Cued Speech (CS) is an innovative visual communication system that integrates lip-reading with hand coding, designed to enhance effective communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) refers to the AI-driven process of automatically recognizing hand gestures…

Cited by 0SourcePDFScholar
2026

Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal Alignment

AAAI 2026technical

Cross-modal alignment is a promising yet challenging task in multimodal learning. Existing methods typically assess it by measuring the cross-modal semantic similarity from both global and local perspectives. However, these methods often neglect their potential interdependence. Specifically, global

Cited by 0SourcePDFScholar
2026

ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object Recognition

CVPR 2026

Recent advances in Remote Sensing Foundation Models (RSFMs) have demonstrated considerable potential for Earth Observation (EO) tasks. While adopting natural image foundation models (e.g., DINO) provides a data-efficient strategy for building RSFMs, their strong generalization capability does not fu

Cited by 0SourcecodeScholar
2026

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

ICLR 2026poster

Deep learning-based respiratory auscultation is currently hindered by two fundamental challenges: (i) inherent information loss, as converting signals into spectrograms discards transient acoustic events and clinical context; (ii) limited data availability, exacerbated by severe class imbalance. To…

Cited by 0SourcecodeScholar
2026

Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object Detection

CVPR 2026

Recent advancements in remote sensing object detection have predominantly focused on oriented bounding box design and small object feature enhancement, while often overlooking the intrinsic geometric properties of remote sensing images, such as rotation invariance and structural symmetry. Many aeria

Cited by 0SourcecodeScholar
2026

SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal Steering

ICML 2026poster

Large Audio–Language Models (LALMs) are becoming essential as a powerful multimodal backbone for real-world applications. However, recent studies show that audio inputs can more easily elicit harmful responses than text, exposing new risks toward deployment. While safety alignment has made initial a…

Cited by 0SourceScholar
2026

Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection

AAAI 2026technical

In this paper, we show that current approaches using large square kernels or transformer-based global modeling aggregate contextual information uniformly across spatial dimensions, leading to feature dilution and localization errors for elongated targets. To mitigate this issue, we propose Strip R-C

Cited by 0SourcePDFScholar
2026

TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability

AAAI 2026technical

Achieving zero-shot adversarial robustness without sacrificing generalization remains challenging for foundation models such as CLIP, especially under large adversarial perturbations. Through empirical analyses, we identify three critical yet overlooked issues: (1) Logit margins exhibit a stable off

Cited by 0SourcePDFScholar
2026

UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation

AAAI 2026technical

Cued Speech (CS) enhances lipreading via hand coding, offering visual phonemic cues that support precise speech perception for the hearing-impaired. The task of CS Video-to-Speech generation (CSV2S) aims to convert CS videos into intelligible speech signals. Most existing research focuses on CS Reco

Cited by 0SourcePDFScholar
2025

Activation Gradient based Poisoned Sample Detection Against Backdoor Attacks

ICLR 2025poster

This work studies the task of poisoned sample detection for defending against data poisoning based backdoor attacks. Its core challenge is finding a generalizable and discriminative metric to distinguish between clean and various types of poisoned samples (e.g., various triggers, various poisoning r…

Cited by 4SourcePDFScholar
2025

AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving

EMNLP 2025

Vision-Language Models (VLMs) show promise for autonomous driving, yet their struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. To overcome this, we introduce AgentThink , a pioneering unified framewor

2025

BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model

NeurIPS 2025poster

Backdoor learning is a critical research topic for understanding the vulnerabilities of deep neural networks. While the diffusion model (DM) has been broadly deployed in public over the past few years, the understanding of its backdoor vulnerability is still in its infancy compared to the extensive…

Cited by 0SourcecodeScholar
2025

CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter

ACL 2025long

Speculative decoding is a powerful technique that accelerates Large Language Model (LLM) inference by leveraging a lightweight speculative draft model. However, existing designs suffers in performance due to misalignment between training and inference. Recent methods have tried to solve this issue b…

2025

Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition

AAAI 2025technical

Wearable Human Activity Recognition (WHAR) is a prominent research area within ubiquitous computing. Multi-sensor synchronous measurement has proven to be more effective for WHAR than using a single sensor. However, existing WHAR methods use shared convolutional kernels for indiscriminate temporal f…

2025

Dual-Process Watermarked Diffusion: Integrating Watermarking With Denoising in Point Clouds

ICASSP 2025accepted

The integration of depth sensing and laser scanning technologies has propelled point cloud data to the forefront of 3D graphical modeling. This paper addresses a critical gap in the literature: the protection of intellectual property in generating point clouds using Diffusion Models (DMs). We introd…

Cited by 0SourceScholar
2025

Failure Cases Are Better Learned But Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial Training

ICCV 2025poster

Adversarial Training (AT) is one of the most effective methods to train robust Deep Neural Networks (DNNs). However, AT creates an inherent trade-off between clean accuracy and adversarial robustness, which is commonly attributed to the more complicated decision boundary caused by the insufficient l…

2025

Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and Regularization

AAAI 2025technical

Audio-Visual Learning (AVL) aims at the audio-visual perception with both audio and vision modalities. AVL also suffers from data insufficiency in many applications as with other unimodal tasks. Concurrently, AVL often needs to continuously learn over time rather than all knowledge simultaneously. C…

Cited by 0SourcePDFScholar
2025

Fusing Pruned and Backdoored Models: Optimal Transport-based Data-free Backdoor Mitigation

AAAI 2025technical

Backdoor attacks present a serious security threat to deep neuron networks (DNNs). Although numerous effective defense techniques have been proposed in recent years, they inevitably rely on the availability of either clean or poisoned data. In contrast, data-free defense techniques have evolved slow…

2025

Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition

ICASSP 2025accepted

Backdoor attacks have posed a significant threat to the security of deep neural networks (DNNs). Despite considerable strides in developing defenses against backdoor attacks in the visual domain, the specialized defenses for the audio domain remain empty. Furthermore, the defenses adapted from the v…

Cited by 0SourceScholar
2025

HiPoser: 3D Human Pose Estimation with Hierarchical Shared Learning at Parts-Level Using Inertial Measurement Units

AAAI 2025technical

This paper considers the challenging problem of 3D Human Pose Estimation (HPE) from a sparse set of Inertial Measurement Units (IMUs). Existing efforts typically reconstruct a pose sequence by either directly tackling whole-body motions or focusing on distinctive spatio-temporal features of local bo…

Cited by 0SourcePDFScholar
2025

Implanting Robust Watermarks in Latent Diffusion Models for Video Generation

ICASSP 2025accepted

In the dynamic realm of digital media, latent diffusion models (LDM) have revolutionized the generation of videos, surpassing the capabilities of traditional generative models. This paper presents Stable Video Signature, a pioneering watermarking framework for LDM in video generation. Addressing the…

Cited by 0SourceScholar
2025

Inter- and Intra-Sentence Cuer-Invariant Representation Learning for Generalizable Cued Speech Recognition

ICASSP 2025accepted

Cued Speech (CS) is a visual coding system that combines lip movements and hand gestures to represent spoken languages for hearing-impaired people. Automatic Cued Speech Recognition (ACSR) is an emerging research topic, but the cuer (i.e., people who perform CS) generalization problem of ACSR remain…

Cited by 0SourceScholar
2025

Learning Class Unique Features in Fine-Grained Visual Classification

ICASSP 2025accepted

A major challenge in Fine-Grained Visual Classification (FGVC) is distinguishing various categories with high inter-class similarity by learning the feature that differentiates the details. Conventional cross-entropy trained Convolutional Neural Network (CNN) fails this challenge as they may suffer…

Cited by 0SourceScholar
2025

Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination Enhancement

NeurIPS 2025poster

Low-light image enhancement (LLIE) faces persistent challenges in balancing reconstruction fidelity with cross-scenario generalization. While existing methods predominantly focus on deterministic pixel-level mappings between paired low/normal-light images, they often neglect the continuous physical…

Cited by 0SourcecodeScholar
2025

MambaTrack: Exploiting Dual-Enhancement for Night UAV Tracking

ICASSP 2025accepted

Night unmanned aerial vehicle (UAV) tracking is impeded by the challenges of poor illumination, with previous daylight-optimized methods demonstrating suboptimal performance in low-light conditions, limiting the utility of UAV applications. To this end, we propose an efficient mamba-based tracker, l…

Cited by 0SourceScholar
2025

MotionComposer: Enhancing Rhythmic Music Generation with Adaptive Retrieval Reference

ICASSP 2025accepted

With the rise of the AIGC era, rhythmic music generation has extensive applications, particularly with the surge in motion video creation. However, generating music that is rhythmically synchronized and stylistically aligned with motion video presents significant challenges. Although existing method…

Cited by 0SourceScholar
2025

Orchestrating Audio: Multi-Agent Framework for Long-Video Audio Synthesis

EMNLP 2025

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an unsolved challenge due to dynamic semantic shifts, audio diver

Cited by 0SourcePDFScholar
2025

Reliable Imputed-Sample Assisted Vertical Federated Learning

ICASSP 2025accepted

Vertical Federated Learning (VFL) is a well-known FL variant that enables multiple parties to collaboratively train a model without sharing their raw data. Existing VFL approaches focus on overlapping samples among different parties, while their performance is constrained by the limited number of th…

Cited by 0SourceScholar
2025

Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion

AAAI 2025technical

Face-based Voice Conversion (FVC) is a novel task that leverages facial images to generate the target speaker's voice style. Previous work has two shortcomings: (1) suffering from obtaining facial embeddings that are well-aligned with the speaker's voice identity information, and (2) inadequacy in d…

Cited by 4SourcePDFScholar
2025

StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding

ICCV 2025poster

Scene text image super-resolution (STISR) focuses on enhancing the clarity and readability of low-resolution text images. Existing methods often rely on text probability distribution priors derived from text recognizers to guide the super-resolution process. While effective in capturing general stru…

2025

SuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor Segmentation

CVPR 2025poster

Multimodal 3D segmentation involves a significant number of 3D convolution operations, which requires substantial computational resources and high-performance computing devices in MRI multimodal brain tumor segmentation. The key challenge in multimodal 3D segmentation is how to minimize network comp…

2025

Teaching Others Teaches Yourself: Semi-supervised Ensembled Pseudo-labeling Method for Image Classification

ICASSP 2025accepted

Semi-supervised methods have recently received significant attention in deep learning because they are able to reduce the dependence on labeled data while ensuring good performance. The pseudo-label based method has been widely used as a classic semi-supervised method but still suffers from the conf…

Cited by 0SourceScholar
2025

Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual Abduction

CVPR 2025poster

Fashion image editing is a valuable tool for designers to convey their creative ideas by visualizing design concepts. With the recent advances in text editing methods, significant progress has been made in fashion image editing. However, they face two key challenges: spurious correlations in trainin…

Cited by 0SourcePDFScholar
2025

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

EMNLP 2025

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has be

2025

Traversal Verification for Speculative Tree Decoding

NeurIPS 2025poster

Speculative decoding is a promising approach for accelerating large language models. The primary idea is to use a lightweight draft model to speculate the output of the target model for multiple subsequent timesteps, and then verify them in parallel to determine whether the drafted tokens should be…

Cited by 0SourceScholar
2025

UEVAVD: A Dataset for Developing UAV's Eye View Active Object Detection

RA-L 2025

Occlusion is a longstanding difficulty that challenges the UAV-based object detection. Many works address this problem by adapting the detection model. However, few of them exploit that the UAV could fundamentally improve detection performance by changing its viewpoint. Active Object Detection (AOD)

Cited by 4SourcecodeScholar
2025

Visual Representation Learning through Causal Intervention for Controllable Image Editing

CVPR 2025highlight

A key challenge for controllable image editing is that visual attributes with semantic meanings are not always independent, resulting in spurious correlations in model training. However, most existing methods ignore such issues, leading to biased causal visual representation learning and unintended…

Cited by 0SourcePDFScholar
2025

When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object Detection

ICCV 2025poster

Few-shot object detection aims to detect novel classes with limited samples. Recent methods have leveraged the rich semantic representations of pretrained vision transformer (ViT) to overcome the limitations of model fine-tuning, thereby improving the performance on novel classes. However, existing…

2025

YOLO-TCT: An Effective Network For Long-Tailed Cervical Cell Detection

ICASSP 2025accepted

The Thinprep Cytologic Test (TCT) is a vital component in the early detection of cervical cancer. However, conventional manual screening methods are hindered by inefficiencies and high levels of subjectivity. This study presents YOLO-TCT, an enhanced YOLOv9 network designed for the automated detecti…

Cited by 0SourceScholar
2024

Bridge to Non-Barrier Communication: Gloss-Prompted Fine-Grained Cued Speech Gesture Generation with Diffusion Model

IJCAI 2024poster

Cued Speech (CS) is an advanced visual phonetic encoding system that integrates lip reading with hand codings, enabling people with hearing impairments to communicate efficiently. CS video generation aims to produce specific lip and gesture movements of CS from audio or text inputs. The main challen…

2024

DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation

ICLR 2024poster

We present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks. DFormer has two new key innovations: 1) Unlike previous works that encode RGB-D information with RGB pretrained backbone, we pretrain the backbone using image-depth pairs from…

Cited by 58SourcePDFScholar
2024

Fall Prediction by a Spatio-Temporal Multi-Channel Causal Model from Wearable Sensors Data

ICASSP 2024accepted

Predicting human falls from wearable devices is a complex task due to the inherent diversity and causality of multivariate physical changes, where each instance exhibits a unique style of motion events and their spatio-temporal causal dependencies. Consequently, we propose a multichannel causal mode…

Cited by 0SourceScholar
2024

Hide in Thicket: Generating Imperceptible and Rational Adversarial Perturbations on 3D Point Clouds

CVPR 2024poster

Adversarial attack methods based on point manipulation for 3D point cloud classification have revealed the fragility of 3D models yet the adversarial examples they produce are easily perceived or defended against. The trade-off between the imperceptibility and adversarial strength leads most point a…

2024

Joint Pre-Encoding Representation and Structure Embedding for Efficient and Low-Resource Knowledge Graph Completion

EMNLP 2024main

Knowledge graph completion (KGC) aims to infer missing or incomplete parts in knowledge graph. The existing models are generally divided into structure-based and description-based models, among description-based models often require longer training and inference times as well as increased memory usa…

2024

Leveraging Noisy Labels of Nearest Neighbors for Label Correction and Sample Selection

ICASSP 2024accepted

Dealing with noisy labels (LNL) emerges as a critical challenge when applying deep learning (DL) in practical settings. Previous methodologies primarily concentrated on harnessing model predictions to mitigate the impact of noisy labels. Nevertheless, their efficacy is strongly contingent on the acc…

Cited by 0SourceScholar
2024

Predicting Fall Events by a Spatio-Temporal Topological Network with Multiple Wearable Sensors

ICASSP 2024accepted

A key challenge in sensor-based fall prediction is the fact that a fall event can often occur in various configurations of fall poses together with their own spatio-temporal dependencies. This leads us to define a spatio-temporal model to explicitly characterize these internal configurations of pose…

Cited by 0SourceScholar
2024

RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision

ICRA 2024poster

3D occupancy prediction holds significant promise in the fields of robot perception and autonomous driving, which quantifies 3D scenes into grid cells with semantic labels. Recent works mainly utilize complete occupancy labels in 3D voxel space for supervision. However, the expensive annotation proc…

Cited by 85SourcecodeScholar
2024

Right this way: Can VLMs Guide Us to See More to Answer Questions?

NeurIPS 2024poster

In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the suff…

2024

SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection

NeurIPS 2024spotlight

Synthetic Aperture Radar (SAR) object detection has gained significant attention recently due to its irreplaceable all-weather imaging capabilities. However, this research field suffers from both limited public datasets (mostly comprising <2K images with only mono-category objects) and inaccessible…

2024

SENCR: A Span Enhanced Two-Stage Network with Counterfactual Rethinking for Chinese NER

AAAI 2024technical

Recently, lots of works that incorporate external lexicon information into character-level Chinese named entity recognition(NER) to overcome the lackness of natural delimiters of words, have achieved many advanced performance. However, obtaining and maintaining high-quality lexicons is costly, espec…

Cited by 4SourcePDFScholar
2024

SurroundSDF: Implicit 3D Scene Understanding Based on Signed Distance Field

CVPR 2024highlight

Vision-centric 3D environment understanding is both vital and challenging for autonomous driving systems. Recently object-free methods have attracted considerable attention. Such methods perceive the world by predicting the semantics of discrete voxel grids but fail to construct continuous and accur…

Cited by 4SourcePDFScholar
2024

Unveiling and Mitigating Backdoor Vulnerabilities based on Unlearning Weight Changes and Backdoor Activeness

NeurIPS 2024poster

The security threat of backdoor attacks is a central concern for deep neural networks (DNNs). Recently, without poisoned data, unlearning models with clean data and then learning a pruning mask have contributed to backdoor defense. Additionally, vanilla fine-tuning with those clean data can help rec…

2024

WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark

NeurIPS 2024poster

Underwater Object Tracking (UOT) is essential for identifying and tracking submerged objects in underwater videos, but existing datasets are limited in scale, diversity of target categories and scenarios covered, impeding the development of advanced tracking algorithms. To bridge this gap, we take t…

2023

Bimodal Fusion Network for Basic Taste Sensation Recognition from Electroencephalography and Electromyography

ICASSP 2023accepted

Taste sensation can be objectively measured using electroencephalography (EEG) or electromyography (EMG). How-ever, it is still challenging to effectively utilize the complementary information from EEG and EMG signals in taste sensation recognition. This paper proposes a bimodal fusion network (Bi-F…

Cited by 0SourceScholar
2023

Evidential Uncertainty and Diversity Guided Active Learning for Scene Graph Generation

ICLR 2023poster

Scene Graph Generation (SGG) has already shown its great potential in various downstream tasks, but it comes at the price of a prohibitively expensive annotation process. To reduce the annotation cost, we propose using Active Learning (AL) for sampling the most informative data. However, directly po…

Cited by 17SourcePDFScholar
2023

Global Balanced Experts for Federated Long-Tailed Learning

ICCV 2023poster

Federated learning (FL) is a prevalent distributed machine learning approach that enables collaborative training of a global model across multiple devices without sharing local data. However, the presence of long-tailed data can negatively deteriorate the model's performance in real-world FL applica…

Cited by 14PDFcodeScholar
2023

Learning Symbolic Models for Graph-structured Physical Mechanism

ICLR 2023poster

Graph-structured physical mechanisms are ubiquitous in real-world scenarios, thus revealing underneath formulas is of great importance for scientific discovery. However, classical symbolic regression methods fail on this task since they can only handle input-output pairs that are not graph-structure…

Cited by 14SourcePDFScholar
2023

Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection With Single Point Supervision

CVPR 2023poster

Training a convolutional neural network (CNN) to detect infrared small targets in a fully supervised manner has gained remarkable research interests in recent years, but is highly labor expensive since a large number of per-pixel annotations are required. To handle this problem, in this paper, we ma…

2023

Memory-Augmented Contrastive Learning for Talking Head Generation

ICASSP 2023accepted

Given one reference facial image and a piece of speech as input, talking head generation aims to synthesize a realistic-looking talking head video. However, generating a lip-synchronized video with natural head movements is challenging. The same speech clip can generate multiple possible lip and hea…

Cited by 0SourceScholar
2023

Multi-Scale Visual Servoing Framework for Optical Microscopy Based on SIFT Matching

RA-L 2023

This letter introduces an innovative multi-scale visual servoing framework for optical microscopy, engineered to automatically reposition the microscope for high-magnification target view across multiple magnifications, thereby facilitating repetitive and accurate histologic biopsies. The framework

Cited by 7SourceScholar
2023

Preserving Structural Consistency in Arbitrary Artist and Artwork Style Transfer

AAAI 2023technical

Deep generative models are effective in style transfer. Previous methods learn one or several specific artist-style from a collection of artworks. These methods not only homogenize the artist-style of different artworks of the same artist but also lack generalization for the unseen artists. To solv…

Cited by 5SourcePDFScholar
2023

Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory Inversion

ICASSP 2023accepted

Acoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent scenario. Most current works focus on extracting features directly…

Cited by 0SourceScholar
2022

"Restore Globally, Refine Locally: A Mask-Guided Scheme to Accelerate Super-Resolution Networks"

ECCV 2022poster

"Single image super-resolution (SR) has been boosted by deep convolutional neural networks with growing model complexity and computational costs. To deploy existing SR networks onto edge devices, it is necessary to accelerate them for large image (4K) processing. The different areas in an image ofte…

2022

Acoustic-to-Articulatory Inversion Based on Speech Decomposition and Auxiliary Feature

ICASSP 2022accepted

Acoustic-to-articulatory inversion (AAI) is to obtain the movement of articulators from speech signals. Until now, achieving a speaker-independent AAI remains a challenge given the limited data. Besides, most current works only use audio speech as input, causing an inevitable performance bottleneck.…

Cited by 0SourceScholar
2022

Anti-Forgery: Towards a Stealthy and Robust DeepFake Disruption Attack via Adversarial Perceptual-aware Perturbations

IJCAI 2022poster

DeepFake is becoming a real risk to society and brings potential threats to both individual privacy and political security due to the DeepFaked multimedia are realistic and convincing. However, the popular DeepFake passive detection is an ex-post forensics countermeasure and failed in blocking the d…

2022

Boosting Black-Box Attack With Partially Transferred Conditional Adversarial Distribution

CVPR 2022poster

This work studies black-box adversarial attacks against deep neural networks (DNNs), where the attacker can only access the query feedback returned by the attacked DNN model, while other information such as model parameters or the training datasets are unknown. One promising approach to improve atta…

Cited by 49PDFcodeScholar
2022

DKNAS: A Practical Deep Keypoint Extraction Framework Based on Neural Architecture Search

ICRA 2022poster

Keypoint extraction including both keypoint detection and description is a fundamental step in a wide range of geometric multimedia applications. In recent years, many learning-based approaches for keypoint extraction emerge and achieve promising results. However, they usually design network archite…

Cited by 1SourceScholar
2022

Decoupling Makes Weakly Supervised Local Feature Better

CVPR 2022poster

Weakly supervised learning can help local feature methods to overcome the obstacle of acquiring a large-scale dataset with densely labeled correspondences. However, since weak supervision cannot distinguish the losses caused by the detection and description steps, directly conducting weakly supervis…

Cited by 61PDFcodeScholar
2022

Dynamic Binary Neural Network by Learning Channel-Wise Thresholds

ICASSP 2022accepted

Binary neural networks (BNNs) constrain weights and activations to +1 or -1 with limited storage and computational cost, which is hardware-friendly for portable devices. Recently, BNNs have achieved remarkable progress and been adopted into various fields. However, the performance of BNNs is sensiti…

Cited by 0SourceScholar
2022

Efficient Video Transformers with Spatial-Temporal Token Selection

ECCV 2022poster

"Video transformers have achieved impressive results on major video recognition benchmarks, however they suffer from high computational cost. In this paper, we present STTS, a token selection framework that dynamically selects a few informative tokens in both temporal and spatial dimensions conditio…

2022

Explainable Survival Analysis with Convolution-Involved Vision Transformer

AAAI 2022technical

Image-based survival prediction models can facilitate doctors in diagnosing and treating cancer patients. With the advance of digital pathology technologies, the big whole slide images (WSIs) provide increasing resolution and more details for diagnosis. However, the gigabyte-size WSIs would make mos…

2022

Highly-Efficient Incomplete Large-Scale Multi-View Clustering With Consensus Bipartite Graph

CVPR 2022poster

Multi-view clustering has received increasing attention due to its effectiveness in fusing complementary information without manual annotations. Most previous methods hold the assumption that each instance appears in all views. However, it is not uncommon to see that some views may contain some miss…

Cited by 143PDFcodeScholar
2022

Learnable Lookup Table for Neural Network Quantization

CVPR 2022poster

Neural network quantization aims at reducing bit-widths of weights and activations for memory and computational efficiency. Since a linear quantizer (i.e., round(*) function) cannot well fit the bell-shaped distributions of weights and activations, many existing methods use pre-defined functions (e.…

Cited by 65PDFScholar
2022

Residual-Guided Personalized Speech Synthesis based on Face Image

ICASSP 2022accepted

Previous works derive personalized speech features by training the model on a large dataset composed of his/her audio sounds. It was reported that face information has a strong link with the speech sound. Thus in this work, we innovatively extract personalized speech features from human faces to syn…

Cited by 0SourceScholar
2021

A New Tubular Structure Tracking Algorithm Based On Curvature-Penalized Perceptual Grouping

ICASSP 2021accepted

In this paper, we propose a new minimal path-based framework for minimally interactive tubular structure tracking in conjunction with a perceptual grouping scheme. The minimal path models have shown great advantages in tubular structures tracing. However, they suffer from shortcuts or short branches…

Cited by 0SourceScholar
2021

Autonomous Navigation of an Ultrasound Probe Towards Standard Scan Planes with Deep Reinforcement Learning

ICRA 2021poster

Autonomous ultrasound (US) acquisition is an important yet challenging task, as it involves interpretation of the highly complex and variable images and their spatial relationships. In this work, we propose a deep reinforcement learning framework to autonomously control the 6-D pose of a virtual US…

Cited by 66SourceScholar
2021

Exploring Inter-Channel Correlation for Diversity-Preserved Knowledge Distillation

ICCV 2021poster

Knowledge Distillation has shown very promising ability in transferring learned representation from the larger model (teacher) to the smaller one (student). Despite many efforts, prior methods ignore the important role of retaining inter-channel correlation of features, leading to the lack of captur…

Cited by 125PDFcodeScholar
2021

From Manual Operation to Collaborative Robot Assembly: An Integrated Model of Productivity and Ergonomic Performance

RA-L 2021

Manufacturing systems involve machines and people. Both productivity and ergonomic performance are of significant importance in manufacturing. However, there is no integrated model to analyze them simultaneously. To bridge this gap, a unified model is introduced to evaluate the productivity and ergo

Cited by 26SourceScholar
2021

Generalized Coherent Point Drift With Multi-Variate Gaussian Distribution and Watson Distribution

RA-L 2021

This letter introduces a novel rigid point set registration (PSR) approach that accurately aligns the pre-operative space and the intra-operative space together in the scenario of computer-assisted orthopedic surgery (CAOS). Motivated by considering anisotropic positional localization noise and util

Cited by 6SourceScholar
2021

Group Whitening: Balancing Learning Efficiency and Representational Capacity

CVPR 2021poster

Batch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model's learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population…

Cited by 24PDFcodeScholar
2021

Machine Learning in Manufacturing Ergonomics: Recent Advances, Challenges, and Opportunities

RA-L 2021

The rapid development of machine learning (ML) technology has introduced substantial impact on ergonomics research in manufacturing. Numerous studies and practices have been carried out to apply ML techniques to address manufacturing ergonomics issues, which has brought extensive opportunities as we

Cited by 15SourceScholar
2021

One Pass Late Fusion Multi-view Clustering

ICML 2021spotlight

Existing late fusion multi-view clustering (LFMVC) optimally integrates a group of pre-specified base partition matrices to learn a consensus one. It is then taken as the input of the widely used k-means to generate the cluster labels. As observed, the learning of the consensus partition matrix and…

Cited by 127SourcePDFScholar
2021

One-Pass Multi-View Clustering for Large-Scale Data

ICCV 2021poster

Existing non-negative matrix factorization based multi-view clustering algorithms compute multiple coefficient matrices respect to different data views, and learn a common consensus concurrently. The final partition is always obtained from the consensus with classical clustering techniques, such as…

Cited by 113PDFcodeScholar
2021

PiPo-Net: A Semi-automatic and Polygon-based Annotation Method for Pathological Images

IROS 2021poster

Metastatic involvement of lymph nodes is one of the most important prognostic variables for many cancers. Several deep learning based algorithms have been developed to segment metastatic regions in pathological images to help predict prognosis. However, the training of these methods requires a large…

Cited by 5SourceScholar
2021

Pixel Difference Networks for Efficient Edge Detection

ICCV 2021poster

Recently, deep Convolutional Neural Networks (CNNs) can achieve human-level performance in edge detection with the rich and abstract edge representation capacities. However, the high performance of CNN based edge detection is achieved with a large pretrained CNN backbone, which is memory and energy…

Cited by 452PDFcodeScholar
2021

Reciprocally Rotating Magnetic Actuation and Automatic Trajectory Following for Wireless Capsule Endoscopy

ICRA 2021poster

Active wireless capsule endoscopy (WCE) under magnetic actuation is a promising technology to reduce the inspection time and relieve the burden of physicians. In this paper, we propose a reciprocally rotating magnetic actuation method for trajectory following of a capsule and develop its dynamic mod…

Cited by 4SourceScholar
2021

Self-Supervised Depth Estimation Via Implicit Cues from Videos

ICASSP 2021accepted

In self-supervised monocular depth estimation, the depth discontinuity and motion objects' artifacts are still challenging problems. Existing self-supervised methods usually utilize two views to train the depth estimation network and use one single view to make predictions. Compared with static view…

Cited by 0SourceScholar
2020

Dynamic Group Convolution for Accelerating Convolutional Neural Networks

ECCV 2020poster

Replacing normal convolutions with group convolutions can significantly increase the computational efficiency of modern deep convolutional networks, which has been widely adopted in compact network architecture designs. However, existing group convolutions undermine the original network structures b…

2020

JGR-P2O: Joint Graph Reasoning based Pixel-to-Offset Prediction Network for 3D Hand Pose Estimation from a Single Depth Image

ECCV 2020poster

State-of-the-art single depth image-based 3D hand pose estimation methods are based on dense predictions, including voxel-to-voxel predictions, point-to-point regression, and pixel-wise estimations. Despite the good performance, those methods have a few issues in nature, such as the poor trade-off b…

2020

Layer-wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs

ECCV 2020poster

Conditioning analysis uncovers the landscape of an optimization objective by exploring the spectrum of its curvature matrix. This has been well explored theoretically for linear models. We extend this analysis to deep neural networks (DNNs) in order to investigate their learning dynamics. To this en…

Cited by 12SourcePDFScholar
2020

Learning Attentive and Hierarchical Representations for 3D Shape Recognition

ECCV 2020poster

This paper proposes a novel method for 3D shape representation learning, namely Hyperbolic Embedded Attentive Representation (HEAR). Different from existing multi-view based methods, HEAR develops a unified framework to address both multi-view redundancy and single-view incompleteness. Specifically,…

Cited by 36SourcePDFScholar
2020

Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification

CVPR 2020poster

Video-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), t…

Cited by 189PDFcodeScholar
2020

On the Number of Linear Regions of Convolutional Neural Networks

ICML 2020poster

One fundamental problem in deep learning is understanding the outstanding performance of deep Neural Networks (NNs) in practice. One explanation for the superiority of NNs is that they can realize a large class of complicated functions, i.e., they have powerful expressivity. The expressivity of a Re…

Cited by 98SourcePDFScholar
2020

Region Graph Embedding Network for Zero-Shot Learning

ECCV 2020poster

Most of the existing Zero-Shot Learning (ZSL) approaches learn direct embeddings from global features or image parts (regions) to the semantic space, which, however, fail to capture the appearance relationships between different local regions within a single image. In this paper, to model the relati…

Cited by 195SourcePDFScholar
2020

Set and Rebase: Determining the Semantic Graph Connectivity for Unsupervised Cross-Modal Hashing

IJCAI 2020poster

The label-free nature of unsupervised cross-modal hashing hinders models from exploiting the exact semantic data similarity. Existing research typically simulates the semantics by a heuristic geometric prior in the original feature space. However, this introduces heavy bias into the model as the ori…

Cited by 0SourcePDFScholar
2019

Attentive Region Embedding Network for Zero-Shot Learning

CVPR 2019poster

Zero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of t…

Cited by 351PDFScholar
2019

Building Detail-Sensitive Semantic Segmentation Networks With Polynomial Pooling

CVPR 2019poster

Semantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classif…

Cited by 34PDFScholar
2019

Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical Images

CVPR 2019poster

Medical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires imag…

Cited by 327PDFScholar
2019

Deep Sketch-Shape Hashing With Segmented 3D Stochastic Viewing

CVPR 2019poster

Sketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH),…

Cited by 48PDFScholar
2019

Iterative Normalization: Beyond Standardization Towards Efficient Whitening

CVPR 2019poster

Batch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies…

Cited by 183PDFcodeScholar
2019

RANet: Ranking Attention Network for Fast Video Object Segmentation

ICCV 2019poster

Despite online learning (OL) techniques have boosted the performance of semi-supervised video object segmentation (VOS) methods, the huge time costs of OL greatly restricts their practicality. Matching based and propagation based methods run at a faster speed by avoiding OL techniques. However, they…

Cited by 274PDFcodeScholar
2019

Sparse Subspace Clustering for Evolving Data Streams

ICASSP 2019accepted

The data streams arising in many applications can be modeled as a union of low-dimensional subspaces known as multi-subspace data streams (MSDSs). Clustering MSDSs according to their underlying low-dimensional subspaces is a challenging problem which has not been resolved satisfactorily by existing…

Cited by 0SourceScholar
2019

Two Generator Game: Learning to Sample via Linear Goodness-of-Fit Test

NeurIPS 2019poster

Learning the probability distribution of high-dimensional data is a challenging problem. To solve this problem, we formulate a deep energy adversarial network (DEAN), which casts the energy model learned from real data into an optimization of a goodness-of-fit (GOF) test statistic. DEAN can be inter…

Cited by 6SourcePDFScholar
2018

Automatic Temporal Segmentation of Hand Movements for Hand Positions Recognition in French Cued Speech

ICASSP 2018accepted

In the context of Cued Speech (CS) recognition, the recognition of lips and hand movements is a key task. As we know, a good temporal segmentation is necessary for the supervised recognition system. However, lips and hand streams cannot share the same temporal segmentation since they are not synchro…

Cited by 0SourceScholar
2018

Deep Multi-Task Learning to Recognise Subtle Facial Expressions of Mental States

ECCV 2018poster

Facial expression recognition is a topical task. However, very little research investigates subtle expression recognition, which is important for mental activity analysis, deception detection, etc. We address subtle expression recognition through convolutional neural networks (CNNs) by developing mu…

Cited by 55SourcePDFScholar
2018

Generative Domain-Migration Hashing for Sketch-to-Image Retrieval

ECCV 2018poster

Due to the succinct nature of free-hand sketch drawings, sketch-based image retrieval (SBIR) has abundant practical use cases in consumer electronics. However, SBIR remains a long-standing unsolved problem mainly due to the significant discrepancy between the sketch domain and the image domain. In t…

2018

High-Speed Light Field Image Formation Analysis Using Wavefield Modeling with Flexible Sampling

ICASSP 2018accepted

Understanding the image formation inside plenoptic cameras is significant for the investigations of improving the low spatial resolution. Most researches explore the image formation from the perspective of geometric optics. However, as the hardware components in combination with low-aperture optical…

Cited by 0SourceScholar
2018

Highly-Economized Multi-View Binary Compression for Scalable Image Clustering

ECCV 2018poster

How to economically cluster large-scale multi-view images is a long-standing problem in computer vision. To tackle this challenge, this paper introduces a novel approach named Highly-economized Scalable Image Clustering (HSIC) that radically surpasses conventional image clustering methods via binary…

Cited by 55SourcePDFScholar
2018

Super Wide Regression Network for Unsupervised Cross-Database Facial Expression Recognition

ICASSP 2018accepted

Unsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature dist…

Cited by 0SourceScholar
2018

TBN: Convolutional Neural Network with Ternary Inputs and Binary Weights

ECCV 2018poster

Despite the remarkable success of Convolutional Neural Networks (CNNs) on generalized visual tasks, high computational and memory costs restrict their comprehensive applications on consumer electronics (e.g., portable or smart wearable devices). Recent advancements in binarized networks have demonst…

2018

Unsupervised Cross-Corpus Speech Emotion Recognition Using Domain-Adaptive Subspace Learning

ICASSP 2018accepted

In this paper, we investigate an interesting problem, i.e., unsupervised cross-corpus speech emotion recognition (SER), in which the training and testing speech signals come from two different speech emotion corpora. Meanwhile, the training speech signals are labeled, while the label information of…

Cited by 0SourceScholar
2017

Binary Coding for Partial Action Analysis With Limited Observation Ratios

CVPR 2017poster

Traditional action recognition methods aim to recognize actions with complete observations/executions. However, it is often difficult to capture fully executed actions due to occlusions, interruptions, etc. Meanwhile, action prediction/recognition in advance based on partial observations is essentia…

Cited by 34PDFScholar
2017

Deep Binaries: Encoding Semantic-Rich Cues for Efficient Textual-Visual Cross Retrieval

ICCV 2017poster

Cross-modal hashing is usually regarded as an effective technique for large-scale textual-visual cross retrieval, where data from different modalities are mapped into a shared Hamming space for matching. Most of the traditional textual-visual binary encoding methods only consider holistic image repr…

Cited by 61PDFScholar
2017

Deep Sketch Hashing: Fast Free-Hand Sketch-Based Image Retrieval

CVPR 2017spotlight

Free-hand sketch-based image retrieval (SBIR) is a specific cross-view retrieval task, in which queries are abstract and ambiguous sketches while the retrieval database is formed with natural images. Work in this area mainly focuses on extracting representative and shared features for sketches and n…

Cited by 319PDFcodeScholar
2017

Fast Person Re-Identification via Cross-Camera Semantic Binary Transformation

CVPR 2017poster

Numerous methods have been proposed for person re-identification, most of which however neglect the matching efficiency. Recently, several hashing based approaches have been developed to make re-identification more scalable for large-scale gallery sets. Despite their efficiency, these works ignore c…

Cited by 94PDFScholar
2017

From Zero-Shot Learning to Conventional Supervised Classification: Unseen Visual Data Synthesis

CVPR 2017poster

Robust object recognition systems usually rely on powerful feature extraction mechanisms from a large number of real images. However, in many realistic applications, collecting sufficient images for ever-growing new classes is unattainable. In this paper, we propose a new Zero-shot learning (ZSL) fr…

Cited by 180PDFScholar
2017

Preliminary study on magnetic tracking based navigation for wire-driven flexible robot

IROS 2017poster

Flexible manipulator enables curvilinear accessibility through small incisions or natural orifices for minimally invasive surgery and diagnosis, which makes it a good choice for minimally invasive surgery. In order to control the robot precisely and safely, the real-time position and shape informati…

Cited by 9SourceScholar
2017

Zero-Shot Action Recognition With Error-Correcting Output Codes

CVPR 2017poster

Recently, zero-shot action recognition (ZSAR) has emerged with the explosive growth of action categories. In this paper, we explore ZSAR from a novel perspective by adopting the Error-Correcting Output Codes (dubbed ZSECOC). Our ZSECOC equips the conventional ECOC with the additional capability of Z…

Cited by 186PDFScholar