← Search

Bin Liu

88 accepted papers

2026

D2ACE: Multi-Label Batch Selection Guided by Dual Dynamics and Adaptive Correlation Enhancement

IJCAI 2026

Batch selection is crucial for improving both training efficiency and predictive performance in deep multi-label classification (MLC). Existing batch selection methods typically rely on a single metric to assess instance importance and use static label weights to distinguish label significance, negl

Cited by 0Scholar
2026

Discriminative Graph Embedding Framework via Label-Free Marginal Fisher Analysis

AAAI 2026technical

Marginal Fisher Analysis (MFA) is a classical dimensionality reduction (DR) method that leverages dual graphs to capture intra-class compactness and inter-class separability. However, MFA’s reliance on high-quality labels limits its practical application. For another, existing unsupervised DR method

Cited by 0SourcePDFScholar
2026

EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding

AAAI 2026technical

In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with both accurate layout and high fidelity to the text description.(e.g., spatial relationship), and grounding the image robustly at the same time. We believe that

Cited by 0SourcePDFScholar
2026

Federated Incomplete Multi-View Clustering with Tensorized Low-Rank Constraint

AAAI 2026technical

Federated Multi-View Clustering has gained increasing attention for its ability to discover complementary clustering structures of distributed multi-view data while preserving data privacy. However, real-world clients often only have access to partial views, and the view incompleteness poses great c

Cited by 0SourcePDFScholar
2026

LAKAN: LANDMARK-ASSISTED ADAPTIVE KOLMOGOROV-ARNOLD NETWORK FOR FACE FORGERY DETECTION

ICASSP 2026poster

The rapid development of deepfake generation techniques necessitates robust face forgery detection algorithms. While methods based on Convolutional Neural Networks (CNNs) and Transformers are effective, there is still room for improvement in modeling the highly complex and non-linear nature of forge…

Cited by 0SourcePDFScholar
2026

MFEN: Multi-Frequency Expert Network for Visible-Infrared Person Re-ID

CVPR 2026

Visible-infrared person re-identification (VI-ReID) is challenging due to the large modality discrepancy between visible and infrared images. We contend that this discrepancy is largely related to differing lighting conditions, including differences in light wavelength and light source type. Recentl

Cited by 0SourceScholar
2026

MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing

IJCAI 2026

Recent writing assistants are increasingly shifting from passive, prompt-driven interaction to proactive, suggestion-based completion, which integrates localized continuations into the writing flow and reduces coordination burden. However, existing evaluations simply focus on output quality, failing

Cited by 0Scholar
2026

R$^2$TUA: Reconstruction-residual Based Targeted and Untargeted Attack Against Text-Image Person Re-Identification

CVPR 2026

Text-Image Person Re-Identification (TI-ReID) is widely deployed in intelligent surveillance. Built on deep neural networks and vision-language models, TI-ReID models inherit vulnerabilities to adversarial attacks, posing security risks. Yet its security remains less explored than retrieval accuracy

Cited by 0SourceScholar
2026

Rethinking Cross-Modal Anchor Alignment for Mitigating Error Accumulation

CVPR 2026

Mitigating noisy correspondence in cross-modal matching poses a serious challenge due to the problem of error accumulation. Existing methods primarily attribute this accumulation to errors caused by noisy sample pairs. However, a novel source of error from clean sample pairs (also termed anchor pair

Cited by 0SourceScholar
2026

SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splatting

CVPR 2026

3D Gaussian representations have emerged as a powerful paradigm for digital head modeling, achieving photorealistic quality with real-time rendering. However, intuitive and interactive creation or editing of 3D Gaussian head models remains challenging. Although 2D sketches provide an ideal interacti

Cited by 0SourceScholar
2025

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

ICML 2025oral

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suff…

2025

Batch Selection for Multi-Label Classification Guided by Uncertainty and Dynamic Label Correlations

AAAI 2025technical

The accuracy of deep neural networks is significantly influenced by the effectiveness of mini-batch construction during training. In single-label scenarios, such as binary and multi-class classification tasks, it has been demonstrated that batch selection algorithms preferring samples with higher un…

2025

Bridging the Sky and Ground: Towards View-Invariant Feature Learning for Aerial-Ground Person Re-Identification

ICCV 2025poster

Aerial-Ground Person Re-Identification (AG-ReID) is a practical yet challenging task that involves cross-platform matching between aerial and ground cameras. Existing person Re-Identification (Re-ID) methods are primarily designed for homogeneous camera settings, such as ground-to-ground or aerial-t…

Cited by 0SourcePDFScholar
2025

CMGait: Enhancing Cross-Modality Gait Recognition between LiDAR and RGB through Contrastive Identity-consistent Feature Aggregation

ICASSP 2025accepted

Combination usage of LiDAR and RGB cameras for gait recognition can achieve cross space recognition and privacy protection. In addition, the widespread application of LiDAR cameras with 3D geometry information and the large amount of RGB gaits has led to the demand for cross-modality gait recognitio…

Cited by 0SourceScholar
2025

FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and Segmentation

ICCV 2025poster

Zero-shot anomaly detection (ZSAD) requires detection models trained using auxiliary data to detect anomalies without any training sample in a target dataset. It is challenging since the models need to generalize to anomalies across different domains. Recently, CLIP-based anomaly detection methods,…

Cited by 0SourcePDFScholar
2025

KLMN: Knowledge distillation based lightweight multi-clue image forgery detection and localization

ICASSP 2025accepted

Current image forensics methods often utilize image features from various frequency domains. However, the effective use of these features frequently depends on complex network architectures and a large number of parameters. In this paper, we introduce a lightweight Multi-Clue image forgery detection…

Cited by 0SourceScholar
2025

Listen, Watch, and Learn to Feel: Retrieval-Augmented Emotion Reasoning for Compound Emotion Generation

ACL 2025finding

The ability to comprehend human emotion using multimodal large language models (MLLMs) is essential for advancing human-AI interaction and multimodal sentiment analysis. While psychology theory-based human annotations have contributed to multimodal emotion tasks, the subjective nature of emotional p…

Cited by 0SourcePDFScholar
2025

OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition

ICML 2025poster

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-apprai…

2025

Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation Learning

AAAI 2025technical

In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. For example, MAMP shows that i…

Cited by 0SourcePDFScholar
2025

Think before Recommendation: Autonomous Reasoning-enhanced Recommender

NeurIPS 2025poster

The core task of recommender systems is to learn user preferences from historical user-item interactions. With the rapid development of large language models (LLMs), recent research has explored leveraging the reasoning capabilities of LLMs to enhance rating prediction tasks. However, existing disti…

Cited by 0SourceScholar
2025

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

IJCAI 2025

In real applications, person re-identification (ReID) expects to retrieve the target person at any time, including both daytime and nighttime, ranging from short-term to long-term. However, existing ReID tasks and datasets cannot meet this requirement, as they are constrained by available time and o

2025

Training an Anti-KD Model that Cannot Teach Students via Similarity Disruption

ICASSP 2025accepted

Knowledge Distillation (KD) aims to enhance the performance of student models by transferring knowledge from teacher models. While reaping the benefits of KD, the intellectual property risks associated with it cannot be ignored. Even if models are released without training data or provided as a serv…

Cited by 0SourceScholar
2025

Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region Matching

AAAI 2025technical

Open-vocabulary semantic segmentation (OVSS) aims to segment images of arbitrary categories specified by class labels. While previous approaches relied on extensive image-text pairs or dense semantic annotations, recent training-free methods attempted to overcome these limitations by constructing se…

Cited by 0SourcePDFScholar
2025

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection

NeurIPS 2025poster

A primary impediment to scaling reinforcement learning (RL) for large language model (LLM) training is the substantial computational cost, predominantly arising from the necessity of multi-sampling for policy optimization and evaluation. This underscores the critical yet challenging nature of effici…

Cited by 0SourceScholar
2025

UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype Discovery

CVPR 2025poster

Recent advancements in in-context segmentation generalists have demonstrated significant success in performing various image segmentation tasks using a limited number of labeled example images. However, real-world applications present challenges due to the variability of support examples, which ofte…

Cited by 0SourcePDFScholar
2025

VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making

EMNLP 2025

Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm. We show that this paradigm fundamentally limits performance in multi-input multi-output (MI

2024

Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning

ICASSP 2024accepted

Automatic anime sketch colorization aims to generate a color image from a sketch image, which is challenging due to limited structure and semantic understanding, leading to constrained style, and semantic color inconsistency. In this paper, we introduce a sketch to color diffusion model with semanti…

Cited by 0SourceScholar
2024

Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning

AAAI 2024technical

Open-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the adv…

2024

DSIS: A Novel (K, N) Threshold Deniable Secret Image Sharing Scheme with Lossless Recovery

ICASSP 2024accepted

Secret image sharing (SIS) schemes have undergone significant development. However, to the best of our knowledge, none of the existing schemes has considered the deniable property during secret sharing. This presents a problem when we need to share secret images through an untrusted and supervised c…

Cited by 0SourceScholar
2024

Enhancing Semi-supervised Domain Adaptation via Effective Target Labeling

AAAI 2024technical

Existing semi-supervised domain adaptation (SSDA) models have exhibited impressive performance on the target domain by effectively utilizing few labeled target samples per class (e.g., 3 samples per class). To guarantee an equal number of labeled target samples for each class, however, they require…

2024

Exploiting Modality-Specific Features for Multi-Modal Manipulation Detection and Grounding

ICASSP 2024accepted

AI-synthesized text and images have gained significant attention, particularly due to the widespread dissemination of multi-modal manipulations on the internet, which has resulted in numerous negative impacts on society. Existing methods for multi-modal manipulation detection and grounding primarily…

Cited by 0SourceScholar
2024

Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents

IJCAI 2024poster

Recent studies have uncovered the potential of Large Language Models (LLMs) in addressing complex sequential decision-making tasks through the provision of high-level instructions. However, LLM-based agents lack specialization in tackling specific target problems, particularly in real-time dynamic e…

2024

Llama SLayer 8B: Shallow Layers Hold the Key to Knowledge Injection

EMNLP 2024finding

As a manner to augment pretrained large language models (LLM), knowledge injection is critical to develop vertical domain large models and has been widely studied. While most current approaches, including parameter-efficient fine-tuning (PEFT) and block expansion methods, uniformly apply knowledge a…

2024

MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators

AAAI 2024technical

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion directly from textual action descriptions, they often support only…

2024

Pseudo Labels Regularization for Imbalanced Partial-Label Learning

ICASSP 2024accepted

Partial-label learning (PLL) is an important branch of weakly supervised learning where the single ground truth resides in a set of candidate labels, while the research rarely considers the label imbalance. A recent study for imbalanced PLL propose that the combinatorial challenge of partial-label l…

Cited by 0SourceScholar
2024

SE-SIS: Shadow-Embeddable Lossless Secret Image Sharing for Greyscale Images

ICASSP 2024accepted

Secret image sharing (SIS) has made significant progress in research and has found wide applications. However, we note that shadows of traditional SIS contain a large amount of redundancy. A novel Shadow-Embeddable Secret Image Sharing scheme (SE-SIS) leveraging the redundancy in the shadows is prop…

Cited by 0SourceScholar
2024

TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target Detection

AAAI 2024technical

Infrared small target detection (ISTD) is critical to national security and has been extensively applied in military areas. ISTD aims to segment small target pixels from background. Most ISTD networks focus on designing feature extraction blocks or feature fusion modules, but rarely describe the IST…

Cited by 15SourcePDFScholar
2024

Towards More Unified In-context Visual Understanding

CVPR 2024poster

The rapid advancement of large language models (LLMs) has accelerated the emergence of in-context learning (ICL) as a cutting-edge approach in the natural language processing domain. Recently ICL has been employed in visual understanding tasks such as semantic segmentation and image captioning yield…

Cited by 12SourcePDFScholar
2024

Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identification

AAAI 2024technical

Text-to-Image person re-identification (TI-ReID) aims to retrieve the images of target identity according to the given textual description. The existing methods in TI-ReID focus on aligning the visual and textual modalities through contrastive feature alignment or reconstructive masked language mode…

Cited by 14SourcePDFScholar
2024

View Crafting For Instance-Level Representation from Scene Images

ICASSP 2024accepted

Existing image-level self-supervised learning (SSL) methods pre-trained on natural scene data can have difficulty in adating to dense prediction tasks. However, scene images contain multiple varied instances. We devise two techniques to craft high-quality scene and instance views for instance-level…

Cited by 0SourceScholar
2023

A Reduction-based Framework for Sequential Decision Making with Delayed Feedback

NeurIPS 2023poster

We study stochastic delayed feedback in general single-agent and multi-agent sequential decision making, which includes bandits, single-agent Markov decision processes (MDPs), and Markov games (MGs). We propose a novel reduction-based framework, which turns any multi-batched algorithm for sequential…

Cited by 7SourcePDFScholar
2023

ALIM: Adjusting Label Importance Mechanism for Noisy Partial Label Learning

NeurIPS 2023poster

Noisy partial label learning (noisy PLL) is an important branch of weakly supervised learning. Unlike PLL where the ground-truth label must conceal in the candidate label set, noisy PLL relaxes this constraint and allows the ground-truth label may not be in the candidate label set. To address this c…

2023

BAUENet: Boundary-Aware Uncertainty Enhanced Network for Infrared Small Target Detection

ICASSP 2023accepted

Infrared small target detection (ISTD) is indispensable in remote sensing and military surveillance. Existing ISTD methods can discover regularly-shaped and clear objects well, but tend to overlook the tough-to-detect ones, such as targets with irregular shapes or blurry boundaries, causing inaccura…

Cited by 0SourceScholar
2023

BBDM: Image-to-Image Translation With Brownian Bridge Diffusion Models

CVPR 2023poster

Image-to-image translation is an important and challenging problem in computer vision and image processing. Diffusion models(DM) have shown great potentials for high-quality image synthesis, and have gained competitive performance on the task of image-to-image translation. However, most of the exist…

2023

Dual-Feature Enhancement for Weakly Supervised Temporal Action Localization

ICASSP 2023accepted

Weakly-supervised Temporal Action Localization (WTAL) aims at localizing actions in untrimmed videos with only video-level labels. Most existing methods embrace a "localization by classification" paradigm and adopt a model that pre-trained with recognition task for feature extraction. The gap betwee…

Cited by 0SourceScholar
2023

Dual-Uncertainty Guided Curriculum Learning and Part-Aware Feature Refinement for Domain Adaptive Person Re-Identification

ICASSP 2023accepted

Unsupervised Domain Adaptative person re-identification (UDA ReID) aims to transfer the knowledge of pre-trained model from labeled source domain to unlabeled target domain. Although the current clustering-based methods have achieved promising success, they neglect the tolerance of the model to cope…

Cited by 0SourceScholar
2023

Evopose: A Recursive Transformer for 3D Human Pose Estimation with Kinematic Structure Priors

ICASSP 2023accepted

Transformer is popular in recent 3D human pose estimation, which utilizes long-term modeling to lift 2D keypoints into the 3D space. However, current transformer-based methods do not fully exploit the prior knowledge of the human skeleton provided by the kinematic structure. In this paper, we propos…

Cited by 0SourceScholar
2023

Fluid Dynamics-Inspired Network for Infrared Small Target Detection

IJCAI 2023poster

Most infrared small target detection (ISTD) networks focus on building effective neural blocks or feature fusion modules but none describes the ISTD process from the image evolution perspective. The directional evolution of image pixels influenced by convolution, pooling and surrounding pixels is an…

Cited by 12SourcePDFScholar
2023

Interpret ESG Rating’s Impact on the Industrial Chain Using Graph Neural Networks

IJCAI 2023poster

We conduct a quantitative analysis of the development of the industry chain from the environmental, social, and governance (ESG) perspective, which is an overall measure of sustainability. Factors that may impact the performance of the industrial chain have been studied in the literature, such as g…

Cited by 9SourcePDFScholar
2023

VRA: Variational Rectified Activation for Out-of-distribution Detection

NeurIPS 2023poster

Out-of-distribution (OOD) detection is critical to building reliable machine learning systems in the open world. Researchers have proposed various strategies to reduce model overconfidence on OOD data. Among them, ReAct is a typical and effective technique to deal with model overconfidence, which tr…

2022

Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification

ECCV 2022poster

"Graph-based models have achieved great success in person re-identification tasks recently, which compute the graph topology structure (affinities) among different people first and then pass the information across them to achieve stronger features. But we find existing graph-based methods in the vis…

Cited by 51SourcePDFScholar
2022

End-to-End Network Based on Transformer for Automatic Detection of Covid-19

ICASSP 2022accepted

The novel coronavirus disease (COVID-19) was declared a pandemic by the World Health Organization. The cumulative number of deaths is more than 4.8 million. Epidemiology experts concur that mass testing is essential for isolating infected individuals, contact tracing, and slowing the progression of…

Cited by 0SourceScholar
2021

Diverse Semantic Image Synthesis via Probability Distribution Modeling

CVPR 2021poster

Semantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level multimodal results, still remains a challenge. In this paper,…

Cited by 86PDFcodeScholar
2021

ISNet: Integrate Image-Level and Semantic-Level Context for Semantic Segmentation

ICCV 2021poster

Co-occurrent visual pattern makes aggregating contextual information a common paradigm to enhance the pixel representation for semantic image segmentation. The existing approaches focus on modeling the context from the perspective of the whole image, i.e., aggregating the image-level contextual info…

Cited by 91PDFcodeScholar
2021

Improve Unsupervised Pretraining for Few-Label Transfer

ICCV 2021poster

Unsupervised pretraining has achieved great success and many recently works have shown unsupervised pretraining can achieve comparable or even slightly better transfer performance than supervised pretraining on downstream target datasets. But in this paper, we find this conclusion may not hold when…

Cited by 17PDFScholar
2021

Joint Color-irrelevant Consistency Learning and Identity-aware Modality Adaptation for Visible-infrared Cross Modality Person Re-identification

AAAI 2021technical

Visible-infrared cross modality person re-identification (VI-ReID) is a core but challenging technology in the 24-hours intelligent surveillance system. How to eliminate the large modality gap lies in the heart of VI-ReID. Conventional methods mainly focus on directly aligning the heterogeneous moda…

Cited by 97SourcePDFScholar
2021

Multi-Scale and Multi-Region Facial Discriminative Representation for Automatic Depression Level Prediction

ICASSP 2021accepted

Physiological studies have shown that differences in facial activities between depressed patients and normal individuals are manifested in different local facial regions and the durations of these activities are not the same. But most previous works extract features from the entire facial region at…

Cited by 0SourceScholar
2021

Multimodal Cross- and Self-Attention Network for Speech Emotion Recognition

ICASSP 2021accepted

Speech Emotion Recognition (SER) requires a thorough understanding of both the linguistic content of an utterance (i.e., textual information) and how the speaker utters it (i.e., acoustic information). The one vital challenge in SER is how to effectively fuse these two kinds of information. In this…

Cited by 0SourceScholar
2020

Cross-Modality Person Re-Identification With Shared-Specific Feature Transfer

CVPR 2020poster

Cross-modality person re-identification (cm-ReID) is a challenging but key technology for intelligent video analysis. Existing works mainly focus on learning modality-shared representation by embedding different modalities into a same feature space, lowering the upper bound of feature distinctivenes…

Cited by 434PDFScholar
2020

Density-Aware Graph for Deep Semi-Supervised Visual Recognition

CVPR 2020poster

Semi-supervised learning (SSL) has been extensively studied to improve the generalization ability of deep neural networks for visual recognition. To involve the unlabelled data, most existing SSL methods are based on common density-based cluster assumption: samples lying in the same high-density reg…

Cited by 35PDFScholar
2020

Learning distributed sentence vectors with bi-directional 3D convolutions

COLING 2020main

We propose to learn distributed sentence representation using text’s visual features as input. Different from the existing methods that render the words or characters of a sentence into images separately, we further fold these images into a 3-dimensional sentence tensor. Then, multiple 3-dimensional…

2020

Multimodal Transformer Fusion for Continuous Emotion Recognition

ICASSP 2020accepted

Multimodal fusion increases the performance of emotion recognition because of the complementarity of different modalities. Compared with decision level and feature level fusion, model level fusion makes better use of the advantages of deep neural networks. In this work, we utilize the Transformer mo…

Cited by 0SourceScholar
2020

Negative Margin Matters: Understanding Margin in Few-shot Classification

ECCV 2020poster

In this paper, we unconventionally propose to adopt appropriate negative-margin to softmax loss for few-shot classification, which surprisingly works well for the open-set scenarios of few-shot classification. We then provide the intuitive explanation and the theoretical proof to understand why nega…

2020

Parametric Instance Classification for Unsupervised Visual Feature learning

NeurIPS 2020poster

This paper presents parametric instance classification (PIC) for unsupervised visual feature learning. Unlike the state-of-the-art approaches which do instance discrimination in a dual-branch non-parametric fashion, PIC directly performs a one-branch parametric instance classification, revealing a s…

2020

Spatial-Temporal Feature Aggregation Network For Video Object Detection

ICASSP 2020accepted

Video object detection is a challenging problem in computer vision. In this paper, we propose a novel spatial-temporal feature aggregation network to deal with this issue. Specifically, we present a novel instance-level feature aggregation module as complementary to traditional pixel-level feature a…

Cited by 0SourceScholar
2019

Dynamic Ensemble Modeling Approach to Nonstationary Neural Decoding in Brain-Computer Interfaces

NeurIPS 2019poster

Brain-computer interfaces (BCIs) have enabled prosthetic device control by decoding motor movements from neural activities. Neural signals recorded from cortex exhibit nonstationary property due to abrupt noises and neuroplastic changes in brain activities during motor control. Current state-of-the-…

Cited by 26SourcePDFScholar
2019

Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword Spotting

ICASSP 2019accepted

Keyword spotting (KWS) system constitutes a critical component of human-computer interfaces, which detects the specific keyword from a continuous stream of audio. The goal of KWS is providing a high detection accuracy at a low false alarm rate while having small memory and computation requirements.…

Cited by 17SourceScholar
2018

Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training

ICASSP 2018accepted

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of…

Cited by 0SourceScholar
2018

HashGAN: Deep Learning to Hash With Pair Conditional Wasserstein GAN

CVPR 2018poster

Deep learning to hash improves image retrieval performance by end-to-end representation learning and hash coding from training data with pairwise similarity information. Subject to the scarcity of similarity information that is often expensive to collect for many application domains, existing deep l…

Cited by 131SourcePDFScholar
2017

A novel pitch extraction based on jointly trained deep BLSTM Recurrent Neural Networks with bottleneck features

ICASSP 2017accepted

Pitch is an important characteristic of speech and is useful for many applications. However, it is still challenging to estimate pitch in strong noise. In this paper, we propose a joint training approach to determinate pitch. First, a Bidirectional Long Short-Term Memory Recurrent Neural Networks (B…

Cited by 0SourceScholar
2017

Online Multi-Object Tracking Using CNN-Based Single Object Tracker With Spatial-Temporal Attention Mechanism

ICCV 2017poster

In this paper, we propose a CNN-based framework for online MOT. This framework utilizes the merits of single object trackers in adapting appearance models and searching for target in the next frame. Simply applying single object tracker for MOT will encounter the problem in computational efficiency…

Cited by 486PDFScholar
2016

Extraction of tongue contour in real-time magnetic resonance imaging sequences

ICASSP 2016accepted

Real-time magnetic resonance imaging (rtMRI) is becoming a practical tool in speech production research and language pathology observation. It is still a challenge to extract the tongue contour accurately in rtMRI sequences, since tongue is a soft tissue and often touches other organs such as lips a…

Cited by 0SourceScholar
2015

Estimate articulatory MRI series from acoustic signal using deep architecture

ICASSP 2015accepted

This paper presents our work on acoustic-to-articulatory inversion mapping, in which, the articulatory data is the MRI series for articulators on mid-sagittal plan. Deep architectures based on restricted Boltzmann machine (RBM) and linear regression are employed to construct the audio-visual mapping…

Cited by 0SourceScholar