ICASSP 2022 Accepted Papers
The full list of 1,864 papers accepted at ICASSP 2022 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- 3D Texture Super Resolution via the Rendering Loss
- 3d Cross-Scale Feature Transformer Network for Brain Mr Image Super-Resolution
- 4D Convolutional Neural Networks for Multi-Spectral and Multi-Temporal Remote Sensing Data Classification
- A Bayesian Permutation Training Deep Representation Learning Method for Speech Enhancement with Variational Autoencoder
- A Benchmark of State-of-the-Art Sound Event Detection Systems Evaluated on Synthetic Soundscapes
- A Bert Based Joint Learning Model with Feature Gated Mechanism for Spoken Language Understanding
- A Bridge between Features and Evidence for Binary Attribute-Driven Perfect Privacy
- A Byzantine-Resilient Dual Subgradient Method for Vertical Federated Learning
- A CRLB Analysis of AoA Estimation Using Bluetooth 5
- A Channel Attention Based MLP-Mixer Network for Motor Imagery Decoding With EEG
- A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction
- A Closer Look at Autoencoders for Unsupervised Anomaly Detection
- A Clustering-based ML Scheme for Capacity Approaching Soft Level Sensing in 3D TLC NAND
- A Commonsense Knowledge Enhanced Network with Retrospective Loss for Emotion Recognition in Spoken Dialog
- A Communication Efficient Quasi-Newton Method for Large-Scale Distributed Multi-Agent Optimization
- A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion
- A Complex Spectral Mapping with Inplace Convolution Recurrent Neural Networks For Acoustic Echo Cancellation
- A Configurable Multilingual Model is All You Need to Recognize All Languages
- A Convex Formulation for the Robust Estimation of Multivariate Exponential Power Models
- A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
- A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech Recording
- A Data-Driven Cognitive Salience Model for Objective Perceptual Audio Quality Assessment
- A Data-Driven Quantization Design for Distributed Testing Against Independence with Communication Constraints
- A Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation
- A Differentiable Optimisation Framework for The Design of Individualised DNN-based Hearing-Aid Strategies
- A Dilated Residual Vision Transformer for Atrial Fibrillation Detection from Stacked Time-Frequency ECG Representations
- A Domain Transfer Based Data Augmentation Method for Automated Respiratory Classification
- A Dynamic Reweighting Strategy For Fair Federated Learning
- A Fast and Efficient Network for Single Image Shadow Detection
- A Few-Sample Strategy for Guitar Tablature Transcription Based on Inharmonicity Analysis and Playability Constraints
- A Frame Loss of Multiple Instance Learning for Weakly Supervised Sound Event Detection
- A Framework for Private Communication with Secret Block Structure
- A Gaussian Mixture Model for Dialogue Generation with Dynamic Parameter Sharing Strategy
- A General Framework For Incomplete Cross-Modal Retrieval With Missing Labels And Missing Modalities
- A Generalized Hierarchical Nonnegative Tensor Decomposition
- A Generalized Kernel Risk Sensitive Loss for Robust Two-Dimensional Singular Value Decomposition
- A Generic Method to Estimate Camera Extrinsic Parameters
- A Glance-and-Gaze Network for Respiratory Sound Classification
- A Global to Local Guiding Network for Missing Data Imputation
- A Graph Attention Interactive Refine Framework with Contextual Regularization for Jointing Intent Detection and Slot Filling
- A Hybrid Approach to Combine Wireless and Earcup Microphones for ANC Headphones with Error Separation Module
- A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal Coding
- A Knowledge/Data Enhanced Method for Joint Event and Temporal Relation Extraction
- A Light Weight Model for Video Shot Occlusion Detection
- A Lightweight Instrument-Agnostic Model for Polyphonic Note Transcription and Multipitch Estimation
- A Lightweight Self-Supervised Training Framework for Monocular Depth Estimation
- A Likelihood Ratio Based Domain Adaptation Method for E2E Models
- A Low-Parametric Model for Bit-Rate Estimation of VVC Residual Coding
- A Maximal Correlation Approach to Imposing Fairness in Machine Learning
- A Melody-Unsupervision Model for Singing Voice Synthesis
- A Method For Estimating The Grouping Of Participants In Classroom Group Work Using Only Audio Information
- A Method for Detecting Coronary Artery Disease using Noisy Ultrashort Electrocardiogram Recordings
- A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter IT
- A Minimally Supervised Approach for Medical Image Quality Assessment in Domain Shift Settings
- A Model for Assessor Bias in Automatic Pronunciation Assessment
- A Multi Domain Knowledge Enhanced Matching Network for Response Selection in Retrieval-Based Dialogue Systems
- A Multi-Resolution Low-Rank Tensor Decomposition
- A Multi-Task Learning Framework for Chinese Medical Procedure Entity Normalization
- A Multi-Task Learning Method for Weakly Supervised Sound Event Detection
- A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video Enhancement
- A Multitask Learning Framework for Speaker Change Detection with Content Information from Unsupervised Speech Decomposition
- A Mutual Learning Framework for Few-Shot Sound Event Detection
- A Neural Network-based Howling Detection Method for Real-Time Communication Applications
- A Neural Prosody Encoder for End-to-End Dialogue Act Classification
- A New Coprime-Array-based Configuration with Augmented Degrees of Freedom and Reduced Mutual Coupling
- A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets
- A New Deep Learning Method for Multispectral Image Time Series Completion Using Hyperspectral Data
- A New Framework for Multiple Deep Correlation Filters Based Object Tracking
- A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech Recognition
- A Non-Convex Proximal Approach for Centroid-Based Classification
- A Non-Hierarchical Attention Network with Modality Dropout for Textual Response Generation in Multimodal Dialogue Systems
- A Nonlinear Steerable Complex Wavelet Decomposition of Images
- A Note on Totally Symmetric Equi-Isoclinic Tight Fusion Frames
- A Novel 1D State Space for Efficient Music Rhythmic Analysis
- A Novel Angular Estimation Method in the Presence of Nonuniform Noise
- A Novel Convolutional Neural Network Based on Adaptive Multi-Scale Aggregation and Boundary-Aware for Lateral Ventricle Segmentation on MR images
- A Novel Lightweight Network for Fast Monocular Depth Estimation
- A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive Networks
- A Novel Negative ℓ1 Penalty Approach for Multiuser One-Bit Massive MIMO Downlink with PSK Signaling
- A Novel Part Feature Integration and Fusion Method for Fine-Grained Vehicle Recognition
- A Novel Sequential Monte Carlo Framework for Predicting Ambiguous Emotion States
- A Novel Unsupervised Autoencoder-Based HFOs Detector in Intracranial EEG Signals
- A Performance Analysis for Multi-Ris-Assisted Full Duplex Wireless Communication System
- A Pre-Trained Audio-Visual Transformer for Emotion Recognition
- A Priori SNR Estimation for Speech Enhancement Based on PESQ-Induced Reinforcement Learning
- A Question-Oriented Propagation Network for News Reading Comprehension
- A Remedy For Distributional Shifts Through Expected Domain Translation
- A Robust Contrastive Alignment Method for Multi-Domain Text Classification
- A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature
- A Robust Object Segmentation Network for UnderWater Scenes
- A Self-Supervised Pre-Training Framework for Vision-Based Seizure Classification
- A Semi-Handcrafted Keypoint Detector with Discriminative Feature Encoding
- A Set-Theoretic Approach to Mimo Detection
- A Simple Formula for the Moments of Unitarily Invariant Matrix Distributions
- A Simple Graph Neural Network via Layer Sniffer
- A Simple Hybrid Filter Pruning for Efficient Edge Inference
- A Slide-Save Based Framework for Multi-Source DOA Extraction with Closely Spaced Sources
- A Stimuli-Relevant Directed Dependency Index for Time Series
- A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network Pruning
- A Study of The Robustness of Raw Waveform Based Speaker Embeddings Under Mismatched Conditions
- A Study on the Efficacy of Model Pre-Training In Developing Neural Text-to-Speech System
- A Style Transfer Mapping and Fine-Tuning Subject Transfer Framework Using Convolutional Neural Networks for Surface Electromyogram Pattern Recognition
- A Test for Conditional Correlation Between Random Vectors Based on Weighted U-Statistics
- A Time Domain Progressive Learning Approach with SNR Constriction for Single-Channel Speech Enhancement and Recognition
- A Time Encoding Approach to Training Spiking Neural Networks
- A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection
- A Trainable Bounded Denoiser Using Double Tight Frame Network for Snapshot Compressive Imaging
- A Training Framework for Stereo-Aware Speech Enhancement Using Deep Neural Networks
- A Transfer Learning Approach for Pronunciation Scoring
- A Two-Stage Contrastive Learning Framework For Imbalanced Aerial Scene Recognition
- A Two-Stage U-Net for High-Fidelity Denoising of Historical Recordings
- A Two-Step Approach to Leverage Contextual Data: Speech Recognition in Air-Traffic Communications
- A Two-Step Backward Compatible Fullband Speech Enhancement System
- A Two-Stream Information Fusion Approach to Abnormal Event Detection in Video
- A Unified Two-Stage Model for Separating Superimposed Images
- A Universal Ordinal Regression for Assessing Phoneme-Level Pronunciation
- A Variational Bayesian Approach to Learning Latent Variables for Acoustic Knowledge Transfer
- A Wavelet-Based Dual-Stream Network for Underwater Image Enhancement
- A free lunch from ViT: adaptive attention multi-scale fusion Transformer for fine-grained visual recognition
- A-PixelHop: A Green, Robust and Explainable Fake-Image Detector
- AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks
- ACP: Adaptive Channel Pruning for Efficient Neural Networks
- ADA-VAD: Unpaired Adversarial Domain Adaptation for Noise-Robust Voice Activity Detection
- ADD 2022: the first Audio Deep Synthesis Detection Challenge
- ADIMA: Abuse Detection In Multilingual Audio
- ADMM-DAD Net: A Deep Unfolding Network for Analysis Compressed Sensing
- ADT: Anti-Deepfake Transformer
- AECMOS: A Speech Quality Assessment Metric for Echo Impairment
- AIMNet: Adaptive Image-Tag Merging Network For Automatic Medical Report Generation
- AISHELL-NER: Named Entity Recognition from Chinese Speech
- ALSNet: A Dilated 1-D CNN for Identifying ALS from Raw EMG Signal
- APPLADE: Adjustable Plug-and-Play Audio Declipper Combining DNN with Sparse Optimization
- ARM 4-BIT PQ: SIMD-Based Acceleration for Approximate Nearest Neighbor Search on ARM
- ASR Error Correction with Dual-Channel Self-Supervised Learning
- ASR-Aware End-to-End Neural Diarization
- ASSEM-VC: Realistic Voice Conversion by Assembling Modern Speech Synthesis Techniques
- Accelerated Intravascular Ultrasound Imaging using Deep Reinforcement Learning
- Accelerating ILL-Conditioned Robust Low-Rank Tensor Regression
- Access Control for Privacy-Preserving Gaussian Process Regression
- Accurate Inference of Unseen Combinations of Multiple Rootcauses with Classifier Ensemble
- Accurate Instance Segmentation Via Collaborative Learning
- Accurate Multiscale Selective Fusion of CT and Video Images for Real-Time Endoscopic Camera 3D Tracking in Robotic Surgery
- Accurate and Resource-Efficient Lipreading with Efficientnetv2 and Transformers
- Acoustic Application of Phase Reconstruction Algorithms in Optics
- Acoustic Comparison of Physical Vocal Tract Models with Hard and Soft Walls
- Acoustic Imaging Aboard The International Space Station (ISS): Challenges and Preliminary Results
- Acoustic-to-Articulatory Inversion Based on Speech Decomposition and Auxiliary Feature
- Ada-JSR: Sample Efficient Adaptive Joint Support Recovery From Extremely Compressed Measurement Vectors
- Ada-STNet: A Dynamic AdaBoost Spatio-Temporal Network for Traffic Flow Prediction
- AdaPID: An Adaptive PID Optimizer for Training Deep Neural Networks
- Adapting Speech Separation to Real-World Meetings using Mixture Invariant Training
- Adaptive Actor-Critic Bilateral Filter
- Adaptive Attention Graph Capsule Network
- Adaptive Diffusion with Compressed Communication
- Adaptive Discounting of Implicit Language Models in RNN-Transducers
- Adaptive Group Testing with Mismatched Models
- Adaptive Identification of Underwater Acoustic Channel with a Mix of Static and Time-Varying Parameters
- Adaptive Intra-Group Aggregation for Co-Saliency Detection
- Adaptive Matching Strategy for Multi-Target Multi-Camera Tracking
- Adaptive Node Participation for Straggler-Resilient Federated Learning
- Adaptive Pseudo Labeling for Source-Free Domain Adaptation in Medical Image Segmentation
- Adaptive Variational Nonlinear Chirp Mode Decomposition
- Adaptive Weighted Network With Edge Enhancement Module For Monocular Self-Supervised Depth Estimation
- Adaptive Wireless Power Allocation with Graph Neural Networks
- AdderIC: Towards Low Computation Cost Image Compression
- Adjacency Pairs-Aware Hierarchical Attention Networks for Dialogue Intent Classification
- Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy
- AdverFacial: Privacy-Preserving Universal Adversarial Perturbation Against Facial Micro-Expression Leakages
- AdverSparse: An Adversarial Attack Framework for Deep Spatial-Temporal Graph Neural Networks
- Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator
- Adversarial Examples Detection Based on Error Level Analysis and Space Mapping
- Adversarial Examples for Image Cropping in Social Media
- Adversarial Input Ablation for Audio-Visual Learning
- Adversarial Learning Enhancement for 3D Human Pose and Shape Estimation
- Adversarial Learning in Transformer Based Neural Network in Radio Signal Classification
- Adversarial Linear Quadratic Regulator under Falsified Actions
- Adversarial Mask Transformer for Sequential Learning
- Adversarial Robustness by Design Through Analog Computing And Synthetic Gradients
- Adversarial Sample Detection for Speaker Verification by Neural Vocoders
- Adversary Distillation for One-Shot Attacks on 3D Target Tracking
- Advin: Automatically Discovering Novel Domains and Intents from User Text Utterances
- Aerial Base Station Placement Leveraging Radio Tomographic Maps
- Against Backdoor Attacks In Federated Learning With Differential Privacy
- Agcyclegan: Attention-Guided Cyclegan for Single Underwater Image Restoration
- Airborne Mimo Radar Transmit-Receive Design Under Spectral Constraint in Signal-Dependent Clutter
- Alarm Sound Detection Using Topological Signal Processing
- Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech Recognition
- All-Neural Beamformer for Continuous Speech Separation
- Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech Enhancement
- Ambiguity Modelling with Label Distribution Learning for Music Classification
- Amicable Examples for Informed Source Separation
- Amicable Examples for Informed Source Separation
- An Accelerated Rank-(L, L, 1, 1) Block Term Decomposition Of Multi-Subject Fmri Data Under Spatial Orthonormality Constraint
- An Adapter Based Pre-Training for Efficient and Scalable Self-Supervised Speech Representation Learning
- An Adaptive Orientational Beamforming Technique for Narrowband Interference Rejection
- An Anomaly Detection Method Based on Self-Supervised Learning with Soft Label Assignment for Defect Visual Inspection
- An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings
- An Asymptotically Optimal Approximation of the Conditional Mean Channel Estimator Based on Gaussian Mixture Models
- An Audio-Saliency Masking Transformer for Audio Emotion Classification in Movies
- An Effective Steganalysis for Robust Steganography with Repetitive JPEG Compression
- An Efficient DP-SGD Mechanism for Large Scale NLU Models
- An Efficient Framework for Detection and Recognition of Numerical Traffic Signs
- An Efficient Method For Generic Dsp Implementation Of Dilated Convolution
- An Efficient Method for Model Pruning Using Knowledge Distillation with Few Samples
- An Embarrassingly Simple Model for Dialogue Relation Extraction
- An End-to-End Chinese Text Normalization Model Based on Rule-Guided Flat-Lattice Transformer
- An End-to-End Deep Learning Framework For Multiple Audio Source Separation And Localization
- An End-to-End Deep Learning Speech Coding and Denoising Strategy for Cochlear Implants
- An Enhanced Deep Learning Approach for Tectonic Fault and Fracture Extraction in Very High Resolution Optical Images
- An Error Correction Scheme for Improved Air-Tissue Boundary in Real-Time MRI Video for Speech Production
- An Experimental Study on Transferring Data-Driven Image Compressive Sensing to Bioelectric Signals
- An Exploration of Hubert with Large Number of Cluster Units and Model Assessment Using Bayesian Information Criterion
- An Implicit Gradient-Type Method for Linearly Constrained Bilevel Problems
- An Information Maximization Based Blind Source Separation Approach for Dependent and Independent Sources
- An Investigation of Streaming Non-Autoregressive sequence-to-sequence Voice Conversion
- An Investigation of the Effectiveness of Phase for Audio Classification
- An Online Throughput Maximization Algorithm for Green Coordinated Multi-Point Systems
- An Overview of the FIRST ICASSP Special Session on Computer Audition for Healthcare
- Analyzing The Robustness of Unsupervised Speech Recognition
- Annihilation Filter Approach for Estimating Graph Dynamics from Diffusion Processes
- Anno-MI: A Dataset of Expert-Annotated Counselling Dialogues
- Anomalous Sound Detection Using Spectral-Temporal Information Fusion
- Applying Deep Learning to Known-Plaintext Attack on Chaotic Image Encryption Schemes
- Applying Differential Privacy to Tensor Completion
- Approaches Toward Physical and General Video Anomaly Detection
- Approximating The Likelihood Ratio in Linear-Gaussian State-Space Models for Change Detection
- Architecture for Variable Bitrate Neural Speech Codec with Configurable Computation Complexity
- Are GAN-based morphs threatening face recognition?
- Asd-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers
- Atomic Norm Based Localization and Orientation Estimation for Millimeter-Wave MIMO OFDM Systems
- Attachment Recognition in School-Age Children: A Multimodal Approach Based on Language and Paralanguage Analysis
- Attention Back-End for Automatic Speaker Verification with Multiple Enrollment Utterances
- Attention Guided Invariance Selection for Local Feature Descriptors
- Attention Probe: Vision Transformer Distillation in the Wild
- Attention-Based Dual-Stream Vision Transformer for Radar Gait Recognition
- Attention-Based Fusion for Bone-Conducted and Air-Conducted Speech Enhancement in the Complex Domain
- Attention-based Adversarial Partial Domain Adaptation
- Attentional Gated Res2net for Multivariate Time Series Classification
- Attentionpit: Soft Permutation Invariant Training for Audio Source Separation with Attention Mechanism
- Attentive Max Feature Map and Joint Training for Acoustic Scene Classification
- Attenuation Of Acoustic Early Reflections In Television Studios Using Pretrained Speech Synthesis Neural Network
- Attributable Watermarking of Speech Generative Models
- Attribute-Conditioned Face Swapping Network for Low-Resolution Images
- Audio Deepfake Detection System with Neural Stitching for ADD 2022
- Audio Peak Reduction Using a Synced allpass Filter
- Audio Signal Processing for Telepresence Based on Wearable Array in Noisy and Dynamic Scenes
- Audio-Text Retrieval in Context
- Audio-To-Symbolic Arrangement Via Cross-Modal Music Representation Learning
- Audio-Visual Multi-Channel Speech Separation, Dereverberation and Recognition
- Audio-Visual Object Classification for Human-Robot Collaboration
- Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher Learning
- Audio-Visual Tracking of Multiple Speakers Via a PMBM Filter
- Audio-Visual Wake Word Spotting System for MISP Challenge 2021
- Audioclip: Extending Clip to Image, Text and Audio
- Auditory-Based Data Augmentation for end-to-end Automatic Speech Recognition
- Augmentation Strategy Optimization for Language Understanding
- Augmenting Molecular Deep Generative Models with Topological Data Analysis Representations
- Automated Audio Captioning Using Transfer Learning and Reconstruction Latent Space Similarity Regularization
- Automated Prosody Classification for Oral Reading Fluency with Quadratic Kappa Loss and Attentive X-Vectors
- Automatic Assessment of the Degree of Clinical Depression from Speech Using X-Vectors
- Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks
- Automatic Depression Detection: an Emotional Audio-Textual Corpus and A Gru/Bilstm-Based Model
- Automatic Depression Level Assessment from Speech By Long-Term Global Information Embedding
- Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional Network
- Autoregressive Variational Autoencoder with a Hidden Semi-Markov Model-Based Structured Attention for Speech Synthesis
- AuxFormer: Robust Approach to Audiovisual Emotion Recognition
- Auxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarization
- Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive Learning
- Axonal Delay as a Short-Term Memory for Feed Forward Deep Spiking Neural Networks
- BNU: A Balance-Normalization-Uncertainty Model for Incremental Event Detection
- BSOLO: Boundary-Aware One-Stage Instance Segmentation SOLO
- Balanced Ranking and Sorting For Class Incremental Object Detection
- Balanced Stripe-Wise Pruning In The Filter
- Bayesian Continual Imputation and Prediction For Irregularly Sampled Time Series Data
- Being Greedy Does Not Hurt: Sampling Strategies for End-To-End Speech Recognition
- Best of Both Worlds: Multi-Task Audio-Visual Automatic Speech Recognition and Active Speaker Detection
- Bi-Directional Modality Fusion Network For Audio-Visual Event Localization
- Bi-Directional Normalization and Color Attention-Guided Generative Adversarial Network for Image Enhancement
- BiP-Net: Bidirectional Perspective Strategy Based Arbitrary-Shaped Text Detection Network
- Bilevel Learning of ℓ1 Regularizers with Closed-Form Gradients (BLORC)
- Bilingual End-to-End ASR with Byte-Level Subwords
- Binary Dense Predictors for Human Pose Estimation Based on Dynamic Thresholds and Filtering
- Blind Equalization of Moving Average Channels Over Galois Fields
- Blind Extraction of Equitable Partitions from Graph Signals
- Blind Modulo Analog-to-Digital Conversion of Vector Processes
- Blind Reverberation Time Estimation in Dynamic Acoustic Conditions
- Blind Separation of Linear-Quadratic Mixtures of Mutually Independent and Autocorrelated Sources
- Blind Source Separation via a Weak Exclusion Principle
- Blind Unmixing Using A Double Deep Image Prior
- Block-Activated Algorithms For Multicomponent Fully Nonsmooth Minimization
- Block-Coordinate Frank-Wolfe Algorithm And Convergence Analysis For Semi-Relaxed Optimal Transport Problem
- Block-Sparse Adversarial Attack to Fool Transformer-Based Text Classifiers
- Bloom-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech Enhancement
- Bona Fide Riesz Projections for Density Estimation
- Boost Ensemble Learning for Classification of CTG SIGNALS
- Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation Model
- Bounded Simplex-Structured Matrix Factorization
- Bounding Box Distribution Learning and Center Point Calibration for Robust Visual Tracking
- Building Robust Spoken Language Understanding by Cross Attention Between Phoneme Sequence and ASR Hypothesis
- Bundle ICP with Virtual Depth for Hand-Held 3d Scanner
- Bytecover2: Towards Dimensionality Reduction of Latent Embedding for Efficient Cover Song Identification
- Byzantine-Resilient Decentralized Collaborative Learning
- Byzantine-Resilient Decentralized Resource Allocation
- Byzantine-Robust Aggregation with Gradient Difference Compression and Stochastic Variance Reduction for Federated Learning
- Byzantine-Robust Federated Deep Deterministic Policy Gradient
- Byzantine-Robust and Communication-Efficient Distributed Non-Convex Learning Over Non-IID Data
- CDMA: Cross-Domain Distance Metric Adaptation for Speaker Verification
- CDX-NET: Cross-Domain Multi-Feature Fusion Modeling Via Deep Neural Networks for Multivariate Time Series Forecasting in AIOps
- CF-Net: Complementary Fusion Network for Rotation Invariant Point Cloud Completion
- CLIPCAM: A Simple Baseline For Zero-Shot Text-Guided Object And Action Localization
- CLseg: Contrastive Learning of Story Ending Generation
- CNN-Aided Factor Graphs with Estimated Mutual Information Features for Seizure Detection
- CNN-Transformer with Self-Attention Network for Sound Event Detection
- CPD Computation via Recursive Eigenspace Decompositions
- CPT: Cross-Modal Prefix-Tuning for Speech-To-Text Translation
- CRPN: Distinguish Novel Categories Via Class-Relevant Region Proposal Network for Few-Shot Object Detection
- CS-GResNet: A Simple and Highly Efficient Network for Facial Expression Recognition
- CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization
- CSI Clustering with Variational Autoencoding
- Cache: Modeling Contribution-Aware Context Hierarchically for Long-Range Dialogue State Tracking
- Caching Networks: Capitalizing on Common Speech for ASR
- Call-Sign Recognition and Understanding for Noisy Air-Traffic Transcripts Using Surveillance Information
- Camera Calibration Through Camera Projection Loss
- Can Audio Captions Be Evaluated With Image Caption Metrics?
- Capitalization Normalization for Language Modeling with an Accurate and Efficient Hierarchical RNN Model
- Carina - A Corpus of Aligned German Read Speech Including Annotations
- Cascade Multi-Channel Noise Reduction and Acoustic Feedback Cancellation
- Cascading Bandit Under Differential Privacy
- Category-Adapted Sound Event Enhancement with Weakly Labeled Data
- Category-Adaptive Domain Adaptation for Semantic Segmentation
- Causal Alignment Based Fault Root Causes Localization for Wireless Network
- Causal Linear Topological Filters Over A 2-Simplex
- Cell-Free Massive Mimo: Exploiting The Wax Decomposition
- Channel Redundancy and Overlap in Convolutional Neural Networks with Channel-Wise NNK Graphs
- Channel-Wise AV-Fusion Attention for Multi-Channel Audio-Visual Speech Recognition
- Characterizing the Adversarial Vulnerability of Speech self-Supervised Learning
- Chinese Spelling Text Generation of Mathematical Formulas
- Chunkfusion: A Learning-Based RGB-D 3D Reconstruction Framework Via Chunk-Wise Integration
- Classical-To-Quantum Transfer Learning for Spoken Command Recognition Based on Quantum Neural Networks
- Climate and Weather: Inspecting Depression Detection via Emotion Recognition
- Cloning One's Voice Using Very Limited Data in the Wild
- Closed-Form Single Source Direction-of-Arrival Estimator Using First-Order Relative Harmonic Coefficients
- Closing the Sim-to-Real Gap in Guided Wave Damage Detection with Adversarial Training of Variational Auto-Encoders
- Clustering Complex Subspaces in Large Dimensions
- Clustering and Separating Similarities for Deep Unsupervised Hashing
- Cmri2spec: Cine MRI Sequence to Spectrogram Synthesis via A Pairwise Heterogeneous Translator
- Co-Attention-Guided Bilinear Model for Echo-Based Depth Estimation
- Coarray Manifold Separation In The Spherical Harmonics Domain For Enhanced Source Localization
- Coarse-To-Fine Unsupervised Change Detection for Remote Sensing Images Via Object-Based MRF and Inception UNET
- Cognitive Coding Of Speech
- Collaborative Object Detectors Adaptive to Bandwidth and Computation
- Combating False Sense of Security: Breaking the Defense of Adversarial Training Via Non-Gradient Adversarial Attack
- Combining Multiple Style Transfer Networks and Transfer Learning For LGE-CMR Segmentation
- Combining Unsupervised and Text Augmented Semi-Supervised Learning For Low Resourced Autoregressive Speech Recognition
- Communication-Efficient Distributed MAX-VAR Generalized CCA via Error Feedback-Assisted Quantization
- Communication-Efficient Online Federated Learning Framework for Nonlinear Regression
- Comparison of Boundary Artifact Removal Methods in Coding of Generalized Cubemap Projection Using VVC
- Competitive Multi-Agent Reinforcement Learning with Self-Supervised Representation
- Complex IRM-Aware Training for Voice Activity Detection Using Attention Model
- Complex-Valued Spatial Autoencoders for Multichannel Speech Enhancement
- Composing Graphical Models with Generative Adversarial Networks for EEG Signal Modeling
- Compressed Data Sharing Based On Information Bottleneck Model
- Compressing Transformer-Based ASR Model by Task-Driven Loss and Attention-Based Multi-Level Feature Distillation
- Compression-Aware Projection with Greedy Dimension Reduction for Convolutional Neural Network Activations
- Compressive Phase Retrieval Based On Sparse Latent Generative Priors
- Compressive Scanning Transmission Electron Microscopy
- Computationally Efficient Fixed-Filter ANC for Speech Based on Long-Term Prediction for Headphone Applications
- Conditional Diffusion Probabilistic Model for Speech Enhancement
- Conditionally Factorized Variational Bayes with Importance Sampling
- Coneface: Approximate Pairwise Loss for Face Recognition
- Confidence Estimation for Speech Emotion Recognition Based on the Relationship Between Emotion Categories and Primitives
- Confidence-Aware Multi-Teacher Knowledge Distillation
- Conformer-Based Hybrid ASR System For Switchboard Dataset
- Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks
- Conformer-Based Speech Recognition with Linear Nyström Attention and Rotary Position Embedding
- Conjugate Augmented Spatial-Temporal Near-Field Sources Localization with Cross Array
- Connecting Targets via Latent Topics And Contrastive Learning: A Unified Framework For Robust Zero-Shot and Few-Shot Stance Detection
- Considering User Agreement in Learning to Predict the Aesthetic Quality
- Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI
- Constant Q Cepstral coefficients for classification of normal vs. Pathological infant cry
- Container Localisation and Mass Estimation with an RGB-D Camera
- Content Preserving Scale Space Network for Fast Image Restoration from Noisy-Blurry Pairs
- Context Modeling with Evidence Filter for Multiple Choice Question Answering
- Context-Adaptive Document-Level Neural Machine Translation
- Context-Aware Graph-Based Self-Supervised Learning of Whole Slide Images
- Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing
- Contextual Adapters for Personalized Speech Recognition in Neural Transducers
- Continual Learning Using Lattice-Free MMI for Speech Recognition
- Continual Self-Training With Bootstrapped Remixing For Speech Enhancement
- Continuous Speech Separation with Recurrent Selective Attention Network
- Continuous Streaming Multi-Talker ASR with Dual-Path Transducers
- Contrastive Heartbeats: Contrastive Learning for Self-Supervised ECG Representation and Phenotyping
- Contrastive Knowledge Graph Attention Network for Request-Based Recipe Recommendation
- Contrastive Prediction Strategies for Unsupervised Segmentation and Categorization of Phonemes and Words
- Contrastive Predictive Coding for Anomaly Detection of Fetal Health from the Cardiotocogram
- Contrastive Sensor Transformer for Predictive Maintenance of Industrial Assets
- Contrastive Siamese Network for Semi-Supervised Speech Recognition
- Contrastive Translation Learning For Medical Image Segmentation
- Contrastive-mixup Learning for Improved Speaker Verification
- Controllable Speech Representation Learning Via Voice Conversion and AIC Loss
- Controlled Sensing and Anomaly Detection Via Soft Actor-Critic Reinforcement Learning
- Controlling Smart Propagation Environments: Long-Term Versus Short-Term Phase Shift Optimization
- Controlling The Fréchet Variance Improves Batch Normalization on the Symmetric Positive Definite Manifold
- Conversational Speech Recognition by Learning Conversation-Level Characteristics
- Convex Clustering for Autocorrelated Time Series
- Convmixer: Feature Interactive Convolution with Curriculum Learning for Small Footprint and Noisy Far-Field Keyword Spotting
- Convoluational Transformer With Adaptive Position Embedding For Covid-19 Detection From Cough Sounds
- Convolutional Beamspace Using IIR Filters
- Convolutional Filtering in Simplicial Complexes
- Convolutional ISTA Network with Temporal Consistency Constraints for Video Reconstruction from Event Cameras
- Convolutional Weighted Minimum Mean Square Error Filter for Joint Source Separation and Dereverberation
- Coughtrigger: Earbuds IMU Based Cough Detection Activator Using An Energy-Efficient Sensitivity-Prioritized Time Series Classifier
- Counting the Number of Different Scaling Exponents in Multivariate Scale-Free Dynamics: Clustering by Bootstrap in the Wavelet Domain
- Coupled Feature Learning Via Structured Convolutional Sparse Coding for Multimodal Image Fusion
- Cramer-Rao Bound Analysis of Distributed DOA Estimation Exploiting Mixed-Precision Covariance Matrix
- Cramer-Rao Bound for the Time-Varying Poisson
- Cramér-Rao Bound and Antenna Selection Optimization for Dual Radar-Communication Design
- Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for the M2met Challenge
- Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion Segmentation
- Cross-Domain Speech Enhancement with a Neural Cascade Architecture
- Cross-Layer Aggregation with Transformers for Multi-Label Image Classification
- Cross-Modal Knowledge Distillation For Vision-To-Sensor Action Recognition
- Cross-Modal Knowledge Distillation in Multi-Modal Fake News Detection
- Cross-Speaker Style Transfer for Text-to-Speech Using Data Augmentation
- Cross-Target Stance Detection Via Refined Meta-Learning
- Csenet: Complex Squeeze-and-Excitation Network for Speech Depression Level Prediction
- Curriculum Optimization for Low-Resource Speech Recognition
- Custom Attribution Loss for Improving Generalization and Interpretability of Deepfake Detection
- Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction Loss
- Customizable End-To-End Optimization Of Online Neural Network-Supported Dereverberation For Hearing Devices
- Cut And Continuous Paste Towards Real-Time Deep Fall Detection
- Cyber-Threat Propagation over Network-Slicing Architectures
- DAM-GAN : Image Inpainting Using Dynamic Attention Map Based on Fake Texture Detection
- DCNGAN: A Deformable Convolution-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed Video
- DCSN: Deformable Convolutional Semantic Segmentation Neural Network for Non-Rigid Scenes
- DGC-Vector: A New Speaker Embedding for Zero-Shot Voice Conversion
- DHWP: Learning High-Quality Short Hash Codes Via Weight Pruning
- DMANET: Deep Learning-Based Differential Microphone Arrays for Multi-Channel Speech Separation
- DNN Based Multiframe Single-Channel Noise Reduction Filters
- DOA M-Estimation Using Sparse Bayesian Learning
- DOMAINDESC: Learning Local Descriptors With Domain Adaptation
- DP-DWA: Dual-Path Dynamic Weight Attention Network With Streaming Dfsmn-San For Automatic Speech Recognition
- DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation and Extraction
- DPT-FSNet: Dual-Path Transformer Based Full-Band and Sub-Band Fusion Network for Speech Enhancement
- DRC-NET: Densely Connected Recurrent Convolutional Neural Network for Speech Dereverberation
- DRVC: A Framework of Any-to-Any Voice Conversion with Self-Supervised Learning
- Data Agnostic Filter Gating For Efficient Deep Networks
- Data Augmentation for Long-Tailed and Imbalanced Polyphone Disambiguation in Mandarin
- Data Efficient Support Vector Machine Training Using the Minimum Description Length Principle
- Data Incubation - Synthesizing Missing Data for Handwriting Recognition
- Data Shapley Value for Handling Noisy Labels: An Application in Screening Covid-19 Pneumonia from Chest CT Scans
- Data-Driven Algorithms for Gaussian Measurement Matrix Design in Compressive Sensing
- Data-Driven Approach for the Floquet Propagator Inverse Problem Solution
- Data-Driven Optimization for Zero-Delay Lossy Source Coding with Side Information
- Data-Driven Spatially Dependent PDE Identification
- Decentralized Bilevel Optimization for Personalized Client Learning
- Decentralized Learning in the Presence of Low-Rank Noise
- Deep Actor-Critic for Continuous 3D Motion Control in Mobile Relay Beamforming Networks
- Deep Adaptation Control for Acoustic Echo Cancellation
- Deep Adaptive Aec: Hybrid of Deep Learning and Adaptive Acoustic Echo Cancellation
- Deep Augmented Music Algorithm for Data-Driven Doa Estimation
- Deep Deterministic Independent Component Analysis for Hyperspectral Unmixing
- Deep Hashing with Hash Center Update for Efficient Image Retrieval
- Deep Impulse Responses: Estimating and Parameterizing Filters with Deep Networks
- Deep Initialization for Guaranteed Unimodular Quadratic Programming
- Deep Iterative Phase Retrieval for Ptychography
- Deep Joint Source-Channel Coding for Wireless Image Transmission with Adaptive Rate Control
- Deep Kernel Learning Networks with Multiple Learning Paths
- Deep Learning Based Off-Angle Iris Recognition
- Deep Learning Based Passive Beamforming for IRS-Assisted Monostatic Backscatter Systems
- Deep Learning for Location Based Beamforming with Nlos Channels
- Deep Learning for Prominence Detection In Children's Read Speech
- Deep Learning on the Sphere for Multi-model Ensembling of Significant Wave Height
- Deep Markov Clustering for Panoptic Segmentation
- Deep Neural Network (DNN) Audio Coder Using A Perceptually Improved Training Method
- Deep Object Detection with Example Attribute Based Prediction Modulation
- Deep Performer: Score-to-Audio Music Performance Synthesis
- Deep Piecewise Hashing for Efficient Hamming Space Retrieval
- Deep Proximal Unfolding For Image Recovery from Under-Sampled Channel Data in Intravascular Ultrasound
- Deep Rank Cross-Modal Hashing with Semantic Consistent for Image-Text Retrieval
- Deep Residual Echo Suppression and Noise Reduction: A Multi-Input FCRN Approach in a Hybrid Speech Enhancement System
- Deep Scale-Aware Image Smoothing
- Deep Sequential Beamformer Learning for Multipath Channels in Mmwave Communication Systems
- Deep Spatio-Temporal Wind Power Forecasting
- Deep Temporal Interpolation of Radar-Based Precipitation
- Deep Video Inpainting Guided by Audio-Visual Self-Supervision
- Deep Video Inpainting Localization Using Spatial and Temporal Traces
- Deep-Learning-Assisted Configuration of Reconfigurable Intelligent Surfaces in Dynamic Rich-Scattering Environments
- Deep-MLE: Fusion between a Neural Network and MLE for A Single Snapshot DOA Estimation
- DeepGBASS: Deep Guided Boundary-Aware Semantic Segmentation
- DeepHull: Fast Convex Hull Approximation in High Dimensions
- Deepchorus: A Hybrid Model of Multi-Scale Convolution And Self-Attention for Chorus Detection
- Deepfake Speech Detection Through Emotion Recognition: A Semantic Approach
- Deepfilternet: A Low Complexity Speech Enhancement Framework for Full-Band Audio Based On Deep Filtering
- Defending Against Universal Attack Via Curvature-Aware Category Adversarial Training
- Deformable Convolution Dense Network for Compressed Video Quality Enhancement
- Deformable VisTR: Spatio Temporal Deformable Attention for Video Instance Segmentation
- Delay-Oriented Distributed Scheduling Using Graph Neural Networks
- Deliberation of Streaming RNN-Transducer by Non-Autoregressive Decoding
- Delta Distancing: A Lifting Approach to Localizing Items from User Comparisons
- Dementia Detection by Fusing Speech and Eye-Tracking Representation
- Demon: Improved Neural Network Training With Momentum Decay
- Denoising-Guided Deep Reinforcement Learning For Social Recommendation
- Denoising-Oriented Deep Hierarchical Reinforcement Learning for Next-Basket Recommendation⋆
- Depth Pruning with Auxiliary Networks for Tinyml
- Depth Removal Distillation for RGB-D Semantic Segmentation
- Depth-Based Ensemble Learning Network For Face Anti-Spoofing
- Deriving Explainable Discriminative Attributes Using Confusion About Counterfactual Class
- Design of Real-Time System Based on Machine Learning for Snoring and OSA Detection
- Designing a QAM Signal Detector for Massive Mimo Systems via PS-ADMM Approach
- Detail Generation and Fusion Networks for Image Inpainting
- Detecting Anomaly in Chemical Sensors via Regularized Contrastive Learning
- Detecting Backdoor Attacks against Point Cloud Classifiers
- Detection of COPD Exacerbation from Speech: Comparison of Acoustic Features and Deep Learning Based Speech Breathing Models
- Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough Audio
- Determining Joint Periodicities in Multi-Time Data with Sampling Uncertainties
- Determining the best Acoustic Features for Smoker Identification
- Deterministic Transform Based Weight Matrices for Neural Networks
- Dictionary Learning with Uniform Sparse Representations for Anomaly Detection
- Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic Sounds
- Differentiable Programming A La Moreau
- Differentiable Wavetable Synthesis
- Differentiate-and-Fire Time-Encoding of Finite-Rate-of-Innovation Signals
- Difficulty-Aware Neural Band-to-Piano Score Arrangement based on Note- and Statistic-Level Criteria
- Dilated Convolutional Neural Network-Based Deep Reference Picture Generation for Video Compression
- Direct Design of Biquad Filter Cascades with Deep Learning by Sampling Random Polynomials
- Direct Localization: An Ising Model Approach
- Direct Noisy Speech Modeling for Noisy-To-Noisy Voice Conversion
- Discourse-Level Prosody Modeling with a Variational Autoencoder for Non-Autoregressive Expressive Speech Synthesis
- Discrete Multi-Kernel K-Means with Diverse and Optimal Kernel Learning
- Disentangled Feature-Guided Multi-Exposure High Dynamic Range Imaging
- Disentangled Speaker Embedding for Robust Speaker Verification
- Disentangling Content and Fine-Grained Prosody Information Via Hybrid ASR Bottleneck Features for Voice Conversion
- Dispeech: A Synthetic Toy Dataset for Speech Disentangling
- Distilhubert: Speech Representation Learning by Layer-Wise Distillation of Hidden-Unit Bert
- Distributed Audio-Visual Parsing Based On Multimodal Transformer and Deep Joint Source Channel Coding
- Distributed Graph Learning With Smooth Data Priors
- Distributed Hybrid Beamforming for Mmwave Cell-Free Massive MIMO
- Distributed Image Transmission Using Deep Joint Source-Channel Coding
- Distributed Label Dequantized Gaussian Process Latent Variable Model for Multi-View Data Integration
- Distributed Link Sparsification for Scalable Scheduling Using Graph Neural Networks
- Distributed Particle Filters for State Tracking on the Stiefel Manifold Using Tangent Space Statistics
- Distribution Augmentation for Low-Resource Expressive Text-To-Speech
- Distribution Learning for Age Estimation from Speech
- Divergence-Guided Feature Alignment for Cross-Domain Object Detection
- Diverse Audio Captioning Via Adversarial Training
- Diversity-Controllable and Accurate Audio Captioning Based on Neural Condition
- Dnsmos P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
- Do You Live a Healthy Life? Analyzing Lifestyle by Visual Life Logging
- Doa Estimation Via Coarray Tensor Completion with Missing Slices
- Document-Level Event Extraction via Human-Like Reading Process
- Domain Adaptation for Speaker Recognition in Singing and Spoken Voice
- Domain Adaptation via Mutual Information Maximization for Handwriting Recognition
- Domain Decomposition Algorithms for Real-Time Homogeneous Diffusion Inpainting in 4K
- Domain Generalized Few-Shot Image Classification via Meta Regularization Network
- Domain Robust Deep Embedding Learning for Speaker Recognition
- Domain-Agnostic Meta-Learning for Cross-Domain Few-Shot Classification
- Domain-Invariant Feature Learning for Cross Corpus Speech Emotion Recognition
- Domain-Invariant Representation Learning from EEG with Private Encoders
- Don't Separate, Learn To Remix: End-To-End Neural Remixing With Joint Optimization
- Don't Speak Too Fast: The Impact of Data Bias on Self-Supervised Speech Models
- Double Closed-Loop Network for Image Deblurring
- Double Noise Mean Teacher Self-Ensembling Model for Semi-Supervised Tumor Segmentation
- Double-RIS Versus Single-RIS Aided Systems: Tensor-Based Mimo Channel Estimation and Design Perspectives
- Downstream Augmentation Generation For Contrastive Learning
- Dual Active Noise Control with Common Sensors
- Dual Attention Pooling Network for Recording Device Classification Using Neutral and Whispered Speech
- Dual Graph Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
- Dual Path Graph Convolutional Networks
- Dual-Attention Network for Few-Shot Segmentation
- Dual-Branch Attention-In-Attention Transformer for Single-Channel Speech Enhancement
- Dual-Domain Low-Rank Fusion Deep Metric Learning for Off-the-Person ECG Biometrics
- Duration Modeling of Neural TTS for Automatic Dubbing
- DynSNN: A Dynamic Approach to Reduce Redundancy in Spiking Neural Networks
- Dynamic Binary Neural Network by Learning Channel-Wise Thresholds
- Dynamic Multi-Scale Loss Balance for Object Detection
- Dynamic Point Cloud Interpolation
- Dynamic Portfolio Cuts: A Spectral Approach to Graph-Theoretic Diversification
- Dynamic Resource Optimization for Adaptive Federated Learning Empowered by Reconfigurable Intelligent Surfaces
- Dynamic Sliding Window for Realtime Denoising Networks
- Dynamic Texture Recognition Using PDV Hashing and Dictionary Learning on Multi-Scale Volume Local Binary Pattern
- Dynamically Pruning Segformer for Efficient Semantic Segmentation
- Dynimp: Dynamic Imputation for Wearable Sensing Data through Sensory and Temporal Relatedness
- Dysfluency Classification in Stuttered Speech Using Deep Learning for Real-Time Applications
- EAD-Conformer: a Conformer-Based Encoder-Attention-Decoder-Network for Multi-Task Audio Source Separation
- EMGSE: Acoustic/EMG Fusion for Multimodal Speech Enhancement
- EMOQ-TTS: Emotion Intensity Quantization for Fine-Grained Controllable Emotional Text-to-Speech
- ER-PIQA: A Task-Guided Pedestrian Image Quality Assessment Via Embedding Reconstruction
- ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnet
- Echo-Aware Adaptation of Sound Event Localization and Detection in Unknown Environments
- Eco-Fedsplit: Federated Learning with Error-Compensated Compression
- Economics of Semantic Communication System in Wireless Powered Internet of Things
- Edge Sampling of Graphs Based on Edge Smoothness
- Effect of Noise Suppression Losses on Speech Distortion and ASR Performance
- Effective and Inconspicuous Over-the-Air Adversarial Examples with Adaptive Filtering
- Efficient Adapter Transfer of Self-Supervised Speech Models for Automatic Speech Recognition
- Efficient Identity-Based Chameleon Hash for Mobile Devices
- Efficient Monaural Speech Separation with Multiscale Time-Delay Sampling
- Efficient Sequence Training of Attention Models Using Approximative Recombination
- Efficient Two-Stage Beam Training and Channel Estimation for Ris-Aided Mmwave Systems Via Fast Alternating Least Squares
- Efficient Universal Shuffle Attack for Visual Object Tracking
- Efficient and Stable Information Directed Exploration for Continuous Reinforcement Learning
- Efficiently and Globally Solving Joint Beamforming and Compression Problem in the Cooperative Cellular Network Via Lagrangian Duality
- Embedding Signals on Graphs with Unbalanced Diffusion Earth Mover's Distance
- Embedding and Beamforming: All-Neural Causal Beamformer for Multichannel Speech Enhancement
- Emotionflow: Capture the Dialogue Level Emotion Transitions
- Enabling On-Device Training of Speech Recognition Models With Federated Dropout
- Encrypted Image Visual Security Index via Non-Local Recognizable Degree Evaluation
- Encryption Resistant Deep Neural Network Watermarking
- End-To-End Alexa Device Arbitration
- End-To-End Deep Learning-Based Adaptation Control for Frequency-Domain Adaptive System Identification
- End-To-End Multi-Modal Speech Recognition with Air and Bone Conducted Speech
- End-To-End Music Remastering System Using Self-Supervised And Adversarial Training
- End-To-End Neural Coreference Resolution Revisited: A Simple Yet Effective Baseline
- End-To-End Speech Recognition with Joint Dereverberation of Sub-Band Autoregressive Envelopes
- End-to-End ASR-Enhanced Neural Network for Alzheimer's Disease Diagnosis
- End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression
- End-to-End Keyword Spotting Using Neural Architecture Search and Quantization
- End-to-End Low Resource Keyword Spotting Through Character Recognition and Beam-Search Re-Scoring
- End-to-End Network Based on Transformer for Automatic Detection of Covid-19
- End-to-End Neural Speech Coding for Real-Time Communications
- End-to-End Speech Recognition from Federated Acoustic Models
- End-to-End Speech Summarization Using Restricted Self-Attention
- Endpoint Detection for Streaming End-to-End Multi-Talker ASR
- Energy Alignment for Bias Rectification in Class Incremental Learning
- Enhance Rnnlms with Hierarchical Multi-Task Learning for ASR
- Enhancing Affective Representations Of Music-Induced Eeg Through Multimodal Supervision And Latent Domain Adaptation
- Enhancing Class Understanding Via Prompt-Tuning For Zero-Shot Text Classification
- Enhancing Contextual Encoding With Stage-Confusion and Stage-Transition Estimation for EEG-Based Sleep Staging
- Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation Generation
- Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion Recognition
- Enhancing Prototypical Few-Shot Learning By Leveraging The Local-Level Strategy
- Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-Based Multi-Modal Context Modeling
- Enhancing Utility In The Watchdog Privacy Mechanism
- Enhancing and Dissecting Crowd Counting by Synthetic Data
- Enrich Features for Few-Shot Point Cloud Classification
- Entrainment Analysis for Assessment of Autistic Speech Prosody Using Bottleneck Features of Deep Neural Network
- Environmental Sound Extraction Using Onomatopoeic Words
- Epileptic Spike Detection by Recurrent Neural Networks with Self-Attention Mechanism
- Equal Loss: A Simple Loss Function for Noise Robust Learning
- Estimating the Confidence of Speech Spoofing Countermeasure
- Estimation Of Channels In Systems With Intelligent Reflecting Surfaces
- Estimation of the Admittance Matrix in Power Systems Under Laplacian and Physical Constraints
- Evaluation of Orthogonal Chirp Division Multiplexing for Automotive Integrated Sensing and Communications
- Evaluation of Video Coding for Machines without Ground Truth
- Event-Based Multimodal Spiking Neural Network with Attention Mechanism
- Evolutionary Neural Architecture Design of Liquid State Machine for Image Classification
- Exact Partitioning of High-Order Planted Models with A Tensor Nuclear Norm Constraint
- Exact Sparse Super-Resolution Via Model Aggregation
- Expectation Consistent Plug-and-Play for MRI
- Experimental Investigation on STFT Phase Representations for Deep Learning-Based Dysarthric Speech Detection
- Experts Versus All-Rounders: Target Language Extraction for Multiple Target Languages
- Explainable Artificial Intelligence for Authorship Attribution on Social Media
- Explainable Fact-Checking Through Question Answering
- Explaining Deep Learning Models for Spoofing and Deepfake Detection with Shapley Additive Explanations
- Explicitly Modeling Importance and Coherence for Timeline Summarization
- Exploiting Annotators' Typed Description of Emotion Perception to Maximize Utilization of Ratings for Speech Emotion Recognition
- Exploiting Caption Diversity for Unsupervised Video Summarization
- Exploiting Cross Domain Acoustic-to-Articulatory Inverted Features for Disordered Speech Recognition
- Exploiting Hybrid Models of Tensor-Train Networks For Spoken Command Recognition
- Exploiting Language Model For Efficient Linguistic Steganalysis
- Explore Relative and Context Information with Transformer for Joint Acoustic Echo Cancellation and Speech Enhancement
- Exploring Auditory Acoustic Features for The Diagnosis of Covid-19
- Exploring Category Consistency for Weakly Supervised Semantic Segmentation
- Exploring Complementarity of Global and Local Spatiotemporal Information for Fake Face Video Detection
- Exploring Deeper Graph Convolutions for Semi-Supervised Node Classification
- Exploring Dementia Detection from Speech: Cross Corpus Analysis
- Exploring Dual Stream Global Information For Image Captioning
- Exploring Effective Data Utilization for Low-Resource Speech Recognition
- Exploring Heterogeneous Characteristics of Layers in ASR Models for More Efficient Training
- Exploring Machine Speech Chain For Domain Adaptation
- Exploring Non-Autoregressive End-to-End Neural Modeling for English Mispronunciation Detection and Diagnosis
- Exploring Transferability Measures and Domain Selection in Cross-Domain Slot Filling
- Exploring Transformer's Potential on Automatic Piano Transcription
- Exploring the Effect of ℓ0/ℓ2 Regularization in Neural Network Pruning using the LC Toolkit
- Extended Graph Temporal Classification for Multi-Speaker End-to-End ASR
- Extending the Use of MDL for High-Dimensional Problems: Variable Selection, Robust Fitting, and Additive Modeling
- Extracting and Distilling Direction-Adaptive Knowledge for Lightweight Object Detection in Remote Sensing Images
- Extreme-Point Pursuit for Unit-Modulus Optimization
- Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated Faces
- FAZ-BV: A Diabetic Macular Ischemia Grading Framework Combining Faz Attention Network and Blood Vessel Enhancement Filters
- FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network
- FDSNeT: An Accurate Real-Time Surface Defect Segmentation Network
- FINT: Field-Aware Interaction Neural Network for Click-Through Rate Prediction
- FOV-Based Coding Optimization for 360-Degree Virtual Reality Videos
- FRCRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement
- FRE-GAN 2: Fast and Efficient Frequency-Consistent Audio Synthesis
- FSM: Feature Sampling Module for Object Detection
- FSOINET: Feature-Space Optimization-Inspired Network For Image Compressive Sensing
- Factorized Neural Transducer for Efficient Language Model Adaptation
- Fairness-Aware Selective Sampling on Attributed Graphs
- Fake Audio Detection Based On Unsupervised Pretraining Models
- Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech Recognition
- Fast Fault Diagnosis Method Of Rolling Bearings In Multi-Sensor Measurement Enviroment
- Fast Graph Sampling for Short Video Summarization Using Gershgorin Disc Alignment
- Fast Learning of Fast Transforms, with Guarantees
- Fast Low Rank Column-Wise Compressive Sensing For Accelerated Dynamic MRI
- Fast Multiscale Diffusion On Graphs
- Fast Task-Specific Adaptation in Spoken Language Assessment with Meta-Learning
- Fast Video Object Segmentation via Dynamic YOLACT
- Fast and Stable Convergence of Online SGD for CV@R-Based Risk-Aware Learning
- Fast-Rir: Fast Neural Diffuse Room Impulse Response Generator
- Fast-Slow Transformer for Visually Grounding Speech
- FastAudio: A Learnable Audio Front-End For Spoof Speech Detection
- Feature Augmentation Learning for Few-Shot Palmprint Image Recognition With Unconstrained Acquisition
- Feature Imitating Networks
- Feature Space Message Passing Network for Medical Image Semantic Segmentation
- Feature-Based Sensing Matrix Design for Analog to Information Converters
- FedClean: A Defense Mechanism against Parameter Poisoning Attacks in Federated Learning
- Federated Learning Challenges and Opportunities: An Outlook
- Federated Multi-Armed Bandit Via Uncoordinated Exploration
- Federated Over-Air Robust Subspace Tracking from Missing Data
- Federated Self-Supervised Learning for Acoustic Event Classification
- Federated Self-Training for Data-Efficient Audio Recognition
- Federated Stochastic Gradient Descent Begets Self-Induced Momentum
- Few-Shot Gaze Estimation with Model Offset Predictors
- Few-Shot Generation By Modeling Stereoscopic Priors
- Few-Shot Learning with Improved Local Representations via Bias Rectify Module
- Few-Shot Musical Source Separation
- Few-Shot Object Detection with Local Correspondence RPN and Attentive Head
- Few-Shot One-Class Domain Adaptation Based On Frequency For Iris Presentation Attack Detection
- Filteraugment: An Acoustic Environmental Data Augmentation Method
- Find The Way Back: Invertible Kernel Estimator For Blind Image Super-Resolution
- Fine-Grained Dynamic Loss for Accurate Single-Image Super-Resolution
- Fine-Grained Style Control In Transformer-Based Text-To-Speech Synthesis
- Fine-Tuning Wav2Vec2 for Speaker Recognition
- Fldp: Flexible Strategy For Local Differential Privacy
- Floor Plan Reconstruction with High-Precision Rf-Based Tracking
- Flow-Based Fast Multichannel Nonnegative Matrix Factorization for Blind Source Separation
- Flow-Based Point Cloud Completion Network with Adversarial Refinement
- FlowDT: A Flow-Aware Digital Twin for Computer Networks
- Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using Transformers
- Fostering The Robustness Of White-Box Deep Neural Network Watermarks By Neuron Alignment
- Fracture Detection and Localization in Chest X-Rays Using Semi-Supervised Learning with Dynamic Sharpening
- Fraug: A Frame Rate Based Data Augmentation Method for Depression Detection from Speech Signals
- Free Lunch for Cross-Domain Occluded Face Recognition without Source Data
- Frequency-Specific Non-Linear Granger Causality in a Network of Brain Signals
- From Bottom-Up To Top-Down: Characterization Of Training Process In Gaze Modeling
- From Shallow to Deep: Compositional Reasoning over Graphs for Visual Question Answering
- Frontend Attributes Disentanglement for Speech Emotion Recognition
- FullSubNet+: Channel Attention Fullsubnet with Complex Spectrograms for Speech Enhancement
- Fusing ASR Outputs in Joint Training for Speech Emotion Recognition
- Fusion and Orthogonal Projection for Improved Face-Voice Association
- Fusion of Modulation Spectral and Spectral Features with Symptom Metadata for Improved Speech-Based Covid-19 Detection
- Fusion-Id: A Photoplethysmography and Motion Sensor Fusion Biometric Authenticator With Few-Shot on-Boarding
- GAZEATTENTIONNET: Gaze Estimation with Attentions
- GOS: A Large-Scale Annotated Outdoor Scene Synthetic Dataset
- GPU-Accelerated Forward-Backward Algorithm with Application to Lattice-Free MMI
- Gan-Based Joint Activity Detection and Channel Estimation for Grant-Free Random Access
- Ganet: Unary Attention Reaches Pairwise Attention Via Implicit Group Clustering in Light-Weight CNNs
- Gated Multimodal Fusion with Contrastive Learning for Turn-Taking Prediction in Human-Robot Dialogue
- Generalization Ability of MOS Prediction Networks
- Generalized Autocorrelation Analysis for Multi-Target Detection
- Generalized Face Anti-Spoofing via Cross-Adversarial Disentanglement with Mixing Augmentation
- Generalized Matching Pursuits for the Sparse Optimization of Separable Objectives
- Generalized Sliced Probability Metrics
- Generalized Time Domain Velocity Vector
- Generalized Zero-Shot Learning Using Conditional Wasserstein Autoencoder
- Generating Disentangled Arguments with Prompts: A Simple Event Extraction Framework That Works
- Generation for Unsupervised Domain Adaptation: A Gan-Based Approach for Object Classification with 3D Point Cloud Data
- Generation of Personal Sound Fields in Reverberant Environments Using Interframe Correlation
- Generative Adversarial Network Including Referring Image Segmentation For Text-Guided Image Manipulation
- Genre-Conditioned Acoustic Models for Automatic Lyrics Transcription of Polyphonic Music
- Genre-Conditioned Long-Term 3D Dance Generation Driven by Music
- Geometric Low-Rank Tensor Approximation for Remotely Sensed Hyperspectral And Multispectral Imagery Fusion
- Glassoformer: A Query-Sparse Transformer for Post-Fault Power Grid Voltage Prediction
- Global Evolution Neural Network for Segmentation of Remote Sensing Images
- Global Optimization Solution for Dynamic Adaptive 360-Degree Streaming
- Global-Local Feature Enhancement Network for Robust Object Detection using mmWave Radar and Camera
- Goal-Oriented Communication for Edge Learning Based On the Information Bottleneck
- Gradient Staleness in Asynchronous Optimization Under Random Communication Delays
- Gradient Variance Loss for Structure-Enhanced Image Super-Resolution
- Gradient-Weighted Class Activation Mapping for Spatio Temporal Graph Convolutional Network
- Gradual Surrogate Gradient Learning in Deep Spiking Neural Networks
- Graph Attentive Feature Aggregation for Text-Independent Speaker Verification
- Graph Convolution for Re-Ranking in Person Re-Identification
- Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting Data
- Graph Convolutional Networks With Autoencoder-Based Compression And Multi-Layer Graph Learning
- Graph Fine-Grained Contrastive Representation Learning
- Graph Learning Based Autoencoder for Hyperspectral Band Selection
- Graph Learning From Multivariate Dependent Time Series Via A Multi-Attribute Formulation
- Graph Learning Information Criterion
- Graph-Based Point Cloud Denoising Using Shape-Aware Consistency For Free-Viewpoint Video
- Graph-Structured Sparse Regularization Via Convex Optimization
- Graphon-Aided Joint Estimation of Multiple Graphs
- Grassmannian Dimensionality Reduction Using Triplet Margin Loss for Ume Classification of 3d Point Clouds
- Gridless DOA Estimation Under the Multi-Frequency Model
- Group-Wise Feature Selection for Supervised Learning
- HBP: An Efficient Block Permutation Solver Using Hungarian Algorithm and Spectrogram Inpainting for Multichannel Audio Source Separation
- HGCN: Harmonic Gated Compensation Network for Speech Enhancement
- HIRL: Hybrid Image Restoration Based on Hierarchical Deep Reinforcement Learning via Two-Step Analysis
- HOQRI: Higher-Order QR Iteration for Scalable Tucker Decomposition
- HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection
- Half Inverted Nested Arrays with Large Hole-Free Fourth-Order Difference Co-Arrays
- Hand Gesture Recognition Using Temporal Convolutions and Attention Mechanism
- Harmonic Gated Compensation Network Plus for ICASSP 2022 DNS Challenge
- Harmonic and Percussive Sound Separation Based on Mixed Partial Derivative of Phase Spectrogram
- Harmonicity Plays a Critical Role in DNN Based Versus in Biologically-Inspired Monaural Speech Segregation Systems
- Harvesting Partially-Disjoint Time-Frequency Information for Improving Degenerate Unmixing Estimation Technique
- Have Best of Both Worlds: Two-Pass Hybrid and E2E Cascading Framework for Speech Recognition
- Heart Rate and Oxygen Saturation Estimation from Facial Video with Multimodal Physiological Data Generation
- Heterogeneous Graph Node Classification With Multi-Hops Relation Features
- Heuristic Dropout: An Efficient Regularization Method for Medical Image Segmentation Models
- HiFi-SVC: Fast High Fidelity Cross-Domain Singing Voice Conversion
- HiFiDenoise: High-Fidelity Denoising Text to Speech with Adversarial Networks
- Hierarchical Classification of Singing Activity, Gender, and Type in Complex Music Recordings
- Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword Units
- Hierarchical Deep Learning Model with Inertial and Physiological Sensors Fusion for Wearable-Based Human Activity Recognition
- Hierarchical Feature Aggregation Network for Deep Image Compression
- Hierarchical Graph-Based Neural Network for Singing Melody Extraction
- Hierarchical Prosody Modeling and Control in Non-Autoregressive Parallel Neural TTS
- Hierarchical Signal Fusion Network for Pulsar Detection with Phase-Correlation and Signal Attentions
- Hierarchical and Multi-View Dependency Modelling Network for Conversational Emotion Recognition
- High-Dimensional Sparse Bayesian Learning without Covariance Matrices
- High-Fidelity Portrait Editing Via Exploring Differentiable Guided Sketches from the Latent Space
- High-Quality Self-Supervised Snapshot Hyperspectral Imaging
- Histogram-Guided Semantic-Aware Colorization
- Histokt: Cross Knowledge Transfer in Computational Pathology
- Hodgelets: Localized Spectral Representations of Flows On Simplicial Complexes
- Holistic Semi-Supervised Approaches for EEG Representation Learning
- How Can a Cognitive Radar Mask its Cognition?
- How Neural Processes Improve Graph Link Prediction
- How Secure Are The Adversarial Examples Themselves?
- Human Decision Making with Bounded Rationality
- Human Emotion Recognition Using Multi-Modal Biological Signals Based On Time Lag-Considered Correlation Maximization
- Hybrid Attention-Based Prototypical Networks for Few-Shot Sound Classification
- Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration
- Hybrid Weighting Loss for Precipitation Nowcasting from Radar Images
- Hybrid sub-word segmentation for handling long tail in morphologically rich low resource languages
- Hypergraph-Based Reinforcement Learning for Stock Portfolio Selection
- Hypergraphs with Edge-Dependent Vertex Weights: Spectral Clustering Based on the 1-Laplacian
- Hyperspectral Image Classification Based on Co-Learning Through Dual-Architecture Ensemble
- Hyperspectral Image Super-Resolution with Deep Priors and Degradation Model Inversion
- ICASSP 2022 Acoustic Echo Cancellation Challenge
- ICASSP 2022 L3DAS22 Challenge: Ensemble of Resnet-Conformers with Ambisonics Data Augmentation for Sound Event Localization and Detection
- ICASSP-SPGC 2022: Root Cause Analysis for Wireless Network Fault Localization
- IMPQ: Reduced Complexity Neural Networks Via Granular Precision Assignment
- ISDA: Position-Aware Instance Segmentation with Deformable Attention
- ISOMETRIC MT: Neural Machine Translation for Automatic Dubbing
- ISTFTNET: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform
- Icassp 2022 Deep Noise Suppression Challenge
- Identification of Pulse Streams Of Unknown Shape From Time Encoding Machine Samples
- Image Denoising with Deep Unfolding And Normalizing Flows
- Image Steganalysis with Convolutional Vision Transformer
- Image-Text Alignment and Retrieval Using Light-Weight Transformer
- Image-to-Graph Transformers for Chemical Structure Recognition
- Image-to-Video Re-Identification via Mutual Discriminative Knowledge Transfer
- Importance Sampling Cams For Weakly-Supervised Segmentation
- Importance of Switch Optimization Criterion in Switching WPE Dereverberation
- Importantaug: A Data Augmentation Agent for Speech
- Improve Few-Shot Voice Cloning Using Multi-Modal Learning
- Improve Image Captioning Via Relation Modeling
- Improved Beamforming Encoding for Joint Radar and Communication
- Improved Language Identification Through Cross-Lingual Self-Supervised Learning
- Improved Meta Learning for Low Resource Speech Recognition
- Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology
- Improved Simulation of Realistically-Spatialised Simultaneous Speech Using Multi-Camera Analysis in The Chime-5 Dataset
- Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing
- Improving Actor-Critic Reinforcement Learning Via Hamiltonian Monte Carlo Method
- Improving Adversarial Waveform Generation Based Singing Voice Conversion with Harmonic Signals
- Improving Anomaly Detection with a Self-Supervised Task Based on Generative Adversarial Network
- Improving BCI-based Color Vision Assessment Using Gaussian Process Regression
- Improving Biomedical Named Entity Recognition with a Unified Multi-Task MRC Framework
- Improving Bird Classification with Unsupervised Sound Separation
- Improving Brain Decoding Methods and Evaluation
- Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models
- Improving Character Error Rate is Not Equal to Having Clean Speech: Speech Enhancement for ASR Systems with Black-Box Acoustic Models
- Improving Class Activation Map for Weakly Supervised Object Localization
- Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition
- Improving Contextual Coherence in Variational Personalized and Empathetic Dialogue Agents
- Improving Cross-Lingual Speech Synthesis with Triplet Training Scheme
- Improving Cross-Modal Understanding in Visual Dialog Via Contrastive Learning
- Improving Dialogue Generation via Proactively Querying Grounded Knowledge
- Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head Attention
- Improving Dynamic Graph Convolutional Network with Fine-Grained Attention Mechanism
- Improving Emotional Speech Synthesis by Using SUS-Constrained VAE and Text Encoder Aggregation
- Improving End-To-End Speech Translation Model with Bert-Based Contextual Information
- Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection
- Improving End-to-end Models for Set Prediction in Spoken Language Understanding
- Improving Factored Hybrid HMM Acoustic Modeling without State Tying
- Improving Fairness in Speaker Verification via Group-Adapted Fusion Network
- Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward Network
- Improving Feature Generalizability with Multitask Learning in Class Incremental Learning
- Improving Generalization of Deep Networks for Estimating Physical Properties of Containers and Fillings
- Improving Inference for Spatial Signals by Contextual False Discovery Rates
- Improving Joint Sparse Hyperspectral Unmixing by Simultaneously Clustering Pixels According To Their Mixtures
- Improving Lyrics Alignment Through Joint Pitch Detection
- Improving Maximum Likelihood Difference Scaling Method To Measure Inter Content Scale
- Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction
- Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language Models
- Improving Phase-Rectified Signal Averaging for Fetal Heart Rate Analysis
- Improving Phonetic Realizations in its by Using Phoneme-Aligned Graphemes
- Improving Pseudo-Label Training For End-To-End Speech Recognition Using Gradient Mask
- Improving Recognition-Synthesis Based any-to-one Voice Conversion with Cyclic Training
- Improving Reference-Based Image Colorization For Line Arts Via Feature Aggregation And Contrastive Learning
- Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer Supervision
- Improving Separation-Based Speaker Diarization Via Iterative Model Refinement And Speaker Embedding Based Post-Processing
- Improving Source Separation by Explicitly Modeling Dependencies between Sources
- Improving Spoken Language Understanding by Enhancing Text Representation
- Improving The Latency And Quality Of Cascaded Encoders
- Improving Ultrasound Image Classification with Local Texture Quantisation
- Improving the Classification of Phonetic Segments from Raw Ultrasound Using Self-Supervised Learning and Hard Example Mining
- Improving the Fusion of Acoustic and Text Representations in RNN-T
- In Pursuit of Preserving the Fidelity of Adversarial Images
- Incipient Fault Severity Estimation Using Local Mahalanobis Distance
- Incoherent Synthesis of Sparse Broadband Arrays based on a Parameter-Free Subspace Clustering
- Incorporating End-to-End Framework Into Target-Speaker Voice Activity Detection
- Incorporating Gaze Behavior Using Joint Embedding With Scene Context for Driver Takeover Detection
- Increasing Loudness in Audio Signals: A Perceptually Motivated Approach to Preserve Audio Quality
- Incremental Context Aware Attentive Knowledge Tracing
- Incremental User Embedding Modeling for Personalized Text Classification
- Independent Vector Analysis Based Subgroup Identification from Multisubject fMRI Data
- Individualized Hear-Through For Acoustic Transparency Using PCA-Based Sound Pressure Estimation At The Eardrum
- Infant Crying Detection In Real-World Environments
- Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
- Inferring Camera Intrinsics Based on Surfaces of Revolution: A Single Image Geometric Network Approach for Camera Calibration
- Information Theoretic Limits For Standard and One-Bit Compressed Sensing with Graph-Structured Sparsity
- Informative Attention Supervision for Grounded Video Description
- Initialization-Free Implicit-Focusing (IF2) for Wideband Direction-of-Arrival Estimation
- Injecting Text and Cross-Lingual Supervision in Few-Shot Learning from Self-Supervised Models
- Instantaneous Linear Dimensionality Reduction of Multichannel Time-Series Signal for Array Signal Processing
- Integer-Only Zero-Shot Quantization for Efficient Speech Recognition
- Integrated Sensing and Communications Via 5G NR Waveform: Performance Analysis
- Integrating Dependency Tree into Self-Attention for Sentence Representation
- Integrating Multiple ASR Systems into NLP Backend with Attention Fusion
- Integrating Pretrained Language Model for Dialogue Policy Evaluation
- Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement
- Integrating Text Inputs for Training and Adapting RNN Transducer ASR Models
- Integration of Anomaly Machine Sound Detection into Active Noise Control to Shape the Residual Sound
- Integration of Pre-Trained Networks with Continuous Token Interface for End-to-End Spoken Language Understanding
- Intelligent Wi-Fi Based Child Presence Detection System
- Interactive Feature Fusion for End-to-End Noise-Robust Speech Recognition
- Interactive Multi-Level Prosody Control for Expressive Speech Synthesis
- Intermix: An Interference-Based Data Augmentation and Regularization Technique for Automatic Deep Sound Classification
- Internet Streaming Audio Based Speech Reception Threshold Measurement in Cochlear Implant Users
- Interpretable Image Classification Using Sparse Oblique Decision Trees
- Interpreting Intermediate Convolutional Layers In Unsupervised Acoustic Word Classification
- Inverse Imaging with Generative Priors Via Langevin Dynamics
- Investigating Robustness of Biological vs. Backprop Based Learning
- Investigating Self-Supervised Learning for Speech Enhancement and Separation
- Investigating Sequence-Level Normalisation For CTC-Like End-to-End ASR
- Investigating the Potential of Auxiliary-Classifier Gans for Image Classification in Low Data Regimes
- Investigation And Comparison of Optimization Methods for Variational Autoencoder-Based Underdetermined Multichannel Source Separation
- Investigation of Robustness of Hubert Features from Different Layers to Domain, Accent and Language Variations
- Invisible and Efficient Backdoor Attacks for Compressed Deep Neural Networks
- Is Cross-Attention Preferable to Self-Attention for Multi-Modal Emotion Recognition?
- Iterative Channel Estimation and Data Detection Algorithm For OTFS Modulation
- Iterative Learning for Distorted Image Restoration
- Iterative Re-weighted Least Squares Algorithms for Non-negative Sparse and Group-sparse Recovery
- Iterative Self Knowledge Distillation - from Pothole Classification to Fine-Grained and Covid Recognition
- ItôWave: Itô Stochastic Differential Equation is all You Need for Wave Generation
- JE2NET: Joint Exploitation and Exploration in Reinforcement Learning Based Image Restoration
- Jmpnet: Joint Motion Prediction for Learning-Based Video Compression
- Joint Beam Selection and Precoding Based on Differential Evolution for Millimeter-Wave Massive MIMO Systems
- Joint Calibration and Mapping of Satellite Altimetry Data Using Trainable Variational Models
- Joint Centrality Estimation and Graph Identification from Mixture of Low Pass Graph Signals
- Joint Dual-Domain Matrix Factorization for ECG Biometric Recognition
- Joint Ego-Noise Suppression and Keyword Spotting on Sweeping Robots
- Joint Far- and Near-End Speech Intelligibility Enhancement Based on the Approximated Speech Intelligibility Index
- Joint Global-Local Alignment for Domain Adaptive Semantic Segmentation
- Joint Hypoglycemia Prediction and Glucose Forecasting via Deep Multi-Task Learning
- Joint Inference of Multiple Graphs with Hidden Variables from Stationary Graph Signals
- Joint Learning for Addressee Selection and Response Generation in Multi-Party Conversation
- Joint Learning of Feature Extraction and Cost Aggregation for Semantic Correspondence
- Joint Magnitude Estimation and Phase Recovery Using Cycle-In-Cycle GAN for Non-Parallel Speech Enhancement
- Joint Model Order Estimation for Multiple Tensors with A Coupled Mode and Applications to the Joint Decomposition of EEG, MEG Magnetometer, and Gradiometer Tensors
- Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
- Joint Multiple Intent Detection and Slot Filling Via Self-Distillation
- Joint Normality Test Via Two-Dimensional Projection
- Joint Radar-Communications Processing from A Dual-Blind Deconvolution Perspective
- Joint Source Localization and Association Through Overcomplete Representation Under Multipath Propagation Environment
- Joint Speech Recognition and Audio Captioning
- Joint Temporal Convolutional Networks and Adversarial Discriminative Domain Adaptation for EEG-Based Cross-Subject Emotion Recognition
- Joint Unsupervised and Supervised Training for Multilingual ASR
- Joint and Adversarial Training with ASR for Expressive Speech Synthesis
- K-Converter: An Unsupervised Singing Voice Conversion System
- KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE Using Mel-Spectrograms
- Kernel Estimation Network for Blind Super-Resolution
- Key-Sparse Transformer for Multimodal Speech Emotion Recognition
- Knowledge Augmented Bert Mutual Network in Multi-Turn Spoken Dialogues
- Knowledge Distillation for Neural Transducers from Large Self-Supervised Pre-Trained Models
- Knowledge Distillation from Language Model to Acoustic Model: A Hierarchical Multi-Task Learning Approach
- Knowledge Transfer from Large-Scale Pretrained Language Models to End-To-End Speech Recognizers
- L-SpEx: Localized Target Speaker Extraction
- L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment
- LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech
- LERPS: Lighting Estimation and Relighting for Photometric Stereo
- LETR: A Lightweight and Efficient Transformer for Keyword Spotting
- LIGHT-SERNET: A Lightweight Fully Convolutional Neural Network for Speech Emotion Recognition
- LMS and NLMS Algorithms for the Identification of Impulse Responses with Intrinsic Symmetric or Antisymmetric Properties
- LPC Augment: an LPC-based ASR Data Augmentation Algorithm for Low and Zero-Resource Children's Dialects
- LRPD: Large Replay Parallel Dataset
- Label Propagation Across Graphs: Node Classification Using Graph Neural Tangent Kernels
- Label-Aware Ranked Loss for Robust People Counting Using Automotive In-Cabin Radar
- Label-Occurrence-Balanced Mixup for Long-Tailed Recognition
- Language Adaptive Cross-Lingual Speech Representation Learning with Sparse Sharing Sub-Networks
- Large-Scale ASR Domain Adaptation Using Self- and Semi-Supervised Learning
- Large-Scale Independent Component Analysis By Speeding Up Lie Group Techniques
- Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
- Latent Space Slicing for Enhanced Entropy Modeling In Learning-Based Point Cloud Geometry Compression
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.