ICASSP 2021 Accepted Papers
The full list of 1,712 papers accepted at ICASSP 2021 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- "You Should Probably Read This": Hedge Detection in Text
- (W)Earable Microphone Array and Ultrasonic Echo Localization for Coarse Indoor Environment Mapping
- 2D-FRFT Based Frequency Shift-Invariant Digital Image Encryption
- 3D Multizone Soundfield Reproduction in a Reverberant Environment Using Intensity Matching Method
- A Bayesian Inference Approach for Location-Based Micro Motions using Radio Frequency Sensing
- A Bayesian Interpretation of the Light Gated Recurrent Unit
- A Better and Faster end-to-end Model for Streaming ASR
- A Bias-Reducing Loss Function for CT Image Denoising
- A Capsule Network Based Approach for Detection of Audio Spoofing Attacks
- A Causal Deep Learning Framework for Classifying Phonemes in Cochlear Implants
- A Chapter-Wise Understanding System for Text-To-Speech in Chinese Novels
- A Classifier for Improving Cause and Effect in SSVEP-based BCIs for Individuals with Complex Communication Disorders
- A Closed-Loop Gain-Control Feedback Model for The Medial Efferent System of The Descending Auditory Pathway
- A Closer Look at Audio-Visual Multi-Person Speech Recognition and Active Speaker Selection
- A Co-Interactive Transformer for Joint Slot Filling and Intent Detection
- A Color Doppler Processing Engine with an Adaptive Clutter Filter for Portable Ultrasound Imaging Devices
- A Compact Joint Distillation Network for Visual Food Recognition
- A Comparative Study of Acoustic and Linguistic Features Classification for Alzheimer's Disease Detection
- A Comparison Study on Infant-Parent Voice Diarization
- A Comparison of Convolutional Neural Networks for Glottal Closure Instant Detection from Raw Speech
- A Comparison of Discrete Latent Variable Models for Speech Representation Learning
- A Comparison of Methods for OOV-Word Recognition on a New Public Dataset
- A Consensus Equilibrium Solution For Deep Image Prior Powered By Red
- A Convex Penalty for Block-Sparse Signals with Unknown Structures
- A Correntropy Based Algorithm for Robust Localization in Wireless Networks
- A Curated Dataset of Urban Scenes for Audio-Visual Scene Analysis
- A DNN Autoencoder for Automotive Radar Interference Mitigation
- A Decentralized Variance-Reduced Method for Stochastic Optimization Over Directed Graphs
- A Deep Reinforcement Learning Approach To Audio-Based Navigation In A Multi-Speaker Environment
- A Deep Spatio-Temporal Model for EEG-Based Imagined Speech Recognition
- A Diffusion FXLMS Algorithm for Multi-Channel Active Noise Control and Variable Spatial Smoothing
- A Dynamical Systems Perspective on Online Bayesian Nonparametric Estimators with Adaptive Hyperparameters
- A Fast Randomized Adaptive CP Decomposition For Streaming Tensors
- A Fast and Efficient Network for Single Image Deraining
- A Features Decoupling Method for Multiple Manipulations Identification in Image Operation Chains
- A Flow-Based Neural Network for Time Domain Speech Enhancement
- A Framework for Pruning Deep Neural Networks Using Energy-Based Models
- A Further Study of Unsupervised Pretraining for Transformer Based Speech Recognition
- A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks
- A General Network Architecture for Sound Event Localization and Detection Using Transfer Learning and Recurrent Neural Network
- A Global Cayley Parametrization of Stiefel Manifold for Direct Utilization of Optimization Mechanisms Over Vector Spaces
- A Global-Local Attention Framework for Weakly Labelled Audio Tagging
- A Graph Learning Algorithm Based On Gaussian Markov Random Fields And Minimax Concave Penalty
- A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery
- A High-Frame-Rate Eye-Tracking Framework for Mobile Devices
- A Homogeneity-Based Multiscale Hyperspectral Image Representation for Sparse Spectral Unmixing
- A Hybrid Approach to Coded Compressed Sensing Where Coupling Takes Place Via the Outer Code
- A Hybrid CNN-BiLSTM Voice Activity Detector
- A Hybrid Feature Enhancement Method for Gl And Segmentation In Histopathology Images
- A Joint Convolutional and Spatial Quad-Directional LSTM Network for Phase Unwrapping
- A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech
- A Large-Dimensional Analysis of Symmetric SNE
- A Large-Scale Chinese Long-Text Extractive Summarization Corpus
- A Layered Embedding-Based Scheme to Cope with Intra-Frame Distortion Drift In IPM-Based HEVC Steganography
- A Low-Complexity Admm-Based Massive Mimo Detectors Via Deep Neural Networks
- A Low-Complexity MIMO Dual Function Radar Communication System via One-Bit Sampling
- A Meta-Learning Framework for Few-Shot Classification of Remote Sensing Scene
- A Method for Determining Periodically Time-Varying Bias and Its Applications in Acoustic Feedback Cancellation
- A Modulation-Domain Loss for Neural-Network-Based Real-Time Speech Enhancement
- A Multi-Channel Temporal Attention Convolutional Neural Network Model for Environmental Sound Classification
- A Multi-Layer Multi-Channel Attentive Network for Gender and Age Recognition
- A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced Data
- A Multi-View Approach to Audio-Visual Speaker Verification
- A Multiple Access Channel Game Using Latency Metric
- A Neural Acoustic Echo Canceller Optimized Using An Automatic Speech Recognizer and Large Scale Synthetic Data
- A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background Music
- A New Automotive Radar 4D Point Clouds Detector by Using Deep Learning
- A New DCASE 2017 Rare Sound Event Detection Benchmark Under Equal Training Data: CRNN With Multi-Width Kernels
- A New Framework Based on Transfer Learning for Cross-Database Pneumonia Detection
- A New High Quality Trajectory Tiling Based Hybrid TTS In Real Time
- A New Tubular Structure Tracking Algorithm Based On Curvature-Penalized Perceptual Grouping
- A Noise-Robust Signal Processing Strategy for Cochlear Implants Using Neural Networks
- A Novel Attention-Based Gated Recurrent Unit and its Efficacy in Speech Emotion Recognition
- A Novel Bayesian Approach for the Two-Dimensional Harmonic Retrieval Problem
- A Novel Convolutional Neural Network Model to Remove Muscle Artifacts from EEG
- A Novel NMF-HMM Speech Enhancement Algorithm Based on Poisson Mixture Model
- A Novel Viewport-Adaptive Motion Compensation Technique for Fisheye Video
- A Novel end-to-end Speech Emotion Recognition Network with Stacked Transformer Layers
- A Parallel Algorithm for Phase Retrieval with Dictionary Learning
- A Parallelizable Lattice Rescoring Strategy with Neural Language Models
- A Parametric Unconstrained Binaural Beamformer Based Noise Reduction and Spatial Cue Preservation for Hearing-Assistive Devices
- A Partially Collapsed Gibbs Sampler for Unsupervised Nonnegative Sparse Signal Restoration
- A Partially-Relaxed Robust DOA Estimator Under Non-Gaussian Low-Rank Interference and Noise
- A Patient-Invariant Model for Freezing of Gait Detection Aided by Wavelet Decomposition
- A Periodic Frame Learning Approach for Accurate Landmark Localization in M-Mode Echocardiography
- A Plug and Play Fast Intersection Over Union Loss for Boundary Box Regression
- A Plug-and-Play Deep Image Prior
- A Probabilistic Model for Segmentation of Ambiguous 3D Lung Nodule
- A Progressive Learning Approach to Adaptive Noise and Speech Estimation for Speech Enhancement and Noisy Speech Recognition
- A Quantitative Analysis Of The Robustness Of Neural Networks For Tabular Data
- A Quantitative Metric for Privacy Leakage in Federated Learning
- A Quaternion-Valued Variational Autoencoder
- A Rank-Constrained Clustering Algorithm with Adaptive Embedding
- A Ranked Similarity Loss Function with pair Weighting for Deep Metric Learning
- A ReLU Dense Layer to Improve the Performance of Neural Networks
- A Real-Time Speaker Diarization System Based on Spatial Spectrum
- A Robust Copula Model for Radar-Based Landmine Detection
- A Robust and Efficient Multi-Scale Seasonal-Trend Decomposition
- A Robust to Noise Adversarial Recurrent Model for Non-Intrusive Load Monitoring
- A Sample-Efficient Scheme for Channel Resource Allocation in Networked Estimation
- A Scale Invariant Measure of Flatness for Deep Network Minima
- A Secure Searchable Image Retrieval Scheme with Correct Retrieval Identity
- A Sequential Contrastive Learning Framework for Robust Dysarthric Speech Recognition
- A Short Tutorial on The Weisfeiler-Lehman Test And Its Variants
- A Simplified Wiener Beamformer Based on Covariance Matrix Modelling
- A Sparse Coding Approach to Automatic Diet Monitoring with Continuous Glucose Monitors
- A Stage Match for Query-by-Example Spoken Term Detection Based On Structure Information of Query
- A Structure-Guided and Sparse-Representation-Based 3d Seismic Inversion Method
- A Time-Domain Convolutional Recurrent Network for Packet Loss Concealment
- A Triplet Appearance Parsing Network for Person Re-Identification
- A Two-Stage Approach to Device-Robust Acoustic Scene Classification
- A Two-Stage Deep Modeling Approach to Articulatory Inversion
- A Tyler-Type Estimator of Location and Scatter Leveraging Riemannian Optimization
- A Unified Approach to Translate Classical Bandit Algorithms to Structured Bandits
- A Universal Bert-Based Front-End Model for Mandarin Text-To-Speech Synthesis
- A Wireless Reference Active Noise Control Headphone Using Coherence Based Selection Technique
- ADAPT-Then-Combine Full Waveform Inversion for Distributed Subsurface Imaging In Seismic Networks
- ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech Separation
- ADMM-Based ML Decoding: from Theory to Practice
- AEC in A Netshell: on Target and Topology Choices for FCRN Acoustic Echo Cancellation
- AISpeech-SJTU ASR System for the Accented English Speech Recognition Challenge
- AISpeech-SJTU Accent Identification System for the Accented English Speech Recognition Challenge
- ASR N-Best Fusion Nets
- ASV-SUBTOOLS: Open Source Toolkit for Automatic Speaker Verification
- ATVIO: Attention Guided Visual-Inertial Odometry
- Absolute 3d Pose Estimation and Length Measurement of Severely Deformed Fish from Monocular Videos in Longline Fishing
- Accdoa: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization And Detection
- Accelerating Auxiliary Function-Based Independent Vector Analysis
- Accelerating Frank-Wolfe with Weighted Average Gradients
- Acoustic Analysis and Dataset of Transitions Between Coupled Rooms
- Acoustic Echo Cancellation with the Dual-Signal Transformation LSTM Network
- Acoustic Reflectors Localization from Stereo Recordings Using Neural Networks
- Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's Disease
- Acoustic-to-Articulatory Inversion for Dysarthric Speech by Using Cross-Corpus Acoustic-Articulatory Data
- Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource Languages
- Action State Update Approach to Dialogue Management
- Active Estimation From Multimodal Data
- Active Privacy-Utility Trade-Off Against A Hypothesis Testing Adversary
- Acute Lymphoblastic Leukemia Detection Based on Adaptive Unsharpening and Deep Learning
- Ada-Sise: Adaptive Semantic Input Sampling for Efficient Explanation of Convolutional Neural Networks
- Adaptable Ensemble Distillation
- Adaptable Multi-Domain Language Model for Transformer ASR
- Adaptive Bi-Directional Attention: Exploring Multi-Granularity Representations for Machine Reading Comprehension
- Adaptive Contention Window Design Using Deep Q-Learning
- Adaptive Dual Tree Structure For Screen Content Coding
- Adaptive Feature Weight Learning For Robust Clustering Problem with Sparse Constraint
- Adaptive GOP Size Decision for Multi-Pass Video Coding Based on Hidden Markov Model
- Adaptive Importance Sampling Via Auto-Regressive Generative Models and Gaussian Processes
- Adaptive Multi-Domain Learning for Outdoor 3d Human Pose and Shape Estimation
- Adaptive Quantization of Model Updates for Communication-Efficient Federated Learning
- Adaptive RF Fingerprint Decomposition in Micro UAV Detection based on Machine Learning
- Adaptive Re-Balancing Network with Gate Mechanism for Long-Tailed Visual Question Answering
- Adaptive Real-Time Filter for Partially-Observed Boolean Dynamical Systems
- Adaptive Subsampling of Multidomain Signals with Product Graphs
- Adaspeech 2: Adaptive Text to Speech with Untranscribed Data
- Admm-Based Fast Algorithm for Robust Multi-Group Multicast Beamforming
- Advances in Morphological Neural Networks: Training, Pruning and Enforcing Shape Constraints
- Advancing RNN Transducer Technology for Speech Recognition
- Adversarial Attacks on Audio Source Separation
- Adversarial Attacks on Coarse-to-Fine Classifiers
- Adversarial Attacks on Object Detectors with Limited Perturbations
- Adversarial Defense for Automatic Speaker Verification by Cascaded Self-Supervised Learning Models
- Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training
- Adversarial Examples Detection Beyond Image Space
- Adversarial Generative Distance-Based Classifier for Robust Out-of-Domain Detection
- Adversarial Learning via Probabilistic Proximity Analysis
- Adversarially Robust Classification Based on GLRT
- Affine Projection Subspace Tracking
- Again-VC: A One-Shot Voice Conversion Using Activation Guidance and Adaptive Instance Normalization
- Age-VOX-Celeb: Multi-Modal Corpus for Facial and Speech Estimation
- Agent-Environment Network for Temporal Action Proposal Generation
- Aggregation Architecture and all-to-one Network for Real-Time Semantic Segmentation
- Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval
- Aligning Sets of Temporal Signals with Riemannian Geometry and Koopman Operator
- Aligning the training and evaluation of unsupervised text style Transfer
- All For One And One For All: Improving Music Separation By Bridging Networks
- Allocating DNN Layers Computation Between Front-End Devices and The Cloud Server for Video Big Data Processing
- Alternating Projections Gridless Covariance-Based Estimation For DOA
- Amplitude Matching: Majorization-Minimization Algorithm for Sound Field Control Only with Amplitude Constraint
- An ADMM Based Network for Hyperspectral Unmixing Tasks
- An Accuracy Network Anomaly Detection Method Based on Ensemble Model
- An Actor-Critic Reinforcement Learning Approach to Minimum age of Information Scheduling in Energy Harvesting Networks
- An Adaptive Discriminant and Sparsity Feature Descriptor for Finger Vein Recognition
- An Adaptive Multi-Scale and Multi-Level Features Fusion Network with Perceptual Loss for Change Detection
- An Adaptive Non-Linear Process for Under-Determined Virtual Microphone Beamforming
- An Adaptive Part-Based Model For Person Re-Identification
- An Adaptive Pyramid Single-View Depth Lookup Table Coding Method
- An Adaptive Regularization Approach to Portfolio Optimization
- An Asymptotically Pointwise Optimal Procedure For Sequential Joint Detection And Estimation
- An Asynchronous WFST-Based Decoder for Automatic Speech Recognition
- An Attention Based Wavelet Convolutional Model for Visual Saliency Detection
- An Attention Model for Hypernasality Prediction in Children with Cleft Palate
- An Attention-Seq2Seq Model Based on CRNN Encoding for Automatic Labanotation Generation from Motion Capture Data
- An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification
- An Efficient Active Set Algorithm for Covariance Based Joint Data and Activity Detection for Massive Random Access with Massive MIMO
- An Efficient Algorithm For Device Detection And Channel Estimation In Asynchronous IOT Systems
- An Efficient Alternating Direction Method for Graph Learning from Smooth Signals
- An Efficient Linear Programming Rounding-and-Refinement Algorithm for Large-Scale Network Slicing Problem
- An Efficient Paper Anti-Counterfeiting Method Based on Microstructure Orientation Estimation
- An Empirical Study of End-To-End Simultaneous Speech Translation Decoding Strategies
- An Empirical Study of Visual Features for DNN Based Audio-Visual Speech Enhancement in Multi-Talker Environments
- An Empirical Study on Task-Oriented Dialogue Translation
- An End-To-End Actor-Critic-Based Neural Coreference Resolution System
- An End-To-End Non-Intrusive Model for Subjective and Objective Real-World Speech Assessment Using a Multi-Task Framework
- An End-to-End Speech Accent Recognition Method Based on Hybrid CTC/Attention Transformer ASR
- An Extension of Sparse Audio Declipper to Multiple Measurement Vectors
- An F-Test for Polynomial Frequency Modulation
- An Hrnet-Blstm Model With Two-Stage Training For Singing Melody Extraction
- An Improved Data Driven Dynamic SIRD Model for Predictive Monitoring of COVID-19
- An Improved Deep Relation Network for Action Recognition in Still Images
- An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
- An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
- An Investigation of End-to-End Models for Robust Speech Recognition
- An Investigation of Using Hybrid Modeling Units for Improving End-to-End Speech Recognition System
- An Iterative Framework for Self-Supervised Deep Speaker Representation Learning
- An Optimal Stochastic Compositional Optimization Method with Applications to Meta Learning
- An Order-Optimal Adaptive Test Plan for Noisy Group Testing Under Unknown Noise Models
- Analog Beamforming With Antenna Selection For Large-Scale Antenna Arrays
- Analysing Bias in Spoken Language Assessment Using Concept Activation Vectors
- Analysis of X-Vectors for Low-Resource Speech Recognition
- Analysis of the but Diarization System for Voxconverse Challenge
- Angle-of-Arrival (AoA) Factorization in Multipath Channels
- Antenna Selection for Massive MIMO Systems Based on POMDP Framework
- Any-to-One Sequence-to-Sequence Voice Conversion Using Self-Supervised Discrete Speech Representations
- Application-Layer DDOS Attacks with Multiple Emulation Dictionaries
- Approximate Weighted C R Coded Matrix Multiplication
- Arrays of First-Order Steerable Differential Microphones
- Arrhythmia Classification with Heartbeat-Aware Transformer
- Artificially Synthesising Data for Audio Classification and Segmentation to Improve Speech and Music Detection in Radio Broadcast
- Assessment of Bipolar Disorder Using Heterogeneous Data of Smartphone-Based Digital Phenotyping
- Assisted Learning: Cooperative AI with Autonomy
- Asymptotic Distribution of Generalized Likelihood Ratio Test Under Model Misspecification With Application to Cooperative Radar-Communications
- Asynchronous Acoustic Echo Cancellation Over Wireless Channels
- Attack on Practical Speaker Verification System Using Universal Adversarial Perturbations
- Attacking and Defending Behind A Psychoacoustics-Based Captcha
- Attention Enhanced Spatial Temporal Neural Network For HRRP Recognition
- Attention Is All You Need In Speech Separation
- Attention on Attention Sparse Dense Convolutional Network for Financial Signal Processing
- Attention-Based Multi-Encoder Automatic Pronunciation Assessment
- Attention-Embedded Decomposed Network with Unpaired CT Images Prior for Metal Artifact Reduction
- Attention-Guided Second-Order Pooling Convolutional Networks
- AttentionLite: Towards Efficient Self-Attention Models for Vision
- Attentive Semantic Exploring for Manipulated Face Detection
- Attribute Decomposition for Flow-Based Domain Mapping
- Audio Dequantization Using (Co)Sparse (Non)Convex Methods
- Audio-Visual Event Recognition Through the Lens of Adversary
- Audio-Visual Speech Enhancement Method Conditioned in the Lip Motion and Speaker-Discriminative Embeddings
- Audio-Visual Speech Inpainting with Deep Learning
- Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss
- Audiovisual Highlight Detection in Videos
- Auditory Filterbanks Benefit Universal Sound Source Separation
- Augmented Gaussian Linear Mixture Model for Spectral Variability in Hyperspectral Unmixing
- Augmenting Transferred Representations for Stock Classification
- AutoKWS: Keyword Spotting with Differentiable Architecture Search
- Autoencoder for Vibrotactile Signal Compression
- Automated Multi-Organ Segmentation in Pet Images Using Cascaded Training of a 3d U-Net and Convolutional Autoencoder
- Automatic And Perceptual Discrimination Between Dysarthria, Apraxia of Speech, and Neurotypical Speech
- Automatic Dysarthric Speech Detection Exploiting Pairwise Distance-Based Convolutional Neural Networks
- Automatic Elicitation Compliance for Short-Duration Speech Based Depression Detection
- Automatic Fine-Grained Localization of Utility Pole Landmarks on Distributed Acoustic Sensing Traces Based on Bilinear Resnets
- Automatic Multitrack Mixing With A Differentiable Mixing Console Of Neural Audio Effects
- Automatic Order Selection in Autoregressive Modeling with Application in EEG Sleep-Stage Classification
- Automatic Registration and Clustering of Time Series
- Autoregressive Fast Multichannel Nonnegative Matrix Factorization For Joint Blind Source Separation And Dereverberation
- B-Small: A Bayesian Neural Network Approach to Sparse Model-Agnostic Meta-Learning
- BLSTM-Based Confidence Estimation for End-to-End Speech Recognition
- BW-EDA-EEND: streaming END-TO-END Neural Speaker Diarization for a Variable Number of Speakers
- Backdoor Attack Against Speaker Verification
- Baitradar: A Multi-Model Clickbait Detection Algorithm Using Deep Learning
- Bandwidth Extension is All You Need
- Banraw: Band-Limited Radar Waveform Design Via Phase Retrieval
- Bayes-Optimal Methods for Finding the Source of a Cascade
- Bayesian Estimation of a Tail-Index with Marginalized Threshold
- Bayesian Massive MIMO Channel Estimation with Parameter Estimation Using Low-Resolution ADCs
- Bayesian Multiple Change-Point Detection of Propagating Events
- Bayesian Transformer Language Models for Speech Recognition
- Beam Focusing for Multi-User MIMO Communications with Dynamic Metasurface Antennas
- Beamforming for Bidirectional Mimo Full Duplex Under the Joint Sum Power and Per Antenna Power Constraints
- Benign Overfitting in Binary Classification of Gaussian Mixtures
- Bi-APC: Bidirectional Autoregressive Predictive Coding for Unsupervised Pre-Training and its Application to Children's ASR
- Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech Synthesis
- Bidirectional Focused Semantic Alignment Attention Network for Cross-Modal Retrieval
- Bifocal Neural ASR: Exploiting Keyword Spotting for Inference Optimization
- Binary Control and Digital-to-Analog Conversion Using Composite NUV Priors and Iterative Gaussian Message Passing
- Bishift-Net for Image Inpainting
- Bit Constrained Communication Receivers In Joint Radar Communications Systems
- Blend-Res2net: Blended Representation Space by Transformation of Residual Mapping with Restrained Learning for Time Series Classification
- Blind Amplitude Estimation of Early Room Reflections Using Alternating Least Squares
- Blind Carbon Copy on Dirty Paper: Seamless Spectrum Underlay via Canonical Correlation Analysis
- Blind Deinterleaving of Signals in Time Series with Self-Attention Based Soft Min-Cost Flow Learning
- Blind Extraction of Moving Audio Source in a Challenging Environment Supported by Speaker Identification Via X-Vectors
- Blind Extraction of Moving Sources via Independent Component and Vector Analysis: Examples
- Blind Image Quality Evaluator with Scale Robustness
- Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source Separation
- Block Kalman Filter: An Asymptotic Block Particle Filter in the Linear Gaussian Case
- Bluetooth Low Energy and CNN-Based Angle of Arrival Localization in Presence of Rayleigh Fading
- Boosting Low-Resource Intent Detection with in-Scope Prototypical Networks
- Branchy-GNN: A Device-Edge Co-Inference Framework for Efficient Point Cloud Processing
- Bridging Unpaired Facial Photos and Sketches by Line-Drawings
- Bytecover: Cover Song Identification Via Multi-Loss Training
- Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation
- CASS-NAT: CTC Alignment-Based Single Step Non-Autoregressive Transformer for Speech Recognition
- CDPAM: Contrastive Learning for Perceptual Audio Similarity
- CMIM: Cross-Modal Information Maximization For Medical Imaging
- CNN-Based Spoken Term Detection and Localization without Dynamic Programming
- CNR-IEMN: A Deep Learning Based Approach to Recognise Covid-19 from CT-Scan
- COOPNet: Multi-Modal Cooperative Gender Prediction in Social Media User Profiling
- CSPN: Multi-Scale Cascade Spatial Pyramid Network for Object Detection
- Cam: Context-Aware Masking for Robust Speaker Verification
- Camera Calibration with Pose Guidance
- Camp: A Two-Stage Approach to Modelling Prosody in Context
- Canet: Context-Aware Loss for Descriptor Learning
- Canonical Polyadic Tensor Decomposition With Low-Rank Factor Matrices
- Capturing Banding in Images: Database Construction and Objective Assessment
- Capturing Multi-Resolution Context by Dilated Self-Attention
- Capturing Temporal Dependencies Through Future Prediction for CNN-Based Audio Classifiers
- Cascade Attention Fusion for Fine-Grained Image Captioning Based on Multi-Layer LSTM
- Cascaded All-Pass Filters with Randomized Center Frequencies and Phase Polarity for Acoustic and Speech Measurement and Data Augmentation
- Cascaded Encoders for Unifying Streaming and Non-Streaming ASR
- Cascaded Models with Cyclic Feedback for Direct Speech Translation
- Cascaded Time + Time-Frequency Unet For Speech Enhancement: Jointly Addressing Clipping, Codec Distortions, And Gaps
- Catiloc: Camera Image Transformer for Indoor Localization
- Centrality Based Number of Cluster Estimation in Graph Clustering
- Cgan-Net: Class-Guided Asymmetric Non-Local Network for Real-Time Semantic Segmentation
- Channel Attention Residual U-Net for Retinal Vessel Segmentation
- Channel-Wise Mix-Fusion Deep Neural Networks for Zero-Shot Learning
- Characterization of Mems Microphone Sensitivity and Phase Distributions with Applications in Array Processing
- Checking PRNU Usability on Modern Devices
- Cif-Based Collaborative Decoding for End-to-End Contextual Speech Recognition
- Class Aware Robust Training
- Class-Conditional Defense GAN Against End-To-End Speech Attacks
- Class-Imbalanced Classifiers Using Ensembles of Gaussian Processes And Gaussian Process Latent Variable Models
- Classification of Expert-Novice Level Using Eye Tracking And Motion Data via Conditional Multimodal Variational Autoencoder
- Classifying Speech Intelligibility Levels of Children in Two Continuous Speech Styles
- Close-Talking Recording with Planarly Distributed Microphones
- Clustering A Collection of Networks With Mixtures of L1-Sparse Graphical Models
- Co-Attentional Transformers for Story-Based Video Understanding
- Co-Capsule Networks Based Knowledge Transfer for Cross-Domain Recommendation
- Coarse-To-Careful: Seeking Semantic-Related Knowledge for Open-Domain Commonsense Question Answering
- Code-Switch Speech Rescoring with Monolingual Data
- Codebook Design for Dual-Polarized Ultra-Massive Mimo Communications at Millimeter Wave and Terahertz Bands
- Cognitive Memory Constrained Human Decision Making based on Multi-source Information
- Cold Start Revisited: A Deep Hybrid Recommender with Cold-Warm Item Harmonization
- Collaborative Inference via Ensembles on the Edge
- Collaborative Intelligence: Challenges and Opportunities
- Collaborative Learning to Generate Audio-Video Jointly
- Combined Differential Beamforming With Uniform Linear Microphone Arrays
- Combining Adaptive Filtering And Complex-Valued Deep Postfiltering For Acoustic Echo Cancellation
- Combining Dynamic Image and Prediction Ensemble for Cross-Domain Face Anti-Spoofing
- Communication Over Block Fading Channels - An Algorithmic Perspective On Optimal Transmission Schemes
- Communication-Cost Aware Microphone Selection for Neural Speech Enhancement with Ad-Hoc Microphone Arrays
- Compact Graph Architecture for Speech Emotion Recognition
- Comparative Study of Different Epoch Extraction Methods for Speech Associated with Voice Disorders
- Comparison of Deep Co-Training and Mean-Teacher Approaches for Semi-Supervised Audio Tagging
- Complex Ratio Masking For Singing Voice Separation
- Complex-Valued Vs. Real-Valued Neural Networks for Classification Perspectives: An Example on Non-Circular Data
- Compositional Embedding Models for Speaker Identification and Diarization with Simultaneous Speech From 2+ Speakers
- Compressed Representation of Cepstral Coefficients via Recurrent Neural Networks for Informed Speech Enhancement
- Compressing Deep Neural Networks for Efficient Speech Enhancement
- Compressing Local Descriptor Models for Mobile Applications
- Compressive Signal Recovery Under Sensing Matrix Errors Combined With Unknown Measurement Gains
- Compressive Wideband Spectrum Sensing and Carrier Frequency Estimation with Unknown Mimo Channels
- Computationally Efficient DNN-Based Approximation of an Auditory Model for Applications in Speech Processing
- Confidence Estimation for Attention-Based Sequence-to-Sequence Models for Speech Recognition
- Constant Approximation Algorithm for Minimizing Concave Impurity
- Constrained Tensor Decomposition for 2d DOA Estimation In Transmit Beamspace Mimo Radar with Subarrays
- Construction of Unit-Norm Tight Frame Based Preconditioner for Sparse Coding
- Construction of a Large-Scale Japanese ASR Corpus on TV Recordings
- Contact Tracing Enhances the Efficiency of Covid-19 Group Testing
- Content-Aware Speaker Embeddings for Speaker Diarisation
- Context-Aware Prosody Correction for Text-Based Speech Editing
- Context-Aware Speech Stress Detection in Hospital Workers Using Bi-LSTM Classifiers
- Continuous Cnn For Nonuniform Time Series
- Continuous Face Aging Generative Adversarial Networks
- Continuous Speech Separation with Conformer
- Continuous-Time Self-Attention in Neural Differential Equation
- Contrastive Embeddind Learning Method for Respiratory Sound Classification
- Contrastive Learning of General-Purpose Audio Representations
- Contrastive Predictive Coding Supported Factorized Variational Autoencoder For Unsupervised Learning Of Disentangled Speech Representations
- Contrastive Self-Supervised Learning for Text-Independent Speaker Verification
- Contrastive Self-Supervised Learning for Wireless Power Control
- Contrastive Semi-Supervised Learning for ASR
- Contrastive Separative Coding for Self-Supervised Representation Learning
- Contrastive Unsupervised Learning for Speech Emotion Recognition
- Control Architecture of the Double-Cross-Correlation Processor for Sampling-Rate-Offset Estimation in Acoustic Sensor Networks
- Controlled Testing and Isolation for Suppressing Covid-19
- Convergence Analysis of the Graph-Topology-Inference Kernel LMS Algorithm
- Conversational Query Rewriting with Self-Supervised Learning
- Convex Neural Autoregressive Models: Towards Tractable, Expressive, and Theoretically-Backed Models for Sequential Forecasting and Generation
- Convolutional Dropout and Wordpiece Augmentation for End-to-End Speech Recognition
- Convolutional Neural Network-Aided Bit-Flipping for Belief Propagation Decoding of Polar Codes
- Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation
- Cooperative Parameter Tracking on the Unit Sphere Using Distributed Adapt-Then-Combine Particle Filters and Parallel Transport
- Cooperative Scenarios for Multi-Agent Reinforcement Learning in Wireless Edge Caching
- CopyPaste: An Augmentation Method for Speech Emotion Recognition
- Correlation-Based Robust Linear Regression with Iterative Outlier Removal
- Corrupted Contextual Bandits: Online Learning with Corrupted Context
- Cost Affinity Learning Network for Stereo Matching
- Coughwatch: Real-World Cough Detection using Smartwatches
- Count And Separate: Incorporating Speaker Counting For Continuous Speaker Separation
- Count Sketch with Zero Checking: Efficient Recovery of Heavy Components
- Covid-19 Diagnostic Using 3d Deep Transfer Learning for Classification of Volumetric Computerised Tomography Chest Scans
- Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder
- Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model Interaction
- Cross-Corpus Speech Emotion Recognition Using Joint Distribution Adaptive Regression
- Cross-Domain Semi-Supervised Deep Metric Learning for Image Sentiment Analysis
- Cross-Domain Sentiment Classification with Contrastive Learning and Mutual Information Maximization
- Cross-Modal Knowledge Distillation For Fine-Grained One-Shot Classification
- Cross-Modal Representation Reconstruction for Zero-Shot Classification
- Cross-Modal Spectrum Transformation Network for Acoustic Scene Classification
- Cross-Silo Federated Training in the Cloud with Diversity Scaling and Semi-Supervised Learning
- Cross-Teager Energy Cepstral Coefficients for Replay Spoof Detection on Voice Assistants
- Crowd Counting Via Multi-Level Regression With Latent Gaussian Maps
- Crowdsourcing Approach for Subjective Evaluation of Echo Impairment
- Crypto-Oriented Neural Architecture Design
- Ct-Caps: Feature Extraction-Based Automated Framework for Covid-19 Disease Identification From Chest Ct Scans Using Capsule Networks
- Cue-Preserving MMSE Filter with Bayesian SNR Marginalization for Binaural Speech Enhancement
- Cycle Generative Adversarial Network Approaches to Produce Novel Portable Chest X-Rays Images for Covid-19 Diagnosis
- D-VDAMP: Denoising-Based Approximate Message Passing for Compressive MRI
- DAG-GAN: Causal Structure Learning with Generative Adversarial Nets
- DBnet: Doa-Driven Beamforming Network for end-to-end Reverberant Sound Source Separation
- DCASENET: An Integrated Pretrained Deep Neural Network for Detecting and Classifying Acoustic Scenes and Events
- DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning
- DEEPTALK: Vocal Style Encoding for Speaker Recognition and Speech Synthesis
- DFDM: A Deep Feature Decoupling Module for Lung Nodule Segmentation
- DHASP: Differentiable Hearing Aid Speech Processing
- DHCN: Deep Hierarchical Context Networks For Image Annotation
- DNANet: Dense Nested Attention Network for Single Image Dehazing
- DO as I Mean, Not as I Say: Sequence Loss Training for Spoken Language Understanding
- DP-SIGNSGD: When Efficiency Meets Privacy and Robustness
- DP-VTON: Toward Detail-Preserving Image-Based Virtual Try-on Network
- DURAS: Deep Unfolded Radar Sensing Using Doppler Focusing
- Data Augmentation with Signal Companding for Detection of Logical Access Attacks
- Data Discovery Using Lossless Compression-Based Sparse Representation
- Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain
- Data-Driven Adaptive Network Resource Slicing for Multi-Tenant Networks
- Data-Efficient Framework for Real-World Multiple Sound Source 2d Localization
- Decentralized Deep Learning Using Momentum-Accelerated Consensus
- Decentralized Motion Inference and Registration of Neuropixel Data
- Decentralized Optimization Over Noisy, Rate-Constrained Networks: How We Agree By Talking About How We Disagree
- Decentralized Optimization on Time-Varying Directed Graphs Under Communication Constraints
- Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech Recognition
- Decision Tree Based Inter Partition Termination For Av1 Encoding
- Decoding Music Attention from "EEG Headphones": A User-Friendly Auditory Brain-Computer Interface
- Decoding Neural Representations of Rhythmic Sounds From Magnetoencephalography
- Decomposing Textures using Exponential Analysis
- Decouple the High-Frequency and Low-Frequency Information of Images for Semantic Segmentation
- Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech Recognition
- Deep Active Learning Approach to Adaptive Beamforming for mmWave Initial Alignment
- Deep Adversarial Quantization Network for Cross-Modal Retrieval
- Deep Auto-Encoding and Biohashing for Secure Finger Vein Recognition
- Deep Color Constancy Using Temporal Gradient Under Ac Light Sources
- Deep Convolutional Gaussian Processes for Mmwave Outdoor Localization
- Deep Convolutional and Recurrent Networks for Polyphonic Instrument Classification from Monophonic Raw Audio Waveforms
- Deep Deterministic Information Bottleneck with Matrix-Based Entropy Functional
- Deep Ensemble Siamese Network For Incremental Signal Classification
- Deep Generative Demixing: Error Bounds for Demixing Subgaussian Mixtures of Lipschitz Signals
- Deep Generative Model Learning For Blind Spectrum Cartography with NMF-Based Radio Map Disaggregation
- Deep Hashing for Motion Capture Data Retrieval
- Deep Learning Architectural Designs for Super-Resolution Of Noisy Images
- Deep Learning Based Hybrid Precoding in Dual-Band Communication Systems
- Deep Learning for Linear Inverse Problems Using the Plug-and-Play Priors Framework
- Deep Learning-Based Cross-Layer Resource Allocation for Wired Communication Systems
- Deep Lung Auscultation Using Acoustic Biomarkers for Abnormal Respiratory Sound Event Detection
- Deep Multi-Frame MVDR Filtering for Single-Microphone Speech Enhancement
- Deep Multiway Canonical Correlation Analysis For Multi-Subject Eeg Normalization
- Deep Neural Network Based Cough Detection Using Bed-Mounted Accelerometer Measurements
- Deep Neural Network Embeddings for the Estimation of the Degree of Sleepiness
- Deep Neural Networks with Flexible Complexity While Training Based on Neural Ordinary Differential Equations
- Deep Residual Echo Suppression With A Tunable Tradeoff Between Signal Distortion And Echo Suppression
- Deep S3PR: Simultaneous Source Separation and Phase Retrieval Using Deep Generative Models
- Deep Semi-Supervised Metric Learning Via Identification of Manifold Memberships
- Deep Transform and Metric Learning Networks
- Deep Unfolding Network for Block-Sparse Signal Recovery
- Deep Weighted MMSE Downlink Beamforming
- DeepF0: End-To-End Fundamental Frequency Estimation for Music and Speech Signals
- Deepemocluster: a Semi-Supervised Framework for Latent Cluster Representation of Speech Emotions
- Deepnodule: Multi-Task Learning of Segmentation Bootstrap for Pulmonary Nodule Detection
- Deficient Basis Estimation of Noise Spatial Covariance Matrix for Rank-Constrained Spatial Covariance Matrix Estimation Method in Blind Speech Extraction
- Demystifying Model Averaging for Communication-Efficient Federated Matrix Factorization
- Denoispeech: Denoising Text to Speech with Frame-Level Noise Modeling
- Dense Attention Module for Accurate Pulmonary Nodule Detection
- Dense Feature Pyramid Grids Network for Single Image Deraining
- Densely Connected Multi-Stage Model with Channel Wise Subband Feature for Real-Time Speech Enhancement
- Dependence-Guided Multi-View Clustering
- Depression Detection by Analysing Eye Movements on Emotional Images
- Design of Graph Signal Sampling Matrices for Arbitrary Signal Subspaces
- Designing Random FM Radar Waveforms with Compact Spectrum
- Detecting Acoustic Reflectors Using A Robot's Ego-Noise
- Detecting Adversarial Attacks on Audiovisual Speech Recognition
- Detecting Alzheimer's Disease from Speech Using Neural Networks with Bottleneck Features and Data Augmentation
- Detecting Covid-19 and Community Acquired Pneumonia Using Chest CT Scan Images With Deep Learning
- Detecting Signal Corruptions in Voice Recordings For Speech Therapy
- Detection Of Malicious DNS and Web Servers using Graph-Based Approaches
- Detection of Audio-Video Synchronization Errors Via Event Detection
- Detection of Covid-19 Through the Analysis of Vocal Fold Oscillations
- Detection of Post-Traumatic Stress Disorder Using Learned Time-Frequency Representations from Pupillometry
- Developing Real-Time Streaming Transformer Transducer for Speech Recognition on Large-Scale Dataset
- Development of the Cuhk Elderly Speech Recognition System for Neurocognitive Disorder Detection Using the Dementiabank Corpus
- Diagnosing Covid-19 from CT Images Based on an Ensemble Learning Framework
- Dian: Duration Informed Auto-Regressive Network for Voice Cloning
- Didispeech: A Large Scale Mandarin Speech Corpus
- Differentiable Signal Processing With Black-Box Audio Effects
- Differential Chaos Shift Keying-Based Wireless Power Transfer
- Differential Convolution Feature Guided Deep Multi-Scale Multiple Instance Learning for Aerial Scene Classification
- Dimension Selected Subspace Clustering
- Direction Of Arrival Estimation For Non-Coherent Sub-Arrays Via Joint Sparse And Low-Rank Signal Recovery
- Direction Preserving Wind Noise Reduction Of B-Format Signals
- Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
- Directional Sparse Filtering Using Weighted Lehmer Mean for Blind Separation of Unbalanced Speech Mixtures
- Discrete Cosine Transform Based Causal Convolutional Neural Network for Drift Compensation in Chemical Sensors
- Discriminability of Single-Layer Graph Neural Networks
- Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTS
- Disentanglement for Audio-Visual Emotion Recognition Using Multitask Setup
- Disentangling Subject-Dependent/-Independent Representations for 2D Motion Retargeting
- Distributed Scheduling Using Graph Neural Networks
- Distributed Speech Separation in Spatially Unconstrained Microphone Arrays
- Distribution-Aware Hierarchical Weighting Method for Deep Metric Learning
- Divide and Conquer: One-bit MIMO-OFDM Detection by Inexact Expectation Maximization
- Dnsmos: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
- DoA estimation of a hidden RF source exploiting simple backscatter radio tags
- Domain Adaptation for Learning Generator From Paired Few-Shot Data
- Domain-Adversarial Autoencoder with Attention Based Feature Level Fusion for Speech Emotion Recognition
- Domain-Aware Neural Language Models for Speech Recognition
- Domestic Activities Clustering From Audio Recordings Using Convolutional Capsule Autoencoder Network
- Don't Look Back: An Online Beat Tracking Method Using RNN and Enhanced Particle Filtering
- Don't Shoot Butterfly with Rifles: Multi-Channel Continuous Speech Separation with Early Exit Transformer
- Double Multi-Head Attention for Speaker Verification
- Double-DCCCAE: Estimation of Body Gestures From Speech Waveform
- Double-Linear Thompson Sampling for Context-Attentive Bandits
- Drawgan: Text to Image Synthesis with Drawing Generative Adversarial Networks
- Drawing Order Recovery from Trajectory Components
- Dual Metric Discriminator for Open Set Video Domain Adaptation
- Dual-Path Modeling for Long Recording Speech Separation in Meetings
- Dual-Stream Network Based On Global Guidance for Salient Object Detection
- Dualformer: A Unified Bidirectional Sequence-to-Sequence Learning
- Dynamic Curriculum Learning via Data Parameters for Noise Robust Keyword Spotting
- Dynamic Graph Learning Based on Graph Laplacian
- Dynamic Graph Modeling Of Simultaneous EEG And Eye-Tracking Data For Reading Task Identification
- Dynamic Point Cloud Compression Using A Cuboid Oriented Discrete Cosine Based Motion Model
- Dynamic Resource Optimization for Adaptive Federated Learning at the Wireless Network Edge
- Dynamic Sparsity Neural Networks for Automatic Speech Recognition
- Dynamic Texture Recognition via Nuclear Distances on Kernelized Scattering Histogram Spaces
- EADNet: Efficient Asymmetric Dilated Network For Semantic Segmentation
- ECCL: Explicit Correlation-Based Convolution Boundary Locator for Moment Localization
- ECG Heart-Beat Classification Using Multimodal Image Fusion
- EEG-Based Emotion Classification Using Graph Signal Processing
- EKFNet: Learning System Noise Statistics from Measurement Data
- Eat: Enhanced ASR-TTS for Self-Supervised Speech Recognition
- Echo State Speech Recognition
- Edge-Aware Multi-Scale Progressive Colorization
- Effect of Language Proficiency on Subjective Evaluation of Noise Suppression Algorithms
- Effect of Noise and Model Complexity on Detection of Amyotrophic Lateral Sclerosis and Parkinson's Disease Using Pitch and MFCC
- Effect of Video Pixel-Binning on Source Attribution of Mixed Media
- Effective Rank-Based Estimation of the Coherent-to-Diffuse Power Ratio
- Efficient Adversarial Audio Synthesis VIA Progressive Upsampling
- Efficient Client Contribution Evaluation for Horizontal Federated Learning
- Efficient End-to-End Audio Embeddings Generation for Audio Classification on Target Applications
- Efficient Face Manipulation Via Deep Feature Disentanglement And Reintegration Net
- Efficient Knowledge Distillation for RNN-Transducer Models
- Efficient Long Periodic Binary Sequence Designs for Automotive Radar
- Efficient Migration to the Next Generation of Networks Based on Digital Annealing
- Efficient Multi-Objective GANs for Image Restoration
- Efficient Network Protection Games Against Multiple Types Of Strategic Attackers
- Efficient Power Allocation Using Graph Neural Networks and Deep Algorithm Unfolding
- Efficient Real-Time Video Stabilization with a Novel Least Squares Formulation
- Efficient Speech Emotion Recognition Using Multi-Scale CNN and Attention
- Efficient Training Data Generation for Phase-Based DOA Estimation
- Efficient Use of End-to-End Data in Spoken Language Processing
- Ego-Based Entropy Measures for Structural Representations on Graphs
- Ego-GNNs: Exploiting Ego Structures in Graph Neural Networks
- Elbert: Fast Albert with Confidence-Window Based Early Exit
- Elliptical Shape Recovery from Blurred Pixels Using Deep Learning
- Embedding Semantic Hierarchy in Discrete Optimal Transport for Risk Minimization
- Emformer: Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech Recognition
- Emotion Controllable Speech Synthesis Using Emotion-Unlabeled Dataset with the Assistance of Cross-Domain Speech Emotion Recognition
- Emotion Recognition by Fusing Time Synchronous and Time Asynchronous Representations
- Empirically Accelerating Scaled Gradient Projection Using Deep Neural Network for Inverse Problems in Image Processing
- Enabling Efficient and Expressive Spatial Keyword Queries On Encrypted Data
- Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone Classification
- End To End Learning For Convolutive Multi-Channel Wiener Filtering
- End-2-End Modeling of Speech and Gait from Patients with Parkinson's Disease: Comparison Between High Quality Vs. Smartphone Data
- End-To-End Audio-Visual Speech Recognition with Conformers
- End-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings
- End-To-End Multi-Accent Speech Recognition with Unsupervised Accent Modelling
- End-To-End Speaker Diarization as Post-Processing
- End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend
- End-to-End Learning of Variational Models and Solvers for the Resolution of Interpolation Problems
- End-to-End Lyrics Recognition with Voice to Singing Style Transfer
- End-to-End Multi-Channel Transformer for Speech Recognition
- End-to-End Multilingual Automatic Speech Recognition for Less-Resourced Languages: The Case of Four Ethiopian Languages
- End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained Features
- End-to-End Text-to-Speech Using Latent Duration Based on VQ-VAE
- End-to-End anti-spoofing with RawNet2
- End2End Acoustic to Semantic Transduction
- Energy Efficiency Optimization Technique for SWIPT-Enabled Multi-Group Multicasting Systems with Heterogeneous Users
- Energy Minimization for Federated Learning with IRS-Assisted Over-the-Air Computation
- Enhanced Automotive Target Detection through Radar and Communications Sensor Fusion
- Enhanced Blind Calibration of Uniform Linear Arrays with One-Bit Quantization by Kullback-Leibler Divergence Covariance Fitting
- Enhanced Standard Esprit For Overcoming Imperfections In DOA Estimation
- Enhancing Audio Augmentation Methods with Consistency Learning
- Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation
- Enhancing Deep Paraphrase Identification via Leveraging Word Alignment Information
- Enhancing Image Steganography Via Stego Generation And Selection
- Enhancing Model Robustness by Incorporating Adversarial Knowledge into Semantic Representation
- Enhancing Multi-Channel Eeg Classification with Gramian Temporal Generative Adversarial Networks
- Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders
- Ensemble Combination between Different Time Segmentations
- Ensemble Distillation Approaches for Grammatical Error Correction
- Ensure: Ensemble Stein's Unbiased Risk Estimator for Unsupervised Learning
- Environment-Independent Wi-Fi Human Activity Recognition with Adversarial Network
- Error Estimates in Second-Order Continuous-Time Sigma-Delta Modulators
- Error-Driven Fixed-Budget ASR Personalization for Accented Speakers
- Error-Driven Pruning of Language Models for Virtual Assistants
- Estimating Fiedler Value on Large Networks Based on Random Walk Observations
- Estimating Severity of Depression From Acoustic Features and Embeddings of Natural Speech
- Estimation of Groundwater Storage Variations in Indus River Basin Using Grace Data
- Estimation of Microphone Clusters in Acoustic Sensor Networks Using Unsupervised Federated Learning
- Estimation of Visual Features of Viewed Image From Individual and Shared Brain Information Based on FMRI Data Using Probabilistic Generative Model
- Evaluation and Comparison of Three Source Direction-of-Arrival Estimators Using Relative Harmonic Coefficients
- Event-Driven Modulo Sampling
- Evolutionary Quantization of Neural Networks with Mixed-Precision
- Evolving Quantized Neural Networks for Image Classification Using A Multi-Objective Genetic Algorithm
- Exact Linear Convergence Rate Analysis for Low-Rank Symmetric Matrix Completion via Gradient Descent
- Expediting discovery in Neural Architecture Search by Combining Learning with Planning
- Exploiting Non-Negative Matrix Factorization for Binaural Sound Localization in the Presence of Directional Interference
- Exploiting the Dual-Tree Complex Wavelet Transform for Ship Wake Detection in SAR Imagery
- Exploring Automatic COVID-19 Diagnosis via Voice and Symptoms from Crowdsourced Data
- Exploring Visual-Audio Composition Alignment Network for Quality Fashion Retrieval in Video
- Exploring the application of synthetic audio in training keyword spotters
- Exploring the use of Common Label Set to Improve Speech Recognition of Low Resource Indian Languages
- Exposing GAN-Generated Faces Using Inconsistent Corneal Specular Highlights
- Extended Object Tracking With Automotive Radar Using B-Spline Chained Ellipses Model
- Extending Music Based On Emotion And Tonality Via Generative Adversarial Network
- Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical Speech
- Extending the Reverse JPEG Compatibility Attack to Double Compressed Images
- F-Net: Fusion Neural Network for Vehicle Trajectory Prediction in Autonomous Driving
- FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection
- FMA-ETA: Estimating Travel Time Entirely Based on FFN with Attention
- FPGA Hardware Design for Plenoptic 3D Image Processing Algorithm Targeting a Mobile Application
- FWB-Net: Front White Balance Network for Color Shift Correction in Single Image Dehazing Via Atmospheric Light Estimation
- Factorized CRF with Batch Normalization Based on the Entire Training Data
- Failure Prediction by Confidence Estimation of Uncertainty-Aware Dirichlet Networks
- Fast DCTTS: Efficient Deep Convolutional Text-to-Speech
- Fast Decentralized Linear Functions Via Successive Graph Shift Operators
- Fast Graph Kernel with Optical Random Features
- Fast Hierarchy Preserving Graph Embedding via Subspace Constraints
- Fast Inverse Mapping of Face GANs
- Fast Local Representation Learning with Adaptive Anchor Graph
- Fast Manifold Landmarking Using Extreme Eigen-Pairs
- Fast Threshold Optimization for Multi-Label Audio Tagging Using Surrogate Gradient Learning
- Fast and Provable Robust PCA VIA Normalized Coherence Pursuit
- Fast and Robust ADMM for Blind Super-Resolution
- Fast and Robust Stratified Self-Calibration Using Time-Difference-Of-Arrival Measurements
- Fast: Feature Aggregation for Detecting Salient Object in Real-Time
- FastEmit: Low-Latency Streaming ASR with Sequence-Level Emission Regularization
- Fastpitch: Parallel Text-to-Speech with Pitch Prediction
- Fcl-Taco2: Towards Fast, Controllable and Lightweight Text-to-Speech Synthesis
- Fden: Mining Effective Information of Features in Detecting Network Anomalies
- Feature Integration via Semi-Supervised Ordinally Multi-Modal Gaussian Process Latent Variable Model
- Feature Redundancy Mining: Deep Light-Weight Image Super-Resolution Model
- Feature Reuse for a Randomization Based Neural Network
- Federated Acoustic Modeling for Automatic Speech Recognition
- Federated Algorithm with Bayesian Approach: Omni-Fedge
- Federated Dropout Learning for Hybrid Beamforming with Spatial Path Index Modulation in Multi-User Mmwave-Mimo Systems
- Federated Learning from Big Data Over Networks
- Federated Learning with Local Differential Privacy: Trade-Offs Between Privacy, Utility, and Communication
- Federated Marginal Personalization for ASR Rescoring
- Few-Shot Continual Learning for Audio Classification
- Few-Shot Image Classification with Multi-Facet Prototypes
- Few-Shot Learning for Ct Scan Based Covid-19 Diagnosis
- Few-Shot Learning for Decoding Surface Electromyography for Hand Gesture Recognition
- Fiber-Sampled Stochastic Mirror Descent for Tensor Decomposition with β-Divergence
- Fine-Grained Mri Reconstruction Using Attentive Selection Generative Adversarial Networks
- Fine-Grained Pose Temporal Memory Module for Video Pose Estimation and Tracking
- Fine-Tuning of Pre-Trained End-to-End Speech Recognition with Generative Adversarial Networks
- First-Order Fast Algorithm for Structurally Optimal Multi-Group Multicast Beamforming in Large-Scale Systems
- Flow-Based Self-Supervised Density Estimation for Anomalous Sound Detection
- Focus on the Present: A Regularization Method for the ASR Source-Target Attention Layer
- Focusing-Based Wideband Adaptive Beamforming Using Covariance Matrix Reconstruction
- Fontnet: On-Device Font Understanding and Prediction Pipeline
- FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances
- Forensicability of Deep Neural Network Inference Pipelines
- Four-Dimensional High-Resolution Automotive Radar Imaging Exploiting Joint Sparse-Frequency and Sparse-Array Design
- Fourier Transformation Autoencoders for Anomaly Detection
- Foveal Avascular Zone Segmentation of Octa Images Using Deep Learning Approach with Unsupervised Vessel Segmentation
- Fragmentvc: Any-To-Any Voice Conversion by End-To-End Extracting and Fusing Fine-Grained Voice Fragments with Attention
- Frame Rate Up-Conversion Using Key Point Agnostic Frequency-Selective Mesh-to-Grid Resampling
- Frame-Rate-Aware Aggregation for Efficient Video Super-Resolution
- Frequency-Temporal Attention Network for Singing Melody Extraction
- Full-Duplex Multifunction Transceiver with Joint Constant Envelope Transmission and Wideband Reception
- Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
- Fully-Neural Approach to Vehicle Weighing and Strain Prediction on Bridges Using Wireless Accelerometers
- Fundamental Frequency Feature Normalization and Data Augmentation for Child Speech Recognition
- Fundamental Trade-Offs in Noisy Super-Resolution with Synthetic Apertures
- Fusing Information Streams in End-to-End Audio-Visual Speech Recognition
- Fusing Multitask Models by Recursive Least Squares
- Fusion-Based Digital Image Correlation Framework for Strain Measurement
- G-Arrays: Geometric Arrays for Efficient Point Cloud Processing
- GAN-Based Out-of-Domain Detection Using Both In-Domain and Out-of-Domain Samples
- GDTW: A Novel Differentiable DTW Loss for Time Series Tasks
- GTA-Net: Gradual Temporal Aggregation Network for Fast Video Deraining
- Gate Trimming: One-Shot Channel Pruning for Efficient Convolutional Neural Networks
- Gating Feature Dense Network for Single Anisotropic Mr Image Super-Resolution
- Gaussian Kernelized Self-Attention for Long Sequence Data and its Application to CTC-Based Speech Recognition
- Gaussian Process Temporal-Difference Learning with Scalability and Worst-Case Performance Guarantees
- General Total Variation Regularized Sparse Bayesian Learning for Robust Block-Sparse Signal Recovery
- Generalized Knowledge Distillation from an Ensemble of Specialized Teachers Leveraging Unsupervised Neural Clustering
- Generalized Polytopic Matrix Factorization
- Generalized Thinned Coprime Array for DOA Estimation
- Generating Empathetic Responses by Injecting Anticipated Emotion
- Generating Human Readable Transcript for Automatic Speech Recognition with Pre-Trained Language Model
- Generating Natural Questions from Images for Multimodal Assistants
- Generative Information Fusion
- Generative Speech Coding with Predictive Variance Regularization
- Geom-Spider-EM: Faster Variance Reduced Stochastic Expectation Maximization for Nonconvex Finite-Sum Optimization
- Geometric Scattering Attention Networks
- Geometry Consistency Of Augmented Reality Based On Semantics
- Global-Localized Agent Graph Convolution for Multi-Agent Reinforcement Learning
- Globally Optimal Beamforming for Rate Splitting Multiple Access
- Gps-Denied Navigation Using Sar Images And Neural Networks
- Gradual Federated Learning Using Simulated Annealing
- Gramian-Based Adaptive Combination Policies for Diffusion Learning Over Networks
- Granger Causality Based Directional Phase-Amplitude Coupling Measure
- Graph Attention Networks for Speaker Verification
- Graph Attention and Interaction Network With Multi-Task Learning for Fact Verification
- Graph Embedding using Multi-Layer Adjacent Point Merging Model
- Graph Enhanced Query Rewriting for Spoken Language Understanding System
- Graph Frequency Analysis of COVID-19 Incidence to Identify County-Level Contagion Patterns in the United States
- Graph Learning Under Spectral Sparsity Constraints
- Graph Neural Network for Large-Scale Network Localization
- Graph Neural Networks for Decentralized Controllers
- Graph Signal Compression via Task-Based Quantization
- Graph Signal Denoising Using Nested-Structured Deep Algorithm Unrolling
- Graph Signal Denoising Via Unrolling Networks
- Graph-Adaptive Incremental Learning Using an Ensemble of Gaussian Process Experts
- Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections Segmentation
- Graph-Homomorphic Perturbations for Private Decentralized Learning
- Graphcomm: A Graph Neural Network Based Method for Multi-Agent Reinforcement Learning
- Graphnet: Graph Clustering with Deep Neural Networks
- Graphon and Graph Neural Network Stability
- Graphspeech: Syntax-Aware Graph Attention Network for Neural Speech Synthesis
- Grid Optimization for Matrix-Based Source Localization Under Inhomogeneous Sensor Topology
- Guaranteed Reconstruction from Integrate-and-Fire Neurons with Alpha Synaptic Activation
- Guided Variational Autoencoder for Speech Enhancement with a Supervised Classifier
- H-GPR: A Hybrid Strategy for Large-Scale Gaussian Process Regression
- HCAG: A Hierarchical Context-Aware Graph Attention Model for Depression Detection
- HCGM-Net: A Deep Unfolding Network for Financial Index Tracking
- HFGCNET: High-Frequency Graph Reasoning for Finer Semantic Image Segmentation
- HIGCNN: Hierarchical Interleaved Group Convolutional Neural Networks for Point Clouds Analysis
- HOCA: Higher-Order Channel Attention for Single Image Super-Resolution
- HSAN: A Hierarchical Self-Attention Network for Multi-Turn Dialogue Generation
- HVS-Based Perceptual Color Compression of Image Data
- Handling Class Imbalance in Low-Resource Dialogue Systems by Combining Few-Shot Classification and Interpolation
- Handwritten Digits Reconstruction from Unlabelled Embeddings
- Hardware Implementation of Iterative Projection-Aggregation Decoding of Reed-Muller Codes
- Head-Synchronous Decoding for Transformer-Based Streaming ASR
- HebbNet: A Simplified Hebbian Learning Framework to do Biologically Plausible Learning
- Heterogeneous two-Stream Network with Hierarchical Feature Prefusion for Multispectral Pan-Sharpening
- Hidden Markov Model Diarisation with Speaker Location Information
- Hide Chopin in the Music: Efficient Information Steganography Via Random Shuffling
- Hierarchical Attention Fusion for Geo-Localization
- Hierarchical Attention-Based Temporal Convolutional Networks for Eeg-Based Emotion Recognition
- Hierarchical Bit-Wise Differential Coding (HBDC) of Point Cloud Attributes
- Hierarchical Coded Elastic Computing
- Hierarchical Context Guided Aggregation Network for Stereo Matching
- Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion Recognition
- Hierarchical Pose Classification for Infant Action Analysis and Mental Development Assessment
- Hierarchical Recurrent Neural Network for Handwritten Strokes Classification
- Hierarchical Refined Attention for Scene Text Recognition
- Hierarchical Similarity Learning for Language-Based Product Image Retrieval
- Hierarchical Speaker-Aware Sequence-to-Sequence Model for Dialogue Summarization
- Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge Distillation
- High Accuracy Tracking of Targets Using Massive MIMO
- High Fidelity Speech Regeneration with Application to Speech Enhancement
- High-Frequency Adversarial Defense for Speech and Audio
- High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC
- High-Throughput VLSI Architecture for Soft-Decision Decoding with ORBGRAND
- Highly Efficient Protection of Biometric Face Samples with Selective JPEG2000 Encryption
- History Utterance Embedding Transformer LM for Speech Recognition
- How Convolutional Neural Networks Deal with Aliasing
- How Phonotactics Affect Multilingual and Zero-Shot ASR Performance
- How Similar or Different is Rakugo Speech Synthesizer to Professional Performers?
- How to Make Text-to-Speech System Pronounce "Voldemort": an Experimental Approach of Foreign Word Phonemization in Vietnamese
- How to Use Time Information Effectively? Combining with Time Shift Module for Lipreading
- Hubert: How Much Can a Bad Teacher Benefit ASR Pre-Training?
- Human-Aware Coarse-to-Fine Online Action Detection
- Human-Centered Favorite Music Classification Using EEG-Based Individual Music Preference Via Deep Time-Series CCA
- Human-Expert-Level Brain Tumor Detection Using Deep Learning with Data Distillation And Augmentation
- Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme Perception
- Hybrid Analog-Digital MIMO Radar Receivers With Bit-Limited ADCs
- Hybrid Beamforming for Wideband OFDM Dual Function Radar Communications
- Hybrid Model for Network Anomaly Detection with Gradient Boosting Decision Trees and Tabtransformer
- Hyperspectral Image Super-Resolution Via Adjacent Spectral Fusion Strategy
- Hypothesis Stitcher for End-to-End Speaker-Attributed ASR on Long-Form Multi-Talker Recordings
- ICA with Orthogonality Constraint: Identifiability And A New Efficient Algorithm
- ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results
- ICASSP 2021 Acoustic Echo Cancellation Challenge: Integrated Adaptive Echo Cancellation with Time Alignment and Deep Learning-Based Residual Echo Plus Noise Suppression
- ICASSP 2021 Deep Noise Suppression Challenge
- ICASSP 2021 Deep Noise Suppression Challenge: Decoupling Magnitude and Phase Optimization with a Two-Stage Deep Network
- ICI-Aware Parameter Estimation for Mimo-Ofdm Radar via Apes Spatial Filtering
- Identification of Deep Breath While Moving Forward Based on Multiple Body Regions and Graph Signal Analysis
- Identification of Uterine Contractions by An Ensemble of Gaussian Processes
- Identifying First-Order Lowpass Graph Signals Using Perron Frobenius Theorem
- Identifying Spammers to Boost Crowdsourced Classification
- Image Coding For Machines: an End-To-End Learned Approach
- Image Coding with Neural Network-Based Colorization
- Image Denoising Based on Correlation Adaptive Sparse Modeling
- Image Generation Based on Texture Guided VAE-AGAN for Regions of Interest Detection in Remote Sensing Images
- Image Steganography Based on Iterative Adversarial Perturbations Onto a Synchronized-Directions Sub-Image
- Image Super-Resolution Using Multi-Resolution Attention Network
- Image-Assisted Transformer in Zero-Resource Multi-Modal Translation
- Impact of Sound Duration and Inactive Frames on Sound Event Detection Performance
- Impact of Speaking Rate on the Source Filter Interaction in Speech: A Study
- Implicit HRTF Modeling Using Temporal Convolutional Networks
- Improved Atomic Norm Based Channel Estimation for Time-Varying Narrowband Leaked Channels
- Improved Data Selection for Domain Adaptation in ASR
- Improved Intra Mode Coding Beyond Av1
- Improved Mask-CTC for Non-Autoregressive End-to-End ASR
- Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer
- Improved Probabilistic Context-Free Grammars for Passwords Using Word Extraction
- Improved Robustness to Disfluencies in Rnn-Transducer Based Speech Recognition
- Improved Step-Size Schedules for Noisy Gradient Methods
- Improved Supervised Training of Physics-Guided Deep Learning Image Reconstruction with Multi-Masking
- Improvements to Prosodic Alignment for Automatic Dubbing
- Improving Audio Anomalies Recognition Using Temporal Convolutional Attention Networks
- Improving Automatic Drum Transcription Using Large-Scale Audio-to-Midi Aligned Data
- Improving Cross-Domain Slot Filling with Common Syntactic Structure
- Improving Deep Learning Sound Events Classifiers Using Gram Matrix Feature-Wise Correlations
- Improving Dialogue Response Generation Via Knowledge Graph Filter
- Improving Entity Recall in Automatic Speech Recognition with Neural Embeddings
- Improving Event Detection by Exploiting Label Hierarchy
- Improving Identification of System-Directed Speech Utterances by Deep Learning of ASR-Based Word Embeddings and Confidence Metrics
- Improving Intraoperative Liver Registration in Image-Guided Surgery with Learning-Based Reconstruction
- Improving Memory Banks for Unsupervised Learning with Large Mini-Batch, Consistency and Hard Negative Mining
- Improving Multimodal Speech Enhancement by Incorporating Self-Supervised and Curriculum Learning
- Improving NER in Social Media via Entity Type-Compatible Unknown Word Substitution
- Improving Naturalness and Controllability of Sequence-to-Sequence Speech Synthesis by Learning Local Prosody Representations
- Improving Neural Text Normalization with Partial Parameter Generator and Pointer-Generator Network
- Improving Pronunciation Assessment Via Ordinal Regression with Anchored Reference Samples
- Improving Prosody Modelling with Cross-Utterance Bert Embeddings for End-to-End Speech Synthesis
- Improving RNN Transducer Modeling for Small-Footprint Keyword Spotting
- Improving RNN Transducer with Target Speaker Extraction and Neural Uncertainty Estimation
- Improving Reconstruction Loss Based Speaker Embedding in Unsupervised and Semi-Supervised Scenarios
- Improving Sound Event Detection Metrics: Insights from DCASE 2020
- Improving Speaker Verification in Reverberant Environments
- Improving Stability of Adversarial Li-ion Cell Usage Data Generation using Generative Latent Space Modelling
- Improving Streaming Automatic Speech Recognition with Non-Streaming Model Distillation on Unsupervised Data
- Improving The Robustness Of Right Whale Detection In Noisy Conditions Using Denoising Autoencoders And Augmented Training
- Improving Ultrasound Tongue Contour Extraction Using U-Net and Shape Consistency-Based Regularizer
- Improving the Classification of Rare Chords With Unlabeled Data
- Improving the Energy-Efficiency of a Kalman Filter Using Unreliable Memories
- Imrnet: An Iterative Motion Compensation and Residual Reconstruction Network for Video Compressed Sensing
- In Situ Calibration of Cross-Sensitive Sensors in Mobile Sensor Arrays Using Fast Informed Non-Negative Matrix Factorization
- In-Bed Pressure-Based Pose Estimation Using Image Space Representation Learning
- Incomplete Multi-View Subspace Clustering with Low-Rank Tensor
- Incorporate Maximum Mean Discrepancy in Recurrent Latent Space for Sequential Generative Model
- Incorporating Syntactic and Phonetic Information into Multimodal Word Embeddings Using Graph Convolutional Networks
- Incorporating Uncertainty In Data Labeling Into Detection of Brain Interictal Epileptiform Discharges From EEG Using Weighted optimization
- Independent Sign Language Recognition with 3d Body, Hands, and Face Reconstruction
- Independent Vector Analysis Using Semi-Parametric Density Estimation via Multivariate Entropy Maximization
- Inertial Proximal Deep Learning Alternating Minimization for Efficient Neutral Network Training
- Inferring High-Resolutional Urban Flow With Internet Of Mobile Things
- Information Decoding and SDR Implementation of DFRC Systems without Training Signals
- Information and Regularization in Restricted Boltzmann Machines
- Injecting Word Information with Multi-Level Word Adapter for Chinese Spoken Language Understanding
- Instance Segmentation with the Number of Clusters Incorporated in Embedding Learning
- Instrument Classification of Solo Sheet Music Images
- Integer Carrier Frequency Offset Estimation in OFDM with Zadoff-Chu Sequences
- Integrated Classification and Localization of Targets Using Bayesian Framework In Automotive Radars
- Integrated Grad-Cam: Sensitivity-Aware Visual Explanation of Deep Convolutional Networks Via Integrated Gradient-Based Scoring
- Integrating Deep Learning with First-Order Logic Programmed Constraints for Zero-Day Phishing Attack Detection
- Integrating End-to-End Neural and Clustering-Based Diarization: Getting the Best of Both Worlds
- Integrating Subgraph-Aware Relation and Direction Reasoning for Question Answering
- Interference Analysis in Reconfigurable Intelligent Surface-Assisted Multiple-Input Multiple-Output Systems
- Intermediate Loss Regularization for CTC-Based Speech Recognition
- Internal Language Model Training for Domain-Adaptive End-To-End Speech Recognition
- Interpolation of Irregularly Sampled Frequency Response Functions Using Convolutional Neural Networks
- Interpreting Glottal Flow Dynamics for Detecting Covid-19 From Voice
- Introducing Deep Reinforcement Learning to Nlu Ranking Tasks
- Investigating Local and Global Information for Automated Audio Captioning with Transfer Learning
- Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech
- Investigating the Efficacy of Music Version Retrieval Systems for Setlist Identification
- Investigation of Fast and Efficient Methods for Multi-Speaker Modeling and Speaker Adaptation
- Iterative Geometry Calibration from Distance Estimates for Wireless Acoustic Sensor Networks
- Iterative Reweighted Algorithms for Joint User Identification and Channel Estimation in Spatially Correlated Massive MTC
- Jamming Strategy Generation for Hidden Communication Modes Via Graph Convolution Networks
- Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language Switching
- Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme Conversion
- Joint Channel, Data, and Phase-Noise Estimation in MIMO-OFDM Systems Using a Tensor Modeling Approach
- Joint Communications with FH-MIMO Radar Systems: An Extended Signaling Strategy
- Joint Coupled Transform Learning Framework for Multimodal Image Super-Resolution
- Joint Dereverberation and Separation With Iterative Source Steering
- Joint Intent Detection and Slot Filling Based on Continual Learning Model
- Joint Learning of Image Aesthetic Quality Assessment and Semantic Recognition Based on Feature Enhancement
- Joint Localization and Predictive Beamforming in Vehicular Networks: Power Allocation Beyond Water-Filling
- Joint Masked CPC And CTC Training For ASR
- Joint Maximum Likelihood Estimation of Power Spectral Densities and Relative Acoustic Transfer Functions for Acoustic Beamforming
- Joint Multi-Pitch Detection and Score Transcription for Polyphonic Piano Music
- Joint Optimization for Full-Duplex Cellular Communications Via Intelligent Reflecting Surface
- Joint Optimization of Spectrally Co-Existing Multi-Carrier Radar and Communication Systems in Cluttered Environments
- Joint Reinforcement Learning and Game Theory Bitrate Control Method for 360-Degree Dynamic Adaptive Streaming
- Jointly Trained Transformers Models for Spoken Language Translation
- KAN: Knowledge-Augmented Networks for Few-Shot Learning
- Kalman Filter Based MIMO CSI Phase Recovery for COTS Wifi Devices
- Kalman Optimizer for Consistent Gradient Descent
- Kalmannet: Data-Driven Kalman Filtering
- Karaoke Key Recommendation Via Personalized Competence-Based Rating Prediction
- Kernearl-Based Lifelong Policy Gradient Reinforcement Learning
- Kernel Orthogonal Nonnegative Matrix Factorization: Application to Multispectral Document Image Decomposition
- Kernel Regression on Graphs in Random Fourier Features Space
- Kernel-Interpolation-Based Filtered-X Least Mean Square for Spatial Active Noise Control In Time Domain
- Kld Minimization-Based Constrained Measurement Filtering For Two-Step TDOA Indoor Tracking
- Knowledge Distillation for Improved Accuracy in Spoken Question Answering
- Knowledge Reasoning for Semantic Segmentation
- Knowledge Transfer for Efficient on-Device False Trigger Mitigation
- Knowledge-Based Chat Detection with False Mention Discrimination
- L-Red: Efficient Post-Training Detection of Imperceptible Backdoor Attacks Without Access to the Training Set
- LIFI: Towards Linguistically Informed Frame Interpolation
- LSSED: A Large-Scale Dataset and Benchmark for Speech Emotion Recognition
- LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation
- Label-Aware Text Representation for Multi-Label Text Classification
- Language Model is all You Need: Natural Language Understanding as Question Answering
- Language-Sensitive Music Emotion Recognition Models: are We Really There Yet?
- Laplacian Regularized Tensor Low-Rank Minimization for Hyperspectral Snapshot Compressive Imaging
- Large Margin Training Improves Language Models for ASR
- Lasaft: Latent Source Attentive Frequency Transformation For Conditioned Source Separation
- Latent Space Motion Analysis for Collaborative Intelligence
- Lattice-Free Mmi Adaptation of Self-Supervised Pretrained Acoustic Models
- Layer-Wise Interpretation of Deep Neural Networks using Identity Initialization
- Leaky Integrator Dynamical Systems and Reachable Sets
- Learned Decimation for Neural Belief Propagation Decoders : Invited Paper
- Learned Transferable Architectures Can Surpass Hand-Designed Architectures for Large Scale Speech Recognition
- Learning Audio Embeddings with User Listening Data for Content-Based Music Recommendation
- Learning Audio-Visual Correlations From Variational Cross-Modal Generation
- Learning Binary Semantic Embedding for Breast Histology Image Classification and Retrieval
- Learning Bollobás-Riordan Graphs Under Partial Observability
- Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags
- Learning Discriminative Features for Semi-Supervised Anomaly Detection
- Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial Training
- Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm
- Learning Double-Compression Video Fingerprints Left From Social-Media Platforms
- Learning From Heterogeneous Eeg Signals with Differentiable Channel Reordering
- Learning Integrodifferential Models for Image Denoising
- Learning Mixed Membership from Adjacency Graph Via Systematic Edge Query: Identifiability and Algorithm
- Learning Model-Blind Temporal Denoisers without Ground Truths
- Learning On Heterogeneous Graphs Using High-Order Relations
- Learning Optimal Lattice Codes for MIMO Communications
- Learning Pose-Adaptive Lip Sync with Cascaded Temporal Convolutional Network
- Learning Representation of Multi-Scale Object for Fine-Grained Image Retrieval
- Learning Separable Time-Frequency Filterbanks for Audio Classification
- Learning Sparse Graph Laplacian with K Eigenvector Prior via Iterative Glasso and Projection
- Learning Sparsifying Transforms for Image Reconstruction in Electrical Impedance Tomography
- Learning Word-Level Confidence for Subword End-To-End ASR
- Learning a Sparse Generative Non-Parametric Supervised Autoencoder
- Learning a Tree of Neural Nets
- Learning the Relevant Substructures for Tasks on Graph Data
- Learning to Continuously Optimize Wireless Resource in Episodically Dynamic Environment
- Learning to Estimate Kernel Scale and Orientation of Defocus Blur with Asymmetric Coded Aperture
- Learning to Select Context in a Hierarchical and Global Perspective for Open-Domain Dialogue Generation
- Learning to Select for Mimo Radar Based on Hybrid Analog-Digital Beamforming
- Learning-Based Lossless Compression of 3D Point Cloud Geometry
- Length No Longer Matters: A Real Length Adaptive Arrhythmia Classification Model with Multi-Scale Convolution
- Less is More: Improved RNN-T Decoding Using Limited Label Context and Path Merging
- Leveraging A Multiple-Strain Model with Mutations in Analyzing the Spread of Covid-19
- Leveraging Acoustic and Linguistic Embeddings from Pretrained Speech and Language Models for Intent Classification
- Leveraging the Structure of Musical Preference in Content-Aware Music Recommendation
- Light Field Style Transfer with Local Angular Consistency
- Light-TTS: Lightweight Multi-Speaker Multi-Lingual Text-to-Speech
- Lightspeech: Lightweight and Fast Text to Speech with Neural Architecture Search
- Lightweight Dual-Task Networks For Crowd Counting In Aerial Images
- Lightweight Human Pose Estimation under Resource-Limited Scenes
- Lightweight Non-Local Network for Image Super-Resolution
- Lightweight and Accurate Single Image Super-Resolution with Channel Segregation Network
- Lightweight and Interpretable Neural Modeling of an Audio Distortion Effect Using Hyperconditioned Differentiable Biquads
- Linear Computation Coding
- Linear Multichannel Blind Source Separation based on Time-Frequency Mask Obtained by Harmonic/Percussive Sound Separation
- Litesing: Towards Fast, Lightweight and Expressive Singing Voice Synthesis
- Locally Optimal Detection of Stochastic Targeted Universal Adversarial Perturbations
- Long-Short Temporal Modeling for Efficient Action Recognition
- Looking Through Walls: Inferring Scenes from Video-Surveillance Encrypted Traffic
- Loopnet: Musical Loop Synthesis Conditioned on Intuitive Musical Parameters
- Low Complexity SLM for OFDMA System with Implicit Side Information
- Low Complexity Secure P-Tensor Product Compressed Sensing Reconstruction Outsourcing and Identity Authentication in Cloud
- Low Latency Online Blind Source Separation Based on Joint Optimization with Blind Dereverberation
- Low Mutual Coupling Sparse Array Design Using ULA Fitting
- Low Resource Audio-To-Lyrics Alignment from Polyphonic Music Recordings
- Low-Complexity Parameter Learning for OTFS Modulation Based Automotive Radar
- Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On Percepnet
- Low-Dimensional Denoising Embedding Transformer for ECG Classification
- Low-Latency Polar Decoder Using Overlapped SCL Processing
- Low-Rank and Sparse Decomposition for Joint DOA Estimation and Contaminated Sensors Detection with Sparsely Contaminated Arrays
- Low-Rank on Graphs Plus Temporally Smooth Sparse Decomposition for Anomaly Detection in Spatiotemporal Data
- Low-Resource Expressive Text-To-Speech Using Data Augmentation
- Ltaf-Net: Learning Task-Aware Adaptive Features and Refining Mask for Few-Shot Semantic Segmentation
- MAEC: Multi-Instance Learning with an Adversarial Auto-Encoder-Based Classifier for Speech Emotion Recognition
- MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-Training
- MBNET: MOS Prediction for Synthesized Speech with Mean-Bias Network
- MCR-NET: A Multi-Step Co-Interactive Relation Network for Unanswerable Questions on Machine Reading Comprehension
- MPDNet: A 3D Missing Part Detection Network Based on Point Cloud Segmentation
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.