ICASSP 2020 Accepted Papers
The full list of 1,849 papers accepted at ICASSP 2020 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- Language-Agnostic Multilingual Modeling
- Laplace State Space Filter with Exact Inference and Moment Matching
- Large Dimensional Asymptotics of Multi-Task Learning
- Large-Context Pointer-Generator Networks for Spoken-to-Written Style Conversion
- Large-Scale Fading Precoding for Maximizing the Product of SINRs
- Large-Scale Time Series Clustering with k-ARs
- Large-Scale Unsupervised Pre-Training for End-to-End Spoken Language Understanding
- Large-Scale Weakly-Supervised Content Embeddings for Music Recommendation and Tagging
- Latency-Minimized Design of secure transmissions in UAV-Aided Communications
- Latent Fused Lasso
- Lattice-Based Improvements for Voice Triggering Using Graph Neural Networks
- Layer-Normalized LSTM for Hybrid-Hmm and End-To-End ASR
- Learn-By-Calibrating: Using Calibration As A Training Objective
- Learned Lossless Image Compression with A Hyperprior and Discretized Gaussian Mixture Likelihoods
- Learning A Common Granger Causality Network Using A Non-Convex Regularization
- Learning Asr-Robust Contextualized Embeddings for Spoken Language Understanding
- Learning Based Reconfigurable Sub-nyquist Sampling Framework for Ultra-wideband Angular Sensing
- Learning Blind Denoising Network for Noisy Image Deblurring
- Learning Data Representation and Emotion Assessment from Physiological Data
- Learning Differentiable Sparse and Low Rank Networks for Audio-Visual Object Localization
- Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions
- Learning Domain Invariant Representations for Child-Adult Classification from Speech
- Learning Eating Environments Through Scene Clustering
- Learning Endmember Dynamics in Multitemporal Hyperspectral Data Using A State-Space Model Formulation
- Learning Geometric Features with Dual-stream CNN for 3D Action Recognition
- Learning Graph Influence from Social Interactions
- Learning Local Structure of Representative Points for Point Cloud Classification and Semantic Segmentation
- Learning Multi-Scale Attentive Features for Series Photo Selection
- Learning Network Representation Through Reinforcement Learning
- Learning Noise Invariant Features Through Transfer Learning For Robust End-to-End Speech Recognition
- Learning Partial Differential Equations From Data Using Neural Networks
- Learning Perception and Planning With Deep Active Inference
- Learning Plug-And-Play Proximal Quasi-Newton Denoisers
- Learning Product Graphs from Multidomain Signals
- Learning Recurrent Neural Network Language Models With Context-Sensitive Label Smoothing for Automatic Speech Recognition
- Learning Sampling and Model-Based Signal Recovery for Compressed Sensing MRI
- Learning Semi-Supervised Anonymized Representations by Mutual Information
- Learning Signed Graphs from Data
- Learning Spatio-Temporal Convolutional Network for Real-Time Object Tracking
- Learning Spatio-Temporal Representations With Temporal Squeeze Pooling
- Learning Spectral-Spatial Prior Via 3DDNCNN for Hyperspectral Image Deconvolution
- Learning Task-Based Analog-to-Digital Conversion for MIMO Receivers
- Learning With Out-of-Distribution Data for Audio Classification
- Learning a Generic Adaptive Wavelet Shrinkage Function for Denoising
- Learning a Representation for Cover Song Identification Using Convolutional Neural Network
- Learning a Subword Inventory Jointly with End-to-End Automatic Speech Recognition
- Learning connectivity and higher-order interactions in radial distribution grids
- Learning from Dances: Pose-Invariant Re-Identification for Multi-Person Tracking
- Learning the Helix Topology of Musical Pitch
- Learning the Spatio-Temporal Dynamics of Physical Processes from Partial Observations
- Learning to Characterize Adversarial Subspaces
- Learning to Detect Keyword Parts and Whole by Smoothed Max Pooling
- Learning to Estimate Driver Drowsiness from Car Acceleration Sensors Using Weakly Labeled Data
- Learning to Fool the Speaker Recognition
- Learning to Generate Diverse Questions from Keywords
- Learning to Rank Music Tracks Using Triplet Loss
- Learning to Separate Sounds from Weakly Labeled Scenes
- Learning-Aided Content Placement in Caching-Enabled fog Computing Systems Using Thompson Sampling
- Learning-Based Content Caching and User Clustering: A Deep Deterministic Policy Gradient Approach
- Least-Squares DOA Estimation with an Informed Phase Unwrapping and Full Bandwidth Robustness
- Levenberg-Marquardt and Line-Search Extended Kalman Smoothers
- Leveraging Cuboids for Better Motion Modeling in High Efficiency Video Coding
- Leveraging Gans to Improve Continuous Path Keyboard Input Models
- Leveraging Ordinal Regression With Soft Labels For 3d Head Pose Estimation From Point Sets
- Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems
- Libri-Adapt: a New Speech Dataset for Unsupervised Domain Adaptation
- Libri-Light: A Benchmark for ASR with Limited or No Supervision
- Lie Group State Estimation via Optimal Transport
- Lifter Training and Sub-Band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials
- Light-Field Reconstruction and Depth Estimation from Focal Stack Images Using Convolutional Neural Networks
- Lightdet: A Lightweight and Accurate Object Detection Network
- Lightweight Hardware Implementation of VVC Transform Block for ASIC Decoder
- Lightweight V-Net for Liver Segmentation
- Lightweight and Efficient End-To-End Speech Recognition Using Low-Rank Transformer
- Limitations of Weak Labels for Embedding and Tagging
- Line Spectral Estimation with Palindromic Kernels
- Linear Model-Based Intra Prediction in VVC Test Model
- Linear Speedup in Saddle-Point Escape for Decentralized Non-Convex Optimization
- Linear Thompson Sampling Under Unknown Linear Constraints
- Lipreading Using Temporal Convolutional Networks
- Load Management with Predictions of Solar Energy Production for Cloud Data Centers
- Local Key Estimation In Classical Music Recordings: A Cross-Version Study on Schubert's Winterreise
- Local-Global Feature for Video-Based One-Shot Person Re-Identification
- Location-Relative Attention Mechanisms for Robust Long-Form Speech Synthesis
- Look Globally, Age Locally: Face Aging With an Attention Mechanism
- Lookahead Converges to Stationary Points of Smooth Non-convex Functions
- Looking Enhances Listening: Recovering Missing Speech Using Images
- Low Complexity NLMS for Multiple Loudspeaker Acoustic ECHO Canceller Using Relative Loudspeaker Transfer Functions
- Low Complexity Single Image Super-Resolution with Channel Splitting and Fusion Network
- Low Mutual and Average Coherence Dictionary Learning Using Convex Approximation
- Low Rank Activations for Tensor-Based Convolutional Sparse Coding
- Low-Complexity 5g Slam with CKF-PHD Filter
- Low-Complexity Accurate Mmwave Positioning for Single-Antenna Users Based on Angle-of-Departure and Adaptive Beamforming
- Low-Complexity Compressed Alignment-Aided Compressive Analysis for Real-Time Electrocardiography Telemonitoring
- Low-Complexity Fixed-Point Convolutional Neural Networks For Automatic Target Recognition
- Low-Complexity LSTM-Assisted Bit-Flipping Algorithm For Successive Cancellation List Polar Decoder
- Low-Complexity Levenberg-Marquardt Algorithm for Tensor Canonical Polyadic Decomposition
- Low-Complexity and Reliable Transforms for Physical Unclonable Functions
- Low-Frequency Compensated Synthetic Impulse Responses For Improved Far-Field Speech Recognition
- Low-Latency Lightweight Streaming Speech Recognition with 8-Bit Quantized Simple Gated Convolutional Neural Networks
- Low-Latency Single Channel Speech Enhancement Using U-Net Convolutional Neural Networks
- Low-Rank Approximation of Matrices Via A Rank-Revealing Factorization with Randomization
- Low-Rank Gradient Approximation for Memory-Efficient on-Device Training of Deep Neural Network
- Low-Rank MMWAVE MIMO Channel Estimation in One-Bit Receivers
- Low-Rank Tensor Ring Model for Completing Missing Visual Data
- Low-Rank Toeplitz Matrix Estimation Via Random Ultra-Sparse Rulers
- Low-Tubal-Rank Tensor Recovery From One-Bit Measurements
- Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers
- Lqaid: Localized Quality Aware Image Denoising Using Deep Convolutional Neural Networks
- Lupulus: A Flexible Hardware Accelerator for Neural Networks
- M-Estimators of Scatter with Eigenvalue Shrinkage
- MDR-SURV: A Multi-Scale Deep Learning-Based Radiomics for Survival Prediction in Pulmonary Malignancies
- ML and EM Estimation of Sampling Intervals of Sensor Devices
- MMSE-Based Channel Estimation for Hybrid Beamforming Massive MIMO with Correlated Channels
- MSPNET: Multi-Supervised Parallel Network for Crowd Counting
- Mahalanobis Distance Based Adversarial Network for Anomaly Detection
- Manet: Multi-Scale Aggregated Network For Light Field Depth Estimation
- Mango: A Python Library for Parallel Hyperparameter Tuning
- Manifold Gradient Descent Solves Multi-Channel Sparse Blind Deconvolution Provably and Efficiently
- Many-To-Many Voice Conversion Using Conditional Cycle-Consistent Adversarial Networks
- Mask-Dependent Phase Estimation for Monaural Speaker Separation
- Masking and Inpainting: A Two-Stage Speech Enhancement Approach for Low SNR and Non-Stationary Noise
- Matching Pursuit Based Dynamic Phase-Amplitude Coupling Measure
- Maximally Energy-Concentrated Differential Window for Phase-Aware Signal Processing Using Instantaneous Frequency
- Maximum Likelihood Estimation of the Interference-Plus-Noise Cross Power Spectral Density Matrix for Own Voice Retrieval
- Maximum Likelihood Multi-Speaker Direction of Arrival Estimation Utilizing a Weighted Histogram
- Maxpolynomial Division with Application To Neural Network Simplification
- Media Classification with Bayesian Optimization and Vapnik-Chervonenkis (VC) Bounds
- Mellotron: Multispeaker Expressive Voice Synthesis by Conditioning on Rhythm, Pitch and Global Style Tokens
- Mental Fatigue Prediction from Multi-Channel ECOG Signal
- Message Transmission ThroughUnderspread Time-Varying Linear Channels
- Meta Learning for End-To-End Low-Resource Speech Recognition
- Meta Metric Learning for Highly Imbalanced Aerial Scene Classification
- Meta-Learning Extractors for Music Source Separation
- Meta-Learning for Robust Child-Adult Classification from Speech
- Meta-Learning to Communicate: Fast End-to-End Training for Fading Channels
- Metric Learning with Background Noise Class for Few-Shot Detection of Rare Sound Events
- Metric Representations of Networks: A Uniqueness Result
- Minimal Adversarial Perturbations in Mobile Health Applications: The Epileptic Brain Activity Case Study
- Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR
- Mining Effective Negative Training Samples for Keyword Spotting
- Mirrored Arrays for Direction-of-Arrival Estimation
- Misspecified Cramer-Rao Bound For Delay Estimation with a Mismatched Waveform: A Case Study
- Mixture Factorized Auto-Encoder for Unsupervised Hierarchical Deep Factorization of Speech Signal
- Mixup Multi-Attention Multi-Tasking Model for Early-Stage Leukemia Identification
- Mixup-breakdown: A Consistency Training Method for Improving Generalization of Speech Separation Models
- MoGA: Searching Beyond Mobilenetv3
- Mobility-Aware Beam Steering in Metasurface-Based Programmable Wireless Environments
- Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
- Model Order Selection in DoA Scenarios via Cross-entropy Based Machine Learning Techniques
- Modeling Behavior as Mutual Dependency between Physiological Signals and Indoor Location in Large-Scale Wearable Sensor Study
- Modeling Behavioral Consistency in Large-Scale Wearable Recordings of Human Bio-Behavioral Signals
- Modeling Piece-Wise Stationary Time Series
- Modeling Plate and Spring Reverberation Using A DSP-Informed Deep Neural Network
- Modeling Uncertainty in Predicting Emotional Attributes from Spontaneous Speech
- Modeling the Environment in Deep Reinforcement Learning: The Case of Energy Harvesting Base Stations
- Modelling Sea Clutter In Sar Images Using Laplace-Rician Distribution
- Monaural Speech Enhancement Using Intra-Spectral Recurrent Layers in the Magnitude and Phase Responses
- Motion Dynamics Improve Speaker-Independent Lipreading
- Motion Feedback Design for Video Frame Interpolation
- Mspec-Net : Multi-Domain Speech Conversion Network
- Mt-Gcn For Multi-Label Audio Tagging With Noisy Labels
- Multi Image Depth from Defocus Network with Boundary Cue for Dual Aperture Camera
- Multi-Agent Deep Reinforcement Learning For Distributed Handover Management In Dense MmWave Networks
- Multi-Branch Learning for Weakly-Labeled Sound Event Detection
- Multi-Channel Speech Source Separation and Dereverberation With Sequential Integration of Determined and Underdetermined Models
- Multi-Conditioning and Data Augmentation Using Generative Noise Model for Speech Emotion Recognition in Noisy Conditions
- Multi-Depth Computational Periscopy with an Ordinary Camera
- Multi-Head Attention for Speech Emotion Recognition with Auxiliary Learning of Gender Recognition
- Multi-Label Consistent Convolutional Transform Learning: Application to Non-Intrusive Load Monitoring
- Multi-Label Sound Event Retrieval Using A Deep Learning-Based Siamese Structure With A Pairwise Presence Matrix
- Multi-Layer Content Interaction Through Quaternion Product for Visual Question Answering
- Multi-Level Deep Neural Network Adaptation for Speaker Verification Using MMD and Consistency Regularization
- Multi-Microphone Complex Spectral Mapping for Speech Dereverberation
- Multi-Modal Self-Supervised Pre-Training for Joint Optic Disc and Cup Segmentation in Eye Fundus Images
- Multi-MotifGAN (MMGAN): Motif-Targeted Graph Generation And Prediction
- Multi-Patch Aggregation Models for Resampling Detection
- Multi-Polarization Information Fusion for Object Contour Display in Passive Millimeter-Wave and Terahertz Security Imaging
- Multi-Resolution Multi-Head Attention in Deep Speaker Embedding
- Multi-Resolution Overlapping Stripes Network for Person Re-Identification
- Multi-Scale Deep Feature Fusion for Vehicle Re-Identification
- Multi-Scale Octave Convolutions for Robust Speech Recognition
- Multi-Scale Residual Network for Image Classification
- Multi-Speaker and Multi-Domain Emotional Voice Conversion Using Factorized Hierarchical Variational Autoencoder
- Multi-Stage Residual Hiding for Image-Into-Audio Steganography
- Multi-Step Online Unsupervised Domain Adaptation
- Multi-Task Center-Of-Pressure Metrics Estimation from Skeleton Using Graph Convolutional Network
- Multi-Task Learning Via SA-FPN and EJ-Head
- Multi-Task Learning for Speaker Verification and Voice Trigger Detection
- Multi-Task Learning for Voice Trigger Detection
- Multi-Task Learning in Autonomous Driving Scenarios Via Adaptive Feature Refinement Networks
- Multi-Task Self-Supervised Learning for Robust Speech Recognition
- Multi-Time-Scale Convolution for Emotion Recognition from Speech Audio Signals
- Multi-View Bayesian Generative Model for Multi-Subject FMRI Data on Brain Decoding of Viewed Image Categories
- Multi-View Clustering Via Mixed Embedding Approximation
- Multi-View Shape Estimation of Transparent Containers
- Multi-View Wasserstein Discriminant Analysis with Entropic Regularized Wasserstein Distance
- Multi-Way Multi-View Deep Autoencoder for Image Feature Learning with Multi-Level Graph Regularization
- Multi-constraint Spectral Co-design for Colocated MIMO Radar and MIMO Communications
- Multichannel Active Noise Control with Spatial Derivative Constraints to Enlarge the Quiet Zone
- Multichannel Signal Classification Using Vector Autoregression
- Multichannel Signal Processing for Road Surface Identification
- Multigraph Spectral Clustering for Joint Content Delivery and Scheduling in Beam-Free Satellite Communications
- Multilinear Generalized Singular Value Decomposition (Ml-gsvd) with Application to Coordinated Beamforming in Multi-user Mimo Systems
- Multilingual Acoustic Word Embedding Models for Processing Zero-resource Languages
- Multilingual Grapheme-To-Phoneme Conversion with Byte Representation
- Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing
- Multimodal Learning for Classroom Activity Detection
- Multimodal Speaker Diarization of Real-World Meetings Using D-Vectors With Spatial Features
- Multimodal Transformer Fusion for Continuous Emotion Recognition
- Multimodal Violence Detection in Videos
- Multiple Points Input For Convolutional Neural Networks in Replay Attack Detection
- Multispectral Fusion of RGB and NIR Images Using Weighted Least Squares and Alternating Guidance
- Multistate Encoding with End-To-End Speech RNN Transducer Network
- Multitaper Spectral Granger Causality with Application to Ssvep
- Multitask Learning and Multistage Fusion for Dimensional Audiovisual Emotion Recognition
- Multitask Learning with Capsule Networks for Speech-to-Intent Applications
- Multiuser Massive Mimo Downlink Precoding Using Second-Order Spatial Sigma-Delta Modulation
- Multivariate Tropical Regression and Piecewise-Linear Surface Fitting
- Mutual-Information-Based Sensor Placement for Spatial Sound Field Recording
- Nasil: Neural Architecture Search with Imitation Learning
- Near Capacity RCQD Constellations for PAPR Reduction of OFDM Systems
- Near-Optimal Interference Exploitation 1-Bit Massive MIMO Precoding Via Partial Branch-and-Bound
- Nearest Kronecker Product Decomposition Based Normalized Least Mean Square Algorithm
- Neural Attentive Multiview Machines
- Neural Coding Strategies for Event-Based Vision Data
- Neural Lattice Search for Speech Recognition
- Neural Network Training with Approximate Logarithmic Computations
- Neural Network Wiretap Code Design for Multi-Mode Fiber Optical Channels
- Neural Oracle Search on N-BEST Hypotheses
- Neural Percussive Synthesis Parameterised by High-Level Timbral Features
- Neural Time Warping for Multiple Sequence Alignment
- Neutral to Lombard Speech Conversion with Deep Learning
- New Metrics for Evaluating the Accuracy of Fundamental Frequency Estimation Approaches in Musical Signals
- No-Regret Non-Convex Online Meta-Learning
- Node-Asynchronous Spectral Clustering On Directed Graphs
- Noise-Robust Key-Phrase Detectors for Automated Classroom Feedback
- Non-Experts or Experts? Statistical Analyses of MOS using DSIS method
- Non-Gaussian BLE-Based Indoor Localization Via Gaussian Sum Filtering Coupled with Wasserstein Distance
- Non-Griffin-Lim Type Signal Recovery from Magnitude Spectrogram
- Non-Local Nested Residual Attention Network for Stereo Image Super-Resolution
- Non-Uniform Video Time-Lapse Method Based on Motion Scenario and Stabilization Constraint
- Non-parametric Community Change-points Detection in Streaming Graph Signals
- Noncoherent Maximum-Likelihood Detection for Ambient Backscattering Communications Over Ambient OFDM Signals
- Nonlinear Spatial Filtering for Multichannel Speech Enhancement in Inhomogeneous Noise Fields
- Normalized Least-Mean-Square Algorithms with Minimax Concave Penalty
- OH, JEEZ! or UH-HUH? A Listener-Aware Backchannel Predictor on ASR Transcriptions
- OOV Recovery with Efficient 2nd Pass Decoding and Open-vocabulary Word-level RNNLM Rescoring for Hybrid ASR
- Object Detection and 3d Estimation Via an FMCW Radar Using a Fully Convolutional Network
- Object Detection with Color and Depth Images with Multi-Reduced Region Proposal Network and Multi-Pooling
- Object Surface Estimation from Radar Images
- Objective Bayesian Detection Under Spatially Correlated Gaussian Observations for Multi-Antenna Cognitive Radio Network
- On Binary Sequence Set Design with Applications to Automotive Radar
- On Cramér-Rao Lower Bounds with Random Equality Constraints
- On Design of Optimal Smart Meter Privacy Control Strategy Against Adversarial Map Detection
- On Distributed Stochastic Gradient Algorithms for Global Optimization
- On Distributed Stochastic Gradient Descent for Nonconvex Functions in the Presence of Byzantines
- On Divergence Approximations for Unsupervised Training of Deep Denoisers Based on Stein's Unbiased Risk Estimator
- On End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments
- On Exponentially Consistency of Linkage-Based Hierarchical Clustering Algorithm Using Kolmogrov-Smirnov Distance
- On Harmonic Approximations of Inharmonic Signals
- On Measuring Doppler Shifts between Tags in a Backscattering Tag-to-Tag Network with Applications in Tracking
- On Modeling ASR Word Confidence
- On Network Science and Mutual Information for Explaining Deep Neural Networks
- On Polar Coding For Finite Blocklength Secret Key Generation Over Wireless Channels
- On Regularization Parameter for L0-Sparse Covariance Fitting Based DOA Estimation
- On Robust Variance Filtering and Change of Variance Detection
- On The Choice of Graph Neural Network Architectures
- On The Degrees Of Freedom in Total Variation Minimization
- On The Frequency Domain Detection of High Dimensional Time Series
- On The Impact of Language Familiarity in Talker Change Detection
- On The Stability of Polynomial Spectral Graph Filters
- On Throughput of Millimeter Wave MIMO Systems with Low Resolution ADCs
- On the Byzantine Robustness of Clustered Federated Learning
- On the Effect of Reflectance on Phasor Field Non-Line-of-Sight Imaging
- On the Importance of Vocal Tract Constriction for Speaker Characterization: The Whispered Speech Study
- On the Limit Distribution of the Canonical Correlation Coefficients Between the Past and the Future of a High-Dimensional White Noise
- On the Opportunistic use of Commercial Ku and Ka Band Satcom Networks for Rain Rate Estimation: Potentials and Critical Issues
- On the Use of Rényi Entropy for Optimal Window Size Computation in the Short-Time Fourier Transform
- On-The-Fly Feature Selection and Classification with Application to Civic Engagement Platforms
- One-Bit Compressed Sensing Using Generative Models
- One-Bit DoA Estimation via Sparse Linear Arrays
- One-Bit Normalized Scatter Matrix Estimation For Complex Elliptically Symmetric Distributions
- One-Bit Sampling in Fractional Fourier Domain
- One-Shot Parametric Audio Production Style Transfer with Application to Frequency Equalization
- One-Shot Voice Conversion Using Star-Gan
- One-Shot Voice Conversion by Vector Quantization
- Online Channel Estimation for Hybrid Beamforming Architectures
- Online Community Detection by Spectral Cusum
- Online Graph Topology Inference with Kernels For Brain Connectivity Estimation
- Online Positron Emission Tomography By Online Portfolio Selection
- Online Tensor Completion and Free Submodule Tracking With The T-SVD
- Open Set Video Camera Model Verification
- Opendenoising: An Extensible Benchmark for Building Comparative Studies of Image Denoisers
- Opportunistic use of GNSS Signals to Characterize the Environment by Means of Machine Learning Based Processing
- Optimal Design of Energy-Efficient Cell-Free Massive Mimo: Joint Power Allocation and Load Balancing
- Optimal Joint Channel Estimation and Data Detection by L1-norm PCA for Streetscape IoT
- Optimal Laplacian Regularization for Sparse Spectral Community Detection
- Optimal Power Flow Using Graph Neural Networks
- Optimal Sampling Rate and Bandwidth of Bandlimited Signals - an Algorithmic Perspective
- Optimal Transport Based Change Point Detection and Time Series Segment Clustering
- Optimal Transport Structure of CycleGAN for Unsupervised Learning for Inverse Problems
- Optimal Window Design for Joint Spatial-Spectral Domain Filtering of Signals on the Sphere
- Optimal Window Design for W-OFDM
- Optimized Sensor Selection for Joint Radar-communication Systems
- Optimized Single Carrier Transceiver for Future Sub-TeraHertz Applications
- Optimizing Backscattering Coefficient Design for Minimizing BER at Monostatic MIMO reader
- Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge
- Optimum Kernel Particle Filter for Asymmetric Laplace Noise
- Ordinal Learning for Emotion Recognition in Customer Service Calls
- Orthogonal Training for Text-Independent Speaker Verification
- Overcoming High Nanopore Basecaller Error Rates for DNA Storage via Basecaller-Decoder Integration and Convolutional Codes
- Overdetermined Independent Vector Analysis
- Overlap Local-SGD: An Algorithmic Approach to Hide Communication Delays in Distributed SGD
- Overlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech Detection
- Overlapped State Hidden Semi-Markov Model for Grouped Multiple Sequences
- PAGAN: A Phase-Adapted Generative Adversarial Networks for Speech Enhancement
- PEVD-Based Speech Enhancement in Reverberant Environments
- Paco and Paco-Dct: Patch Consensus and Its Application To Inpainting
- Pan: Phoneme-Aware Network for Monaural Speech Enhancement
- Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram
- Parallelizing Adam Optimizer with Blockwise Model-Update Filtering
- Parameter Estimation of In-City Frontal Rainfall Propagation
- Parsing Map Guided Multi-Scale Attention Network For Face Hallucination
- Partial AUC Optimization Based Deep Speaker Embeddings with Class-Center Learning for Text-Independent Speaker Verification
- Particle Filter with Rejection Control and Unbiased Estimator of the Marginal Likelihood
- Particle Filtering on the Complex Stiefel Manifold with Application to Subspace Tracking
- Particle Group Metropolis Methods for Tracking the Leaf Area Index
- Passive Intelligent Surface Assisted MIMO Powered Sustainable IoT
- Patch-Level Selection and Breadth-First Prediction Strategy for Reversible Data Hiding
- Pathloss Prediction using Deep Learning with Applications to Cellular Optimization and Efficient D2D Link Scheduling
- Peer To Peer Offloading With Delayed Feedback: An Adversary Bandit Approach
- Perception-Distortion Trade-Off with Restricted Boltzmann Machines
- Perceptual loss function for neural modeling of audio systems
- Performance Analysis for Path Attenuation Estimation of Microwave Signals Due to Rainfall and Beyond
- Performance Bounds for Displaced Sensor Automotive Radar Imaging
- Performance Comparison of Lossless Compression Strategies for Dynamic Vision Sensor Data
- Performance Study of a Convolutional Time-Domain Audio Separation Network for Real-Time Speech Denoising
- Person Identification Using Deep Convolutional Neural Networks on Short-Term Signals from Wearable Sensors
- Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks
- Phoneme Boundary Detection Using Learnable Segmental Features
- Phonetic Feedback for Speech Enhancement with and Without Parallel Speech Data
- Phylogenetic Minimum Spanning Tree Reconstruction Using Autoencoders
- Pitch Estimation Via Self-Supervision
- Pitchnet: Unsupervised Singing Voice Conversion with Pitch Adversarial Network
- Pixel-Level Self-Paced Learning For Super-Resolution
- Pixel-Wise Linear/Nonlinear Nonnegative Matrix Factorization for Unmixing of Hyperspectral Data
- Playing Technique Recognition by Joint Time-Frequency Scattering
- Polarization Parameters Estimation with Scalar Sensor Arrays
- Polarizing Front Ends for Robust Cnns
- Polyphonic Sound Event Detection Using Transposed Convolutional Recurrent Neural Network
- Portfolio Cuts: A Graph-Theoretic Framework to Diversification
- Position Constraint Loss For Fashion Landmark Estimation
- Positive Semidefinite Matrix Factorization: A Link to Phase Retrieval And A Block Gradient Algorithm
- Power Spectrum Optimization for Capacity of the Extended Spectrum Hybrid Fiber Coax Network
- Pre-Training for Query Rewriting in a Spoken Language Understanding System
- Preconditioned Ghost Imaging Via Sparsity Constraint
- Preconditioning ADMM for Fast Decentralized Optimization
- Predicting Performance Outcome with a Conversational Graph Convolutional Network for Small Group Interactions
- Predicting Word Error Rate for Reverberant Speech
- Prediction of Individual Progression Rate in Parkinson's Disease Using Clinical Measures and Biomechanical Measures of Gait and Postural Stability
- Prediction of Voicing and the F0 Contour from Electromagnetic Articulography Data for Articulation-to-Speech Synthesis
- Prediction oof Vessel Trajectories From AIS Data Via Sequence-To-Sequence Recurrent Neural Networks
- Preference-Aware Mask for Session-Based Recommendation with Bidirectional Transformer
- Preservation of Anomalous Subgroups On Variational Autoencoder Transformed Data
- Primal-Dual Stochastic Subgradient Method For Log-Determinant Optimization
- Primary Path Estimator Based on Individual Secondary Path for ANC Headphones
- Principal Angle Detector for Subspace Signal with Structured Unknown Interference
- Principle-Inspired Multi-Scale Aggregation Network for Extremely Low-Light Image Enhancement
- Privacy Aware Acoustic Scene Synthesis Using Deep Spectral Feature Inversion
- Privacy-Aware Quickest Change Detection
- Privacy-Preserving Image Sharing Via Sparsifying Layers on Convolutional Groups
- Privacy-Preserving Pattern Recognition Using Encrypted Sparse Representations in L0 Norm Minimization
- Privacy-Preserving Phishing Web Page Classification Via Fully Homomorphic Encryption
- Private FL-GAN: Differential Privacy Synthetic Data Generation Based on Federated Learning
- Probabilistic Filter and Smoother for Variational Inference of Bayesian Linear Dynamical Systems
- Processing Convolutional Neural Networks on Cache
- Programmable Dataflow Accelerators: A 5G OFDM Modulation/Demodulation Case Study
- Progressive Multi-Target Network Based Speech Enhancement with Snr-Preselection for Robust Speaker Diarization
- Projected Weight Regularization to Improve Neural Network Generalization
- Projection Free Dynamic Online Learning
- Propeller Noise Detection with Deep Learning
- Prototypical Networks for Small Footprint Text-Independent Speaker Verification
- Proximal Distance Algorithm for Nonconvex QCQP with Beamforming Applications
- Proximal Multitask Learning Over Distributed Networks with Jointly Sparse Structure
- Pseudo Labeling and Negative Feedback Learning for Large-Scale Multi-Label Domain Classification
- Pseudo Likelihood Correction Technique for Low Resource Accented ASR
- Pyannote.Audio: Neural Building Blocks for Speaker Diarization
- Q-GADMM: Quantized Group ADMM for Communication Efficient Decentralized Machine Learning
- Q-Learning Based Predictive Relay Selection for Optimal Relay Beamforming
- QOS-Aware Flow Control for Power-Efficient Data Center Networks with Deep Reinforcement Learning
- Quality-of-Service Prediction for Physical-layer Security via Secrecy Maps
- Quantized Tensor Robust Principal Component Analysis
- Quartznet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions
- Quickest Change Detection In Anonymous Heterogeneous Sensor Networks
- Quickest Detection of Growing Dynamic Anomalies in Networks
- REV-AE: A Learned Frame Set for Image Reconstruction
- ROIMIX: Proposal-Fusion Among Multiple Images for Underwater Object Detection
- Rate Assignment in 360-Degree Video Tiled Streaming Using Random Forest Regression
- Rate-Invariant Autoencoding of Time-Series
- Raw Waveform Based End-to-end Deep Convolutional Network for Spatial Localization of Multiple Acoustic Sources
- Ray Separation and Source Depth Estimation Based on Sound Pressure Field Transformation
- Re-Translation Strategies for Long Form, Simultaneous, Spoken Language Translation
- Real-Time Binaural Speech Separation with Preserved Spatial Cues
- Real-Time Hand Gesture Recognition Using Temporal Muscle Activation Maps of Multi-Channel Semg Signals
- Real-Time Implementation Aspects of Large Intelligent Surfaces
- Real-Time Speech Enhancement Using Equilibriated RNN
- Real-Time Task Offloading for Large-Scale Mobile Edge Computing
- Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems
- Realizability of Planar Point Embeddings from Angle Measurements
- Receiver Design and AGC optimization with Self Interference Induced Saturation
- Receptive Field Pyramid Network for Object Detection
- Reconstruction of Fri Signals Using Deep Neural Network Approaches
- Recurrent Neural Audiovisual Word Embeddings for Synchronized Speech and Real-Time Mri
- Recursive Prediction of Graph Signals With Incoming Nodes
- Reduced-Complexity Singular Value Decomposition For Tucker Decomposition: Algorithm And Hardware
- Redundant Convolutional Network With Attention Mechanism For Monaural Speech Enhancement
- Reflectance-Guided, Contrast-Accumulated Histogram Equalization
- Regression Before Classification for Temporal Action Detection
- Regularized Beamformer for the Spherical Microphone Array to Cope with the White Noise Amplification
- Regularized Fast Multichannel Nonnegative Matrix Factorization with ILRMA-Based Prior Distribution of Joint-Diagonalization Process
- Regularized Partial Phase Synchrony Index Applied to Dynamical Functional Connectivity Estimation
- Reinforced Depth-Aware Deep Learning for Single Image Dehazing
- Relative Cost Based Model Selection for Sparse High-Dimensional Linear Regression Models
- Reliable and Secure Transmission for Future Networks
- Residual Attention Network for Wavelet Domain Super-Resolution
- Residual Recurrent Neural Network for Speech Enhancement
- Resilient Distributed Recovery of Large Fields
- Resilient to Byzantine Attacks Finite-Sum Optimization Over Networks
- Resource Management in the Multibeam NOMA-based Satellite Downlink
- Resting-State EEG-Based Biometrics with Signals Features Extracted by Multivariate Empirical Mode Decomposition
- Rethinking Retinal Landmark Localization as Pose Estimation: Naïve Single Stacked Network for Optic Disk and Fovea Detection
- Retinal Vessel Segmentation via a Semantics and Multi-Scale Aggregation Network
- Retrieving Vocal-Tract Resonance and anti-Resonance From High-Pitched Vowels Using a Rahmonic Subtraction Technique
- Revealing Backdoors, Post-Training, in DNN Classifiers via Novel Inference on Optimized Perturbations Inducing Group Misclassification
- Revealing Hidden Drawings in Leonardo's 'the Virgin of the Rocks' from Macro X-Ray Fluorescence Scanning Data through Element Line Localisation
- Reversal No Longer Matters: Attention-Based Arrhythmia Detection with Lead-Reversal ECG Data
- Revisit of Estimate Sequence for Accelerated Gradient Methods
- Revisiting Fast Spectral Clustering with Anchor Graph
- Rgb-D Based Multi-Modal Deep Learning for Face Identification
- Riemannian Framework for Robust Covariance Matrix Estimation in Spiked Models
- Riemannian Geometry and Cramér-rao Bound for Blind Separation of Gaussian Sources
- Risk Convergence of Centered Kernel Ridge Regression with Large Dimensional Data
- Rnn-Transducer with Stateless Prediction Network
- Robust CFAR Radar Detection Using a K-nearest Neighbors Rule
- Robust Covariance Matrix Estimation and Portfolio Allocation: The Case of Non-Homogeneous Assets
- Robust Frequency-Domain Recursive Least M-Estimate Adaptive Filter For Acoustic System Identification
- Robust Full-Fov Depth Estimation in Tele-Wide Camera System
- Robust Fundamental Frequency Estimation in Coloured Noise
- Robust Global Optimized Affine Registration Method for Microscopic Images of Biological Tissue
- Robust Hybrid Beamforming for Satellite-Terrestrial Integrated Networks
- Robust Likelihood Ratio Test Using α-Divergence
- Robust Low Rate Speech Coding Based on Cloned Networks and Wavenet
- Robust Marine Buoy Placement for Ship Detection Using Dropout K-Means
- Robust Matrix Completion via ℓP-Greedy Pursuits
- Robust Multi-Channel Speech Recognition Using Frequency Aligned Network
- Robust Music Estimation Under Array Response Uncertainty
- Robust Online Matrix Completion with Gaussian Mixture Model
- Robust Online Mirror Saddle-Point Method for Constrained Resource Allocation
- Robust Parameter Estimation of Contaminated Damped Exponentials
- Robust Phase Retrieval with Outliers
- Robust Pricing Mechanism for Resource Sustainability Under Privacy Constraint in Competitive Online Learning Multi-Agent Systems
- Robust Rank Constrained Sparse Learning: A Graph-Based Method for Clustering
- Robust Speaker Recognition Using Unsupervised Adversarial Invariance
- Robust Symbol-Level Precoding Via Autoencoder-Based Deep Learning
- Robust Tdoa Indoor Tracking Using Constrained Measurement Filtering and Grid-Based Filtering
- Robust Transmission Over Channels with Channel Uncertainty: an Algorithmic Perspective
- Robust Unsupervised Audio-Visual Speech Enhancement Using a Mixture of Variational Autoencoders
- Robust Visual Tracking with Context-Based Active Occlusion Recognition
- Robust and Computationally-Efficient Anomaly Detection Using Powers-Of-Two Networks
- Robust and steerable kronecker product differential beamforming With rectangular microphone arrays
- Robustness Assessment of Automatic Reinke's Edema Diagnosis Systems
- Robustness of Sparse Bayesian Learning in Correlated Environments
- S-DOD-CNN: Doubly Injecting Spatially-Preserved Object Information for Event Recognition
- SDTCN: Similarity Driven Transmission Computing Network for Image Dehazing
- SECL-UMons Database for Sound Event Classification and Localization
- SED-MDD: Towards Sentence Dependent End-To-End Mispronunciation Detection and Diagnosis
- SLOGD: Speaker Location Guided Deflation Approach to Speech Separation
- SNDCNN: Self-Normalizing Deep CNNs with Scaled Exponential Linear Units for Speech Recognition
- SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds
- SSGD: Sparsity-Promoting Stochastic Gradient Descent Algorithm for Unbiased Dnn Pruning
- SSTNet: Detecting Manipulated Faces Through Spatial, Steganalysis and Temporal Features
- Saliency-Based Image Contrast Enhancement with Reversible Data Hiding
- Salient Object Detection Based On Image Bit-Map
- Sampling Classes of Non-Bandlimited Signals Using Integrate-and-Fire Devices: Average Case Analysis
- Sampling Strategies for GAN Synthetic Data
- Sampling of Surfaces and Learning Functions in High Dimensions
- Scalable Detection and Tracking of Extended Objects
- Scalable Kernel Learning Via the Discriminant Information
- Scalable Learning-Based Sampling Optimization for Compressive Dynamic MRI
- Scalable Multilingual Frontend for TTS
- Scalpnet: Detection of Spatiotemporal Abnormal Intervals in Epileptic EEG Using Convolutional Neural Networks
- Scene Text Recognition with Temporal Convolutional Encoder
- Scene-Dependent Acoustic Event Detection with Scene Conditioning and Fake-Scene-Conditioned Loss
- SeCoST: : Sequential Co-Supervision for Large Scale Weakly Labeled Audio Event Detection
- Secure Face Recognition in Edge and Cloud Networks: From the Ensemble Learning Perspective
- Secure Identification for Gaussian Channels
- Secure Symbol-Level Miso Precoding
- Selection-Channel-Aware Reverse JPEG Compatibility for Highly Reliable Steganalysis of JPEG Images
- Selective Attention Encoders by Syntactic Graph Convolutional Networks for Document Summarization
- Selective Convolutional Network: An Efficient Object Detector with Ignoring Background
- Self-Adaptive Feature Fool
- Self-Attention and Retrieval Enhanced Neural Networks for Essay Generation
- Self-Attentive Sentimental Sentence Embedding for Sentiment Analysis
- Self-Driven Graph Volterra Models for Higher-Order Link Prediction
- Self-Paced Probabilistic Principal Component Analysis For Data With Outliers
- Self-Supervised Adversarial Training
- Self-Supervised Deep Learning for Fisheye Image Rectification
- Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech Enhancement
- Self-Supervised Learning for Audio-Visual Speaker Diarization
- Self-Supervised Learning for ECG-Based Emotion Recognition
- Self-Training for End-to-End Speech Recognition
- Semantic Augmentation Hashing for Zero-Shot Image Retrieval
- Semanticgan: Generative Adversarial Networks For Semantic Image To Photo-Realistic Image Translation
- Semi-Implicit Stochastic Recurrent Neural Networks
- Semi-Regular Geometric Kernel Encoding & Reconstruction for Video Compression
- Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis
- Semi-Supervised Learning for Text Classification by Layer Partitioning
- Semi-Supervised Learning of Processes Over Multi-Relational Graphs
- Semi-Supervised Optimal Transport Methods for Detecting Anomalies
- Semi-Supervised Sentence Classification Based on User Polarity in the Social Scenarios
- Semi-Supervised Speaker Adaptation for End-to-End Speech Synthesis with Pretrained Models
- Sensor Selection for Model-Free Source Localization: where Less is More
- Separable Optimization for Joint Blind Deconvolution and Demixing
- Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech Recognition
- Sequence-To-Subsequence Learning With Conditional Gan For Power Disaggregation
- Sequence-to-Sequence Automatic Speech Recognition with Word Embedding Regularization and Fused Decoding
- Sequence-to-Sequence Labanotation Generation Based on Motion Capture Data
- Sequence-to-Sequence Singing Synthesis Using the Feed-Forward Transformer
- Sequential Deep Unrolling With Flow Priors For Robust Video Deraining
- Sequential IoT Data Augmentation Using Generative Adversarial Networks
- Sequential Joint Detection and Estimation with an Application to Joint Symbol Decoding and Noise Power Estimation
- Sequential Methods for Detecting a Change in the Distribution of an Episodic Process
- Sequential Semi-Orthogonal Multi-Level NMF with Negative Residual Reduction for Network Embedding
- Sequential Vessel Trajectory Identification Using Truncated Viterbi Algorithm
- Shadow Removal of Text Document Images by Estimating Local and Global Background Colors
- Shape From Bandwidth: Central Projection Case
- Short and Squeezed: Accelerating the Computation of Antisparse Representations with Safe Squeezing
- Sight to Sound: An End-to-End Approach for Visual Piano Transcription
- Signal Clustering With Class-Independent Segmentation
- Signal Sensing and Reconstruction Paradigms for a Novel Multi-Source Static Computed Tomography System
- Signal-Aware Broadband DOA Estimation Using Attention Mechanisms
- Similarity Learning For Cover Song Identification Using Cross-Similarity Matrices of Multi-Level Deep Sequences
- Simple Caching Schemes for Non-homogeneous MISO Cache-Aided Communication via Convexity
- Simplified Dynamic SC-Flip Polar Decoding
- Simultaneous Separation and Transcription of Mixtures with Multiple Polyphonic and Percussive Instruments
- Singing Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders
- Single Frequency Filter Bank Based Long-Term Average Spectra for Hypernasality Detection and Assessment in Cleft Lip and Palate Speech
- Single-Channel Speech Separation Integrating Pitch Information Based on a Multi Task Learning Framework
- Single-Shot Real-Time Multiple-Path Time-of-Flight Depth Imaging for Multi-Aperture and Macro-Pixel Sensors
- Sketchppnet: A Joint Pixel and Point Convolutional Neural Network For Low Resolution Sketch Image Recognition
- SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation
- Slicenet: Slice-Wise 3D Shapes Reconstruction from Single Image
- Slow-Time MIMO-FMCW Automotive Radar Detection with Imperfect Waveform Separation
- Small Energy Masking for Improved Neural Network Training for End-To-End Speech Recognition
- Small-Footprint Keyword Spotting on Raw Audio Data with Sinc-Convolutions
- Smoothing Graph Signals via Random Spanning Forests
- Snorer Diarisation Based On Deep Neural Network Embeddings
- Social Data Assisted Multi-Modal Video Analysis For Saliency Detection
- Social Learning with Partial Information Sharing
- Soft-Output Finite Alphabet Equalization for mmWave Massive MIMO
- Solving Missing-Annotation Object Detection with Background Recalibration Loss
- Solving Non-Convex Non-Differentiable Min-Max Games Using Proximal Gradient Method
- Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks
- Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene Labels
- Sound Event Detection in Synthetic Domestic Environments
- Sound Event Localization Based on Sound Intensity Vector Refined by Dnn-Based Denoising and Source Separation
- Sound Texture Synthesis Using RI Spectrograms
- Source Coding of Audio Signals with a Generative Model
- Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech Recognition
- Source Enumeration via Toeplitz Matrix Completion
- Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis
- Space Filling Curves for MRI Sampling
- Sparse Beamspace Equalization for Massive MU-MIMO MMWave Systems
- Sparse Branch and Bound for Exact Optimization of L0-Norm Penalized Least Squares
- Sparse CSP Algorithm via Joint Spatio-Temporal Filtering
- Sparse Convolutional Beamforming for Wireless Ultrasound
- Sparse Directed Graph Learning for Head Movement Prediction in 360 Video Streaming
- Sparse Low-redundancy Linear Array with Uniform Sum Co-array
- Sparse Modeling on Distributed Encryption Data
- Sparse Recovery with Non-Linear Fourier Features
- Spatial Active Noise Control Based on Kernel Interpolation with Directional Weighting
- Spatial Attention for Far-Field Speech Recognition with Deep Beamforming Neural Networks
- Spatial Attentional Bilinear 3D Convolutional Network for Video-Based Autism Spectrum Disorder Detection
- Spatial Gating Strategies for Graph Recurrent Neural Networks
- Spatial and Temporal Smoothing for Covariance Estimation in Super-Resolution Angle Estimation in Automotive Radars
- Spatial-Temporal Feature Aggregation Network For Video Object Detection
- Spatially Adaptive Intra Mode Pre-Selection for ERP 360 Video Coding
- Spatially Guided Independent Vector Analysis
- Spatio-Temporal and Geometry Constrained Network for Automobile Visual Odometry
- Speaker Adaptation of a Multilingual Acoustic Model for Cross-Language Synthesis
- Speaker Augmentation for Low Resource Speech Recognition
- Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network
- Speaker Diarization with Region Proposal Network
- Speaker Diarization with Session-Level Speaker Embedding Refinement Using Graph Neural Networks
- Speaker Embeddings Incorporating Acoustic Conditions for Diarization
- Speaker Independence of Neural Vocoders and Their Effect on Parametric Resynthesis Speech Enhancement
- Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding Extraction
- Speaker-Aware Training of Attention-Based End-to-End Speech Recognition Using Neural Speaker Embeddings
- Speaker-Invariant Affective Representation Learning via Adversarial Training
- Speakerfilter: Deep Learning-Based Target Speaker Extraction Using Anchor Speech
- Specaugment on Large Scale Datasets
- Spectrogram Analysis Via Self-Attention for Realizing Cross-Model Visual-Audio Generation
- Spectrograms Fusion with Minimum Difference Masks Estimation for Monaural Speech Dereverberation
- Spectrum Allocation in Wireless Networks for Crowd Labelling
- Speech Breathing Estimation Using Deep Learning Methods
- Speech Emotion Recognition with Dual-Sequence LSTM Architecture
- Speech Emotion Recognition with Local-Global Aware Deep Representation Learning
- Speech Enhancement Using Self-Adaptation and Multi-Head Self-Attention
- Speech Intelligibility Enhancement by Equalization for in-Car Applications
- Speech Recognition Model Compression
- Speech Sentiment Analysis via Pre-Trained Features from End-to-End ASR Models
- Speech Synthesis Using EEG
- Speech-Based Parameter Estimation of an Asymmetric Vocal Fold Oscillation Model and its Application in Discriminating Vocal Fold Pathologies
- Speech-Driven Facial Animation Using Polynomial Fusion of Features
- Speech-To-Singing Conversion in an Encoder-Decoder Framework
- Spherical Large Intelligent Surfaces
- Spherical Video Coding with Geometry and Region Adaptive Transform Domain Temporal Prediction
- Spiking Neural Networks Trained With Backpropagation for Low Power Neuromorphic Implementation of Voice Activity Detection
- Spoken Document Retrieval Leveraging Bert-Based Modeling and Query Reformulation
- Spoken Language Acquisition Based on Reinforcement Learning and Word Unit Segmentation
- Srzoo: An Integrated Repository For Super-Resolution Using Deep Learning
- Stability of Graph Neural Networks to Relative Perturbations
- Stabilizing Multi-Agent Deep Reinforcement Learning by Implicitly Estimating Other Agents' Behaviors
- Stable Training of Dnn for Speech Enhancement Based on Perceptually-Motivated Black-Box Cost Function
- Stacked Pooling for Boosting Scale Invariance of Crowd Counting
- Staged Training Strategy and Multi-Activation for Audio Tagging with Noisy and Sparse Multi-Label Data
- Stargan for Emotional Speech Conversion: Validated by Data Augmentation of End-To-End Emotion Recognition
- State-Based Transcription of Components of Carnatic Music
- State-Space Gaussian Process for Drift Estimation in Stochastic Differential Equations
- Static Visual Spatial Priors for DoA Estimation
- Statistical Signal Processing Approach for Rain Estimation Based on Measurements from Network Management Systems
- Statistics Pooling Time Delay Neural Network Based on X-Vector for Speaker Verification
- Steepening Squared Error Function Facilitates Online Adaptation of Gaussian Scales
- Steganography and its Detection in JPEG Images Obtained with the "TRUNC" Quantizer
- Stochastic Admm For Byzantine-Robust Distributed Learning
- Stochastic Geometry Planning of Electric Vehicles Charging Stations
- Stochastic Graph Neural Networks
- Stochastic Ml Estimation for Hyperspectral Unmixing Under Endmember Variability and Nonlinear Models
- Stochastic Multi-Scale Aggregation Network for Crowd Counting
- Stock Movement Prediction That Integrates Heterogeneous Data Sources Using Dilated Causal Convolution Networks with Attention
- Storing Digital Data Into DNA: A Comparative Study Of Quaternary Code Construction
- Strategic Attention Learning for Modality Translation
- Streaming Automatic Speech Recognition with the Transformer Model
- Structural Sparsification for Far-Field Speaker Recognition with Intel® Gna
- Structured Citation Trend Prediction Using Graph Neural Networks
- Structured Sparse Attention for end-to-end Automatic Speech Recognition
- Study of Closed Phase Resonance Bandwidths for Oral and Nasal Tracts Using Zero Time Windowing
- Study of Formant Modification for Children ASR
- Sub-Dip: Optimization On A Subspace With Deep Image Prior Regularization And Application To Superresolution
- Subject Transfer Framework Based on Source Selection and Semi-Supervised Style Transfer Mapping for Semg Pattern Recognition
- Subjective Quality Estimation Using PESQ For Hands-Free Terminals
- Submodular Rank Aggregation on Score-Based Permutations for Distributed Automatic Speech Recognition
- Subspace-Based Speech Correlation Vector Estimation for Single-Microphone Multi-Frame MVDR Filtering
- Super-Resolution of 3D Color Point Clouds Via Fast Graph Total Variation
- Super-Resolution with Noisy Measurements: Reconciling Upper and Lower Bounds
- Superpixel Segmentation Via Convolutional Neural Networks with Regularized Information Maximization
- Supervised Canonical Correlation Analysis of Data on Symmetric Positive Definite Manifolds by Riemannian Dimensionality Reduction
- Supervised Deep Hashing for Efficient Audio Event Retrieval
- Supervised Encoding for Discrete Representation Learning
- Supervised Graph Representation Learning for Modeling the Relationship between Structural and Functional Brain Connectivity
- Supervised Online Diarization with Sample Mean Loss for Multi-Domain Data
- Synchronous Transformers for end-to-end Speech Recognition
- Synthesizing Engaging Music Using Dynamic Models of Statistical Surprisal
- Synthetic Crowd and Pedestrian Generator for Deep Learning Problems
- Synthetic Data Generation Through Statistical Explosion: Improving Classification Accuracy of Coronary Artery Disease Using PPG
- Synthetic Speech References for Automatic Pathological Speech Intelligibility Assessment
- T-GSA: Transformer with Gaussian-Weighted Self-Attention for Speech Enhancement
- TDMF: Task-Driven Multilevel Framework for End-to-End Speaker Verification
- TOSO: Student's-T Distribution Aided One-Stage Orientation Target Detection in Remote Sensing Images
- Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization System
- Talker-Independent Speaker Separation in Reverberant Conditions
- Target Parameter Estimation via One-Bit PMCW Radar
- Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
- Teacher-Student Training For Robust Tacotron-Based TTS
- Teaching Signals and Systems - A First Course in Signal Processing
- Temporal Coding in Spiking Neural Networks with Alpha Synaptic Function
- Tensor Decomposition-based Beamspace Esprit Algorithm for Multidimensional Harmonic Retrieval
- Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train Network
- Tensorflow Audio Models in Essentia
- Texception: A Character/Word-Level Deep Learning Model for Phishing URL Detection
- Text Adaptation for Speaker Verification with Speaker-Text Factorized Embeddings
- Text-Independent Speaker Verification with Adversarial Learning on Short Utterances
- Text-To-Image Synthesis Method Evaluation Based On Visual Patterns
- The Compressed Nested Array for Underdetermined DOA Estimation by Fourth-order Difference Coarrays
- The Discrete Stockwell Transforms for Infinite-Length Signals and Their Real-Time Implementations
- The Effect of Data Augmentation on Classification of Atrial Fibrillation in Short Single-Lead ECG Signals Using Deep Neural Networks
- The Effect of Power Allocation on Visible Light Communication Using Commercial Phosphor-Converted Led Lamp for Indirect Illumination
- The Empirical Duality Gap of Constrained Statistical Learning
- The Fifthnet Chroma Extractor
- The Fractional Quaternion Fourier Number Transform
- The Graphon Fourier Transform
- The Matched Reassigned Cross-Spectrogram for Phase Estimation
- The Open Brands Dataset: Unified Brand Detection and Recognition at Scale
- The Picasso Algorithm for Bayesian Localization Via Paired Comparisons in a Union of Subspaces Model
- The Processing of Mandarin Chinese Tonal Alternations in Contexts: An Eye-Tracking Study
- The Role of Annotation Fusion Methods in the Study of Human-Reported Emotion Experience During Music Listening
- The Rwth Asr System for Ted-Lium Release 2: Improving Hybrid Hmm With Specaugment
- The Sound of My Voice: Speaker Representation Loss for Target Voice Separation
- Theoretical Analysis of Multi-Carrier Agile Phased Array Radar
- Theoretical Performance Bound of Uplink Channel Estimation Accuracy in Massive MIMO
- This Dataset Does Not Exist: Training Models from Generated Images
- Threshold-Adjusted ORB Strategies with Genetic Algorithm and Protective Closing Strategy on Taiwan Futures Market
- Time Difference of Arrival Estimation from Frequency-Sliding Generalized Cross-Correlations Using Convolutional Neural Networks
- Time Domain Velocity Vector for Retracing the Multipath Propagation
- Time Reversal Based Robust Gesture Recognition Using Wifi
- Time-Domain Audio Source Separation Based on Wave-U-Net Combined with Discrete Wavelet Transform
- Time-Domain Neural Network Approach for Speech Bandwidth Extension
- Time-Frequency Analysis of Unimodal Sensory Processing In Autism Spectrum Disorder
- Time-Frequency Feature Decomposition Based on Sound Duration for Acoustic Scene Classification
- Time-Frequency Loss for CNN Based Speech Super-Resolution
- Time-Predictable Software-Defined Architecture with Sdf-Based Compiler Flow for 5g Baseband Processing
- Time-Scale Synthesis for Locally Stationary Signals
- Toward Better Speaker Embeddings: Automated Collection of Speech Samples From Unknown Distinct Speakers
- Towards Blind Quality Assessment of Concert Audio Recordings Using Deep Neural Networks
- Towards Data-Efficient Modeling for Wake Word Spotting
- Towards Decoding Selective Attention from Single-Trial EEG Data in Cochlear Implant users based on Deep Neural Networks
- Towards Fast and Accurate Streaming End-To-End ASR
- Towards High-Performance Object Detection: Task-Specific Design Considering Classification and Localization Separation
- Towards Linking the Lakh and IMSLP Datasets
- Towards Multilingual Sign Language Recognition
- Towards Pose-Invariant Lip-Reading
- Towards Real-Time Single-Channel Singing-Voice Separation with Pruned Multi-Scaled Densenets
- Towards Real-Time, Multi-View Video Stereopsis
- Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
- Towards a New Understanding of the Training of Neural Networks with Mislabeled Training Data
- Towards an Efficient and General Framework of Robust Training for Graph Neural Networks
- Towards an Intelligent Microscope: Adaptively Learned Illumination for Optimal Sample Classification
- Trace Norm Generative Adversarial Networks for Sensor Generation and Feature Extraction
- Tracing Network Evolution Using The Parafac2 Model
- Track-Before-Detect for Sub-Nyquist Radar
- Tracking to Improve Detection Quality in Lidar For Autonomous Driving
- Training ASR Models By Generation of Contextual Information
- Training Code-Switching Language Model with Monolingual Data
- Training Deep Spiking Neural Networks for Energy-Efficient Neuromorphic Computing
- Training Keyword Spotters with Limited and Synthesized Speech Data
- Training LSTM for Unsupervised Anomaly Detection Without A Priori Knowledge
- Training Spoken Language Understanding Systems with Non-Parallel Speech and Text
- Transfer Learning from Youtube Soundtracks to Tag Arctic Ecoacoustic Recordings
- Transferable Policies for Large Scale Wireless Networks with Graph Neural Networks
- Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation
- Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
- Transformer VAE: A Hierarchical Model for Structure-Aware and Interpretable Music Representation Learning
- Transformer-Based Acoustic Modeling for Hybrid Speech Recognition
- Transformer-Based Online CTC/Attention End-To-End Speech Recognition Architecture
- Transformer-Based Text-to-Speech with Weighted Forced Attention
- Transforming Seismocardiograms Into Electrocardiograms by Applying Convolutional Autoencoders
- Translation of a Higher Order Ambisonics Sound Scene Based on Parametric Decomposition
- Transmit Beamforming Design with Received-Interference Power Constraints: The Zero-Forcing Relaxation
- Transmit Beampattern Shaping via Waveform Design in Cognitive Mimo Radar
- Trapezoidal Segment Sequencing: A Novel Approach for Fusion of Human-Produced Continuous Annotations
- Tree of Shapes Cut for Material Segmentation Guided by a Design
- Triggerless Random Interleaved Sampling
- Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms
- Triplet Loss Feature Aggregation for Scalable Hash
- Truth-to-Estimate Ratio Mask: A Post-Processing Method for Speech Enhancement Direct at Low Signal-to-Noise Ratios
- Ts-Fen: Probing Feature Selection Strategy for Face Anti-Spoofing
- Two-Element Biomimetic Antenna Array Design and Performance
- Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition
- Two-Step Sound Source Separation: Training On Learned Latent Targets
- Two-dimensional DOA Estimation for Coprime Planar Array: A Coarray Tensor-based Solution
- UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation
- Uncertainties in Short Commercial Microwave Links Fading Due to Rain
- Uncertainty Quantification for Remaining Useful Lifetime Prediction with Multi-Channel Sensory Data
- Underwater Tracking Based on the Sum-Product Algorithm Enhanced by a Neural Network Detections Classifier
- Unified Signal Compression Using Generative Adversarial Networks
- Universal Phone Recognition with a Multilingual Allophone System
- Unseen Face Presentation Attack Detection with Hypersphere Loss
- Unsupervised Auto-Encoding Multiple-Object Tracker for Constraint-Consistent Combinatorial Problem
- Unsupervised Change Detection for Multimodal Remote Sensing Images via Coupled Dictionary Learning and Sparse Coding
- Unsupervised Content-Preserved Adaptation Network for Classification of Pulmonary Textures from Different CT Scanners
- Unsupervised Domain Adaptation for Semantic Segmentation with Symmetric Adaptation Consistency
- Unsupervised Feature Enhancement for Speaker Verification
- Unsupervised Image-to-Image Translation Via Fair Representation of Gender Bias
- Unsupervised Key Hand Shape Discovery of Sign Language Videos with Correspondence Sparse Autoencoders
- Unsupervised Multiple Source Localization Using Relative Harmonic Coefficients
- Unsupervised Neural Mask Estimator for Generalized Eigen-Value Beamforming Based Asr
- Unsupervised Person Re-Identification Using Multi-Branch Feature Compensation Network and Link-Based Cluster Dissimilarity Metric
- Unsupervised Pre-Training of Bidirectional Speech Encoders via Masked Reconstruction
- Unsupervised Pretraining Transfers Well Across Languages
- Unsupervised Speaker Adaptation Using Attention-Based Speaker Memory for End-to-End ASR
- Unsupervised Style and Content Separation by Minimizing Mutual Information for Speech Synthesis
- Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function
- Unsupervised Variational Bayesian Kalman Filtering For Large-Dimensional Gaussian Systems
- Upgrade Methods for Stratified Sensor Network Self-Calibration
- Upgrading CRFS to JRFS and its Benefits to Sequence Modeling and Labeling
- Upscaling Vector Approximate Message Passing
- Urtis: a Small 3d Imaging Sonar Sensor for Robotic Applications
- Using Automatic Speech Recognition and Speech Synthesis to Improve the Intelligibility of Cochlear Implant users in Reverberant Listening Environments
- Using Intelligent Reflecting Surfaces for Rank Improvement in MIMO Communications
- Using Panoramic Videos for Multi-Person Localization and Tracking In A 3D Panoramic Coordinate
- Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker Adaptation
- Using Separate Losses for Speech and Noise in Mask-Based Speech Enhancement
- Using Speech Synthesis to Train End-To-End Spoken Language Understanding Models
- Using Vaes and Normalizing Flows for One-Shot Text-To-Speech Synthesis of Expressive Speech
- Using X-Vectors to Automatically Detect Parkinson's Disease from Speech
- Utterance-Level Sequential Modeling for Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit
- VAMP with Vector-Valued Diagonalization
- Vapar Synth - A Variational Parametric Model for Audio Synthesis
- Variable Bitrate Image Compression with Quality Scaling Factors
- Variable Metric Proximal Gradient Method with Diagonal Barzilai-Borwein Stepsize
- Variable Projection for Multiple Frequency Estimation
- Variational Student: Learning Compact and Sparser Networks In Knowledge Distillation Framework
- Versatile Video Coding and Super-Resolution for Efficient Delivery of 8k Video with 4k Backward-Compatibility
- Vggsound: A Large-Scale Audio-Visual Dataset
- ViMo: Vital Sign Monitoring Using Commodity Millimeter Wave Radio
- Video Deblurring Via 3d CNN and Fourier Accumulation Learning
- Video Frame Interpolation Via Exceptional Motion-Aware Synthesis
- Video Frame Interpolation Via Residue Refinement
- Video Question Generation via Semantic Rich Cross-Modal Self-Attention Networks Learning
- View-Angle Invariant Object Monitoring Without Image Registration
- Visually Guided Self Supervised Learning of Speech Representations
- Vocal Tract Articulatory Contour Detection in Real-Time Magnetic Resonance Images Using Spatio-Temporal Context
- Voice Conversion with Transformer Network
- Voice based classification of patients with Amyotrophic Lateral Sclerosis, Parkinson's Disease and Healthy Controls with CNN-LSTM using transfer learning
- Voiceai Systems to NIST Sre19 Evaluation: Robust Speaker Recognition on Conversational Telephone Speech
- Volume Reconstruction for Light Field Microscopy
- WHAMR!: Noisy and Reverberant Single-Channel Speech Separation
- WaveFFJORD: FFJORD-Based Vocoder for Statistical Parametric Speech Synthesis
- Wawenets: A No-Reference Convolutional Waveform-Based Approach to Estimating Narrowband and Wideband Speech Quality
- Weakly Labelled Audio Tagging Via Convolutional Networks with Spatial and Channel-Wise Attention
- Weakly Supervised Crowd-Wise Attention For Robust Crowd Counting
- Weakly Supervised Segmentation Guided Hand Pose Estimation During Interaction with Unknown Objects
- Weakly Supervised Semantic Segmentation For Remote Sensing Hyperspectral Imaging
- Weakly-Supervised Sound Event Detection with Self-Attention
- Weight Sharing and Deep Learning for Spectral Data
- Weighted Gradient Coding with Leverage Score Sampling
- Weighted Krylov-Levenberg-Marquardt Method for Canonical Polyadic Tensor Decomposition
- Weighted Null Vector Initialization and its Application to Phase Retrieval
- Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement
- What Does a Network Layer Hear? Analyzing Hidden Representations of End-to-End ASR Through Speech Synthesis
- What Makes the Sound?: A Dual-Modality Interacting Network for Audio-Visual Event Localization
- What did your adversary believeƒ Optimal Filtering and Smoothing in Counter-Adversarial Autonomous Systems
- What is best for spoken language understanding: small but task-dependant embeddings or huge but out-of-domain embeddings?
- Whosecough: In-the-Wild Cougher Verification Using Multitask Learning
- Wideband Channel Tracking for Millimeter Wave Massive Mimo Systems with Hybrid Beamforming Reception
- Wideband Direction of Arrival Estimation with Sparse Linear Arrays
- Wind: Wasserstein Inception Distance For Evaluating Generative Adversarial Network Performance
- Wirtinger Flow Algorithms for Phase Retrieval from Binary Measurements
- Witchcraft: Efficient PGD Attacks with Random Step Size
- Within-Sample Variability-Invariant Loss for Robust Speaker Recognition Under Noisy Environments
- X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker Recognition
- XceptionTime: Independent Time-Window Xceptiontime Architecture for Hand Gesture Classification
- Xpsnr: A Low-Complexity Extension of The Perceptually Weighted Peak Signal-To-Noise Ratio For High-Resolution Video Quality Assessment
- Y-Net: Multi-Scale Feature Aggregation Network With Wavelet Structure Similarity Loss Function For Single Image Dehazing
- Zero-Crossing Precoding with Maximum Distance to the Decision Threshold for Channels with 1-Bit Quantization and Oversampling
- Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings
- dMazeRunner: Optimizing Convolutions on Dataflow Accelerators
- β-NMF and Sparsity Promoting Regularizations for Complex Mixture Unmixing. Application to 2D HSQC NMR
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.