ICASSP 2020 Accepted Papers
The full list of 1,849 papers accepted at ICASSP 2020 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- 1.5GBIT/S 4.9W Hyperspectral Image Encoders on a Low-Power Parallel Heterogeneous Processing Platform
- 2D-to-2D Mask Estimation for Speech Enhancement Based on Fully Convolutional Neural Network
- 3-D Acoustic Modeling for Far-Field Multi-Channel Speech Recognition
- 3D Unknown View Tomography Via Rotation Invariants
- 3d Deformation Signature for Dynamic Face Recognition
- A BI-Model Approach for Handling Unknown Slot Values in Dialogue State Tracking
- A Bidirectional Context Propagation Network for Urine Sediment Particle Detection in Microscopic Images
- A Bin Encoding Training of a Spiking Neural Network Based Voice Activity Detection
- A Comparative Study of Estimating Articulatory Movements from Phoneme Sequences and Acoustic Features
- A Comparative Study of Western and Chinese Classical Music Based on Soundscape Models
- A Comparison of Pooling Methods on LSTM Models for Rare Acoustic Event Classification
- A Complexity Efficient DMT-Optimal Tree Pruning Based Sphere Decoding
- A Composite DNN Architecture for Speech Enhancement
- A Comprehensive Framework for 2D-JND Extension to 360-DEG Images
- A Comprehensive Study of Residual CNNS for Acoustic Modeling in ASR
- A Computationally Light Algorithm for Bayesian Speech Enhancement with SNR Marginalization
- A Connected Auto-Encoders Based Approach for Image Separation with Side Information: With Applications to Art Investigation
- A Constrained Maximum Likelihood Estimator of Speech and Noise Spectra with Application to Multi-Microphone Noise Reduction
- A Cross-Task Transfer Learning Approach to Adapting Deep Speech Enhancement Models to Unseen Background Noise Using Paired Senone Classifiers
- A DSP Acceleration Framework For Software-Defined Radios On X86 64
- A Data Efficient End-to-End Spoken Language Understanding Architecture
- A Dataset for Measuring Reading Levels In India At Scale
- A Deep Gradient Boosting Network for Optic Disc and Cup Segmentation
- A Deep Learning Approach to Object Affordance Segmentation
- A Deep Learning Architecture for Epileptic Seizure Classification Based on Object and Action Recognition
- A Deep Multimodal Approach for Map Image Classification
- A Deep Neural Network-Driven Feature Learning Method for Polyphonic Acoustic Event Detection from Real-Life Recordings
- A Dense U-Net with Cross-Layer Intersection for Detection and Localization of Image Forgery
- A Dialogical Emotion Decoder for Speech Motion Recognition in Spoken Dialog
- A Differential Approach for Rain Field Tomographic Reconstruction Using Microwave Signals from Leo Satellites
- A Discriminative Condition-Aware Backend for Speaker Verification
- A Dual-Staged Context Aggregation Method towards Efficient End-to-End Speech Enhancement
- A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker Tracking
- A Fast Non-Contact Vital Signs Detection Method Based on Regional Hidden Markov Model in A 77ghz Lfmcw Radar System
- A Fast Proximal Point Algorithm for Generalized Graph Laplacian Learning
- A Fast Reduced-Rank Sound Zone Control Algorithm Using The Conjugate Gradient Method
- A Fast Sparse Covariance-Based Fitting Method for DOA Estimation via Non-Negative Least Squares
- A Fast and Accurate Frequent Directions Algorithm for Low Rank Approximation via Block Krylov Iteration
- A Fast and Accurate Super-Resolution Network Using Progressive Residual Learning
- A Fifo Based Accelerator for Convolutional Neural Networks
- A Forward-Backward Algorithm for Reweighted Procedures: Application to Radio-Astronomical Imaging
- A Framework for Parameters Estimation of Image Operator Chain
- A Framework for the Robust Evaluation of Sound Event Detection
- A Frequency-Domain BSS Method Based on ℓ1 Norm, Unitary Constraint, and Cayley Transform
- A Gated Hypernet Decoder for Polar Codes
- A General Difficulty Control Algorithm for Proof-of-Work Based Blockchains
- A General Test for the Linear Structure of Covariance Matrices of Gaussian Populations
- A Generalization of Principal Component Analysis
- A Generalized Framework for Domain Adaptation of PLDA in Speaker Recognition
- A Geometric Approach for Unsupervised Similarity Learning
- A Graph Network Model for Distributed Learning with Limited Bandwidth Links and Privacy Constraints
- A Greedy Sparse Approximation Algorithm Based On L1-Norm Selection Rules
- A Hardware Architecture For Reconfigurable Intelligent Surfaces with Minimal Active Elements for Explicit Channel Estimation
- A Hierarchical Model for Dialog Act Recognition Considering Acoustic and Lexical Context Information
- A Hierarchical Tracker for Multi-Domain Dialogue State Tracking
- A Hybrid Approach for Thermographic Imaging With Deep Learning
- A Hybrid Model for Bipolar Disorder Classification from Visual Information
- A Hybrid Structural Sparse Error Model for Image Deblocking
- A Hybrid Text Normalization System Using Multi-Head Self-Attention For Mandarin
- A Large-Scale Deep Architecture for Personalized Grocery Basket Recommendations
- A Learning Approach to Cooperative Communication System Design
- A Lightweight Multi-Label Segmentation Network for Mobile Iris Biometrics
- A Linear Time Partitioning Algorithm for Frequency Weighted Impurity Functions
- A Low-Complexity Map Detector for Distributed Networks
- A Low-Dimensionality Method for Data-Driven Graph Learning
- A Low-Latency Successive Cancellation Hybrid Decoder for Convolutional Polar Codes
- A Low-Resolution ADC Proof-of-Concept Development for a Fully-Digital Millimeter-wave Joint Communication-Radar
- A Maximum Likelihood Approach to Multi-Objective Learning Using Generalized Gaussian Distributions for Dnn-Based Speech Enhancement
- A Memory Augmented Architecture for Continuous Speaker Identification in Meetings
- A Method for Millimeter-Wave Imaging of Concealed Objects Via De-Aliasing
- A Minimal Personalization of Dynamic Binaural Synthesis with Mixed Structural Modeling and Scattering Delay Networks
- A Model of Double Descent for High-Dimensional Logistic Regression
- A Model-Based Deep Network for MRI Reconstruction Using Approximate Message Passing Algorithm
- A Model-Free Approach to Distributed Transmit Beamforming
- A Moment-Based Approach for Guaranteed Tensor Decomposition
- A Monte Carlo Search-Based Triplet Sampling Method for Learning Disentangled Representation of Impulsive Noise on Steering Gear
- A Multi-Dilation and Multi-Resolution Fully Convolutional Network for Singing Melody Extraction
- A Multi-Phase Gammatone Filterbank for Speech Separation Via Tasnet
- A Multi-Scaled Receptive Field Learning Approach for Medical Image Segmentation
- A Multichannel Kalman-Based Wiener Filter Approach for Speaker Interference Reduction in Meetings
- A Multitaper Reassigned Spectrogram for Increased Time-Frequency Localization Precision
- A Neural Document Language Modeling Framework for Spoken Document Retrieval
- A Neural Network Based on First Principles
- A Neural Network for Monaural Intrusive Speech Intelligibility Prediction
- A Neural Network-Based Spike Sorting Feature Map That Resolves Spike Overlap in the Feature Space
- A New Application of Ultrasound Signal Processing for Archaeological Ceramic Classification
- A New Multihypothesis Prediction Scheme for Compressed Video Sensing Reconstruction
- A New Perspective for Flexible Feature Gathering in Scene Text Recognition Via Character Anchor Pooling
- A New Sampling Scheme for Distributed Blind Spectrum Sensing Using Energy Detectors
- A New Variational Method for Deep Supervised Semantic Image Hashing
- A Noninvasive Method to Detect Diabetes Mellitus and Lung Cancer Using the Stacked Sparse autoencoder
- A Novel Approach for Intelligibility Assessment in Dysarthric Subjects
- A Novel Method for Obtaining Diffuse Field Measurements for Microphone Calibration
- A Novel Moving Sparse Array Geometry with Increased Degrees of Freedom
- A Novel Pruning Approach for Bagging Ensemble Regression Based on Sparse Representation
- A Novel Rank Selection Scheme in Tensor Ring Decomposition Based on Reinforcement Learning for Deep Neural Networks
- A Novel Saliency-Driven Oil Tank Detection Method for Synthetic Aperture Radar Images
- A Novel Two-Pathway Encoder-Decoder Network for 3D Face Reconstruction
- A Partial Relaxation DOA Estimator Based on Orthogonal Matching Pursuit
- A Particle Gibbs Sampling Approach to Topology Inference in Gene Regulatory Networks
- A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech Recognition
- A Priori Estimates of the Generalization Error for Autoencoders
- A Probabilistic Scheme for Representation Learning with Radial Transform Images
- A Prototypical Triplet Loss for Cover Detection
- A Proximal Dual Consensus Method for Linearly Coupled Multi-Agent Non-Convex Optimization
- A Random Gossip BMUF Process for Neural Language Modeling
- A Real Time Implementation of a Bayer Domain Image Deblurring Core for Optical Blur Compensation
- A Real-Time Deep Network for Crowd Counting
- A Recurrent Variational Autoencoder for Speech Enhancement
- A Recursive Bayesian Solution for the Excess Over Threshold Distribution with Stochastic Parameters
- A Recursive Edge Detector For Color Filter Array Image
- A Regularized Attention Mechanism for Graph Attention Networks
- A Return to Dereverberation in the Frequency Domain Using a Joint Learning Approach
- A Robust Audio-Visual Speech Enhancement Model
- A Robust Speaker Clustering Method Based on Discrete Tied Variational Autoencoder
- A Segmentation Based Robust Deep Learning Framework for Multimodal Retinal Image Registration
- A Self-Attentive Emotion Recognition Network
- A Semi-Supervised Approach For Identifying Abnormal Heart Sounds Using Variational Autoencoder
- A Semi-Supervised Rank Tracking Algorithm For On-Line Unmixing Of Hyperspectral Images
- A Sequence Matching Network for Polyphonic Sound Event Localization and Detection
- A Siamese Content-Attentive Graph Convolutional Network for Personality Recognition Using Physiology
- A Simple But Effective Bert Model for Dialog State Tracking on Resource-Limited Systems
- A Simple Derivation of AMP and its State Evolution via First-Order Cancellation
- A Simple and Efficient Iterative Method for Toa Localization
- A Single-RF Architecture for Multiuser Massive MIMO Via Reflecting Surfaces
- A Sparse Linear Array Approach in Automotive Radars Using Matrix Completion
- A Stacked-Autoencoder Based End-to-End Learning Framework for Decode-and-Forward Relay Networks
- A Streaming On-Device End-To-End Model Surpassing Server-Side Conventional Model Quality and Latency
- A Study of Child Speech Extraction Using Joint Speech Enhancement and Separation in Realistic Conditions
- A Study of Generalization of Stochastic Mirror Descent Algorithms on Overparameterized Nonlinear Models
- A Study on the Transferability of Adversarial Attacks in Sound Event Classification
- A Switching Transmission Game with Latency as the User's Communication Utility
- A Theoretical Basis for Practitioners Heuristic 1/N and Long-Only Quintile Portfolio
- A Time-Based Sampling Framework for Finite-Rate-of-Innovation Signals
- A Time-Frequency Network with Channel Attention and Non-Local Modules for Artificial Bandwidth Extension
- A Unified Sequence-to-Sequence Front-End Model for Mandarin Text-to-Speech Synthesis
- A Variational Bayesian Approach for Multichannel Through-Wall Radar Imaging with Low-Rank and Sparse Priors
- A Visual-Pilot Deep Fusion for Target Speech Separation in Multitalker Noisy Environment
- A Whiteness Test Based on the Spectral Measure of Large Non-Hermitian Random Matrices
- A WiFi-Based Passive Fall Detection System
- A Zeroth-Order Learning Algorithm for Ergodic Optimization of Wireless Systems with no Models and no Gradients
- A multi-view approach for Mandarin non-native mispronunciation verification
- A-CRNN: A Domain Adaptation Model for Sound Event Detection
- ADI17: A Fine-Grained Arabic Dialect Identification Dataset
- ADMM-Based One-Bit Quantized Signal Detection for Massive MIMO Systems With Hardware Impairments
- ADRN: Attention-Based Deep Residual Network for Hyperspectral Image Denoising
- AL2: Progressive Activation Loss for Learning General Representations in Classification Neural Networks
- APB2FACE: Audio-Guided Face Reenactment with Auxiliary Pose and Blink Signals
- ASR Error Correction and Domain Adaptation Using Machine Translation
- ASR is All You Need: Cross-Modal Distillation for Lip Reading
- AV(SE)2: Audio-Visual Squeeze-Excite Speech Enhancement
- Accelerating Distributed Deep Learning By Adaptive Gradient Quantization
- Accelerating Linear Algebra Kernels on a Massively Parallel Reconfigurable Architecture
- Accent Estimation of Japanese Words from Their Surfaces and Romanizations for Building Large Vocabulary Accent Dictionaries
- Accounting for Microprosody in Modeling Intonation
- Accuracy-Robustness Trade-Off for Positively Weighted Neural Networks
- Accurate 6D Object Pose Estimation by Pose Conditioned Mesh Reconstruction
- Accurate Localization of AUV in Motion by Explicit Solution Using Time Delays
- Accurate Semidefinite Relaxation Method for 3-D Rigid Body Localization Using AOA
- Accurate and Scalable Version Identification Using Musically-Motivated Embeddings
- Achieving Fully-Digital Performance by Hybrid Analog/Digital Beamforming in Wide-Band Massive-Mimo Systems
- Achieving the Capacity of the DNA Storage Channel
- Acoustic Matching By Embedding Impulse Responses
- Acoustic Model Adaptation for Presentation Transcription and Intelligent Meeting Assistant Systems
- Acoustic Scene Classification Using Deep Residual Networks with Late Fusion of Separated High and Low Frequency Paths
- Acoustic Scene Classification for Mismatched Recording Devices Using Heated-Up Softmax and Spectrum Correction
- Action-Manipulation Attacks on Stochastic Bandits
- Active Control of Line Spectral Noise with Simultaneous Secondary Path Modeling Without Auxiliary Noise
- Active Learning with Unsupervised Ensembles of Classifiers
- Active Noise Control Over Multiple Regions: Performance Analysis
- Active Semi-Supervised Learning for Diffusions on Graphs
- Acu-Net: A 3D Attention Context U-Net for Multiple Sclerosis Lesion Segmentation
- Adaptation and Learning in Multi-Task Decision Systems
- Adaptation of RNN Transducer with Text-To-Speech Technology for Keyword Spotting
- Adaptive Blind Audio Source Extraction Supervised By Dominant Speaker Identification Using X-Vectors
- Adaptive Distributed Stochastic Gradient Descent for Minimizing Delay in the Presence of Stragglers
- Adaptive Elastic Loss Based on Progressive Inter-Class Association for Cervical Histology Image Segmentation
- Adaptive Knowledge Distillation Based on Entropy
- Adaptive Matched Filter using Non-Target Free Training Data
- Adaptive Normalization for Forecasting Limit Order Book Data Using Convolutional Neural Networks
- Adaptive Prediction of Financial Time-Series for Decision-Making Using A Tensorial Aggregation Approach
- Adaptive Region Aggregation Network: Unsupervised Domain Adaptation with Adversarial Training for ECG Delineation
- Adaptive Sequential Interpolator Using Active Learning for Efficient Emulation of Complex Systems
- Adaptive Subspace Detectors for off-grid Mismatched Targets
- Addressing Accent Mismatch In Mandarin-English Code-Switching Speech Recognition
- Addressing Challenges in Building Web-Scale Content Classification Systems
- Addressing The Confounds Of Accompaniments In Singer Identification
- Addressing the Polysemy Problem in Language Modeling with Attentional Multi-Sense Embeddings
- AdvMS: A Multi-Source Multi-Cost Defense Against Adversarial Attacks
- Adversarial Anomaly Detection for Marked Spatio-Temporal Streaming Data
- Adversarial Attacks on Deep Unfolded Networks for Sparse Coding
- Adversarial Attacks on GMM I-Vector Based Speaker Verification Systems
- Adversarial Detection of Counterfeited Printable Graphical Codes: Towards "Adversarial Games" In Physical World
- Adversarial Example Detection by Classification for Deep Speech Recognition
- Adversarial Mixup Synthesis Training for Unsupervised Domain Adaptation
- Adversarial Multi-Task Learning for Speaker Normalization in Replay Detection
- Adversarial Networks for Secure Wireless Communications
- Adversarial Text Image Super-Resolution using Sinkhorn Distance
- Adversarial Video Compression Guided by Soft Edge Detection
- Age of Information with Finite Horizon and Partial Updates
- Age-Based Scheduling Policy for Federated Learning in Mobile Edge Networks
- Aipnet: Generative Adversarial Pre-Training of Accent-Invariant Networks for End-To-End Speech Recognition
- Algorithmic Exploration of American English Dialects
- Alignment-Length Synchronous Decoding for RNN Transducer
- Aligntts: Efficient Feed-Forward Text-to-Speech System Without Explicit Alignment
- All In One Network for Driver Attention Monitoring
- All You Need is a Second Look: Towards Tighter Arbitrary Shape Text Detection
- Allocation of Computing Tasks In Distributed MEC Servers Co-Powered By Renewable Sources And The Power Grid
- Alternative Half-Sample Interpolation Filters for Versatile Video Coding
- An Acoustic Modelling Based Remote Error Sensing Approach for Quiet Zone Generation in a Noisy Environment
- An Adaptive Linear Estimator Based Approach to Bi-Directional Motion Compensated Prediction
- An Alternative Signature Design Using L1 Principal Components for Spread-Spectrum Steganography
- An Analysis of Speech Enhancement and Recognition Losses in Limited Resources Multi-Talker Single Channel Audio-Visual ASR
- An Analytical Solution to Jacobsen Estimator for Windowed Signals
- An Attention Enhanced Multi-Task Model for Objective Speech Assessment in Real-World Environments
- An Attention-Based Joint Acoustic and Text on-Device End-To-End Model
- An Early Termination Scheme for Successive Cancellation List Decoding of Polar Codes
- An Easy-to-Implement Framework of Fast Subspace Clustering For Big Data Sets
- An Efficient Augmented Lagrangian-Based Method for Linear Equality-Constrained Lasso
- An Efficient Methodology to De-Anonymize the 5G-New Radio Physical Downlink Control Channel
- An Empirical Bayes Approach to Partially Labeled and Shuffled Data Sets
- An Empirical Study of Conv-Tasnet
- An Empirical Study of Transformer-Based Neural Language Model Adaptation
- An Empirical Study on Acoustic Feedback Path Across Hearing Aid Users
- An Enhanced Decoding Algorithm for Coded Compressed Sensing
- An Ensemble Based Approach for Generalized Detection of Spoofing Attacks to Automatic Speaker Recognizers
- An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal Scales
- An Improved Frame-Unit-Selection Based Voice Conversion System Without Parallel Training Data
- An Improved Selective Active Noise Control Algorithm Based on Empirical Wavelet Transform
- An Improved Solution to the Frequency-Invariant Beamforming with Concentric Circular Microphone Arrays
- An LSTM Based Architecture to Relate Speech Stimulus to Eeg
- An LSTM-Based Dynamic Chord Progression Generation System for Interactive Music Performance
- An Odorant Encoding Machine for Sampling, Reconstruction and Robust Representation of Odorant Identity
- An Online Kernel Scalar Quantization Scheme for Signal Classification
- An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation
- An Ontology-Aware Framework for Audio Event Classification
- An Optimal Channel Estimation Scheme for Intelligent Reflecting Surfaces Based on a Minimum Variance Unbiased Estimator
- An Optimal Symmetric Threshold Strategy for Remote Estimation Over The Collision Channel
- An Unsupervised Retinal Vessel Extraction and Segmentation Method Based On a Tube Marked Point Process Model
- Analysis of Acoustic Features for Speech Sound Based Classification of Asthmatic and Healthy Subjects
- Analyzing ASR Pretraining for Low-Resource Speech-to-Text Translation
- Anefficient Alternative to Network Pruning Through Ensemble Learning
- Angular Discriminative Deep Feature Learning for Face Verification
- Anomalous Sound Detection Based on Interpolation Deep Neural Network
- Anomaly Detection for Time Series Using VAE-LSTM Hybrid Model
- Anomaly Detection in Mixed Time-Series Using A Convolutional Sparse Representation With Application To Spacecraft Health Monitoring
- Anomaly Detection with Training Data in Hyperspectral Imagery
- Anomalydae: Dual Autoencoder for Anomaly Detection on Attributed Networks
- Anti-Jamming Routing For Internet of Satellites: a Reinforcement Learning Approach
- Anytime Minibatch with Delayed Gradients: System Performance and Convergence Analysis
- Application Informed Motion Signal Processing for Finger Motion Tracking Using Wearable Sensors
- Approaching Optimal Embedding In Audio Steganography With GAN
- Approximate Bayesian Computation with the Sliced-Wasserstein Distance
- Approximate Inference by Kullback-Leibler Tensor Belief Propagation
- Arnet: Attention-Based Refinement Network for Few-Shot Semantic Segmentation
- Array-Geometry-Aware Spatial Active Noise Control Based on Direction-of-Arrival Weighting
- Arsm Gradient Estimator for Supervised Learning to Rank
- Artificial Bandwidth Extension Using Conditional Variational Auto-encoders and Adversarial Learning
- Assessing the Scope of Generalized Countermeasures for Anti-Spoofing
- Assimilation-Based Learning of Chaotic Dynamical Systems from Noisy and Partial Data
- Asymptotic Stochastic Analysis of Partially Relaxed DML
- Asymptotically Optimal Blind Calibration of Acoustic Vector Sensor Uniform Linear Arrays
- Asynchrounous Decentralized Learning of a Neural Network
- Atomic Norm Based Localization of Far-Field and Near-Field Signals with Generalized Symmetric Arrays
- Atomic Norm Denoising In Blind Two-Dimensional Super-Resolution
- Atrial Fibrillation Risk Prediction from Electrocardiogram and Related Health Data with Deep Neural Network
- Attention Driven Fusion for Multi-Modal Emotion Recognition
- Attention Guided Region Division for Crowd Counting
- Attention Mechanism Enhanced Kernel Prediction Networks for Denoising of Burst Images
- Attention-Based ASR with Lightweight and Dynamic Convolutions
- Attention-Based Curiosity-Driven Exploration in Deep Reinforcement Learning
- Attention-Based Gated Scaling Adaptive Acoustic Model for CTC-Based Speech Recognition
- Attention-Guided Deraining Network Via Stage-Wise Learning
- Attention-Mask Dense Merger (Attendense) Deep HDR for Ghost Removal
- Attentional Fused Temporal Transformation Network for Video Action Recognition
- Attentive Cutmix: An Enhanced Data Augmentation Approach for Deep Learning Based Image Classification
- Attentive Item2vec: Neural Attentive User Representations
- Attentive Modality Hopping Mechanism for Speech Emotion Recognition
- Audio Codec Enhancement with Generative Adversarial Networks
- Audio Feature Extraction for Vehicle Engine Noise Classification
- Audio Sound Determination Using Feature Space Attention Based Convolution Recurrent Neural Network
- Audio-Assisted Image Inpainting for Talking Faces
- Audio-Attention Discriminative Language Model for ASR Rescoring
- Audio-Based Auto-Tagging With Contextual Tags for Music
- Audio-Based Detection of Explicit Content in Music
- Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT
- Audio-Visual Recognition of Overlapped Speech for the LRS2 Dataset
- Auditory Model Based Subsetting of Head-Related Transfer Function Datasets
- Auglabel: Exploiting Word Representations to Augment Labels for Face Attribute Classification
- Augmentation Data Synthesis Via Gans: Boosting Latent Fingerprint Reconstruction
- Augmented Grad-CAM: Heat-Maps Super Resolution Through Augmentation
- Augmenting Molecular Images with Vector Representations as a Featurization Technique for Drug Classification
- Auto-Fas: Searching Lightweight Networks for Face Anti-Spoofing
- Automatic Classification of Volumes of Water Using Swallow Sounds from Cervical Auscultation
- Automatic Data Augmentation Via Deep Reinforcement Learning for Effective Kidney Tumor Segmentation
- Automatic Epileptic Seizure Onset-Offset Detection Based On CNN in Scalp EEG
- Automatic Event Detection of REM Sleep Without Atonia From Polysomnography Signals Using Deep Neural Networks
- Automatic Fluency Evaluation of Spontaneous Speech Using Disfluency-Based Features
- Automatic Identification of Speakers From Head Gestures in a Narration
- Automatic Lyrics Alignment and Transcription in Polyphonic Music: Does Background Music Help?
- Automatic Prediction of Suicidal Risk in Military Couples Using Multimodal Interaction Cues from Couples Conversations
- Automatic and Simultaneous Adjustment of Learning Rate and Momentum for Stochastic Gradient-based Optimization Methods
- Automotive Collision Risk Estimation Under Cooperative Sensing
- Automotive Radar Signal Interference Mitigation Using RNN with Self Attention
- Autoregressive Parameter Estimation with Dnn-Based Pre-Processing
- Auxiliary Capsules for Natural Language Understanding
- Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection
- BBA-NET: A Bi-Branch Attention Network For Crowd Counting
- BBAND INDEX: A NO-REFERENCE BANDING ARTIFACT PREDICTOR
- BOFFIN TTS: Few-Shot Speaker Adaptation by Bayesian Optimization
- BP-VB-EP Based Static and Dynamic Sparse Bayesian Learning with Kronecker Structured Dictionaries
- Back-And-Forth Prediction for Deep Tensor Compression
- Back-to-Back Butterfly Network, an Adaptive Permutation Network for New Communication Standards
- Balanced Binary Neural Networks with Gated Residual
- Balancing Rates and Variance via Adaptive Batch-Sizes in First-Order Stochastic Optimization
- Bandit Sampling for Faster Activity and Data Detection in Massive Random Access
- Bandwidth Extension of Musical Audio Signals With No Side Information Using Dilated Convolutional Neural Networks
- Bangla Voice Command Recognition in end-to-end System Using Topic Modeling based Contextual Rescoring
- Batman: Bayesian Target Modelling For Active Inference
- Bayesian Estimation of Plda with Noisy Training Labels, with Applications to Speaker Verification
- Bayesian Multiple Change-Point Detection with Limited Communication
- Beam Elimination Based on Sequentially Estimated a Posteriori Probabilities of Winning
- Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain Beamformer
- Beamformed Feature for Learning-based Dual-channel Speech Separation
- Beamforming Design for High-Resolution Low-Intensity Focused Ultrasound Neuromodulation
- Beamforming in Intelligent Environments based on Ultra-Massive MIMO Platforms in Millimeter Wave and Terahertz Bands
- Bert is Not All You Need for Commonsense Inference
- Better Safe Than Sorry: Risk-Aware Nonlinear Bayesian Estimation
- Beyond the Dcase 2017 Challenge on Rare Sound Event Detection: A Proposal for a More Realistic Training and Test Framework
- Bilateral Recurrent Network for Single Image Deraining
- Binary Probability Model for Learning Based Image Compression
- Binaural Audio Source Remixing with Microphone Array Listening Devices
- Bio-Mimetic Attentional Feedback in Music Source Separation
- Bipartite Belief Propagation Polar Decoding With Bit-Flipping
- Bit Allocation for Multi-Task Collaborative Intelligence
- Blaster: An Off-Grid Method for Blind and Regularized Acoustic Echoes Retrieval
- Blind Bounded Source Separation Using Neural Networks with Local Learning Rules
- Blind Hyperspectral Unmixing using Dual Branch Deep Autoencoder with Orthogonal Sparse Prior
- Blind Inference of Centrality Rankings from Graph Signals
- Blind Multi-Spectral Image Pan-Sharpening
- Blind Source Separation of Graph Signals
- Blood Pressure Estimation From PPG Signals Using Convolutional Neural Networks And Siamese Network
- Body Movement Generation for Expressive Violin Performance Applying Neural Networks
- Boosted Locality Sensitive Hashing: Discriminative Binary Codes for Source Separation
- Breathing and Speech Planning in Spontaneous Speech Synthesis
- Bridging Mixture Density Networks with Meta-Learning for Automatic Speaker Identification
- Bringing in the Outliers: A Sparse Subspace Clustering Approach to Learn a Dictionary of Mouse Ultrasonic Vocalizations
- Building Firmly Nonexpansive Convolutional Neural Networks
- But System for the Second Dihard Speech Diarization Challenge
- Byzantine-Robust Decentralized Stochastic Optimization
- C3DVQA: Full-Reference Video Quality Assessment with 3D Convolutional Neural Network
- CAD-AEC: Context-Aware Deep Acoustic Echo Cancellation
- CGCNN: Complex Gabor Convolutional Neural Network on Raw Speech
- CIF: Continuous Integrate-And-Fire for End-To-End Speech Recognition
- CLCNET: Deep Learning-Based Noise Reduction for Hearing aids using Complex Linear Coding
- CN-Celeb: A Challenging Chinese Speaker Recognition Dataset
- CNN-Based Analog CSI Feedback in FDD MIMO-OFDM Systems
- CORRGAN: Sampling Realistic Financial Correlation Matrices Using Generative Adversarial Networks
- CP-GAN: Context Pyramid Generative Adversarial Network for Speech Enhancement
- CPWC: Contextual Point Wise Convolution for Object Recognition
- CS-R-FCN: Cross-Supervised Learning for Large-Scale Object Detection
- Camera Configuration Design in Cooperative Active Visual 3d Reconstruction: A Statistical Approach
- Can every analog system be simulated on a digital computer?
- Capacity of the Erasure Shuffling Channel
- Cartoon-Texture Decomposition-Based Variational Pansharpening
- Cell-Phone Classification: A Convolutional Neural Network Approach Exploiting Electromagnetic Emanations
- Challenges and Perspectives in Neuromorphic-based Visual IoT Systems and Networks
- Channel Adversarial Training for Speaker Verification and Diarization
- Channel Attention Based Generative Network for Robust Visual Tracking
- Channel Charting: an Euclidean Distance Matrix Completion Perspective
- Channel Covariance Estimation in Multiuser Massive Mimo Systems with an Approach Based on Infinite Dimensional Hilbert Spaces
- Channel Invariant Speaker Embedding Learning with Joint Multi-Task and Adversarial Training
- Channel Selection over Riemannian Manifold with Non-Stationarity Consideration for Brain-Computer Interface Applications
- Channel-Attention Dense U-Net for Multichannel Speech Enhancement
- Characterisation of a Snapshot Fourier Transform Imaging Spectrometer Based on an Array of Fabry-Perot Interferometers
- Characterizing Speech Adversarial Examples Using Self-Attention U-Net Enhancement
- Chirping up the Right Tree: Incorporating Biological Taxonomies into Deep Bioacoustic Classifiers
- Classification of Depth and Surface Edges with Deep Features
- Classification of Epileptic IEEG Signals by CNN and Data Augmentation
- Classification of High-Dimensional Motor Imagery Tasks Based on An End-To-End Role Assigned Convolutional Neural Network
- Classify and Explain: An Interpretable Convolutional Neural Network For Lung Cancer Diagnosis
- Classifying Anomalies for Network Security
- Classifying Partially Labeled Networked Data VIA Logistic Network Lasso
- Clock Synchronization Over Networks Using Sawtooth Models
- Clotho: an Audio Captioning Dataset
- Cloud-Driven Multi-Way Multiple-Antenna Relay Systems: Best-User-Link Selection and Joint Mmse Detection
- Clustering of Nonnegative Data and an Application to Matrix Completion
- Clutter Identification Based on Sparse Recovery and L1-Type Probabilistic Distance Measures
- Cochlear Signal Processing: A Platform for Learning the Fundamentals of Digital Signal Processing
- Code-Switched Speech Synthesis Using Bilingual Phonetic Posteriorgram with Only Monolingual Corpora
- Coded Illumination and Multiplexing for Lensless Imaging
- Cogans For Unsupervised Visual Speech Adaptation To New Speakers
- Coincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervision
- Color Stabilization for Multi-Camera Light-Field Imaging
- Color and Angular Reconstruction of Light Fields from Incomplete-Color Coded Projections
- Colour Compression of Plenoptic Point Clouds Using Raht-Klt with Prior Colour Clustering and Specular/Diffuse Component Separation
- Combining Acoustics, Content and Interaction Features to Find Hot Spots in Meetings
- Combining CGAN and Mil for Hotspot Segmentation in Bone Scintigraphy
- Combining Deep Embeddings of Acoustic and Articulatory Features for Speaker Identification
- Communication Constrained Learning with Uncertain Models
- Commuting Conditional GANS for Multi-Modal Fusion
- Compare Learning: Bi-Attention Network for Few-Shot Learning
- Comparison of Glottal Closure Instants Detection Algorithms for Emotional Speech
- Comparison of User Models Based on GMM-UBM and I-Vectors for Speech, Handwriting, and Gait Assessment of Parkinson's Disease Patients
- Complex Pairwise Activity Analysis Via Instance Level Evolution Reasoning
- Complex Trainable Ista for Linear and Nonlinear Inverse Problems
- Complex Transformer: A Framework for Modeling Complex-Valued Sequence
- Complexity Reduction Methods for Index Modulation Based Dual-Function Radar Communication Systems
- Composite Dynamic Texture Synthesis Using Hierarchical Linear Dynamical System
- Compressed Sensing Based Channel Estimation and Open-loop Training Design for Hybrid Analog-digital Massive MIMO Systems
- Compressing Flow Fields with Edge-Aware Homogeneous Diffusion Inpainting
- Compressive 2-d Off-grid DOA Estimation for Propeller Cavitation Localization
- Compressive Adaptive Bilateral Filtering
- Computability of the Peak Value of Bandlimited Signals
- Computation of "Best" Interpolants in the Lp Sense
- Computing Hilbert Transform and Spectral Factorization for Signal Spaces of Smooth Functions
- Concentration-Based Polynomial Calculations on Nicked DNA
- Conditional Density Driven Grid Design in Point-Mass Filter
- Conditional Domain Adversarial Transfer for Robust Cross-Site ADHD Classification Using Functional MRI
- Conditional Mutual Information Neural Estimator
- Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks
- Confirmnet: Convolutional Firmnet and Application to Image Denoising and Inpainting
- Consensus-Based Distributed Clustering for IoT
- Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural Networks
- Constant Envelope Massive MIMO-OFDM Precoding: an Improved Formulation and Solution
- Constant-Envelope Precoding for Satellite Systems
- Constrained Spectral Clustering for Dynamic Community Detection
- Content Based Singing Voice Extraction from a Musical Mixture
- Content Vs Context: How About "Walking Hand-In-Hand" For Image Clustering?
- Context and Uncertainty Modeling for Online Speaker Change Detection
- Continual Learning Through One-Class Classification Using VAE
- Continual Learning for Infinite Hierarchical Change-Point Detection
- Continuous Speech Separation: Dataset and Analysis
- Control of Linear Dynamical Systems Using Sparse Inputs
- Controllable Time-Delay Transformer for Real-Time Punctuation Prediction and Disfluency Detection
- Controlling the Perceived Sound Quality for Dialogue Enhancement With Deep Learning
- Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T Distribution
- Converting Written Language to Spoken Language with Neural Machine Translation for Language Modeling
- Convex Optimisation-Based Privacy-Preserving Distributed Average Consensus in Wireless Sensor Networks
- Convolutional Beamspace for Array Signal Processing
- Cooperative Learning VIA Federated Distillation OVER Fading Channels
- Corrdrop: Correlation Based Dropout for Convolutional Neural Networks
- Correction of Automatic Speech Recognition with Transformer Sequence-To-Sequence Model
- Correlated Multi-Armed Bandits with A Latent Random Source
- Cost Aware Adversarial Learning
- Counting Dense Objects in Remote Sensing Images
- Coupled Training of Sequence-to-Sequence Models for Accented Speech Recognition
- Cra: A Generic Compression Ratio Adapter for End-To-End Data-Driven Image Compressive Sensing Reconstruction Frameworks
- Cramer-Rao Bound on DOA Estimation of Finite Bandwidth Signals Using a Moving Sensor
- Cramér-Rao Bounds for Flaw Localization in Subsampled Multistatic Multichannel Ultrasound Ndt Data
- Crnn-Ctc Based Mandarin Keywords Spotting
- Cross Image Cubic Interpolator for Spatially Varying Exposures
- Cross Lingual Transfer Learning for Zero-Resource Domain Adaptation
- Cross-Domain Adaptation for Biometric Identification Using Photoplethysmogram
- Cross-Domain Joint Dictionary Learning for ECG Reconstruction from PPG
- Cross-Lingual Topic Prediction For Speech Using Translations
- Cross-Speaker Silent-Speech Command Word Recognition Using Electro-Optical Stomatography
- Cross-Stained Segmentation from Renal Biopsy Images Using Multi-Level Adversarial Learning
- Cross-VAE: Towards Disentangling Expression from Identity For Human Faces
- Cross-View Attention Network for Breast Cancer Screening from Multi-View Mammograms
- Crowdsourcing-Based Ranking Aggregation for Person Re-Identification
- Cumulant Slice Reconstruction from Compressive Measurements and Its Application to Line Spectrum Estimation
- D-SLAM: Diffusion Source Localization and Trajectory Mapping
- D2NA: Day-To-Night Adaptation for Vision based Parking Management System
- DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks
- DGAN: Disentangled Representation Learning for Anisotropic BRDF Reconstruction
- DNN-Based Speech Presence Probability Estimation for Multi-Frame Single-Microphone Speech Enhancement
- DNN-Based Speech Recognition for Globalphone Languages
- DNN-Chip Predictor: An Analytical Performance Predictor for DNN Accelerators with Various Dataflows and Hardware Architectures
- DNN-based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone Arrays
- DNN-based Mask Estimation Integrating Spectral and Spatial Features for Robust Beamforming
- DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source Separation
- DOA Estimation in Systems with Nonlinearities for MMWAVE Communications
- DOA Tracking Via Signal-Subspace Projector Update
- Damage-Sensitive and Domain-Invariant Feature Extraction for Vehicle-Vibration-Based Bridge Health Monitoring
- Data Augmentation Using Empirical Mode Decomposition on Neural Networks to Classify Impact Noise in Vehicle
- Data Selection Kernel Conjugate Gradient Algorithm
- Data-Driven Harmonic Filters for Audio Representation Learning
- Data-Driven Model Set Design for Model Averaged Particle Filter
- Data-Driven Wind Speed Estimation Using Multiple Microphones
- Deblurring And Super-Resolution Using Deep Gated Fusion Attention Networks For Face Images
- Decentralized Min-Max Optimization: Formulations, Algorithms and Applications in Network Poisoning Attack
- Decentralized Optimization with Non-Identical Sampling in Presence of Stragglers
- Decentralized Stochastic Non-Convex Optimization over Weakly Connected Time-Varying Digraphs
- Decentralized expected consistent signal recovery for quantization Measurements
- Decidable Variable-Rate Dataflow for Heterogeneous Signal Processing Systems
- Decoding 5G-NR Communications VIA Deep Learning
- Decoding Movement Imagination and Execution From Eeg Signals Using Bci-Transfer Learning Method Based on Relation Network
- Decomposed Cyclegan for Single Image Deraining With Unpaired Data
- Deep Audio-Visual Speech Separation with Attention Mechanism
- Deep Autotuner: A Pitch Correcting Network for Singing Performances
- Deep Casa for Talker-independent Monaural Speech Separation
- Deep Clustering for Domain Adaptation
- Deep Clusteringwith Concrete K-Means
- Deep Contextualized Acoustic Representations for Semi-Supervised Speech Recognition
- Deep Encoded Linguistic and Acoustic Cues for Attention Based End to End Speech Emotion Recognition
- Deep Exposure Fusion with Deghosting via Homography Estimation and Attention Learning
- Deep Flow Collaborative Network for Online Visual Tracking
- Deep Geometric Knowledge Distillation with Graphs
- Deep Image Deblurring Using Local Correlation Block
- Deep James-Stein Neural Networks For Brain-Computer Interfaces
- Deep Joint Source-Channel Coding for Wireless Image Retrieval
- Deep Joint Source-Channel Coding of Images with Feedback
- Deep Learning Abilities to Classify Intricate Variations in Temporal Dynamics of Multivariate Time Series
- Deep Learning Based Prediction of Hypernasality for Clinical Applications
- Deep Learning for Robust Power Control for Wireless Networks
- Deep Learning-Based Beam Alignment in Mmwave Vehicular Networks
- Deep Matrix Completion on Graphs: Application in Drug Target Interaction Prediction
- Deep Meta-Relation Network for Visual Few-Shot Learning
- Deep Metric Learning Based On Center-Ranked Loss for Gait Recognition
- Deep Monocular Video Depth Estimation Using Temporal Attention
- Deep Multi-Region Hashing
- Deep Multi-Scale Gabor Wavelet Network for Image Restoration
- Deep Neural Network Based Matrix Completion for Internet of Things Network Localization
- Deep Neural Networks Based Automatic Speech Recognition for Four Ethiopian Languages
- Deep Product Quantization Module for Efficient Image Retrieval
- Deep Rainrate Estimation from Highly Attenuated Downlink Signals of Ground-Based Communications Satellite Terminals
- Deep Residual Network for MSFA Raw Image Denoising
- Deep Soft Interference Cancellation for MIMO Detection
- Deep Speech Extraction with Time-Varying Spatial Filtering Guided By Desired Direction Attractor
- Deep-Neural-Network Based Fall-Back Mechanism in Interference-Aware Receiver Design
- Deep-SST-Eddies: A Deep Learning Framework to Detect Oceanic Eddies in Sea Surface Temperature Images
- Defending Graph Convolutional Networks Against Adversarial Attacks
- Defense Against Adversarial Attacks on Spoofing Countermeasures of ASV
- Deliberation Model Based Two-Pass End-To-End Speech Recognition
- Demystifying TasNet: A Dissecting Approach
- Denoising of Event-Based Sensors with Spatial-Temporal Correlation
- Dense Mapping of Intracellular Diffusion and Drift from Single-Particle Tracking Data Analysis
- Dense Residual Network for Retinal Vessel Segmentation
- Densely Connected Neural Network with Dilated Convolutions for Real-Time Speech Enhancement in The Time Domain
- Depth Estimation From Single Image Through Multi-Path-Multi-Rate Diverse Feature Extractor
- Depth Map Fingerprinting and Splicing Detection
- Depthwise-STFT Based Separable Convolutional Neural Networks
- Deriving Compact Feature Representations Via Annealed Contraction
- Design Considerations for Hypothesis Rejection Modules in Spoken Language Understanding Systems
- Design of A Convergence-Aware Based Expectation Propagation Algorithm for Uplink Mimo Scma Systems
- Design-Gan: Cross-Category Fashion Translation Driven By Landmark Attention
- Detect Insider Attacks Using CNN in Decentralized Optimization
- Detecting Adversarial Attacks In Time-Series Data
- Detecting Autism Spectrum Disorder Using Topological Data Analysis
- Detecting Emotion Primitives from Speech and Their Use in Discerning Categorical Emotions
- Detecting Mismatch Between Text Script and Voice-Over Using Utterance Verification Based on Phoneme Recognition Ranking
- Detecting Multiple Speech Disfluencies Using a Deep Residual Network with Bidirectional Long Short-Term Memory
- Detection Of S1 And S2 Locations In Phonocardiogram Signals Using Zero Frequency Filter
- Detection and Analysis of T/D Deletion in Librispeech
- Detection of Adversarial Attacks and Characterization of Adversarial Subspace
- Detection of Malicious Vbscript Using Static and Dynamic Analysis with Recurrent Deep Learning
- Detection of Mild Dyspnea from Pairs of Speech Recordings
- Detection of Speech Events and Speaker Characteristics through Photo-Plethysmographic Signal Neural Processing
- Determined Source Separation Using the Sparsity of Impulse Responses
- Deterministic Feature Decoupling by Surfing Invariance Manifolds
- Dfsmn-San with Persistent Memory Model for Automatic Speech Recognition
- Diacritic-Level Pronunciation Analysis Using Phonological Features
- Diagonalizable Shift and Filters for Directed Graphs Based on the Jordan-Chevalley Decomposition
- Dialogue History Integration into End-to-End Signal-to-Concept Spoken Language Understanding Systems
- Differentiable Branching In Deep Networks for Fast Inference
- Digital Watermarking For Protecting Audio Classification Datasets
- Dilated Convolutional Neural Networks for Panoramic Image Saliency Prediction
- Discovering Causalities from Cardiotocography Signals using Improved Convergent Cross Mapping with Gaussian Processes
- Discrete Wasserstein Autoencoders for Document Retrieval
- Discriminant Generative Adversarial Networks with its Application to Equipment Health Classification
- Discriminant and Sparsity Based Least Squares Regression with l1 Regularization for Feature Representation
- Disentangled Multidimensional Metric Learning for Music Similarity
- Disentangled Speech Embeddings Using Cross-Modal Self-Supervision
- Disentangling Controllable Object Through Video Prediction Improves Visual Reinforcement Learning
- Disentangling Timbre and Singing Style with Multi-Singer Singing Synthesis System
- Dispersive Grid-free Orthogonal Matching Pursuit for Modal Estimation in Ocean Acoustics
- Distilling Attention Weights for CTC-Based ASR Systems
- Distributed Detection of Sparse Signals with 1-Bit Data in Two-Level Two-Degree Tree-Structured Sensor Networks
- Distributed Equalization and Power Allocation For Multi-Carrier Bidirectional Filter-and-Forward Relay Networks
- Distributed Non-Orthogonal Pilot Design for Multi-Cell Massive Mimo Systems
- Distributed Quantization for Sparse Time Sequences
- Distributed Tensor Completion Over Networks
- Distributed Tracking and Circumnavigation Using Bearing Measurements
- Distributed Verification of Belief Precisions Convergence in Gaussian Belief Propagation
- Distributed Wave-Domain Active Noise Control Based on the Diffusion Strategy
- Distribution of the Product of a Complex Gaussian Matrix and Vector and Its Sum with a Complex Gaussian Vector
- Divergence-Based Adaptive Extreme Video Completion
- Diversity and Sparsity: A New Perspective on Index Tracking
- Domain Adaptation for Generalization of Face Presentation Attack Detection in Mobile Settengs with Minimal Information
- Domain Robust, Fast, and Compact Neural Language Models
- Drift Detection and Correction Post-Tracking
- Drss-Based Localisation Using Weighted Instrumental Variables and Selective Power Measurement
- Dual-Path RNN: Efficient Long Sequence Modeling for Time-Domain Single-Channel Speech Separation
- Duration Robust Weakly Supervised Sound Event Detection
- Dyna-Bolt: Domain Adaptive Binary Factorization Of Current Waveforms For Energy Disaggregation
- Dynamic Attack Scoring Using Distributed Local Detectors
- Dynamic Channel Pruning For Correlation Filter Based Object Tracking
- Dynamic Metasurface Antennas for Bit-Constrained MIMO-OFDM Receivers
- Dynamic Oversampling in 1-Bit Quantized Asynchronous Large-Scale Multiple-Antenna Systems for Sustainable Iot Networks
- Dynamic Resource Allocation for Wireless Edge Machine Learning with Latency And Accuracy Guarantees
- Dynamic Resource Optimization and Altitude Selection in Uav-Based Multi-Access Edge Computing
- Dynamic Temporal Residual Learning for Speech Recognition
- Dynamic Variational Autoencoders for Visual Process Modeling
- Dynamically Modulated Deep Metric Learning for Visual Search
- Dysarthric Speech Recognition with Lattice-Free MMI
- E2E-SINCNET: Toward Fully End-To-End Speech Recognition
- ECG Heartbeat Classification Based on Multi-Scale Wavelet Convolutional Neural Networks
- EDNFC-Net: Convolutional Neural Network with Nested Feature Concatenation for Nuclei-Instance Segmentation
- EMET: Embeddings from Multilingual-Encoder Transformer for Fake News Detection
- EPI-Neighborhood Distribution Based Light Field Depth Estimation
- ESRGAN+ : Further Improving Enhanced Super-Resolution Generative Adversarial Network
- Edgefool: an Adversarial Image Enhancement Filter
- Eeg Connectivity - Informed Cooperative Adaptive Line Enhancer for Recognition of Brain State
- Eeg Feature Selection Using Orthogonal Regression: Application to Emotion Recognition
- Effect of Choice of Probability Distribution, Randomness, and Search Methods for Alignment Modeling in Sequence-to-Sequence Text-to-Speech Synthesis Using Hard Alignment
- Effect of Frication Duration and Formant Transitions on the Perception of Fricatives in VCV Utterances
- Effect of Undersampling on Non-Negative Blind Deconvolution with Autoregressive Filters
- Effective Approximate Maximum Likelihood Estimation of Angles of Arrival for Non-Coherent Sub-Arrays
- Effective Approximation of Bandlimited Signals and Their Samples
- Effective Pipeline for Compressing Deep Object Detectors
- Effective Wavenet Adaptation for Voice Conversion with Limited Data
- Effectiveness of Random Deep Feature Selection for Securing Image Manipulation Detectors Against Adversarial Examples
- Effectiveness of Self-Supervised Pre-Training for ASR
- Effects of Spectral Tilt on Listeners' Preferences And Intelligibility
- Efficient Algorithm to Implement Sliding Singular Spectrum Analysis with Application to Biomedical Signal Denoising
- Efficient Belief Propagation for Graph Matching
- Efficient Bird Sound Detection on the Bela Embedded System
- Efficient Constrained Encoders Correcting a Single Nucleotide Edit in DNA Storage
- Efficient Decoupled Neural Architecture Search by Structure And Operation Sampling
- Efficient Deep Learning-Based Lossy Image Compression Via Asymmetric Autoencoder and Pruning
- Efficient Estimation of Mixing Matrix Using a Two-sensor Array
- Efficient Image Super Resolution Via Channel Discriminative Deep Neural Network Pruning
- Efficient Multichannel Nonlinear Acoustic Echo Cancellation Based on a Cooperative Strategy
- Efficient Scene Text Detection with Textual Attention Tower
- Efficient Shallow Wavenet Vocoder Using Multiple Samples Output Based on Laplacian Distribution and Linear Prediction
- Efficient Super-Resolution Two-Dimensional Harmonic Retrieval Via Enhanced Low-Rank Structured Covariance Reconstruction
- Efficient Techniques For in-band System Information Broadcast in Multi-Cell Massive Mimo
- Efficient Trainable Front-Ends for Neural Speech Enhancement
- Efficient and Scalable Neural Residual Waveform Coding with Collaborative Quantization
- Electric Analog Circuit Design with Hypernetworks And A Differential Simulator
- Electro-Magnetic Side-Channel Attack Through Learned Denoising and Classification
- Eliminating Out-Of-Cell Interference in Cellular Massive Mimo with a Single Additional Transceiver
- Embedded Large-Scale Handwritten Chinese Character Recognition
- Emotional Speech Synthesis with Rich and Granularized Control
- Emotional Voice Conversion Using Multitask Learning with Text-To-Speech
- Empirical Sure-Guided Microscopy Super-Resolution Image Reconstruction from Confocal Multi-Array Detectors
- Encoder-Recurrent Decoder Network for Single Image Dehazing
- Encoding Temporal Information For Automatic Depression Recognition From Facial Analysis
- Encoding and Decoding Mixed Bandlimited Signals Using Spiking Integrate-and-Fire Neurons
- End to End Speech Recognition Error Prediction with Sequence to Sequence Learning
- End-To-End Accent Conversion Without Using Native Utterances
- End-To-End Auditory Object Recognition Via Inception Nucleus
- End-To-End Generation of Talking Faces from Noisy Speech
- End-To-End Multi-Speaker Speech Recognition With Transformer
- End-To-End Multi-Talker Overlapping Speech Recognition
- End-To-End Non-Negative Autoencoders for Sound Source Separation
- End-To-End Spoken Language Understanding Without Matched Language Speech Model Pretraining Data
- End-To-End Voice Conversion Via Cross-Modal Knowledge Distillation for Dysarthric Speech Reconstruction
- End-end Speech-to-Text Translation with Modality Agnostic Meta-Learning
- End-to-End Architectures for ASR-Free Spoken Language Understanding
- End-to-End Articulatory Modeling for Dysarthric Articulatory Attribute Detection
- End-to-End Automatic Speech Recognition Integrated with CTC-Based Voice Activity Detection
- End-to-End Code-Switching TTS with Cross-Lingual Language Model
- End-to-End Multi-Person Audio/Visual Automatic Speech Recognition
- End-to-End Speech Translation with Self-Contained Vocabulary Manipulation
- End-to-End Training of Time Domain Audio Separation and Recognition
- End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation
- EnerGAN: A GENERATIVE ADVERSARIAL NETWORK FOR ENERGY DISAGGREGATION
- Energy Disaggregation Using Fractional Calculus
- Energy Disaggregation from Low Sampling Frequency Measurements Using Multi-Layer Zero Crossing Rate
- Energy Efficient Acceleration Of Floating Point Applications Onto CGRA
- Energy-Efficient 3D UAV Trajectory Design for Data Collection in Wireless Sensor Networks
- Energy-Efficient Bit Allocation for Resolution-Adaptive ADC in Multiuser Large-Scale MIMO Systems: Global Optimality
- Enhance Part-Based Model for Person Re-Identification with Fused Multi-Scale Features
- Enhance feature representation of electroencephalogram for Seizure detection
- Enhanced Action Tubelet Detector for Spatio-Temporal Video Action Detection
- Enhanced Adversarial Strategically-Timed Attacks Against Deep Reinforcement Learning
- Enhanced Method of Audio Coding Using CNN-Based Spectral Recovery with Adaptive Structure
- Enhanced Mixture Population Monte Carlo Via Stochastic Optimization and Markov Chain Monte Carlo Sampling
- Enhanced Non-Local Cascading Network with Attention Mechanism for Hyperspectral Image Denoising
- Enhanced Safety of Autonomous Driving by Incorporating Terrestrial Signals of Opportunity
- Enhancement of Coded Speech Using a Mask-Based Post-Filter
- Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature Learning
- Enhancing the Labelling of Audio Samples for Automatic Instrument Classification Based on Neural Networks
- Ensemble Network For Ranking Images Based On Visual Appeal
- Environment-Aware Reconfigurable Noise Suppression
- Epigraphical Reformulation for Non-Proximable Mixed Norms
- Epoch Estimation from a Speech Signal Using Gammatone Wavelets in a Scattering Network
- Equalization of OFDM Waveforms with Insufficient Cyclic Prefix
- Ernet Family: Hardware-Oriented Cnn Models For Computational Imaging Using Block-Based Inference
- Error Analysis Applied to End-to-End Spoken Language Understanding
- Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
- Estimating Centrality Blindly From Low-Pass Filtered Graph Signals
- Estimating Structural Missing Values Via Low-Tubal-Rank Tensor Completion
- Estimating the Degree of Sleepiness by Integrating Articulatory Feature Knowledge in Raw Waveform Based CNNS
- Estimation of Information in Parallel Gaussian Channels via Model Order Selection
- Estimation of Post-Nonlinear Causal Models Using Autoencoding Structure
- Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates
- Evaluating Voice Conversion-Based Privacy Protection against Informed Attackers
- Evaluation of Deep-Learning-Based Voice Activity Detectors and Room Impulse Response Models in Reverberant Environments
- Evaluation of Joint Auditory Attention Decoding and Adaptive Binaural Beamforming Approach for Hearing Devices with Attention Switching
- Evaluation of Sensor Self-Noise In Binaural Rendering of Spherical Microphone Array Signals
- Event-Driven Signal Processing with Neuromorphic Computing Systems
- Exact Sparse Nonnegative Least Squares
- Exocentric to Egocentric Image Generation Via Parallel Generative Adversarial Network
- Experiments in Creating Online Course Content for Signal Processing Education
- Exploitation of 3D City Maps for Hybrid 5G RTT and GNSS Positioning Simulations
- Exploiting Channel Locality for Adaptive Massive MIMO Signal Detection
- Exploiting Commutativity Condition for CP Decomposition Via Approximate Simultaneous Diagonalization
- Exploiting Periodicity Features for Joint Detection and DOA Estimation of Speech Sources Using Convolutional Neural Networks
- Exploiting Rays in Blind Localization of Distributed Sensor Arrays
- Exploiting Sparsity for Robust Sensor Network Localization in Mixed LOS/NLOS Environments
- Exploiting Two-Dimensional Symmetry and Unimodality for Model-Free Source Localization in Harsh Environment
- Exploiting Vocal Tract Coordination Using Dilated CNNS For Depression Detection In Naturalistic Environments
- Exploration Methodology for BTI-Induced Failures on RRAM-Based Edge AI Systems
- Exploring A Zero-Order Direct Hmm Based on Latent Attention for Automatic Speech Recognition
- Exploring Appropriate Acoustic and Language Modelling Choices for Continuous Dysarthric Speech Recognition
- Exploring Bio-Behavioral Signal Trajectories of State Anxiety During Public Speaking
- Exploring Energy Efficient Quantum-resistant Signal Processing Using Array Processors
- Exploring Entity-Level Spatial Relationships for Image-Text Matching
- Exploring Pre-Training with Alignments for RNN Transducer Based End-to-End Speech Recognition
- Exposure Interpolation Via Hybrid Learning
- Expression-Guided EEG Representation Learning for Emotion Recognition
- Extended Cyclic Coordinate Descent for Robust Row-Sparse Signal Reconstruction in the Presence of Outliers
- Extended Object Tracking Using Hierarchical Truncation Measurement Model with Automotive Radar
- Extracting Unit Embeddings Using Sequence-To-Sequence Acoustic Models for Unit Selection Speech Synthesis
- Extrapolated Alternating Algorithms for Approximate Canonical Polyadic Decomposition
- F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional Autoencoder
- FCEM: A Novel Fast Correlation Extract Model For Real Time Steganalysis Of VoIP Stream Via Multi-Head Attention
- FDDWNet: A Lightweight Convolutional Neural Network for Real-Time Semantic Segmentation
- FIR Filter Design and Implementation for Phase-Based Processing
- FIR Filtering of Discontinuous Signals: A Random-Stratified Sampling Approach
- Face Feature Recovery via Temporal Fusion for Person Search
- Facial Emotion Recognition Using Light Field Images with Deep Attention-Based Bidirectional LSTM
- Facial Feature Embedded Cyclegan For Vis-Nir Translation
- Far-Field Location Guided Target Speech Extraction Using End-to-End Speech Recognition Objectives
- Fast Acoustic Scattering Using Convolutional Neural Networks
- Fast Block-Sparse Estimation for Vector Networks
- Fast Clustering With Co-Clustering Via Discrete Non-Negative Matrix Factorization for Image Identification
- Fast Direction-of-arrival Estimation of Multiple Targets Using Deep Learning and Sparse Arrays
- Fast Domain Adaptation for Goal-Oriented Dialogue Using a Hybrid Generative-Retrieval Transformer
- Fast Independent Vector Extraction by Iterative SINR Maximization
- Fast Intent Classification for Spoken Language Understanding Systems
- Fast Lattice-Free Keyword Filtering for Accelerated Spoken Term Detection
- Fast Optical System Identification by Numerical Interferometry
- Fast Single-View 3D Object Reconstruction with Fine Details Through Dilated Downsample and Multi-Path Upsample Deep Neural Network
- Fast Start-Up Algorithm for Adaptive Noise Cancellers with Novel SNR Estimation and Stepsize Control
- Fast Training of Deep Neural Networks for Speech Recognition
- Fast and Accurate Embedded DCNN for Rgb-D Based Sign Language Recognition
- Fast and High-Quality Singing Voice Synthesis System Based on Convolutional Neural Networks
- Fast and Stable Blind Source Separation with Rank-1 Updates
- Faster-Than-Nyquist Signaling Via Spatiotemporal Symbol-Level Precoding for Multi-User MISO Redundant Transmissions
- Favorable Propagation and Linear Multiuser Detection for Distributed Antenna Systems
- Feature Affine Projection Algorithms
- Feature Drift Resilient Tracking of The Carotid Artery Wall Using Unscented Kalman Filtering With Data Fusion
- Feature Enhancement with Deep Feature Losses for Speaker Verification
- Feature Selection Under Orthogonal Regression with Redundancy Minimizing
- Federated Classification with Low Complexity Reproducing Kernel Hilbert Space Representations
- Federated Learning with Mutually Cooperating Devices: A Consensus Approach Towards Server-Less Model Optimization
- Federated Learning with Quantization Constraints
- Federated Neuromorphic Learning of Spiking Neural Networks for Low-Power Edge Intelligence
- Federated Truth Inference over Distributed Crowdsourcing Platforms
- Federating Solar, Storage and Communications in the Electric Grid and Internet of things
- Feedback Recurrent Autoencoder
- Feedback Turbo Autoencoder
- Few-Shot Acoustic Event Detection Via Meta Learning
- Few-Shot Sound Event Detection
- Fg2seq: Effectively Encoding Knowledge for End-To-End Task-Oriented Dialog
- Filterbank Design for End-to-end Speech Separation
- Filtering Out Time-Frequency Areas Using Gabor Multipliers
- Fine-Grained Action Recognition on a Novel Basketball Dataset
- Fine-Grained Giant Panda Identification
- Finite Sample Deviation and Variance Bounds for First Order Autoregressive Processes
- Fixed Smooth Convolutional Layer for Avoiding Checkerboard Artifacts in CNNS
- Fixed-Point Optimization of Transformer Neural Network
- Flexibly-tunable bitcube-based perceptual encryption within jpeg compression
- Flow-TTS: A Non-Autoregressive Network for Text to Speech Based on Flow
- Focus on Semantic Consistency for Cross-Domain Crowd Understanding
- Focusing on Attention: Prosody Transfer and Adaptative Optimization Strategy for Multi-Speaker End-to-End Speech Synthesis
- Forecasting Multi-Dimensional Processes Over Graphs
- Forecasting Sparse Traffic Congestion Patterns Using Message-Passing RNNS
- Foreground Signature Extraction for an Intimate Mixing Model in Hyperspectral Image Classification
- Formulating Divergence Framework for Multiclass Motor Imagery EEG Brain Computer Interface
- Forward-Backward Splitting for Optimal Transport Based Problems
- Fourier Phase Retrieval with Arbitrary Reference Signal
- Fourth Order Cumulant Based Active Direction of Arrival Estimation Using Coprime Arrays
- Fractional Fourier Transform Based QRS Complex Detection in ECG Signal
- Frame-Based Overlapping Speech Detection Using Convolutional Neural Networks
- Frame-Level MMI as A Sequence Discriminative Training Criterion for LVCSR
- Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances
- Frequency Diverse Array Radar: A Closed-Form Solution to Design Weights for Desired Beampattern
- Frequency and Temporal Convolutional Attention for Text-Independent Speaker Recognition
- Frequency-Dependent Directional Feedback Delay Network
- From Symbols to Signals: Symbolic Variational Autoencoders
- From Unsupervised Machine Translation to Adversarial Text Generation
- From Video Game to Real Robot: The Transfer Between Action Spaces
- Full Reference Video Quality Measures Improvement Using Neural Networks
- Full-Reference Speech Quality Estimation with Attentional Siamese Neural Networks
- Full-Sum Decoding for Hybrid Hmm Based Speech Recognition Using LSTM Language Model
- Fully Convolutional Recurrent Networks for Speech Enhancement
- Fully Learnable Front-End for Multi-Channel Acoustic Modeling Using Semi-Supervised Learning
- Fully Pipelined Iteration Unrolled Decoders the Road to TB/S Turbo Decoding
- Fully-Hierarchical Fine-Grained Prosody Modeling For Interpretable Speech Synthesis
- Fully-Neural Approach to Heavy Vehicle Detection on Bridges Using a Single Strain Sensor
- Fusion Approaches for Emotion Recognition from Speech Using Acoustic and Text-Based Features
- Fusionndvi: A Novel Fusion Method for NDVI in Remote Sensing
- G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR
- GCI Detection from Raw Speech Using a Fully-Convolutional Network
- GFCN: A New Graph Convolutional Network Based on Parallel Flows
- GFNet: A Lightweight Group Frame Network for Efficient Human Action Recognition
- Gait Phase Segmentation Using Weighted Dynamic Time Warping and K-Nearest Neighbors Graph Embedding
- Gated Attentive Convolutional Network Dialogue State Tracker
- Gated Mechanism for Attention Based Multi Modal Sentiment Analysis
- Gated Multi-Layer Convolutional Feature Extraction Network for Robust Pedestrian Detection
- Gaussian Lpcnet for Multisample Speech Synthesis
- Gaussian Process Imputation of Multiple Financial Series
- Gaussian Processes Over Graphs
- Gender Differences on the Perception and Production of Utterances with Willingness and Reluctance in Chinese
- Generalized Coherence-Based Signal Enhancement
- Generalized Graph Spectral Sampling with Stochastic Priors
- Generalized Kernel-Based Dynamic Mode Decomposition
- Generalized Linear Bandits with Safety Constraints
- Generalized Spatial Modulation for Wireless Terabits Systems Under Sub-THZ Channel With RF Impairments
- Generating Diverse and Natural Text-to-Speech Samples Using a Quantized Fine-Grained VAE and Autoregressive Prosody Prior
- Generating Empathetic Responses by Looking Ahead the User's Sentiment
- Generating Multilingual Voices Using Speaker Space Translation Based on Bilingual Speaker Data
- Generating Synthetic Audio Data for Attention-Based Speech Recognition Systems
- Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition Models
- Generative Adversarial Networks for Graph Data Imputation from Signed Observations
- Generative Pre-Training for Speech with Autoregressive Predictive Coding
- Genetic Algorithm Optimized Support Vector Machine in NOMA-based Satellite Networks with Imperfect CSI
- Geometrically Constrained Independent Vector Analysis for Directional Speech Enhancement
- Geometry Constrained Progressive Learning for Lstm-Based Speech Enhancement
- Global Structure Graph Guided Fine-Grained Vehicle Recognition
- Global Traffic State Recovery VIA Local Observations with Generative Adversarial Networks
- Global and Local Discriminative Patches Exploiting for Action Recognition
- Gpu-Accelerated Viterbi Exact Lattice Decoder for Batched Online and Offline Speech Recognition
- Gradient Delay Analysis in Asynchronous Distributed Optimization
- Gradient-Based Algorithm with Spatial Regularization for Optimal Sensor Placement
- Graph Auto-Encoder for Graph Signal Denoising
- Graph Construction from Data by Non-Negative Kernel Regression
- Graph Convolutional Neural Networks to Classify Whole Slide Images
- Graph Metric Learning via Gershgorin Disc Alignment
- Graph Neural Net Using Analytical Graph Filters and Topology Optimization for Image Denoising
- Graph Vertex Sampling with Arbitrary Graph Signal Hilbert Spaces
- GraphTTS: Graph-to-Sequence Modelling in Neural Text-to-Speech
- Graphem: EM Algorithm for Blind Kalman Filtering Under Graphical Sparsity Constraints
- Graphical Evolutionary Game Theoretic Analysis of Super Users in Information Diffusion
- Gray-Scale Image Colorization Using Cycle-Consistent Generative Adversarial Networks with Residual Structure Enhancer
- Greedy Hybrid Rate Adaptation in Dynamic Wireless Communication Environment
- Greedy Sparse Array Design for Optimal Localization under Spatially Prioritized Source Distribution
- Group-Utility Metric for Efficient Sensor Selection and Removal in LCMV Beamformers
- Guided Learning for Weakly-Labeled Semi-Supervised Sound Event Detection
- Gyroscope Aided Video Stabilization Using Nonlinear Regression on Special Orthogonal Group
- H-Vectors: Utterance-Level Speaker Embedding Using a Hierarchical Attention Model
- HDMFH: Hypergraph Based Discrete Matrix Factorization Hashing for Multimodal Retrieval
- HGFM : A Hierarchical Grained and Feature Model for Acoustic Emotion Recognition
- HI-MIA: A Far-Field Text-Dependent Speaker Verification Database and the Baselines
- HKA: A Hierarchical Knowledge Attention Mechanism for Multi-Turn Dialogue System
- HPRNN: A Hierarchical Sequence Prediction Model for Long-Term Weather Radar Echo Extrapolation
- Hand-3d-Studio: A New Multi-View System for 3d Hand Reconstruction
- Harmonic/Percussive Sound Separation and Spectral Complexity Reduction of Music Signals for Cochlear Implant Listeners
- Harmonics Based Representation in Clarinet Tone Quality Evaluation
- Headless Horseman: Adversarial Attacks on Transfer Learning Models
- Hearing aid Research Data Set for Acoustic Environment Recognition
- Height and Weight Estimation from Unconstrained Images
- Heterogeneous Domain Generalization Via Domain Mixup
- Hidden Markov Models for Sepsis Detection in Preterm Infants
- Hierarchical Attention Transfer Networks for Depression Assessment from Speech
- Hierarchical Caching via Deep Reinforcement Learning
- Hierarchical Federated Learning ACROSS Heterogeneous Cellular Networks
- Hierarchical Sequence Representation with Graph Network
- High Dynamic Range Imaging Using Deep Image Priors
- High-Accuracy Classification of Attention Deficit Hyperactivity Disorder with L2, 1-Norm Linear Discriminant Analysis
- High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
- High-Dimensional Neural Feature Using Rectified Linear Unit And Random Matrix Instance
- High-Resolution Attention Network with Acoustic Segment Model for Acoustic Scene Classification
- Hijacking Tracker: A Powerful Adversarial Attack on Visual Tracking
- How Much Self-Attention Do We Need? Trading Attention for Feed-Forward Layers
- How confident are you? Exploring the role of fillers in the automatic prediction of a speaker's confidence
- Human-Machine Collaboration for Medical Image Segmentation
- Humangan: Generative Adversarial Network With Human-Based Discriminator And Its Evaluation In Speech Perception Modeling
- Humbug Zooniverse: A Crowd-Sourced Acoustic Mosquito Dataset
- Hybrid Active Contour Driven by Double-Weighted Signed Pressure Force for Image Segmentation
- Hybrid Autoregressive Transducer (HAT)
- Hybrid Deep-Semantic Matrix Factorization for Tag-Aware Personalized Recommendation
- Hybrid Neural-Parametric F0 Model for Singing Synthesis
- Hybrid Precoding for Secure Transmission in Reflect-Array-Assisted Massive MIMO Systems
- Hydranet: A Real-Time Waveform Separation Network
- I-Vector Transformation Using K-Nearest Neighbors for Speaker Verification
- IQ-STAN: Image Quality Guided Spatio-Temporal Attention Network for License Plate Recognition
- Identification of Essential Proteins Using A Novel Multi-Objective Optimization Method
- Identifying Truthful Language in Child Interviews
- Image Fusion using Joint Sparse Representations and Coupled Dictionary Learning
- Image Processing in DNA
- Image Recovery from Rotational And Translational Invariants
- Image Restoration Via Data-Dependent Proximal Averaged Optimization
- Image Segmentation Based Privacy-Preserving Human Action Recognition for Anomaly Detection
- Image Super-Resolution Using Residual Global Context Network
- Impact of a Shift-Invariant Harmonic Phase Model in Fully Parametric Harmonic Voice Representation and Time/Frequency Synthesis
- Improved End-To-End Spoken Utterance Classification with a Self-Attention Acoustic Classifier
- Improved Large-Margin Softmax Loss for Speaker Diarisation
- Improved Nearest Neighbor Density-Based Clustering Techniques with Application to Hyperspectral Images
- Improved Probability Modelling for Exception Handling in Lossless Screen Content Coding
- Improved Real-Time Visual Tracking via Adversarial Learning
- Improved Speaker Independent Dysarthria Intelligibility Classification Using Deepspeech Posteriors
- Improving Auditory Attention Decoding Performance of Linear and Non-Linear Methods using State-Space Model
- Improving Automated Segmentation of Radio Shows with Audio Embeddings
- Improving Convergent Cross Mapping for Causal Discovery with Gaussian Processes
- Improving Cross-Dataset Performance of Face Presentation Attack Detection Systems Using Face Recognition Datasets
- Improving Deep CNN Networks with Long Temporal Context for Text-Independent Speaker Verification
- Improving Deep Learning Classification of JPEG2000 Images Over Bandlimited Networks
- Improving Device Directedness Classification of Utterances With Semantic Lexical Features
- Improving Efficiency in Large-Scale Decentralized Distributed Training
- Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced Transformer
- Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density Network
- Improving Language Identification for Multilingual Speakers
- Improving Music Transcription by Pre-Stacking A U-Net
- Improving Noise Robust Automatic Speech Recognition with Single-Channel Time-Domain Enhancement Network
- Improving Proper Noun Recognition in End-To-End Asr by Customization of the Mwer Loss Criterion
- Improving Prosody with Linguistic and Bert Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTS
- Improving Reverberant Speech Training Using Diffuse Acoustic Simulation
- Improving Robustness of Deep Learning Based Monaural Speech Enhancement Against Processing Artifacts
- Improving Sample-Efficiency in Reinforcement Learning for Dialogue Systems by Using Trainable-Action-Mask
- Improving Sequence-To-Sequence Speech Recognition Training with On-The-Fly Data Augmentation
- Improving Singing Voice Separation with the Wave-U-Net Using Minimum Hyperspherical Energy
- Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam
- Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster Information
- Improving Speech Recognition Using Consistent Predictions on Synthesized Speech
- Improving Spoken Question Answering Using Contextualized Word Representation
- Improving Universal Sound Separation Using Sound Classification
- Improving Voice Separation by Incorporating End-To-End Speech Recognition
- Improving the Chronological Sorting of Images through Occlusion: A Study on the Notre-Dame Cathedral Fire
- Improving the Performance of Transformer Based Low Resource Speech Recognition for Indian Languages
- Improving the Scalability of Deep Reinforcement Learning-Based Routing with Control on Partial Nodes
- Impulse Response Data Augmentation and Deep Neural Networks for Blind Room Acoustic Parameter Estimation
- In-Domain and Out-of-Domain Data Augmentation to Improve Children's Speaker Verification System in Limited Data Scenario
- In-Network Caching for Hybrid Satellite-Terrestrial Networks Using Deep Reinforcement Learning
- Incorporating Written Domain Numeric Grammars into End-To-End Contextual Speech Recognition Systems for Improved Recognition of Numeric Sequences
- Incremental Semi-Supervised Learning for Multi-Genre Speech Recognition
- Independent Language Modeling Architecture for End-To-End ASR
- Individual Distance-Dependent HRTFS Modeling Through A Few Anthropometric Measurements
- Indoor Altitude Estimation of Unmanned Aerial Vehicles Using a Bank of Kalman Filters
- Indoor Heading Direction Estimation Using Rf Signals
- Indylstms: Independently Recurrent LSTMS
- Inferring Dynamic Group Leadership Using Sequential Bayesian Methods
- Information Flow Optimization in Inference Networks
- Information Maximized Variational Domain Adversarial Learning for Speaker Verification
- Information Theoretic Approach for Waveform Design in Coexisting MIMO Radar and MIMO Communications
- Instance-based Model Adaptation for Direct Speech Translation
- Instant Adaptive Learning: An Adaptive Filter Based Fast Learning Model Construction for Sensor Signal Time Series Classification on Edge Devices
- Integrating Discrete and Neural Features Via Mixed-Feature Trans-Dimensional Random Field Language Models
- Integration of Multi-Look Beamformers for Multi-Channel Keyword Spotting
- Intelligent Reflecting Surface for Massive Device Connectivity: Joint Activity Detection and Channel Estimation
- Intelligent Student Behavior Analysis System for Real Classrooms
- Intensity-Image Reconstruction for Event Cameras Using Convolutional Neural Network
- Interpolation and Range Extrapolation of Sound Source Directivity Based on a Spherical Wave Propagation Model
- Interpretability-Guided Convolutional Neural Networks for Seismic Fault Segmentation
- Interpretable Machine Learning In Sustainable Edge Computing: A Case Study of Short-Term Photovoltaic Power Output Prediction
- Interpretable Self-Attention Temporal Reasoning for Driving Behavior Understanding
- Interrupted and Cascaded Permutation Invariant Training for Speech Separation
- Intra Frame Rate Control for Versatile Video Coding with Quadratic Rate-Distortion Modelling
- Inverse Multiple Scattering with Phaseless Measurements
- Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement
- Investigating Generalization in Neural Networks Under Optimally Evolved Training Perturbations
- Investigation of Methods to Improve the Recognition Performance of Tamil-English Code-Switched Data in Transformer Framework
- Investigation of Specaugment for Deep Speaker Embedding Learning
- JPEG Steganography with Side Information from the Processing Pipeline
- Jhu-HLTCOE System for the Voxsrc Speaker Recognition Challenge
- Joint Beamforming and Reverberation Cancellation Using a Constrained Kalman Filter With Multichannel Linear Prediction
- Joint Blind Calibration and Time-Delay Estimation for Multiband Ranging
- Joint Coding and Modulation in the Ultra-Short Blocklength Regime for Bernoulli-Gaussian Impulsive Noise Channels Using Autoencoders
- Joint Contextual Modeling for ASR Correction and Language Understanding
- Joint Enhancement And Denoising of Low Light Images Via JND Transform
- Joint Estimation Of Acoustic Parameters From Single-Microphone Speech Observations
- Joint Frequency Domain Channel Estimation and Equalization Based on Expectation Propagation for Single Carrier Transmissions
- Joint Learning of Assignment and Representation for Biometric Group Membership
- Joint Learning of Cartesian under Sampling Andre Construction for Accelerated MRI
- Joint Multitarget Tracking and Dynamic Network Localization in the Underwater Domain
- Joint Optimization of Sampling Patterns and Deep Priors for Improved Parallel MRI
- Joint Phoneme Alignment and Text-Informed Speech Separation on Highly Corrupted Speech
- Joint Phoneme-Grapheme Model for End-To-End Speech Recognition
- Joint Resource Allocation and Routing for Service Function Chaining with In-Subnetwork Processing
- Joint Scheduling and Beamforming for Delay Sensitive Traffic with Priorities and Deadlines
- Joint Semi-Supervised Feature Auto-Weighting and Classification Model for EEG-Based Cross-Subject Sleep Quality Evaluation
- Joint Source-Channel Coding and Bayesian Message Passing Detection for Grant-Free Radio Access in IoT
- Joint Sparse Recovery Using Deep Unfolding With Application to Massive Random Access
- Joint Training of Deep Neural Networks for Multi-Channel Dereverberation and Speech Source Separation
- Jointly Optimal Dereverberation and Beamforming
- Just Noticeable Distortion Based Perceptually Lossless Intra Coding
- K-Autoencoders Deep Clustering
- K-Space Trajectory Design for Reduced MRI Scan Time
- KALM: Key Area Localization Mechanism for Abnormality Detection in Musculoskeletal Radiographs
- Kernel Computations from Large-Scale Random Features Obtained by Optical Processing Units
- Kernel Ridge Regression with Autocorrelation Prior: Optimal Model and Cross-Validation
- Key Action and Joint CTC-Attention based Sign Language Recognition
- Keyword Search for Sign Language
- Knowledge Distillation and Random Erasing Data Augmentation for Text-Dependent Speaker Verification
- Knowledge Enhanced Latent Relevance Mining for Question Answering
- Korean Singing Voice Synthesis Based on Auto-Regressive Boundary Equilibrium Gan
- L-Vector: Neural Label Embedding for Domain Adaptation
- L1-Norm Higher-Order Orthogonal Iterations for Robust Tensor Analysis
- LEt-SNE: A Hybrid Approach to Data Embedding and Visualization Of Hyperspectral Imagery
- LSTM-Based One-Pass Decoder for Low-Latency Streaming
- Label Propagation Adaptive Resonance Theory for Semi-Supervised Continuous Learning
- Label Reuse for Efficient Semi-Supervised Learning
- Lai-Net: Local-Ancestry Inference with Neural Networks
- Lance: efficient low-precision quantized winograd convolution for neural networks based on graphics processing units
- Language Independent Gender Identification from Raw Waveform Using Multi-Scale Convolutional Neural Networks
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.