ICASSP 2019 Accepted Papers
The full list of 1,730 papers accepted at ICASSP 2019 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- "Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification
- 1-D Convolutional Neural Networks for Signal Processing Applications
- 2.5D Multizone Reproduction with Active Control of Scattered Sound Fields
- 2D Beamforming on Sparse Arrays with Sparse Bayesian Learning
- 3D AOA Target Tracking with Two-step Instrumental-variable Kalman Filtering
- 3D Coprime Arrays in Sparse Sensing
- 3D Point Cloud Denoising via Deep Neural Network Based Local Surface Estimation
- 3D Reconstruction Using Single-photon Lidar Data Exploiting the Widths of the Returns
- 3D Visual Speech Animation Using 2D Videos
- A Bayesian Approach to Inter-task Fusion for Speaker Recognition
- A Bayesian Attention Neural Network Layer for Speaker Recognition
- A Bayesian Framework for Intent Prediction in Object Tracking
- A Calibrated Learning Approach to Distributed Power Allocation in Small Cell Networks
- A Cascade of CNN and LSTM Network with 3D Anchors for Mitotic Cell Detection in 4D Microscopic Image
- A Case of Distributed Optimization in Adversarial Environment
- A Characterization of Stochastic Mirror Descent Algorithms and Their Convergence Properties
- A Clustering Approach to Construct Multi-scale Overcomplete Dictionaries for ECG Modeling
- A Compact Framework for Voice Conversion Using Wavenet Conditioned on Phonetic Posteriorgrams
- A Comparison of Five Multiple Instance Learning Pooling Functions for Sound Event Detection with Weak Labeling
- A Comparison of Lattice-free Discriminative Training Criteria for Purely Sequence-trained Neural Network Acoustic Models
- A Convex Lifting Approach to Image Phase Unwrapping
- A Data-centric Approach to Unsupervised Texture Segmentation Using Principle Representative Patterns
- A Data-selective LS Solution to TDOA-based Source Localization
- A Deep Dictionary Model to Preserve and Disentangle Key Features in a Signal
- A Deep Dual-path Network for Improved Mammogram Image Processing
- A Deep Generative Model of Speech Complex Spectrograms
- A Deep Learning Based Binaural Speech Enhancement Approach with Spatial Cues Preservation
- A Deep Neural Network Based End to End Model for Joint Height and Age Estimation from Short Duration Speech
- A Deep Neural Network Based Maneuvering-target Tracking Algorithm
- A Deep-narma Filter for Unusual Behavior Detection from Visual, Thermal and Wireless Signals
- A Denoising Autoencoder for Speaker Recognition. Results on the MCE 2018 Challenge
- A Deterministic Annealing Approach to Switched Predictor Design for Adaptive Compression Systems
- A Differential-geometric Approach for Globally Solving a Non-convex, Discontinuous Depth Estimation Problem for Plenoptic Camera Images
- A Discrete Signal Processing Framework for Meet/join Lattices with Applications to Hypergraphs and Trees
- A Double K-best Viterbi-sphere Decoder for Trellis-coded Generalized Spatial Modulation with Multiple Code Rates
- A Double-cross-correlation Processor for Blind Sampling Rate Offset Estimation in Acoustic Sensor Networks
- A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from Speech
- A Fair and Scalable Power Control Scheme in Multi-cell Massive MIMO
- A Fast Method of Computing Persistent Homology of Time Series Data
- A Fast and Robust Paradigm for Fourier Compressed Sensing Based on Coded Sampling
- A Format-compliant Selective Secret 3D Object Sharing Scheme Based on Shamir's Scheme
- A Fully Convolutional Neural Network for Complex Spectrogram Processing in Speech Enhancement
- A Fuzzy-based Two-stage Biometric Sample Quality Evaluation System
- A Geometry-aware Framework for Compressing 3D Mesh Textures
- A Gridless CS Approach for Channel Estimation in Hybrid Massive MIMO Systems
- A Hierarchical Decoding Model for Spoken Language Understanding from Unaligned Data
- A Hierarchical Neural Summarization Framework for Spoken Documents
- A Highly Adaptive Acoustic Model for Accurate Multi-dialect Speech Recognition
- A History-based Stopping Criterion in Recursive Bayesian State Estimation
- A Hybrid Method for Blind Estimation of Frequency Dependent Reverberation Time Using Speech Signals
- A Joint Auditory Attention Decoding and Adaptive Binaural Beamforming Algorithm for Hearing Devices
- A LSTM and CNN Based Assemble Neural Network Framework for Arrhythmias Classification
- A Large Scale Analysis of Logistic Regression: Asymptotic Performance and New Insights
- A Learning Approach for Wavelet Design
- A Learning Approach to Wireless Information and Power Transfer Signal and System Design
- A Learning Based Depth Estimation Framework for 4D Densely and Sparsely Sampled Light Fields
- A Learning-based Framework for Line-spectra Super-resolution
- A Low-latency Sparse-Winograd Accelerator for Convolutional Neural Networks
- A Map Framework for Support Recovery of Sparse Signals Using Orthogonal Least Squares
- A Markerless Body Motion Capture System for Character Animation Based on Multi-view Cameras
- A Method Based on L-bfgs to Solve Constrained Complex-valued Ica
- A Modified Frank-wolfe Algorithm for Tensor Factorization with Unimodal Signals
- A Multi-radar Joint Beamforming Method
- A Multi-spike Approach for Robust Sound Recognition
- A Music Structure Informed Downbeat Tracking System Using Skip-chain Conditional Random Fields and Deep Learning
- A Needle in a Haystack? Harnessing Onomatopoeia and User-specific Stylometrics for Authorship Attribution of Micro-messages
- A Neighbor-aware Approach for Image-text Matching
- A Neural Network Based Ranking Framework to Improve ASR with NLU Related Knowledge Deployed
- A New Fusion Framework for Multimodal Medical Image Based on GRWT
- A New Quadratic Matrix Inequality Approach to Robust Adaptive Beamforming for General-rank Signal Model
- A New Separation Method for Galaxy Spectra Based on Data Fusion between Two Grism Orders in Slitless Spectroscopy
- A New Spatial Steganographic Scheme by Modeling Image Residuals with Multivariate Gaussian Model
- A Noise Robust Hearable Device with an Adaptive Noise Canceller and Its DSP Implementation
- A Non-convex Approach to Non-negative Super-resolution: Theory and Algorithm
- A Novel Approach for Feedforward Control of Noise in Ducts Using Simplified Multichannel Inverse Filters
- A Novel Approximate Lloyd-max Quantizer and Its Analysis
- A Novel Binaural Beamforming Scheme with Low Complexity Minimizing Binaural-cue Distortions
- A Novel Deep Hashing Method with Top Similarity for Image Retrieval
- A Novel Deterministic Sensing Matrix Based on Kasami Codes for Cluster Structured Sparse Signals
- A Novel Fractional Order Derivate Based Log-demons with Driving Force for High Accurate Image Registration
- A Novel Framework for Designing Directional Linear Transforms with Application to Video Compression
- A Novel Framework of Hand Localization and Hand Pose Estimation
- A Novel Progressive Gaussian Approximate Filter with Variable Step Size Based on a Variational Bayesian Approach
- A Novel Repetition Normalized Adversarial Reward for Headline Generation
- A Novel Resource-aware Tensor Decomposition Design Based on Reinforcement Learning
- A Novel Super-resolution Method Based on Patch Reconstruction with Simk Clustering and Nonlinear Mapping
- A Novel Three Dimensional Probability-based Classifier for Improving Motor Imagery-based BCI
- A Penalized Autoencoder Approach for Nonlinear Independent Component Analysis
- A Pipeline for Lung Tumor Detection and Segmentation from CT Scans Using Dilated Convolutional Neural Networks
- A Pitch-aware Approach to Single-channel Speech Separation
- A Privacy-preserving Diffusion Strategy over Multitask Networks
- A Projected Newton-type Algorithm for Nonnegative Matrix Factorization with Model Order Selection
- A Proper Version of Synthesis-based Sparse Audio Declipper
- A Quantitative Comparison of Epoch Extraction Algorithms for Telephone Speech
- A Quasi-Newton Algorithm on the Orthogonal Manifold for NMF with Transform Learning
- A Recurrent Graph Neural Network for Multi-relational Data
- A Recursive Bayesian Model for Extreme Values
- A Recursive Least-squares Algorithm Based on the Nearest Kronecker Product Decomposition
- A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification
- A Riemannian Distance Approach to MIMO Radar Signal Design
- A Robust Text-independent Speaker Verification Method Based on Speech Separation and Deep Speaker
- A Rotation Invariant HOG Descriptor for Tire Pattern Image Classification
- A Sequential Guiding Network with Attention for Image Captioning
- A Simple Bound on the BER of the Map Decoder for Massive MIMO Systems
- A Simple Way of Multimodal and Arbitrary Style Transfer
- A Sparse Encoding and Phaseless Decoding Approach for Fast Mmwave Beam Alignment
- A Sparsity Measure for Echo Density Growth in General Environments
- A Spectral Glottal Flow Model for Source-filter Separation of Speech
- A Spectral-change-aware Loss Function for DNN-based Speech Separation
- A Spectro-temporal Technique for Estimating Aperiodicity and Voiced/unvoiced Decision Boundaries of Speech Signals
- A Spelling Correction Model for End-to-end Speech Recognition
- A Spiking Neural Network Approach to Auditory Source Lateralisation
- A Spiking Neural Network with Local Learning Rules Derived from Nonnegative Similarity Matching
- A Streamlined Encoder/decoder Architecture for Melody Extraction
- A Study on How Pre-whitening Influences Fundamental Frequency Estimation
- A Study on Robustness of Articulatory Features for Automatic Speech Recognition of Neutral and Whispered Speech
- A Subband Energy Modification Method for Elevation Control in Median Plane
- A Subject-to-subject Transfer Learning Framework Based on Jensen-Shannon Divergence for Improving Brain-computer Interface
- A Tighter Bayesian CramÉR-rao Bound
- A Time-frequency Based Multivariate Phase-amplitude Coupling Measure
- A Topology-aware Coding Framework for Distributed Graph Processing
- A Training Method Using DNN-guided Layerwise Pretraining for Deep Gaussian Processes
- A Training Procedure for Quantum Random Vector Functional-link Networks
- A Two-class Hyper-spherical Autoencoder for Supervised Anomaly Detection
- A Two-stage Single-channel Speaker-dependent Speech Separation Approach for Chime-5 Challenge
- A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models
- A Unified Framework for Neural Speech Separation and Extraction
- A Unified Neural Architecture for Instrumental Audio Tasks
- A Variational Adaptive Population Importance Sampler
- A Variational Bayes Approach to Adaptive Channel-gain Cartography
- A Video Camera Model Identification System Using Deep Learning and Fusion
- A Vocoder Based Method for Singing Voice Extraction
- A Weight-shared Dual-branch Convolutional Neural Network for Unsupervised Dense Depth Prediction and Camera Motion Estimation
- A Windowed Digraph Fourier Transform
- ADMM for ND Line Spectral Estimation Using Grid-free Compressive Sensing from Multiple Measurements with Applications to DOA Estimation
- ADMM-based Beamforming Optimization for Physical Layer Security in a Full-duplex Relay System
- ADMM-based Bipartite Graph Approximation
- APE-GAN: Adversarial Perturbation Elimination with GAN
- ATTS2S-VC: Sequence-to-sequence Voice Conversion with Attention and Context Preservation Mechanisms
- AUC Optimization for Deep Learning Based Voice Activity Detection
- Accelerating Iterative Hard Thresholding for Low-rank Matrix Completion via Adaptive Restart
- Accuracy Evaluation Based on Simulation for Finite Precision Systems Using Inferential Statistics
- Accurate Reconstruction of Finite Rate of Innovation Signals on the Sphere
- Accurate Target Annotation in 3D from Multimodal Streams
- Accurate Vehicle Detection Using Multi-camera Data Fusion and Machine Learning
- Acoustic Equalization for Headphones Using a Fixed Feed-forward Filter
- Acoustic Event Detection from Weakly Labeled Data Using Auditory Salience
- Acoustic Impulse Responses for Wearable Audio Devices
- Acoustic Modeling for Distant Multi-talker Speech Recognition with Single- and Multi-channel Branches
- Acoustic Modeling for Overlapping Speech Recognition: Jhu Chime-5 Challenge System
- Acoustic Scene Generation with Conditional Samplernn
- Acoustic and Lexical Sentiment Analysis for Customer Service Calls
- Acoustically Grounded Word Embeddings for Improved Acoustics-to-word Speech Recognition
- Active Anomaly Detection with Switching Cost
- Active Learning for Efficient Audio Annotation and Classification with a Large Amount of Unlabeled Data
- Active Learning with Label Proportions
- Active Noise Control Using Finite Element-based Virtual Sensors
- Active Sampling for Approximately Bandlimited Graph Signals
- Ad-net: Attention Guided Network for Optical Flow Estimation Using Dilated Convolution
- AdaFlow: Domain-adaptive Density Estimator with Application to Anomaly Detection and Unpaired Cross-domain Translation
- Adaptation of Multiple Sound Source Localization Neural Networks with Weak Supervision and Domain-adversarial Training
- Adapting End-to-end Neural Speaker Verification to New Languages and Recording Conditions with Adversarial Training
- Adaptive Adjustment with Semantic Feature Space for Zero-shot Recognition
- Adaptive Blind Sparse Source Separation Based on Shear and Givens Rotations
- Adaptive Brightness Learning for Active Object Recognition
- Adaptive Dereverberation Using Multi-channel Linear Prediction with Deficient Length Filter
- Adaptive Filtering for Event Recognition from Noisy Signal: an Application to Earthquake Detection
- Adaptive Graph Formulation for 3D Shape Representation
- Adaptive Reduced-Dimensional Beamspace Beamformer Design by Analogue Beam Selection
- Adaptive Scenario Discovery for Crowd Counting
- Adaptive Sensing Matrix Design for Greedy Algorithms in Mmv Compressive Sensing
- Adaptive Subspace Detector in High Dimensional Space with Insufficient Training Data
- Adaptive Waveform Design for Automotive Joint Radar-communications System
- Adaptively Weighted Multi-task Learning Using Inverse Validation Loss
- Adversarial Examples for Improving End-to-end Attention-based Small-footprint Keyword Spotting
- Adversarial Inpainting of Medical Image Modalities
- Adversarial Learning of Label Dependency: A Novel Framework for Multi-class Classification
- Adversarial Learning-based Data Augmentation for Rotation-robust Human Tracking
- Adversarial Multi-label Prediction for Spoken and Visual Signal Tagging
- Adversarial Multi-task Deep Features and Unsupervised Back-end Adaptation for Language Recognition
- Adversarial Multi-user Bandits for Uncoordinated Spectrum Access
- Adversarial Speaker Adaptation
- Adversarial Speaker Verification
- Adversarial Training of End-to-end Speech Recognition Using a Criticizing Language Model
- Adversarial Variational Bayes Methods for Tweedie Compound Poisson Mixed Models
- Adversarial Watermarking to Attack Deep Neural Networks
- Adversarially Trained Autoencoders for Parallel-data-free Voice Conversion
- Adversarially-enriched Acoustic Code Vector Learned from Out-of-context Affective Corpus for Robust Emotion Recognition
- Aggregation Graph Neural Networks
- Aggregation and Embedding for Group Membership Verification
- Air-tissue Boundary Segmentation in Real Time Magnetic Resonance Imaging Video Using a Convolutional Encoder-decoder Network
- Algebraic Solution for Tdoa Localization in Modified Polar Representation
- Algebraically-initialized Expectation Maximization for Header-free Communication
- All for One: Frame-wise Rank Loss for Improving Video-based Person Re-identification
- All-neural Online Source Separation, Counting, and Diarization for Meeting Analysis
- Alternately Guided Depth Super-resolution Using Weighted Least Squares and Zero-order Reverse Filtering
- Alternating Direction Method of Multipliers for Semi-blind Astronomical Image Deconvolution
- Alternating Phase Projected Gradient Descent with Generative Priors for Solving Compressive Phase Retrieval
- An Admm Algorithm for Peak Transmission Energy Minimization in Symbol-level Precoding
- An Algorithm Unrolling Approach to Deep Image Deblurring
- An Alternating Projection Algorithm for Approximate Simultaneous Diagonalization
- An Analysis of Noise-aware Features in Combination with the Size and Diversity of Training Data for DNN-based Speech Enhancement
- An Antipodally Symmetric Optimal Dimensionality Sampling on the Sphere
- An Attention-aware Bidirectional Multi-residual Recurrent Neural Network (Abmrnn): A Study about Better Short-term Text Classification
- An Attention-based Neural Network Approach for Single Channel Speech Enhancement
- An Attribute-invariant Variational Learning for Emotion Recognition Using Physiology
- An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal Module
- An Autoencoder-based Approach for Recognizing Null Class in Activities of Daily Living In-the-wild via Wearable Motion Sensors
- An Educational Tool for Hearing Aid Compression Fitting via a Web-based Adjusted Smartphone App
- An Efficiency-improved Tdoa-based Direct Position Determination Method for Multiple Sources
- An Efficient Algorithm for Hyperspectral Image Clustering
- An Empirical Study of Speech Processing in the Brain by Analyzing the Temporal Syllable Structure in Speech-input Induced EEG
- An End-to-end Network to Synthesize Intonation Using a Generalized Command Response Model
- An Enhanced Hierarchical Extreme Learning Machine with Random Sparse Matrix Based Autoencoder
- An Ensemble of Deep Recurrent Neural Networks for P-wave Detection in Electrocardiogram
- An Event-contrastive Connectome Network for Automatic Assessment of Individual Face Processing and Memory Ability
- An Heterogeneous Compiler of Dataflow Programs for Zynq Platforms
- An Image Coding Approach Based on Mixture-of-experts Regression Using Epanechnikov Kernel
- An Improved Air Tissue Boundary Segmentation Technique for Real Time Magnetic Resonance Imaging Video Using Segnet
- An Improved Approach to Weakly Supervised Semantic Segmentation
- An Improved Low Rank Detector in the High Dimensional Regime
- An Improved Method for Parametric Spectral Estimation
- An Improved Uncertainty Propagation Method for Robust I-vector Based Speaker Recognition
- An Information-theoretic Approach for Automatically Determining the Number of State Groups When Aggregating Markov Chains
- An Integrated Framework for Field Recording, Localization, Classification and Annotation of Birdsongs Using Robot Audition Techniques - Harkbird 2.0
- An Interaction-aware Attention Network for Speech Emotion Recognition in Spoken Dialogs
- An Investigation of Multilingual ASR Using End-to-end LF-MMI
- An Iterative Time Domain Denoising Method
- An LS Localisation Method for Massive MIMO Transmission Systems
- An Online Multiple-speaker DOA Tracking Using the CappÉ-Moulines Recursive Expectation-maximization Algorithm
- An Unsupervised Learning Approach to Neural-net-supported Wpe Dereverberation
- Analysis Dictionary Learning: an Efficient and Discriminative Solution
- Analysis and Mitigation of Vocal Effort Variations in Speaker Recognition
- Analysis of Broadband GEVD-based Blind Source Separation
- Analysis of Coprime Arrays on Moving Platform
- Analysis of Information Diffusion with Irrational Users: A Graphical Evolutionary Game Approach
- Analysis of Multichannel Virtual Sensing Active Noise Control to Overcome Spatial Correlation and Causality Constraints
- Analysis of Reverberation via Teager Energy Features for Replay Spoof Speech Detection
- Analysis of Sparse-integer Measurement Matrices in Compressive Sensing
- Analytic Properties of Downsampling for Bandlimited Signals
- Analyzing Human Reaction Time for Talker Change Detection
- Analyzing Uncertainties in Speech Recognition Using Dropout
- Anarchic Urban Expansion Detection and Monitoring with Integration of Expert Knowledge
- Anomaly Detection Based on an Ensemble of Dereverberation and Anomalous Sound Extraction
- Anomaly Detection and Tracking Based on Mean-Reverting Processes with Unknown Parameters
- Anomaly Detection in Raw Audio Using Deep Autoregressive Networks
- Anomaly Detection in Single Subject vs Group Using Manifold Learning
- Anomaly Imaging for Structural Health Monitoring Exploiting Clustered Sparsity
- Artificial Bandwidth Extension Using a Conditional Generative Adversarial Network with Discriminative Training
- Asymmetric Cyclegan for Unpaired NIR-to-RGB Face Image Translation
- Asymptotic Kullback-Leibler Increment to Characterize Experiment-induced Stress
- Asymptotic Performance of Linear Discriminant Analysis with Random Projections
- Asymptotically Optimal Quickest Change Detection under a Nuisance Change
- Asymptotically Optimal Recovery of Gaussian Sources from Noisy Stationary Mixtures: the Least-noisy Maximally-separating Solution
- Asynchronous Neighbor Discovery Using Coupled Compressive Sensing
- Atom Selection in Continuous Dictionaries: Reconciling Polar and SVD Approximations
- Atrous Convolution for Binary Semantic Segmentation of Lung Nodule
- Attacks on Digital Watermarks for Deep Neural Networks
- Attention in Recurrent Neural Networks for Ransomware Detection
- Attention-augmented End-to-end Multi-task Learning for Emotion Prediction from Speech
- Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic Scenes
- Attention-based Graph Convolutional Network for Recommendation System
- Attention-based Transfer Learning for Brain-computer Interface
- Attention-based Wavenet Autoencoder for Universal Voice Conversion
- Attentive Adversarial Learning for Domain-invariant Training
- Attentive Filtering Networks for Audio Replay Attack Detection
- Attitude Recognition Using Multi-resolution Cochleagram Features
- Audio Caption: Listen and Tell
- Audio Coding Based on Spectral Recovery by Convolutional Neural Network
- Audio Feature Generation for Missing Modality Problem in Video Action Recognition
- Audio Replay Spoof Attack Detection Using Segment-based Hybrid Feature and DenseNet-LSTM Network
- Audio Texture Synthesis with Random Neural Networks: Improving Diversity and Quality
- Audio Watermarking over the Air with Modulated Self-correlation
- Audio-based Identification of Beehive States
- Audio-linguistic Embeddings for Spoken Sentences
- Audiovisual Speaker Conversion: Jointly and Simultaneously Transforming Facial Expression and Acoustic Characteristics
- Auditory Inspired Spatial Differentiation for Replay Spoofing Attack Detection
- Augmented Time-frequency Mask Estimation in Cluster-based Source Separation Algorithms
- Augmenting Dysphonia Voice Using Fourier-based Synchrosqueezing Transform for a CNN Classifier
- Auralization of Omnidirectional Room Impulse Responses Based on the Spatial Decomposition Method and Synthetic Spatial Data
- Autoencoding HRTFS for DNN Based HRTF Personalization Using Anthropometric Features
- Automatic Assessment of Spoken Language Proficiency of Non-native Children
- Automatic Diagnosis of Alzheimer's Disease Using Neural Network Language Models
- Automatic Grammar Augmentation for Robust Voice Command Recognition
- Automatic Grammatical Error Detection of Non-native Spoken Learner English
- Automatic Kernel Weighting for Multikernel Adaptive Filtering: Multiscale Aspects
- Automatic Lyrics-to-audio Alignment on Polyphonic Music Using Singing-adapted Acoustic Models
- Automatic Radar-based Gesture Detection and Classification via a Region-based Deep Convolutional Neural Network
- Automatic Segmentation of Common Carotid Artery in Longitudinal Mode Ultrasound Images Using Active Oblongs
- Automatic Segmentation of Nuclei in Histopathology Images Using Encoding-decoding Convolutional Neural Networks
- Automatic Segmentation of Optic Disc Using Affine Snakes in Gradient Vector Field
- Automatic Singing Evaluation without Reference Melody Using Bi-dense Neural Network
- Automatic Singing Transcription Based on Encoder-decoder Recurrent Neural Networks with a Weakly-supervised Attention Mechanism
- Automatic Transcription of Diatonic Harmonica Recordings
- Automating the Classification of Urban Issue Reports: an Optimal Stopping Approach
- Autonomous Detection and Disambiguation of Martian Ion Trails Using Geometric Signal Processing Techniques
- BLHUC: Bayesian Learning of Hidden Unit Contributions for Deep Neural Network Speaker Adaptation
- BLP - Boundary Likelihood Pinpointing Networks for Accurate Temporal Action Localization
- Background Adaptation for Improved Listening Experience in Broadcasting
- Baseline Wander Removal and Isoelectric Correction in Electrocardiograms Using Clustering
- Bayesian Drum Transcription Based on Nonnegative Matrix Factor Decomposition with a Deep Score Prior
- Bayesian Fusion of Asynchronous Inertial, Speed and Position Data for Object Tracking
- Bayesian Image Restoration under Poisson Noise and Log-concave Prior
- Bayesian Neural Networks for Sparse Coding
- Bayesian Non-parametric Multi-source Modelling Based Determined Blind Source Separation
- Bayesian and Gaussian Process Neural Networks for Large Vocabulary Continuous Speech Recognition
- Bayesian-optimized Bidirectional LSTM Regression Model for Non-intrusive Load Monitoring
- Beamformer Design under Time-correlated Interference and Online Implementation: Brain-activity Reconstruction from EEG
- Beamforming Design for Coexistence of Full-duplex Multi-cell MU-MIMO Cellular Network and MIMO Radar
- Beamforming Optimization for Intelligent Reflecting Surface with Discrete Phase Shifts
- Belief Condensation Filtering for RSSI-Based State Estimation in Indoor Localization
- Beyond Word-level to Sentence-level Sentiment Analysis for Financial Reports
- Bhattacharyya Distance-based Transfer Learning for a Hybrid EEG-FTCD Brain-computer Interface
- Bi-directional Lattice Recurrent Neural Networks for Confidence Estimation
- Bias Mitigation Post-processing for Individual and Group Fairness
- Bidirectional Quaternion Long Short-term Memory Recurrent Neural Networks for Speech Recognition
- Bilinear Dictionary Update via Linear Least Squares
- Bilinear Representation for Language-based Image Editing Using Conditional Generative Adversarial Networks
- Binary Classification Only from Unlabeled Data by Iterative Unlabeled-unlabeled Classification
- Binaural Beamforming Based on Automatic Interferer Selection
- Blind Calibration of Sparse Arrays for DOA Estimation with Analog and One-bit Measurements
- Blind Denoising of Mixed Gaussian-impulse Noise by Single CNN
- Blind Motion Deblurring via Inceptionresdensenet by Using GAN Model
- Blind Quality Assessment for 3D-synthesized Images by Measuring Geometric Distortions and Image Complexity
- Blind Quality Evaluator for Screen Content Images via Analysis of Structure
- Blind Room Volume Estimation from Single-channel Noisy Speech
- Blind Signal Processing for Time-varying Convolutive Mixing Systems Based on Sequence Estimation on Partly Smooth Manifolds
- Blind Super-resolution in Two-dimensional Parameter Space
- Block Alternating Optimization for Non-convex Min-max Problems: Algorithms and Applications in Signal Processing and Communications
- Block-randomized Stochastic Proximal Gradient for Constrained Low-rank Tensor Factorization
- Bluetooth Based Indoor Localization Using Triplet Embeddings
- Boolean CP Decomposition of Binary Tensors: Uniqueness and Algorithm
- Bootstrap-based Bias Reduction for the Estimation of the Self-similarity Exponents of Multivariate Time Series
- Bootstrapping Graph Convolutional Neural Networks for Autism Spectrum Disorder Classification
- Bootstrapping Single-channel Source Separation via Unsupervised Spatial Clustering on Stereo Mixtures
- Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker Verification
- Boundary Information Matters More: Accurate Temporal Action Detection with Temporal Boundary Network
- Brain Correlates of Task-Load and Dementia Elucidation with Tensor Machine Learning Using Oddball BCI Paradigm
- Breast Cancer Detection Based on Merging Four Modes MRI Using Convolutional Neural Networks
- Breast Cancer Image Classification on WSI with Spatial Correlations
- Burst-survive Temporal Matching Kernel with Fibonacci Periods
- Bytes Are All You Need: End-to-end Multilingual Speech Recognition and Synthesis with Bytes
- Byzantine-resilient Distributed Large-scale Matrix Completion
- CAN: Contextual Aggregating Network for Semantic Segmentation
- CMU Wilderness Multilingual Speech Dataset
- CNN Based Two-stage Multi-resolution End-to-end Model for Singing Melody Extraction
- CNN-RNN-CTC Based End-to-end Mispronunciation Detection and Diagnosis
- COLA: Communication-censored Linearized ADMM for Decentralized Consensus Optimization
- CONV-codes: Audio Hashing for Bird Species Classification
- COVER: A Cluster-based Variance Reduced Method for Online Learning
- CRF-based Single-stage Acoustic Modeling with CTC Topology
- Can Automatic Facial Expression Analysis Be Used for Treatment Outcome Estimation in Schizophrenia?
- Can We Predict Self-reported Customer Satisfaction from Interactions?
- Can We Use Speaker Recognition Technology to Attack Itself? Enhancing Mimicry Attacks Using Automatic Target Speaker Selection
- Canonical Correlation Based Feature Extraction with Application to Anomaly Detection in Electric Appliances
- Canonical Polyadic Decomposition of a Tensor That Has Missing Fibers: A Monomial Factorization Approach
- Capsule Networks for Brain Tumor Classification Based on MRI Images and Coarse Tumor Boundaries
- Capsule-forensics: Using Capsule Networks to Detect Forged Images and Videos
- Cascaded Point Network for 3D Hand Pose Estimation
- Casting to Corpus: Segmenting and Selecting Spontaneous Dialogue for Tts with a Cnn-lstm Speaker-dependent Breath Detector
- Causality and Robustness in the Remote Sensing of Acoustic Pressure, with Application to Local Active Sound Control
- Centroid-based Deep Metric Learning for Speaker Recognition
- Chained Compressed Sensing for Iot Node Security
- Channel Adversarial Training for Cross-channel Text-independent Speaker Recognition
- Channel Estimation Using 1-Bit Quantization and Oversampling for Large-scale Multiple-antenna Systems
- Channel Estimation and Low-complexity Beamforming Design for Passive Intelligent Surface Assisted MISO Wireless Energy Transfer
- Channel Impulsive Noise Mitigation for Linear Video Coding Schemes
- Channel Protection Using Random Modulation
- Class-conditional Embeddings for Music Source Separation
- Classification of Bioacoustic Signals with Tangent Singular Spectrum Analysis
- Classification of Chinese Dialect Regions from L2 English Speech
- Classification of Hyperspectral and Lidar with Deep Rotation Forest
- Cleaning Adversarial Perturbations via Residual Generative Network for Face Verification
- Clonability of Anti-counterfeiting Printable Graphical Codes: A Machine Learning Approach
- Clustering by Orthogonal Non-negative Matrix Factorization: A Sequential Non-convex Penalty Approach
- Co-attention Network and Low-rank Bilinear Pooling for Aspect Based Sentiment Analysis
- Co-clustering of High-order Data via Regularized Tucker Decompositions
- Co-prime Circular Microphone Arrays and Their Application to Direction of Arrival Estimation of Speech Sources
- Coding Tree Early Termination for Fast HEVC Transrating Based on Random Forests
- Cognitive-driven Binaural LCMV Beamformer Using EEG-based Auditory Attention Decoding
- Coherent Radar Imaging Using Unsynchronized Distributed Antennas
- Collaboration between Bordeaux-inp and Utp, from Research to Education, in the Field of Signal Processing
- Collaborative Sensor Caching via Sequential Compressed Sensing
- Combating Jamming in Wireless Networks: A Bayesian Game with Jammer's Channel Uncertainty
- Combining Linear Estimation with Scalar Widely Linear Estimation
- Combining Linear Spatial Filtering and Non-linear Parametric Processing for High-quality Spatial Sound Capturing
- Combining Matrix Design for 2D DoA Estimation with Compressive Antenna Arrays Using Stochastic Gradient Descent
- Combining Phone Posteriorgrams from Strong and Weak Recognizers for Automatic Speech Assessment of People with Aphasia
- Common Mode Patterns for Supervised Tensor Subspace Learning
- Common Randomized Shortest Paths (C-RSP): A Simple Yet Effective Framework for Multi-view Graph Embedding
- Communication Standards for Distributed Renewable Energy Sources Integration in Future Electricity Distribution Networks
- Communications under the Constraint of de-chirp Channel Distortion
- Community Detection in Sparse Realistic Graphs: Improving the Bethe Hessian
- Community Inference from Graph Signals with Hidden Nodes
- Compact Convolutional Recurrent Neural Networks via Binarization for Speech Emotion Recognition
- Compact Network for Speakerbeam Target Speaker Extraction
- Comparing CQT and Reassignment Based Chroma Features for Template-based Automatic Chord Recognition
- Comparison of Data Augmentation and Adaptation Strategies for Code-switched Automatic Speech Recognition
- Complementary Sequence Encoding for 1D and 2D Constant-modulus OFDM Transmission at Millimeter Wave Frequencies
- Complementary Siamese Networks for Robust Visual Tracking
- Completely Blind Image Quality Assessment Using Latent Quality Factor from Image Local Structure Representation
- Complex Spectral Mapping with a Convolutional Recurrent Network for Monaural Speech Enhancement
- Component Fusion: Learning Replaceable Language Model Component for End-to-end Speech Recognition System
- Composite Singer Arrays with Hole-free Coarrays and Enhanced Robustness
- Compound Variational Auto-encoder
- Compressed Randomized Utv Decompositions for Low-rank Matrix Approximations in Data Science
- Compressing Deep Neural Networks Using Toeplitz Matrix: Algorithm Design and Fpga Implementation
- Compression Improvement via Reference Organization for 2D-multiview Content
- Compressive Sensing with Applications to Millimeter-wave Architectures
- Compressive Single-pixel Fourier Transform Imaging Using Structured Illumination
- Computation Scheduling for Distributed Machine Learning with Straggling Workers
- Computational Analysis of Gaze Behavior in Autism During Interaction with Virtual Agents
- Computational Cognitive Assessment: Investigating the Use of an Intelligent Virtual Agent for the Detection of Early Signs of Dementia
- Computing the Largest Eigenvalue Distribution for Non-central Wishart Matrices
- Concrete: A Per-layer Configurable Framework for Evaluating DNN with Approximate Operators
- Condition-transforming Variational Autoencoder for Conversation Response Generation
- Conditional Teacher-student Learning
- Connectionist Temporal Localization for Sound Event Detection with Sequential Labeling
- Consensus-based Distributed Total Least-squares Estimation Using Parametric Semidefinite Programming
- Consistency Constrained Reconstruction of Depth Maps from Epipolar Plane Images
- Constructing and Compressing Frames in Blockchain-based Verifiable Multi-party Computation
- Construction of Overcomplete Multiscale Dictionary of Slepian Functions on the Sphere
- Content Adaptive Wavelet Lifting for Scalable Lossless Video Coding
- Content Placement Learning for Success Probability Maximization in Wireless Edge Caching Networks
- Context Modelling Using Hierarchical Attention Networks for Sentiment and Self-assessed Emotion Detection in Spoken Narratives
- Context-aware Deep Learning for Multi-modal Depression Detection
- Context-aware Neural-based Dialog Act Classification on Automatically Generated Transcriptions
- Contextual Out-of-domain Utterance Handling with Counterfeit Data Augmentation
- Contextual Speech Recognition with Difficult Negative Training Examples
- Continual Learning for Anomaly Detection with Variational Autoencoder
- Control Aware Communication Design for Time Sensitive Wireless Systems
- Convergence Bounds for Compressed Gradient Methods with Memory Based Error Compensation
- Convex Combination of Constraint Vectors for Set-membership Affine Projection Algorithms
- Convex Energy Optimization of Streaming Applications for MPSoCs
- Convex Relaxations of Convolutional Neural Nets
- Convexity-edge-preserving Signal Recovery with Linearly Involved Generalized Minimax Concave Penalty Function
- Convolutional Dictionary Regularizers for Tomographic Inversion
- Convolutional Neural Network on Embedded Platform for People Presence Detection in Low Resolution Thermal Images
- Convolutional Neural Networks for Video Intra Prediction Using Cross-component Adaptation
- Convolutional-sparse-coded Dynamic Mode Decomposition and Its Application to River State Estimation
- Cooperative Deep Reinforcement Learning for Multiple-group NB-IoT Networks Optimization
- Cooperative Detection via Direct Localization in Mobile Multi-agent Networks
- Coordinated Pilot Design for Massive MIMO
- Coprime Array Design with Minimum Lag Redundancy
- Coupled Ista Network for Multi-modal Image Super-resolution
- Coupled Tensor Low-rank Multilinear Approximation for Hyperspectral Super-resolution
- CramÉr-rao Bound for DOA Estimators under the Partial Relaxation Framework
- Crime Event Embedding with Unsupervised Feature Selection
- Cross Evaluation of Speech Enhancement Methods under Different Noise Conditions
- Cross Modal Audio Search and Retrieval with Joint Embeddings Based on Text and Audio
- Cross-culture Multimodal Emotion Recognition with Adversarial Learning
- Cross-gender Voice Conversion with Constant F0-Ratio and Average Background Conversion Model
- Cross-language Speech Dependent Lip-synchronization
- Cross-lingual Speech-based Tobi Label Generation Using Bidirectional Lstm
- Cross-lingual Text-independent Speaker Verification Using Unsupervised Adversarial Discriminative Domain Adaptation
- Cross-lingual Transfer Learning for Spoken Language Understanding
- Cross-lingual Voice Conversion with Bilingual Phonetic Posteriorgram and Average Modeling
- Cross-view Identical Part Area Alignment for Person Re-identification
- Crownn: Human-in-the-loop Network with Crowd-generated Inputs
- Curiosity-driven Reinforcement Learning for Dialogue Management
- Cutensor-tubal: Optimized GPU Library for Low-tubal-rank Tensors
- Cycle-GANs for Domain Adaptation of Acoustic Features for Speaker Recognition
- Cycle-consistency Training for End-to-end Speech Recognition
- Cycle-consistent Adversarial Networks for Non-parallel Vocal Effort Based Speaking Style Conversion
- Cyclegan Bandwidth Extension Acoustic Modeling for Automatic Speech Recognition
- Cyclegan-VC2: Improved Cyclegan-based Non-parallel Voice Conversion
- D2PGGAN: Two Discriminators Used in Progressive Growing of GANS
- DADA: Deep Adversarial Data Augmentation for Extremely Low Data Regime Classification
- DNN Training Based on Classic Gain Function for Single-channel Speech Enhancement and Recognition
- DNN-based Emotion Recognition Based on Bottleneck Acoustic Features and Lexical Features
- DNN-based Speaker-adaptive Postfiltering with Limited Adaptation Data for Statistical Speech Synthesis Systems
- DOD-CNN: Doubly-injecting Object Information for Event Recognition
- DSNET: Accelerate Indoor Scene Semantic Segmentation
- DSSLIC: Deep Semantic Segmentation-based Layered Image Compression
- Data Augmentation Strategies for Neural Network F0 Estimation
- Data Augmentation for Low Resource Sentiment Analysis Using Generative Adversarial Networks
- Data Driven Vessel Trajectory Forecasting Using Stochastic Generative Models
- Data Efficient Voice Cloning for Neural Singing Synthesis
- Data Poisoning Attacks against MRMR
- Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement
- Data-driven Simulation Using the Nuclear Norm Heuristic
- Data-selective LMS-Newton and LMS-Quasi-Newton Algorithms
- Dead Time Compensation for High-flux Depth Imaging
- Decision Feedback Semi-blind Estimation Algorithm for Specular OFDM Channels
- Decoding Homomorphically Encrypted Flac Audio without Decryption
- Decoupling Category-wise Independence and Relevance with Self-attention for Multi-label Image Classification
- Deep CNN for Wideband Mmwave Massive Mimo Channel Estimation Using Frequency Correlation
- Deep Complex-valued Neural Beamformers
- Deep Convolutional Feature Histograms for Visual Object Tracking
- Deep Convolutional Robust PCA with Application to Ultrasound Imaging
- Deep Counting Model Extensions with Segmentation for Person Detection
- Deep Embeddings for Rare Audio Event Detection with Imbalanced Data
- Deep Graph Regularized Learning for Binary Classification
- Deep Griffin-Lim Iteration
- Deep Hidden Analysis: A Statistical Framework to Prune Feature Maps
- Deep Hybrid Networks Based Response Selection for Multi-turn Dialogue Systems
- Deep Joint Source-channel Coding for Wireless Image Transmission
- Deep Latent Factor Model for Predicting Drug Target Interactions
- Deep Learning Based Online Power Control for Large Energy Harvesting Networks
- Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective
- Deep Learning Based on Orthogonal Approximate Message Passing for CP-Free OFDM
- Deep Learning Features for Robust Detection of Acoustic Events in Sleep-disordered Breathing
- Deep Learning Propagation Models over Irregular Terrain
- Deep Learning for Classroom Activity Detection from Audio
- Deep Learning for Fast Adaptive Beamforming
- Deep Learning for Minimal-context Block Tracking through Side-channel Analysis
- Deep Learning for Super-resolution Vascular Ultrasound Imaging
- Deep Learning for Tube Amplifier Emulation
- Deep Learning for Vertex Reconstruction of Neutrino-nucleus Interaction Events with Combined Energy and Time Data
- Deep Learning the EEG Manifold for Phonological Categorization from Active Thoughts
- Deep Neural Networks for Appliance Transient Classification
- Deep Neural Networks for Low-resolution Photon-limited Imaging
- Deep Polyphonic ADSR Piano Note Transcription
- Deep Ptych: Subsampled Fourier Ptychography Using Generative Priors
- Deep Quantization for MIMO Channel Estimation
- Deep Recurrent Neural Networks with Layer-wise Multi-head Attentions for Punctuation Restoration
- Deep Reinforcement Learning for Financial Trading Using Price Trailing
- Deep Reinforcement Learning-based Rate Adaptation for Adaptive 360-Degree Video Streaming
- Deep Signal Recovery with One-bit Quantization
- Deep Speaker Embedding Learning with Multi-level Pooling for Text-independent Speaker Verification
- Deep Speaker Representation Using Orthogonal Decomposition and Recombination for Speaker Verification
- Deep Spline Networks with Control of Lipschitz Regularity
- Deep Synthesizer Parameter Estimation
- Deep Temporal Logistic Bag-of-features for Forecasting High Frequency Limit Order Book Time Series
- Deep Temporal Pyramid Design for Action Recognition
- Deep Variational Filter Learning Models for Speech Recognition
- Deepwalk-assisted Graph PCA (DGPCA) for Language Networks
- DenXFPN: Pulmonary Pathologies Detection Based on Dense Feature Pyramid Networks
- Denoising Convolutional Autoencoder Based B-mode Ultrasound Tongue Image Feature Extraction
- Denoising Gravitational Waves with Enhanced Deep Recurrent Denoising Auto-encoders
- Denoising of 3D Point Clouds Constructed from Light Fields
- Dense Multimodal Fusion for Hierarchically Joint Representation
- Densely Connected Network with Time-frequency Dilated Convolution for Speech Enhancement
- Deriving Spectro-temporal Properties of Hearing from Speech Data
- Design of Optimal Linear Differential Microphone Arrays Based Array Geometry Optimization
- Designing (In)finite-alphabet Sequences via Shaping the Radar Ambiguity Function
- Designing Sar Images Change-point Estimation Strategies Using an Mse Lower Bound
- Designing an Effective Metric Learning Pipeline for Speaker Diarization
- Detectability of Denial-of-service Attacks on Communication Systems
- Detecting Attention Shift from Neural Response Based on Beat-frequency-modulated Musical Excerpts
- Detecting Cyber Attacks Using Anomaly Detection with Explanations and Expert Feedback
- Detecting Gas Vapor Leaks through Uncalibrated Sensor Based CPS
- Detecting and Classifying Rail Corrugation Based on Axle Bearing Vibration
- Detection and Amplification of Molecular Signals Using Cooperating Nano-devices
- Detection and Estimation of Delays in Bivariate Self-similarity: Bootstrapped Complex Wavelet Coherence
- Detection of Grid Voltage Anomalies via Broadband Subspace Decompositon
- Detection of Non Random Phase Signal in Additive Noise with Surrogate Analysis
- Detection of Pilot-hopping Sequences for Grant-free Random Access in Massive Mimo Systems
- Detection of Real-world Fights in Surveillance Videos
- Detection of Row-sparse Matrices with Row-structure Constraints
- Detection of Sleep Apnea/hypopnea Events Using Synchrosqueezed Wavelet Transform
- Detection of Voice Transformation Spoofing Based on Dense Convolutional Network
- Development and Evaluation of Japanese Text-to-speech Middleware for 32-Bit Microcontrollers
- Dialogue State Tracking with Convolutional Semantic Taggers
- Differentiable Consistency Constraints for Improved Deep Speech Enhancement
- Differentially Private Compressive K-means
- Differentially Private Greedy Decision Forest
- Diffusion Learning in Non-convex Environments
- Digitally Annealed Solution for the Vertex Cover Problem with Application in Cyber Security
- Dilated Residual Network with Multi-head Self-attention for Speech Emotion Recognition
- Dimensional Analysis of Laughter in Female Conversational Speech
- Direct Estimation of Weights and Efficient Training of Deep Neural Networks without SGD
- Direct-to-reverberant Energy Ratio Estimation Based on Interaural Coherence and a Joint ITD/ILD Model
- Direction Preserving Wiener Matrix Filtering for Ambisonic Input-output Systems
- Directional Interference Suppression Using a Spatial Relative Transfer Function Feature
- Discovering Optimal Variable-length Time Series Motifs in Large-scale Wearable Recordings of Human Bio-behavioral Signals
- Discrete Constant Envelope Transceiver Design for Multiuser Massive MIMO Downlink
- Discriminate Natural versus Loudspeaker Emitted Speech
- Discriminative Feature Selection Guided Deep Canonical Correlation Analysis
- Discriminative Features Reconstruction Network for Semantic Segmentation
- Discriminative Saliency-pose-attention Covariance for Action Recognition
- Discriminative Video Representation with Temporal Order for Micro-expression Recognition
- Discriminatively Re-trained I-vector Extractor for Speaker Recognition
- Disentangling Correlated Speaker and Noise for Speech Synthesis via Data Augmentation and Adversarial Factorization
- Disjunct Matrices for Compressed Sensing
- Distance-dependent Modeling of Head-related Transfer Functions
- Distributed Bayesian Estimation with Low-rank Data: Application to Solar Array Processing
- Distributed Convex Optimization with Limited Communications
- Distributed Deep Learning Strategies for Automatic Speech Recognition
- Distributed Differentially-private Canonical Correlation Analysis
- Distributed Gradient Descent with Coded Partial Gradient Computations
- Distributed Inference over Networks under Subspace Constraints
- Distributed Joint Transmitter Design and Selection Using Augmented Admm
- Distributed Network Caching via Dynamic Programming
- Distributed Noncoherent Transmit Beamforming for Dense Small Cell Networks
- Distributed Power Allocation for Spectral Coexisting Multistatic Radar and Communication Systems Based on Stackelberg Game
- Distributed Quickest Detection of Significant Events in Networks
- Distributed Signal Recovery Based on In-network Subspace Projections
- Distributed Tracking of Maneuvering Target: A Finite-time Algorithm
- Distributed UAV Placement Optimization for Cooperative Line-of-sight MIMO Communications
- Distribution Preserving Network Embedding
- Divergence Based Weighting for Information Channels in Deep Convolutional Neural Networks for Bird Audio Detection
- Diving Deep onto Discriminative Ensemble of Histological Hashing & Class-Specific Manifold Learning for Multi-class Breast Carcinoma Taxonomy
- Dnn-based Spectral Enhancement for Neural Waveform Generators with Low-bit Quantization
- Domain Adaptation Using Riemannian Geometry of Spd Matrices
- Domain Adversarial Training for Improving Keyword Spotting Performance of ESL Speech
- Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain
- Domain Mismatch Robust Acoustic Scene Classification Using Channel Information Conversion
- Drawing Order Recovery for Handwriting Chinese Characters
- Dual Domain Learning of Optimal Resource Allocations in Wireless Systems
- Dual-modality Seq2Seq Network for Audio-visual Event Localization
- Dual-stream CNN for Structured Time Series Classification
- Dynamic Joint PHY-MAC Waveform Design for IoT Connectivity
- Dynamic Joint Resource Allocation and User Assignment in Multi-access Edge Computing
- Dynamic Metasurfaces for Massive MIMO Networks
- Dynamic Point Cloud Geometry Compression via Patch-wise Polynomial Fitting
- Dynamic Resource Optimization for Decentralized Signal Estimation in Energy Harvesting Wireless Sensor Networks
- Dynamic Selection of Classifiers for Fusing Imbalanced Heterogeneous Data
- Dynamic Temporal Alignment of Speech to Lips
- Dynamic Texture Recognition Using 3D Random Features
- Dynamic Weight Alignment for Temporal Convolutional Neural Networks
- Dynamical Component Analysis (DYCA) and Its Application on Epileptic EEG
- Dynamically Context-sensitive Time-decay Attention for Dialogue Modeling
- E-CNN: Accurate Spherical Camera Rotation Estimation via Uniformization of Distorted Optical Flow Fields
- EDUQA: Educational Domain Question Answering System Using Conceptual Network Mapping
- EE-AE: An Exclusivity Enhanced Unsupervised Feature Learning Approach
- EMG Wrist-hand Motion Recognition System for Real-time Embedded Platform
- ENF Signal Extraction for Rolling-shutter Videos Using Periodic Zero-padding
- ENLLVM: Ensemble Based Nonlinear Bayesian Filtering Using Linear Latent Variable Models
- Early Wildfire Smoke Detection Based on Motion-based Geometric Image Transformation and Deep Convolutional Generative Adversarial Networks
- Eco-panda: A Computationally Economic, Geometrically Converging Dual Optimization Method on Time-varying Undirected Graphs
- Edge Detection Evaluation: A New Normalized Figure of Merit
- Effect of Data Reduction on Sequence-to-sequence Neural TTS
- Effective and Stable Neuron Model Optimization Based on Aggregated CMA-ES
- Effects of Lombard Reflex on the Performance of Deep-learning-based Audio-visual Speech Enhancement Systems
- Efficient Arabic Emotion Recognition Using Deep Neural Networks
- Efficient Belief Propagation Detection Based on Channel Hardening for Massive MIMO
- Efficient Constrained Signal Reconstruction by Randomized Epigraphical Projection
- Efficient Convolutional Neural Network Weight Compression for Space Data Classification on Multi-fpga Platforms
- Efficient Indoor Localization via Reinforcement Learning
- Efficient Keyword Spotting Using Dilated Convolutions and Gating
- Efficient Lossless Compression Scheme for Multi-channel ECG Signal
- Efficient Multi-agent Cooperative Navigation in Unknown Environments with Interlaced Deep Reinforcement Learning
- Efficient Nonlinear Acoustic Echo Cancellation by Dual-stage Multi-channel Kalman Filtering
- Efficient Randomized Defense against Adversarial Attacks in Deep Convolutional Neural Networks
- Efficient Sampling through Variable Splitting-inspired Bayesian Hierarchical Models
- Efficient Signal Reconstruction via Distributed Least Square Optimization on a Systolic FPGA Architecture
- Efficient Stochastic Subgradient Descent Algorithms for High-dimensional Semi-sparse Graphical Model Selection
- Ego-motion Estimation for Low-cost Freehand Ultrasound Scanner
- Embedding Physical Augmentation and Wavelet Scattering Transform to Generative Adversarial Networks for Audio Classification with Limited Training Resources
- Empirical Evaluation and Combination of Punctuation Prediction Models Applied to Broadcast News
- Empirical Wavelet Transform Based Lung Sound Removal from Phonocardiogram Signal for Heart Sound Segmentation
- Encrypted Speech Recognition Using Deep Polynomial Networks
- End-to-end Anchored Speech Recognition
- End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features
- End-to-end Binaural Sound Localisation from the Raw Waveform
- End-to-end Change Detection Using a Symmetric Fully Convolutional Network for Landslide Mapping
- End-to-end Code-switched TTS with Mix of Monolingual Recordings
- End-to-end Contextual Speech Recognition Using Class Language Models and a Token Passing Decoder
- End-to-end Dysarthric Speech Recognition Using Multiple Databases
- End-to-end Feedback Loss in Speech Chain Framework via Straight-through Estimator
- End-to-end Language Recognition Using Attention Based Hierarchical Gated Recurrent Unit Models
- End-to-end Lyrics Alignment for Polyphonic Music Using an Audio-to-character Recognition Model
- End-to-end Monaural Multi-speaker ASR System without Pretraining
- End-to-end Sound Source Separation Conditioned on Instrument Labels
- End-to-end Speech Recognition Using a High Rank LSTM-CTC Based Model
- End-to-end Speech Recognition with Adaptive Computation Steps
- End-to-end Streaming Keyword Spotting
- Energy Blowup of Sampling-based Approximation Methods
- Energy Minimization of Multi-user Latency-constrained Binary Computation Offloading
- Energy-efficient Design for Underlay Cognitive Radio Using Improper Signaling
- English Broadcast News Speech Recognition by Humans and Machines
- Enhanced Hierarchical Music Structure Annotations via Feature Level Similarity Fusion
- Enhanced Recurrent Neural Network for Combining Static and Dynamic Features for Credit Card Default Prediction
- Enhanced Streaming Based Subspace Clustering Applied to Acoustic Scene Data Clustering
- Enhanced Virtual Singers Generation by Incorporating Singing Dynamics to Personalized Text-to-speech-to-singing
- Enhancing Acoustic Sensory Responsiveness by Exploiting Bio-inspired Feedback Computation
- Enhancing Beamformed Fingerprint Outdoor Positioning with Hierarchical Convolutional Neural Networks
- Enhancing HEVC Spatial Prediction by Context-based Learning
- Enhancing Hybrid Self-attention Structure with Relative-position-aware Bias for Speech Synthesis
- Enhancing Music Features by Knowledge Transfer from User-item Log Data
- Enhancing Sound Texture in CNN-based Acoustic Scene Classification
- Ensemble Additive Margin Softmax for Speaker Verification
- Entropy-regularized Optimal Transport Generative Models
- Epoch Extraction from Speech Signals Using Temporal and Spectral Cues by Exploiting Harmonic Structure of Impulse-like Excitations
- Equation-Error Model Based Active Noise Cancellation Systems
- Error Bounds for Spectral Clustering over Samples from Spherical Gaussian Mixture Models
- Estimating Viewed Image Categories from Human Brain Activity via Semi-supervised Fuzzy Discriminative Canonical Correlation Analysis
- Estimating the Number of Correlated Components Based on Random Projections
- Estimation of Gaze Region Using Two Dimensional Probabilistic Maps Constructed Using Convolutional Neural Networks
- Estimation of Guitar String, Fret and Plucking Position Using Parametric Pitch Estimation
- Estimation of Network Processes via Blind Graph Multi-filter Identification
- Estimation of Sampling Frequency Mismatch between Distributed Asynchronous Microphones under Existence of Source Movements with Stationary Time Periods Detection
- Estimation of Widely Factorizable Hypercomplex Signals with Uncertain Observations
- Evaluating Salience Representations for Cross-modal Retrieval of Western Classical Music Recordings
- Evaluation Measures for Depression Prediction and Affective Computing
- Evaluation of Non-intrusive Load Monitoring Algorithms for Appliance-level Anomaly Detection
- Evaluation of Source-wise Missing Data Techniques for the Prediction of Parkinson's Disease Using Smartphones
- Event-driven Pipeline for Low-latency Low-compute Keyword Spotting and Speaker Verification System
- Every Rating Matters: Joint Learning of Subjective Labels and Individual Annotators for Speech Emotion Classification
- Evolutionary Subspace Clustering: Discovering Structure in Self-expressive Time-series Data
- Exact Discrete-time Realizations of the Gammatone Filter
- Exact Distribution and High-dimensional Asymptotics for Improperness Test of Complex Signals
- Exact Recovery by Semidefinite Programming in the Binary Stochastic Block Model with Partially Revealed Side Information
- Exact Sparse Signal Recovery via Orthogonal Matching Pursuit with Prior Information
- Exactly Decoding a Vector through Relu Activation
- Excess Cyclic Prefix Window Optimization for High Doppler MIMO OFDM Capacity
- Expectation-propagation Algorithms for Linear Regression with Poisson Noise: Application to Photon-limited Spectral Unmixing
- Exploiting Acoustic and Lexical Properties of Phonemes to Recognize Valence from Speech
- Exploiting Uncertainty of Deep Neural Networks for Improving Segmentation Accuracy in MRI Images
- Exploring Attention Mechanism for Acoustic-based Classification of Speech Utterances into System-directed and Non-system-directed
- Exploring Complex Time-series Representations for Riemannian Machine Learning of Radar Data
- Exploring Deep Complex Networks for Complex Spectrogram Enhancement
- Exploring Retraining-free Speech Recognition for Intra-sentential Code-switching
- Exponential Collapse of Social Beliefs over Weakly-connected Heterogeneous Networks
- Exposing Deep Fakes Using Inconsistent Head Poses
- Expression-identity Fusion Network for Facial Expression Recognition
- Extraction of Independent Vector Component from Underdetermined Mixtures through Block-wise Determined Modeling
- F0 Contour Estimation Using Phonetic Feature in Electrolaryngeal Speech Enhancement
- FRI Modelling of Fourier Descriptors
- FRI Sensing: Sampling Images along Unknown Curves
- Face Landmark-based Speaker-independent Audio-visual Speech Enhancement in Multi-talker Environments
- Facial Micro-expression Spotting and Recognition Using Time Contrasted Feature with Visual Memory
- Factors Affecting Enf Based Time-of-recording Estimation for Video
- Fast Coding Unit Decision for Intra Screen Content Coding Based on Ensemble Learning
- Fast Compressive Sensing Recovery Using Generative Models with Structured Latent Variables
- Fast Edge Preserving 2D Smoothing Filter Using Indicator Function
- Fast First-order Methods for the Massive Robust Multicast Beamforming Problem with Interference Temperature Constraints
- Fast Implementation of Double-coupled Nonnegative Canonical Polyadic Decomposition
- Fast Inter-prediction Based on Decision Trees for AV1 Encoding
- Fast MVAE: Joint Separation and Classification of Mixed Sources Based on Multichannel Variational Autoencoder with Auxiliary Classifier
- Fast Optimization of Boolean Quadratic Functions via Iterative Submodular Approximation and Max-flow
- Fast Sampling of Graph Signals with Noise via Neumann Series Conversion
- Fast and Communication-efficient Distributed Pca
- Fast and Efficient Distributed Matrix-vector Multiplication Using Rateless Fountain Codes
- Fast and Global Optimal Nonconvex Matrix Factorization via Perturbed Alternating Proximal Point
- FastMNMF: Joint Diagonalization Based Accelerated Algorithms for Multichannel Nonnegative Matrix Factorization
- Feature Selection for Mutlti-labeled Variables via Dependency Maximization
- Federated Learning for Keyword Spotting
- Feedback-controlled Channel Estimation with Low-resolution Adcs in Multiuser Mimo Systems
- Feedforward Spatial Active Noise Control Based on Kernel Interpolation of Sound Field
- Fine-tuning Approach to NIR Face Recognition
- Fire Detection from Images Using Faster R-CNN and Multidimensional Texture Analysis
- Fire Detection in H.264 Compressed Video
- FirmNet: A Sparsity Amplified Deep Network for Solving Linear Inverse Problems
- Flexible Design of Finite Blocklength Wiretap Codes by Autoencoders
- Flexible Non-negative Matrix Factorization with Adaptively Learned Graph Regularization
- Flexible Spectral Precoding for Sidelobe Suppression in OFDM Systems
- Forked Recurrent Neural Network for Hand Gesture Classification Using Inertial Measurement Data
- Formant-gaps Features for Speaker Verification Using Whispered Speech
- Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition
- Frequency Separation Method Based on Sparse Coding
- Frequency-domain Adaptive Filtering: from Real to Hypercomplex Signal Processing
- Frequency-domain Based Waveform Design for Binary Extended-target Classification
- Frequency-domain Decoupling for MIMO-GFDM Spatial Multiplexing
- Frequency-selective Hybrid Precoding and Combining for Mmwave Mimo Systems with Per-antenna Power Constraints
- From Gene Expression to Drug Response: A Collaborative Filtering Approach
- From Local to Global Subspace Clustering for Image Data
- From TV-L1 to Gated Recurrent Nets
- Fully Data-driven Convolutional Filters with Deep Learning Models for Epileptic Spike Detection
- Fully Supervised Speaker Diarization
- Function Designable Beamformer Based on Probabilistic Assumptions on Filter and Its Auxiliary Variables
- Fundamental Frequency Contour Classification: A Comparison between Hand-crafted and CNN-based Features
- Furcax: End-to-end Monaural Speech Separation Based on Deep Gated (De)convolutional Neural Networks with Adversarial Example Training
- FuseLoc: A CCA Based Information Fusion for Indoor Localization Using CSI Phase and Amplitude of Wifi Signals
- Fusing Eigenvalues
- Fuzzy Personalized Scoring Model for Recommendation System
- GPU-based Implementation of Belief Propagation Decoding for Polar Codes
- Gaussian Process Lstm Recurrent Neural Network Language Models for Speech Recognition
- Gaussian-constrained Training for Speaker Verification
- Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and Challenge
- Generalized Boundary Detection Using Compression-based Analytics
- Generalized Dantzig Selector for Low-tubal-rank Tensor Recovery
- Generalized Distributed Dual Coordinate Ascent in a Tree Network for Machine Learning
- Generalized Interval Valued Nonnegative Matrix Factorization
- Generalized and Differential Likelihood Ratio Tests with Quantum Signal Processing
- Generating Pseudo-relevant Representations for Spoken Document Retrieval
- Generative Adversarial Networks Based Error Concealment for Low Resolution Video
- Generative Adversarial Speaker Embedding Networks for Domain Robust End-to-end Speaker Verification
- Generative Graph Convolutional Network for Growing Graphs
- Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking
- Geometric Invariants for Sparse Unknown View Tomography
- Geometry of Deep Learning for Magnetic Resonance Fingerprinting
- Global Energy Efficiency Maximization in Non-orthogonal Interference Networks
- Global and Local Mode-domain Adaptive Algorithms for Spatial Active Noise Control Using Higher-order Sources
- Globally Convergent Accelerated Proximal Alternating Maximization Method for L1-Principal Component Analysis
- Glottal Instants Extraction from Speech Signal Using Generative Adversarial Network
- Glottographic and Aerodynamic Analysis on Consonant Aspiration and Onset F0 in Mandarin Chinese
- Gradient Image Super-resolution for Low-resolution Image Recognition
- Gradient-based Active Learning Query Strategy for End-to-end Speech Recognition
- Graph Filtering with Multiple Shift Matrices
- Graph Regularized Nonnegative Tucker Decomposition for Tensor Data Representation
- Graph Signal Representation with Wasserstein Barycenters
- Graph Signal Sampling via Reinforcement Learning
- Graph Spectral Clustering of Convolution Artefacts in Radio Interferometric Images
- Graph Spectral Domain Blind Watermarking
- Graph-based RGB-D Image Segmentation Using Color-directional-region Merging
- Graphic Delay Equalizer
- Graphical Lasso for High-dimensional Complex Gaussian Graphical Model Selection
- Grayscale-thermal Tracking via Canonical Correlation Analysis Based Inverse Sparse Representation
- Gridless Angle and Range Estimation for FDA-MIMO Radar Based on Decoupled Atomic Norm Minimization
- Gridless DOA Estimation via. Alternating Projections
- Gridless Super-resolution Doa Estimation with Unknown Mutual Coupling
- Group Action Equivariance and Generalized Convolution in Multi-layer Neural Networks
- Group Sparsity Based Target Localization for Distributed Sensor Array Networks
- Guided-spatio-temporal Filtering for Extracting Sound from Optically Measured Images Containing Occluding Objects
- HCU400: an Annotated Dataset for Exploring Aural Phenomenology through Causal Uncertainty
- HMM-based Approaches to Model Multichannel Information in Sign Language Inspired from Articulatory Features-based Speech Processing
- Hadamard Product Perspective on Source Resolvability of Spatial-smoothing-based Subspace Methods
- Hand Graph Representations for Unsupervised Segmentation of Complex Activities
- Hardware-friendly LDPC Decoding Scheduling for 5G HARQ Applications
- Hardware-oriented Memory-limited Online Fastica Algorithm and Hardware Architecture for Signal Separation
- Harmonic-band Complex Wavelet Transform Audio Analysis and Synthesis
- Head Related Impulse Response Interpolation and Extrapolation Using Deep Belief Networks
- Hearing Aid-controlled Beamformer for Binaural Speech Enhancement Using a Model-based Approach
- Heart Rate Estimation from Phonocardiogram Signals Using Non-negative Matrix Factorization
- Heterogeneous Information Fusion for Multitarget Tracking Using the Sum-product Algorithm
- Hierarchical Residual-pyramidal Model for Large Context Based Media Presence Detection
- Hierarchical Two-level Modelling of Emotional States in Spoken Dialog Systems
- Hierarchy-aware Loss Function on a Tree Structured Label Space for Audio Event Detection
- High Accuracy Image Rotation and Scale Estimation Using Radon Transform and Sub-pixel Shift Estimation
- High-quality Speech Coding with Sample RNN
- Higher-order Nonnegative CANDECOMP/PARAFAC Tensor Decomposition Using Proximal Algorithm
- Homography Estimation Based on Error Elliptical Distribution
- Horizontal 3D Sound Field Recording and 2.5D Synthesis with Omni-directional Circular Arrays
- Hotword Cleaner: Dual-microphone Adaptive Noise Cancellation with Deferred Filter Coefficients for Robust Keyword Spotting
- How Transferable Are Features in Convolutional Neural Network Acoustic Models across Languages?
- How Video Object Tracking Is Affected by In-capture Distortions?
- How to Globally Solve Non-convex Optimization Problems Involving an Approximate ℓ0 Penalization
- How to Improve Your Speaker Embeddings Extractor in Generic Toolkits
- Human Behaviour Recognition Using Wifi Channel State Information
- Human Perception Oriented Image Enhancement
- Hybrid Beamforming with Sub-arrayed MIMO Radar: Enabling Joint Sensing and Communication at mmWave Band
- Hybrid Beamforming: Where Should the Analog Power Amplifiers Be Placed?
- Hybrid Deep Neural Network Model for Remaining Useful Life Estimation
- Hybrid Sparse Array Design for Under-determined Models
- Hyper-parameter Learning for Sparse Structured Probabilistic Models
- Hypercomplex Low Rank Matrix Completion with Non-negative Constraints via Convex Optimization
- Hyperspectral Image Super-resolution Using Generative Adversarial Network and Residual Learning
- Hyperspectral and Multispectral Data Fusion by a Regularization Considering
- ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance Analysis
- Identifying Structural Brain Networks from Functional Connectivity: A Network Deconvolution Approach
- Image Captioning with Two Cascaded Agents
- Image Compression Using GMM Model Optimization
- Image Correction in Emission Tomography Using Deep Convolution Neural Network
- Image Demosaicking via Chrominance Images with Parallel Convolutional Neural Networks
- Image Reconstruction by Orthogonal Moments Derived by the Parity of Polynomials
- Image Reflection Removal Using the Wasserstein Generative Adversarial Network
- Image Restoration Using Total Variation Regularized Deep Image Prior
- Image Super-resolution via Deep Aggregation Network
- Imitation Refinement for X-ray Diffraction Signal Processing
- Immersive Audio Coding for Virtual Reality Using a Metadata-assisted Extension of the 3GPP EVS Codec
- Imperceptible Audio Communication
- Implementing Prosodic Phrasing in Chinese End-to-end Speech Synthesis
- Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality
- Importance of Analytic Phase of the Speech Signal for Detecting Replay Attacks in Automatic Speaker Verification Systems
- Improper Gaussian Signaling for the Two-user Broadcast Channel Treating Interference as Noise
- Improve Diverse Text Generation by Self Labeling Conditional Variational Auto Encoder
- Improved Estimation of the Distance between Covariance Matrices
- Improved Gesture Recognition Based on sEMG Signals and TCN
- Improved Hyperspectral Unmixing with Endmember Variability Parametrized Using an Interpolated Scaling Tensor
- Improved Latency-communication Trade-off for Map-shuffle-reduce Systems with Stragglers
- Improved Measurement Noise Covariance Estimation for N-channel Feedback Cancellation Based on the Frequency Domain Adaptive Kalman Filter
- Improved Metrical Alignment of Midi Performance Based on a Repetition-aware Online-adapted Grammar
- Improved Multipath Time Delay Estimation Using Cepstrum Subtraction
- Improvements to N-gram Language Model Using Text Generated from Neural Language Model
- Improvements to the Matching Projection Decoding Method for Ambisonic System with Irregular Loudspeaker Layouts
- Improving ASR Robustness to Perturbed Speech Using Cycle-consistent Generative Adversarial Networks
- Improving Audio-visual Speech Recognition Performance with Cross-modal Student-teacher Training
- Improving Binaural Ambisonics Decoding by Spherical Harmonics Domain Tapering and Coloration Compensation
- Improving CTC Using Stimulated Learning for Sequence Modeling
- Improving Children Speech Recognition through Feature Learning from Raw Speech Signal
- Improving Content-based Audio Retrieval by Vocal Imitation Feedback
- Improving Deep Models of Speech Quality Prediction through Voice Activity Detection and Entropy-based Measures
- Improving Emotion Classification through Variational Inference of Latent Variables
- Improving End-to-end Speech Recognition with Pronunciation-assisted Sub-word Modeling
- Improving Eye Movement Biometrics Using Remote Registration of Eye Blinking Patterns
- Improving Facial Attractiveness Prediction via Co-attention Learning
- Improving Graph Trend Filtering with Non-convex Penalties
- Improving Human-computer Interaction in Low-resource Settings with Text-to-phonetic Data Augmentation
- Improving Layer Trajectory LSTM with Future Context Frames
- Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning
- Improving Object Detection with Relation Graph Inference
- Improving Sequence-to-sequence Voice Conversion by Adding Text-supervision
- Improving Speech Emotion Recognition with Unsupervised Representation Learning on Unlabeled Speech
- Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
- Improving the Prediction of Therapist Behaviors in Addiction Counseling by Exploiting Class Confusions
- Improving the Rate-distortion Model of HEVC Intra by Integrating the Maximum Absolute Error
- In Search of the Optimal Walsh-hadamard Transform for Streamed Parallel Processing
- In-Car Driver Authentication Using Wireless Sensing
- Inaudible Speech Watermarking Based on Self-compensated Echo-hiding and Sparse Subspace Clustering
- Incorporate User Representation for Personal Question Answer Selection Using Siamese Network
- Increase Apparent Public Speaking Fluency by Speech Augmentation
- Incremental Binarization on Recurrent Neural Networks for Single-channel Source Separation
- Incremental Transfer Learning in Two-pass Information Bottleneck Based Speaker Diarization System for Meetings
- Indoor Time Reversal Wireless Communication: Experimental Results for Localization and Signal Coverage
- Inductive Conformal Predictor for Sparse Coding Classifiers: Applications to Image Classification
- Inference about Causality from Cardiotocography Signals Using Gaussian Processes
- Inferring Private Information in Wireless Sensor Networks
- Information Constrained Control for Visual Detection of Important Areas
- Information Theoretic Lower Bound of Restricted Isometry Property Constant
- Information-bottleneck Based on the Jensen-shannon Divergence with Applications to Pairwise Clustering
- Informed Ego-noise Suppression Using Motor Data-driven Dictionaries
- Inpainting in Omnidirectional Images for Privacy Protection
- Integrating Spectrotemporal Context into Features Based on Auditory Perception for Classification-based Speech Separation
- Intelligent Environments Based on Ultra-massive Mimo Platforms for Wireless Communication in Millimeter Wave and Terahertz Bands
- Inter- and Intra- Patient ECG Heartbeat Classification for Arrhythmia Detection: A Sequence to Sequence Deep Learning Approach
- Interactive Deep Colorization Using Simultaneous Global and Local Inputs
- Interactive Learning of Teacher-student Model for Short Utterance Spoken Language Identification
- Interactive Subjective Study on Picture-level Just Noticeable Difference of Compressed Stereoscopic Images
- Interference Exploitation Precoding for Multi-level Modulations
- Interpolation and Denoising of Graph Signals Using Plug-and-play Admm
- Intonation: A Dataset of Quality Vocal Performances Refined by Spectral Clustering on Pitch Congruence
- Intrasystem Entanglement Generator and Unambiguos Bell States Discriminator on Chip
- Introducing Machine Learning in Undergraduate DSP Classes
- Introducing Undergraduates to Pattern Recognition and Machine Learning through Speech Processing
- Introducing the Orthogonal Periodic Sequences for the Identification of Functional Link Polynomial Filters
- Investigating Context Features Hidden in End-to-end TTS
- Investigating Domain Sensitivity of DNN Embeddings for Speaker Recognition Systems
- Investigating End-to-end Speech Recognition for Mandarin-english Code-switching
- Investigating the Effects of Word Substitution Errors on Sentence Embeddings
- Investigation into Joint Optimization of Single Channel Speech Enhancement and Acoustic Modeling for Robust ASR
- Investigation of Enhanced Tacotron Text-to-speech Synthesis Systems with Self-attention for Pitch Accent Language
- Investigation of Modeling Units for Mandarin Speech Recognition Using Dfsmn-ctc-smbr
- Investigation of Sampling Techniques for Maximum Entropy Language Modeling Training
- Investigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic Models
- Investigation on Neural Bandwidth Extension of Telephone Speech for Improved Speaker Recognition
- Investigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic Features
- Ising Model Formulation of Outlier Rejection, with Application in WiFi Based Positioning
- Ising-dropout: A Regularization Method for Training and Compression of Deep Neural Networks
- Iterative Approximation of Analytic Eigenvalues of a Parahermitian Matrix EVD
- Iterative Hessian Sketch with Momentum
- Iterative Mirror Decomposition for Signal Representation
- Iteratively Reweighted Linear Least Squares for Frequency Estimation in Unbalanced Three-phase Power System
- Iteratively Reweighted Penalty Alternating Minimization Methods with Continuation for Image Deblurring
- JSR-Net: A Deep Network for Joint Spatial-radon Domain CT Reconstruction from Incomplete Data
- Joint Acoustic and Class Inference for Weakly Supervised Sound Event Detection
- Joint Codebook Design for Multi-cell Noma
- Joint Design for MIMO Radar and Downlink Communication Systems Coexistence
- Joint Endpointing and Decoding with End-to-end Models
- Joint Estimation of RETF Vector and Power Spectral Densities for Speech Enhancement Based on Alternating Least Squares
- Joint On-line Learning of a Zero-shot Spoken Semantic Parser and a Reinforcement Learning Dialogue Manager
- Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASR
- Joint Optimization of Quantization and Structured Sparsity for Compressed Deep Neural Networks
- Joint Separation and Dereverberation of Reverberant Mixtures with Multichannel Variational Autoencoder
- Joint Structured Graph Learning and Clustering Based on Concept Factorization
- Joint Structured Graph Learning and Unsupervised Feature Selection
- Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr
- Joint Transaction Transmission and Channel Selection in Cognitive Radio Based Blockchain Networks: A Deep Reinforcement Learning Approach
- Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model
- Jointly Predicting Future Sequence and Steering Angles for Dynamic Driving Scenes
- Jointly Sparse Convolutional Neural Networks in Dual Spatial-winograd Domains
- Just Noticeable Difference Model for Asymmetrically Distorted Stereoscopic Images
- K-edge Coded Apertures for Compressive Spectral X-ray Tomography
- Kernel Random Matrices of Large Concentrated Data: the Example of GAN-Generated Images
- Kernel Regression for Graph Signal Prediction in Presence of Sparse Noise
- Knowledge Distillation Using Output Errors for Self-attention End-to-end Models
- Knowledge Distillation for Recurrent Neural Network Language Modeling with Trust Regularization
- Knowledge Distillation for Small Foot-print Deep Speaker Embedding
- Kullback-Leibler Divergence Frequency Warping Scale for Acoustic Scene Classification Using Convolutional Neural Network
- L2 Learners' Emotion Production in Video Dubbing Practices
- LMS to Deep Learning: How DSP Analysis Adds Depth to Learning
- LMS: Past, Present and Future
- LPCNET: Improving Neural Speech Synthesis through Linear Prediction
- Labelled Non-zero Particle Flow for SMC-PHD Filtering
- LadderNet: Knowledge Transfer Based Viewpoint Prediction in 360◦ Video
- Langevin-based Strategy for Efficient Proposal Adaptation in Population Monte Carlo
- Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition
- Language Person Search with Mutually Connected Classification Loss
- Language-invariant Bottleneck Features from Adversarial End-to-end Acoustic Models for Low Resource Speech Recognition
- Lapped Transforms: A Graph-based Extension
- Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder Models
- Large-pose Face Alignment via Shape-aware Heatmap
- Latency Driven Fronthaul Bandwidth Allocation and Cooperative Beamforming for Cache-enabled Cloud-based Small Cell Networks
- Latent Heterogeneous Multilayer Community Detection
- Latent Representation Learning for Artificial Bandwidth Extension Using a Conditional Variational Auto-encoder
- Latent Schatten TT Norm for Tensor Completion
- Layer-wise Deep Neural Network Pruning via Iteratively Reweighted Optimization
- Layout-aware Subfigure Decomposition for Complex Figures in the Biomedical Literature
- Learned in Speech Recognition: Contextual Acoustic Word Embeddings
- Learning Affective Correspondence between Music and Image
- Learning Bandwidth Expansion Using Perceptually-motivated Loss
- Learning Comment Generation by Leveraging User-generated Data
- Learning Compact Partial Differential Equations for Color Images with Efficiency
- Learning Convolutional Neural Networks with Deep Part Embeddings
- Learning Discriminative Features from Spectrograms Using Center Loss for Speech Emotion Recognition
- Learning Discriminative Features in Sequence Training without Requiring Framewise Labelled Data
- Learning Discriminative Finger-knuckle-print Descriptor
- Learning Disentangled Representation in Latent Stochastic Models: A Case Study with Image Captioning
- Learning Dynamic Stream Weights for Linear Dynamical Systems Using Natural Evolution Strategies
- Learning Efficient Sparse Structures in Speech Recognition
- Learning Efficient Tensor Representations with Ring-structured Networks
- Learning Laplacian Matrix from Bandlimited Graph Signals
- Learning Latent Representations for Style Control and Transfer in End-to-end Speech Synthesis
- Learning Low Rank and Sparse Models via Robust Autoencoders
- Learning Motion Disfluencies for Automatic Sign Language Segmentation
- Learning Pose-aware 3D Reconstruction via 2D-3D Self-consistency
- Learning Requirements for Stealth Attacks
- Learning Search Path for Region-level Image Matching
- Learning Semantic-preserving Space Using User Profile and Multimodal Media Content from Political Social Network
- Learning Shallow Neural Networks via Provable Gradient Descent with Random Initialization
- Learning Shared Vector Representations of Lyrics and Chords in Music
- Learning Sheaf Laplacians from Smooth Signals
- Learning Similarity-specific Dictionary for Zero-shot Fine-grained Recognition
- Learning Sound Event Classifiers from Web Audio with Noisy Labels
- Learning Spatially-correlated Temporal Dictionaries for Calcium Imaging
- Learning Stochastic Representations of Geophysical Dynamics
- Learning Temporal Information from Spatial Information Using CapsNets for Human Action Recognition
- Learning Voice Source Related Information for Depression Detection
- Learning by Inertia: Self-supervised Monocular Visual Odometry for Road Vehicles
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.