ICASSP 2023 Accepted Papers
The full list of 2,718 papers accepted at ICASSP 2023 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- "Prediction of Sleepiness Ratings from Voice by Man and Machine": A Perceptual Experiment Replication Study
- -Complexity Low-Rank Approximation SVD for Massive Matrix in Tensor Train Format
- 2DSBG: A 2d Semi Bi-Gaussian Filter Adapted for Adjacent and Multi-Scale Line Feature Detection
- 3D Audio Signal Processing Systems for Speech Enhancement and Sound Localization and Detection
- 3D Point Cloud Completion Based on Multi-Scale Degradation
- 6G Integrated Sensing and Communication - Sensing Assisted Environmental Reconstruction and Communication
- A 3D-Assisted Framework to Evaluate the Quality of Head Motion Replication by Reenactment DEEPFAKE Generators
- A Bandit Online Convex Optimization Approach To Distributed Energy Management In Networked Systems
- A Bayesian Perspective for Determinant Minimization Based Robust Structured Matrix Factorization
- A Bayesian Perspective on Noise2Noise: Theory and Extensions
- A Benchmark for Evaluating Robustness of Spoken Language Understanding Models in Slot Filling
- A Bidirectional Joint Model for Spoken Language Understanding
- A Causal Convolutional Approach for Packet Loss Concealment in Low Powered Devices
- A Closer Look At Scoring Functions And Generalization Prediction
- A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at Scale
- A Compensated Shrinkage Affine Projection Algorithm for Debiased Sparse Adaptive Filtering
- A Comprehensive Comparison of Projections in Omnidirectional Super-Resolution
- A Computationally Efficient Algorithm for Distributed Adaptive Signal Fusion Based on Fractional Programs
- A Content Adaptive Learnable "Time-Frequency" Representation for audio Signal Processing
- A Content-Based Multi-Scale Network for Single Image Super-Resolution
- A Context-Aware Computational Approach for Measuring Vocal Entrainment in Dyadic Conversations
- A Contrastive Embedding-Based Domain Adaptation Method for Lung Sound Recognition in Children Community-Acquired Pneumonia
- A Contrastive Framework to Enhance Unsupervised Sentence Representation Learning
- A Contrastive Knowledge Transfer Framework for Model Compression and Transfer Learning
- A Controllable Lifestyle Simulator for Use in Deep Reinforcement Learning Algorithms
- A Critical Look at Recent Trends in Compression of Channel State Information
- A DNN Based Normalized Time-Frequency Weighted Criterion for Robust Wideband DoA Estimation
- A DNN-Based Hearing-Aid Strategy For Real-Time Processing: One Size Fits All
- A Database for Multi-Modal Short Video Quality Assessment
- A Dataset for Audio-Visual Sound Event Detection in Movies
- A Deep Disentangled Approach for Interpretable Hyperspectral Unmixing
- A Deep Fusion Rule for Infrared and Visible Image Fusion: Feature Communication for Importance Assessment
- A Deep Temporal Factor Analysis Method for Large Scale Financial Portfolio Selection
- A Discriminative Multi-Channel Noise Feature Representation Model for Image Manipulation Localization
- A Distributed Adaptive Algorithm for Non-Smooth Spatial Filtering Problems
- A Dual-Branch Adaptive Distribution Fusion Framework for Real-World Facial Expression Recognition
- A Dual-Path Transformer Network for Scene Text Detection
- A Dynamic Cross-Scale Transformer with Dual-Compound Representation for 3D Medical Image Segmentation
- A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding
- A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville Distribution
- A Few Shot Learning of Singing Technique Conversion Based on Cycle Consistency Generative Adversarial Networks
- A Flow-Guided Non-Local Alignment Network for Video Compressive Sensing Reconstruction
- A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement
- A Frequency-Domain Recursive Least-Squares Adaptive Filtering Algorithm Based On A Kronecker Product Decomposition
- A Frequency-Weighted Leaky Fxlms Algorithm with Application to Feedback Active Noise Control Systems
- A Fusion-Based and Multi-Layer Method for Low Light Image Enhancement
- A Game of Snakes and Gans
- A Gaussian Latent Variable Model for Incomplete Mixed Type Data
- A Generalized Subspace Distribution Adaptation Framework for Cross-Corpus Speech Emotion Recognition
- A Geometric Surrogate for Simulation Calibration
- A Graph Neural Network Multi-Task Learning-Based Approach for Detection and Localization of Cyberattacks in Smart Grids
- A Hierarchical Regression Chain Framework for Affective Vocal Burst Recognition
- A Highly Interpretable Deep Equilibrium Network for Hyperspectral Image Deconvolution
- A Holistic Cascade System, Benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation
- A Hybrid Deep Neural Network for Nonlinear Causality Analysis in Complex Industrial Control System
- A Knowledge-Driven Vowel-Based Approach of Depression Classification from Speech Using Data Augmentation
- A Large-Scale Pretrained Deep Model for Phishing URL Detection
- A Learnable Spatial Mapping for Decoding the Directional Focus of Auditory Attention Using EEG
- A Lightweight Convolutional Neural Network using Feature Filtering Module
- A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech Enhancement
- A Low-Latency Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation
- A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids
- A Magnetic Framelet-Based Convolutional Neural Network for Directed Graphs
- A Mathematical Model for Neuronal Activity and Brain Information Processing Capacity
- A Memory-Free Evolving Bipolar Neural Network for Efficient Multi-Label Stream Learning
- A Meta-Gnn Approach to Personalized Seizure Detection and Classification
- A Method of Constructing and Automatically Labeling Radio Frequency Signal Training Dataset for UAV
- A Model-Based Hearing Compensation Method Using a Self-Supervised Framework
- A Momentum Two-Gradient Direction Algorithm with Variable Step Size Applied to Solve Practical Output Constraint Issue for Active Noise Control
- A Multi-Channel Aggregation Framework for Object Detection in Large-Scale SAR Image
- A Multi-Modal Approach For Context-Aware Network Traffic Classification
- A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement
- A Multi-Signal Perception Network for Textile Composition Identification
- A Multi-Stage Hierarchical Relational Graph Neural Network for Multimodal Sentiment Analysis
- A Multi-Stage Low-Latency Enhancement System for Hearing Aids
- A Multi-Stage Triple-Path Method For Speech Separation in Noisy and Reverberant Environments
- A Mutual Implicit Sentiment Analysis Model with Bundle-Aware Contrastive Learning
- A Nested Ensemble Method to Bilevel Machine Learning
- A New Approach to Extract Fetal Electrocardiogram Using Affine Combination of Adaptive Filters
- A New Personalized Efficacy Atlas for Pallidal Deep Brain Stimulation
- A New Probabilistic Distance Metric with Application in Gaussian Mixture Reduction
- A New Semi-Supervised Classification Method Using a Supervised Autoencoder for Biomedical Applications
- A Novel Approach Based on Voronoï Cells to Classify Spectrogram Zeros of Multicomponent Signals
- A Novel Cross-Component Context Model for End-to-End Wavelet Image Coding
- A Novel Efficient Multi-View Traffic-Related Object Detection Framework
- A Novel Extrapolation Technique to Accelerate WMMSE
- A Novel Heart Rate Estimation Method Exploiting Heartbeat Second Harmonic Reconstruction Via Millimeter Wave Radar
- A Novel Metric For Evaluating Audio Caption Similarity
- A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVC
- A Novel State Connection Strategy for Quantum Computing to Represent and Compress Digital Images
- A Novel Transformer-Based Pipeline for Lung Cytopathological Whole Slide Image Classification
- A Parallel Attention Mechanism for Image Manipulation Detection and Localization
- A Patient Invariant Model Towards the Prediction of Freezing of Gait
- A Perceptual Neural Audio Coder with a Mean-Scale Hyperprior
- A Person Identification System for the ICASSP 2023 e-Prevention Challenge
- A Perturbation-Based Policy Distillation Framework with Generative Adversarial Nets
- A Phoneme-Informed Neural Network Model For Note-Level Singing Transcription
- A Physically Explainable Framework for Human-Related Anomaly Detection
- A Point is A Wave: Point-Wave Network for Place Recognition
- A Practical Distributed Active Noise Control Algorithm Overcoming Communication Restrictions
- A Principled Approach to Model Validation in Domain Generalization
- A Privacy-Preserving Trajectory Mining Model
- A Probabilistic Framework for Pruning Transformers Via a Finite Admixture of Keys
- A Processing Framework to Access Large Quantities of Whispered Speech Found in ASMR
- A Progressive Neural Network for Acoustic Echo Cancellation
- A Prototypical Semantic Decoupling Method via Joint Contrastive Learning for Few-Shot Named Entity Recognition
- A Proximal Approach to IVA-G with Convergence Guarantees
- A Quantum Approach for Stochastic Constrained Binary Optimization
- A Quantum Kernel Learning Approach to Acoustic Modeling for Spoken Command Recognition
- A Radar-Jammer Zero-Sum Repeated Bayesian Game
- A Reality Check and a Practical Baseline for Semantic Speech Embedding
- A Robust Kalman Filter Based Approach for Indoor Robot Positionning with Multi-Path Contaminated UWB Data
- A Role Engineering Approach Based on Spectral Clustering Analysis for Restful Permissions in Cloud
- A Sentiment and Syntactic-Aware Graph Convolutional Network for Aspect-Level Sentiment Classification
- A Sidecar Separator Can Convert A Single-Talker Speech Recognition System to A Multi-Talker One
- A Simple Scheme for Coupled Factorization for Hyperspectral Super-Resolution: Exploiting Sparsity in an Easy Way
- A Simple Yet Effective Approach to Structured Knowledge Distillation
- A Simulation-Based Framework for Urban Traffic Accident Detection
- A Slot-Shared Span Prediction-Based Neural Network for Multi-Domain Dialogue State Tracking
- A Spatial-Temporal ECG Emotion Recognition Model Based on Dynamic Feature Fusion
- A Spatio-Temporal Decomposition Network for Compressed Video Quality Enhancement
- A Speech Representation Anonymization Framework via Selective Noise Perturbation
- A Statistical Interpretation of the Maximum Subarray Problem
- A Study of Audio Mixing Methods for Piano Transcription in Violin-Piano Ensembles
- A Study on Bias and Fairness in Deep Speaker Recognition
- A Study on the Integration of Pipeline and E2E SLU Systems for Spoken Semantic Parsing Toward Stop Quality Challenge
- A Study on the Invariance in Security Whatever the Dimension of Images for the Steganalysis by Deep-Learning
- A Synthetic Corpus Generation Method for Neural Vocoder Training
- A Targeted Sampling Strategy for Compressive Cryo Focused Ion Beam Scanning Electron Microscopy
- A Template Matching Approach for Reference Picture Padding in Video Coding
- A Token-Level Contrastive Framework for Sign Language Translation
- A Topic-Enhanced Approach for Emotion Distribution Forecasting in Conversations
- A Transformer-Based E2E SLU Model for Improved Semantic Parsing
- A Two-Branch Network for Video Anomaly Detection with Spatio-Temporal Feature Learning
- A Two-Stage System for Spoken Language Understanding
- A Unified One-Shot Prosody and Speaker Conversion System with Self-Supervised Discrete Speech Units
- A Unified Uncertainty-Aware Exploration: Combining Epistemic and Aleatory Uncertainty
- A Unitary Transform Based Generalized Approximate Message Passing
- A Variational Inequality Model for Learning Neural Networks
- A Video Anomaly Detection Framework Based on Appearance-Motion Semantics Representation Consistency
- A Wavelet Scattering Approach for Load Identification with Limited Amount of Training Data
- A non-contact SpO2 estimation using video magnification and infrared data
- A2S-NAS: Asymmetric Spectral-Spatial Neural Architecture Search for Hyperspectral Image Classification
- A3S: Adversarial Learning of Semantic Representations for Scene-Text Spotting
- ACE-VC: Adaptive and Controllable Voice Conversion Using Explicitly Disentangled Self-Supervised Speech Representations
- ACF: Aligned Contrastive Finetuning For Language and Vision Tasks
- AD-YOLO: You Look Only Once in Training Multiple Sound Event Localization and Detection
- ADHD Classification with Biomarker Identification Using a Triplet Loss Attention Auto-Encoding Network
- AE-Flow: Autoencoder Normalizing Flow
- AERO: Audio Super Resolution in the Spectral Domain
- AMC-Net: An Effective Network for Automatic Modulation Classification
- AMPose: Alternately Mixed Global-Local Attention Model for 3D Human Pose Estimation
- APGP: Accuracy-Preserving Generative Perturbation for Defending Against Model Cloning Attacks
- ASSD: Synthetic Speech Detection in the AAC Compressed Domain
- AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer
- AURA: Privacy-Preserving Augmentation to Improve Test Set Diversity in Speech Enhancement
- AV-TAD: Audio-Visual Temporal Action Detection With Transformer
- AVES: Animal Vocalization Encoder Based on Self-Supervision
- Absolute Decision Corrupts Absolutely: Conservative Online Speaker Diarisation
- Abstract Representation for Multi-Intent Spoken Language Understanding
- Abusive Activity Detection with Multi-Modality Based on Convolutional Neural Network
- Accelerated Distributed Stochastic Non-Convex Optimization over Time-Varying Directed Networks
- Accelerated Massive MIMO Detector Based on Annealed Underdamped Langevin Dynamics
- Accelerating Matrix Trace Estimation by Aitken's Δ2 Process
- Accelerating RNN-T Training and Inference Using CTC Guidance
- Accidental Learners: Spoken Language Identification in Multilingual Self-Supervised Models
- Achievable Error Exponents for Almost Fixed-Length M-Ary Hypothesis Testing
- Achieving Fair Speech Emotion Recognition via Perceptual Fairness
- Acoustic Source Localization in the Spherical Harmonics Domain Exploiting Low-Rank Approximations
- Acoustically-Driven Phoneme Removal that Preserves Vocal Affect Cues
- Active Beam Tracking with Reconfigurable Intelligent Surface
- Active IRS-Assisted MIMO Channel Estimation and Prediction
- Active Learning for Efficient Few-Shot Classification
- Active Learning of non-Semantic Speech Tasks with Pretrained models
- Active Noise Control over 3D Space: A Realistic Error Microphone Geometry Design
- Active Perception System for Enhanced Visual Signal Recovery Using Deep Reinforcement Learning
- Active Selection of Source Patients in Transfer Learning for Epileptic Seizure Detection Using Riemannian Manifold
- Active Subsampling Using Deep Generative Models by Maximizing Expected Information Gain
- Activity-Informed Industrial Audio Anomaly Detection Via Source Separation
- AdapITN: A Fast, Reliable, and Dynamic Adaptive Inverse Text Normalization
- Adaptable End-to-End ASR Models Using Replaceable Internal LMs and Residual Softmax
- Adapted Multimodal Bert with Layer-Wise Fusion for Sentiment Analysis
- Adapter Tuning With Task-Aware Attention Mechanism
- Adapting Exploratory Behaviour in Active Inference for Autonomous Driving
- Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker Embeddings
- Adapting a Self-Supervised Speech Representation for Noisy Speech Emotion Recognition by Using Contrastive Teacher-Student Learning
- Adaptive Axonal Delays in Feedforward Spiking Neural Networks for Accurate Spoken Word Recognition
- Adaptive CSI Feedback with Hidden Semantic Information Transfer
- Adaptive Data Augmentation for Contrastive Learning
- Adaptive Eccm for Mitigating Smart Jammers
- Adaptive Endpointing with Deep Contextual Multi-Armed Bandits
- Adaptive Filtering Algorithms For Set-Valued Observations-Symmetric Measurement Approach To Unlabeled And Anonymized Data
- Adaptive Gaussian Nested Filter for Parameter Estimation and State Tracking in Dynamical Systems
- Adaptive Knowledge Distillation Between Text and Speech Pre-Trained Models
- Adaptive Large Margin Fine-Tuning For Robust Speaker Verification
- Adaptive Mask Co-Optimization for Modal Dependence in Multimodal Learning
- Adaptive Multi-Corpora Language Model Training for Speech Recognition
- Adaptive Noise Canceller Algorithm with SNR-Based Stepsize and Data-Dependent Averaging
- Adaptive Non-Local Generative Adversarial Networks for Low-Dose CT Image Denoising
- Adaptive Scale and Spatial Aggregation for Real-Time Object Detection
- Adaptive Semantic Fusion Framework for Unsupervised Monocular Depth Estimation
- Adaptive Simulated Annealing Through Alternating Rényi Divergence Minimization
- Adaptive Step-Size Methods for Compressed SGD
- Adaptive Submanifold-Preserving Sparse Regression for Feature Selection And Multiclass Classification
- Adaptive Time-Scale Modification for Improving Speech Intelligibility Based On Phoneme Clustering For Streaming Services
- Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech Activity
- Adversarial Attacks on Genotype Sequences
- Adversarial Contrastive Distillation with Adaptive Denoising
- Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition
- Adversarial Guitar Amplifier Modelling with Unpaired Data
- Adversarial Network Pruning by Filter Robustness Estimation
- Adversarial Permutation Invariant Training for Universal Sound Separation
- Adversarially Robust Fairness-Aware Regression
- Affinity Learning With Blind-Spot Self-Supervision for Image Denoising
- Agile Radio Map Prediction Using Deep Learning
- Aiding Speech Harmonic Recovery in DNN-Based Single Channel Noise Reduction Using Cepstral Excitation Manipulation (CEM) Components
- Aleatoric Uncertainty Estimation of Overnight Sleep Statistics Through Posterior Sampling Using Conditional Normalizing Flows
- Algebraic Convolutional Filters on Lie Group Algebras
- Align, Write, Re-Order: Explainable End-to-End Speech Translation via Operation Sequence Generation
- Alignment Entropy Regularization
- Alternating Constrained Minimization Based Approximate Message Passing
- Alternating Phase Langevin Sampling with Implicit Denoiser Priors for Phase Retrieval
- Amicable Aid: Perturbing Images to Improve Classification Performance
- An ASR-Free Fluency Scoring Approach with Self-Supervised Learning
- An Adapter Based Multi-Label Pre-Training for Speech Separation and Enhancement
- An Adaptive DFE Using Light-Pattern-Protection Algorithm in 12 NM CMOS Technology
- An Adaptive Enhancement Method for Gastrointestinal Low-Light Images of Capsule Endoscope
- An Adaptive Plug-and-Play Network for Few-Shot Learning
- An Analysis of Degenerating Speech Due to Progressive Dysarthria on ASR Performance
- An Antispoofing Approach in Biometric Authentication System for a Smartcard
- An Application of Quantum Mechanics to Attention Methods in Computer Vision
- An Approach to Ontological Learning from Weak Labels
- An Asynchronous Updating Reinforcement Learning Framework for Task-Oriented Dialog System
- An Attention-Based Approach to Hierarchical Multi-Label Music Instrument Classification
- An Augmented Gaussian Sum Filter through a mixture Decomposition
- An Auto-Encoder Based Method for Camera Fingerprint Compression
- An Automotive Radar Dataset For Object Classification
- An Edge Alignment-Based Orientation Selection Method for Neutron Tomography
- An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies
- An Efficient Beam-Sharing Algorithm for RIS-aided Simultaneous Wireless Information and Power Transfer Applications
- An Efficient Relay Selection Scheme for Relay-assisted HARQ
- An Empirical Study and Improvement for Speech Emotion Recognition
- An Empirical Study of Backdoor Attacks on Masked Auto Encoders
- An Empirical Study on Speech Restoration Guided by Self-Supervised Speech Representation
- An End-to-End Framework for Partial View-Aligned Clustering with Graph Structure
- An End-to-End Neural Network for Image-to-Audio Transformation
- An Evaluation Platform to Scope Performance of Synthetic Environments in Autonomous Ground Vehicles Simulation
- An Experimental Study on Sound Event Localization and Detection Under Realistic Testing Conditions
- An Implicit Gradient Method for Constrained Bilevel Problems Using Barrier Approximation
- An Improved Optimal Transport Kernel Embedding Method with Gating Mechanism for Singing Voice Separation and Speaker Identification
- An Interpretable Model Using Evidence Information for Multi-Hop Question Answering Over Long Texts
- An Isotropy Analysis for Self-Supervised Acoustic Unit Embeddings on the Zero Resource Speech Challenge 2021 Framework
- An Online Algorithm for Chance Constrained Resource Allocation
- An Online Algorithm for Contrastive Principal Component Analysis
- Analysing Diffusion-based Generative Approaches Versus Discriminative Approaches for Speech Restoration
- Analysing Discrete Self Supervised Speech Representation For Spoken Language Modeling
- Analysing the Masked Predictive Coding Training Criterion for Pre-Training a Speech Representation Model
- Analysis Of Noisy-Target Training For Dnn-Based Speech Enhancement
- Analysis and Re-Synthesis of Natural Cricket Sounds Assessing the Perceptual Relevance of Idiosyncratic Parameters
- Analysis and Transformation of Voice Level in Singing Voice
- Analyzing Acoustic Word Embeddings from Pre-Trained Self-Supervised Speech Models
- Anchored Speech Recognition with Neural Transducers
- Ancient Chinese Word Segmentation and Part-of-Speech Tagging Using Distant Supervision
- Angle-Of-Arrival Target Tracking Using A Mobile Uav In External Signal-Denied Environment
- Animal Re-Identification Algorithm for Posture Diversity
- Anomalous Signal Detection for Cyber-Physical Systems Using Interpretable Causal Neural Network
- Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining
- Anomaly Detection in Optical Spectra VIA Joint Optimization
- Antenna Impedance Estimation in Correlated Rayleigh Fading Channels
- Any-to-Any Voice Conversion with F0 and Timbre Disentanglement and Novel Timbre Conditioning
- Applying Independent Vector Analysis on EEG-Based Motor Imagery Classification
- Applying Symmetrical Component Transform for Industrial Appliance Classification in Non-Intrusive Load Monitoring
- Approximation Error Back-Propagation for Q-Function in Scalable Reinforcement Learning with Tree Dependence Structure
- Aprogressive Image Dehazing Framework with inter and Intra Contrastive Learning
- Articulation GAN: Unsupervised Modeling of Articulatory Learning
- Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization
- Assessing the Robustness of Deep Learning-Assisted Pathological Image Analysis Under Practical Variables of Imaging System
- Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting A Calibrated External Microphone Array
- Associative Learning Network for Coherent Visual Storytelling
- Asymmetric Polynomial Loss for Multi-Label Classification
- Asymptotic Bias and Variance of Kernel Ridge Regression
- Asymptotic Distribution of Stochastic Mirror Descent Iterates in Average Ensemble Models
- Asymptotically Optimal Nonparametric Classification Rules for Spike Train Data
- Asynchronous Federated Learning for Real-Time Multiple Licence Plate Recognition Through Semantic Communication
- Asynchronous Social Learning
- Attention Based Relation Network for Facial Action Units Recognition
- Attention Localness in Shared Encoder-Decoder Model For Text Summarization
- Attention Mixup: An Accurate Mixup Scheme Based On Interpretable Attention Mechanism for Multi-Label Audio Classification
- Attention-Guided Deep Learning Framework For Movement Quality Assessment
- Audio Barlow Twins: Self-Supervised Audio Representation Learning
- Audio Coding With Unified Noise Shaping And Phase Contrast Control
- Audio Cross Verification Using Dual Alignment Likelihood Ratio Test
- Audio Quality Assessment of Vinyl Music Collections Using Self-Supervised Learning
- Audio Signal Enhancement with Learning from Positive and Unlabeled Data
- Audio-Driven Facial Landmark Generation in Violin Performance using 3DCNN Network with Self Attention Model
- Audio-Driven High Definetion and Lip-Synchronized Talking Face Generation Based on Face Reenactment
- Audio-Driven Talking Head Video Generation with Diffusion Model
- Audio-Text Models Do Not Yet Leverage Natural Language
- Audio-Visual Inpainting: Reconstructing Missing Visual Information with Sound
- Audio-Visual Speaker Diarization in the Framework of Multi-User Human-Robot Interaction
- Audio-Visual Speech Enhancement with a Deep Kalman Filter Generative Model
- Audio-to-Intent Using Acoustic-Textual Subword Representations from End-to-End ASR
- Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec
- AugTarget Data Augmentation for Infrared Small Target Detection
- Augmentation Robust Self-Supervised Learning for Human Activity Recognition
- Augmenting Transformer-Transducer Based Speaker Change Detection with Token-Level Training Loss
- Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
- AutoGCF: Personalized Aggregation on Neural Graph Collaborative Filtering
- Automatic Camera Pose Estimation by Key-Point Matching of Reference Objects
- Automatic Classification of Vocal Intensity Category from Speech
- Automatic Error Detection in Integrated Circuits Image Segmentation: A Data-Driven Approach
- Automatic Segmentation of Nasopharyngeal Carcinoma in CT Images Using Dual Attention and Edge Detection
- Automatic Severity Classification of Dysarthric Speech by Using Self-Supervised Model with Multi-Task Learning
- Autonomous Navigation of a Robotic Swarm in Space Exploration Missions
- Autonomous Soundscape Augmentation with Multimodal Fusion of Visual and Participant-Linked Inputs
- Autotts: End-to-End Text-to-Speech Synthesis Through Differentiable Duration Modeling
- Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal Processing
- Auxiliary Pooling Layer For Spoken Language Understanding
- Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker Extraction
- Avoid Overthinking in Self-Supervised Models for Speech Recognition
- BATT: Backdoor Attack with Transformation-Based Triggers
- BAUENet: Boundary-Aware Uncertainty Enhanced Network for Infrared Small Target Detection
- BEANS: The Benchmark of Animal Sounds
- BECTRA: Transducer-Based End-To-End ASR with Bert-Enhanced Encoder
- BER-Aware Dynamic Resource Management for Edge-Assisted Goal-Oriented Communications
- BHE-DARTS: Bilevel Optimization Based on Hypergradient Estimation for Differentiable Architecture Search
- BIRD-PCC: Bi-Directional Range Image-Based Deep Lidar Point Cloud Compression
- BISVP: Building Footprint Extraction Via Bidirectional Serialized Vertex Prediction
- BTS-E: Audio Deepfake Detection Using Breathing-Talking-Silence Encoder
- Backdoor Attack Against Automatic Speaker Verification Models in Federated Learning
- Backdoor Defense via Suppressing Model Shortcuts
- Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation
- Background-Weakening Consistency Regularization for Semi-Supervised Video Action Detection
- BadRes: Reveal the Backdoors Through Residual Connection
- Bag of Tricks with Quantized Convolutional Neural Networks for Image Classification
- Bagging R-CNN: Ensemble for Object Detection in Complex Traffic Scenes
- Balanced Deep CCA for Bird Vocalization Detection
- Balanced Mixup Loss for Long-Tailed Visual Recognition
- Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation
- Batch Normalization Damages Federated Learning on NON-IID Data: Analysis and Remedy
- Batch-Ensemble Stochastic Neural Networks for Out-of-Distribution Detection
- Bayesian Cramér-Rao Bound Estimation With Score-Based Models
- Bayesian Methods for Optical Flow Estimation Using a Variational Approximation, with Applications to Ultrasound
- Bayesian Network Modeling and Prediction of Transitions Within the Homelessness System
- Bayesian Optimization with Ensemble Learning Models and Adaptive Expected Improvement
- Beamformer-Guided Target Speaker Extraction
- Beamforming Optimization in RIS-Aided Mimo Systems Under Multiple-Reflection Effects
- Bebert: Efficient And Robust Binary Ensemble Bert
- Benchmark of Physiological Model Based and Deep Learning Based Remote Photoplethysmography in Automotive Applications
- Benchmarking Convolutional Neural Network Inference on Low-Power Edge Devices
- Benchmarking Cross-Domain Face Recognition with Avatars, Caricatures and Sketches
- Benchmarking White Blood Cell Classification under Domain Shift
- Bert is Robust! A Case Against Word Substitution-Based Adversarial Attacks
- Better Together: Dialogue Separation and Voice Activity Detection for Audio Personalization in TV
- Beyond Neural-on-Neural Approaches to Speaker Gender Protection
- Beyond Rate Coding: Signal Coding and Reconstruction Using Lean Spike Trains
- Bias Identification with RankPix Saliency
- Bias Reduced Semidefinite Relaxation Method for Multistatic Localization in the Absence of Transmitter Position And Its Synchronization
- Bilateral Coarse-to-Fine Network for Point Cloud Completion
- Bimodal Fusion Network for Basic Taste Sensation Recognition from Electroencephalography and Electromyography
- Binary Image Fast Perfect Recovery from Sparse 2D-DFT Coefficients
- Binary Sequence Set Optimization for CDMA Applications via Mixed-Integer Quadratic Programming
- Binauralization Robust To Camera Rotation Using 360° Videos
- Biologically-Inspired Continual Learning of Human Motion Sequences
- Bipartite Graph Convolutional Networks with Adversarial Domain Transfer
- Bit Error and Block Error Rate Training for ML-Assisted Communication
- Blind Acoustic Room Parameter Estimation Using Phase Features
- Blind Estimation of Audio Processing Graph
- Blind Polynomial Regression
- Blind Source Counting and Separation with Relative Harmonic Coefficients
- Block-Based Color Constancy: The Deviation of Salient Pixels
- Blood Oxygen Saturation Estimation from Facial Video Via DC and AC Components of Spatio-Temporal Map
- Body Prior Guided Graph Convolutional Neural Network for Skeleton-Based Action Recognition
- Boosting Bert Subnets with Neural Grafting
- Boosting Face Recognition Performance with Synthetic Data and Limited Real Data
- Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning
- Boosting No-Reference Super-Resolution Image Quality Assessment with Knowledge Distillation and Extension
- Boosting Person Re-Identification with Viewpoint Contrastive Learning and Adversarial Training
- Boosting Prompt-Based Few-Shot Learners Through Out-of-Domain Knowledge Distillation
- Boosting Semi-Supervised Federated Learning with Model Personalization and Client-Variance-Reduction
- Boosting Signal Modulation Few-Shot Learning with Pre-Transformation
- Boosting Transferability of Adversarial Example via an Enhanced Euler's Method
- Boosting the Accuracy of SRAM-Based in-Memory Architectures Via Maximum Likelihood-Based Error Compensation Method
- Boundary Cue Guidance and Contextual Feature Mining for Glass Segmentation
- Brain Network Features Differentiate Intentions from Different Emotional Expressions of the Same Text
- Brainnetformer: Decoding Brain Cognitive States with Spatial-Temporal Cross Attention
- Breaking the Trade-Off in Personalized Speech Enhancement With Cross-Task Knowledge Distillation
- BreathIE: Estimating Breathing Inhale Exhale Ratio Using Motion Sensor Data from Consumer Earbuds
- Bridging Speech and Textual Pre-Trained Models With Unsupervised ASR
- Building Blocks for a Complex-Valued Transformer Architecture
- Building Change Detection Using Cross-Temporal Feature Interaction Network
- Building Keyword Search System from End-To-End Asr Systems
- Burst Perception-Distortion Tradeoff: Analysis and Evaluation
- Bytecover3: Accurate Cover Song Identification On Short Queries
- Byzantine-Robust and Communication-Efficient Personalized Federated Learning
- C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object Detection
- C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval
- CADET: Control-Aware Dynamic Edge Computing for Real-Time Target Tracking in UAV Systems
- CAENet: Using Collaborative Attention Transformer and Add-Boost Strategy for Single Image Deraining
- CAN2V: Can-Bus Data-Based Seq2seq Model for Vehicle Velocity Prediction
- CANDY: Category-Kernelized Dynamic Convolution for Instance Segmentation
- CANet: Curved Guide Line Network with Adaptive Decoder for Lane Detection
- CAT: Causal Audio Transformer for Audio Classification
- CB-Conformer: Contextual Biasing Conformer for Biased Word Recognition
- CC-PoseNet: Towards Human Pose Estimation in Crowded Classrooms
- CD-FSOD: A Benchmark For Cross-Domain Few-Shot Object Detection
- CDHD: Contrastive Dreamer for Hint Distillation
- CF-VTON: Multi-Pose Virtual Try-on with Cross-Domain Fusion
- CFFMixer: Multi-Dimensional Feature Fusion for Object Detection
- CLAP Learning Audio Concepts from Natural Language Supervision
- CLIP4VideoCap: Rethinking Clip for Video Captioning with Multiscale Temporal Fusion and Commonsense Knowledge
- CLMAE: A Liter and Faster Masked Autoencoders
- CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing
- CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech Synthesis
- CNEG-VC: Contrastive Learning Using Hard Negative Example In Non-Parallel Voice Conversion
- CNN Filter for RPR-Based SR in VVC with Wavelet Decomposition
- CNN Filter for Super-Resolution with RPR Functionality in VVC
- CO-NET: Classification-Oriented Point Cloud Sampling via Informative Feature Learning and Non-Overlapped Local Adjustment
- CONSEN: Complementary and Simultaneous Ensemble for Alzheimer's Disease Detection and MMSE Score Prediction
- CORSD: Class-Oriented Relational Self Distillation
- COVID-19 Detection from Speech in Noisy Conditions
- CPA: Compressed Private Aggregation for Scalable Federated Learning Over Massive Networks
- CPD-GAN: Cascaded Pyramid Deformation GAN for Pose Transfer
- CRFAST: Clip-Based Reference-Guided Facial Image Semantic Transfer
- CROSSSPEECH: Speaker-Independent Acoustic Representation for Cross-Lingual Speech Synthesis
- CSM In Motion Vector Steganalysis: The Effect of Coders on Motion Vectors in H.264 Video Encoding
- CTCBERT: Advancing Hidden-Unit Bert with CTC Objectives
- CTTSR: A Hybrid CNN-Transformer Network for Scene Text Image Super-Resolution
- Calibrating AI Models for Few-Shot Demodulation VIA Conformal Prediction
- Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?
- Can Spoofing Countermeasure And Speaker Verification Systems Be Jointly Optimised?
- Cancelling Intermodulation Distortions for Otoacoustic Emission Measurements with Earbuds
- Capacity Maximization for Active RIS Assisted Outdoor-to-Indoor Communication System
- Capturing Cross-Scale Disparity for Stereo Image Super-Resolution
- Cardiac Disease Diagnosis on Imbalanced Electrocardiography Data Through Optimal Transport Augmentation
- Cascading and Direct Approaches to Unsupervised Constituency Parsing on Spoken Sentences
- Causal Discovery and Causal Inference Based Counterfactual Fairness in Machine Learning
- Central Nodes Detection from Partially Observed Graph Signals
- Centralized Cascade Multi-Channel Noise Reduction and Acoustic Feedback Cancellation in a Wireless Acoustic Sensor And Actuator Network
- Centroid Distance Distillation for Effective Rehearsal in Continual Learning
- Certified Robustness of Quantum Classifiers Against Adversarial Examples Through Quantum Noise
- Change Point Detection with Neural Online Density-Ratio Estimator
- Channel Estimation in Massive MIMO with Heavy-Tailed Noise: Gaussian-Mixture Versus Cauchy Models
- Channel Estimation with Tightly-Coupled Antenna Arrays
- Channel State Information-Free Artificial Noise-Aided Location-Privacy Enhancement
- Channel-Driven Decentralized Bayesian Federated Learning for Trustworthy Decision Making in D2D Networks
- Choice Fusion As Knowledge For Zero-Shot Dialogue State Tracking
- Chord-Conditioned Melody Harmonization With Controllable Harmonicity
- Class-Aware Contextual Information for Semantic Segmentation
- Class-Aware Shared Gaussian Process Dynamic Model
- Class-Guided Triple Head Prediction Network for Long-Tail Object Detection
- Class-Incremental Learning on Multivariate Time Series Via Shape-Aligned Temporal Distillation
- ClassA Entropy for the Analysis of Structural Complexity of Physiological Signals
- Classification of Synthetic Facial Attributes by Means of Hybrid Classification/Localization Patch-Based Analysis
- Classification of the Cervical Vertebrae Maturation (CVM) Stages Using the Tripod Network
- Classification via Subspace Learning Machine (SLM): Methodology and Performance Evaluation
- Classification-Based Dynamic Network for Efficient Super-Resolution
- Classifying Non-Individual Head-Related Transfer Functions with A Computational Auditory Model: Calibration And Metrics
- Classifying Pathological Images Based on Multi-Instance Learning and End-to-End Attention Pooling
- Clean Sample Guided Self-Knowledge Distillation for Image Classification
- Cleanformer: A Multichannel Array Configuration-Invariant Neural Enhancement Frontend for ASR in Smart Speakers
- Clicker: Attention-Based Cross-Lingual Commonsense Knowledge Transfer
- Client Selection for Generalization in Accelerated Federated Learning: A Bandit Approach
- Clustered Greedy Algorithm For Large-Scale Sensor Selection
- Clustering-Based Supervised Contrastive Learning for Identifying Risk Items on Heterogeneous Graph
- Co-Design for Mimo Radar and Mimo Communication Aided by Reconfigurable Intelligent Surface
- Co-Operative CNN for Visual Saliency Prediction on WCE Images
- Coarse-To-Fine Knowledge Selection for Document Grounded Dialogs
- Coarse-to-Fine Covid-19 Segmentation via Vision-Language Alignment
- Cochlear Decomposition: A Novel Bio-Inspired Multiscale Analysis Framework
- Cocktail Hubert: Generalized Self-Supervised Pre-Training for Mixture and Single-Source Speech
- Code-Enhanced Fine-Grained Semantic Matching For Tag Recommendation In Software Information Sites
- Code-Switching Speech Synthesis Based on Self-Supervised Learning and Domain Adaptive Speaker Encoder
- Code-Switching Text Generation and Injection in Mandarin-English ASR
- Codebook-Based User Tracking in IRS-Assisted mmWave Communication Networks
- Coded Matrix Computations for D2D-Enabled Linearized Federated Learning
- Codes Correcting Burst and Arbitrary Erasures for Reliable and Low-Latency Communication
- Cold Diffusion for Speech Enhancement
- Collaborative Audio-Visual Event Localization Based on Sequential Decision and Cross-Modal Consistency
- Color Guided Depth Map Super-Resolution with Nonlocla Autoregres-Sive Modeling
- Column-Based Matrix Approximation with Quasi-Polynomial Structure
- Combining Dual-Tree Wavelet Analysis and Proximal Optimization for Anisotropic Scale-Free Texture Segmentation
- Combining Loss Reweighting and Sample Resampling for Long-Tailed Instance Segmentation
- Combining the Silhouette and Skeleton Data for Gait Recognition
- Commdre: Document-Level Relation Extraction with Self-Supervised Commonsense Learning
- Communication-Constrained Exchange of Zeroth-Order Information with Application to Collaborative Target Tracking
- Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization
- Comparative Layer-Wise Analysis of Self-Supervised Speech Models
- Comparative Study of IRS Assisted Opportunistic Communications Over i.i.d. and los channels
- Comparing Decentralized Gradient Descent Approaches and Guarantees
- Comparison of Soft and Hard Target RNN-T Distillation for Large-Scale ASR
- Compensatory Debiasing For Gender Imbalances In Language Models
- Complementary Learning System Based Intrinsic Reward in Reinforcement Learning
- Compose & Embellish: Well-Structured Piano Performance Generation via A Two-Stage Approach
- Composition of Motion from Video Animation Through Learning Local Transformations
- Comprehensive Complexity Assessment of Emerging Learned Image Compression on CPU and GPU
- Compressed Distributed Regression over Adaptive Networks
- Compressed-Sensing-Based 3D Localization with Distributed Passive Reconfigurable Intelligent Surfaces
- Compressing Cross-Domain Representation via Lifelong Knowledge Distillation
- Compressive Channel Estimation for IRS-Aided Millimeter-Wave Systems via Two-Stage Lamp Network
- Compressive Estimation of Near Field Channels for Ultra Massive-Mimo Wideband THz Systems
- Compressive Sensing with Tensorized Autoencoder
- Conditional Conformer: Improving Speaker Modulation For Single And Multi-User Speech Enhancement
- Conditional LS-GAN Based Skylight Polarization Image Restoration and Application in Meridian Localization
- Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
- Confidence-Based Event-Centric Online Video Question Answering on a Newly Constructed ATBS Dataset
- Conformer-Based Target-Speaker Automatic Speech Recognition For Single-Channel Audio
- Consistent Estimators of a New Class of Covariance Matrix Distances in the Large Dimensional Regime
- Constrained Dynamical Neural ODE for Time Series Modelling: A Case Study on Continuous Emotion Prediction
- Constrained Independent Component Analysis Based on Entropy Bound Minimization for Subgroup Identification from Multi-subject fMRI Data
- Constrained non-negative PARAFAC2 for electromyogram separation
- Content-Insensitive Dynamic Lip Feature Extraction for Visual Speaker Authentication Against Deepfake Attacks
- Context-Aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis
- Context-Aware Face Clustering with Graph Convolutional Networks
- Context-Aware Fine-Tuning of Self-Supervised Speech Models
- Context-Aware end-to-end ASR Using Self-Attentive Embedding and Tensor Fusion
- Contextual Similarity is More Valuable Than Character Similarity: An Empirical Study for Chinese Spell Checking
- Contextually-Rich Human Affect Perception Using Multimodal Scene Information
- Continilm: A Continual Learning Scheme for Non-Intrusive Load Monitoring
- Continual Cell Instance Segmentation of Microscopy Images
- Continual Learning for On-Device Speech Recognition Using Disentangled Conformers
- Continuous Action Space-Based Spoken Language Acquisition Agent Using Residual Sentence Embedding and Transformer Decoder
- Continuous Descriptor-Based Control for Deep Audio Synthesis
- Continuous Interaction with A Smart Speaker via Low-Dimensional Embeddings of Dynamic Hand Pose
- Continuous Learning for Blind Image Quality Assessment with Contrastive Transformer
- Contrast-PLC: Contrastive Learning for Packet Loss Concealment
- Contrastive Domain Adaptation Via Delimitation Discriminator
- Contrastive Learning at the Relation and Event Level for Rumor Detection
- Contrastive Learning of Functionality-Aware Code Embeddings
- Contrastive Learning of Sentence Embeddings in Product Search
- Contrastive Learning with Dialogue Attributes for Neural Dialogue Generation
- Contrastive Learning-Based Audio to Lyrics Alignment for Multiple Languages
- Contrastive Representation Learning for Acoustic Parameter Estimation
- Contrastive Self-Supervised Learning for Automated Multi-Modal Dance Performance Assessment
- Contrastive Speech Mixup for Low-Resource Keyword Spotting
- Controllable Music Inpainting with Mixed-Level and Disentangled Representation
- Convergence Analysis of Graphical Game-Based Nash Q-Learning using the Interaction Detection Signal of N-Step Return
- Convergence of Stochastic PDMM
- Conversation-Oriented ASR with Multi-Look-Ahead CBS Architecture
- Conversational Text-to-SQL: An Odyssey into State-of-the-Art and Challenges Ahead
- Convex Optimization of Deep Polynomial and ReLU Activation Neural Networks
- Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification
- Convolutional Filtering on Sampled Manifolds
- Convolutional Recurrent MetriCGAN With Spectral Dimension Compression For Full-Band Speech Enhancement
- Convolutional Recurrent Neural Networks for the Classification of Cetacean Bioacoustic Patterns
- Convolutive NTF for Ambisonic Source Separation under Reverberant Conditions
- Cooperative Five Degrees Of Freedom Motion Estimation For A Swarm Of Autonomous Vehicles
- Core: Transferable Long-Range Time Series Forecasting Enhanced by Covariates-Guided Representation
- Cosmopolite Sound Monitoring (CoSMo): A Study of Urban Sound Event Detection Systems Generalizing to Multiple Cities
- Cough Detection Using Millimeter-Wave Fmcw Radar
- Could the BubbleView Metaphor be used to Infer Visual Attention on 3D Graphical Content?
- Counterfactual Explanation for Multivariate Times Series Using A Contrastive Variational Autoencoder
- Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval
- Coupled CP Tensor Decomposition with Shared and Distinct Components for Multi-Task Fmri Data Fusion
- Cov Loss: Covariance-Based Loss for Deep Face Recognition
- Covariance Regularization for Probabilistic Linear Discriminant Analysis
- Cramér-Rao Bound on Lie Groups with Observations on Lie Groups: Application to SE(2)
- Cross Modality Knowledge Distillation for Robust Pedestrian Detection in Low Light and Adverse Weather Conditions
- Cross-Device Federated Learning for Mobile Health Diagnostics: A First Study on COVID-19 Detection
- Cross-Domain Diffusion Based Speech Enhancement for Very Noisy Speech
- Cross-Domain Learning with Normalizing Flow
- Cross-Domain Object Classification Via Successive Subspace Alignment
- Cross-Head Supervision for Crowd Counting with Noisy Annotations
- Cross-Lingual Alzheimer's Disease Detection Based on Paralinguistic and Pre-Trained Features
- Cross-Lingual Transfer Learning for Alzheimer's Detection from Spontaneous Speech
- Cross-Modal Adversarial Contrastive Learning for Multi-Modal Rumor Detection
- Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification
- Cross-Modal Fusion Techniques for Utterance-Level Emotion Recognition from Text and Speech
- Cross-Modal Matching and Adaptive Graph Attention Network for RGB-D Scene Recognition
- Cross-Modal Mutual Learning for Cued Speech Recognition
- Cross-Modal Optical Flow Estimation via Modality Compensation and Alignment
- Cross-Modality depth Estimation via Unsupervised Stereo RGB-to-infrared Translation
- Cross-Site Generalization for Imbalanced Epileptic Classification
- Cross-Speaker Emotion Transfer by Manipulating Speech Style Latents
- Cross-Subject Mental Fatigue Detection based on Separable Spatio-Temporal Feature Aggregation
- Cross-Training: A Semi-Supervised Training Scheme for Speech Recognition
- Cross-Utterance ASR Rescoring with Graph-Based Label Propagation
- CryoSWD: Sliced Wasserstein Distance Minimization for 3D Reconstruction in Cryo-electron Microscopy
- Cumulative Attention Based Streaming Transformer ASR with Internal Language Model Joint Training and Rescoring
- Customized Automatic Face Beautification
- Cutting Through the Noise: An Empirical Comparison of Psycho-Acoustic and Envelope-based Features for Machinery Fault Detection
- CyFi-TTS: Cyclic Normalizing Flow with Fine-Grained Representation for End-to-End Text-to-Speech
- CyPMLI: WISL-Minimized Unimodular Sequence Design via Power Method-Like Iterations
- D-3DLD: Depth-Aware Voxel Space Mapping for Monocular 3D Lane Detection with Uncertainty
- D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object Detection
- D2Former: A Fully Complex Dual-Path Dual-Decoder Conformer Network Using Joint Complex Masking and Complex Spectral Mapping for Monaural Speech Enhancement
- D2Q-DETR: Decoupling and Dynamic Queries for Oriented Object Detection with Transformers
- DAIS: The Delft Database of EEG Recordings of Dutch Articulated and Imagined Speech
- DASA: Difficulty-Aware Semantic Augmentation for Speaker Verification
- DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks
- DB-UNet: MLP Based Dual Branch UNet for Accurate Vessel Segmentation in OCTA Images
- DDN: Dynamic Aggregation Enhanced Dual-Stream Network for Medical Image Classification
- DEHRFormer: Real-Time Transformer for Depth Estimation and Haze Removal from Varicolored Haze Scenes
- DGN: Descriptor Generation Network for Feature Matching in Monocular Endoscopy 3D Reconstruction
- DL-NET: Dilation Location Network for Temporal Action Detection
- DMFormer: Closing the gap Between CNN and Vision Transformers
- DMSA: Dynamic Multi-Scale Unsupervised Semantic Segmentation Based On Adaptive Affinity
- DO-FAM: Disentangled Non-Linear Latent Navigation For Facial Attribute Manipulation
- DPP-Based Client Selection for Federated Learning with NON-IID DATA
- DQFORMER: Dynamic Query Transformer for Lane Detection
- DRL Path Planning for UAV-Aided V2X Networks: Comparing Discrete to Continuous Action Spaces
- DSPGAN: A Gan-Based Universal Vocoder for High-Fidelity TTS by Time-Frequency Domain Supervision from DSP
- DST: Deformable Speech Transformer for Emotion Recognition
- DTTR: Detecting Text with Transformers
- DVQVC: An Unsupervised Zero-Shot Voice Conversion Framework
- DWFormer: Dynamic Window Transformer for Speech Emotion Recognition
- Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis
- DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech
- Dasformer: Deep Alternating Spectrogram Transformer For Multi/Single-Channel Speech Separation
- Data Augmentation Based On Invariant Shape Blending For Deep Learning Classification
- Data Driven Joint Sensor Fusion and Regression Based on Geometric Mean Squared Error
- Data Leakage in Cross-Modal Retrieval Training: A Case Study
- Data-Aware Zero-Shot Neural Architecture Search for Image Recognition
- Data-Driven Graph Convolutional Neural Networks for Power System Contingency Analysis
- Data-Driven Quickest Change Detection in Markov Models
- Data2vec-Aqc: Search for the Right Teaching Assistant in the Teacher-Student Training Setup
- Database-Aware ASR Error Correction for Speech-to-SQL Parsing
- Dataset Balancing Can Hurt Model Performance
- De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition
- Decaying Contrast for Fine-Grained Video Representation Learning
- Decoding Auditory EEG Responses Using an Adapted Wavenet
- Decoding Musical Pitch from Human Brain Activity with Automatic Voxel-Wise Whole-Brain FMRI Feature Selection
- DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene Classification
- Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking
- Decontamination Transformer For Blind Image Inpainting
- Decorrelating Language Model Embeddings for Speech-Based Prediction of Cognitive Impairment
- Decoupled Non-Parametric Knowledge Distillation for end-to-End Speech Translation
- Decoupled Visual Causality for Robust Detection
- Deep AHS: A Deep Learning Approach to Acoustic Howling Suppression
- Deep Adaptive Superpixels For Hadamard Single Pixel Imaging In Near-Infrared Spectrum
- Deep Architecture for DOA Trajectory Localization
- Deep Autoencoding One-Class time Series Anomaly Detection
- Deep Born Operator Learning for Reflection Tomographic Imaging
- Deep Double Self-Expressive Subspace Clustering
- Deep Feature Aggregation for Lightweight Single Image Super-Resolution
- Deep Fusion of Multi-Object Densities Using Transformer
- Deep Generative Fixed-Filter Active Noise Control
- Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion Recognition
- Deep Learning Sparse Array Design Using Binary Switching Configurations
- Deep Learning for Lagrangian Drift Simulation at The Sea Surface
- Deep Learning-Based Compressive Sampling Optimization in Massive MIMO Systems
- Deep Learning-Based Path Loss Prediction for Outdoor Wireless Communication Systems
- Deep Learning-Based Stereo Camera Multi-Video Synchronization
- Deep Low Light Image Enhancement Via Multi-Scale Recursive Feature Enhancement and Curve Adjustment
- Deep Manifold Graph Auto-Encoder For Attributed Graph Embedding
- Deep Network Series for Large-Scale High-Dynamic Range Imaging
- Deep Neural Mel-Subband Beamformer for in-Car Speech Separation
- Deep Plug-and-Play for Tensor Robust Principal Component Analysis
- Deep Probabilistic Model for Lossless Scalable Point Cloud Attribute Compression
- Deep Proximal Gradient Method for Learned Convex Regularizers
- Deep Quantigraphic Image Enhancement via Comparametric Equations
- Deep Reinforcement Learning for Green UAV-Assisted Data Collection
- Deep Root Music Algorithm for Data-Driven Doa Estimation
- Deep Spatio-Temporal Multiplex Graph Learning for Cardiac Imaging Classification
- Deep Spectrum Cartography Using Quantized Measurements
- Deep Subband Network for Joint Suppression of Echo, Noise and Reverberation in Real-Time Fullband Speech Communication
- Deep Survival Analysis and Counterfactual Inference Using Balanced Representations
- Deep Triple-Supervision Learning Unannotated Surgical Endoscopic Video Data for Monocular Dense Depth Estimation
- Deep Unfolded Tensor Robust PCA With Self-Supervised Learning
- Deep Unfolding-Enabled Hybrid Beamforming Design for mmWave Massive MIMO Systems
- Deep-Unfolded Adaptive Projected Subgradient Method For Mimo Detection
- Deep3DSketch: 3D Modeling from Free-Hand Sketches with View- and Structural-Aware Adversarial Training
- Deepspace: Dynamic Spatial and Source CUE Based Source Separation for Dialog Enhancement
- Defending Against Universal Patch Attacks by Restricting Token Attention in Vision Transformers
- Defense Against Black-Box Adversarial Attacks Via Heterogeneous Fusion Features
- Deformable Cross Attention for Learning Optical Flow
- Deformable Temporal Convolutional Networks for Monaural Noisy Reverberant Speech Separation
- Delay-Aware Backpressure Routing Using Graph Neural Networks
- Delay-Penalized Transducer for Low-Latency Streaming ASR
- Delivering Speaking Style in Low-Resource Voice Conversion with Multi-Factor Constraints
- Dense Adversarial Transfer Learning Based On Class-Invariance
- Densitytoken: Weakly-Supervised Crowd Counting with Density Classification
- Depth Estimation for a Single Omnidirectional Image with Reversed-Gradient Warming-up Thresholds Discriminator
- DepthFormer: Multimodal Positional Encodings and Cross-Input Attention for Transformer-based Segmentation Networks
- Dereverberation in Acoustic Sensor Networks Using weighted Prediction Error with Microphone-Dependent Prediction Delays
- Design Choices for Learning Embeddings from Auxiliary Tasks for Domain Generalization in Anomalous Sound Detection
- Design and Performance of the Low-Power Noise Reduction Algorithm of the Med-El Sonnet 2™ Cochlear Implant Audio Processor
- Designing A 3d-Aware Stylenerf Encoder for Face Editing
- Designing Transformer Networks for Sparse Recovery of Sequential Data Using Deep Unfolding
- Designing and Evaluating Speech Emotion Recognition Systems: A Reality Check Case Study with IEMOCAP
- Detail-Aware Uncalibrated Photometric Stereo
- Detecting Malicious Migration on Edge to Prevent Running Data Leakage
- Detecting Out-of-Distribution Examples Via Class-Conditional Impressions Reappearing
- Detection of Real-Time Deepfakes in Video Conferencing with Active Probing and Corneal Reflection
- Dewarping Documents Using C2 Continuous Boundary Estimation
- Diabetic Retinopathy Grading with Weakly-Supervised Lesion Priors
- Diagonal State Space Augmented Transformers for Speech Recognition
- Dialog Act Guided Contextual Adapter for Personalized Speech Recognition
- DialogMI: A Dialogue Model Based on Enhancing Dialogue Mutual Information
- Dialogue Context Modelling for Action Item Detection: Solution for ICASSP 2023 Mug Challenge Track 5
- Dialogue System with Missing Observation
- Dictionary Learning on Graph Data with Weisfieler-Lehman Sub-Tree Kernel and Ksvd
- DiffPhase: Generative Diffusion-Based STFT Phase Retrieval
- DiffVoice: Text-to-Speech with Latent Diffusion
- Difference Coarrays of Rational Arrays
- Difference Guided VHR Remote Sensing Image Change Detection
- Differentiable Adaptive Short-Time Fourier Transform with Respect to the Window Length
- Differential Analysis for Networks Obeying Conservation Laws
- Difficulty-Aware Data Augmentor for Scene Text Recognition
- Diffroll: Diffusion-Based Generative Music Transcription with Unsupervised Pretraining Capability
- Diffusion Motion: Generate Text-Guided 3D Human Motion by Diffusion Model
- Diffusion Probabilistic Modeling for Fine-Grained Urban Traffic Flow Inference with Relaxed Structural Constraint
- Diffusion-Based Generative Speech Source Separation
- Diffusion-Based Sound Source Localization Using Networks of Planar Microphone Arrays
- Diffusionnet: An Efficient Framework to Classify Single-Molecule Images with Latent Entropy Minimization
- Digital Phenotype Representation by Statistical, Information Theory, Data-Driven Approach with Digital Health Data
- Direct Position Determination with One-Bit Signal for Multiple Targets
- Direction Aware Positional and Structural Encoding for Directed Graph Neural Networks
- Direction-of-Arrival Estimation Using Gaussian Process Interpolation
- DisCoHead: Audio-and-Video-Driven Talking Head Generation by Disentangled Control of Head Pose and Facial Expressions
- Disambiguation of Cognitive Impairment Diagnosis with EEG-Based Dual-Contrastive Learning
- Discriminative Speaker Representation Via Contrastive Learning with Class-Aware Attention in Angular Space
- Discriminative Vector Learning with Application to Single Channel Speech Separation
- Disentangled Feature Learning for Real-Time Neural Speech Coding
- Disentangled Training with Adversarial Examples for Robust Small-Footprint Keyword Spotting
- Disentangled and Robust Representation Learning for Bragging Classification in Social Media
- Disentangling Speech from Surroundings with Neural Embeddings
- Disentangling the Horowitz Factor: Learning Content and Style From Expressive Piano Performance
- Distance-Based Online Label Inference Attacks Against Split Learning
- Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker Verification
- Distill-Quantize-Tune - Leveraging Large Teachers for Low-Footprint Efficient Multilingual NLU on Edge
- Distinguishable Speaker Anonymization Based on Formant and Fundamental Frequency Scaling
- Distortion-Aware Convolutional Neural Network-Based Interpolation Filter for AVS3
- Distributed Adaptive Norm Estimation for Blind System Identification in Wireless Sensor Networks
- Distributed Admm with Limited Communications Via Deep Unfolding
- Distributed Bayesian Tracking on the Special Euclidean Group Using Lie Algebra Parametric Approximations
- Distributed Gaussian Process Hyperparameter Optimization for Multi-Agent Systems
- Distributed Online Learning With Adversarial Participants In An Adversarial Environment
- Distributed Quantum Sensing Network with Geographically Constrained Measurement Strategies
- Distributed Signal Processing for Out-of-System Interference Suppression in Cell-Free Massive MIMO
- Distributionally Robust Multiclass Classification and Applications in Deep Image Classifiers
- Divcon: Learning Concept Sequences for Semantically Diverse Image Captioning
- Diverse and Vivid Sound Generation from Text Descriptions
- Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm Regularization
- Do Coarser Units Benefit Cluster Prediction-Based Speech Pre-Training?
- Do Prosody Transfer Models Transfer Prosodyƒ
- DocRED-FE: A Document-Level Fine-Grained Entity and Relation Extraction Dataset
- Does Human Speech Follow Benford's Law?
- Does Your Model Think Like an Engineer? Explainable AI for Bearing Fault Detection with Deep Learning
- Does a Quieter City Mean Fewer Complaints? The Sounds of New York City During Covid-19 Lockdown
- Domain Adaptation with External Off-Policy Acoustic Catalogs for Scalable Contextual End-to-End Automated Speech Recognition
- Domain Adaptation without Catastrophic Forgetting on a Small-Scale Partially-Labeled Corpus for Speech Emotion Recognition
- Domain Generalized Fundus Image Segmentation via Dual-Level Mixing
- Domain and Language Adaptation Using Heterogeneous Datasets for Wav2vec2.0-Based Speech Recognition of Low-Resource Language
- Doppler-Coded Joint Division Multiple Access Waveform for Automotive MIMO Radar
- Double Compression Detection Based on the De-Blocking Filtering of HEVC Videos
- Downlink Covariance Estimation in URA FDD Massive MIMO Systems
- Drone-vs-Bird Detection Grand Challenge at ICASSP2023
- Drone-vs-Bird: Drone Detection Using YOLOv7 with CSRT Tracker
- Dual Collaborative Visual-Semantic Mapping for Multi-Label Zero-Shot Image Recognition
- Dual Meta Calibration Mix for Improving Generalization in Meta-Learning
- Dual Path Modeling for Semantic Matching by Perceiving Subtle Conflicts
- Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition
- Dual-Based Online Learning of Dynamic Network Topologies
- Dual-Cycle: Self-Supervised Dual-View Fluorescence Microscopy Image Reconstruction using CycleGAN
- Dual-Feature Enhancement for Weakly Supervised Temporal Action Localization
- Dual-Head Fusion Network for Image Enhancement
- Dual-Path Cross-Modal Attention for Better Audio-Visual Speech Extraction
- Dual-Path Dilated Convolutional Recurrent Network with Group Attention for Multi-Channel Speech Enhancement
- Dual-Stage Graph Convolution Network With Graph Learning For Traffic Prediction
- Dual-Stream Siamese Vision Transformer With Mutual Attention For Radar Gait Verification
- Dual-Uncertainty Guided Curriculum Learning and Part-Aware Feature Refinement for Domain Adaptive Person Re-Identification
- Dual-Use Signal Design for MIMO Radcom with Inter-Pulse Index Modulation
- Dual-graph co-representation learning for knowledge-Graph Enhanced Recommendation
- Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech
- DyLiteRADHAR: Dynamic Lightweight Slowfast Network for Human Activity Recognition Using MMWAVE Radar
- Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross Entropy
- Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR
- Dynamic Distributed Convex Optimization "Over-The-Air" In Decentralized Wireless Networks
- Dynamic Fair Node Representation Learning
- Dynamic Independent Component Extraction with Blending Mixing Vector: Lower Bound on Mean Interference-to-Signal Ratio
- Dynamic Local and Global Context Exploration for Small Object Detection
- Dynamic Multi-View Scene Reconstruction Using Neural Implicit Surface
- Dynamic Scalable Self-Attention Ensemble for Task-Free Continual Learning
- Dynamic Selection of p-norm in Linear Adaptive Filtering via online Kernel-based Reinforcement Learning
- Dynamic Signed Graph Learning
- Dynamic Speech Endpoint Detection with Regression Targets
- Dynamic Split Computing for Efficient Deep EDGE Intelligence
- Dynamic TF-TDNN: Dynamic Time Delay Neural Network Based on Temporal-Frequency Attention for Dialect Recognition
- Dynamic Vehicle Graph Interaction for Trajectory Prediction Based on Video Signals
- E-Branchformer-Based E2E SLU Toward Stop on-Device Challenge
- E-Prevention: The ICASSP-2023 Challenge on Person Identification and Relapse Detection from Continuous Recordings of Biosignals
- E2E Segmentation in a Two-Pass Cascaded Encoder ASR Model
- EBEN: Extreme Bandwidth Extension Network Applied To Speech Signals Captured With Noise-Resilient Body-Conduction Microphones
- ECG Artifact Removal from Single-Channel Surface EMG Using Fully Convolutional Networks
- ECGT2T: Towards Synthesizing Twelve-Lead Electrocardiograms from Two Asynchronous Leads
- EEG Emotion Recognition Via Ensemble Learning Representations
- EEG2IMAGE: Image Reconstruction from EEG Brain Signals
- EGAN: A Neural Excitation Generation Model Based on Generative Adversarial Networks with Harmonics and Noise Input
- EH-Enabled Distributed Detection Over Temporally Correlated Markovian MIMO Channels
- EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text Detection
- EMC2-Net: Joint Equalization and Modulation Classification Based on Constellation Network
- EMCLR: Expectation Maximization Contrastive Learning Representations
- EMIX: A Data Augmentation Method for Speech Emotion Recognition
- ERBNet: An Effective Representation Based Network for Unbiased Scene Graph Generation
- ERSAM: Neural Architecture Search for Energy-Efficient and Real-Time Social Ambiance Measurement
- ESCL: Equivariant Self-Contrastive Learning for Sentence Representations
- Early Detection of Cognitive Decline Using Voice Assistant Commands
- Effect of Lossy Compression Algorithms on Face Image Quality and Recognition
- Effective Graph-Based Modeling of Articulation Traits for Mispronunciation Detection and Diagnosis
- Effective Training of RNN Transducer Models on Diverse Sources of Speech and Text Data
- Effectiveness of Inter- and Intra-Subarray Spatial Features for Acoustic Scene Classification
- Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages
- Effectiveness of Text, Acoustic, and Lattice-Based Representations in Spoken Language Understanding Tasks
- Efficent Large-Scale Multi-Unimodular Waveform Design with Good Correlation Properties via Direct Phase Optimizations
- Efficient Compressed Video Action Recognition Via Late Fusion with a Single Network
- Efficient Data Loading with Quantum Autoencoder
- Efficient Domain Adaptation for Speech Foundation Models
- Efficient Feature Extraction for Non-Maximum Suppression in Visual Person Detection
- Efficient Feature Fusion for Learning-Based Photometric Stereo
- Efficient Implementation of Robust CUSUM Algorithm to Characterize Nanogaps Measurements with Heavy-Tailed Noise
- Efficient Intelligibility Evaluation Using Keyword Spotting: A Study on Audio-Visual Speech Enhancement
- Efficient Large-Scale Audio Tagging Via Transformer-to-CNN Knowledge Distillation
- Efficient Learning of Balanced Signature Graphs
- Efficient Monaural Speech Enhancement with Universal Sample Rate Band-Split RNN
- Efficient Multi-Scale Attention Module with Cross-Spatial Learning
- Efficient Online Convolutional Dictionary Learning Using Approximate Sparse Components
- Efficient Personalized Federated Learning on Selective Model Training
- Efficient Practices for Profile-to-Frontal Face Synthesis and Recognition
- Efficient Privacy Preserving Graph Neural Network for Node Classification
- Efficient Protein Structural Class Prediction Via Chaos Game Representation and Recurrent Neural Networks
- Efficient Quantized Constant Envelope Precoding for Multiuser Downlink Massive MIMO Systems
- Efficient Siamese Network for UAV Tracking
- Efficient Similarity-Based Passive Filter Pruning for Compressing CNNS
- Efficient Speech Quality Assessment Using Self-Supervised Framewise Embeddings
- Efficient Speech Translation with Dynamic Latent Perceivers
- Efficient Stuttering Event Detection Using Siamese Networks
- Efficient Super-Resolution for Compression Of Gaming Videos
- Efficient Uncertainty Estimation with Gaussian Process for Reliable Dialog Response Retrieval
- Efficient and Effective Multi-Camera Pose Estimation with Weighted M-Estimate Sample Consensus
- EfficientSpeech: An On-Device Text to Speech Model
- Efficiently Fusing Sparse Lidar for Enhanced Self-Supervised Monocular Depth Estimation
- Egocentric Action Anticipation for Personal Health
- Egocentric Audio-Visual Noise Suppression
- Eigen-Decomposition-Free Directed Graph Sampling via Gershgorin Disc Alignment
- Elastic Graph Transformer Networks for EEG-Based Emotion Recognition
- Electric Network Frequency Detection Using Least Absolute Deviations
- Element Selection with Wide Class of Optimization Criteria Using Non-Convex Sparse Optimization
- Elliptical Wishart Distribution: Maximum Likelihood Estimator from Information Geometry
- Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System
- Embrace Smaller Attention: Efficient Cross-Modal Matching with Dual Gated Attention Fusion
- Emodiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
- Emotion Recognition in Conversation from Variable-Length Context
- Empathetic Response Generation via Emotion Cause Transition Graph
- Enabling Large-Scale Image Search with Co-Attention Mechanism
- Encoder-Decoder Graph Convolutional Network for Automatic Timed-Up-and-Go and Sit-to-Stand Segmentation
- End-to-End Amp Modeling: from Data to Controllable Guitar Amplifier Models
- End-to-End Classification of Cell-Cycle Stages with Center-Cell Focus Tracker Using Recurrent Neural Networks
- End-to-End Neural Audio Coding in the MDCT Domain
- End-to-End Non-Autoregressive Image Captioning
- End-to-End Spoken Language Understanding Using Joint CTC Loss and Self-Supervised, Pretrained Acoustic Encoders
- End-to-End Spoken Language Understanding with Tree-Constrained Pointer Generator
- End-to-End Unsupervised Sketch to Image Generation
- End-to-End Word-Level Disfluency Detection and Classification in Children's Reading Assessment
- Energy Efficiency Maximization in RIS-aided Networks with Global Reflection Constraints
- Energy Regularized RNNS for solving non-stationary Bandit problems
- Enhance Transferability of Adversarial Examples with Model Architecture
- Enhanced Coprime Array Configuration for DoA Estimation of Non-Circular Signals
- Enhanced Dcf Tracker Regularized by Reliable Sample Construction
- Enhanced Embeddings in Zero-Shot Learning for Environmental Audio
- Enhanced GM-PHD Filter for Real Time Satellite Multi-Target Tracking
- Enhanced Low-Resolution LiDAR-Camera Calibration via Depth Interpolation and Supervised Contrastive Learning
- Enhancement of Text-Predicting Style Token With Generative Adversarial Network for Expressive Speech Synthesis
- Enhancing Multimodal Alignment with Momentum Augmentation for Dense Video Captioning
- Enhancing Ontology Translation Through Cross-Lingual Agreement
- Enhancing Representation Learning with Deep Classifiers in Presence of Shortcut
- Enhancing Robustness and Imperceptibility of Blind Watermarking with Improved Message Processor
- Enhancing Spatio-Spectral Regularization by Structure Tensor Modeling for Hyperspectral Image Denoising
- Enhancing Speech-To-Speech Translation with Multiple TTS Targets
- Enhancing Unsupervised Speech Recognition with Diffusion GANS
- Enhancing and Adversarial: Improve ASR with Speaker Labels
- Enhancing the Accuracy of Resistive In-Memory Architectures using Adaptive Signal Processing
- Enhancing the Efficiency of WMMSE and FP for Beamforming by Minorization-Maximization
- Enhancing the Vocal Range of Single-Speaker Singing Voice Synthesis with Melody-Unsupervised Pre-Training
- Enlightening the Student in Knowledge Distillation
- Enrollment Rate Prediction in Clinical Trials based on CDF Sketching and Tensor Factorization tools
- Ensemble Graph Q-Learning for Large Scale Networks
- Ensemble Knowledge Distillation of Self-Supervised Speech Models
- Ensemble Prosody Prediction For Expressive Speech Synthesis
- Ensemble and Personalized Transformer Models for Subject Identification and Relapse Detection in E-Prevention Challenge
- Ensemble of Deep Neural Network Models for MOS Prediction
- Entropy Based Feature Regularization to Improve Transferability of Deep Learning Models
- Epic-Sounds: A Large-Scale Dataset of Actions that Sound
- Epilepsy Detection Grand Challenge
- Equivalence of Aperture Reduction in Element Space and Constrained Combination of DFT Beams in Beamspace
- Error Analysis of Convolutional Beamspace Algorithms
- Estimating Acoustic Direction of Arrival Using a Single Structural Sensor on a Resonant Surface
- Estimating Inharmonic Signals with Optimal Transport Priors
- Estimating Normalized Graph Laplacians in Financial Markets
- Estimating Shapley Values of Training Utterances for Automatic Speech Recognition Models
- Estimating Uncertainty On Video Quality Metrics
- Estimating and Analyzing Neural Information flow using Signal Processing on Graphs
- Estimation of Cardiac Fibre Direction Based on Activation Maps
- Estimation of High-Dimensional Differential Graphs from Multi-Attribute Data
- Estimation of Time-Varying Graph Topologies from Graph Signals
- Estimation of Visual Contents from Human Brain Signals via VQA Based on Brain-Specific Attention
- Euro: Espnet Unsupervised ASR Open-Source Toolkit
- Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding
- Evaluating Speech-Phoneme Alignment and its Impact on Neural Text-To-Speech Synthesis
- Evaluating Variants of wav2vec 2.0 on Affective Vocal Burst Tasks
- Evaluation of Categorical Generative Models - Bridging the Gap Between Real and Synthetic Data
- Event-Based Visual Microphone
- Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
- Evopose: A Recursive Transformer for 3D Human Pose Estimation with Kinematic Structure Priors
- Expectation Propagation on Factor Graphs Based on Matrix Decomposition
- Explainable audio Classification of Playing Techniques with Layer-wise Relevance Propagation
- Explanations for Automatic Speech Recognition
- Explicit Ziv-Zakai Bound For Multiple Sources Doa Estimation
- Explicit and Implicit Knowledge Distillation via Unlabeled Data
- Exploiting 3D Human Recovery for Action Recognition with Spatio-Temporal Bifurcation Fusion
- Exploiting CCTV Cameras for Hand Hygiene Recognition in ICU
- Exploiting Interactivity and Heterogeneity for Sleep Stage Classification Via Heterogeneous Graph Neural Network
- Exploiting Modality-Invariant Feature for Robust Multimodal Emotion Recognition with Missing Modalities
- Exploiting Multi-Decision and Deep Refinement for Ultrasound Image Segmentation
- Exploiting One-Class Classification Optimization Objectives for Increasing Adversarial Robustness
- Exploiting PRNU and Linear Patterns in Forensic Camera Attribution under Complex Lens Distortion Correction
- Exploiting Prompt Learning with Pre-Trained Language Models for Alzheimer's Disease Detection
- Exploiting Sparse Recovery Algorithms for Semi-Supervised Training of Deep Neural Networks for Direction-of-Arrival Estimation
- Exploiting Spatial Information with the Informed Complex-Valued Spatial Autoencoder for Target Speaker Extraction
- Exploiting Speaker Embeddings for Improved Microphone Clustering and Speech Separation in ad-hoc Microphone Arrays
- Exploiting Virtual Array Diversity for Accurate Radar Detection
- Exploration Into Translation-Equivariant Image Quantization
- Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models
- Exploring Approaches to Multi-Task Automatic Synthesizer Programming
- Exploring Attention Mechanisms for Multimodal Emotion Recognition in an Emergency Call Center Corpus
- Exploring Binary Classification Loss for Speaker Verification
- Exploring Complementary Features in Multi-Modal Speech Emotion Recognition
- Exploring Instance Relation for Decentralized Multi-Source Domain Adaptation
- Exploring Language-Agnostic Speech Representations Using Domain Knowledge for Detecting Alzheimer's Dementia
- Exploring Progressive Hybrid-Degraded Image Processing for Homography Estimation
- Exploring Self-Supervised Pre-Trained ASR Models for Dysarthric and Elderly Speech Recognition
- Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting
- Exploring Subgroup Performance in End-to-End Speech Models
- Exploring Universal Singing Speech Language Identification Using Self-Supervised Learning Based Front-End Features
- Exploring Vision Transformer Layer Choosing for Semantic Segmentation
- Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition
- Exploring the Role of Fricatives in Classifying Healthy Subjects and Patients with Amyotrophic Lateral Sclerosis and Parkinson's Disease
- Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features
- Extended Expectation Maximization for Under-Fitted Models
- Extended Kalman Filter for Graph Signals in Nonlinear Dynamic Systems
- Extracting the Brain-Like Representation by an Improved Self-Organizing Map for Image Classification
- Extreme Audio Time Stretching Using Neural Synthesis
- F-PABEE: Flexible-Patience-Based Early Exiting For Single-Label and Multi-Label Text Classification Tasks
- F0 Estimation From Telephone Speech Using Deep Feature Loss
- FAPM: Fast Adaptive Patch Memory for Real-Time Industrial Anomaly Detection
- FCIR: Rethink Aerial Image Super Resolution with Fourier Analysis
- FED-3DA: A Dynamic and Personalized Federated Learning Framework
- FEW-Shot Continual Learning with Weight Alignment and Positive Enhancement for Bioacoustic Event Detection
- FFEDCL: Fair Federated Learning with Contrastive Learning
- FFFN: Fashion Feature Fusion Network by Co-Attention Model for Fashion Recommendation
- FNeural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated full- and sub-band Modeling
- Face Recognition on Point Cloud with Cgan-Top for Denoising
- Facial Texure Perceiver: Towards High-Fidelity Facial Texture Recovery with Input-Level Inductive Biased Perceiver IO
- Factorized AED: Factorized Attention-Based Encoder-Decoder for Text-Only Domain Adaptive ASR
- Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers
- Factorized Projection-Domain Spatio-Temporal Regularization for Dynamic Tomography
- False Alarm Regulation for Off-Grid Target Detection With The Matched Filter
- Fan-Net: Fourier-Based Adaptive Normalization for Cross-Domain Stroke Lesion Segmentation
- Fast 3D Human Pose Estimation Using RF Signals
- Fast Convolution Algorithm for Real-Valued Finite Length Sequences
- Fast Cross-Correlation for TDoA Estimation on Small Aperture Microphone Arrays
- Fast Low-Latency Convolution by Low-Rank Tensor Approximation
- Fast Multiscale 3D Reconstruction Using Single-Photon Lidar Data
- Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector Analysis
- Fast Robust Principle Component Analysis Using Gauss-Newton Iterations
- Fast Single-Person 2D Human Pose Estimation Using Multi-Task Convolutional Neural Networks
- Fast Yet Effective Speech Emotion Recognition with Self-Distillation
- Fast and Accurate Factorized Neural Transducer for Text Adaption of End-to-End Speech Recognition Models
- Fast and Efficient Speech Enhancement with Variational Autoencoders
- Fast and Exact Enumeration of Deep Networks Partitions Regions
- Fast and Parallel Decoding for Transducer
- Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
- Faster Than Fast: Accelerating the Griffin-Lim Algorithm
- Feature Selection and Text Embedding for Detecting Dementia from Spontaneous Cantonese
- Feature Space Recovery for Incomplete Multi-View Clustering
- Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound Classification
- FedAudio: A Federated Learning Benchmark for Audio Tasks
- FedEEG: Federated EEG Decoding Via inter-Subject Structure Matching
- FedPrompt: Communication-Efficient and Privacy-Preserving Prompt Tuning in Federated Learning
- FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event Classification
- FedSD: A New Federated Learning Structure Used in Non-iid Data
- FedVMR: A New Federated Learning Method for Video Moment Retrieval
- Federated Intelligent Terminals Facilitate Stuttering Monitoring
- Federated Learning for ASR Based on wav2vec 2.0
- Federated Self-Learning with Weak Supervision for Speech Recognition
- Federated Semi-Supervised Learning for Object Detection in Autonomous Driving
- Few but Informative Local Hash Code Matching for Image Retrieval
- Filler Word Detection with Hard Category Mining and Inter-Category Focal Loss
- Filter Pruning Via Filters Similarity in Consecutive Layers
- Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting
- FindAdaptNet: Find and Insert Adapters by Learned Layer Importance
- Finding Optimal Numerical Format for Sub-8-Bit Post-Training Quantization of Vision Transformers
- Fine-Grained Blind Face Inpainting with 3D Face Component Disentanglement
- Fine-Grained Emotional Control of Text-to-Speech: Learning to Rank Inter- and Intra-Class Emotion Intensities
- Fine-Grained Private Knowledge Distillation
- Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and Understanding
- Finer-Grained Decomposition for Parallel Quantum Mimo Processing
- Fixed-Point Quantization Aware Training for on-Device Keyword-Spotting
- Flexible Beam Design for Vital Sign Monitoring Using a Phased Array Equipped With Double-Phase Shifters
- Flow-Guided Deformable Alignment Network with Self-Supervision for Video Inpainting
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.