← All conferences

ICASSP 2023 Accepted Papers

The full list of 2,718 papers accepted at ICASSP 2023 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. "Prediction of Sleepiness Ratings from Voice by Man and Machine": A Perceptual Experiment Replication Study
  2. -Complexity Low-Rank Approximation SVD for Massive Matrix in Tensor Train Format
  3. 2DSBG: A 2d Semi Bi-Gaussian Filter Adapted for Adjacent and Multi-Scale Line Feature Detection
  4. 3D Audio Signal Processing Systems for Speech Enhancement and Sound Localization and Detection
  5. 3D Point Cloud Completion Based on Multi-Scale Degradation
  6. 6G Integrated Sensing and Communication - Sensing Assisted Environmental Reconstruction and Communication
  7. A 3D-Assisted Framework to Evaluate the Quality of Head Motion Replication by Reenactment DEEPFAKE Generators
  8. A Bandit Online Convex Optimization Approach To Distributed Energy Management In Networked Systems
  9. A Bayesian Perspective for Determinant Minimization Based Robust Structured Matrix Factorization
  10. A Bayesian Perspective on Noise2Noise: Theory and Extensions
  11. A Benchmark for Evaluating Robustness of Spoken Language Understanding Models in Slot Filling
  12. A Bidirectional Joint Model for Spoken Language Understanding
  13. A Causal Convolutional Approach for Packet Loss Concealment in Low Powered Devices
  14. A Closer Look At Scoring Functions And Generalization Prediction
  15. A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at Scale
  16. A Compensated Shrinkage Affine Projection Algorithm for Debiased Sparse Adaptive Filtering
  17. A Comprehensive Comparison of Projections in Omnidirectional Super-Resolution
  18. A Computationally Efficient Algorithm for Distributed Adaptive Signal Fusion Based on Fractional Programs
  19. A Content Adaptive Learnable "Time-Frequency" Representation for audio Signal Processing
  20. A Content-Based Multi-Scale Network for Single Image Super-Resolution
  21. A Context-Aware Computational Approach for Measuring Vocal Entrainment in Dyadic Conversations
  22. A Contrastive Embedding-Based Domain Adaptation Method for Lung Sound Recognition in Children Community-Acquired Pneumonia
  23. A Contrastive Framework to Enhance Unsupervised Sentence Representation Learning
  24. A Contrastive Knowledge Transfer Framework for Model Compression and Transfer Learning
  25. A Controllable Lifestyle Simulator for Use in Deep Reinforcement Learning Algorithms
  26. A Critical Look at Recent Trends in Compression of Channel State Information
  27. A DNN Based Normalized Time-Frequency Weighted Criterion for Robust Wideband DoA Estimation
  28. A DNN-Based Hearing-Aid Strategy For Real-Time Processing: One Size Fits All
  29. A Database for Multi-Modal Short Video Quality Assessment
  30. A Dataset for Audio-Visual Sound Event Detection in Movies
  31. A Deep Disentangled Approach for Interpretable Hyperspectral Unmixing
  32. A Deep Fusion Rule for Infrared and Visible Image Fusion: Feature Communication for Importance Assessment
  33. A Deep Temporal Factor Analysis Method for Large Scale Financial Portfolio Selection
  34. A Discriminative Multi-Channel Noise Feature Representation Model for Image Manipulation Localization
  35. A Distributed Adaptive Algorithm for Non-Smooth Spatial Filtering Problems
  36. A Dual-Branch Adaptive Distribution Fusion Framework for Real-World Facial Expression Recognition
  37. A Dual-Path Transformer Network for Scene Text Detection
  38. A Dynamic Cross-Scale Transformer with Dual-Compound Representation for 3D Medical Image Segmentation
  39. A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding
  40. A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville Distribution
  41. A Few Shot Learning of Singing Technique Conversion Based on Cycle Consistency Generative Adversarial Networks
  42. A Flow-Guided Non-Local Alignment Network for Video Compressive Sensing Reconstruction
  43. A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement
  44. A Frequency-Domain Recursive Least-Squares Adaptive Filtering Algorithm Based On A Kronecker Product Decomposition
  45. A Frequency-Weighted Leaky Fxlms Algorithm with Application to Feedback Active Noise Control Systems
  46. A Fusion-Based and Multi-Layer Method for Low Light Image Enhancement
  47. A Game of Snakes and Gans
  48. A Gaussian Latent Variable Model for Incomplete Mixed Type Data
  49. A Generalized Subspace Distribution Adaptation Framework for Cross-Corpus Speech Emotion Recognition
  50. A Geometric Surrogate for Simulation Calibration
  51. A Graph Neural Network Multi-Task Learning-Based Approach for Detection and Localization of Cyberattacks in Smart Grids
  52. A Hierarchical Regression Chain Framework for Affective Vocal Burst Recognition
  53. A Highly Interpretable Deep Equilibrium Network for Hyperspectral Image Deconvolution
  54. A Holistic Cascade System, Benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation
  55. A Hybrid Deep Neural Network for Nonlinear Causality Analysis in Complex Industrial Control System
  56. A Knowledge-Driven Vowel-Based Approach of Depression Classification from Speech Using Data Augmentation
  57. A Large-Scale Pretrained Deep Model for Phishing URL Detection
  58. A Learnable Spatial Mapping for Decoding the Directional Focus of Auditory Attention Using EEG
  59. A Lightweight Convolutional Neural Network using Feature Filtering Module
  60. A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech Enhancement
  61. A Low-Latency Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation
  62. A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids
  63. A Magnetic Framelet-Based Convolutional Neural Network for Directed Graphs
  64. A Mathematical Model for Neuronal Activity and Brain Information Processing Capacity
  65. A Memory-Free Evolving Bipolar Neural Network for Efficient Multi-Label Stream Learning
  66. A Meta-Gnn Approach to Personalized Seizure Detection and Classification
  67. A Method of Constructing and Automatically Labeling Radio Frequency Signal Training Dataset for UAV
  68. A Model-Based Hearing Compensation Method Using a Self-Supervised Framework
  69. A Momentum Two-Gradient Direction Algorithm with Variable Step Size Applied to Solve Practical Output Constraint Issue for Active Noise Control
  70. A Multi-Channel Aggregation Framework for Object Detection in Large-Scale SAR Image
  71. A Multi-Modal Approach For Context-Aware Network Traffic Classification
  72. A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement
  73. A Multi-Signal Perception Network for Textile Composition Identification
  74. A Multi-Stage Hierarchical Relational Graph Neural Network for Multimodal Sentiment Analysis
  75. A Multi-Stage Low-Latency Enhancement System for Hearing Aids
  76. A Multi-Stage Triple-Path Method For Speech Separation in Noisy and Reverberant Environments
  77. A Mutual Implicit Sentiment Analysis Model with Bundle-Aware Contrastive Learning
  78. A Nested Ensemble Method to Bilevel Machine Learning
  79. A New Approach to Extract Fetal Electrocardiogram Using Affine Combination of Adaptive Filters
  80. A New Personalized Efficacy Atlas for Pallidal Deep Brain Stimulation
  81. A New Probabilistic Distance Metric with Application in Gaussian Mixture Reduction
  82. A New Semi-Supervised Classification Method Using a Supervised Autoencoder for Biomedical Applications
  83. A Novel Approach Based on Voronoï Cells to Classify Spectrogram Zeros of Multicomponent Signals
  84. A Novel Cross-Component Context Model for End-to-End Wavelet Image Coding
  85. A Novel Efficient Multi-View Traffic-Related Object Detection Framework
  86. A Novel Extrapolation Technique to Accelerate WMMSE
  87. A Novel Heart Rate Estimation Method Exploiting Heartbeat Second Harmonic Reconstruction Via Millimeter Wave Radar
  88. A Novel Metric For Evaluating Audio Caption Similarity
  89. A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVC
  90. A Novel State Connection Strategy for Quantum Computing to Represent and Compress Digital Images
  91. A Novel Transformer-Based Pipeline for Lung Cytopathological Whole Slide Image Classification
  92. A Parallel Attention Mechanism for Image Manipulation Detection and Localization
  93. A Patient Invariant Model Towards the Prediction of Freezing of Gait
  94. A Perceptual Neural Audio Coder with a Mean-Scale Hyperprior
  95. A Person Identification System for the ICASSP 2023 e-Prevention Challenge
  96. A Perturbation-Based Policy Distillation Framework with Generative Adversarial Nets
  97. A Phoneme-Informed Neural Network Model For Note-Level Singing Transcription
  98. A Physically Explainable Framework for Human-Related Anomaly Detection
  99. A Point is A Wave: Point-Wave Network for Place Recognition
  100. A Practical Distributed Active Noise Control Algorithm Overcoming Communication Restrictions
  101. A Principled Approach to Model Validation in Domain Generalization
  102. A Privacy-Preserving Trajectory Mining Model
  103. A Probabilistic Framework for Pruning Transformers Via a Finite Admixture of Keys
  104. A Processing Framework to Access Large Quantities of Whispered Speech Found in ASMR
  105. A Progressive Neural Network for Acoustic Echo Cancellation
  106. A Prototypical Semantic Decoupling Method via Joint Contrastive Learning for Few-Shot Named Entity Recognition
  107. A Proximal Approach to IVA-G with Convergence Guarantees
  108. A Quantum Approach for Stochastic Constrained Binary Optimization
  109. A Quantum Kernel Learning Approach to Acoustic Modeling for Spoken Command Recognition
  110. A Radar-Jammer Zero-Sum Repeated Bayesian Game
  111. A Reality Check and a Practical Baseline for Semantic Speech Embedding
  112. A Robust Kalman Filter Based Approach for Indoor Robot Positionning with Multi-Path Contaminated UWB Data
  113. A Role Engineering Approach Based on Spectral Clustering Analysis for Restful Permissions in Cloud
  114. A Sentiment and Syntactic-Aware Graph Convolutional Network for Aspect-Level Sentiment Classification
  115. A Sidecar Separator Can Convert A Single-Talker Speech Recognition System to A Multi-Talker One
  116. A Simple Scheme for Coupled Factorization for Hyperspectral Super-Resolution: Exploiting Sparsity in an Easy Way
  117. A Simple Yet Effective Approach to Structured Knowledge Distillation
  118. A Simulation-Based Framework for Urban Traffic Accident Detection
  119. A Slot-Shared Span Prediction-Based Neural Network for Multi-Domain Dialogue State Tracking
  120. A Spatial-Temporal ECG Emotion Recognition Model Based on Dynamic Feature Fusion
  121. A Spatio-Temporal Decomposition Network for Compressed Video Quality Enhancement
  122. A Speech Representation Anonymization Framework via Selective Noise Perturbation
  123. A Statistical Interpretation of the Maximum Subarray Problem
  124. A Study of Audio Mixing Methods for Piano Transcription in Violin-Piano Ensembles
  125. A Study on Bias and Fairness in Deep Speaker Recognition
  126. A Study on the Integration of Pipeline and E2E SLU Systems for Spoken Semantic Parsing Toward Stop Quality Challenge
  127. A Study on the Invariance in Security Whatever the Dimension of Images for the Steganalysis by Deep-Learning
  128. A Synthetic Corpus Generation Method for Neural Vocoder Training
  129. A Targeted Sampling Strategy for Compressive Cryo Focused Ion Beam Scanning Electron Microscopy
  130. A Template Matching Approach for Reference Picture Padding in Video Coding
  131. A Token-Level Contrastive Framework for Sign Language Translation
  132. A Topic-Enhanced Approach for Emotion Distribution Forecasting in Conversations
  133. A Transformer-Based E2E SLU Model for Improved Semantic Parsing
  134. A Two-Branch Network for Video Anomaly Detection with Spatio-Temporal Feature Learning
  135. A Two-Stage System for Spoken Language Understanding
  136. A Unified One-Shot Prosody and Speaker Conversion System with Self-Supervised Discrete Speech Units
  137. A Unified Uncertainty-Aware Exploration: Combining Epistemic and Aleatory Uncertainty
  138. A Unitary Transform Based Generalized Approximate Message Passing
  139. A Variational Inequality Model for Learning Neural Networks
  140. A Video Anomaly Detection Framework Based on Appearance-Motion Semantics Representation Consistency
  141. A Wavelet Scattering Approach for Load Identification with Limited Amount of Training Data
  142. A non-contact SpO2 estimation using video magnification and infrared data
  143. A2S-NAS: Asymmetric Spectral-Spatial Neural Architecture Search for Hyperspectral Image Classification
  144. A3S: Adversarial Learning of Semantic Representations for Scene-Text Spotting
  145. ACE-VC: Adaptive and Controllable Voice Conversion Using Explicitly Disentangled Self-Supervised Speech Representations
  146. ACF: Aligned Contrastive Finetuning For Language and Vision Tasks
  147. AD-YOLO: You Look Only Once in Training Multiple Sound Event Localization and Detection
  148. ADHD Classification with Biomarker Identification Using a Triplet Loss Attention Auto-Encoding Network
  149. AE-Flow: Autoencoder Normalizing Flow
  150. AERO: Audio Super Resolution in the Spectral Domain
  151. AMC-Net: An Effective Network for Automatic Modulation Classification
  152. AMPose: Alternately Mixed Global-Local Attention Model for 3D Human Pose Estimation
  153. APGP: Accuracy-Preserving Generative Perturbation for Defending Against Model Cloning Attacks
  154. ASSD: Synthetic Speech Detection in the AAC Compressed Domain
  155. AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer
  156. AURA: Privacy-Preserving Augmentation to Improve Test Set Diversity in Speech Enhancement
  157. AV-TAD: Audio-Visual Temporal Action Detection With Transformer
  158. AVES: Animal Vocalization Encoder Based on Self-Supervision
  159. Absolute Decision Corrupts Absolutely: Conservative Online Speaker Diarisation
  160. Abstract Representation for Multi-Intent Spoken Language Understanding
  161. Abusive Activity Detection with Multi-Modality Based on Convolutional Neural Network
  162. Accelerated Distributed Stochastic Non-Convex Optimization over Time-Varying Directed Networks
  163. Accelerated Massive MIMO Detector Based on Annealed Underdamped Langevin Dynamics
  164. Accelerating Matrix Trace Estimation by Aitken's Δ2 Process
  165. Accelerating RNN-T Training and Inference Using CTC Guidance
  166. Accidental Learners: Spoken Language Identification in Multilingual Self-Supervised Models
  167. Achievable Error Exponents for Almost Fixed-Length M-Ary Hypothesis Testing
  168. Achieving Fair Speech Emotion Recognition via Perceptual Fairness
  169. Acoustic Source Localization in the Spherical Harmonics Domain Exploiting Low-Rank Approximations
  170. Acoustically-Driven Phoneme Removal that Preserves Vocal Affect Cues
  171. Active Beam Tracking with Reconfigurable Intelligent Surface
  172. Active IRS-Assisted MIMO Channel Estimation and Prediction
  173. Active Learning for Efficient Few-Shot Classification
  174. Active Learning of non-Semantic Speech Tasks with Pretrained models
  175. Active Noise Control over 3D Space: A Realistic Error Microphone Geometry Design
  176. Active Perception System for Enhanced Visual Signal Recovery Using Deep Reinforcement Learning
  177. Active Selection of Source Patients in Transfer Learning for Epileptic Seizure Detection Using Riemannian Manifold
  178. Active Subsampling Using Deep Generative Models by Maximizing Expected Information Gain
  179. Activity-Informed Industrial Audio Anomaly Detection Via Source Separation
  180. AdapITN: A Fast, Reliable, and Dynamic Adaptive Inverse Text Normalization
  181. Adaptable End-to-End ASR Models Using Replaceable Internal LMs and Residual Softmax
  182. Adapted Multimodal Bert with Layer-Wise Fusion for Sentiment Analysis
  183. Adapter Tuning With Task-Aware Attention Mechanism
  184. Adapting Exploratory Behaviour in Active Inference for Autonomous Driving
  185. Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker Embeddings
  186. Adapting a Self-Supervised Speech Representation for Noisy Speech Emotion Recognition by Using Contrastive Teacher-Student Learning
  187. Adaptive Axonal Delays in Feedforward Spiking Neural Networks for Accurate Spoken Word Recognition
  188. Adaptive CSI Feedback with Hidden Semantic Information Transfer
  189. Adaptive Data Augmentation for Contrastive Learning
  190. Adaptive Eccm for Mitigating Smart Jammers
  191. Adaptive Endpointing with Deep Contextual Multi-Armed Bandits
  192. Adaptive Filtering Algorithms For Set-Valued Observations-Symmetric Measurement Approach To Unlabeled And Anonymized Data
  193. Adaptive Gaussian Nested Filter for Parameter Estimation and State Tracking in Dynamical Systems
  194. Adaptive Knowledge Distillation Between Text and Speech Pre-Trained Models
  195. Adaptive Large Margin Fine-Tuning For Robust Speaker Verification
  196. Adaptive Mask Co-Optimization for Modal Dependence in Multimodal Learning
  197. Adaptive Multi-Corpora Language Model Training for Speech Recognition
  198. Adaptive Noise Canceller Algorithm with SNR-Based Stepsize and Data-Dependent Averaging
  199. Adaptive Non-Local Generative Adversarial Networks for Low-Dose CT Image Denoising
  200. Adaptive Scale and Spatial Aggregation for Real-Time Object Detection
  201. Adaptive Semantic Fusion Framework for Unsupervised Monocular Depth Estimation
  202. Adaptive Simulated Annealing Through Alternating Rényi Divergence Minimization
  203. Adaptive Step-Size Methods for Compressed SGD
  204. Adaptive Submanifold-Preserving Sparse Regression for Feature Selection And Multiclass Classification
  205. Adaptive Time-Scale Modification for Improving Speech Intelligibility Based On Phoneme Clustering For Streaming Services
  206. Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech Activity
  207. Adversarial Attacks on Genotype Sequences
  208. Adversarial Contrastive Distillation with Adaptive Denoising
  209. Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition
  210. Adversarial Guitar Amplifier Modelling with Unpaired Data
  211. Adversarial Network Pruning by Filter Robustness Estimation
  212. Adversarial Permutation Invariant Training for Universal Sound Separation
  213. Adversarially Robust Fairness-Aware Regression
  214. Affinity Learning With Blind-Spot Self-Supervision for Image Denoising
  215. Agile Radio Map Prediction Using Deep Learning
  216. Aiding Speech Harmonic Recovery in DNN-Based Single Channel Noise Reduction Using Cepstral Excitation Manipulation (CEM) Components
  217. Aleatoric Uncertainty Estimation of Overnight Sleep Statistics Through Posterior Sampling Using Conditional Normalizing Flows
  218. Algebraic Convolutional Filters on Lie Group Algebras
  219. Align, Write, Re-Order: Explainable End-to-End Speech Translation via Operation Sequence Generation
  220. Alignment Entropy Regularization
  221. Alternating Constrained Minimization Based Approximate Message Passing
  222. Alternating Phase Langevin Sampling with Implicit Denoiser Priors for Phase Retrieval
  223. Amicable Aid: Perturbing Images to Improve Classification Performance
  224. An ASR-Free Fluency Scoring Approach with Self-Supervised Learning
  225. An Adapter Based Multi-Label Pre-Training for Speech Separation and Enhancement
  226. An Adaptive DFE Using Light-Pattern-Protection Algorithm in 12 NM CMOS Technology
  227. An Adaptive Enhancement Method for Gastrointestinal Low-Light Images of Capsule Endoscope
  228. An Adaptive Plug-and-Play Network for Few-Shot Learning
  229. An Analysis of Degenerating Speech Due to Progressive Dysarthria on ASR Performance
  230. An Antispoofing Approach in Biometric Authentication System for a Smartcard
  231. An Application of Quantum Mechanics to Attention Methods in Computer Vision
  232. An Approach to Ontological Learning from Weak Labels
  233. An Asynchronous Updating Reinforcement Learning Framework for Task-Oriented Dialog System
  234. An Attention-Based Approach to Hierarchical Multi-Label Music Instrument Classification
  235. An Augmented Gaussian Sum Filter through a mixture Decomposition
  236. An Auto-Encoder Based Method for Camera Fingerprint Compression
  237. An Automotive Radar Dataset For Object Classification
  238. An Edge Alignment-Based Orientation Selection Method for Neutron Tomography
  239. An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies
  240. An Efficient Beam-Sharing Algorithm for RIS-aided Simultaneous Wireless Information and Power Transfer Applications
  241. An Efficient Relay Selection Scheme for Relay-assisted HARQ
  242. An Empirical Study and Improvement for Speech Emotion Recognition
  243. An Empirical Study of Backdoor Attacks on Masked Auto Encoders
  244. An Empirical Study on Speech Restoration Guided by Self-Supervised Speech Representation
  245. An End-to-End Framework for Partial View-Aligned Clustering with Graph Structure
  246. An End-to-End Neural Network for Image-to-Audio Transformation
  247. An Evaluation Platform to Scope Performance of Synthetic Environments in Autonomous Ground Vehicles Simulation
  248. An Experimental Study on Sound Event Localization and Detection Under Realistic Testing Conditions
  249. An Implicit Gradient Method for Constrained Bilevel Problems Using Barrier Approximation
  250. An Improved Optimal Transport Kernel Embedding Method with Gating Mechanism for Singing Voice Separation and Speaker Identification
  251. An Interpretable Model Using Evidence Information for Multi-Hop Question Answering Over Long Texts
  252. An Isotropy Analysis for Self-Supervised Acoustic Unit Embeddings on the Zero Resource Speech Challenge 2021 Framework
  253. An Online Algorithm for Chance Constrained Resource Allocation
  254. An Online Algorithm for Contrastive Principal Component Analysis
  255. Analysing Diffusion-based Generative Approaches Versus Discriminative Approaches for Speech Restoration
  256. Analysing Discrete Self Supervised Speech Representation For Spoken Language Modeling
  257. Analysing the Masked Predictive Coding Training Criterion for Pre-Training a Speech Representation Model
  258. Analysis Of Noisy-Target Training For Dnn-Based Speech Enhancement
  259. Analysis and Re-Synthesis of Natural Cricket Sounds Assessing the Perceptual Relevance of Idiosyncratic Parameters
  260. Analysis and Transformation of Voice Level in Singing Voice
  261. Analyzing Acoustic Word Embeddings from Pre-Trained Self-Supervised Speech Models
  262. Anchored Speech Recognition with Neural Transducers
  263. Ancient Chinese Word Segmentation and Part-of-Speech Tagging Using Distant Supervision
  264. Angle-Of-Arrival Target Tracking Using A Mobile Uav In External Signal-Denied Environment
  265. Animal Re-Identification Algorithm for Posture Diversity
  266. Anomalous Signal Detection for Cyber-Physical Systems Using Interpretable Causal Neural Network
  267. Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining
  268. Anomaly Detection in Optical Spectra VIA Joint Optimization
  269. Antenna Impedance Estimation in Correlated Rayleigh Fading Channels
  270. Any-to-Any Voice Conversion with F0 and Timbre Disentanglement and Novel Timbre Conditioning
  271. Applying Independent Vector Analysis on EEG-Based Motor Imagery Classification
  272. Applying Symmetrical Component Transform for Industrial Appliance Classification in Non-Intrusive Load Monitoring
  273. Approximation Error Back-Propagation for Q-Function in Scalable Reinforcement Learning with Tree Dependence Structure
  274. Aprogressive Image Dehazing Framework with inter and Intra Contrastive Learning
  275. Articulation GAN: Unsupervised Modeling of Articulatory Learning
  276. Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization
  277. Assessing the Robustness of Deep Learning-Assisted Pathological Image Analysis Under Practical Variables of Imaging System
  278. Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting A Calibrated External Microphone Array
  279. Associative Learning Network for Coherent Visual Storytelling
  280. Asymmetric Polynomial Loss for Multi-Label Classification
  281. Asymptotic Bias and Variance of Kernel Ridge Regression
  282. Asymptotic Distribution of Stochastic Mirror Descent Iterates in Average Ensemble Models
  283. Asymptotically Optimal Nonparametric Classification Rules for Spike Train Data
  284. Asynchronous Federated Learning for Real-Time Multiple Licence Plate Recognition Through Semantic Communication
  285. Asynchronous Social Learning
  286. Attention Based Relation Network for Facial Action Units Recognition
  287. Attention Localness in Shared Encoder-Decoder Model For Text Summarization
  288. Attention Mixup: An Accurate Mixup Scheme Based On Interpretable Attention Mechanism for Multi-Label Audio Classification
  289. Attention-Guided Deep Learning Framework For Movement Quality Assessment
  290. Audio Barlow Twins: Self-Supervised Audio Representation Learning
  291. Audio Coding With Unified Noise Shaping And Phase Contrast Control
  292. Audio Cross Verification Using Dual Alignment Likelihood Ratio Test
  293. Audio Quality Assessment of Vinyl Music Collections Using Self-Supervised Learning
  294. Audio Signal Enhancement with Learning from Positive and Unlabeled Data
  295. Audio-Driven Facial Landmark Generation in Violin Performance using 3DCNN Network with Self Attention Model
  296. Audio-Driven High Definetion and Lip-Synchronized Talking Face Generation Based on Face Reenactment
  297. Audio-Driven Talking Head Video Generation with Diffusion Model
  298. Audio-Text Models Do Not Yet Leverage Natural Language
  299. Audio-Visual Inpainting: Reconstructing Missing Visual Information with Sound
  300. Audio-Visual Speaker Diarization in the Framework of Multi-User Human-Robot Interaction
  301. Audio-Visual Speech Enhancement with a Deep Kalman Filter Generative Model
  302. Audio-to-Intent Using Acoustic-Textual Subword Representations from End-to-End ASR
  303. Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec
  304. AugTarget Data Augmentation for Infrared Small Target Detection
  305. Augmentation Robust Self-Supervised Learning for Human Activity Recognition
  306. Augmenting Transformer-Transducer Based Speaker Change Detection with Token-Level Training Loss
  307. Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
  308. AutoGCF: Personalized Aggregation on Neural Graph Collaborative Filtering
  309. Automatic Camera Pose Estimation by Key-Point Matching of Reference Objects
  310. Automatic Classification of Vocal Intensity Category from Speech
  311. Automatic Error Detection in Integrated Circuits Image Segmentation: A Data-Driven Approach
  312. Automatic Segmentation of Nasopharyngeal Carcinoma in CT Images Using Dual Attention and Edge Detection
  313. Automatic Severity Classification of Dysarthric Speech by Using Self-Supervised Model with Multi-Task Learning
  314. Autonomous Navigation of a Robotic Swarm in Space Exploration Missions
  315. Autonomous Soundscape Augmentation with Multimodal Fusion of Visual and Participant-Linked Inputs
  316. Autotts: End-to-End Text-to-Speech Synthesis Through Differentiable Duration Modeling
  317. Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal Processing
  318. Auxiliary Pooling Layer For Spoken Language Understanding
  319. Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker Extraction
  320. Avoid Overthinking in Self-Supervised Models for Speech Recognition
  321. BATT: Backdoor Attack with Transformation-Based Triggers
  322. BAUENet: Boundary-Aware Uncertainty Enhanced Network for Infrared Small Target Detection
  323. BEANS: The Benchmark of Animal Sounds
  324. BECTRA: Transducer-Based End-To-End ASR with Bert-Enhanced Encoder
  325. BER-Aware Dynamic Resource Management for Edge-Assisted Goal-Oriented Communications
  326. BHE-DARTS: Bilevel Optimization Based on Hypergradient Estimation for Differentiable Architecture Search
  327. BIRD-PCC: Bi-Directional Range Image-Based Deep Lidar Point Cloud Compression
  328. BISVP: Building Footprint Extraction Via Bidirectional Serialized Vertex Prediction
  329. BTS-E: Audio Deepfake Detection Using Breathing-Talking-Silence Encoder
  330. Backdoor Attack Against Automatic Speaker Verification Models in Federated Learning
  331. Backdoor Defense via Suppressing Model Shortcuts
  332. Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation
  333. Background-Weakening Consistency Regularization for Semi-Supervised Video Action Detection
  334. BadRes: Reveal the Backdoors Through Residual Connection
  335. Bag of Tricks with Quantized Convolutional Neural Networks for Image Classification
  336. Bagging R-CNN: Ensemble for Object Detection in Complex Traffic Scenes
  337. Balanced Deep CCA for Bird Vocalization Detection
  338. Balanced Mixup Loss for Long-Tailed Visual Recognition
  339. Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation
  340. Batch Normalization Damages Federated Learning on NON-IID Data: Analysis and Remedy
  341. Batch-Ensemble Stochastic Neural Networks for Out-of-Distribution Detection
  342. Bayesian Cramér-Rao Bound Estimation With Score-Based Models
  343. Bayesian Methods for Optical Flow Estimation Using a Variational Approximation, with Applications to Ultrasound
  344. Bayesian Network Modeling and Prediction of Transitions Within the Homelessness System
  345. Bayesian Optimization with Ensemble Learning Models and Adaptive Expected Improvement
  346. Beamformer-Guided Target Speaker Extraction
  347. Beamforming Optimization in RIS-Aided Mimo Systems Under Multiple-Reflection Effects
  348. Bebert: Efficient And Robust Binary Ensemble Bert
  349. Benchmark of Physiological Model Based and Deep Learning Based Remote Photoplethysmography in Automotive Applications
  350. Benchmarking Convolutional Neural Network Inference on Low-Power Edge Devices
  351. Benchmarking Cross-Domain Face Recognition with Avatars, Caricatures and Sketches
  352. Benchmarking White Blood Cell Classification under Domain Shift
  353. Bert is Robust! A Case Against Word Substitution-Based Adversarial Attacks
  354. Better Together: Dialogue Separation and Voice Activity Detection for Audio Personalization in TV
  355. Beyond Neural-on-Neural Approaches to Speaker Gender Protection
  356. Beyond Rate Coding: Signal Coding and Reconstruction Using Lean Spike Trains
  357. Bias Identification with RankPix Saliency
  358. Bias Reduced Semidefinite Relaxation Method for Multistatic Localization in the Absence of Transmitter Position And Its Synchronization
  359. Bilateral Coarse-to-Fine Network for Point Cloud Completion
  360. Bimodal Fusion Network for Basic Taste Sensation Recognition from Electroencephalography and Electromyography
  361. Binary Image Fast Perfect Recovery from Sparse 2D-DFT Coefficients
  362. Binary Sequence Set Optimization for CDMA Applications via Mixed-Integer Quadratic Programming
  363. Binauralization Robust To Camera Rotation Using 360° Videos
  364. Biologically-Inspired Continual Learning of Human Motion Sequences
  365. Bipartite Graph Convolutional Networks with Adversarial Domain Transfer
  366. Bit Error and Block Error Rate Training for ML-Assisted Communication
  367. Blind Acoustic Room Parameter Estimation Using Phase Features
  368. Blind Estimation of Audio Processing Graph
  369. Blind Polynomial Regression
  370. Blind Source Counting and Separation with Relative Harmonic Coefficients
  371. Block-Based Color Constancy: The Deviation of Salient Pixels
  372. Blood Oxygen Saturation Estimation from Facial Video Via DC and AC Components of Spatio-Temporal Map
  373. Body Prior Guided Graph Convolutional Neural Network for Skeleton-Based Action Recognition
  374. Boosting Bert Subnets with Neural Grafting
  375. Boosting Face Recognition Performance with Synthetic Data and Limited Real Data
  376. Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning
  377. Boosting No-Reference Super-Resolution Image Quality Assessment with Knowledge Distillation and Extension
  378. Boosting Person Re-Identification with Viewpoint Contrastive Learning and Adversarial Training
  379. Boosting Prompt-Based Few-Shot Learners Through Out-of-Domain Knowledge Distillation
  380. Boosting Semi-Supervised Federated Learning with Model Personalization and Client-Variance-Reduction
  381. Boosting Signal Modulation Few-Shot Learning with Pre-Transformation
  382. Boosting Transferability of Adversarial Example via an Enhanced Euler's Method
  383. Boosting the Accuracy of SRAM-Based in-Memory Architectures Via Maximum Likelihood-Based Error Compensation Method
  384. Boundary Cue Guidance and Contextual Feature Mining for Glass Segmentation
  385. Brain Network Features Differentiate Intentions from Different Emotional Expressions of the Same Text
  386. Brainnetformer: Decoding Brain Cognitive States with Spatial-Temporal Cross Attention
  387. Breaking the Trade-Off in Personalized Speech Enhancement With Cross-Task Knowledge Distillation
  388. BreathIE: Estimating Breathing Inhale Exhale Ratio Using Motion Sensor Data from Consumer Earbuds
  389. Bridging Speech and Textual Pre-Trained Models With Unsupervised ASR
  390. Building Blocks for a Complex-Valued Transformer Architecture
  391. Building Change Detection Using Cross-Temporal Feature Interaction Network
  392. Building Keyword Search System from End-To-End Asr Systems
  393. Burst Perception-Distortion Tradeoff: Analysis and Evaluation
  394. Bytecover3: Accurate Cover Song Identification On Short Queries
  395. Byzantine-Robust and Communication-Efficient Personalized Federated Learning
  396. C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object Detection
  397. C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval
  398. CADET: Control-Aware Dynamic Edge Computing for Real-Time Target Tracking in UAV Systems
  399. CAENet: Using Collaborative Attention Transformer and Add-Boost Strategy for Single Image Deraining
  400. CAN2V: Can-Bus Data-Based Seq2seq Model for Vehicle Velocity Prediction
  401. CANDY: Category-Kernelized Dynamic Convolution for Instance Segmentation
  402. CANet: Curved Guide Line Network with Adaptive Decoder for Lane Detection
  403. CAT: Causal Audio Transformer for Audio Classification
  404. CB-Conformer: Contextual Biasing Conformer for Biased Word Recognition
  405. CC-PoseNet: Towards Human Pose Estimation in Crowded Classrooms
  406. CD-FSOD: A Benchmark For Cross-Domain Few-Shot Object Detection
  407. CDHD: Contrastive Dreamer for Hint Distillation
  408. CF-VTON: Multi-Pose Virtual Try-on with Cross-Domain Fusion
  409. CFFMixer: Multi-Dimensional Feature Fusion for Object Detection
  410. CLAP Learning Audio Concepts from Natural Language Supervision
  411. CLIP4VideoCap: Rethinking Clip for Video Captioning with Multiscale Temporal Fusion and Commonsense Knowledge
  412. CLMAE: A Liter and Faster Masked Autoencoders
  413. CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing
  414. CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech Synthesis
  415. CNEG-VC: Contrastive Learning Using Hard Negative Example In Non-Parallel Voice Conversion
  416. CNN Filter for RPR-Based SR in VVC with Wavelet Decomposition
  417. CNN Filter for Super-Resolution with RPR Functionality in VVC
  418. CO-NET: Classification-Oriented Point Cloud Sampling via Informative Feature Learning and Non-Overlapped Local Adjustment
  419. CONSEN: Complementary and Simultaneous Ensemble for Alzheimer's Disease Detection and MMSE Score Prediction
  420. CORSD: Class-Oriented Relational Self Distillation
  421. COVID-19 Detection from Speech in Noisy Conditions
  422. CPA: Compressed Private Aggregation for Scalable Federated Learning Over Massive Networks
  423. CPD-GAN: Cascaded Pyramid Deformation GAN for Pose Transfer
  424. CRFAST: Clip-Based Reference-Guided Facial Image Semantic Transfer
  425. CROSSSPEECH: Speaker-Independent Acoustic Representation for Cross-Lingual Speech Synthesis
  426. CSM In Motion Vector Steganalysis: The Effect of Coders on Motion Vectors in H.264 Video Encoding
  427. CTCBERT: Advancing Hidden-Unit Bert with CTC Objectives
  428. CTTSR: A Hybrid CNN-Transformer Network for Scene Text Image Super-Resolution
  429. Calibrating AI Models for Few-Shot Demodulation VIA Conformal Prediction
  430. Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?
  431. Can Spoofing Countermeasure And Speaker Verification Systems Be Jointly Optimised?
  432. Cancelling Intermodulation Distortions for Otoacoustic Emission Measurements with Earbuds
  433. Capacity Maximization for Active RIS Assisted Outdoor-to-Indoor Communication System
  434. Capturing Cross-Scale Disparity for Stereo Image Super-Resolution
  435. Cardiac Disease Diagnosis on Imbalanced Electrocardiography Data Through Optimal Transport Augmentation
  436. Cascading and Direct Approaches to Unsupervised Constituency Parsing on Spoken Sentences
  437. Causal Discovery and Causal Inference Based Counterfactual Fairness in Machine Learning
  438. Central Nodes Detection from Partially Observed Graph Signals
  439. Centralized Cascade Multi-Channel Noise Reduction and Acoustic Feedback Cancellation in a Wireless Acoustic Sensor And Actuator Network
  440. Centroid Distance Distillation for Effective Rehearsal in Continual Learning
  441. Certified Robustness of Quantum Classifiers Against Adversarial Examples Through Quantum Noise
  442. Change Point Detection with Neural Online Density-Ratio Estimator
  443. Channel Estimation in Massive MIMO with Heavy-Tailed Noise: Gaussian-Mixture Versus Cauchy Models
  444. Channel Estimation with Tightly-Coupled Antenna Arrays
  445. Channel State Information-Free Artificial Noise-Aided Location-Privacy Enhancement
  446. Channel-Driven Decentralized Bayesian Federated Learning for Trustworthy Decision Making in D2D Networks
  447. Choice Fusion As Knowledge For Zero-Shot Dialogue State Tracking
  448. Chord-Conditioned Melody Harmonization With Controllable Harmonicity
  449. Class-Aware Contextual Information for Semantic Segmentation
  450. Class-Aware Shared Gaussian Process Dynamic Model
  451. Class-Guided Triple Head Prediction Network for Long-Tail Object Detection
  452. Class-Incremental Learning on Multivariate Time Series Via Shape-Aligned Temporal Distillation
  453. ClassA Entropy for the Analysis of Structural Complexity of Physiological Signals
  454. Classification of Synthetic Facial Attributes by Means of Hybrid Classification/Localization Patch-Based Analysis
  455. Classification of the Cervical Vertebrae Maturation (CVM) Stages Using the Tripod Network
  456. Classification via Subspace Learning Machine (SLM): Methodology and Performance Evaluation
  457. Classification-Based Dynamic Network for Efficient Super-Resolution
  458. Classifying Non-Individual Head-Related Transfer Functions with A Computational Auditory Model: Calibration And Metrics
  459. Classifying Pathological Images Based on Multi-Instance Learning and End-to-End Attention Pooling
  460. Clean Sample Guided Self-Knowledge Distillation for Image Classification
  461. Cleanformer: A Multichannel Array Configuration-Invariant Neural Enhancement Frontend for ASR in Smart Speakers
  462. Clicker: Attention-Based Cross-Lingual Commonsense Knowledge Transfer
  463. Client Selection for Generalization in Accelerated Federated Learning: A Bandit Approach
  464. Clustered Greedy Algorithm For Large-Scale Sensor Selection
  465. Clustering-Based Supervised Contrastive Learning for Identifying Risk Items on Heterogeneous Graph
  466. Co-Design for Mimo Radar and Mimo Communication Aided by Reconfigurable Intelligent Surface
  467. Co-Operative CNN for Visual Saliency Prediction on WCE Images
  468. Coarse-To-Fine Knowledge Selection for Document Grounded Dialogs
  469. Coarse-to-Fine Covid-19 Segmentation via Vision-Language Alignment
  470. Cochlear Decomposition: A Novel Bio-Inspired Multiscale Analysis Framework
  471. Cocktail Hubert: Generalized Self-Supervised Pre-Training for Mixture and Single-Source Speech
  472. Code-Enhanced Fine-Grained Semantic Matching For Tag Recommendation In Software Information Sites
  473. Code-Switching Speech Synthesis Based on Self-Supervised Learning and Domain Adaptive Speaker Encoder
  474. Code-Switching Text Generation and Injection in Mandarin-English ASR
  475. Codebook-Based User Tracking in IRS-Assisted mmWave Communication Networks
  476. Coded Matrix Computations for D2D-Enabled Linearized Federated Learning
  477. Codes Correcting Burst and Arbitrary Erasures for Reliable and Low-Latency Communication
  478. Cold Diffusion for Speech Enhancement
  479. Collaborative Audio-Visual Event Localization Based on Sequential Decision and Cross-Modal Consistency
  480. Color Guided Depth Map Super-Resolution with Nonlocla Autoregres-Sive Modeling
  481. Column-Based Matrix Approximation with Quasi-Polynomial Structure
  482. Combining Dual-Tree Wavelet Analysis and Proximal Optimization for Anisotropic Scale-Free Texture Segmentation
  483. Combining Loss Reweighting and Sample Resampling for Long-Tailed Instance Segmentation
  484. Combining the Silhouette and Skeleton Data for Gait Recognition
  485. Commdre: Document-Level Relation Extraction with Self-Supervised Commonsense Learning
  486. Communication-Constrained Exchange of Zeroth-Order Information with Application to Collaborative Target Tracking
  487. Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization
  488. Comparative Layer-Wise Analysis of Self-Supervised Speech Models
  489. Comparative Study of IRS Assisted Opportunistic Communications Over i.i.d. and los channels
  490. Comparing Decentralized Gradient Descent Approaches and Guarantees
  491. Comparison of Soft and Hard Target RNN-T Distillation for Large-Scale ASR
  492. Compensatory Debiasing For Gender Imbalances In Language Models
  493. Complementary Learning System Based Intrinsic Reward in Reinforcement Learning
  494. Compose & Embellish: Well-Structured Piano Performance Generation via A Two-Stage Approach
  495. Composition of Motion from Video Animation Through Learning Local Transformations
  496. Comprehensive Complexity Assessment of Emerging Learned Image Compression on CPU and GPU
  497. Compressed Distributed Regression over Adaptive Networks
  498. Compressed-Sensing-Based 3D Localization with Distributed Passive Reconfigurable Intelligent Surfaces
  499. Compressing Cross-Domain Representation via Lifelong Knowledge Distillation
  500. Compressive Channel Estimation for IRS-Aided Millimeter-Wave Systems via Two-Stage Lamp Network
  501. Compressive Estimation of Near Field Channels for Ultra Massive-Mimo Wideband THz Systems
  502. Compressive Sensing with Tensorized Autoencoder
  503. Conditional Conformer: Improving Speaker Modulation For Single And Multi-User Speech Enhancement
  504. Conditional LS-GAN Based Skylight Polarization Image Restoration and Application in Meridian Localization
  505. Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
  506. Confidence-Based Event-Centric Online Video Question Answering on a Newly Constructed ATBS Dataset
  507. Conformer-Based Target-Speaker Automatic Speech Recognition For Single-Channel Audio
  508. Consistent Estimators of a New Class of Covariance Matrix Distances in the Large Dimensional Regime
  509. Constrained Dynamical Neural ODE for Time Series Modelling: A Case Study on Continuous Emotion Prediction
  510. Constrained Independent Component Analysis Based on Entropy Bound Minimization for Subgroup Identification from Multi-subject fMRI Data
  511. Constrained non-negative PARAFAC2 for electromyogram separation
  512. Content-Insensitive Dynamic Lip Feature Extraction for Visual Speaker Authentication Against Deepfake Attacks
  513. Context-Aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis
  514. Context-Aware Face Clustering with Graph Convolutional Networks
  515. Context-Aware Fine-Tuning of Self-Supervised Speech Models
  516. Context-Aware end-to-end ASR Using Self-Attentive Embedding and Tensor Fusion
  517. Contextual Similarity is More Valuable Than Character Similarity: An Empirical Study for Chinese Spell Checking
  518. Contextually-Rich Human Affect Perception Using Multimodal Scene Information
  519. Continilm: A Continual Learning Scheme for Non-Intrusive Load Monitoring
  520. Continual Cell Instance Segmentation of Microscopy Images
  521. Continual Learning for On-Device Speech Recognition Using Disentangled Conformers
  522. Continuous Action Space-Based Spoken Language Acquisition Agent Using Residual Sentence Embedding and Transformer Decoder
  523. Continuous Descriptor-Based Control for Deep Audio Synthesis
  524. Continuous Interaction with A Smart Speaker via Low-Dimensional Embeddings of Dynamic Hand Pose
  525. Continuous Learning for Blind Image Quality Assessment with Contrastive Transformer
  526. Contrast-PLC: Contrastive Learning for Packet Loss Concealment
  527. Contrastive Domain Adaptation Via Delimitation Discriminator
  528. Contrastive Learning at the Relation and Event Level for Rumor Detection
  529. Contrastive Learning of Functionality-Aware Code Embeddings
  530. Contrastive Learning of Sentence Embeddings in Product Search
  531. Contrastive Learning with Dialogue Attributes for Neural Dialogue Generation
  532. Contrastive Learning-Based Audio to Lyrics Alignment for Multiple Languages
  533. Contrastive Representation Learning for Acoustic Parameter Estimation
  534. Contrastive Self-Supervised Learning for Automated Multi-Modal Dance Performance Assessment
  535. Contrastive Speech Mixup for Low-Resource Keyword Spotting
  536. Controllable Music Inpainting with Mixed-Level and Disentangled Representation
  537. Convergence Analysis of Graphical Game-Based Nash Q-Learning using the Interaction Detection Signal of N-Step Return
  538. Convergence of Stochastic PDMM
  539. Conversation-Oriented ASR with Multi-Look-Ahead CBS Architecture
  540. Conversational Text-to-SQL: An Odyssey into State-of-the-Art and Challenges Ahead
  541. Convex Optimization of Deep Polynomial and ReLU Activation Neural Networks
  542. Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification
  543. Convolutional Filtering on Sampled Manifolds
  544. Convolutional Recurrent MetriCGAN With Spectral Dimension Compression For Full-Band Speech Enhancement
  545. Convolutional Recurrent Neural Networks for the Classification of Cetacean Bioacoustic Patterns
  546. Convolutive NTF for Ambisonic Source Separation under Reverberant Conditions
  547. Cooperative Five Degrees Of Freedom Motion Estimation For A Swarm Of Autonomous Vehicles
  548. Core: Transferable Long-Range Time Series Forecasting Enhanced by Covariates-Guided Representation
  549. Cosmopolite Sound Monitoring (CoSMo): A Study of Urban Sound Event Detection Systems Generalizing to Multiple Cities
  550. Cough Detection Using Millimeter-Wave Fmcw Radar
  551. Could the BubbleView Metaphor be used to Infer Visual Attention on 3D Graphical Content?
  552. Counterfactual Explanation for Multivariate Times Series Using A Contrastive Variational Autoencoder
  553. Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval
  554. Coupled CP Tensor Decomposition with Shared and Distinct Components for Multi-Task Fmri Data Fusion
  555. Cov Loss: Covariance-Based Loss for Deep Face Recognition
  556. Covariance Regularization for Probabilistic Linear Discriminant Analysis
  557. Cramér-Rao Bound on Lie Groups with Observations on Lie Groups: Application to SE(2)
  558. Cross Modality Knowledge Distillation for Robust Pedestrian Detection in Low Light and Adverse Weather Conditions
  559. Cross-Device Federated Learning for Mobile Health Diagnostics: A First Study on COVID-19 Detection
  560. Cross-Domain Diffusion Based Speech Enhancement for Very Noisy Speech
  561. Cross-Domain Learning with Normalizing Flow
  562. Cross-Domain Object Classification Via Successive Subspace Alignment
  563. Cross-Head Supervision for Crowd Counting with Noisy Annotations
  564. Cross-Lingual Alzheimer's Disease Detection Based on Paralinguistic and Pre-Trained Features
  565. Cross-Lingual Transfer Learning for Alzheimer's Detection from Spontaneous Speech
  566. Cross-Modal Adversarial Contrastive Learning for Multi-Modal Rumor Detection
  567. Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification
  568. Cross-Modal Fusion Techniques for Utterance-Level Emotion Recognition from Text and Speech
  569. Cross-Modal Matching and Adaptive Graph Attention Network for RGB-D Scene Recognition
  570. Cross-Modal Mutual Learning for Cued Speech Recognition
  571. Cross-Modal Optical Flow Estimation via Modality Compensation and Alignment
  572. Cross-Modality depth Estimation via Unsupervised Stereo RGB-to-infrared Translation
  573. Cross-Site Generalization for Imbalanced Epileptic Classification
  574. Cross-Speaker Emotion Transfer by Manipulating Speech Style Latents
  575. Cross-Subject Mental Fatigue Detection based on Separable Spatio-Temporal Feature Aggregation
  576. Cross-Training: A Semi-Supervised Training Scheme for Speech Recognition
  577. Cross-Utterance ASR Rescoring with Graph-Based Label Propagation
  578. CryoSWD: Sliced Wasserstein Distance Minimization for 3D Reconstruction in Cryo-electron Microscopy
  579. Cumulative Attention Based Streaming Transformer ASR with Internal Language Model Joint Training and Rescoring
  580. Customized Automatic Face Beautification
  581. Cutting Through the Noise: An Empirical Comparison of Psycho-Acoustic and Envelope-based Features for Machinery Fault Detection
  582. CyFi-TTS: Cyclic Normalizing Flow with Fine-Grained Representation for End-to-End Text-to-Speech
  583. CyPMLI: WISL-Minimized Unimodular Sequence Design via Power Method-Like Iterations
  584. D-3DLD: Depth-Aware Voxel Space Mapping for Monocular 3D Lane Detection with Uncertainty
  585. D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object Detection
  586. D2Former: A Fully Complex Dual-Path Dual-Decoder Conformer Network Using Joint Complex Masking and Complex Spectral Mapping for Monaural Speech Enhancement
  587. D2Q-DETR: Decoupling and Dynamic Queries for Oriented Object Detection with Transformers
  588. DAIS: The Delft Database of EEG Recordings of Dutch Articulated and Imagined Speech
  589. DASA: Difficulty-Aware Semantic Augmentation for Speaker Verification
  590. DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks
  591. DB-UNet: MLP Based Dual Branch UNet for Accurate Vessel Segmentation in OCTA Images
  592. DDN: Dynamic Aggregation Enhanced Dual-Stream Network for Medical Image Classification
  593. DEHRFormer: Real-Time Transformer for Depth Estimation and Haze Removal from Varicolored Haze Scenes
  594. DGN: Descriptor Generation Network for Feature Matching in Monocular Endoscopy 3D Reconstruction
  595. DL-NET: Dilation Location Network for Temporal Action Detection
  596. DMFormer: Closing the gap Between CNN and Vision Transformers
  597. DMSA: Dynamic Multi-Scale Unsupervised Semantic Segmentation Based On Adaptive Affinity
  598. DO-FAM: Disentangled Non-Linear Latent Navigation For Facial Attribute Manipulation
  599. DPP-Based Client Selection for Federated Learning with NON-IID DATA
  600. DQFORMER: Dynamic Query Transformer for Lane Detection
  601. DRL Path Planning for UAV-Aided V2X Networks: Comparing Discrete to Continuous Action Spaces
  602. DSPGAN: A Gan-Based Universal Vocoder for High-Fidelity TTS by Time-Frequency Domain Supervision from DSP
  603. DST: Deformable Speech Transformer for Emotion Recognition
  604. DTTR: Detecting Text with Transformers
  605. DVQVC: An Unsupervised Zero-Shot Voice Conversion Framework
  606. DWFormer: Dynamic Window Transformer for Speech Emotion Recognition
  607. Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis
  608. DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech
  609. Dasformer: Deep Alternating Spectrogram Transformer For Multi/Single-Channel Speech Separation
  610. Data Augmentation Based On Invariant Shape Blending For Deep Learning Classification
  611. Data Driven Joint Sensor Fusion and Regression Based on Geometric Mean Squared Error
  612. Data Leakage in Cross-Modal Retrieval Training: A Case Study
  613. Data-Aware Zero-Shot Neural Architecture Search for Image Recognition
  614. Data-Driven Graph Convolutional Neural Networks for Power System Contingency Analysis
  615. Data-Driven Quickest Change Detection in Markov Models
  616. Data2vec-Aqc: Search for the Right Teaching Assistant in the Teacher-Student Training Setup
  617. Database-Aware ASR Error Correction for Speech-to-SQL Parsing
  618. Dataset Balancing Can Hurt Model Performance
  619. De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition
  620. Decaying Contrast for Fine-Grained Video Representation Learning
  621. Decoding Auditory EEG Responses Using an Adapted Wavenet
  622. Decoding Musical Pitch from Human Brain Activity with Automatic Voxel-Wise Whole-Brain FMRI Feature Selection
  623. DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene Classification
  624. Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking
  625. Decontamination Transformer For Blind Image Inpainting
  626. Decorrelating Language Model Embeddings for Speech-Based Prediction of Cognitive Impairment
  627. Decoupled Non-Parametric Knowledge Distillation for end-to-End Speech Translation
  628. Decoupled Visual Causality for Robust Detection
  629. Deep AHS: A Deep Learning Approach to Acoustic Howling Suppression
  630. Deep Adaptive Superpixels For Hadamard Single Pixel Imaging In Near-Infrared Spectrum
  631. Deep Architecture for DOA Trajectory Localization
  632. Deep Autoencoding One-Class time Series Anomaly Detection
  633. Deep Born Operator Learning for Reflection Tomographic Imaging
  634. Deep Double Self-Expressive Subspace Clustering
  635. Deep Feature Aggregation for Lightweight Single Image Super-Resolution
  636. Deep Fusion of Multi-Object Densities Using Transformer
  637. Deep Generative Fixed-Filter Active Noise Control
  638. Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion Recognition
  639. Deep Learning Sparse Array Design Using Binary Switching Configurations
  640. Deep Learning for Lagrangian Drift Simulation at The Sea Surface
  641. Deep Learning-Based Compressive Sampling Optimization in Massive MIMO Systems
  642. Deep Learning-Based Path Loss Prediction for Outdoor Wireless Communication Systems
  643. Deep Learning-Based Stereo Camera Multi-Video Synchronization
  644. Deep Low Light Image Enhancement Via Multi-Scale Recursive Feature Enhancement and Curve Adjustment
  645. Deep Manifold Graph Auto-Encoder For Attributed Graph Embedding
  646. Deep Network Series for Large-Scale High-Dynamic Range Imaging
  647. Deep Neural Mel-Subband Beamformer for in-Car Speech Separation
  648. Deep Plug-and-Play for Tensor Robust Principal Component Analysis
  649. Deep Probabilistic Model for Lossless Scalable Point Cloud Attribute Compression
  650. Deep Proximal Gradient Method for Learned Convex Regularizers
  651. Deep Quantigraphic Image Enhancement via Comparametric Equations
  652. Deep Reinforcement Learning for Green UAV-Assisted Data Collection
  653. Deep Root Music Algorithm for Data-Driven Doa Estimation
  654. Deep Spatio-Temporal Multiplex Graph Learning for Cardiac Imaging Classification
  655. Deep Spectrum Cartography Using Quantized Measurements
  656. Deep Subband Network for Joint Suppression of Echo, Noise and Reverberation in Real-Time Fullband Speech Communication
  657. Deep Survival Analysis and Counterfactual Inference Using Balanced Representations
  658. Deep Triple-Supervision Learning Unannotated Surgical Endoscopic Video Data for Monocular Dense Depth Estimation
  659. Deep Unfolded Tensor Robust PCA With Self-Supervised Learning
  660. Deep Unfolding-Enabled Hybrid Beamforming Design for mmWave Massive MIMO Systems
  661. Deep-Unfolded Adaptive Projected Subgradient Method For Mimo Detection
  662. Deep3DSketch: 3D Modeling from Free-Hand Sketches with View- and Structural-Aware Adversarial Training
  663. Deepspace: Dynamic Spatial and Source CUE Based Source Separation for Dialog Enhancement
  664. Defending Against Universal Patch Attacks by Restricting Token Attention in Vision Transformers
  665. Defense Against Black-Box Adversarial Attacks Via Heterogeneous Fusion Features
  666. Deformable Cross Attention for Learning Optical Flow
  667. Deformable Temporal Convolutional Networks for Monaural Noisy Reverberant Speech Separation
  668. Delay-Aware Backpressure Routing Using Graph Neural Networks
  669. Delay-Penalized Transducer for Low-Latency Streaming ASR
  670. Delivering Speaking Style in Low-Resource Voice Conversion with Multi-Factor Constraints
  671. Dense Adversarial Transfer Learning Based On Class-Invariance
  672. Densitytoken: Weakly-Supervised Crowd Counting with Density Classification
  673. Depth Estimation for a Single Omnidirectional Image with Reversed-Gradient Warming-up Thresholds Discriminator
  674. DepthFormer: Multimodal Positional Encodings and Cross-Input Attention for Transformer-based Segmentation Networks
  675. Dereverberation in Acoustic Sensor Networks Using weighted Prediction Error with Microphone-Dependent Prediction Delays
  676. Design Choices for Learning Embeddings from Auxiliary Tasks for Domain Generalization in Anomalous Sound Detection
  677. Design and Performance of the Low-Power Noise Reduction Algorithm of the Med-El Sonnet 2™ Cochlear Implant Audio Processor
  678. Designing A 3d-Aware Stylenerf Encoder for Face Editing
  679. Designing Transformer Networks for Sparse Recovery of Sequential Data Using Deep Unfolding
  680. Designing and Evaluating Speech Emotion Recognition Systems: A Reality Check Case Study with IEMOCAP
  681. Detail-Aware Uncalibrated Photometric Stereo
  682. Detecting Malicious Migration on Edge to Prevent Running Data Leakage
  683. Detecting Out-of-Distribution Examples Via Class-Conditional Impressions Reappearing
  684. Detection of Real-Time Deepfakes in Video Conferencing with Active Probing and Corneal Reflection
  685. Dewarping Documents Using C2 Continuous Boundary Estimation
  686. Diabetic Retinopathy Grading with Weakly-Supervised Lesion Priors
  687. Diagonal State Space Augmented Transformers for Speech Recognition
  688. Dialog Act Guided Contextual Adapter for Personalized Speech Recognition
  689. DialogMI: A Dialogue Model Based on Enhancing Dialogue Mutual Information
  690. Dialogue Context Modelling for Action Item Detection: Solution for ICASSP 2023 Mug Challenge Track 5
  691. Dialogue System with Missing Observation
  692. Dictionary Learning on Graph Data with Weisfieler-Lehman Sub-Tree Kernel and Ksvd
  693. DiffPhase: Generative Diffusion-Based STFT Phase Retrieval
  694. DiffVoice: Text-to-Speech with Latent Diffusion
  695. Difference Coarrays of Rational Arrays
  696. Difference Guided VHR Remote Sensing Image Change Detection
  697. Differentiable Adaptive Short-Time Fourier Transform with Respect to the Window Length
  698. Differential Analysis for Networks Obeying Conservation Laws
  699. Difficulty-Aware Data Augmentor for Scene Text Recognition
  700. Diffroll: Diffusion-Based Generative Music Transcription with Unsupervised Pretraining Capability
  701. Diffusion Motion: Generate Text-Guided 3D Human Motion by Diffusion Model
  702. Diffusion Probabilistic Modeling for Fine-Grained Urban Traffic Flow Inference with Relaxed Structural Constraint
  703. Diffusion-Based Generative Speech Source Separation
  704. Diffusion-Based Sound Source Localization Using Networks of Planar Microphone Arrays
  705. Diffusionnet: An Efficient Framework to Classify Single-Molecule Images with Latent Entropy Minimization
  706. Digital Phenotype Representation by Statistical, Information Theory, Data-Driven Approach with Digital Health Data
  707. Direct Position Determination with One-Bit Signal for Multiple Targets
  708. Direction Aware Positional and Structural Encoding for Directed Graph Neural Networks
  709. Direction-of-Arrival Estimation Using Gaussian Process Interpolation
  710. DisCoHead: Audio-and-Video-Driven Talking Head Generation by Disentangled Control of Head Pose and Facial Expressions
  711. Disambiguation of Cognitive Impairment Diagnosis with EEG-Based Dual-Contrastive Learning
  712. Discriminative Speaker Representation Via Contrastive Learning with Class-Aware Attention in Angular Space
  713. Discriminative Vector Learning with Application to Single Channel Speech Separation
  714. Disentangled Feature Learning for Real-Time Neural Speech Coding
  715. Disentangled Training with Adversarial Examples for Robust Small-Footprint Keyword Spotting
  716. Disentangled and Robust Representation Learning for Bragging Classification in Social Media
  717. Disentangling Speech from Surroundings with Neural Embeddings
  718. Disentangling the Horowitz Factor: Learning Content and Style From Expressive Piano Performance
  719. Distance-Based Online Label Inference Attacks Against Split Learning
  720. Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker Verification
  721. Distill-Quantize-Tune - Leveraging Large Teachers for Low-Footprint Efficient Multilingual NLU on Edge
  722. Distinguishable Speaker Anonymization Based on Formant and Fundamental Frequency Scaling
  723. Distortion-Aware Convolutional Neural Network-Based Interpolation Filter for AVS3
  724. Distributed Adaptive Norm Estimation for Blind System Identification in Wireless Sensor Networks
  725. Distributed Admm with Limited Communications Via Deep Unfolding
  726. Distributed Bayesian Tracking on the Special Euclidean Group Using Lie Algebra Parametric Approximations
  727. Distributed Gaussian Process Hyperparameter Optimization for Multi-Agent Systems
  728. Distributed Online Learning With Adversarial Participants In An Adversarial Environment
  729. Distributed Quantum Sensing Network with Geographically Constrained Measurement Strategies
  730. Distributed Signal Processing for Out-of-System Interference Suppression in Cell-Free Massive MIMO
  731. Distributionally Robust Multiclass Classification and Applications in Deep Image Classifiers
  732. Divcon: Learning Concept Sequences for Semantically Diverse Image Captioning
  733. Diverse and Vivid Sound Generation from Text Descriptions
  734. Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm Regularization
  735. Do Coarser Units Benefit Cluster Prediction-Based Speech Pre-Training?
  736. Do Prosody Transfer Models Transfer Prosodyƒ
  737. DocRED-FE: A Document-Level Fine-Grained Entity and Relation Extraction Dataset
  738. Does Human Speech Follow Benford's Law?
  739. Does Your Model Think Like an Engineer? Explainable AI for Bearing Fault Detection with Deep Learning
  740. Does a Quieter City Mean Fewer Complaints? The Sounds of New York City During Covid-19 Lockdown
  741. Domain Adaptation with External Off-Policy Acoustic Catalogs for Scalable Contextual End-to-End Automated Speech Recognition
  742. Domain Adaptation without Catastrophic Forgetting on a Small-Scale Partially-Labeled Corpus for Speech Emotion Recognition
  743. Domain Generalized Fundus Image Segmentation via Dual-Level Mixing
  744. Domain and Language Adaptation Using Heterogeneous Datasets for Wav2vec2.0-Based Speech Recognition of Low-Resource Language
  745. Doppler-Coded Joint Division Multiple Access Waveform for Automotive MIMO Radar
  746. Double Compression Detection Based on the De-Blocking Filtering of HEVC Videos
  747. Downlink Covariance Estimation in URA FDD Massive MIMO Systems
  748. Drone-vs-Bird Detection Grand Challenge at ICASSP2023
  749. Drone-vs-Bird: Drone Detection Using YOLOv7 with CSRT Tracker
  750. Dual Collaborative Visual-Semantic Mapping for Multi-Label Zero-Shot Image Recognition
  751. Dual Meta Calibration Mix for Improving Generalization in Meta-Learning
  752. Dual Path Modeling for Semantic Matching by Perceiving Subtle Conflicts
  753. Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition
  754. Dual-Based Online Learning of Dynamic Network Topologies
  755. Dual-Cycle: Self-Supervised Dual-View Fluorescence Microscopy Image Reconstruction using CycleGAN
  756. Dual-Feature Enhancement for Weakly Supervised Temporal Action Localization
  757. Dual-Head Fusion Network for Image Enhancement
  758. Dual-Path Cross-Modal Attention for Better Audio-Visual Speech Extraction
  759. Dual-Path Dilated Convolutional Recurrent Network with Group Attention for Multi-Channel Speech Enhancement
  760. Dual-Stage Graph Convolution Network With Graph Learning For Traffic Prediction
  761. Dual-Stream Siamese Vision Transformer With Mutual Attention For Radar Gait Verification
  762. Dual-Uncertainty Guided Curriculum Learning and Part-Aware Feature Refinement for Domain Adaptive Person Re-Identification
  763. Dual-Use Signal Design for MIMO Radcom with Inter-Pulse Index Modulation
  764. Dual-graph co-representation learning for knowledge-Graph Enhanced Recommendation
  765. Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech
  766. DyLiteRADHAR: Dynamic Lightweight Slowfast Network for Human Activity Recognition Using MMWAVE Radar
  767. Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross Entropy
  768. Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR
  769. Dynamic Distributed Convex Optimization "Over-The-Air" In Decentralized Wireless Networks
  770. Dynamic Fair Node Representation Learning
  771. Dynamic Independent Component Extraction with Blending Mixing Vector: Lower Bound on Mean Interference-to-Signal Ratio
  772. Dynamic Local and Global Context Exploration for Small Object Detection
  773. Dynamic Multi-View Scene Reconstruction Using Neural Implicit Surface
  774. Dynamic Scalable Self-Attention Ensemble for Task-Free Continual Learning
  775. Dynamic Selection of p-norm in Linear Adaptive Filtering via online Kernel-based Reinforcement Learning
  776. Dynamic Signed Graph Learning
  777. Dynamic Speech Endpoint Detection with Regression Targets
  778. Dynamic Split Computing for Efficient Deep EDGE Intelligence
  779. Dynamic TF-TDNN: Dynamic Time Delay Neural Network Based on Temporal-Frequency Attention for Dialect Recognition
  780. Dynamic Vehicle Graph Interaction for Trajectory Prediction Based on Video Signals
  781. E-Branchformer-Based E2E SLU Toward Stop on-Device Challenge
  782. E-Prevention: The ICASSP-2023 Challenge on Person Identification and Relapse Detection from Continuous Recordings of Biosignals
  783. E2E Segmentation in a Two-Pass Cascaded Encoder ASR Model
  784. EBEN: Extreme Bandwidth Extension Network Applied To Speech Signals Captured With Noise-Resilient Body-Conduction Microphones
  785. ECG Artifact Removal from Single-Channel Surface EMG Using Fully Convolutional Networks
  786. ECGT2T: Towards Synthesizing Twelve-Lead Electrocardiograms from Two Asynchronous Leads
  787. EEG Emotion Recognition Via Ensemble Learning Representations
  788. EEG2IMAGE: Image Reconstruction from EEG Brain Signals
  789. EGAN: A Neural Excitation Generation Model Based on Generative Adversarial Networks with Harmonics and Noise Input
  790. EH-Enabled Distributed Detection Over Temporally Correlated Markovian MIMO Channels
  791. EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text Detection
  792. EMC2-Net: Joint Equalization and Modulation Classification Based on Constellation Network
  793. EMCLR: Expectation Maximization Contrastive Learning Representations
  794. EMIX: A Data Augmentation Method for Speech Emotion Recognition
  795. ERBNet: An Effective Representation Based Network for Unbiased Scene Graph Generation
  796. ERSAM: Neural Architecture Search for Energy-Efficient and Real-Time Social Ambiance Measurement
  797. ESCL: Equivariant Self-Contrastive Learning for Sentence Representations
  798. Early Detection of Cognitive Decline Using Voice Assistant Commands
  799. Effect of Lossy Compression Algorithms on Face Image Quality and Recognition
  800. Effective Graph-Based Modeling of Articulation Traits for Mispronunciation Detection and Diagnosis
  801. Effective Training of RNN Transducer Models on Diverse Sources of Speech and Text Data
  802. Effectiveness of Inter- and Intra-Subarray Spatial Features for Acoustic Scene Classification
  803. Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages
  804. Effectiveness of Text, Acoustic, and Lattice-Based Representations in Spoken Language Understanding Tasks
  805. Efficent Large-Scale Multi-Unimodular Waveform Design with Good Correlation Properties via Direct Phase Optimizations
  806. Efficient Compressed Video Action Recognition Via Late Fusion with a Single Network
  807. Efficient Data Loading with Quantum Autoencoder
  808. Efficient Domain Adaptation for Speech Foundation Models
  809. Efficient Feature Extraction for Non-Maximum Suppression in Visual Person Detection
  810. Efficient Feature Fusion for Learning-Based Photometric Stereo
  811. Efficient Implementation of Robust CUSUM Algorithm to Characterize Nanogaps Measurements with Heavy-Tailed Noise
  812. Efficient Intelligibility Evaluation Using Keyword Spotting: A Study on Audio-Visual Speech Enhancement
  813. Efficient Large-Scale Audio Tagging Via Transformer-to-CNN Knowledge Distillation
  814. Efficient Learning of Balanced Signature Graphs
  815. Efficient Monaural Speech Enhancement with Universal Sample Rate Band-Split RNN
  816. Efficient Multi-Scale Attention Module with Cross-Spatial Learning
  817. Efficient Online Convolutional Dictionary Learning Using Approximate Sparse Components
  818. Efficient Personalized Federated Learning on Selective Model Training
  819. Efficient Practices for Profile-to-Frontal Face Synthesis and Recognition
  820. Efficient Privacy Preserving Graph Neural Network for Node Classification
  821. Efficient Protein Structural Class Prediction Via Chaos Game Representation and Recurrent Neural Networks
  822. Efficient Quantized Constant Envelope Precoding for Multiuser Downlink Massive MIMO Systems
  823. Efficient Siamese Network for UAV Tracking
  824. Efficient Similarity-Based Passive Filter Pruning for Compressing CNNS
  825. Efficient Speech Quality Assessment Using Self-Supervised Framewise Embeddings
  826. Efficient Speech Translation with Dynamic Latent Perceivers
  827. Efficient Stuttering Event Detection Using Siamese Networks
  828. Efficient Super-Resolution for Compression Of Gaming Videos
  829. Efficient Uncertainty Estimation with Gaussian Process for Reliable Dialog Response Retrieval
  830. Efficient and Effective Multi-Camera Pose Estimation with Weighted M-Estimate Sample Consensus
  831. EfficientSpeech: An On-Device Text to Speech Model
  832. Efficiently Fusing Sparse Lidar for Enhanced Self-Supervised Monocular Depth Estimation
  833. Egocentric Action Anticipation for Personal Health
  834. Egocentric Audio-Visual Noise Suppression
  835. Eigen-Decomposition-Free Directed Graph Sampling via Gershgorin Disc Alignment
  836. Elastic Graph Transformer Networks for EEG-Based Emotion Recognition
  837. Electric Network Frequency Detection Using Least Absolute Deviations
  838. Element Selection with Wide Class of Optimization Criteria Using Non-Convex Sparse Optimization
  839. Elliptical Wishart Distribution: Maximum Likelihood Estimator from Information Geometry
  840. Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System
  841. Embrace Smaller Attention: Efficient Cross-Modal Matching with Dual Gated Attention Fusion
  842. Emodiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
  843. Emotion Recognition in Conversation from Variable-Length Context
  844. Empathetic Response Generation via Emotion Cause Transition Graph
  845. Enabling Large-Scale Image Search with Co-Attention Mechanism
  846. Encoder-Decoder Graph Convolutional Network for Automatic Timed-Up-and-Go and Sit-to-Stand Segmentation
  847. End-to-End Amp Modeling: from Data to Controllable Guitar Amplifier Models
  848. End-to-End Classification of Cell-Cycle Stages with Center-Cell Focus Tracker Using Recurrent Neural Networks
  849. End-to-End Neural Audio Coding in the MDCT Domain
  850. End-to-End Non-Autoregressive Image Captioning
  851. End-to-End Spoken Language Understanding Using Joint CTC Loss and Self-Supervised, Pretrained Acoustic Encoders
  852. End-to-End Spoken Language Understanding with Tree-Constrained Pointer Generator
  853. End-to-End Unsupervised Sketch to Image Generation
  854. End-to-End Word-Level Disfluency Detection and Classification in Children's Reading Assessment
  855. Energy Efficiency Maximization in RIS-aided Networks with Global Reflection Constraints
  856. Energy Regularized RNNS for solving non-stationary Bandit problems
  857. Enhance Transferability of Adversarial Examples with Model Architecture
  858. Enhanced Coprime Array Configuration for DoA Estimation of Non-Circular Signals
  859. Enhanced Dcf Tracker Regularized by Reliable Sample Construction
  860. Enhanced Embeddings in Zero-Shot Learning for Environmental Audio
  861. Enhanced GM-PHD Filter for Real Time Satellite Multi-Target Tracking
  862. Enhanced Low-Resolution LiDAR-Camera Calibration via Depth Interpolation and Supervised Contrastive Learning
  863. Enhancement of Text-Predicting Style Token With Generative Adversarial Network for Expressive Speech Synthesis
  864. Enhancing Multimodal Alignment with Momentum Augmentation for Dense Video Captioning
  865. Enhancing Ontology Translation Through Cross-Lingual Agreement
  866. Enhancing Representation Learning with Deep Classifiers in Presence of Shortcut
  867. Enhancing Robustness and Imperceptibility of Blind Watermarking with Improved Message Processor
  868. Enhancing Spatio-Spectral Regularization by Structure Tensor Modeling for Hyperspectral Image Denoising
  869. Enhancing Speech-To-Speech Translation with Multiple TTS Targets
  870. Enhancing Unsupervised Speech Recognition with Diffusion GANS
  871. Enhancing and Adversarial: Improve ASR with Speaker Labels
  872. Enhancing the Accuracy of Resistive In-Memory Architectures using Adaptive Signal Processing
  873. Enhancing the Efficiency of WMMSE and FP for Beamforming by Minorization-Maximization
  874. Enhancing the Vocal Range of Single-Speaker Singing Voice Synthesis with Melody-Unsupervised Pre-Training
  875. Enlightening the Student in Knowledge Distillation
  876. Enrollment Rate Prediction in Clinical Trials based on CDF Sketching and Tensor Factorization tools
  877. Ensemble Graph Q-Learning for Large Scale Networks
  878. Ensemble Knowledge Distillation of Self-Supervised Speech Models
  879. Ensemble Prosody Prediction For Expressive Speech Synthesis
  880. Ensemble and Personalized Transformer Models for Subject Identification and Relapse Detection in E-Prevention Challenge
  881. Ensemble of Deep Neural Network Models for MOS Prediction
  882. Entropy Based Feature Regularization to Improve Transferability of Deep Learning Models
  883. Epic-Sounds: A Large-Scale Dataset of Actions that Sound
  884. Epilepsy Detection Grand Challenge
  885. Equivalence of Aperture Reduction in Element Space and Constrained Combination of DFT Beams in Beamspace
  886. Error Analysis of Convolutional Beamspace Algorithms
  887. Estimating Acoustic Direction of Arrival Using a Single Structural Sensor on a Resonant Surface
  888. Estimating Inharmonic Signals with Optimal Transport Priors
  889. Estimating Normalized Graph Laplacians in Financial Markets
  890. Estimating Shapley Values of Training Utterances for Automatic Speech Recognition Models
  891. Estimating Uncertainty On Video Quality Metrics
  892. Estimating and Analyzing Neural Information flow using Signal Processing on Graphs
  893. Estimation of Cardiac Fibre Direction Based on Activation Maps
  894. Estimation of High-Dimensional Differential Graphs from Multi-Attribute Data
  895. Estimation of Time-Varying Graph Topologies from Graph Signals
  896. Estimation of Visual Contents from Human Brain Signals via VQA Based on Brain-Specific Attention
  897. Euro: Espnet Unsupervised ASR Open-Source Toolkit
  898. Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding
  899. Evaluating Speech-Phoneme Alignment and its Impact on Neural Text-To-Speech Synthesis
  900. Evaluating Variants of wav2vec 2.0 on Affective Vocal Burst Tasks
  901. Evaluation of Categorical Generative Models - Bridging the Gap Between Real and Synthetic Data
  902. Event-Based Visual Microphone
  903. Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
  904. Evopose: A Recursive Transformer for 3D Human Pose Estimation with Kinematic Structure Priors
  905. Expectation Propagation on Factor Graphs Based on Matrix Decomposition
  906. Explainable audio Classification of Playing Techniques with Layer-wise Relevance Propagation
  907. Explanations for Automatic Speech Recognition
  908. Explicit Ziv-Zakai Bound For Multiple Sources Doa Estimation
  909. Explicit and Implicit Knowledge Distillation via Unlabeled Data
  910. Exploiting 3D Human Recovery for Action Recognition with Spatio-Temporal Bifurcation Fusion
  911. Exploiting CCTV Cameras for Hand Hygiene Recognition in ICU
  912. Exploiting Interactivity and Heterogeneity for Sleep Stage Classification Via Heterogeneous Graph Neural Network
  913. Exploiting Modality-Invariant Feature for Robust Multimodal Emotion Recognition with Missing Modalities
  914. Exploiting Multi-Decision and Deep Refinement for Ultrasound Image Segmentation
  915. Exploiting One-Class Classification Optimization Objectives for Increasing Adversarial Robustness
  916. Exploiting PRNU and Linear Patterns in Forensic Camera Attribution under Complex Lens Distortion Correction
  917. Exploiting Prompt Learning with Pre-Trained Language Models for Alzheimer's Disease Detection
  918. Exploiting Sparse Recovery Algorithms for Semi-Supervised Training of Deep Neural Networks for Direction-of-Arrival Estimation
  919. Exploiting Spatial Information with the Informed Complex-Valued Spatial Autoencoder for Target Speaker Extraction
  920. Exploiting Speaker Embeddings for Improved Microphone Clustering and Speech Separation in ad-hoc Microphone Arrays
  921. Exploiting Virtual Array Diversity for Accurate Radar Detection
  922. Exploration Into Translation-Equivariant Image Quantization
  923. Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models
  924. Exploring Approaches to Multi-Task Automatic Synthesizer Programming
  925. Exploring Attention Mechanisms for Multimodal Emotion Recognition in an Emergency Call Center Corpus
  926. Exploring Binary Classification Loss for Speaker Verification
  927. Exploring Complementary Features in Multi-Modal Speech Emotion Recognition
  928. Exploring Instance Relation for Decentralized Multi-Source Domain Adaptation
  929. Exploring Language-Agnostic Speech Representations Using Domain Knowledge for Detecting Alzheimer's Dementia
  930. Exploring Progressive Hybrid-Degraded Image Processing for Homography Estimation
  931. Exploring Self-Supervised Pre-Trained ASR Models for Dysarthric and Elderly Speech Recognition
  932. Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting
  933. Exploring Subgroup Performance in End-to-End Speech Models
  934. Exploring Universal Singing Speech Language Identification Using Self-Supervised Learning Based Front-End Features
  935. Exploring Vision Transformer Layer Choosing for Semantic Segmentation
  936. Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition
  937. Exploring the Role of Fricatives in Classifying Healthy Subjects and Patients with Amyotrophic Lateral Sclerosis and Parkinson's Disease
  938. Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features
  939. Extended Expectation Maximization for Under-Fitted Models
  940. Extended Kalman Filter for Graph Signals in Nonlinear Dynamic Systems
  941. Extracting the Brain-Like Representation by an Improved Self-Organizing Map for Image Classification
  942. Extreme Audio Time Stretching Using Neural Synthesis
  943. F-PABEE: Flexible-Patience-Based Early Exiting For Single-Label and Multi-Label Text Classification Tasks
  944. F0 Estimation From Telephone Speech Using Deep Feature Loss
  945. FAPM: Fast Adaptive Patch Memory for Real-Time Industrial Anomaly Detection
  946. FCIR: Rethink Aerial Image Super Resolution with Fourier Analysis
  947. FED-3DA: A Dynamic and Personalized Federated Learning Framework
  948. FEW-Shot Continual Learning with Weight Alignment and Positive Enhancement for Bioacoustic Event Detection
  949. FFEDCL: Fair Federated Learning with Contrastive Learning
  950. FFFN: Fashion Feature Fusion Network by Co-Attention Model for Fashion Recommendation
  951. FNeural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated full- and sub-band Modeling
  952. Face Recognition on Point Cloud with Cgan-Top for Denoising
  953. Facial Texure Perceiver: Towards High-Fidelity Facial Texture Recovery with Input-Level Inductive Biased Perceiver IO
  954. Factorized AED: Factorized Attention-Based Encoder-Decoder for Text-Only Domain Adaptive ASR
  955. Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers
  956. Factorized Projection-Domain Spatio-Temporal Regularization for Dynamic Tomography
  957. False Alarm Regulation for Off-Grid Target Detection With The Matched Filter
  958. Fan-Net: Fourier-Based Adaptive Normalization for Cross-Domain Stroke Lesion Segmentation
  959. Fast 3D Human Pose Estimation Using RF Signals
  960. Fast Convolution Algorithm for Real-Valued Finite Length Sequences
  961. Fast Cross-Correlation for TDoA Estimation on Small Aperture Microphone Arrays
  962. Fast Low-Latency Convolution by Low-Rank Tensor Approximation
  963. Fast Multiscale 3D Reconstruction Using Single-Photon Lidar Data
  964. Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector Analysis
  965. Fast Robust Principle Component Analysis Using Gauss-Newton Iterations
  966. Fast Single-Person 2D Human Pose Estimation Using Multi-Task Convolutional Neural Networks
  967. Fast Yet Effective Speech Emotion Recognition with Self-Distillation
  968. Fast and Accurate Factorized Neural Transducer for Text Adaption of End-to-End Speech Recognition Models
  969. Fast and Efficient Speech Enhancement with Variational Autoencoders
  970. Fast and Exact Enumeration of Deep Networks Partitions Regions
  971. Fast and Parallel Decoding for Transducer
  972. Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
  973. Faster Than Fast: Accelerating the Griffin-Lim Algorithm
  974. Feature Selection and Text Embedding for Detecting Dementia from Spontaneous Cantonese
  975. Feature Space Recovery for Incomplete Multi-View Clustering
  976. Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound Classification
  977. FedAudio: A Federated Learning Benchmark for Audio Tasks
  978. FedEEG: Federated EEG Decoding Via inter-Subject Structure Matching
  979. FedPrompt: Communication-Efficient and Privacy-Preserving Prompt Tuning in Federated Learning
  980. FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event Classification
  981. FedSD: A New Federated Learning Structure Used in Non-iid Data
  982. FedVMR: A New Federated Learning Method for Video Moment Retrieval
  983. Federated Intelligent Terminals Facilitate Stuttering Monitoring
  984. Federated Learning for ASR Based on wav2vec 2.0
  985. Federated Self-Learning with Weak Supervision for Speech Recognition
  986. Federated Semi-Supervised Learning for Object Detection in Autonomous Driving
  987. Few but Informative Local Hash Code Matching for Image Retrieval
  988. Filler Word Detection with Hard Category Mining and Inter-Category Focal Loss
  989. Filter Pruning Via Filters Similarity in Consecutive Layers
  990. Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting
  991. FindAdaptNet: Find and Insert Adapters by Learned Layer Importance
  992. Finding Optimal Numerical Format for Sub-8-Bit Post-Training Quantization of Vision Transformers
  993. Fine-Grained Blind Face Inpainting with 3D Face Component Disentanglement
  994. Fine-Grained Emotional Control of Text-to-Speech: Learning to Rank Inter- and Intra-Class Emotion Intensities
  995. Fine-Grained Private Knowledge Distillation
  996. Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and Understanding
  997. Finer-Grained Decomposition for Parallel Quantum Mimo Processing
  998. Fixed-Point Quantization Aware Training for on-Device Keyword-Spotting
  999. Flexible Beam Design for Vital Sign Monitoring Using a Phased Array Equipped With Double-Phase Shifters
  1000. Flow-Guided Deformable Alignment Network with Self-Supervision for Video Inpainting

Looking for submission deadlines instead? See the conference deadline calendar.