← All conferences

ICASSP 2023 Accepted Papers

The full list of 2,718 papers accepted at ICASSP 2023 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. Flowgrad: Using Motion for Visual Sound Source Localization
  2. Flowpose: Conditional Normalizing Flows for 3D Human Pose and Shape Estimation from Monocular Videos
  3. Flowreg: Latent Space Regularization Using Normalizing Flow For Limited Samples Learning
  4. Focusing on Targets for Improving Weakly Supervised Visual Grounding
  5. Forecasting of Breathing Events from Speech for Respiratory Support
  6. Forensics for Adversarial Machine Learning Through Attack Mapping Identification
  7. Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention Mechanism
  8. Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting Diarization
  9. Framewise Multiple Sound Source Localization and Counting Using Binaural Spatial Audio Signals
  10. Framewise Wavegan: High Speed Adversarial Vocoder In Time Domain With Very Low Computational Complexity
  11. Free-View Expressive Talking Head Video Editing
  12. Freevc: Towards High-Quality Text-Free One-Shot Voice Conversion
  13. Frequency Bin-Wise Single Channel Speech Presence Probability Estimation Using Multiple DNNS
  14. Frequency Reciprocal Action and Fusion for Single Image Super-Resolution
  15. Frequency and Scale Perspectives of Feature Extraction
  16. Frequency-Aware Attentional Feature Fusion for Deepfake Detection
  17. Frequency-Selective Hybrid Beamforming For Mmwave Full-Duplex
  18. Fretnet: Continuous-Valued Pitch Contour Streaming For Polyphonic Guitar Tablature Transcription
  19. From Easy to Hard: Two-Stage Selector and Reader for Multi-Hop Question Answering
  20. From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition
  21. Front-End Adapter: Adapting Front-End Input of Speech Based Self-Supervised Learning for Speech Recognition
  22. Full-Band General Audio Synthesis with Score-Based Diffusion
  23. Fully Complex-Valued Deep Learning Model for Visual Perception
  24. Fully Distributed Federated Learning with Efficient Local Cooperations
  25. Fully Unsupervised Topic Clustering of Unlabelled Spoken Audio Using Self-Supervised Representation Learning and Topic Model
  26. G2CNN: Geometric Prior Based GCNN for Single-View 3D Reconstruction with Loop Subdivision
  27. G2PL: Lexicon Enhanced Chinese Polyphone Disambiguation Using Bert Adapter with a New Dataset
  28. GANStrument: Adversarial Instrument Sound Synthesis with Pitch-Invariant Instance Conditioning
  29. GCC-Speaker: Target Speaker Localization with Optimal Speaker-Dependent Weighting in Multi-Speaker Scenarios
  30. GOP-Based Latent Refinement for Learned Video Coding
  31. GSWIN: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
  32. GTN-Bailando: Genre Consistent long-Term 3D Dance Generation Based on Pre-Trained Genre Token Network
  33. GaPP: Multi-Target Tracking with Gaussian Processes
  34. Gaitcotr: Improved Spatial-Temporal Representation for Gait Recognition with a Hybrid Convolution-Transformer Framework
  35. Gaitmixer: Skeleton-Based Gait Representation Learning Via Wide-Spectrum Multi-Axial Mixer
  36. Gated Contextual Adapters For Selective Contextual Biasing In Neural Transducers
  37. Gated Enhanced RPN and Hybrid-View for Few-Shot Object Detection
  38. Gator: Graph-Aware Transformer with Motion-Disentangled Regression for Human Mesh Recovery from a 2D Pose
  39. Gaussian Prior Reinforcement Learning for Nested Named Entity Recognition
  40. Gaussian Process Dynamical Modeling for Adaptive Inference Over Graphs
  41. Gaze Pre-Train For Improving Disparity Estimation Networks
  42. Gct: Gated Contextual Transformer for Sequential Audio Tagging
  43. Gender-Cartoon: Image Cartoonization Method Based on Gender Classification
  44. General Category Network: Handwritten Mathematical Expression Recognition with Coarse-Grained Recognition Task
  45. General or Specific? Investigating Effective Privacy Protection in Federated Learning for Speech Emotion Recognition
  46. Generalized Invariant Matching Property Via Lasso
  47. Generalized Relative Harmonic Coefficients
  48. Generalized Two-Stage Particle Filter for High Dimensions
  49. Generative Model based Highly Efficient Semantic Communication Approach for Image Transmission
  50. Generative Modeling Based Manifold Learning for Adaptive Filtering Guidance
  51. Generic Dependency Modeling for Multi-Party Conversation
  52. Geogcn: Geometric Dual-Domain Graph Convolution Network For Point Cloud Denoising
  53. Geometric Matrix Completion with Collaborative Routing Between Capsules
  54. Geometry-Aware DOA Estimation Using a Deep Neural Network with Mixed-Data Input Features
  55. Gesper: A Unified Framework for General Speech Restoration
  56. Glacier: Glass-Box Transformer for Interpretable Dynamic Neuroimaging
  57. Global HRTF Interpolation Via Learned Affine Transformation of Hyper-Conditioned Features
  58. Global Localisation in Continuous Magnetic Vector Fields Using Gaussian Processes
  59. Global Matching-Optimization Network for Stereo Depth Estimation
  60. Global and Nodal Mutual Information Maximization in Heterogeneous Graphs
  61. Global-Context Aware Generative Protein Design
  62. Gluformer: Transformer-based Personalized glucose Forecasting with uncertainty quantification
  63. Going in Style: Audio Backdoors Through Stylistic Transformations
  64. Good Neighbors are All You Need for Chinese Grapheme-To-Phoneme Conversion
  65. Grad-CAM-Inspired Interpretation of Nearfield Acoustic Holography using Physics-Informed Explainable Neural Network
  66. Grad-StyleSpeech: Any-Speaker Adaptive Text-to-Speech Synthesis with Diffusion Models
  67. Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech Recognition
  68. Graph Based Semantic Ensemble of Riemannian Neural Structured Learning for BCI-EEG Signal Classification
  69. Graph Contrastive Learning with Learnable Graph Augmentation
  70. Graph Learning from Gaussian and Stationary Graph Signals
  71. Graph Neural Networks for Object Type Classification Based on Automotive Radar Point Clouds and Spectra
  72. Graph Neural Networks for Sound Source Localization on Distributed Microphone Networks
  73. Graph Representation Learning For Stroke Recurrence Prediction
  74. Graph Signal Processing For Neurogimaging to Reveal Dynamics of Brain Structure-Function Coupling
  75. Graph Signal Processing for Narrowband Direction of Arrival Estimation
  76. Graph Wavelet-Based Point Cloud Geometric Denoising with Surface-Consistent Non-Negative Kernel Regression
  77. Graph-Based Point Cloud Color Denoising with 3-Dimensional Patch-Based Similarity
  78. Graph-Based Spectro-Temporal Dependency Modeling for Anti-Spoofing
  79. Graph-Graph Context Dependency Attention for Graph Edit Distance
  80. Graphit: Iterative Reweighted ℓ1 Algorithm for Sparse Graph Inference in State-Space Models
  81. Graphmad: Graph Mixup for Data Augmentation Using Data-Driven Convex Clustering
  82. Gridless Target Localization for FDA-Mimo Radar with Sparse Arrays
  83. Group Personalized Federated Learning
  84. Group-Wise Co-Salient Object Detection with Siamese Transformers Via Brownian Distance Covariance Matching
  85. Guide and Select: A Transformer-Based Multimodal Fusion Method for Points of Interest Description Generation
  86. Guided Speech Enhancement Network
  87. HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in Conversation
  88. HARQ Delay Minimization of 5G Wireless Network with Imperfect Feedback
  89. HDNet: Hierarchical Dynamic Network for Gait Recognition using Millimeter-wave radar
  90. HEiMDaL: Highly Efficient Method for Detection and Localization of Wake-Words
  91. HIFI++: A Unified Framework for Bandwidth Extension and Speech Enhancement
  92. HIPI: A Hierarchical Performer Identification Model Based on Symbolic Representation of Music
  93. HITSZ TMG at ICASSP 2023 SPGC Shared Task: Leveraging Pre-Training and Distillation Method for Title Generation with Limited Resource
  94. HPFTN: Hierarchical Progressive Fusion Transformer Network for Video Denoising
  95. HQP-MVS:High-Quality Plane Priors Assisted Multi-View Stereo for Low-Textured Areas
  96. HRTF Field: Unifying Measured HRTF Magnitude Representation with Neural Fields
  97. HTNet: Human Topology aware network for 3d Human pose estimation
  98. HYDRA-HGR: A Hybrid Transformer-Based Architecture for Fusion of Macroscopic and Microscopic Neural Drive Information
  99. Hadamard Layer to Improve Semantic Segmentation
  100. Half-Temporal and Half-Frequency Attention U2Net for Speech Signal Improvement
  101. Halluaudio: Hallucinate Frequency as Concepts For Few-Shot Audio Classification
  102. Hankel Structured Low Rank and Sparse Representation Via L0-Norm Optimization for Compressed Ultrasound Plane Wave Signal Reconstruction
  103. HappyQuokka System for ICASSP 2023 Auditory EEG Challenge
  104. Hardware Friendly Spline Sketched Lidar
  105. Hardware-Limited Non-Uniform Task-Based Quantizers
  106. He-Gan: Differentially Private Gan Using Hamiltonian Monte Carlo Based Exponential Mechanism
  107. HeMPPCAT: Mixtures of Probabilistic Principal Component analysers for data with heteroscedastic noise
  108. Healthcall Corpus and Transformer Embeddings from Healthcare Customer-Agent Conversations
  109. Hearing and Seeing Abnormality: Self-Supervised Audio-Visual Mutual Learning for Deepfake Detection
  110. Heart Rate Estimation and Performance Analysis using MIMO Radar with Dispersed Antennas
  111. Heart Rate Extraction from Abdominal Audio Signals
  112. Hearttoheart: The Arts of Infant Versus Adult-Directed Speech Classification
  113. Heterogeneous Graph Learning for Acoustic Event Classification
  114. Heuristic Masking for Text Representation Pretraining
  115. HiSSNet: Sound Event Detection and Speaker Identification via Hierarchical Prototypical Networks for Low-Resource Headphones
  116. Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline
  117. Hierarchical Diffusion Models for Singing Voice Neural Vocoder
  118. Hierarchical Filtering With Online Learned Priors for ECG Denoising
  119. Hierarchical Graph Learning for Stock Market Prediction Via a Domain-Aware Graph Pooling Operator
  120. Hierarchical Hypergraph Recurrent Attention Network for Temporal Knowledge Graph Reasoning
  121. Hierarchical Interactive Reconstruction Network for Video Compressive Sensing
  122. Hierarchical Multi-Agent Reinforcement Learning with Intrinsic Reward Rectification
  123. Hierarchical Multi-Task Learning for Fabric Component Analysis Based on NIR Spectral Signals
  124. Hierarchical Network with Decoupled Knowledge Distillation for Speech Emotion Recognition
  125. Hierarchical Pronunciation Assessment with Multi-Aspect Attention
  126. Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech Recognition
  127. Hierarchical Spatial-Temporal Transformer with Motion Trajectory for Individual Action and Group Activity Recognition
  128. Hierarchical Spatiotemporal Feature Fusion Network For Video Saliency Prediction
  129. Hierarchical Transformer for Multi-Label Trailer Genre Classification
  130. High Quality Audio Coding with Mdctnet
  131. High-Acoustic Fidelity Text To Speech Synthesis With Fine-Grained Control Of Speech Attributes
  132. High-Dimensional Confidence Regions in Sparse MRI
  133. High-Dynamic Range ADC for Finite-Rate-of-Innovation Signals
  134. High-Frequency Transformer Network Based on Window Cross-Attention for Pansharpening
  135. High-Level Feature Fusion Network for Session-Based Social Recommendation
  136. High-Resolution Embedding Extractor for Speaker Diarisation
  137. High-Resolution Neural Network Processing of LFM Radar Pulses
  138. High-Speed Drone Detection Based On Yolo-V8
  139. Higher-Order Link Prediction Via Learnable Maximum Mean Discrepancy
  140. Higher-Order Sparse Convolutions in Graph Neural Networks
  141. Higher-Order Spatio-Temporal Neural Networks for Covid-19 Forecasting
  142. Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples
  143. Hint-Dynamic Knowledge Distillation
  144. History, Present and Future: Enhancing Dialogue Generation with Few-Shot History-Future Prompt
  145. How to Push the Fastest Model 50x Faster: Streaming Non-Autoregressive Speech Synthesis on Resouce-Limited Devices
  146. HuBERT-AGG: Aggregated Representation Distillation of Hidden-Unit Bert for Robust Speech Recognition
  147. Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-Temporal Masked Transformers
  148. Hybrid Neural Network with Cross- and Self-Module Attention Pooling for Text-Independent Speaker Verification
  149. Hybrid Ris-Assisted Interference Mitigation for Spectrum Sharing
  150. Hybrid Transformers for Music Source Separation
  151. Hybridformer: Improving Squeezeformer with Hybrid Attention and NSR Mechanism
  152. Hyneter: Hybrid Network Transformer for Object Detection
  153. HyperSteg: Hyperbolic Learning for Deep Steganography
  154. Hyperbolic Audio Source Separation
  155. Hypernetwork-Based Adaptive Image Restoration
  156. Hyperspectral Image Denoising Via Nonlocal Rank Residual Modeling
  157. Hypothesis Test for Leakage Detection in Water Pipelines with High-Dimensional Sensor Signals
  158. I Hear Your True Colors: Image Guided Audio Generation
  159. I See What You Hear: A Vision-Inspired Method to Localize Words
  160. I-Tuning: Tuning Frozen Language Models with Image for Lightweight Image Captioning
  161. I3D: Transformer Architectures with Input-Dependent Dynamic Depth for Speech Recognition
  162. IAST: Instance Association Relying on Spatio-Temporal Features for Video Instance Segmentation
  163. ICASSP 2023 Auditory EEG Decoding Challenge
  164. ICASSP 2023 Spoken Language Understanding Grand Challenge
  165. ICCRN: Inplace Cepstral Convolutional Recurrent Neural Network for Monaural Speech Enhancement
  166. ICEL: Learning with Inconsistent Explanations
  167. ICStega: Image Captioning-based Semantically Controllable Linguistic Steganography
  168. IQGAN: Robust Quantum Generative Adversarial Network for Image Synthesis On NISQ Devices
  169. IR-ECG: Invertible Reconstruction of ECG
  170. ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target Detection
  171. ITER-SIS: Robust Unlimited Sampling Via Iterative Signal Sieving
  172. Ideal: Improved Dense Local Contrastive Learning For Semi-Supervised Medical Image Segmentation
  173. Identifiable Bounded Component Analysis Via Minimum Volume Enclosing Parallelotope
  174. Identifying Coordination in a Cognitive Radar Network - A Multi-Objective Inverse Reinforcement Learning Approach
  175. Identifying Entrainment in Task-Oriented Conversations
  176. Identifying Opinion Influencers over Social Networks
  177. Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems
  178. Image Adversarial Steganography Based on Joint Distortion
  179. Image Completion Via Dual-Path Cooperative Filtering
  180. Image Fusion Via Slice-Based Convolutional Sparse Representation
  181. Image Generation is May All You Need for VQA
  182. Image Inpainting with Semantic-Aware Transformer
  183. Image Reconstruction without Explicit Priors
  184. Image Segmentation for Improved Lossless Screen Content Compression
  185. Image Sharing Chain Detection VIA Sequence-To-Sequence Model
  186. Image Source Method Based on the Directional Impulse Responses
  187. Imaginary Voice: Face-Styled Diffusion Model for Text-to-Speech
  188. ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding Inpainting
  189. Immersive Enhancement and Removal of Loudspeaker Sound Using Wireless Assistive Listening Systems and Binaural Hearing Devices
  190. Implementing Continuous HRTF Measurement in Near-Field
  191. Implicit Bayes Adaptation: A Collaborative Transport Approach
  192. Implicit Vehicle Positioning with Cooperative Lidar Sensing
  193. Implicitly Rotation Equivariant Neural Networks
  194. Importance of Different Temporal Modulations of Speech: a Tale of two Perspectives
  195. Improved Acoustic-to-Articulatory Inversion Using Representations from Pretrained Self-Supervised Learning Models
  196. Improved Appliance Transient Feature Extraction Via Template Matching
  197. Improved Belief Propagation Decoding of Turbo Codes
  198. Improved Deep Speaker Localization and Tracking: Revised Training Paradigm and Controlled Latency
  199. Improved Indoor Localization With NLOS Signal Propagations
  200. Improved Mask-Based Neural Beamforming for Multichannel Speech Enhancement by Snapshot Matching Masking
  201. Improved Projection Learning for Lower Dimensional Feature Maps
  202. Improved Small Sample Hypothesis Testing Using the Uncertain Likelihood Ratio
  203. Improved Training Of Mixture-Of-Experts Language GANs
  204. Improved Wifi-Based Respiration Tracking via Contrast Enhancement
  205. Improved Wordpcfg for Passwords with Maximum Probability Segmentation
  206. Improvements to Embedding-Matching Acoustic-to-Word ASR Using Multiple-Hypothesis Pronunciation-Based Embeddings
  207. Improving Accented Speech Recognition with Multi-Domain Training
  208. Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with Transformer
  209. Improving Adversarial Robustness with Hypersphere Embedding and Angular-Based Regularizations
  210. Improving Audio Captioning Using Semantic Similarity Metrics
  211. Improving Automatic Sleep Staging Via Temporal Smoothness Regularization
  212. Improving Bert Fine-Tuning via Stabilizing Cross-Layer Mutual Information
  213. Improving CTC-Based ASR Models With Gated Interlayer Collaboration
  214. Improving Contextual Biasing with Text Injection
  215. Improving Contextual Spelling Correction by External Acoustics Attention and Semantic Aware Data Augmentation
  216. Improving Disfluency Detection with Multi-Scale Self Attention and Contrastive Learning
  217. Improving Dropout in Graph Convolutional Networks for Recommendation via Contrastive Loss
  218. Improving EEG-based Emotion Recognition by Fusing Time-Frequency and Spatial Representations
  219. Improving Electric Load Demand Forecasting with Anchor-Based Forecasting Method
  220. Improving Fairness and Robustness in End-to-End Speech Recognition Through Unsupervised Clustering
  221. Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation
  222. Improving Heart Rate and Heart Rate Variability Estimation from Video Through a HR-RR-Tuned Filter
  223. Improving Image Captioning with Control Signal of Sentence Quality
  224. Improving Knowledge Distillation for Non-Intrusive Load Monitoring Through Explainability Guided Learning
  225. Improving Learning Objectives for Speaker Verification from the Perspective of Score Comparison
  226. Improving Massively Multilingual ASR with Auxiliary CTC Objectives
  227. Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations Perspective
  228. Improving Noisy Student Training on Non-Target Domain Data for Automatic Speech Recognition
  229. Improving Non-Autoregressive Speech Recognition with Autoregressive Pretraining
  230. Improving Occluded Human Pose Estimation Via Linked Joints
  231. Improving Performance of Real-Time Full-Band Blind Packet-Loss Concealment with Predictive Network
  232. Improving Phase-Vocoder-Based Time Stretching by Time-Directional Spectrogram Squeezing
  233. Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis
  234. Improving Retrieval-Based Dialogue System Via Syntax-Informed Attention
  235. Improving Scheduled Sampling for Neural Transducer-Based ASR
  236. Improving Self-Supervised Learning for Audio Representations by Feature Diversity and Decorrelation
  237. Improving Sentence Similarity Estimation for Unsupervised Extractive Summarization
  238. Improving Speech Enhancement via Event-Based Query
  239. Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts
  240. Improving Speech-to-Speech Translation Through Unlabeled Text
  241. Improving Spoken Language Identification with Map-Mix
  242. Improving Text-Audio Retrieval by Text-Aware Attention Pooling and Prior Matrix Revised Loss
  243. Improving Transformer-Based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads
  244. Improving Transformer-Based Networks with Locality for Automatic Speaker Verification
  245. Improving Weakly Supervised Sound Event Detection with Causal Intervention
  246. Improving fast-slow Encoder based Transducer with Streaming Deliberation
  247. Improving the Modality Representation with multi-view Contrastive Learning for Multimodal Sentiment Analysis
  248. Improving the Stochastic Gradient Descent's Test Accuracy by Manipulating the ℓ∞ Norm of its Gradient Approximation
  249. Improving the out-of-Distribution Generalization Capability of Language Models: Counterfactually-Augmented Data is not Enough
  250. In Search of Strong Embedding Extractors for Speaker Diarisation
  251. In-Sensor & Neuromorphic Computing Are all You Need for Energy Efficient Computer Vision
  252. Incorporating Lip Features into Audio-Visual Multi-Speaker DOA Estimation by Gated Fusion
  253. Incorporating Reliability in Graph Information Propagation by Fluid Dynamics Diffusion: A case of Multimodal Semisupervised Deep Learning
  254. Incorporating Uncertainty from Speaker Embedding Estimation to Speaker Verification
  255. Incorporating Visual Information Reconstruction into Progressive Learning for Optimizing audio-visual Speech Enhancement
  256. Independent Vector Analysis with Multivariate Gaussian Model: a Scalable Method by Multilinear Regression
  257. Individual Sub-Band Estimation Approach to Bandwidth Extension and Enhancement of Coded Speech
  258. Inductive Relation Prediction from Relational Paths and Context with Hierarchical Transformers
  259. InfoShape: Task-Based Neural Data Shaping via Mutual Information
  260. Information Extraction from Pill Bottle Images via Text Stitching
  261. Information and Sensing Beamforming Optimization for Multi-User Multi-Target MIMO ISAC Systems
  262. Infrared and Visible Image Fusion by Using Multi-Scale Transformation and Fractional-Order Gradient Information
  263. Inplace Cepstral Speech Enhancement System for the ICASSP 2023 Clarity Challenge
  264. Input-Dependent Dynamical Channel Association For Knowledge Distillation
  265. Instance-Aware Hierarchical Structured Policy for Prompt Learning in Vision-Language Models
  266. Int-GNN: A User Intention Aware Graph Neural Network for Session-Based Recommendation
  267. Integrated Sensing and Full-Duplex Communication: Joint Transceiver Beamforming and Power Allocation
  268. Integrating Syntactic and Semantic Knowledge in AMR Parsing with Heterogeneous Graph Attention Network
  269. Integrating the Sensing and Radio Communications Channel Modelling From Radar Mutual Interference
  270. Intent Does Matter! Propagating High-Order Relations for Exploring Interest Preferences
  271. Inter-Pulse Estimation for Sperm Whale Click Detection
  272. Inter-Scale Sure-Let Denoise with Structured Deep Image Prior: Interpretable Self-Supervised Learning
  273. Inter-Subnet: Speech Enhancement with Subband Interaction
  274. Interaction-Assisted Multi-Modal Representation Learning for Recommendation
  275. Interference Leakage Minimization in RIS-Assisted MIMO Interference Channels
  276. Intermediate Fine-Tuning Using Imperfect Synthetic Speech for Improving Electrolaryngeal Speech Recognition
  277. Intermpl: Momentum Pseudo-Labeling With Intermediate CTC Loss
  278. Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain Adaptation
  279. Interpolation Filter Model For Ramanujan Subspace Signals
  280. Interpolation of Spatial Room Impulse Responses Using Partial Optimal Transport
  281. Interpretability in the Context of Sequential Cost-Sensitive Feature Acquisition
  282. Interpretable Multi-Scale Neural Network for Granger Causality Discovery
  283. Interpretable Nonnegative Incoherent Deep Dictionary Learning for FMRI Data Analysis
  284. Interpretable, Unrolled Deep Radar Beampattern Design
  285. Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
  286. Interweaved Graph and Attention Network for 3D Human Pose Estimation
  287. Introducing Topography in Convolutional Neural Networks
  288. Inv-Senet: Invariant Self Expression Network for Clustering Under Biased Data
  289. Invariant Adversarial Imitation Learning From Visual Inputs
  290. Inverse Quadratic Transform for Minimizing A Sum of Ratios
  291. Inverse Reinforcement Learning with Graph Neural Networks for IoT Resource Allocation
  292. Investigating Content-Aware Neural Text-to-Speech MOS Prediction Using Prosodic and Linguistic Features
  293. Investigating SINDy as a Tool for Causal Discovery in Time Series Signals
  294. Investigation into Phone-Based Subword Units for Multilingual End-to-End Speech Recognition
  295. IoU-Aware Multi-Expert Cascade Network Via Dynamic Ensemble for Long-Tailed Object Detection
  296. Is Multi-Task Learning an Upper Bound for Continual Learning?
  297. Is Quality Enoughƒ Integrating Energy Consumption in a Large-Scale Evaluation of Neural Audio Synthesis Models
  298. Iterative Shallow Fusion of Backward Language Model for End-To-End Speech Recognition
  299. Iterative Water-Filling Power and Subcarrier Allocation for Multicarrier NOMA Downlink
  300. JEIT: Joint End-to-End Model and Internal Language Model Training for Speech Recognition
  301. JNDMix: Jnd-Based Data Augmentation for No-Reference Image Quality Assessment
  302. JPEG Pleno Call for Proposals Responses Quality Assessment
  303. JSV-VC: Jointly Trained Speaker Verification and Voice Conversion Models
  304. Jamming Source Localization Using Augmented Physics-Based Model
  305. Jazznet: A Dataset of Fundamental Piano Patterns for Music Audio Machine Learning Research
  306. Jeffreys Divergence-Based Regularization of Neural Network Output Distribution Applied to Speaker Recognition
  307. Joint Angle and Respiration Estimation for Passive and Device-Free Respiration Monitoring
  308. Joint Ann-SNN Co-training for Object Localization and Image Segmentation
  309. Joint Antenna Selection and Beamforming in Integrated Automotive Radar Sensing-Communications with Quantized Double Phase Shifters
  310. Joint Channel and Direction Estimation for Ground-to-UAV Communications Enabled by a Simultaneous Reflecting and Sensing RIS
  311. Joint Compression and Demosaicking For Satellite Images
  312. Joint Cryo-ET Alignment and Reconstruction with Neural Deformation Fields
  313. Joint Data Association, NLOS Mitigation, and Clutter Suppression for Networked Device-Free Sensing in 6G Cellular Network
  314. Joint Discriminator and Transfer Based Fast Domain Adaptation For End-To-End Speech Recognition
  315. Joint Estimation of Clustered user Activity and Correlated Channels with Unknown Covariance in mMTC
  316. Joint Estimation of DOA and Distance in Noisy Reverberant Conditions
  317. Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
  318. Joint Human Orientation-Activity Recognition Using WIFI Signals for Human-Machine Interaction
  319. Joint Microstrip Selection and Beamforming Design for MmWave Systems with Dynamic Metasurface Antennas
  320. Joint Millimeter-Wave AoD and AoA Estimation Using one OFDM Symbol and Frequency-Dependent Beams
  321. Joint Modeling for ASR Correction and Dialog State Tracking
  322. Joint Modelling of Spoken Language Understanding Tasks with Integrated Dialog History
  323. Joint Multi-Level Feature Network for Lightweight Person Re-Identification
  324. Joint Neural Representation for Multiple Light Fields
  325. Joint Noise Reduction and Listening Enhancement for Full-End Speech Enhancement
  326. Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
  327. Joint Robust Representation And Generalization Enhancement For Cross-Modality Person Re-Identification
  328. Joint Symbol-Level Precoding and Sub-Block-Level RIS Design for Dual-Function Radar-Communications
  329. Joint Training and Decoding for Multilingual End-to-End Simultaneous Speech Translation
  330. Joint Training of Hierarchical GANs and Semantic Segmentation for Expression Translation
  331. Joint Unmixing And Demosaicing Methods For Snapshot Spectral Images
  332. Joint Unsupervised and Supervised Learning for Context-Aware Language Identification
  333. Joint Waveform and Passive Beamformer Design in Multi-IRS-Aided Radar
  334. Jointly Visual- and Semantic-Aware Graph Memory Networks for Temporal Sentence Localization in Videos
  335. K2NN: Self-Supervised Learning with Hierarchical Nearest Neighbors for Remote Sensing
  336. KEPS-NET: Robust Parking slot Detection based Keypoint estimation for High Localization Accuracy
  337. KG-ECO: Knowledge Graph Enhanced Entity Correction For Query Rewriting
  338. Kalmanbot: Kalmannet-Aided Bollinger Bands for Pairs Trading
  339. Kernel Estimation and Deconvolution for Blind Image Super-Resolution
  340. Kernel Interpolation of Acoustic Transfer Functions with Adaptive Kernel for Directed and Residual Reverberations
  341. Kernel Ridge Regression for Generalized Graph Signal Processing
  342. Keyword-Specific Acoustic Model Pruning for Open-Vocabulary Keyword Spotting
  343. Knowledge Distillation with Active Exploration and Self-Attention Based Inter-Class Variation Transfer for Image Segmentation
  344. Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured Learning
  345. Knowledge-Augmented Frame Semantic Parsing with Hybrid Prompt-Tuning
  346. Knowledge-Aware Bayesian Co-Attention for Multimodal Emotion Recognition
  347. Knowledge-Aware Few Shot Learning for Event Detection from Short Texts
  348. Knowledge-Aware Graph Convolutional Network with Utterance-Specific Window Search for Emotion Recognition In Conversations
  349. Knowledge-Graph Augmented Music Representation for Genre Classification
  350. LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural Vocoders
  351. LABANet: Lead-Assisting Backbone Attention Network for Oral Multi-Pathology Segmentation
  352. LDTSF: A Label-Decoupling Teacher-Student Framework for Semi-Supervised Echocardiography Segmentation
  353. LE-DTA: Local Extrema Convolution for Drug Target Affinity Prediction
  354. LEAPT: Learning Adaptive Prefix-to-Prefix Translation For Simultaneous Machine Translation
  355. LED: Label Correlation Enhanced Decoder for Multi-Label Text Classification
  356. LGVIT: Local-Global Vision Transformer for Breast Cancer Histopathological Image Classification
  357. LIMI-VC: A Light Weight Voice Conversion Model with Mutual Information Disentanglement
  358. LINK: Linguistic Steganalysis Framework with External Knowledge
  359. LMBAO: A Landmark Map for Bundle Adjustment Odometry in LiDAR SLAM
  360. LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models
  361. LP-IOANet: Efficient High Resolution Document Shadow Removal
  362. LQGNET: Hybrid Model-Based and Data-Driven Linear Quadratic Stochastic Control
  363. LSSED: A Robust Segmentation Network for Inflamed Appendix from CT Images
  364. LSTM-Based Video Quality Prediction Accounting for Temporal Distortions in Videoconferencing Calls
  365. Label-Efficient and Robust Learning from Multiple Experts
  366. Label-Guided Contrastive Learning for Out-of-Domain Detection
  367. Large Covariance Matrix Estimation with Oracle Statistical Rate
  368. Large Dimensional Analysis of LS-SVM Transfer Learning: Application to Polsar Classification
  369. Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
  370. Large-Scale Language Model Rescoring on Long-Form Data
  371. Large-Scale Nonverbal Vocalization Detection Using Transformers
  372. Laryngeal Leukoplakia Classification Via Dense Multiscale Feature Extraction in White Light Endoscopy Images
  373. Lasso-Based Fast Residual Recovery For Modulo Sampling
  374. Last: Scalable Lattice-Based Speech Modelling in Jax
  375. Latent Iterative Refinement for Modular Source Separation
  376. Lattice-Free Sequence Discriminative Training for Phoneme-Based Neural Transducers
  377. LeanSpeech: The Microsoft Lightweight Speech Synthesis System for Limmits Challenge 2023
  378. Learn Topological Representation with Flexible Manifold Layer
  379. Learnable Flow Model Conditioned on Graph Representation Memory for Anomaly Detection
  380. Learnable Frontends That Do Not Learn: Quantifying Sensitivity To Filterbank Initialisation
  381. Learned Generative Misspecified Lower Bound
  382. Learned Kalman Filtering in Latent Space with High-Dimensional Data
  383. Learned Video Coding with Motion Compensation Mixture Model
  384. Learning 3D Human Pose and Shape Estimation Using Uncertainty-Aware Body Part Segmentation
  385. Learning ASR Pathways: A Sparse Multilingual ASR Model
  386. Learning Audio-Visual Dereverberation
  387. Learning Causal Representations for Generalizable Face Anti Spoofing
  388. Learning Cross-Lingual Visual Speech Representations
  389. Learning Cross-Modal Audiovisual Representations with Ladder Networks for Emotion Recognition
  390. Learning Dependencies of Discrete Speech Representations with Neural Hidden Markov Models
  391. Learning Dynamic Graphs under Partial Observability
  392. Learning Environmental Structure Using Acoustic Probes with a Deep Neural Network
  393. Learning Expressive And Generalizable Motion Features For Face Forgery Detection
  394. Learning From Label Proportion with Online Pseudo-Label Decision by Regret Minimization
  395. Learning From Positive and Unlabeled Data Using Observer-GAN
  396. Learning From Single-Expert Annotated Labels for Automatic Sleep Staging
  397. Learning From Yourself: A Self-Distillation Method For Fake Speech Detection
  398. Learning Generalizable Light Field Networks from Few Images
  399. Learning Gradients of Convex Functions with Monotone Gradient Networks
  400. Learning Graph Laplacian from Intrinsic Patterns via Gaussian Process
  401. Learning How to Learn Domain-Invariant Parameters for Domain Generalization
  402. Learning Hybrid Representations of Semantics and Distortion for Blind Image Quality Assessment
  403. Learning Hypergraphs From Signals With Dual Smoothness Prior
  404. Learning Interpretable Filters In Wav-UNet For Speech Enhancement
  405. Learning Properties of Holomorphic Neural Networks of Dual Variables
  406. Learning Quantum Entanglement Distillation With Noisy Classical Communications
  407. Learning Robust Self-Attention Features for Speech Emotion Recognition with Label-Adaptive Mixup
  408. Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion Cues
  409. Learning Silhouettes with Group Sparse Autoencoders
  410. Learning Sparse Alignments via Optimal Transport for Cross-Domain Fake News Detection
  411. Learning Sparse auto-Encoders for Green AI image coding
  412. Learning Speech Representations with Flexible Hidden Feature Dimensions
  413. Learning Supervised Covariation Projection Through General Covariance
  414. Learning Task-Aligned Mask Query for Instance Segmentation
  415. Learning To Generate 3d Representations of Building Roofs Using Single-View Aerial Imagery
  416. Learning To Locate Visual Answer In Video Corpus Using Question
  417. Learning To Regularized Resource Allocation with Budget Constraints
  418. Learning Unbiased Rewards with Mutual Information in Adversarial Imitation Learning
  419. Learning a Weight Map for Weakly-Supervised Localization
  420. Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action Recognition
  421. Learning on Entropy Coded Images with CNN
  422. Learning on Graphs under Label Noise
  423. Learning to Auto-Correct for High-Quality Spectrograms
  424. Learning to Balance the Global Coherence and Informativeness in Knowledge-Grounded Dialogue Generation
  425. Learning to Build Reasoning Chains by Reliable Path Retrieval
  426. Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations
  427. Learning to Explain: a Gradient-based Attribution Method for Interpreting Super-Resolution Networks
  428. Learning to Locate the Text Forgery in Smartphone Screenshots
  429. Learning to Personalize Equalization for High-Fidelity Spatial Audio Reproduction
  430. Learning to Reconnect Interrupted Trajectories for Weakly Supervised Multi-Object Tracking
  431. Learning with Multigraph Convolutional Filters
  432. Learnt Mutual Feature Compression for Machine Vision
  433. Lego-Features: Exporting Modular Encoder Features for Streaming and Deliberation ASR
  434. Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation Types
  435. Level-Line Guided Edge Drawing for Robust Line Segment Detection
  436. Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech Enhancement
  437. Leveraging Label Correlations in a Multi-Label Setting: a Case Study in Emotion
  438. Leveraging Language Embeddings for Cross-Lingual Self-Supervised Speech Representation Learning
  439. Leveraging Large Text Corpora For End-To-End Speech Summarization
  440. Leveraging Multiple Sources in Automatic African American English Dialect Detection for Adults and Children
  441. Leveraging Neural Koopman Operators to Learn Continuous Representations of Dynamical Systems from Scarce Data
  442. Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation Scoring
  443. Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection
  444. Leveraging Pretrained Representations With Task-Related Keywords for Alzheimer's Disease Detection
  445. Leveraging Sparsity with Spiking Recurrent Neural Networks for Energy-Efficient Keyword Spotting
  446. Lexicon-injected Semantic Parsing for Task-Oriented Dialog
  447. LiNuIQA: Lightweight No-Reference Image Quality Assessment Based on Non-Uniform Weighting
  448. LiQuiD-MIMO Radar: Distributed MIMO Radar with Low-Bit Quantization
  449. Light Field Compression Via Compact Neural Scene Representation
  450. Light Projection-Based Physical-World Vanishing Attack Against Car Detection
  451. Light-Weight CNN-Attention Based Architecture for Hand Gesture Recognition Via Electromyography
  452. Light-Weight Sequential SBL Algorithm: An Alternative to OMP
  453. LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech
  454. Lightvessel: Exploring Lightweight Coronary Artery Vessel Segmentation Via Similarity Knowledge Distillation
  455. Lightweight Annotation and Class Weight Training for Automatic Estimation of Alarm Audibility in Noise
  456. Lightweight Feature Encoder for Wake-Up Word Detection Based on Self-Supervised Speech Representation
  457. Lightweight Fisher Vector Transfer Learning for Video Deduplication
  458. Lightweight Machine Learning for Seizure Detection on Wearable Devices
  459. Lightweight Portrait Segmentation Via Edge-Optimized Attention
  460. Lightweight Prosody-TTS for Multi-Lingual Multi-Speaker Scenario
  461. Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
  462. Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-Speech
  463. Line Segment Matching Based on Intersection-Enhanced Point Correspondences
  464. Linear Microphone Array Parallel to the Driving Direction for in-Car Speech Enhancement
  465. Lip-to-Speech Synthesis in the Wild with Multi-Task Learning
  466. Lit the Darkness: Three-Stage Zero-Shot Learning for Low-Light Enhancement with Multi-Neighbor Enhancement Factors
  467. LiteG2P: A Fast, Light and High Accuracy Model for Grapheme-to-Phoneme Conversion
  468. Liveness Score-Based Regression Neural Networks for Face Anti-Spoofing
  469. Local Feature Enhanced Adversarial Network for the Blind Image Quality Assessment
  470. Local Graph-Homomorphic Processing for Privatized Distributed Systems
  471. Local to global prior Learning for blind Unsupervised Image super Resolution
  472. Local-Global Progressive U-Transformers for Accurate Hepatic and Portal Veins Segmentation in Abdominal MR Images
  473. Local-Global Siamese Network with Efficient Inter-Scale Feature Learning for Change Detection in VHR Remote Sensing Images
  474. Locale Encoding for Scalable Multilingual Keyword Spotting Models
  475. Locality Preserving Multiview Graph Hashing For Large Scale Remote Sensing Image Search
  476. Log-Can: Local-Global Class-Aware Network For Semantic Segmentation of Remote Sensing Images
  477. Logo-Former: Local-Global Spatio-Temporal Transformer for Dynamic Facial Expression Recognition
  478. Logovit: Local-Global Vision Transformer for Object Re-Identification
  479. Long Range Imaging Using Multispectral Fusion of RGB and NIR Images
  480. Long-Memory Message-Passing for Spatially Coupled Systems
  481. Long-Short Attention Network For The Spectral Super-Resolution Of Multispectral Images
  482. Long-Tailed Image Recognition with Dynamic Re-Weighting
  483. Long-Tailed Recognition with Causal Invariant Transformation
  484. Long-Term Synchronization of Wireless Acoustic Sensor Networks with Nonpersistent Acoustic Activity Using Coherence State
  485. LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer
  486. Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming Perception
  487. Look and Think: Intrinsic Unification of Self-Attention and Convolution for Spatial-Channel Specificity
  488. Loss Function Design for DNN-Based Sound Event Localization and Detection on Low-Resource Realistic Data
  489. Lost In Translation: Generating Adversarial Examples Robust to Round-Trip Translation
  490. Low Precision Representations for High Dimensional Models
  491. Low in Resolution, High in Precision: UAV Detection with Super-Resolution and Motion Information Extraction
  492. Low-Bitrate Redundancy Coding of Speech Using A Rate-Distortion-Optimized Variational Autoencoder
  493. Low-Complexity Acoustic Echo Cancellation with Neural Kalman Filtering
  494. Low-Dose CT Reconstruction Via Optimization-Inspired GAN
  495. Low-Latency Electrolaryngeal Speech Enhancement Based on Fastspeech2-Based Voice Conversion and Self-Supervised Speech Representation
  496. Low-Rank Constrained Memory Autoencoder for Hyperspectral Anomaly Detection
  497. Low-Rank Plus Sparse Trajectory Decomposition for Direct Exoplanet Imaging
  498. Low-Rank Tensor Decompositions for Quaternion Multiway Arrays
  499. Low-Resource Music Genre Classification with Cross-Modal Neural Model Reprogramming
  500. Lyapunov-Driven Deep Reinforcement Learning for Edge Inference Empowered by Reconfigurable Intelligent Surfaces
  501. M-CTRL: A Continual Representation Learning Framework with Slowly Improving Past Pre-Trained Model
  502. M-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image Retrieval
  503. M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis
  504. M22: Rate-Distortion Inspired Gradient Compression
  505. M2TSR: Multi-Range and Mix-Grained Transformer for Single Image Super-Resolution
  506. M3ST: Mix at Three Levels for Speech Translation
  507. MADI: Inter-Domain Matching and Intra-Domain Discrimination for Cross-Domain Speech Recognition
  508. MAID: A Conditional Diffusion Model for Long Music Audio Inpainting
  509. MASKED-AP: Attention Pyramid Convolutional Neural Network with Mask for Cervical Cell Classification
  510. MAST: Multiscale Audio Spectrogram Transformers
  511. MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And Generalization
  512. MCNET: Fuse Multiple Cues for Multichannel Speech Enhancement
  513. MCNeT: Measurement-Consistent Networks Via A Deep Implicit Layer For Solving Inverse Problems
  514. MDR-MFI:Multi-Branch Decoupled Regression and Multi-Scale Feature Interaction for Partial-to-Partial Cloud Registration
  515. MEET: A Monte Carlo Exploration-Exploitation Trade-Off for Buffer Sampling
  516. MFAT: A Multi-Level Feature Aggregated Transformer for Person Re-Identification
  517. MFCCGAN: A Novel MFCC-Based Speech Synthesizer Using Adversarial Learning
  518. MGAT: Multi-Granularity Attention Based Transformers for Multi-Modal Emotion Recognition
  519. MHLAT: Multi-Hop Label-Wise Attention Model for Automatic ICD Coding
  520. MHSCNET: A Multimodal Hierarchical Shot-Aware Convolutional Network for Video Summarization
  521. MID-Attribute Speaker Generation Using Optimal-Transport-Based Interpolation of Gaussian Mixture Models
  522. MLCGAN: Multi-Lead ECG Synthesis with Multi Label Conditional Generative Adversarial Network
  523. MLP-GAN for Brain Vessel Image Segmentation
  524. MMATR: A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor Methods
  525. MMCosine: Multi-Modal Cosine Loss Towards Balanced Audio-Visual Fine-Grained Learning
  526. MODEFORMER: Modality-Preserving Embedding For Audio-Video Synchronization Using Transformers
  527. MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation
  528. MPS-AMS: Masked Patches Selection and Adaptive Masking Strategy Based Self-Supervised Medical Image Segmentation
  529. MRML: Multimodal Rumor Detection by Deep Metric Learning
  530. MRNET: Multi-Refinement Network for Dual-Pixel Images Defocus Deblurring
  531. MSFORMER: Multi-Scale Transformer with Neighborhood Consensus for Feature Matching
  532. MSN-net: Multi-Scale Normality Network for Video Anomaly Detection
  533. MSNet: A Deep Architecture Using Multi-Sentiment Semantics for Sentiment-Aware Image Style Transfer
  534. MSP-Former: Multi-Scale Projection Transformer for Single Image Desnowing
  535. MTDL-NET: Morphological and Temporal Discriminative Learning for Heartbeat Classification
  536. MTFD: Multi-Teacher Fusion Distillation for Compressed Video Action Recognition
  537. MUG: A General Meeting Understanding and Generation Benchmark
  538. Mabnet: Master Assistant Buddy Network With Hybrid Learning for Image Retrieval
  539. Machine Learning Based Early Debris Detection Using Automotive Low Level Radar Data
  540. Machine Learning-Aided Piece-Wise Modeling Technique of Power Amplifier for Digital Predistortion
  541. Make More of Your Data: Minimal Effort Data Augmentation for Automatic Speech Recognition and Translation
  542. Make Your Enemy Your Friend: Improving Image Rotation Angle Estimation with Harmonics
  543. Making Synchrosqueezing Locally Adaptive in The Time-Frequency Plane
  544. Managing Information Updating with Edge Computing: A Distributed and Learning Approach
  545. Margin-Mixup: A Method for Robust Speaker Verification In Multi-Speaker Audio
  546. MarginNCE: Robust Sound Localization with a Negative Margin
  547. Mask Guided Selective Context Decoding for Handwritten Chinese Text Recognition
  548. Mask the Bias: Improving Domain-Adaptive Generalization of CTC-Based ASR with Internal Language Model Estimation
  549. Maskdul: Data Uncertainty Learning in Masked Face Recognition
  550. Masked Autoencoders are Articulatory Learners
  551. Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input
  552. Masked Spectrogram Prediction for Self-Supervised Audio Pre-Training
  553. Masked Token Similarity Transfer for Compressing Transformer-Based ASR Models
  554. Masking Speech Contents by Random Splicing: is Emotional Expression Preserved?
  555. Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities
  556. Massively Multilingual Shallow Fusion with Large Language Models
  557. Matching-Based Term Semantics Pre-Training for Spoken Patient Query Understanding
  558. Matrix Low-Rank Approximation for Policy Gradient Methods
  559. Matrix Recovery using Deep Generative Priors with Low-Rank Deviations
  560. Matrix Resolvent Eigenembeddings for Dynamic Graphs
  561. Maximum Likelihood Distillation for Robust Modulation Classification
  562. Mcrood: Multi-Class Radar Out-Of-Distribution Detection
  563. Measure and Countermeasure of the Capsulation Attack Against Backdoor-Based Deep Neural Network Watermarks
  564. Measuring Deviation from Stochasticity in Time-Series Using Autoencoder Based Time-Invariant Representation: Application to Black Hole Data
  565. Measuring the Transferability of ℓ∞ Attacks by the ℓ2 Norm
  566. Medleyvox: An Evaluation Dataset for Multiple Singing Voices Separation
  567. Meeting Action Item Detection with Regularized Context Modeling
  568. Memory-Augmented Contrastive Learning for Talking Head Generation
  569. Memory-Augmented U-Transformer For Multivariate Time Series Anomaly Detection
  570. Mendam: Multi-Expert Network with Distribution-Aware Momentum for Long-Tailed Recognition
  571. Meta Learning for Domain Agnostic Soft Prompt
  572. Meta Learning with Adaptive Loss Weight for Low-Resource Speech Recognition
  573. Meta++ Network for Few-Shot Aerospace Crack Segmentation
  574. Meta-Dag: Meta Causal Discovery Via Bilevel Optimization
  575. Meta-Learning for Image-Guided Millimeter-Wave Beam Selection in Unseen Environments
  576. Metric Learning for User-Defined Keyword Spotting
  577. Metric-Oriented Speech Enhancement Using Diffusion Probabilistic Model
  578. Mimo Radar Transmit Beampattern Matching Via Manifold Optimization
  579. Mingling or Misalignment? Temporal Shift for Speech Emotion Recognition with Pre-Trained Representations
  580. Minimising Distortion for GAN-Based Facial Attribute Manipulation
  581. Misspecified Cramér-Rao Bound of RIS-Aided Localization Under Geometry Mismatch
  582. Mitigating Domain Dependency for Improved Speech Enhancement Via SNR Loss Boosting
  583. Mitigating Unintended Memorization in Language Models Via Alternating Teaching
  584. Mixed Far-field and Near-field Source Localization Based on Low-Rank Matrix Reconstruction
  585. Mixed Sample Augmentation for Online Distillation
  586. Mixer: DNN Watermarking using Image Mixup
  587. MoLE : Mixture Of Language Experts For Multi-Lingual Automatic Speech Recognition
  588. Modaldrop: Modality-Aware Regularization for Temporal-Spectral Fusion in Human Activity Recognition
  589. Model Fingerprinting with Benign Inputs
  590. Model-Based Spectral Reconstruction Of Interferometric Acquisitions
  591. Model-Free Learning of Optimal Beamformers for Passive IRS-Assisted Sumrate Maximization
  592. Model-Free Online Learning for Waveform Optimization In Integrated Sensing And Communications
  593. Model-Matching Principle Applied to the Design of an Array-Based All-Neural Binaural Rendering System for Audio Telepresence
  594. Model-based vs. Data-driven Approaches for Predicting Rain-induced Attenuation in Commercial Microwave Links: A Comparative Empirical Study
  595. Modeling Global Latent Semantic in Multi-Turn Conversations with Random Context Reconstruction
  596. Modeling Turn-Taking in Human-To-Human Spoken Dialogue Datasets Using Self-Supervised Features
  597. Modeling the Wave Equation Using Physics-Informed Neural Networks Enhanced With Attention to Loss Weights
  598. Modelling Black-Box Audio Effects with Time-Varying Feature Modulation
  599. Modelling Low-Resource Accents Without Accent-Specific TTS Frontend
  600. Modify: Model-Driven Face Stylization Without Style Images
  601. Modular Conformer Training for Flexible End-to-End ASR
  602. Modulation-Based Center Alignment and Motion Mining for Spatial Temporal Action Detection
  603. Modulo EEG Signal Recovery Using Transformer
  604. Monocular 3D Human Pose Estimation Based on Global Temporal-Attentive and Joints-Attention In Video
  605. More Speaking or More Speakers?
  606. MossFormer: Pushing the Performance Limit of Monaural Speech Separation Using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions
  607. Motion Matters: A Novel Motion Modeling for Cross-View Gait Feature Learning
  608. Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge
  609. Motor Activity Recognition Using Eeg Data and Ensemble of Stacked BLSTM-LSTM Network and Transformer Model
  610. Mouth Breathing Detection Using Audio Captured Through Earbuds
  611. Movienet-PS: A Large-Scale Person Search Dataset in the Wild
  612. Moving Towards Non-Binary Gender Identification Via Analysis of System Errors in Binary Gender Classification
  613. Multi-Agent Adversarial Training Using Diffusion Learning
  614. Multi-Agent Reinforcement Learning for Covert Semantic Communications over Wireless Networks
  615. Multi-Aspect Interest Neighbor-Augmented Network for Next-Basket Recommendation
  616. Multi-Blank Transducers for Speech Recognition
  617. Multi-Carrier Wideband OCDM-Based THZ Automotive Radar
  618. Multi-Channel Audio Signal Generation
  619. Multi-Channel Speaker Extraction with Adversarial Training: The Wavlab Submission to The Clarity ICASSP 2023 Grand Challenge
  620. Multi-Dimensional Frequency Dynamic Convolution with Confident Mean Teacher for Sound Event Detection
  621. Multi-Dimensional Signal Recovery Using Low-Rank Deconvolution
  622. Multi-Dimensional and Multi-Scale Modeling for Speech Separation Optimized by Discriminative Learning
  623. Multi-Functional Reconfigurable Intelligent Surface
  624. Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG Response
  625. Multi-Head Feature Pyramid Networks for Breast Mass Detection
  626. Multi-Head Uncertainty Inference for Adversarial Attack Detection
  627. Multi-Label Temporal Evidential Neural Networks for Early Event Detection
  628. Multi-Layer Feature Division Transferable Adversarial Attack
  629. Multi-Layer Seasonal Perception Network for Time Series Forecasting
  630. Multi-Level Fusion for Burst Super-Resolution with Deep Permutation-Invariant Conditioning
  631. Multi-Lingual Pronunciation Assessment with Unified Phoneme Set and Language-Specific Embeddings
  632. Multi-Local Attention for Speech-Based Depression Detection
  633. Multi-Microphone Speaker Separation by Spatial Regions
  634. Multi-Modal Domain Generalization for Cross-Scene Hyperspectral Image Classification
  635. Multi-Modal Food Classification in a Diet Tracking System with Spoken and Visual Inputs
  636. Multi-Object Localization and Irrelevant-Semantic Separation for Nuclei Segmentation in Histopathology Images
  637. Multi-Observation Hidden Semi-Markov Model for Photoplethysmogram Signal Semantic Segmentation
  638. Multi-Output RNN-T Joint Networks for Multi-Task Learning of ASR and Auxiliary Tasks
  639. Multi-Rate Adaptive Transform Coding for Video Compression
  640. Multi-Resolution Convolutional Dictionary Learning for Riverbed Dynamics Modeling
  641. Multi-Resolution Location-Based Training for Multi-Channel Continuous Speech Separation
  642. Multi-Resolution Sequence Aggregation and Model-Agnostic Framework for Time-Series Forecasting
  643. Multi-Scale Compositional Constraints for Representation Learning on Videos
  644. Multi-Scale Receptive Field Graph Model for Emotion Recognition in Conversations
  645. Multi-Source Templates Learning for Real-Time Aerial Tracking
  646. Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech Recognition
  647. Multi-Speaker End-to-End Multi-Modal Speaker Diarization System for the MISP 2022 Challenge
  648. Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling
  649. Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
  650. Multi-Speaker Speech Synthesis from Electromyographic Signals by Soft Speech Unit Prediction
  651. Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization
  652. Multi-Stage Aggregation Transformer for Medical Image Segmentation
  653. Multi-Stream Facial Adaptive Network for Expression Recognition from a Single Image
  654. Multi-Task Bias-Variance Trade-Off Through Functional Constraints
  655. Multi-Task Sub-Band Network For Deep Residual Echo Suppression
  656. Multi-Task Transformer with Relation-Attention and Type-Attention for Named Entity Recognition
  657. Multi-Temporal Lip-Audio Memory for Visual Speech Recognition
  658. Multi-User Data Detection in Massive MIMO with 1-Bit ADCS
  659. Multi-User Methods for Vibrational Radar Backscatter Communications
  660. Multi-View Graph Regularized Deep Autoencoder-Like NMF Framework
  661. Multi-View K-Means with Laplacian Embedding
  662. Multi-View Learning for Speech Emotion Recognition with Categorical Emotion, Categorical Sentiment, and Dimensional Scores
  663. Multi-View Millimeter-Wave Imaging Over Wireless Cellular Network
  664. Multi-modal ASR error correction with joint ASR error detection
  665. Multicast Beamformer Design for Mimo Coded Caching Systems
  666. Multichannel Time-Encoding of Finite-Rate-of-Innovation Signals
  667. Multilayer Subspace Learning With Self-Sparse Robustness for Two-Dimensional Feature Extraction
  668. Multilevel FISTA for Image Restoration
  669. Multilevel Transformer for Multimodal Emotion Recognition
  670. Multilingual Alzheimer's Dementia Recognition through Spontaneous Speech: A Signal Processing Grand Challenge
  671. Multilingual End-To-End Spoken Language Understanding For Ultra-Low Footprint Applications
  672. Multilingual Query-by-Example Keyword Spotting with Metric Learning and Phoneme-to-Embedding Mapping
  673. Multilingual Word Error Rate Estimation: E-Wer3
  674. Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain Fusion
  675. Multimodal Emotion Recognition Based on Deep Temporal Features Using Cross-Modal Transformer and Self-Attention
  676. Multimodal Facial Action unit Detection with Physiological Signals
  677. Multimodal Knowledge Distillation for Arbitrary-Oriented Object Detection in Aerial Images
  678. Multimodal Microscopy Image Alignment Using Spatial and Shape Information and a Branch-and-Bound Algorithm
  679. Multimodal Propaganda Detection Via Anti-Persuasion Prompt enhanced contrastive learning
  680. Multiple Access Computation Offloading for the K-User Case
  681. Multiple Acoustic Features Speech Emotion Recognition Using Cross-Attention Transformer
  682. Multiple Contrastive Learning for Multimodal Sentiment Analysis
  683. Multiple Domain-Adversarial Ensemble Learning for Domain Generalization
  684. Multiple Signed Graph Learning for Gene Regulatory Network Inference
  685. Multiple Target Measurements: Bayesian Framework for Moving Object Detection in Mimo Radar
  686. Multiresolution Signal Processing of Financial Market Objects
  687. Multiscale Audio Spectrogram Transformer for Efficient Audio Classification
  688. Multispectral Image Fusion based on Super Pixel Segmentation
  689. Multistage Spatial Context Models for Learned Image Compression
  690. Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using Wav2vec 2.0
  691. Multitrack Music Transcription with a Time-Frequency Perceiver
  692. Multitrack Music Transformer
  693. Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects
  694. Music Rearrangement Using Hierarchical Segmentation
  695. Mutual Information Based Reweighting for Precipitation Nowcasting
  696. Mutually Guided Few-Shot Learning For Relational Triple Extraction
  697. MvCo-DoT: Multi-View Contrastive Domain Transfer Network for Medical Report Generation
  698. Möbius Total Variation for Directed Acyclic Graphs
  699. N2MVSNet: Non-Local Neighbors Aware Multi-View Stereo Network
  700. NAS-DYMC: NAS-Based Dynamic Multi-Scale Convolutional Neural Network for Sound Event Detection
  701. NBA-OMP: Near-Field Beam-Split-Aware Orthogonal Matching Pursuit for Wideband THz Channel Estimation
  702. NC-WAMKD: Neighborhood Correction Weight-Adaptive Multi-Teacher Knowledge Distillation for Graph-Based Semi-Supervised Node Classification
  703. NCL: Textual Backdoor Defense Using Noise-Augmented Contrastive Learning
  704. NF-PCAC: Normalizing Flow Based Point Cloud Attribute Compression
  705. NL-DSE: Non-Local Neural Network with Decoder-Squeeze-and-Excitation for Monocular Depth Estimation
  706. NNSVS: A Neural Network-Based Singing Voice Synthesis Toolkit
  707. NRTSI: Non-Recurrent Time Series Imputation
  708. NSV-TTS: Non-Speech Vocalization Modeling And Transfer In Emotional Text-To-Speech
  709. NVOC-22: A Low Cost Mel Spectrogram Vocoder for Mobile Devices
  710. Named Entity Detection and Injection for Direct Speech Translation
  711. Narrow Down Before Selection: A Dynamic Exclusion Model for Multiple-Choice QA
  712. Nasty-SFDA: Source Free Domain Adaptation from a Nasty Model
  713. Native Multi-Band Audio Coding Within Hyper-Autoencoded Reconstruction Propagation Networks
  714. Naturalistic Head Motion Generation from Speech
  715. Navigating and Reaching Therapeutic Goals with Dynamical Systems in Conversation-Based Interventions
  716. Near-field Localization with Dynamic Metasurface Antennas
  717. Neighborhood Information-Based Label Refinement for Person Re-Identification with Label Noise
  718. Nested Attention Network with Graph Filtering for Visual Question and Answering
  719. Networked Policy Gradient Play in Markov Potential Games
  720. Neural Architecture Search with Multimodal Fusion Methods for Diagnosing Dementia
  721. Neural Architecture of Speech
  722. Neural Band-to-Piano Score Arrangement with Stepless Difficulty Control
  723. Neural Diarization with Non-Autoregressive Intermediate Attractors
  724. Neural Feature Predictor and Discriminative Residual Coding for Low-Bitrate Speech Coding
  725. Neural Fourier Shift for Binaural Speech Rendering
  726. Neural Maximum-a-Posteriori Beamforming for Ultrasound Imaging
  727. Neural Mode Estimation
  728. Neural Network Models with Integrated Training and Adaptation For Nonlinear Acoustic System Identification
  729. Neural Networks with Quantization Constraints
  730. Neural Optimization Of Geometry And Fixed Beamformer For Linear Microphone Arrays
  731. Neural Source Coding For Bandwidth-Efficient Brain-Computer Interfacing With Wireless Neuro-Sensor Networks
  732. Neural Speech Phase Prediction Based on Parallel Estimation Architecture and Anti-Wrapping Losses
  733. Neural Transducer Training: Reduced Memory Consumption with Sample-Wise Computation
  734. Neural-AFC: Learning-Based Step-Size Control for Adaptive Feedback Cancellation with Closed-Loop Model Training
  735. Neurally Augmented State Space Model for Simultaneous Communication and Tracking with Low Complexity Receivers
  736. New Interpretable Patterns and Discriminative Features from Brain Functional Network Connectivity using Dictionary Learning
  737. Newton-Based Trainable Learning Rate
  738. Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video Conversation
  739. No Reference Quality Assessment for Screen Content Images Based on Entire and High-Influence Regions
  740. Node-Wise Domain Adaptation Based on Transferable Attention for Recognizing Road Rage via EEG
  741. Noise PSD Insensitive RTF Estimation in a Reverberant and Noisy Environment
  742. Noise-Aware Target Extension with Self-Distillation for Robust Speech Recognition
  743. Noise-Disentanglement Metric Learning for Robust Speaker Verification
  744. Non-Convex Approaches for Low-Rank Tensor Completion under Tubal Sampling
  745. Noncoherent Multiuser Grassmannian Constellations for the Mimo Multiple Access Channel
  746. Nonnegative Block-Term Decomposition with the β-Divergence: Joint Data Fusion and Blind Spectral Unmixing
  747. Nonparallel Emotional Voice Conversion for Unseen Speaker-Emotion Pairs Using Dual Domain Adversarial Network & Virtual Domain Pairing
  748. Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs
  749. Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech
  750. Not All Classes are Equal: Adaptively Focus-Aware Confidence for Semi-Supervised Object Detection
  751. Note and Playing Technique Transcription of Electric Guitar Solos in Real-World Music Performance
  752. Nowcasting of Extreme Precipitation Using Deep Generative Models
  753. Numerical Semantic Modeling for Implicit Discourse Relation Recognition
  754. OAFormer: Learning Occlusion Distinguishable Feature for Amodal Instance Segmentation
  755. OPT: One-shot Pose-Controllable Talking Head Generation
  756. OTW: Optimal Transport Warping for Time Series
  757. Oct Image Blind Despeckling Based on Gradient Guided Filter with Speckle Statistical Prior
  758. On Adversarial Robustness of Audio Classifiers
  759. On Batching Variable Size Inputs for Training End-to-End Speech Enhancement Systems
  760. On Bidirectional Preestimates and Their Application to Identification of fast Time-Varying Systems
  761. On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks
  762. On Crowdsourcing-Design with Comparison Category Rating for Evaluating Speech Enhancement Algorithms
  763. On Designing A 3d Imaging Summer Project For Ontario's High School Students During Covid-19 Pandemic
  764. On Designing Light-Weight Object Trackers Through Network Pruning: Use CNNS or Transformers?
  765. On Minimal Variations for Unsupervised Representation Learning
  766. On Multiple-Input/Binaural-Output Antiphasic Speaker Signal Extraction
  767. On Negative Sampling for Contrastive Audio-Text Retrieval
  768. On Neural Architectures for Deep Learning-Based Source Separation of Co-Channel OFDM Signals
  769. On Out-of-Distribution Detection for Audio with Deep Nearest Neighbors
  770. On Parametric Misspecified Bayesian Cramér-Rao Bound: An Application to Linear/Gaussian Systems
  771. On Super-Resolution with Separation Prior
  772. On The Design and Training Strategies for Rnn-Based Online Neural Speech Separation Systems
  773. On The Detection of Synthetic Images Generated by Diffusion Models
  774. On The Fairness of Multitask Representation Learning
  775. On The Primal and Dual Formulations Of The Discrete Mumford-Shah Functional
  776. On Tracking a Stochastically Time-Varying Subspace
  777. On Unsupervised Uncertainty-Driven Speech Pseudo-Label Filtering and Model Calibration
  778. On Using the UA-Speech and Torgo Databases to Validate Automatic Dysarthric Speech Classification Approaches
  779. On Weighted Cross-Entropy for Label-Imbalanced Separable Data: An Algorithmic-Stability Study
  780. On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition Systems
  781. On the Effectiveness of Monoaural Target Source Extraction for Distant end-to-end Automatic Speech Recognition
  782. On the Importance of Different Cough Phases for COVID-19 Detection
  783. On the Joint Estimation of Phase Noise and time-Varying Channels for OFDM under High-Mobility Conditions
  784. On the Minimum Perimeter Criterion for Bounded Component Analysis
  785. On the Quantization of Recurrent Neural Networks for Smiles Generation
  786. On the Reduction of Large-Scale Room Acoustic Models
  787. On the Relevance of the Differences Between HRTF Measurement Setups for Machine Learning
  788. On the Robustness of Non-Intrusive Speech Quality Model by Adversarial Examples
  789. On the Role of LIP Articulation in Visual Speech Perception
  790. On the Role of Visual Context in Enriching Music Representations
  791. On the Value of Stochastic Side Information in Online Learning
  792. On-the-Fly Text Retrieval for end-to-end ASR Adaptation
  793. Once-for-All Sequence Compression for Self-Supervised Speech Models
  794. One-Shot Action Detection via Attention Zooming In
  795. One-Shot Medical Action Recognition With A Cross-Attention Mechanism And Dynamic Time Warping
  796. One-Shot Neural Band Selection for Spectral Recovery
  797. Online Binaural Speech Separation Of Moving Speakers With A Wavesplit Network
  798. Online Caching with Fetching cost for Arbitrary Demand Pattern: a Drift-Plus-Penalty Approach
  799. Online Edge Flow Prediction Over Expanding Simplicial Complexes
  800. Online Learning-Based Waveform Selection for Improved Vehicle Recognition in Automotive Radar
  801. Online Model Compression for Federated Learning with Large Models
  802. Online Residual-Based Key Frame Sampling with Self-Coach Mechanism and Adaptive Multi-Level Feature Fusion
  803. Online Vector Autoregressive Models Over Expanding Graphs
  804. Ontology-Aware Network for Zero-Shot Sketch-Based Image Retrieval
  805. Open-Set Automatic Target Recognition
  806. Optimal Carrier Frequency Design for Frequency Diverse Array Mimo Radar
  807. Optimal Compression for Minimizing Classification Error Probability: An Information-Theoretic Approach
  808. Optimal Condition Training for Target Source Separation
  809. Optimal Kernel for Real-Time Arbitrary-Shaped Text Detection
  810. Optimal Mixed-ADC Arrangement for DOA Estimation Via CRB Using ULA
  811. Optimal Transport in Diffusion Modeling for Conversion Tasks in Audio Domain
  812. Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification
  813. Optimising Different Feature Types for Inpainting-Based Image Representations
  814. Optimization for Robustness Evaluation Beyond ℓp Metrics
  815. Optimization of Sensor Configurations for Fault Identification in Smart Buildings
  816. Optimization of the Deep Neural Networks for Seizure Detection
  817. Optimized Dithering for Quantization Index Modulation
  818. Optimized Quality Feature Learning for Video Quality Assessment
  819. Optimizing Distributed Multi-Sensor Multi-Target Tracking Algorithm Based On Labeled Multi-Bernoulli Filter
  820. Optimizing Quantum Federated Learning Based on Federated Quantum Natural Gradient Descent
  821. Optimizing Vision Transformers for Medical Image Segmentation
  822. Order Reduction of Multi-Channel FIR Filters by Balanced Truncation
  823. Outlier-Insensitive Kalman Filtering Using NUV Priors
  824. Output-Dependent Gaussian Process State-Space Model
  825. Outside Knowledge Visual Question Answering Version 2.0
  826. Overcoming Posterior Collapse in Variational Autoencoders Via EM-Type Training
  827. Overcoming the Seesaw in Monocular 3D Object Detection Via Language Knowledge Transferring
  828. Overlay Cognitive Radio Using Symbol Level Precoding With Quantized CSI
  829. Overview of the 2023 ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids
  830. Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)
  831. Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality
  832. PAGE: A Position-Aware Graph-Based Model for Emotion Cause Entailment in Conversation
  833. PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker Verification
  834. PCQA-Graphpoint: Efficient Deep-Based Graph Metric for Point Cloud Quality Assessment
  835. PCSalmix: Gradient Saliency-Based Mix Augmentation for Point Cloud Classification
  836. PFT-SSR: Parallax Fusion Transformer for Stereo Image Super-Resolution
  837. PI-Trans: Parallel-Convmlp and Implicit-Transformation Based Gan for Cross-View Image Translation
  838. PMMSD: Development of the Matrix Sentence Intelligibility Dataset for Mandarin with Lombard Effect
  839. PMNet: Large-Scale Channel Prediction System for ICASSP 2023 First Pathloss Radio Map Prediction Challenge
  840. POINTACL: Adversarial Contrastive Learning for Robust Point Clouds Representation Under Adversarial Attack
  841. PQLM - Multilingual Decentralized Portable Quantum Language Model
  842. PRIME: 3D Human Pose and Body Shape Recovery with Perspective Projection
  843. PRRD: Pixel-Region Relation Distillation For Efficient Semantic Segmentation
  844. PU-Edgeformer: Edge Transformer for Dense Prediction in Point Cloud Upsampling
  845. PUFFIN: Pitch-Synchronous Neural Waveform Generation for Fullband Speech on Modest Devices
  846. Paaploss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement
  847. Pair DETR: Toward Faster Convergent DETR
  848. Papez: Resource-Efficient Speech Separation with Auditory Working Memory
  849. Parafac2-Based Coupled Matrix and Tensor Factorizations
  850. Parallel 2D Seismic Ray Tracing Using Cuda on a Jetson Nano
  851. Parallel Sentence-Level Explanation Generation for Real-World Low-Resource Scenarios
  852. Parameter Efficient Transfer Learning for Various Speech Processing Tasks
  853. Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using Adapters
  854. Parasympathetic-Sympathetic Causal Interactions and Perceived Workload for Varying Difficulty Affective Computing Tasks
  855. Partially Adaptive Multichannel Joint Reduction of Ego-Noise and Environmental Noise
  856. Particle Flow Gaussian Sum Particle Filter
  857. Passive Acoustic Tracking of Whales in 3-D
  858. Passive Detection of Rank-One Gaussian Signals for Known Channel Subspaces and Arbitrary Noise
  859. Peak-First CTC: Reducing the Peak Latency of CTC Models by Applying Peak-First Regularization
  860. Perceive and Predict: Self-Supervised Speech Representation Based Loss Functions for Speech Enhancement
  861. Perceptual Analysis of Speaker Embeddings for Voice Discrimination between Machine And Human Listening
  862. Perceptual Quality Assessment for Digital Human Heads
  863. Perceptual-Neural-Physical Sound Matching
  864. Performance Above All? Energy Consumption vs. Performance, a Study on Sound Event Detection with Heterogeneous Data
  865. Performance Comparison of TTS Models for Brazilian Portuguese to Establish a Baseline
  866. Performance of Social Machine Learning Under Limited Data
  867. Performing Neural Architecture Search Without Gradients
  868. Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis
  869. Permutation Invariant Training for Paraphrase Identification
  870. Person Identification with Wearable Sensing Using Missing Feature Encoding and Multi-Stage Modality Fusion
  871. Personalized Federated Learning on Long-Tailed Data via Adversarial Feature Augmentation
  872. Personalized Lightweight Text-to-Speech: Voice Cloning with Adaptive Structured Pruning
  873. Personalized Speech Enhancement Combining Band-Split RNN and Speaker Attentive Module
  874. Personalized Task Load Prediction in Speech Communication
  875. Personalizing Federated Learning with Over-The-Air Computations
  876. Perspective Projection-Based 3d CT Reconstruction from Biplanar X-Rays
  877. Phase Retrieval for Rydberg Quantum Arrays
  878. Phase Unwrapping in Correlated Noise for FMCW Lidar Depth Estimation
  879. Phase-Aware Spoof Speech Detection Based On Res2net with Phase Network
  880. PhaseAug: A Differentiable Augmentation for Speech Synthesis to Simulate One-to-Many Mapping
  881. Phonation Mode Detection in Singing: A Singer Adapted Model
  882. Phoneix: Acoustic Feature Processing Strategy for Enhanced Singing Pronunciation With Phoneme Distribution Predictor
  883. Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme Predictions
  884. Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion Recognition
  885. Phonetic RNN-Transducer for Mispronunciation Diagnosis
  886. Physics-Informed Transfer Learning for Voltage Stability Margin Prediction
  887. Picking the Underused Heads: A Network Pruning Perspective of Attention Head Selection for Fusing Dialogue Coreference Information
  888. Piecewise Position Encoding in Convolutional Neural Network for Cough-Based Covid-19 Detection
  889. Pitch Mark Detection from Noisy Speech Waveform Using Wave-U-Net
  890. Play It Back: Iterative Attention For Audio Recognition
  891. Polarized Signal Singular Spectrum Analysis with Complex SSA
  892. Police: Provably Optimal Linear Constraint Enforcement For Deep Neural Networks
  893. Pondering About Task Spatial Misalignment: Classification-Localization Equilibrated Object Detection
  894. Pooling Strategies for Simplicial Convolutional Networks
  895. Pop2Piano : Pop Audio-Based Piano Cover Generation
  896. Position-Aware Graph-Based Learning of Whole Slide Images
  897. Positive-Pair Redundancy Reduction Regularisation for Speech-Based Asthma Diagnosis Prediction
  898. Possibilistic Bernoulli Filter for Extended Target Tracking
  899. Post-Trained Language Model Adaptive to Extractive Summarization of Long Spoken Documents
  900. Powerful and Extensible WFST Framework for Rnn-Transducer Losses
  901. Practice of the Conformer Enhanced Audio-Visual Hubert on Mandarin and English
  902. Pre-Trained Model Representations and Their Robustness Against Noise for Speech Emotion Analysis
  903. Pre-Training Strategies Using Contrastive Learning and Playlist Information for Music Classification and Similarity
  904. Precognition in Contextual Spoken Language Understanding via Knowledge Distillation
  905. Predicting Brain Age Using Transferable Covariance Neural Networks
  906. Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
  907. Predictive Skim: Contrastive Predictive Coding for Low-Latency Online Speech Separation
  908. Prefallkd: Pre-Impact Fall Detection Via CNN-ViT Knowledge Distillation
  909. Prefix Tuning for Automated Audio Captioning
  910. Prefix-Level Detection and Autocorrection of Keyboard Input Errors
  911. Preformer: Predictive Transformer with Multi-Scale Segment-Wise Correlations for Long-Term Time Series Forecasting
  912. Preserving Background Sound in Noise-Robust Voice Conversion Via Multi-Task Learning
  913. Pretrained Transformers for Seizure Detection
  914. Pretraining Conformer with ASR for Speaker Verification
  915. Prior-Enhanced Temporal Action Localization Using Subject-Aware Spatial Attention
  916. Priv-Aug-Shap-ECGResNet: Privacy Preserving Shapley-Value Attributed Augmented Resnet for Practical Single-Lead Electrocardiogram Classification
  917. Privacy Preserving Face Recognition with Lensless Camera
  918. Privacy-Enhanced Federated Learning Against Attribute Inference Attack for Speech Emotion Recognition
  919. Privacy-Preserving Automatic Speaker Diarization
  920. Privacy-Preserving Occupancy Estimation
  921. Probabilistic Back-ends for Online Speaker Recognition and Clustering
  922. Procontext: Exploring Progressive Context Transformer for Tracking
  923. Procter: Pronunciation-Aware Contextual Adapter For Personalized Speech Recognition In Neural Transducers
  924. Product Graph Learning From Multi-Attribute Graph Signals with Inter-Layer Coupling
  925. Progressive Diversifying Policy for Multi-Agent Reinforcement Learning
  926. Progressive Meta-Pooling Learning for Lightweight Image Classification Model
  927. Progressive Multi-Stage Neural Audio Codec with Psychoacoustic Loss and Discriminator
  928. Progressive Perception Learning for Distribution Modulation in Siamese Tracking
  929. Progressive Refinement Learning Based on Feature Cross Perception for Residential Areas Semantic Segmentation
  930. Projected Hierarchical ALS for Generalized Boolean Matrix Factorization
  931. Promoting Cooperation in Multi-Agent Reinforcement Learning via Mutual Help
  932. Prompt Makes mask Language Models Better Adversarial Attackers
  933. Prompt-Distiller: Few-Shot Knowledge Distillation for Prompt-Based Language Learners with Dual Contrastive Learning
  934. Prompttts: Controllable Text-To-Speech With Text Descriptions
  935. Prosody Is Not Identity: A Speaker Anonymization Approach Using Prosody Cloning
  936. Prosody-Aware Speecht5 for Expressive Neural TTS
  937. Prosody-Controllable Spontaneous TTS with Neural HMMS
  938. Prototype Knowledge Distillation for Medical Segmentation with Missing Modality
  939. Prototype-Based Layered Federated Cross-Modal Hashing
  940. Provable Computational and Statistical Guarantees for Efficient Learning of Continuous-Action Graphical Games
  941. Provably Convergent Plug & Play Linearized ADMM, Applied to Deblurring Spatially Varying Kernels
  942. Prune Then Distill: Dataset Distillation with Importance Sampling
  943. Pseudo Multi-Source Domain Extension and Selective Pseudo-Labeling for Unsupervised Domain Adaptive Medical Image Segmentation
  944. Pseudo-Inverted Bottleneck Convolution for Darts Search Space
  945. Pseudo-Query Generation For Semi-Supervised Visual Grounding With Knowledge Distillation
  946. Pushing the Limits of Self-Supervised Speaker Verification using Regularized Distillation Framework
  947. Pyramid Dynamic Inference: Encouraging Faster Inference Via Early Exit Boosting
  948. Pyramid Spatial Feature Transform and Shared-Offsets Deformable Alignment Based Convolutional Network for HDR Imaging
  949. QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis
  950. QTROJAN: A Circuit Backdoor Against Quantum Neural Networks
  951. Quantifying Catastrophic Forgetting in Continual Federated Learning
  952. Quantile Online Learning for Semiconductor Failure Analysis
  953. Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker Separation
  954. Quantized Precoding and RIS-Assisted Modulation for Integrated Sensing and Communications Systems
  955. Quantpipe: Applying Adaptive Post-Training Quantization For Distributed Transformer Pipelines In Dynamic Edge Environments
  956. Quantum Deep Recurrent Reinforcement Learning
  957. Quantum Graph Transformers
  958. Quantum Transfer Learning Using the Large-Scale Unsupervised Pre-Trained Model Wavlm-Large for Synthetic Speech Detection
  959. Quantum Variational Bayes on Manifolds
  960. Quaternion Orthogonal Transformer for Facial Expression Recognition in the Wild
  961. Query-Utterance Attention With Joint Modeuing For Query-Focused Meeting Summarization
  962. Question Answering System with Sparse and Noisy Feedback
  963. Quickest Change Detection with Leave-one-out Density Estimation
  964. RAT: Radial Attention Transformer for Singing Technique Recognition
  965. RCDPT: Radar-Camera Fusion Dense Prediction Transformer
  966. RD-NAS: Enhancing One-Shot Supernet Ranking Ability Via Ranking Distillation From Zero-Cost Proxies
  967. RDO Candidate Selection for Maximizing Coding Efficiency in a Practical HEVC Encoder
  968. RGB-D Based Pose-Invariant Face Recognition Via Attention Decomposition Module
  969. RIS Reflection and Placement Optimisation for Underlay D2D Communications in Cognitive Cellular Networks
  970. RIS-Aided Wideband DFRC with Reconfigurable Holographic Surface
  971. RL-IFF: Indoor Localization via Reinforcement Learning-Based Information Fusion
  972. RNN-Based Step-Size Estimation for the RLS Algorithm with Application to Acoustic Echo Cancellation
  973. ROI-Based Deep Image Compression with Swin Transformers
  974. Radar Clutter Covariance Estimation: A Nonlinear Spectral Shrinkage Approach
  975. Radio Map Based UAV Target Localization
  976. Radio Sensing with Large Intelligent Surface for 6G
  977. Radio-Astronomy Imaging and Interference Excision Using Tensor Decomposition and Canonical Correlation Analysis
  978. Rain2Avoid: Self-Supervised Single Image Deraining
  979. Raising The Limit of Image Rescaling Using Auxiliary Encoding
  980. Randmasking Augment: A Simple and Randomized Data Augmentation For Acoustic Scene Classification
  981. Random Projector: Efficient Deep Image Prior
  982. Range-ISL Minimization and Spectral Shaping in MIMO Radar Systems via Waveform Design
  983. Rapid Audiometric Evaluation for Personalized Headphone Listening
  984. Rate Region Characterization for Semantics and Bits based Multiuser Communications
  985. Rate Splitting and Precoding Strategies for Multi-User MIMO Broadcast Channels with Common and Private Streams
  986. Rate-Distortion Optimization with Alternative References for UGC Video Compression
  987. Rate-Distortion Optimized Variable-Node-size Trisoup for Point Cloud Coding
  988. Raw Ultrasound-Based Phonetic Segments Classification Via Mask Modeling
  989. Real-Time Audio-Visual End-To-End Speech Enhancement
  990. Real-Time Human Reconstruction Based on Human Pose Prior and Epipolar Refinement
  991. Real-Time MRI Video Synthesis from Time Aligned Phonemes with Sequence-to-Sequence Networks
  992. Real-Time Modelling of Observation Filter in the Remote Microphone Technique for an Active Noise Control Application
  993. Real-Time Multichannel Speech Separation and Enhancement Using a Beamspace-Domain-Based Lightweight CNN
  994. Real-Time Speech Enhancement with Dynamic Attention Span
  995. Real-Time Speech Interruption Analysis: from Cloud to Client Deployment
  996. Real-Time Target Sound Extraction
  997. Real-Time Wireless ECG-Derived Respiration Rate Estimation using an Autoencoder with a DCT Layer
  998. Received Power Maximization with Practical Phase-Dependent Amplitude Response in RIS-Aided OFDM Wireless Communications
  999. Receptive Field Reliant Zero-Cost Proxies for Neural Architecture Search
  1000. Recouple Event Field via Probabilistic Bias for Event Extraction

Looking for submission deadlines instead? See the conference deadline calendar.