← All conferences

ICASSP 2024 Accepted Papers

The full list of 2,679 papers accepted at ICASSP 2024 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. "It os Okay to be Uncommon": Quantizing Sound Event Detection Networks on Hardware Accelerators with Uncommon Sub-Byte Support
  2. 1-D Spatial Attention in Binarized Convolutional Neural Networks
  3. 2D Human Pose Estimation Calibration and Keypoint Visibility Classification
  4. 3-D Near-Field Localization by Jointly Exploiting Spatial and Temporal Information Based on a Nonuniform Cross Array
  5. 3D Automated Quantitative Calculations Based on CT Images of the Hip Joint
  6. 3D Hand Joint and Grasping Estimation for Teleoperation System
  7. 3D Parallelism for Transformers via Integer Programming
  8. 3D Point Cloud Semantic Segmentation Based on Diffusion Model
  9. 3D Pose Estimation from Monocular Video with Camera-Bone Angle Regularization on the Image Feature
  10. 3DSAM: Segment Anything in NeRF
  11. 3M-Transformer: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction
  12. 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
  13. 6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on Self-Motioning Human
  14. A 3D Virtual Try-On Method with Global-Local Alignment and Diffusion Model
  15. A Bayesian Approach to High-Order Link Prediction
  16. A Bi-Pyramid Multimodal Fusion Method for the Diagnosis Of Bipolar Disorders
  17. A Binary BP Decoding Using Posterior Adjustment for Quantum LDPC Codes
  18. A Birgat Model for Multi-Intent Spoken Language Understanding with Hierarchical Semantic Frames
  19. A CCM-Based Joint DOA-Frequency Estimation and Signal Recovery with Efficient Sub-Nyquist Sampling
  20. A Chat about Boring Problems: Studying GPT-Based Text Normalization
  21. A Closer Look at Wav2vec2 Embeddings for On-Device Single-Channel Speech Enhancement
  22. A Codec-Based Approach for Video Life-Cycle Characterization in Social Networks
  23. A Comparative Analysis of Poetry Reading Audio: Singing, Narrating, or Somewhere in Between?
  24. A Comparative Study on Annotation Quality of Crowdsourcing and LLm Via Label Aggregation
  25. A Comparison of Parameter-Efficient ASR Domain Adaptation Methods for Universal Speech and Language Models
  26. A Complete Method for the 3D Reconstruction of Axonal Pathways from 2 Orthogonal 3D OCT Images of the Lamina Cribrosa
  27. A Comprehensive Framework for Occluded Human Pose Estimation
  28. A Computationally Efficient Semi-Blind Source Separation Approach for Nonlinear Echo Cancellation Based on an Element-Wise Iterative Source Steering
  29. A Concept for a Slam Back End Hardware Accelerator
  30. A Contrario Paradigm for Yolo-Based Infrared Small Target Detection
  31. A Convergent Primal-Dual Deep Plug-and-Play Algorithm for Constrained Image Restoration
  32. A Counterfactual Inspired Framework For Quantifying Edge Effects On Gnns Fairness
  33. A Cross Search Method for Data Augmentation in Neural Machine Translation
  34. A Crowdsourcing Approach to Video Quality Assessment
  35. A Deep Representation Learning-Based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder
  36. A DenseNet-Based Method for Decoding Auditory Spatial Attention with EEG
  37. A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Videos
  38. A Detailed Audio-Text Data Simulation Pipeline Using Single-Event Sounds
  39. A Distributed Joint Integrated Probabilistic Data Association (JIPDA) Filter with Soft Object Association
  40. A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection
  41. A Facial Expression Transfer Method Based on 3DMM and Diffusion Models
  42. A Fast Blind Deblurring Algorithm Using Local Gradient Product Prior
  43. A Fast, Performant, Secure Distributed Training Framework For LLM
  44. A Federated Graph to Embedding Approach for Knowledge Graph Completion
  45. A Fine-Grained Attribute Pre-Labeling Method Based on Label Dependency and Feature Similarity Dynamics
  46. A Fine-Grained Tri-Modal Interaction Model for Multimodal Sentiment Analysis
  47. A Flexible Online Framework for Projection-Based Stft Phase Retrieval
  48. A Foundation Model for Music Informatics
  49. A Framework for Portrait Stylization with Skin-Tone Awareness and Nudity Identification
  50. A Fully Differentiable Model for Unsupervised Singing Voice Separation
  51. A General Framework for Rotation Invariant Point Cloud Analysis
  52. A Generative Adversarial Framework for Dialogue Generation with Neural Architecture Search
  53. A Gibbs Sampler for Bayesian Nonparametric State-Space Models
  54. A Graph Neural Network Based Approach for Fault Delineation in Seismic Data using Graph Total Variation and Multigraph
  55. A Graph Neural Network Based Fusion of MRI-Derived Brain Network and Clinical Data for Glioblastoma Survival Prediction
  56. A Graph-Prediction-Based Approach for Debiasing Underreported Data
  57. A Green Learning Approach to Spoofed Speech Detection
  58. A Guided Upsampling Network for Short wave Infrared Images Using Graph Regularization
  59. A Hierarchical Multi-Proxy Loss with Dynamic Main-Proxy for Deep Metric Learning
  60. A Hybrid CNN-Transformer for Focal Liver Lesion Classification
  61. A Hybrid Deep-Online Learning Based Method for Active Noise Control in Wave Domain
  62. A Hybrid Slow-Time Coding Framework for Automotive MIMO Radar
  63. A Joint Data Compression and Time-Delay Estimation Distributed Systems via Extremum Encoding
  64. A Joint Look on Lunar Satellite and Cooperative Surface PNT
  65. A Keyless Extraction Framework Targeting at Deep Learning Based Image-Within-Image Models
  66. A Learning Resource Recommendation Algorithm Based on Online Learning Behavior
  67. A Learning-Based Multi-Node Fusion Positioning Method Using Wearable Inertial Sensors
  68. A Learning-Based System for Automatic Intentional Non-Adherence Detection from Dosing Videos
  69. A Light-Weight State Detection Model for Kalman-Filter-Based Acoustic Feedback Cancellation with Rapid Recovery from Abrupt Path Changes
  70. A Lightweight Change Detection Method Based on Feature Interaction and Transformer for High Resolution Remote Sensing Images
  71. A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity Detection
  72. A Low-Latency Fft-Ifft Cascade Architecture
  73. A Machine-Learning Model for Detecting Depression, Anxiety, and Stress from Speech
  74. A Meta-Preconditioning Approach for Deep Q-Learning
  75. A Method for Bilevel Optimization with Convex Lower-Level Problem
  76. A Method for X-Ray Image Landmarks Localization using Cyclic Coordinate-Guided Strategy
  77. A Modified Cramér-Rao Bound for Discrete-Time Markovian Dynamic Systems
  78. A Multi-Carrier Information Hiding Algorithm Based on Layered Compression of 3d Point Cloud Model
  79. A Multi-Scale Bimodal Fusion Network for Robust and Accurate Online Handwriting Recognition
  80. A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
  81. A Multiscale Objective Function for Camera Color Correction
  82. A Near-Field Source Localization Method for Uniform/Sparse Centrally Symmetric Rectangular Arrays
  83. A Neural Syntax Parser for Coronary Artery Anatomical Labeling in Coronary CT Angiography
  84. A Neurophysiological-Auditory "Listen Receipt" for Communication Enhancement
  85. A New Fourth-Order Sparse Array Generator Based on Sum-Difference Co-Array Analysis
  86. A New Perspective on Understanding Resolution Limit Via an Asymptotic Study of Christoffel-Darboux Kernel Based Spectrum Estimator
  87. A New Pre-Training Paradigm for Offline Multi-Agent Reinforcement Learning with Suboptimal Data
  88. A New Similarity-Based Relational Knowledge Distillation Method
  89. A Novel 3-D Focusing Scheme for Distributed SAR Tomography
  90. A Novel Cascade Instruction Tuning Method for Biomedical NER
  91. A Novel Contrastive Diffusion Graph Convolutional Network for Few-Shot Skeleton-Based Action Recognition
  92. A Novel Cross-Sensor Self-Supervised Learning Method for Rotating Machinery Fault Diagnosis
  93. A Novel Demodulation and Selection Pilot Power Trade-Off for Codebook-Based IRS with Imperfect Channel Estimates
  94. A Novel Discrete Fractional Complex Hadamard Transform for Medical Image Encryption
  95. A Novel Iterative Thresholding Algorithm for Arctangent Regularization Problem
  96. A Novel Local-Global Feature Fusion Framework for Body-Weight Exercise Recognition with Pressure Mapping Sensors
  97. A Novel Medical Image Fusion Framework Integrating Multi-scale Encoder-Decoder with Discrete Wavelet Decomposition
  98. A Novel Multi-Atlas Fusion Model Based On Contrastive Learning For Functional Connectivity Graph Diagnosis
  99. A Novel Multimodal Sentiment Analysis Model Based on Gated Fusion and Multi-Task Learning
  100. A Novel Residual-Guided Learning Method for Image Steganography
  101. A One-Class Approach to Detect Super-Resolution Satellite Imagery with Spectral Features
  102. A PLS-Integrated Lasso Method With Application in Index Tracking
  103. A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications
  104. A Practical Online Multichannel Dereverberation Approach with Data-Reuse Technique
  105. A Prior Driven Semi-Supervised ViTGAN for Image Recolorization
  106. A Probability Gradient Based Approach for Sampling Boundaries of In-Domain Data
  107. A Prompt-Based Method with Multi-View Optimization for Open Relation Extraction
  108. A Property-Guided Diffusion Model For Generating Molecular Graphs
  109. A Real-Time Active Speaker Detection System Integrating an Audio-Visual Signal with a Spatial Querying Mechanism
  110. A Real-Time Lyrics Alignment System Using Chroma and Phonetic Features for Classical Vocal Performance
  111. A Real-Time Video Quality Metric for HTTP Adaptive Streaming
  112. A Reconstruction-Based Feature Adaptation for Anomaly Detection with Self-Supervised Multi-Scale Aggregation
  113. A Reduced-Reference Quality Assessment Metric for Textured Mesh Digital Humans
  114. A Relation-Aware Heterogeneous Graph Transformer on Dynamic Fusion for Multimodal Classification Tasks
  115. A Riemannian-Based Joint Design Framework of Mimo Radar Transmit Waveform And Receive Filter Via Information Theory
  116. A Robust Audio Deepfake Detection System via Multi-View Feature
  117. A Robust GLRT Detector Against Missing Data in Cooperative Sensing
  118. A Robust Pitch-Fusion Model for Speech Emotion Recognition in Tonal Languages
  119. A Robust Quantile Huber Loss with Interpretable Parameter Adjustment in Distributional Reinforcement Learning
  120. A Robust and Scalable Method with an Analytic Solution for Multi-Subject FMRI Data Analysis
  121. A Saliency Enhanced Feature Fusion Based Multiscale RGB-D Salient Object Detection Network
  122. A Scalable Sparse Transformer Model for Singing Melody Extraction
  123. A Self-Supervised Pressure Map Human Keypoint Detection Approch: Optimizing Generalization and Computational Efficiency Across Datasets
  124. A Separation Priority Pipeline for Single-Channel Speech Separation in Noisy Environments
  125. A Sequential Averaging Plug-and-Play Method for Image Restoration Via Fixed-Point Projection
  126. A Simple and Effective Method for Anomaly Detection on Attributed Graphs via Feature Consistency
  127. A Smoothed Bregman Proximal Gradient Algorithm for Decentralized Nonconvex Optimization
  128. A Soft Contrastive Learning-Based Prompt Model for Few-Shot Sentiment Analysis
  129. A Sound Approach: Using Large Language Models to Generate Audio Descriptions for Egocentric Text-Audio Retrieval
  130. A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition
  131. A Speaker Recognition Method Based on Stable Learning
  132. A Spectral Analysis of Graph Neural Networks on Dense and Sparse Graphs
  133. A Statistical Characterization Of Communication Performance In RIS-Aided Networks
  134. A Steered Response Power Approach with Bilinear Prediction-Based Trade-Off Prewhitening for Speaker Localization
  135. A Stochastic Gradient Approach for Communication Efficient Confederated Learning
  136. A Stochastic Proximal WMMSE for Ergodic Sum Rate Maximization
  137. A Study of Mispronunciation Detection and Diagnosis Based on Meta-Learning
  138. A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction
  139. A Study on Combining Non-Parallel and Parallel Methodologies for Mandarin-English Cross-Lingual Voice Conversion
  140. A Study on Graph Embedding for Speaker Recognition
  141. A Study on the Adverse Impact of Synthetic Speech on Speech Recognition
  142. A Supervised Information Enhanced Multi-Granularity Contrastive Learning Framework for EEG Based Emotion Recognition
  143. A Targeted Adversarial Attack Method for Multi-Classification Malicious Traffic Detection
  144. A Transformer Approach for Polyphonic Audio-to-Score Transcription
  145. A Tri-Dynamic Preprocessing Framework for UGC Video Compression
  146. A Two-Stage Dehazing Framework Based on Inverted Image Curve-Enhancement
  147. A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement
  148. A Unified DNN-Based System for Industrial Pipeline Segmentation
  149. A Unified Framework for Multi-Intent Spoken Language Understanding with Prompting
  150. A Unified Front-End Framework for English Text-to-Speech Synthesis
  151. A Unified Loss Function to Tackle Inter-Class and Intra-Class Data Imbalance in Sound Event Detection
  152. A Variable Smoothing for Nonconvexly Constrained Nonsmooth Optimization with Application to Sparse Spectral Clustering
  153. A Wasserstein Graph Distance Based on Distributions of Probabilistic Node Embeddings
  154. A Weighted-Variance Variational Autoencoder Model for Speech Enhancement
  155. AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
  156. ADHD Diagnosis and Biomarker Detection Based on Multimodal Graph Convolutional Neural Network
  157. ADIFT: Zero-Shot Generative Model Adaption Via Adaptive Domain-Invariant Feature Transfer
  158. ADVSV: An Over-the-Air Adversarial Attack Dataset for Speaker Verification
  159. AEAM3D: Adverse Environment-Adaptive Monocular 3D Object Detection via Feature Extraction Regularization
  160. AEGIS-Net: Attention-Guided Multi-Level Feature Aggregation for Indoor Place Recognition
  161. AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
  162. AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh Recovery
  163. ANM-Based Source Localization Under Mixed Field
  164. AQF: Assessing the Quality of Hyperspectral Reconstruction with a Learnable Metric
  165. ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal Prediction
  166. AS-pVAD: A Frame-Wise Personalized Voice Activity Detection Network with Attentive Score Loss
  167. ASPED: An Audio Dataset for Detecting Pedestrians
  168. AUTOSGM: A Unified Lowpass Regularization Framework for Accelerated Learning
  169. AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models
  170. AV2WAV: Diffusion-Based Re-Synthesis from Continuous Self-Supervised Features for Audio-Visual Speech Enhancement
  171. Accelerated Recovery of Spectrally Sparse Signals Viamodified Proximal Gradient in Hankel Space
  172. Accelerating Gradient Descent for Over-Parameterized Asymmetric Low-Rank Matrix Sensing via Preconditioning
  173. Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition
  174. Accurate Gigapixel Crowd Counting by Iterative Zooming and Refinement
  175. Accurate Interpolation of Scattered Data Via Learning Relation Graph
  176. Accurate and Robust Scene Text Recognition via Adversarial Training
  177. Acoustic BPE for Speech Generation with Discrete Tokens
  178. Activation Compression of Graph Neural Networks Using Block-Wise Quantization with Improved Variance Minimization
  179. Active Explainable Recommendation with Limited Labeling Budgets
  180. Active Learning for Sound Event Classification Using Bayesian Neural Networks with Gaussian Variational Posterior
  181. Active Learning with Core-Set Sampling and Scale-Sensitive Loss for 3D Object Detection
  182. Active Noise Control Over 3D Space with A Dynamic Noise Source
  183. Active Noise Control Over A Large Region with Multiple Spherical Microphone Arrays In Wave Domain
  184. Activity Recognition Method Based on Kernel Supervised Laplacian Eigenmaps
  185. AdaFL: Adaptive Client Selection and Dynamic Contribution Evaluation for Efficient Federated Learning
  186. AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
  187. AdaPlus: Integrating Nesterov Momentum and Precise Stepsize Adjustment on Adamw Basis
  188. Adapter-Based Incremental Learning for Face Forgery Detection
  189. Adapting Frechet Audio Distance for Generative Music Evaluation
  190. Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition
  191. Adapting Pitch-Based Self Supervised Learning Models for Tempo Estimation
  192. Adaptive Chroma Block Vector Derivation from Luma for Screen Content Coding
  193. Adaptive Confidence Multi-View Hashing for Multimedia Retrieval
  194. Adaptive Data Augmentation for Aspect Sentiment Quad Prediction
  195. Adaptive Fourier Decomposition Based Signal Extraction on Weak Electromagnetic Field
  196. Adaptive Gaussian Regularization Constrained Sparse Subspace Clustering for Image Segmentation
  197. Adaptive Grid 2-D Direction of Arrival Estimation Method Using an Integrated Dictionary
  198. Adaptive Head Pose Estimation with Real-Time Structured Light
  199. Adaptive Image-Enhanced Knowledge Graph Completion
  200. Adaptive Joint Channel Estimation/Data Detection in Flexible Multicarrier Mimo Systems - A Tensor-Based Approach
  201. Adaptive Kalmannet: Data-Driven Kalman Filter with Fast Adaptation
  202. Adaptive Multi-Armed Bandit Learning for Task Offloading in Mobile Edge Computing
  203. Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance Fields
  204. Adaptive Multi-View Joint Contrastive Learning on Graphs
  205. Adaptive Multiview Community-Preserved Graph Convolutional Network for Multiatlas-Based Functional Connectivity Analysis
  206. Adaptive Order Aggregator and Extractor Graph Neural Network
  207. Adaptive Parameter Sharing for Multi-Agent Reinforcement Learning
  208. Adaptive Pedestrian Trajectory Prediction via Target-Directed Angle Augmentation
  209. Adaptive Prompt Construction Method for Relation Extraction
  210. Adaptive Quantization with Mixed-Precision Based on Low-Cost Proxy
  211. Adaptive Reweighted Sparse Belief Propagation Decoding for Polar Codes
  212. Adaptive Secondary Transform Sets for Video Coding Beyond AV1
  213. Adaptive Sensor Selection with Deterministic Priors for DoA Tracking
  214. Adaptive Spatial-Temporal Hypergraph Fusion Learning for Next POI Recommendation
  215. Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
  216. Adaptive Super Resolution for One-Shot Talking-Head Generation
  217. Adaptive Video Watermarking with Perceptual Guarantee and Efficiency Optimization
  218. Adaptive-Avg-Pooling Based Attention Vision Transformer for Face Anti-Spoofing
  219. Addressing Confounds in Functional Connectivity Analyses of Calcium Imaging
  220. Addressing Data Scarcity in Voice Disorder Detection with Self-Supervised Models
  221. AdvShadow: Evading DeepFake Detection via Adversarial Shadow Attack
  222. AdvTTS: Adversarial Text-to-Speech Synthesis Attack on Speaker Identification Systems
  223. Advancing Acoustic Howling Suppression Through Recursive Training of Neural Networks
  224. Adversarial Domain Adaptation for Classification with Nested Dichotomies
  225. Adversarial Jamming for Autoencoder Distribution Matching
  226. Adversarial Learning on Compressed Posterior Space for Non-Iterative Score-based End-to-End Text-to-Speech
  227. Adversarial Robustness of Convolutional Models Learned in the Frequency Domain
  228. Adversarial Speech for Voice Privacy Protection from Personalized Speech Generation
  229. Aerial-IRS-Assisted Load Balancing In Downlink Networks
  230. Ainur: Harmonizing Speed and Quality in Deep Music Generation Through Lyrics-Audio Embeddings
  231. Align, Adapt and Inject: Audio-Guided Image Generation, Editing and Stylization
  232. All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays
  233. Alleviating Hallucinations Via Supportive Window Indexing in Abstractive Summarization
  234. Alpharotate: A Rotation Detection Benchmark Using Tensorflow
  235. Ambisonics Networks - The Effect of Radial Functions Regularization
  236. An Accurate and Efficient Neural Network for OCTA Vessel Segmentation and a New Dataset
  237. An Active Noise Control System Based On Soundfield Interpolation Using A Physics-Informed Neural Network
  238. An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks
  239. An Adaptive Algorithm for Tracking Third-Order Coupled Canonical Polyadic Decomposition
  240. An Anchor Learning Approach for Citation Field Learning
  241. An Asymptotically Achievable Rate Bound for Establishing High-Fidelity Entanglements in Quantum Networks
  242. An Attention-Enhanced Retentive Broad Learning System for Subject-Generic Emotion Recognition from EEG Signals
  243. An Audio-Textual Diffusion Model for Converting Speech Signals into Ultrasound Tongue Imaging Data
  244. An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
  245. An Efficient Algorithm For Clustered Multi-Task Compressive Sensing
  246. An Efficient Algorithm for Multiuser Sum-Rate Maximization of Large-Scale Active RIS-Aided MIMO System
  247. An Efficient Alternating Riemannian/Projected Gradient Descent Ascent Algorithm for Fair Principal Component Analysis
  248. An Efficient Hierarchical Block Coordinate Descent Method for Time-Varying Graphical Lasso
  249. An Efficient Temporary Deepfake Location Approach Based Embeddings for Partially Spoofed Audio Detection
  250. An Efficient Transformer For Demosaicing Via Compressed Multi-Branch Attention Mechanism
  251. An Efficient and Interpre Table Speech Enhancement Network Via Deep Dictionary Learning
  252. An Empirical Investigation of Domain Adaptation Ability for Chinese Spelling Check Models
  253. An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement
  254. An End-to-End EEG Channel Selection Method with Residual Gumbel Softmax for Brain-Assisted Speech Enhancement
  255. An Error Self-Corrected DOA Estimation Model for Sparse Array Based on ANM
  256. An Experimental Comparison of Multi-View Self-Supervised Methods for Music Tagging
  257. An Experimental Comparison of Noise-Robust Text-To-Speech Synthesis Systems Based On Self-Supervised Representation
  258. An Explainable Proxy Model for Multilabel Audio Segmentation
  259. An Explicit Multi-Modal Fusion Method for Sign Language Translation
  260. An Initial Investigation of Neural Replay Simulator for Over-The-Air Adversarial Perturbations to Automatic Speaker Verification
  261. An Interpretable and Generalizable Speech Detector Based on a CNN-LSTM Framework
  262. An Investigation of Distribution Alignment in Multi-Genre Speaker Recognition
  263. An MVDR-Embedded U-Net Beamformer for Effective and Robust Multichannel Speech Enhancement
  264. An Optimized Interleaved OFDM Chirp Orthogonal Waveform Design for Dechirped Miniature MMW MIMO Radar
  265. An Unsupervised Segmentation of Vocal Breath Sounds
  266. Analysis of High-Order Brain Networks Resolved in Time and Frequency Using CP Decomposition
  267. Analysis of an Elliptic Localization Algorithm Using Fixed Point Iteration
  268. Analysis of the Memorization and Generalization Capabilities of AI Agents: are Continual Learners Robust?
  269. Analysis of the SINR in LEO-PNT Systems with 5G PRS Multiplexing: Integration of PRS and NTN
  270. Analyzing Adversarial Vulnerabilities of Graph Lottery Tickets
  271. Anchor-Guided GAN with Contrastive Loss for Low-Resource Out-of-Domain Detection
  272. Anim-400K: A Large-Scale Dataset for Automated End to End Dubbing of Video
  273. Anomalous Sound Detection by Feature-Level Anomaly Simulation
  274. Anomaly Detection from a Frequency Perspective: M-Band Wavelet Packet Anomaly Detection Network
  275. Anomaly-Aware Semantic Self-Alignment Framework for Video-Based Person Re-Identification
  276. Anonymizing Speaker Voices: Easy to Imitate, Difficult to Recognize?
  277. Anti-Deception Jamming Power Optimization Strategy for Multi-Target Tracking Tasks in Multi-Radar Systems
  278. Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection
  279. Application of SNNS Model Based On Multi-Dimensional Attention In Drone Radio Frequency Signal Classification
  280. Applying Hybrid Quantum LSTM for Indoor Localization Based on RSSI
  281. Arbitrary Style Transfer Based on Content Integrity and Style Consistency Enhancement
  282. Arbitrary Style Transfer with Prototype-Based Channel Alignment
  283. Architecture-Agnostic Iterative Black-Box Certified Defense Against Adversarial Patches
  284. Are Deep Neural Networks Robust to Named Entities? An Adversarial Attack and Defense Perspective
  285. Are SNNs Truly Energy-efficient? - A Hardware Perspective
  286. Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?
  287. Array Geometry Optimization for Region-of-Interest Near-Field Beamforming
  288. Asformer: Learning From Adjacent Scale
  289. Assessing GNSS Carrier-to-Noise-Density Ratio Estimation in The Presence of Meaconer Interference
  290. Assessing Vibroacoustic Sound Massage Through The Biosignal of Human Speech: Evidence of Improved Wellbeing
  291. Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification
  292. Asymptotic Behavior of Super-Resolution Sparse Bayesian Learning
  293. Asymptotically Tight Misspecified Bayesian Cramér-Rao Bound
  294. Asynchronous Diffusion Learning with Agent Subsampling and Local Updates
  295. AttA-NET: Attention Aggregation Network for Audio-Visual Emotion Recognition
  296. AttHear: Explaining Audio Transformers Using Attention-Aware NMF
  297. Attention Decoupling for Query-Based Object Detection
  298. Attention Is All You Need For Blind Room Volume Estimation
  299. Attention-Based Spatial-Frequency Information Network for Underwater Single Image Super-Resolution
  300. Attention-Driven Multichannel Speech Enhancement in Moving Sound Source Scenarios
  301. Attention-Guided Adaptation for Code-Switching Speech Recognition
  302. AttentionLUT: Attention Fusion-Based Canonical Polyadic LUT for Real-Time Image Enhancement
  303. Attr-Int: A Simple and Effective Entity Alignment Framework for Heterogeneous Knowledge Graphs
  304. Attribute-Aware Amplification of Facial Feature Sequences for Facial Emotion Recognition
  305. Attribute-Aware Head Swapping Guided by 3d Modeling
  306. Attribution-Based Scanline Perturbation Attack on 3d Detectors of Lidar Point Clouds
  307. Audio Deepfake Detection With Self-Supervised Wavlm And Multi-Fusion Attentive Classifier
  308. Audio Difference Learning for Audio Captioning
  309. Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
  310. Audio Prompt Tuning for Universal Sound Separation
  311. Audio Transformer for Synthetic Speech Detection via Formant Magnitude and Phase Analysis
  312. Audio-Aided Learning Framework for Image Classification with Limited Training Images
  313. Audio-Free Prompt Tuning for Language-Audio Models
  314. Audio-Journey: Open Domain Latent Diffusion Based Text-To-Audio Generation
  315. Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker Speech
  316. Audio-Visual Child-Adult Speaker Classification in Dyadic Interactions
  317. Audio-Visual Speech Recognition In-The-Wild: Multi-Angle Vehicle Cabin Corpus and Attention-Based Method
  318. Audiosr: Versatile Audio Super-Resolution at Scale
  319. Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain
  320. Auditory Cortex-Inspired Spectral Attention Modulation for Binaural Sound Localization in HRTF Mismatch
  321. AugSumm: Towards Generalizable Speech Summarization Using Synthetic Labels from Large Language Models
  322. Augment on Manifold: Mixup Regularization with UMAP
  323. Augmenting Conformers With Structured State-Space Sequence Models For Online Speech Recognition
  324. Augmenting Transformer Autoencoders with Phenotype Classification for Robust Detection of Psychotic Relapses
  325. AutoCali: Enhancing AoA-based Indoor Localization through Automatic Phase Calibration
  326. AutoFGNN: A Framework for Extracting All Frequency Information from Large-Scale Graphs
  327. AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech Data
  328. AutoSen: Improving Automatic WiFi Human Sensing through Cross-Modal Autoencoder
  329. Automated Labeling of Automotive Radar Azimuth Multipath
  330. Automatic Channel Selection and Spatial Feature Integration for Multi-Channel Speech Recognition Across Various Array Topologies
  331. Automatic Design of Adapter Architectures for Enhanced Parameter-Efficient Fine-Tuning
  332. Automatic Detection Of Sleepiness-Related Syndromes and Symptoms Using Voice and Speech Biomarkers
  333. Automatic Recognition of Gesture Identity and Onset of Cued-Speech
  334. Automatic Speech Recognition Tuned for Child Speech in the Classroom
  335. Automatic Temporal Alignment for Pitch Estimation Evaluation
  336. Automotive Radar Interference Characterization: FMCW or PMCW?
  337. Automotive Radar Interference Mitigation Via SINR Maximization
  338. Automotive Radar Point Cloud Parametric Density Estimation using Camera Images
  339. Autonomous Generative Feature Replay for Non-Exemplar Class-Incremental Learning
  340. Autoregressive 3D Shape Completion via Sphere-Guided Disentangled Representation
  341. Autost: Training-Free Neural Architecture Search For Spiking Transformers
  342. Axis Order Invariance Learned from Point Clouds
  343. BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution
  344. BCC: Bidirectional Consistency Constraint Method for Hierarchical Text Classification
  345. BEVLOC: End-to-End 6-DoF Localization Via Cross-Modality Correlation Under Bird's Eye View
  346. BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation
  347. BFRFormer: Transformer-Based Generator for Real-World Blind Face Restoration
  348. BIGVSAN: Enhancing Gan-Based Neural Vocoders with Slicing Adversarial Network
  349. BNMTrans: A Brain Network Sequence-Driven Manifold-Based Transformer for Cognitive Impairment Detection Using EEG
  350. BPDO: Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text Detection
  351. BRAVEn: Improving Self-supervised pre-training for Visual and Auditory Speech Recognition
  352. BWSNET: Automatic Perceptual Assessment of Audio Signals
  353. Balanced And Discriminative Contrastive Learning For Class-Imbalanced Medical Images
  354. Balanced Learning for Multi-Domain Long-Tailed Speaker Recognition
  355. Balancing Easy and Hard Distortions: A Multi-Rate Knowledge Distillation Strategy for Blind Image Quality Assessment
  356. Balancing Representation Abstractions and Local Details Preservation for 3d Point Cloud Quality Assessment
  357. Balancing Speaker-Rater Fairness for Gender-Neutral Speech Emotion Recognition
  358. Ballistocardiogram-Based Heart Rate Variability Estimation for Stress Monitoring using Consumer Earbuds
  359. Bandwidth-Efficient Inference for Nerual Image Compression
  360. Bass Accompaniment Generation Via Latent Diffusion
  361. Batch Substitution Calibration of a Mems Microphone Array : Impact of Sensor Performance Dispersion on Directivity Estimation
  362. Bayesian Activity Detection for Massive Connectivity in Cell-Free IoT Networks
  363. Bayesian Learning-Based Kalman Smoothing For Linear Dynamical Systems With Unknown Sparse Inputs
  364. Bayesian Optimization with Gaussian Processes for Robust Localization
  365. Bayesian Topology Inference on Partially Known Networks from Input-Output Pairs
  366. Bayesian-Boosted MetaLoc: Efficient Training and Guaranteed Generalization for Indoor Localization
  367. Beamforming Design and Performance Evaluation for RIS-Aided Localization Using LEO Satellite Signals
  368. Beamforming Through Online Convex Combination of Differential Beamformers
  369. Beast: Online Joint Beat and Downbeat Tracking Based on Streaming Transformer
  370. Benchmarking Adversarial Robustness of Image Shadow Removal with Shadow-Adaptive Attacks
  371. Beta Quantile Regression for Robust Estimation of Uncertainty in the Presence of Outliers
  372. Beyond Empirical Windowing: An Attention-Based Approach for Trust Prediction In Autonomous Vehicles
  373. Beyond Simple Text Style Transfer: Unveiling Compound Text Style Transfer with Prompt-Based Pre-Trained Language Models
  374. Beyond the Limit of Weight-Sharing: Pioneering Space-Evolving NAS with Large Language Models
  375. Beyond the Snowfall: Enhancing Snowy Day Object Detection Through Progressive Restoration and Multi-Feature Fusion
  376. Bi-Directional Motion Attention with Contrastive Learning for few-shot Action Recognition
  377. Binary Signal Alignment: Optimal Solution is Polynomial-Time and Linear-Time Solution is Quasi-Optimal
  378. Binaural Angular Separation Network
  379. Binaural Rendering of Heterogeneous Sound Sources with Extent
  380. Binaural Room Transfer Function Interpolation Via System Inversion
  381. Binaural Sound Source Localization Using a Hybrid Time and Frequency Domain Model
  382. Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
  383. Binauralmusic: A Diverse Dataset for Improving Cross-Modal Binaural Audio Generation
  384. Biomimetic Mappings for Active Sonar Object Recognition in Clutter
  385. Blenda: Domain Adaptive Object Detection Through Diffusion-Based Blending
  386. Blind Beamforming for Intelligent Reflecting Surface: A Reinforcement Learning Approach
  387. Blind Deconvolution of Sparse Graph Signals in the Presence of Perturbations
  388. Blind Estimation of Audio Effects Using an Auto-Encoder Approach and Differentiable Digital Signal Processing
  389. Blind Inpainting with Object-Aware Discrimination for Artificial Marker Removal
  390. Blind Separation of Noisy Mixtures Over Galois Fields
  391. Block Adaptive Subspace Pursuit Method for Wall Clutter Mitigation
  392. Boosting Adversarial Robustness Distillation Via Hybrid Decomposed Knowledge
  393. Boosting End-to-End Multilingual Phoneme Recognition Through Exploiting Universal Speech Attributes Constraints
  394. Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori Estimation
  395. Boosting LLMS with Ontology-Aware Prompt for Ner Data Augmentation
  396. Boosting Pruned Networks with Linear Over-Parameterization
  397. Boosting Speech Enhancement with Clean Self-Supervised Features Via Conditional Variational Autoencoders
  398. Boosting Unknown-Number Speaker Separation with Transformer Decoder-Based Attractor
  399. Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer
  400. Boosting Zero-Shot Node Classification via Dependency Capture and Discriminative Feature Learning
  401. Boosting of Implicit Neural Representation-Based Image Denoiser
  402. Bootstrap Predictive Coding: Investigating a Non-Contrastive Self-Supervised Learning Approach
  403. Boundary-Driven Active Learning for Anomaly Detection in Time Series Data Streams
  404. Bounding Box-Guided Pseudo Point Clouds Early-Fusion and Density Optimize for 3D Object Detection
  405. Brain Structure-Function Interaction Network for Fluid Cognition Prediction
  406. BrainFC-CGAN: A Conditional Generative Adversarial Network for Brain Functional Connectivity Augmentation and Aging Synthesis
  407. Branchformer-Based TDNN for Automatic Speaker Verification
  408. Breaking Speaker Recognition with Paddingback
  409. Breaking the Barrier: Selective Uncertainty-Based Active Learning for Medical Image Segmentation
  410. Breast Ultrasound Computer-Aided Diagnosis Using Structure-Aware Triplet Path Networks
  411. Bregman Graph Neural Network
  412. Bridging The Domain Gap Arising from Text Description Differences for Stable Text-To-Image Generation
  413. Bridging the Gap: A Self-Learning Model Using Implicit Knowledge for Chinese Spelling Correction
  414. Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning
  415. Bridging the Gaps of Both Modality and Language: Synchronous Bilingual CTC for Speech Translation and Speech Recognition
  416. Bringing the Discussion of Minima Sharpness to the Audio Domain: A Filter-Normalised Evaluation for Acoustic Scene Classification
  417. Broadband Personal Sound Zone Control in the Presence of Nonlinearities
  418. Buffered Gaussian Modeling for Vectorized HD Map Construction
  419. Build a 50+ Hours Chinese Mandarin Corpus for Children's Speech Recognition
  420. Building Lane-Level Maps from Aerial Images
  421. ByteHum: Fast and Accurate Query-by-Humming in the Wild
  422. C-CLAPA: Improving Text-Audio Cross Domain Retrieval with Captioning and Augmentations
  423. CAG-FPN: Channel Self-Attention Guided Feature Pyramid Network for Object Detection
  424. CAGEN: Controllable Anomaly Generator using Diffusion Model
  425. CALSeg: Improving Calibration of Medical Image Segmentation Via Variational Label Smoothing
  426. CC-DA: Cross-Domain Consistency Data Augmentation for 3D Tumor Segmentation
  427. CDA-MBPO: Corrected Data Aggregation for Model-Based Policy Optimization
  428. CDCNet: A Fast and Lightweight Dehazing Network with Color Distortion Correction
  429. CDUMA: An Adaptive Approach for Mitigating Confounder for MCQA
  430. CED: Consistent Ensemble Distillation for Audio Tagging
  431. CEDNet: A Continuous Emotion Detection Network for Naturalistic Stimuli Using MEG Signals
  432. CEMOAE: A Dynamic Autoencoder with Masked Channel Modeling for Robust EEG-Based Emotion Recognition
  433. CENet: Content-Aware Enhanced Network for Practical Scene Parsing
  434. CGN: A Simple Yet Effective Multi-Channel Gated Network for Long-Term Time Series Forecasting
  435. CIF-RNNT: Streaming ASR Via Acoustic Word Embeddings with Continuous Integrate-and-Fire and RNN-Transducers
  436. CIF-T: A Novel CIF-Based Transducer Architecture for Automatic Speech Recognition
  437. CKT-RCM: Clip-Based Knowledge Transfer and Relational Context Mining for Unbiased Panoptic Scene Graph Generation
  438. CLAF: Contrastive Learning with Augmented Features for Imbalanced Semi-Supervised Learning
  439. CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision
  440. CLIP-Font: Sementic Self-Supervised Few-Shot Font Generation with Clip
  441. CLIP-MSA: Incorporating Inter-Modal Dynamics and Common Knowledge to Multimodal Sentiment Analysis With Clip
  442. CLPSD: Detecting Ethereum Phishing Scams based on Curriculum Learning
  443. CLT: Cooperative Lottery Ticket Hypothesis in Live Streaming Sales Prediction
  444. CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video Parsing
  445. CNFA: Conditional Normalizing Flow for Query-Limited Attack
  446. COLLD: Contrastive Layer-to-Layer Distillation for Compressing Multilingual Pre-Trained Speech Encoders
  447. COLORFLOW: A Conditional Normalizing Flow for Image Colorization
  448. COPHTC: Contrastive Learning with Prompt Tuning for Hierarchical Text Classification
  449. CORAAL QA: A Dataset and Framework for Open Domain Spontaneous Speech Question Answering from Long Audio Files
  450. CPAUG: Refining Copy-Paste Augmentation for Speech Anti-Spoofing
  451. CPMSVD: Cross-Project Multiclass Software Vulnerability Detection Via Fused Deep Feature and Domain Adaptation
  452. CRC-Aided Learned Ensembles of Belief-Propagation Polar Decoders
  453. CROCFUN: Cross-Modal Conditional Fusion Network for Pansharpening
  454. CROSSWORD: A Semantic Approach To Text Compression Via Masking
  455. CReStyler: Text-Guided Single Image Style Transfer Method Based on CNN and Restormer
  456. CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning
  457. CSI-Free Over-The-Air Decentralized Learning Over Frequency Selective Channels
  458. CSNet: Contrastive Siamese Network for Robust SLU
  459. CST-Former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection
  460. CT and MRI Fusion with Anisotropic Guided Filtering
  461. CUTDEM: Depth-Aware Enhanced Multi-View Image Mixing for Light Field Super-Resolution
  462. Camera Calibration using a Single View of a Symmetric Object
  463. Camera-Radar Association for Data Annotation
  464. Can ChatGPT Serve as a Multi-Criteria Decision Maker? A Novel Approach to Supplier Evaluation
  465. Can LLM Find the Green Circle? Investigation and Human-Guided Tool Manipulation for Compositional Generalization
  466. Can Large-Scale Vocoded Spoofed Data Improve Speech Spoofing Countermeasure with a Self-Supervised Front End?
  467. Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
  468. Can We Trust Explainable AI Methods on ASR? An Evaluation on Phoneme Recognition
  469. Can Whisper Perform Speech-Based In-Context Learning?
  470. Caption Unification for Multi-View Lifelogging Images Based on In-Context Learning with Heterogeneous Semantic Contents
  471. Capturing Detail Variations for Lightweight Neural Radiance Fields
  472. Cardinality-Constrained Binary Quadratic Optimization via Extreme Point Pursuit, with Application to the Densest K-Subgraph Problem
  473. CartoonDiff: Training-free Cartoon Image Generation with Diffusion Transformer Models
  474. Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning for Visual Story Synthesis
  475. CausalME: Balancing bi-modalities in Visual Question Answering
  476. Causality-Inspired Single-Source Domain Generalization for Face Anti-Spoofing
  477. Causally Uncovering Bias in Video Micro-Expression Recognition
  478. Center of Pressure Estimation by Analyzing Walking Videos
  479. Changenet: Multi-Temporal Asymmetric Change Detection Dataset
  480. Channel Estimation and Prediction in Wireless Communications Assisted by Semi-Passive RIS
  481. Channel Estimation in Underdetermined Systems Utilizing Variational Autoencoders
  482. Channel-Spatial Transformer for Efficient Image Super-Resolution
  483. Character Attribute Extraction from Movie Scripts Using LLMs
  484. Chat: Cascade Hole-Aware Transformers with Geometric Spatial Consistency for Accurate Monocular Endoscopic Depth Estimation
  485. Child FER: Domain-Agnostic Facial Expression Recognition in Children Using a Secondary Image Diffusion Model
  486. Chunked Attention-Based Encoder-Decoder Model for Streaming Speech Recognition
  487. Circular Decomposition and Cross-Modal Recombination for Multimodal Sentiment Analysis
  488. Class-Incremental Learning for Multi-Label Audio Classification
  489. Class-Wise Buffer Management for Incremental Object Detection: An Effective Buffer Training Strategy
  490. Class: Continual Learning Approach for Speech Super-Resolution
  491. Classification-Oriented Semantic Wireless Communications
  492. Client-Free Federated Unlearning via Training Reconstruction with Anchor Subspace Calibration
  493. Clinical Scores Prediction and Medication Adjustment for Course of Parkinson's Disease
  494. Clip-Based Synergistic Knowledge Transfer for text-based Person Retrieval
  495. Cliprerank: An Extremely Simple Method For Improving Ad-Hoc Video Search
  496. Close-Range Direction of Arrival Estimation in the Presence of Clock Jitter
  497. Co-Occurrence Graph-Enhanced Hierarchical Prediction of ICD Codes
  498. Co-Salient Object Detection via Discriminative Prototypes Contrast
  499. CoQ: AN Empirical Framework for Multi-hop Question Answering Empowered by Large Language Models
  500. CoSLR: Contrastive Chinese Sign Language Recognition with prior knowledge And Multi-Tasks Joint Learning
  501. Coding for the Unsourced B-Channel with Erasures: Enhancing the Linked Loop Code
  502. Cognitive Virtual Sensing Technique for Feedforward Active Noise Control
  503. Collaborative Watermarking for Adversarial Speech Synthesis
  504. Color Agnostic Cross-Spectral Disparity Estimation
  505. Combining Conformer and Dual-Path-Transformer Networks for Single Channel Noisy Reverberant Speech Separation
  506. CommIN: Semantic Image Communications as an Inverse Problem with INN-Guided Diffusion Models
  507. Communication Efficient Private Federated Learning Using Dithering
  508. Communication-Efficient Decentralized Dynamic Kernel Learning
  509. Communication-Efficient Federated Learning Through Adaptive Weight Clustering And Server-Side Distillation
  510. Communication-Efficient Federated Optimization over Semi-Decentralized Networks
  511. Communication-Efficient Laplace Mechanism for Differential Privacy via Random Quantization
  512. Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
  513. Communication-Oriented Automatic Assessment System for Accented Spoken Chinese in Read-Aloud Tasks
  514. Compact and De-Biased Negative Instance Embedding for Multi-Instance Learning on Whole-Slide Image Classification
  515. Comparable Demonstrations Are Important In In-Context Learning: A Novel Perspective On Demonstration Selection
  516. Comparative Study of Tokenization Algorithms for End-to-End Open Vocabulary Keyword Detection
  517. Comparing and Combining Audio Processing and Deep Learning Features for Classification of Heartbeat Sounds
  518. Comparing data-Driven and Handcrafted Features for Dimensional Emotion Recognition
  519. Comparison Of Frequency-Fusion Mechanisms For Binaural Direction-Of-Arrival Estimation For Multiple Speakers
  520. Comparison of Conditions for Omnidirectional Video with Spatial Audio in Terms of Subjective Quality and Impacts on Objective Metrics Resolving Power
  521. Complementary Fusion Network Based on Frequency Hybrid Attention for Pansharpening
  522. Complex Bounded Component Analysis: Identifiability and Algorithm
  523. Complexity Reduction of Template Matching-Based Reference Picture Padding in Video Coding
  524. Complexity Scaling for Speech Denoising
  525. Composite Federated Learning with Heterogeneous Data
  526. Computational Complexity of Asynchronous Policy Iteration for Two-Player Zero-Sum Markov Games
  527. Computing an Entire Solution Path of a Nonconvexly Regularized Convex Sparse Model
  528. Concealing Medical Condition by Node Toggling in ASR for Dementia Patients
  529. Concentrated Reasoning and Unified Reconstruction for Multi-Modal Media Manipulation
  530. Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis
  531. Confidence-Aware Spatial-Temporal Attention Graph Convolutional Network for Skeleton-Based Expert-Novice Level Classification
  532. Conformalized Multimodal Uncertainty Regression and Reasoning
  533. Conformer is All You Need for Visual Speech Recognition
  534. Congestion-Aware Distributed Task Offloading in Wireless Multi-Hop Networks Using Graph Neural Networks
  535. Conjugate Gradient Based Adaptive Algorithm for Nonlinear AEC
  536. Connecting Speech Encoder and Large Language Model for ASR
  537. Considering Temporal Connection between Turns for Conversational Speech Synthesis
  538. Consistent and Relevant: Rethink the Query Embedding in General Sound Separation
  539. Contactless Radar Heart Rate Variability Monitoring Via Deep Spatio-Temporal Modeling
  540. Content-Based Objective Evaluation of Artificially Generated Sign Language Videos
  541. Context-Aware Dual Attention Network for Multimodal Sarcasm Detection
  542. Context-Aware Preference Learning System Based on Dirichlet Process Gaussian Mixture Model
  543. Context-Aware Transformer for Single Image Rain Streaks Removal
  544. Context-Aware and Contrastiveness-Driven Feature Learning for Cross-Domain Few-Shot Hyperspectral Image Classification
  545. Context-Guided and Syntactic Augmented Dual Graph Convolutional Network for Aspect-Based Sentiment Analysis
  546. Contextual Biasing Methods for Improving Rare Word Detection in Automatic Speech Recognition
  547. Contextual Biasing of Named-Entities with Large Language Models
  548. Contextual Human Object Interaction Understanding from Pre-Trained Large Language Model
  549. Contextualized Automatic Speech Recognition With Attention-Based Bias Phrase Boosted Beam Search
  550. Continual Learning with Class-Level Minimally Interfered Update
  551. Continuous Review and Timely Correction: Enhancing the Resistance to Noisy Labels via Self-Not-True Distillation
  552. Contrastive Deep Nonnegative Matrix Factorization For Community Detection
  553. Contrastive Learning for Regression on Hyperspectral Data
  554. Contrastive Learning with Audio Discrimination for Customizable Keyword Spotting in Continuous Speech
  555. Contrastive Learning with Bidirectional Transformers for Knowledge Tracing
  556. Contrastive Learning with High-Quality and Low-Quality Augmented Data for Query-Focused Summarization
  557. Contrastive Loss Based Frame-Wise Feature Disentanglement for Polyphonic Sound Event Detection
  558. Contrastive Speaker Embedding With Sequential Disentanglement
  559. Contrmix: Progressive Mixed Contrastive Learning for Semi-Supervised Medical Image Segmentation
  560. ControlCap: Controllable Captioning via No-Fuss Lexicon
  561. Controllable Prosody Generation with Partial Inputs
  562. Controllable Semantic Linguistic Steganography via Summarization Generation
  563. Controllable Speaking Styles Using A Large Language Model
  564. Convergent Plug-And-Play Using Contractive Denoisers
  565. Conversation Clique-Based Model for Emotion Recognition In Conversation
  566. Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models
  567. Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion
  568. Cooking-Clip: Context-Aware Language-Image Pretraining for Zero-Shot Recipe Generation
  569. Cooperative Sensing Via Matrix Factorization of the Partially Received Sample Covariance Matrix
  570. Coordinate-Based Neural Network for Fourier Phase Retrieval
  571. Core Body Temperature and its Role in Detecting Acute Stress: A Feasibility Study
  572. Corn: Co-Trained Full- and No-Reference Speech Quality Assessment
  573. Corner Detection Based on a Rotation-Invariant and Noise-Insensitive Curvature Measurement
  574. Corpus Synthesis for Zero-Shot ASR Domain Adaptation Using Large Language Models
  575. Correcting Faulty Road Maps by Image Inpainting
  576. Correction Focused Language Model Training For Speech Recognition
  577. Correlation-Based Machine Learning Techniques for Channel Estimation with Fluid Antennas
  578. Cost Aware Untargeted Poisoning Attack Against Graph Neural Networks
  579. Counting Network for Learning from Majority Label
  580. Coupled Block-Term Tensor Decomposition for Near-Field Localization in multi-static MIMO Radar Systems
  581. Coupling Self-Supervised and Supervised Contrastive Learning for Multiple Classification of Cervical Cytological Whole Slide Images
  582. Coverage Analysis For mmWAVE UAV Networks with Static and Dynamic Blockages
  583. Cramer-Rao Bound for Admittance Matrix Estimation under Laplacian Constraints
  584. Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
  585. Credible Teacher for Semi-Supervised Object Detection in Open Scene
  586. Cross Branch Feature Fusion Decoder for Consistency Regularization-Based Semi-Supervised Change Detection
  587. Cross Modal Training for ASR Error Correction with Contrastive Learning
  588. Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
  589. Cross-Age Contrastive Learning for Age-Invariant Face Recognition
  590. Cross-Attention watermarking of Large Language Models
  591. Cross-Camera Human Motion Transfer by Time Series Analysis
  592. Cross-Domain Cross-Task Transfer Mobile Touch-Stroke Authentication
  593. Cross-Image Distillation for Semi-Supervised Semantic Segmentation
  594. Cross-Lingual Learning in Multilingual Scene Text Recognition
  595. Cross-Modal Alignment for End-to-End Spoken Language Understanding Based on Momentum Contrastive Learning
  596. Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing
  597. Cross-Modal Multiscale Difference-Aware Network for Joint Moment Retrieval and Highlight Detection
  598. Cross-Modal Parallel Training for Improving end-to-end Accented Speech Recognition
  599. Cross-Modal Synthesis of Structural MRI and Functional Connectivity Networks via Conditional ViT-GANs
  600. Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake Detection
  601. Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
  602. Cross-Subject EEG Emotion Recognition Based on Interconnected Dynamic Domain Adaptation
  603. Cross-Target Stance Detection by Exploiting Target Analytical Perspectives
  604. Cross-Triggering Issue in Audio Event Detection and Mitigation
  605. Crowd Modeling and Control Via Cooperative Adaptive Filtering
  606. Crowdsourced Multilingual Speech Intelligibility Testing
  607. Crowdsourced and Automatic Speech Prominence Estimation
  608. CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds
  609. Crypto-Mine: Cryptanalysis Via Mutual Information Neural Estimation
  610. Cubic Knowledge Distillation for Speech Emotion Recognition
  611. Cuffless Blood Pressure Estimation Using Magnetic Flux In A Ring Form Factor
  612. Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised Representations
  613. Customising General Large Language Models for Specialised Emotion Recognition Tasks
  614. Customized Treatment Per Pixel for Blind Image Super-Resolution
  615. Cutransnet: Transformers to Make Strong Encoders for Multi-Task Vision Perception of Autonomous Driving
  616. Cyclic Misspecified Cramer-Rao Bound for Periodic Parameter Estimation
  617. D3: Dual-Domain Defenses for Byzantine-Resilient Decentralized Resource Allocation
  618. DACR: Distribution-Augmented Contrastive Reconstruction for Time-Series Anomaly Detection
  619. DAMP: Distribution-Aware Magnitude Pruning for Budget-Sensitive Graph Convolutional Networks
  620. DAP: Domain-Aware Prompt Learning for Vision-and-Language Navigation
  621. DBS: Differentiable Budget-Aware Searching For Channel Pruning
  622. DCL-Net: Dual Contrastive Learning Network for Semi-Supervised Multi-Organ Segmentation
  623. DCS: Debiased Contrastive Learning with Weak Supervision for Time Series Classification
  624. DCTTS: Discrete Diffusion Model with Contrastive Learning for Text-to-Speech Generation
  625. DDD: A Perceptually Superior Low-Response-Time DNN-Based Declipper
  626. DDI-CoCo: A Dataset for Understanding the Effect of Color Contrast in Machine-Assisted Skin Disease Detection
  627. DDN-Net: Deep Residual Shrinkage Denoising Networks with Channel-Wise Adaptively Soft Thresholds for Automated Major Depressive Disorder Identification
  628. DEEPOREDNET: Contrastive Learning-Based Attention-Weighted Dual Channel Residual Network for Ocular Redness Assessment
  629. DEGAN: Discrimination Enhanced GAN for Perceptual-Oriented Super-Resolution
  630. DETS: End-to-End Single-Stage Text-to-Speech Via Hierarchical Diffusion Gan Models
  631. DF-VTON: Dense Flow Guided Virtual Try-On Network
  632. DG-RainDiff: Depth-Guided Dynamic Message Passing Diffusion Model for Mixture of Rain Removal
  633. DGLP: Incorporating Orientation Information for Enhanced Link Prediction in Directed Graphs
  634. DI-MVS: Learning Efficient Multi-View Stereo With Depth-Aware Iterations
  635. DIB-X: Formulating Explainability Principles for a Self-Explainable Model Through Information Theoretic Learning
  636. DIFFSC: Semantic Communication Framework With Enhanced Denoising Through Diffusion Probabilistic Models
  637. DITW: A High-Performance Deep-Independent Template-Based Watermarking
  638. DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
  639. DMEL: The Differentiable Log-Mel Spectrogram as a Trainable Layer in Neural Networks
  640. DMKD: Improving Feature-Based Knowledge Distillation for Object Detection Via Dual Masking Augmentation
  641. DMT: Comprehensive Distillation with Multiple Self-Supervised Teachers
  642. DOA Estimation for Switch-Element Arrays Based on Sparse Representation
  643. DONE: Dynamic Neural Representation Via Hyperplane Neural ODE
  644. DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction
  645. DROPFL: Client Dropout Attacks Against Federated Learning Under Communication Constraints
  646. DRSM: Efficient Neural 4D Decomposition for Dynamic Reconstruction in Stationary Monocular Cameras
  647. DSIS: A Novel (K, N) Threshold Deniable Secret Image Sharing Scheme with Lossless Recovery
  648. DT-NeRF: Decomposed Triplane-Hash Neural Radiance Fields For High-Fidelity Talking Portrait Synthesis
  649. DURRNET: Deep Unfolded Single Image Reflection Removal Network with Joint Prior
  650. Darkshot: Lighting Dark Images with Low-Compute and High-Quality
  651. Data Augmentation via Subgroup Mixup for Improving Fairness
  652. Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-Speech
  653. Data-Aided Channel Estimation Utilizing Gaussian Mixture Models
  654. Data-Driven Convex Regularizers for Inverse Problems
  655. Data-Driven Lattices for Vector Quantization
  656. Data-Free Watermark for Deep Neural Networks by Truncated Adversarial Distillation
  657. Data-Scarce Condition Modeling Requires Model-Based Prior Regularization
  658. Dataset Distillation with Channel Efficient Process
  659. De Novo Molecule Generation with Graph Latent Diffusion Model
  660. Debiasing Recommenders Through Personalized Popularity-Aware Margins
  661. Debris Sensing Based on Leo Constellation: An Intersatellite Channel Parameter Estimation Approach
  662. Decentralized Generalized Approximate Message-Passing for Tree-Structured Networks
  663. Decentralized Low Rank Matrix Recovery from Column-Wise Projections by Alternating GD and Minimization
  664. Decentralizing Coherent Joint Transmission Precoding Via Deterministic Equivalents
  665. Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification
  666. Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
  667. Decoupling and Refilling: A Simple Data Augmentation Method for Aspect Term Extraction
  668. Deep Convolution Network Based Super Resolution DOA Estimation with Toeplitz and Sparse Prior
  669. Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation
  670. Deep INCM Reconstruction for Adaptive Beamforming
  671. Deep Learning AMR Model Inference Acceleration with CFU for Edge Systems
  672. Deep Learning Based Single-Shot Profilometry by Three-Channel Binary-Defocused Projection
  673. Deep Learning Inversion of Ocean Wave Spectrum from SAR Satellite Observations
  674. Deep Manifold Transformation for Protein Representation Learning
  675. Deep Neighbor Layer Aggregation for Lightweight Self-Supervised Monocular Depth Estimation
  676. Deep Neural Network Models Trained with a Fixed Random Classifier Transfer Better Across Domains
  677. Deep Optimization of Relay Networks-Using Relays as Neurons
  678. Deep Plug-and-Play Algorithm for Unsaturated Imaging
  679. Deep Regression for Biological Age Estimation in Multiple Organs: Investigations on 40, 000 Subjects of the UK Biobank
  680. Deep Reinforcement Learning for Energy Minimization in Multi-RIS-Aided Cell-Free MEC Networks
  681. Deep Residual W-Unit Learning with Semantic Embedding for Automatic Pulmonary CT Artery-Vein Separation
  682. Deep Unfolded Annealed Stein Particle Filter for Vehicle Tracking
  683. Deep Unrolling Network for SAR Image Despeckling
  684. Deep Variational Privacy Funnel: General Modeling with Applications in Face Recognition
  685. Deep Versatile Hyperspectral Reconstruction Model from A Snapshot Measurement with Arbitrary Masks
  686. DeepGRE: Global Robustness Evaluation of Deep Neural Networks
  687. Defending against Clean-Image Backdoor Attack in Multi-Label Classification
  688. DefocusSR: An Efficient Framework for Defocus Image Super-Resolution Guided by Depth Information
  689. DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion Prediction
  690. Deformation And Penetration Hybrid Detection-Net For Parcels Inspection In Industrial Supply Chain
  691. Delay Embedding for Matrix Graphical Model Learning from Dependent Data
  692. Delineation of Prostate Cancer Via Enhanced AI-Based Algorithm In Ultrasound Images
  693. Delving Deeper Into Vulnerable Samples in Adversarial Training
  694. Dementia Assessment Using Mandarin Speech with an Attention-Based Speech Recognition Encoder
  695. Denoising Diffusion Probabilistic Models for Action-Conditioned 3D Motion Generation
  696. Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation
  697. Design of Spatial-Slow-Time Constant-Modulus Waveform Transmission and Receive Adaptive Filter for Dual-Function Radar Communications with Reconfigurable Intelligent Surface
  698. Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
  699. Detecting Continuous Gravitational Waves Using Generated Training Data
  700. Detection and Attribution of Models Trained on Generated Data
  701. Detection in Complex Scenes Using Rgb and Depth Multimodal Feature Fusion
  702. Detection of Epileptic Seizures in Long Eeg Recordings Using an Anomaly Detector with Artifact Rejection
  703. Detector Design for Distributed Multichannel Radar Sensors in Colored Interference Environments
  704. Determined BSS by Combination of IVA and DNN via Proximal Average
  705. Diacorrect: Error Correction Back-End for Speaker Diarization
  706. Diagnosis of Autism Spectrum Disorder Based on Contrastive Functional Connectivity Graph Learning Network
  707. Diagonalize Integral Graph by DCT
  708. DialCLIP: Empowering Clip As Multi-Modal Dialog Retriever
  709. Dialog Modeling in Audiobook Synthesis
  710. Diarist: Streaming Speech Translation with Speaker Diarization
  711. Dicetrack: Lightweight Dice Classification on Resource-Constrained Platforms with Optimized Deep Learning Models
  712. Diff-HOD: Diffusion Model for Object Detection in Hazy Weather Conditions
  713. Diff-SV: A Unified Hierarchical Framework for Noise-Robust Speaker Verification Using Score-Based Diffusion Probabilistic Models
  714. DiffDub: Person-Generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-Encoder
  715. DiffRENT: A Diffusion Model for Recording Environment Transfer of Speech
  716. Differentiable Quantum Architecture Search For Job Shop Scheduling Problem
  717. Differentiable Resolution Compression and Alignment for Efficient Video Classification and Retrieval
  718. Differential Beamforming with Null Constraints for Spherical Microphone Arrays
  719. Differentially Private Federated Frank-Wolfe
  720. Diffevent: Event Residual Diffusion for Image Deblurring
  721. Diffradar: High-Quality Mmwave Radar Perception With Diffusion Probabilistic Model
  722. Diffstock: Probabilistic Relational Stock Market Predictions Using Diffusion Models
  723. Diffusion Models for Audio Semantic Communication
  724. Diffusion Optimistic Learning for Min-Max Optimization
  725. Diffusion-Based Adversarial Purification for Robust Deep Mri Reconstruction
  726. Diffusion-Based Pose Refinement and Multi-Hypothesis Generation for 3D Human Pose Estimation
  727. Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
  728. Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
  729. Diffusion-Based Speech Enhancement with a Weighted Generative-Supervised Learning Loss
  730. Diffusioninst: Diffusion Model for Instance Segmentation
  731. Digital Pathology Image Deblurring Via Local Focus Quality Assessment
  732. Digital Task-Oriented Communication with Hardware-Limited Task-Based Quantization
  733. Direct Position Determination by Covariance-Fitting on the Riemannian Manifold of Hermitian Positive Definite Matrices
  734. Directed Scattering for Knowledge Graph-Based Cellular Signaling Analysis
  735. Directional Gain Based Noise Covariance Matrix Estimation for MVDR Beamforming
  736. Discovering Malicious Signatures in Software from Structural Interactions
  737. Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition
  738. Discriminative Semi-Supervised Feature Selection Via a Class-Credible Pseudo-Label Learning Framework
  739. Discriminative Training of VBx Diarization
  740. Disentangle Estimation of Causal Effects from Cross-Silo Data
  741. Disentangled Graph Representation with Contrastive Learning for Rumor Detection
  742. Disentanglement Network: Disentangle the Emotional Features from Acoustic Features for Speech Emotion Recognition
  743. Disentangling the Spectral Properties of the Hodge Laplacian: not all small Eigenvalues are Equal
  744. Distill Vision Transformers to CNNs via Teacher Collaboration
  745. Distilling Distributional Uncertainty from a Gaussian Process
  746. Distilling Hubert with LSTMs via Decoupled Knowledge Distillation
  747. Distributed Decision-Making for Community Structured Networks
  748. Distributed Stochastic Contextual Bandits for Protein Drug Interaction
  749. Distributed Vector Approximate Message Passing
  750. Distribution-Aware Contrastive Learning for Robust Medical Image Segmentation
  751. Diversifying Cross-Domain Few-Shot Learning via Multimodal Image Editing
  752. Diversity-Aware Buffer for Coping with Temporally Correlated Data Streams in Online Test-Time Adaptation
  753. Diversity-Based Core-Set Selection for Text-to-Speech with Linguistic and Acoustic Features
  754. Do Learned Speech Symbols Follow Zipf's Law?
  755. Do Self-Supervised Speech and Language Models Extract Similar Representations as Human Brain?
  756. Does Audio Deepfake Detection Rely on Artifacts?
  757. Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization
  758. Domain Adaptive Graph Classification
  759. Domain Generalization with fourier Transform and soft thresholding
  760. Domain-Adaptive Semantic Segmentation Emerges From Vision-Language Supervised Domain-Debiased Self-Training
  761. Domain-Adaptive and Subgroup-Specific Cascaded Temperature Regression for Out-of-Distribution Calibration
  762. Domain-Slot Aware Contrastive Learning for Improved Dialogue State Tracking
  763. Domain-Wise Invariant Learning for Panoptic Scene Graph Generation
  764. Domaindiff: Boost out-of-Distribution Generalization with Synthetic Data
  765. Double Reverse Regularization Network Based on Self-Knowledge Distillation for SAR Object Classification
  766. Driver Scanpath Prediction Based On Inverse Reinforcement Learning
  767. Drop Sparse Convolution for 3D Object Detection
  768. Dropout Multi-Head Attention for Single Image Super-Resolution
  769. DuNet: A Robust End-to-End Deep Neural Network Framework for Imbalanced Classification
  770. Dual Contrastive Learning Guided Pathological Image Re-Staining
  771. Dual Directional Complementary Gradient Fusion and Deep Refinement for Hyperspectral Image Super Resolution
  772. Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding
  773. Dual Parameter-Efficient Fine-Tuning for Speaker Representation Via Speaker Prompt Tuning and Adapters
  774. Dual Rank-1 Tensor Attention Module for Convolutional Neural Networks
  775. Dual-Channel Unlimited Sampling for Bandpass Signals
  776. Dual-Color Granularity Alignment for Text-Based Person Search
  777. Dual-Mix for Cross-Modal Retrieval with Noisy Labels
  778. Dual-Path Minimum-Phase and All-Pass Decomposition Network for Single Channel Speech Dereverberation
  779. Dual-Stream Contrastive Predictive Network with Joint Handcrafted Feature View for SAR Ship Classification
  780. DualGCN-MIL: Whole Slide Image Classification Based on Double Relationship Graph Learning
  781. Dualvc 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
  782. DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
  783. Dust: Dual-Grained Syntax-Aware Transformer Network for Chinese Named Entity Recognition
  784. Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of a Multilingual ASR Model
  785. Dynamic Bandwidth Variational Mode Decomposition
  786. Dynamic Clustering and Cluster Contrastive Learning for Unsupervised Person Re-Id With Feature Distribution Alignment
  787. Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language Models
  788. Dynamic Frequency Domain Graph Convolutional Network for Traffic Forecasting
  789. Dynamic Label Smoothing Strategy for Biosignal Classification
  790. Dynamic Model Structure Adjustment to Realize Quantum Continual Learning Based on Quantum Data
  791. Dynamic Multi-Scale Context Aggregation for Conversational Aspect-Based Sentiment Quadruple Analysis
  792. Dynamic Mutual-Activated Transformer for Human Motion Prediction
  793. Dynamic Privacy Allocation for Locally Differentially Private Federated Learning with Composite Objectives
  794. Dynamic Random Feature Gaussian Processes for Bayesian Optimization of Time-Varying Functions
  795. Dynamic Replay Training for Class-Incremental Learning
  796. Dynamic Speech Emotion Recognition Using A Conditional Neural Process
  797. Dynamic Video Frame Interpolation with Integrated Difficulty Pre-Assessment
  798. Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark For Speech
  799. EC-NAS: Energy Consumption Aware Tabular Benchmarks for Neural Architecture Search
  800. ECIL-MU: Embedding Based Class Incremental Learning and Machine Unlearning
  801. ECM-OPCC: Efficient Context Model for Octree-Based Point Cloud Compression
  802. ECPNet: An Enhanced Curve Perception Network for Lane Detection
  803. ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
  804. EDM: Synthetic Data from Exemplar Diffusion Model Improves Non-Communicable Diseases Detection
  805. EEG Emotion Recognition Based on Dynamical Graph Attention Network
  806. EEG-Based Fast Auditory Attention Detection in Real-Life Scenarios Using Time-Frequency Attention Mechanism
  807. EK-Net: Real-Time Scene Text Detection with Expand Kernel Distance
  808. EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences
  809. EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild Data
  810. EOFD-Net: Edge Optimization and Feature Denoising for Weakly Supervised Deep Nuclei Segmentation with Point Annotations
  811. EPA: Neural Collapse Inspired Robust Out-of-distribution Detector
  812. ESA: Expert-and-Samples-Aware Incremental Learning Under Longtail Distribution
  813. ESTGN: Enhanced Self-Mined Text Guided Super-Resolution Network for Superior Image Super Resolution
  814. ESVC: Combining Adaptive Style Fusion and Multi-Level Feature Disentanglement for Expressive Singing Voice Conversion
  815. ETP: Learning Transferable ECG Representations via ECG-Text Pre-Training
  816. Early Diagnosing Parkinson's Disease Via a Deep Learning Model Based on Augmented Facial Expression Data
  817. Echocardiography Video Synthesis from End Diastolic Semantic Map Via Diffusion Model
  818. Edge Attention Learning for Efficient Camouflaged Object Detection
  819. Edge Deployable Distributed Evolutionary Optimization based Calibration method for Neural Quantization
  820. Effect of Beampattern on Matrix Completion with Sparse Arrays
  821. Effect of Target Signals and Delays on Spatially Selective Active Noise Control for Open-Fitting Hearables
  822. Effective Connectivity-Based Multi-View Feature Learning Method for Dementia Diagnosis with FNIRS Signal
  823. Effective Image Tampering Localization Via Enhanced Transformer and Co-Attention Fusion
  824. Effective Internal Language Model Training and Fusion for Factorized Transducer Model
  825. Efficient 3D Position Estimation in Badminton Scene
  826. Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
  827. Efficient Adapter Tuning of Pre-Trained Speech Models for Automatic Speaker Verification
  828. Efficient Architecture Search for Real-Time Instance Segmentation
  829. Efficient Black-Box Speaker Verification Model Adaptation With Reprogramming And Backend Learning
  830. Efficient Content Reconstruction for High Dynamic Range Imaging
  831. Efficient Federated Learning with Smooth Aggregation for Non-IID Data from Multiple Edges
  832. Efficient Functional Link Adaptive Filters Based On Nearest Kronecker Product Decomposition
  833. Efficient Fusion of Depth Information for Defocus Deblurring
  834. Efficient Hierarchical Stripe Attention for Lightweight Image Super-Resolution
  835. Efficient High-Performance Bark-Scale Neural Network for Residual Echo and Noise Suppression
  836. Efficient Joint Rectification of Photometric and Geometric Distortions in Document Images
  837. Efficient Learned Image Compression with Selective Kernel Residual Module and Channel-Wise Causal Context Model
  838. Efficient Learning on Successive Test Time Augmentation
  839. Efficient Multi-Channel Speech Enhancement with Spherical Harmonics Injection for Directional Encoding
  840. Efficient Personal Voice Activity Detection with Wake Word Reference Speech
  841. Efficient Point Cloud Attribute Compression Framework using Attribute-Guided Graph Fourier Transform
  842. Efficient Point Cloud Attribute Compression Using Rich Parallelizable Context Model
  843. Efficient Polyp Segmentation via Integrity Learning
  844. Efficient Posenet with Coarse to Fine Transformer
  845. Efficient Quantum Recurrent Reinforcement Learning Via Quantum Reservoir Computing
  846. Efficient Scene Text Image Super-Resolution with Semantic Guidance
  847. Efficient Video and Audio Processing with Loihi 2
  848. EiffHDR: An Efficient Network for Multi-Exposure High Dynamic Range Imaging
  849. Eigendecomposition-Based Spatial-Temporal Attention for Brain Cognitive States Identification
  850. Electroencephalogram Helps Few-Shot Learning
  851. Electroencephalogram Sensor Data Compression Using an Asymmetrical Sparse Autoencoder with a Discrete Cosine Transform Layer
  852. Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders
  853. Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
  854. Elevating Visual Prompting in Transfer Learning Via Pruned Model Ensembles: No Retrain, No Pain
  855. Ellipse Detection Based On Structure-Preserving Anisotropic Edge Extraction
  856. Ellipse Detection Based on Contrast-Guided Arc Enhancement
  857. Embedded Feature Similarity Optimization with Specific Parameter Initialization for 2D/3D Medical Image Registration
  858. Embedded Graph Representation for Inter-Frame Coding of Dynamic Meshes
  859. EmoRED: A Dataset for Relation Extraction in Texts with Emoticons
  860. EmoTVR: A Hybrid Model to Estimate Continuous-Time and Continuous-Level Emotion from Electroencephalography
  861. EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
  862. Emohrnet: High-Resolution Neural Network Based Speech Emotion Recognition
  863. Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
  864. Emotion-Aligned Contrastive Learning Between Images and Music
  865. Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
  866. Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
  867. Employing Real Training Data for Deep Noise Suppression
  868. Empowering Vision-Language Models for Reasoning Ability through Large Language Models
  869. EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
  870. Enabling Device Control Planning Capabilities of Small Language Model
  871. Enabling Orientation-Free Mmwave-Based Vital Sign Sensing with Multi-Domain Signal Analysis
  872. Enabling Secure Wireless Communications via Movable Antennas
  873. Encoder-Minimal and Decoder-Minimal Framework for Remote Sensing Image Dehazing
  874. Encoding Seasonal Climate Predictions with Modular Neural Network
  875. Encoding Time and Energy Model for SVT-AV1 Based on Video Complexity
  876. End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG Signals
  877. End-To-End Real Time Tracking of Children's Reading with Pointer Network
  878. End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image Captioning
  879. End-to-End Learning of Gaussian Mixture Proposals Using Differentiable Particle Filters and Neural Networks
  880. End-to-End Speech Recognition Contextualization with Large Language Models
  881. End-to-End Speech Translation with Mutual Knowledge Distillation
  882. Energy Efficient Wake-Up Solution for Large-Scale Internet of Underwater Things Networks
  883. Energy-Aware Resolution Selection for Per-Title Encoding
  884. Energy-Based Models for Speech Synthesis
  885. Energy-Efficient Decentralized Learning Via Graph Sparsification
  886. Energy-Saving Cell-Free Massive MIMO Precoders with a per-AP Wideband Kronecker Channel Model
  887. Engineering the Neural Collapse Geometry of Supervised-Contrastive Loss
  888. Enhanced Axle-Based Vehicle Classification Using Angle-Based Micro-Doppler Signature
  889. Enhanced Channel Estimation in mm-Wave Mimo Systems Leveraging Integrated Communication and Sensing
  890. Enhanced Color Palette Modeling For Lossless Screen Content Compression
  891. Enhanced Deep Reinforcement Learning for Parcel Singulation in Non-Stationary Environments
  892. Enhanced KPI Anomaly Detection: An Unsupervised Hybrid Model with Dynamic Threshold
  893. Enhanced Low-Rank and Sparse Tucker Decomposition For Image Completion
  894. Enhanced Screen Shooting Resilient Document Watermarking
  895. Enhanced Transfer Learning with Efficient Modeling and Adaptive Fusion of Knowledge Via Prompt Tuning
  896. Enhanced Unsupervised Domain Adaptation with Dual-Attention Between Classification and Domain Alignment
  897. Enhancing Adversarial Robustness of DNNS Via Weight Decorrelation in Training
  898. Enhancing Adversarial Training with Prior Knowledge Distillation for Robust Image Compression
  899. Enhancing Adversarial Transferability in Object Detection with Bidirectional Feature Distortion
  900. Enhancing AoA Estimation Via Phase Modeling of Bluetooth 5 CTE Signals
  901. Enhancing Argumentative Relation Classification by Multi-Granularity Retrieval and Heterogeneous Graph Reasoning
  902. Enhancing Audio Generation Diversity with Visual Information
  903. Enhancing Audio-Visual Question Answering with Missing Modality via Trans-Modal Associative Learning
  904. Enhancing Code-Switching Speech Recognition With Interactive Language Biases
  905. Enhancing Conversation Smoothness in Language Learning Chatbots: An Evaluation of GPT4 for ASR Error Correction
  906. Enhancing Cross-Domain Detection: Adaptive Class-Aware Contrastive Transformer
  907. Enhancing Document-Level Event Extraction via Structure-Aware Heterogeneous Graph with Multi-Granularity Subsentences
  908. Enhancing End-to-End Conversational Speech Translation Through Target Language Context Utilization
  909. Enhancing Event Sequence Modeling with Contrastive Relational Inference
  910. Enhancing Expressiveness in Dance Generation Via Integrating Frequency and Music Style Information
  911. Enhancing GAN Performance Through Neural Architecture Search and Tensor Decomposition
  912. Enhancing Gender Privacy with Photo-Realistic Fusion of Disentangled Spatial Segments
  913. Enhancing Generalization Of Invisible Facial Privacy Cloak Via Gradient Accumulation
  914. Enhancing Generalization in Medical Visual Question Answering Tasks Via Gradient-Guided Model Perturbation
  915. Enhancing Generative Aspect-Based Sentiment Analysis with Relation-Level Supervision and Prompt
  916. Enhancing Healthcare with EOG: A Novel Approach to Sleep Stage Classification
  917. Enhancing Hyperspectral Anomaly Detection by Difference-of-Convex Sparse Anomaly Modeling
  918. Enhancing Image-Text Matching with Adaptive Feature Aggregation
  919. Enhancing Low-Latency Speaker Diarization with Spatial Dictionary Learning
  920. Enhancing Multi-Task Models For Recommendation with Tensor Trace Norm
  921. Enhancing Multilingual Speech Recognition through Language Prompt Tuning and Frame-Level Language Adapter
  922. Enhancing Multilingual TTS with Voice Conversion Based Data Augmentation and Posterior Embedding
  923. Enhancing Noisy Label Learning Via Unsupervised Contrastive Loss with Label Correction Based on Prior Knowledge
  924. Enhancing Note-Level Singing Transcription Model with Unlabeled and Weakly Labeled Data
  925. Enhancing Performance of Coarsened Graphs with Gradient-Matching
  926. Enhancing Pre-Trained ASR System Fine-Tuning for Dysarthric Speech Recognition Using Adversarial Data Augmentation
  927. Enhancing Quantised End-to-End ASR Models Via Personalisation
  928. Enhancing Realism in 3D Facial Animation Using Conformer-Based Generation and Automated Post-Processing
  929. Enhancing Reinforcement Learning via Causally Correct Input Identification and Targeted Intervention
  930. Enhancing Semantic Communication with Deep Generative Models: An Overview
  931. Enhancing Short-and Long-Term Sea Surface Temperature Forecasting with a Static and Dynamic Learnable Personalized Graph Convolution Network
  932. Enhancing Spatial Audio Generation with Source Separation and Channel Panning Loss
  933. Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach
  934. Enhancing Steganography of Generative Image Based on Image Retouching
  935. Enhancing Targeted Transferability VIA Feature Space Fine-Tuning
  936. Enhancing Two-Stage Finetuning for Speech Emotion Recognition Using Adapters
  937. Enhancing Violin Fingering Generation through Audio-Symbolic Fusion
  938. Enhancing the Domain Robustness of Self-Supervised pre-Training with Synthetic Images
  939. Enriching Music Descriptions with A Finetuned-LLM and Metadata for Text-to-Music Retrieval
  940. Entwined Inversion: Tune-Free Inversion For Real Image Faithful Reconstruction and Editing
  941. Environmental Sound Synthesis from Vocal Imitations and Sound Event Labels
  942. Esihgnn: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition
  943. Estimating Directed Spectral Information Flow between Multi-Resolution Time Series
  944. Estimating Exercise-Induced Fatigue from Thermal Facial Images
  945. Estimating Symptoms and Clinical Signs Instead of Disorders: The Path Toward The Clinical Use of Voice and Speech Biomarkers In Psychiatry
  946. Estimation of Impulse Responses for a Moving Source Using Optimal Transport Regularization
  947. Estimation of Spectral Lines Using Expectation Propagation
  948. Evaluation of an Improved Ultrasonic Imaging Helmet for Observing Articulatory Data
  949. Evidence-Aware Multimodal Chinese Social Media Rumor Detection
  950. Evolution Backcasting of Edge Flows From Partial Observations Using Simplicial Vector Autoregressive Models
  951. Exact Classification of NMR Spectra from NMR Signals
  952. Exploiting A Quantum Multiple Kernel Learning Approach For Low-Resource Spoken Command Recognition
  953. Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
  954. Exploiting Modality-Specific Features for Multi-Modal Manipulation Detection and Grounding
  955. Exploiting Spatial-Temporal Data for Sleep Stage Classification via Hypergraph Learning
  956. Exploration of Visual Prompt in Grounded Pre-Trained Open-Set Detection
  957. Exploring Adapters with Conformers for Children's Automatic Speech Recognition
  958. Exploring Consistent Spatio-Temporal Distortion and Stable 3-D DCT Coefficients for Robust Blind Video Watermarking
  959. Exploring Label Hierarchy in Dialogue Intent Classification
  960. Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection
  961. Exploring Latent Cross-Channel Embedding for Accurate 3d Human Pose Reconstruction in a Diffusion Framework
  962. Exploring Meta Information for Audio-Based Zero-Shot Bird Classification
  963. Exploring Multi-Modal Control in Music-Driven Dance Generation
  964. Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning
  965. Exploring Phonetic Context-Aware Lip-Sync for Talking Face Generation
  966. Exploring Self-Explainable Street-Level IP Geolocation with Graph Information Bottleneck
  967. Exploring Self-supervised Contrastive Learning of Spatial Sound Event Representation
  968. Exploring Soft Prompt Initialization Strategy for Few-Shot Continual Text Classification
  969. Exploring Spatio-Temporal Discriminative Cues for Group Activity Recognition Via Contrastive Learning
  970. Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
  971. Exploring Targeted Universal Adversarial Attack for Deep Hashing
  972. Exploring the Utility of Clip Priors for Visual Relationship Prediction
  973. Expression Domain Translation Network for Cross-Domain Head Reenactment
  974. Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
  975. Extending Implicit Neural Representations for Text-to-Image Generation
  976. Extending Large Language Models for Speech and Audio Captioning
  977. Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder Components
  978. Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
  979. Extending Whisper with Prompt Tuning to Target-Speaker ASR
  980. Extension of Clifford Data Regression Methods for Quantum Error Mitigation
  981. External Division of Two Proximity Operators: An Application to Signal Recovery with Structured Sparsity
  982. Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
  983. Extremely Light-Weight Learning Based LDR to PQ HDR Conversion Using Bernstein Curves
  984. Extrinsic Versus App Information Feedback in Turbo Vep Mu-Mimo Receivers: Optimization Via Deep Unfolding
  985. Eye Motion Matters for 3D Face Reconstruction
  986. F1-EV score: Measuring The Likelihood of Estimating a Good Decision Threshold for Semi-Supervised Anomaly Detection
  987. F2GNN: An Adaptive Filter with Feature Segmentation for Graph-Based Fraud Detection
  988. FAMIM: A Novel Frequency-Domain Augmentation Masked Image Model Framework for Domain Generalizable Face Anti-Spoofing
  989. FAVANO: Federated Averaging with Asynchronous Nodes
  990. FCC-MF: Detecting Violence in Audio-Visual Context with Frame-Wise Cluster Contrast and Modality-Stage Flooding
  991. FDA-MIMO Radar Using Ambiguity Function for Target Two-Dimensional Localization
  992. FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth Consistency
  993. FDIG: A Fine-Grained Data Integration Approach for Group Recommendation
  994. FDNet: A Novel Multivariate Time Series Classification Model Through Fusing Feature and Difference
  995. FED-SDS: Adaptive Structured Dynamic Sparsity for Federated Learning Under Heterogeneous Clients
  996. FEDKA: Federated Knowledge Augmentation for Multi-Center Medical Image Segmentation on non-IID Data
  997. FFT-Based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
  998. FIBA: Federated Invisible Backdoor Attack
  999. FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter
  1000. FPGNet: Single Image Deraining with High-Frequency Channel and Frequency Domain Prior Guidance

Looking for submission deadlines instead? See the conference deadline calendar.