← All conferences

ICASSP 2024 Accepted Papers

The full list of 2,679 papers accepted at ICASSP 2024 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. RL-LOGO: Deep Reinforcement Learning Localization for Logo Recognition
  2. RSED: Zero-Shot Relation Triplet Extraction via Relation Selection and Entity Boundary Detection
  3. RTLBP-AN Efficient Local Pattern For Facial Images Retrieval
  4. RVAE-EM: Generative Speech Dereverberation Based On Recurrent Variational Auto-Encoder And Convolutive Transfer Function
  5. RVDNet: A Two-Stage Network for Real-World Video Desnowing with Domain Adaptation
  6. Radar Perception with Scalable Connective Temporal Relations for Autonomous Driving
  7. Radar Recognition in the Wild: Enhancing Radar Emitter Recognition through Auto-Correlation Model-Agnostic Meta Learning
  8. Radardiff: Improving Sea Clutter Suppression Using Diffusion Models for Radar Images
  9. Rademacher Complexity Regularization for Correlation-Based Multiview Representation Learning
  10. Radio Slam with Hybrid Sensing for Mixed Reflection Type Environments
  11. Randomized Maximum Likelihood Via High-Dimensional Bayesian Optimization
  12. Ranking Enhanced Fine-Grained Contrastive Learning for Recommendation
  13. Ranking of Visual Trackers Using Robust Error Norms
  14. Rapid Change Localization in Dynamic Graphical Models
  15. Rapid Hybrid Modular Receive Beamforming Via Learned Optimization
  16. Rate-Quality Based Rate Control Model for Neural Video Compression
  17. Rating-Augmented No-Reference Point Cloud Quality Assessment Using Multi-Task Learning
  18. Read, Spell and Repeat: Scene Text Recognition with Vision-Language Circular Refinement
  19. Real-Oriented Object Detection Driven by Intelligent Stockbreeding
  20. Real-Time Low-Latency Music Source Separation Using Hybrid Spectrogram-Tasnet
  21. Real-Time Multi-Human Parsing on Embedded Devices
  22. Real-Time Privacy-Preserving Fall Risk Assessment with a Single Body-Worn Tracking Camera
  23. Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure
  24. Recap: Retrieval-Augmented Audio Captioning
  25. Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural Networks: from Algorithms to Technology
  26. Recognition-Guided Diffusion Model for Scene Text Image Super-Resolution
  27. Reconstruction of Sound Field Through Diffusion Models
  28. Recovering Missing Node Features with Local Structure-Based Embeddings
  29. Recovering from Privacy-Preserving Masking with Large Language Models
  30. Recursive-Tail-Fista for Sparse Signal Recovery
  31. Redefining Night Vision: The Power of MSR-Driven Neural ISP
  32. Reduced-Dimensional Decomposition and Eigenspace Reconstruction of Coherent Sources with Arbitrary Rectangle Arrays
  33. Reducing the Complexity of Normalizing Flow Architectures for Point Cloud Attribute Compression
  34. Reference Line Network: On Simultaneous Gaussian Line Detection and Connection Graph Inference
  35. Refinement Bird's Eye View Feature for 3D Lane Detection with Dual-Branch View Transformation Module
  36. Refining 3D Human Mesh via Model-Free Offsets Estimation
  37. Refining Text Input For Augmentative and Alternative Communication (AAC) Devices: Analysing Language Model Layers For Optimisation
  38. Reflection Removal Using Recurrent Polarization-to-Polarization Network
  39. Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-Speech
  40. Region-Adaptive Video Sharpening Via Rate-Perception Optimization
  41. Regularized Conditional Alignment for Multi-Domain Text Classification
  42. Reinforcement Learning Compensated Filter for Multi-Agents Cooperative Localization
  43. Reinforcement Learning-Guided Optogenetic Stimulation Policies for Robust Functional Network Discovery
  44. Relational Graph-Bridged Image-Text Interaction: A Novel Method for Multi-Modal Relation Extraction
  45. Remixed2remixed: Domain Adaptation for Speech Enhancement by Noise2noise Learning with Remixing
  46. Renyi Divergences Learning for explainable classification of SAR Image Pairs
  47. Reparameterization Head for Efficient Multi-Input Networks
  48. Representation Learning across Feature and Topology Views with Output Correction for Graph Convolutional Networks
  49. Representation and Boundary Enhancement for Action Segmentation Using Transformer
  50. Repurposing Mu-Mimo Downlink For Joint Wireless Communications And Imaging Via Virtual Users
  51. Residual Dense Swin Transformer for Continuous Depth-Independent Ultrasound Imaging
  52. Residualtransformer: Residual Low-Rank Learning With Weight-Sharing For Transformer Layers
  53. Resource-Constrained Stereo Singing Voice Cancellation
  54. Resource-Efficient Separation Transformer
  55. Retaining Informative Latent Variables in Probabilistic Segmentation
  56. Rethinking Normals: Direction Guided Point Cloud Recognition
  57. Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification
  58. Rethinking Targeted Adversarial Attacks for Neural Machine Translation
  59. Retrieval Augmented End-to-End Spoken Dialog Models
  60. Retrieval-Augmented Text-to-Audio Generation
  61. Retrieval-Generation Synergy Augmented Large Language Models
  62. Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
  63. Reversible Jump Markov Chain Monte Carlo for Pulse Fitting
  64. Revise the NLU: A Prompting Strategy for Robust Dialogue System
  65. Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
  66. Revisiting the Equivalence of In-Context Learning and Gradient Descent: The Impact of Data Distribution
  67. Reweighted Atomic Norm Minimization for One-Bit Multichannel Spectral Compressed Sensing
  68. Riemannian Diffusion Adaptation over Graphs with Application to Online Distributed PCA
  69. Risk-Managed Sparse Index Tracking Via Market Graph Clustering
  70. RoFi: Robust WiFi Intrusion Detection via Distribution Matching
  71. Robust Beamforming for DFRC Systems in Complex Environments
  72. Robust Cross-Domain Speaker Verification with Multi-Level Domain Adapters
  73. Robust Decoding of the Auditory Attention from EEG Recordings Through Graph Convolutional Networks
  74. Robust DoA Estimation from Deep Acoustic Imaging
  75. Robust Face Recognition Based on an Angle-Aware Loss and Masked Autoencoder Pre-Training
  76. Robust Lightweight Depth Estimation Model via Data-Free Distillation
  77. Robust Localization of Key Fob Using Channel Impulse Response of Ultra Wide Band Sensors for Keyless Entry Systems
  78. Robust Low-Rank Correlation Fitting
  79. Robust Near-Field Beamforming for Millimeter Wave Communication System with Aperture Perturbations
  80. Robust Recovery of Joint Sparse Signals via Simultaneous Orthogonal Matching Pursuit
  81. Robust Regression Analysis Based on the K-Divergence
  82. Robust Self-Supervised Learning with Contrast Samples for Natural Language Understanding
  83. Robust Single-Particle Cryo-Em Image Denoising and Restoration
  84. Robust Speaker Personalisation Using Generalized Low-Rank Adaptation for Automatic Speech Recognition
  85. Robust Spoof Speech Detection Based on Multi-Scale Feature Aggregation and Dynamic Convolution
  86. Robust Symbol-Level Precoding via a Symbol-Perturbed Zero-Forcing Structure
  87. Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
  88. Robust and Imperceptible Commercial Camera-Screen Communication with 60Hz Refresh Rate
  89. RobustTSVar: A Robust Time Series Variance Estimation Algorithm
  90. Robustness Against Adversarial Attacks Via Learning Confined Adversarial Polytopes
  91. Robustness Evaluation of Machine Learning Models for Robot Arm Action Recognition in Noisy Environments
  92. Rényi Differential Privacy in the Shuffle Model: Enhanced Amplification Bounds
  93. S-Evaluator: Enhance Factual Consistency Evaluator with Adversarial Data Synthesized by Large Language Model
  94. S2E: Towards an End-to-End Entity Resolution Solution from Acoustic Signal
  95. SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
  96. SADA: Saudi Audio Dataset for Arabic
  97. SADE: A Speaker-Aware Dual Encoding Model Based on Diagbert for Medical Triage and Pre-Diagnosis
  98. SALM: Speech-Augmented Language Model with in-Context Learning for Speech Recognition and Translation
  99. SAM-DEBLUR: Let Segment Anything Boost Image Deblurring
  100. SAM-GEBD: Zero-Cost Approach for Generic Event Boundary Detection
  101. SAM-OCTA: A Fine-Tuning Strategy for Applying Foundation Model OCTA Image Segmentation Tasks
  102. SAM: A Self-Adaptive Attention Module for Context-Aware Recommendation System
  103. SAMF: Small-Area-Aware Multi-Focus Image Fusion for Object Detection
  104. SAMVG: A Multi-Stage Image Vectorization Model with the Segment-Anything Model
  105. SAR2NDVI: Pre-Training for SAR-to-NDVI Image Translation
  106. SASA: Saliency-Aware Self-Adaptive Snapshot Compressive Imaging
  107. SBM: Smoothness-Based Minimization for Domain Generalization
  108. SC-MAD: Mixtures of Higher-Order Networks for Data Augmentation
  109. SCNet: Sparse Compression Network for Music Source Separation
  110. SCORE: Self-Supervised Correspondence Fine-Tuning for Improved Content Representations
  111. SCRN: A Spectrogram Convolutional Recurrent Network for AoA Estimation Using Bluetooth 5
  112. SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in Hubert
  113. SDEMG: Score-Based Diffusion Model for Surface Electromyographic Signal Denoising
  114. SDIF-DA: A Shallow-to-Deep Interaction Framework with Data Augmentation for Multi-Modal Intent Detection
  115. SDRNet: Saliency-Guided Dynamic Restoration Network for Rain and Haze Removal in Nighttime Images
  116. SE-SIS: Shadow-Embeddable Lossless Secret Image Sharing for Greyscale Images
  117. SEA-GNN: Sequence Extension Augmented Graph Neural Network for Sequential Recommendation
  118. SECP: A Speech Enhancement-Based Curation Pipeline for Scalable Acquisition of Clean Speech
  119. SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
  120. SEGLLM: Topic-Oriented Call Segmentation Via LLM-Based Conversation Synthesis
  121. SELM: Speech Enhancement using Discrete Tokens and Language Models
  122. SERC-GCN: Speech Emotion Recognition In Conversation Using Graph Convolutional Networks
  123. SG2SC: A Generative Semantic Communication Framework for Scene Understanding-Oriented Image Transmission
  124. SGM: A Dataset for 3D Garment Reconstruction from Single Hand-Drawn Sketch
  125. SGT: Self-Guided Transformer for Few-Shot Semantic Segmentation
  126. SIANet: Support Information-Aware Network for Category-Agnostic Pose Estimation
  127. SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
  128. SIMFALL: A Data Generator for RF-Based Fall Detection
  129. SIMMKD: Simple Mask-Flow Keypoint Detection for Both Typhoon Detection and Typhoon Eye Location
  130. SJTU-TMQA: A Quality Assessment Database for Static Mesh with Texture Map
  131. SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual Attention
  132. SO-Net: Model-Agnostic Sequential Hand Pose Optimization Framework
  133. SPASE: Spatial Saliency Explanation For Time Series Models
  134. SPATIALCODEC: Neural Spatial Speech Coding
  135. SPCL-MER: Supervised Prototypical Contrastive Learning for Micro-Expression Recognition
  136. SPDG-Net: Semantics Preserving Domain Augmentation through Style Interpolation for Multi-Source Domain Generalization
  137. SPEC-NERF: Multi-Spectral Neural Radiance Fields
  138. SPGFusion: A Semantic Prior Guided Infrared and Visible Image Fusion Network
  139. SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance
  140. SPTESleepNet: Automatic Sleep Staging Model Based On Strip Patch Embeddings And Transformer Encoder
  141. SPY-Watermark: Robust Invisible Watermarking for Backdoor Attack
  142. SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification
  143. SR-VFA: Accurate Self-Refined Face Alignment in Videos
  144. SRECT: Machine-Specific Spatial-Resolution Enhancement in Computed Tomography
  145. SRP-UOD: Multi-Branch Hybrid Network Framework Based on Structural Re-Parameterization for Underwater Small Object Detection
  146. SSHNN: Semi-Supervised Hybrid NAS Network for Echocardiographic Image Segmentation
  147. SSL-Net: A Synergistic Spectral and Learning-Based Network for Efficient Bird Sound Classification
  148. SSR-GPCsT: Deep Learning Models Based on Functional Connectivity Maps in Autism Research
  149. SSTA: Salient Spatially Transformed Attack
  150. STEMGEN: A Music Generation Model That Listens
  151. STREAMVC: Real-Time Low-Latency Voice Conversion
  152. STS-CCL: Spatial-Temporal Synchronous Contextual Contrastive Learning for Urban Traffic Forecasting
  153. STYLECAP: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-Supervised Learning Models
  154. STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
  155. SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
  156. SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker
  157. Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach
  158. Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image Captioning
  159. Sampling and Recovery of Signals Over Product Cell Structures
  160. Sandwiched Lo-Res Simulation for Scalable Flood Modeling
  161. Scalable Ensemble-Based Detection Method Against Adversarial Attacks For Speaker Verification
  162. Scalable Model-Based Gaussian Process Clustering
  163. Scalable and Efficient Speech Enhancement Using Modified Cold Diffusion: A Residual Learning Approach
  164. Scale-Aware Competition Network for Palmprint Recognition
  165. Scale-Free And Task-Generic Attack: Generating Photo-Realistic Adversarial Patterns With Patch Quilting Generator
  166. Scaling Results for Robust Distributed Estimation in Sensor Networks Using Order Statistics
  167. ScanPCGC: Learning-Based Lossless Point Cloud Geometry Compression using Sequential Slice Representation
  168. Scene Sketch-to-Image Synthesis Based on Multi-Object Control
  169. Score Calibration Based on Consistency Measure Factor for Speaker Verification
  170. Score-based Diffusion Models for Photoacoustic Tomography Image Reconstruction
  171. ScoreDec: A Phase-Preserving High-Fidelity Audio Codec with a Generalized Score-Based Diffusion Post-Filter
  172. SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability
  173. Seam Mask Guided Partial Reconstruction with Quantum-Inspired Local Aggregation For Deep Image Stitching
  174. Search Robust and Adaptable Architecture
  175. Search for Gravitational Wave Probes - A Self-Supervised Learning for Pulsars Based on Signal Contexts
  176. Sec2Sec Co-Attention Transformer for Video-Based Apparent Affective Prediction
  177. Sector-Based Interference Cancellation for Robust Keyword Spotting Applications Using an Informed MPDR Beamformer
  178. Secure Energy Efficiency Fairness Maximization in Backscatter Throughput Constrained UAV-Assisted Data Collection
  179. Securely and Efficiently Outsourcing Neural Network Inference via Parallel MSB Extraction
  180. Security Equivalence Assessment between Cloud Standards by Mapping of Control Items
  181. Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model
  182. Seeking Similarities While Removing Differences: Graph Neural Networks Based on Node Correlation
  183. Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change Detection
  184. Segment Anything Model Meets Image Harmonization
  185. Segment then Match: Find the Carrier before Reasoning in Scene-Text VQA
  186. Segmentation-Driven Infrared and Visible Image Fusion Via Transformer-Enhanced Architecture Searching
  187. Segmented Error Minimisation (Semi) for Robust Training of Deep Learning Models with Non-Linear Shifts in Reference Data
  188. Selecting N-Lowest Scores for Training MOS Prediction Models
  189. Selective Domain-Invariant Feature for Generalizable Deepfake Detection
  190. Selective User Forwarded Cell-Free Massive Mimo with Quantized Symbols
  191. Self Knowledge Distillation Based On Layer-Wise Weighted Feature Imitation For Efficient Object Detection
  192. Self-Adaptive Scale Handling for Forecasting Time Series with Scale Heterogeneity
  193. Self-Distilled Dynamic Fusion Network for Language-Based Fashion Retrieval
  194. Self-Knowledge Distillation with Learning from Role-Model Samples
  195. Self-Motion As Supervision For Egocentric Audiovisual Localization
  196. Self-Supervised Adaptive AV Fusion Module for Pre-Trained ASR Models
  197. Self-Supervised Adaptive Pre-Training of Multilingual Speech Models for Language and Dialect Identification
  198. Self-Supervised Cross-Level Consistency Learning For Fundus Image Classification
  199. Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion Recognition
  200. Self-Supervised Dual Generative Networks for Edge-Preserving Image Smoothing
  201. Self-Supervised Face Image Restoration with a One-Shot Reference
  202. Self-Supervised Learning for Anomalous Sound Detection
  203. Self-Supervised Learning for Sleep Stage Classification with Temporal Augmentation and False Negative Suppression
  204. Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
  205. Self-Supervised Multi-Scale Hierarchical Refinement Method for Joint Learning of Optical Flow and Depth
  206. Self-Supervised Path Planning in UAV-Aided Wireless Networks Based on Active Inference
  207. Self-Supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
  208. Self-Supervised Pulse-Aware Interpretable Disentangled ECG Representation Learning
  209. Self-Supervised Reinforcement Learning for Out-of-Distribution Recovery via Auxiliary Reward
  210. Self-Supervised Spatially Variant PSF Estimation for Aberration-Aware Depth-from-Defocus
  211. Self-Supervised Speaker Verification Employing A Novel Clustering Algorithm
  212. Self-Supervised Speaker Verification with Adaptive Threshold and Hierarchical Training
  213. Self-Training Domain Adaptation Via Weight Transmission Between Generators
  214. SemDA: Communication-Efficient Data Aggregation Through Distributed Semantic Transmission
  215. Semantic Distillation and Structural Alignment Network for Fake News Detection
  216. Semantic Enrichment for Video Question Answering with Gated Graph Neural Networks
  217. Semantic Latent Decomposition with Normalizing Flows for Face Editing
  218. Semantic Proximity Alignment: Towards Human Perception-Consistent Audio Tagging by Aligning with Label Text Description
  219. Semantic Reconstruction of Continuous Language from Meg Signals
  220. Semantic Security: A Digital Watermark Method for Image Semantic Preservation
  221. Semantic Segmentation for Multi-Scene Remote Sensing Images with Noisy Labels Based on Uncertainty Perception
  222. Semantic-Enhanced Supervised Contrastive Learning
  223. Semantic-Guided Network with Contrastive Learning for Video Caption
  224. Semantic-Preserving Image Coding Based on Conditional Diffusion Models
  225. Semanticmapper: Region-Specific Domain Adaptation for 3D Shapes Through Lexical Delineation
  226. Semantics Driven Multi-View Knowledge Graph Embedding for Cross-Lingual Entity Alignment
  227. Semi-Autoregressive Streaming ASR with Label Context
  228. Semi-Blind Estimation of Direct-to-Reverberant Energy Ratio Using Residual Energy Test Statistics
  229. Semi-Decoupled 6D Pose Estimation via Multi-Modal Feature Fusion
  230. Semi-Supervised Domain Adaptation for Eeg-Based Sleep Stage Classification
  231. Semi-Supervised Metrics-Based Self-Training Root Cause Analysis for Cloud-Native Systems with Class-Imbalanced Data
  232. Semi-Supervised Sound Event Detection with Local and Global Consistency Regularization
  233. Semi-Supervised Volumetric Medical Image Segmentation via Class Prototype Guided Distribution-Aligned Representation Learning
  234. Sensi-Bert: Towards Sensitivity Driven Fine-Tuning for Parameter-Efficient Language Model
  235. Sensing with Random Signals
  236. Sensing-Aided Communication Channel Estimation with Tensor-Based Moving Target Localization
  237. Sensing-Assisted Distributed User Scheduling and Beamforming in Muli-Cell mmWave Networks
  238. Sequence of Linear Program for Robust Phase Retrieval
  239. Sequential Acquisition of Features and Experts for Datum-Wise Classification
  240. Sequential Detection of Anomalies in Noisy Outputs of an Unknown Function Using Gaussian and Yule-Simon Processes
  241. Sequential Monte Carlo Graph Convolutional Network for Dynamic Brain Connectivity
  242. Sequential Wasserstein Uncertainty Sets for Minimax Robust Online Change Detection
  243. Shapley Value Guided Extractive Text Summarization
  244. Shift Operator and Separation Filter for Different Period Mixed Signals Using Companion Matrix
  245. Shifted-Rectangle-Window Based Transformer for non-Displaced Femoral Neck Fracture Diagnosis
  246. Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter Model
  247. Signal Reconstruction from Nonideal Samples in Fractional Fourier Transform Domain
  248. Signal Transformer: Complex-Valued Attention and Meta-Learning for Signal Recognition
  249. Significant ASR Error Detection for Conversational Voice Assistants
  250. Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search
  251. Similarity Knowledge Distillation with Calibrated Mask
  252. Simple Contrastive Representation Learning for Time Series Forecasting
  253. Simultaneous Interior and Exterior Sound Field Synthesis Using Cylindrical and Spherical Loudspeaker Arrays
  254. Simultaneous Positioning and Tracking Using Dynamic Factor Graphs and Geometric Average Fusion
  255. SingFake: Singing Voice Deepfake Detection
  256. Single Image Reflection removal Using Feature Difference Enhancement
  257. Single and Few-Step Diffusion for Generative Speech Enhancement
  258. Single-Pixel Imaging Of Dynamic Flows Using Neural Ode Regularization
  259. Single-Source Domain Generalization in Fundus Image Segmentation Via Moderating and Interpolating Input Space Augmentation
  260. Situation-Aware Adaptive Transmit Beamforming for Automotive Radars
  261. Situational Signal Processing with Ecological Momentary Assessment: Leveraging Environmental Context for Cochlear Implant Users
  262. Sketch-Based 3D Shape Retrieval With Multi-View Fusion Transformer
  263. Sketched Column-Based Matrix Approximation With Side Information
  264. SkillNet-X: A Multilingual Multitask Model with Sparsely Activated Skills
  265. Skin Tone Disentanglement in 2D Makeup Transfer With Graph Neural Networks
  266. Skip-Step Contrastive Predictive Coding for Time Series Anomaly Detection
  267. SlideSpeech: A Large Scale Slide-Enriched Audio-Visual Corpus
  268. Slowfast Network for Continuous Sign Language Recognition
  269. Small Object Detection on the Water Surface Based on Radar and Camera Fusion
  270. Small-Footprint Automatic Speech Recognition System using Two-Stage Transfer Learning based Symmetrized Ternary Weight Network
  271. Small-Footprint Convolutional Neural Network with Reduced Feature Map for Voice Activity Detection
  272. Smooth Start: A Unified Approach for Gradual Transition from Cold to Old in Recommender Systems
  273. Snapshot Prompt Ensemble for Parameter-Efficient Soft Prompt Transfer
  274. Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs Detection
  275. Social Learning with Adaptive Models
  276. Social Lode: Human Trajectory Prediction with Latent Odes
  277. Sod-Uav: Small Object Detection For Unmanned Aerial Vehicle Images Via Improved Yolov7
  278. Soft Alignment of Modality Space for End-to-End Speech Translation
  279. Soft Dynamic Time Warping with Variable Step Weights
  280. Soft Image Segmentation Using Gradient Graph Laplacian Regularizer
  281. Solution and Analysis For 3-D Localization In Closed-Form Integrating Sa and TDOA Measurements
  282. Sorting, Reasoning, and Extraction: An Easy-to-Hard Reasoning Framework for Document-Level Event Argument Extraction
  283. SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
  284. Source-Free Domain Adaptation for Millimeter Wave Radar Based Human Activity Recognition
  285. Source-Free Online Domain Adaptive Semantic Segmentation of Satellite Images Under Image Degradation
  286. SourceP: Detecting Ponzi Schemes on Ethereum with Source Code
  287. Space-Time Adaptive Processing for Radars in Connected and Automated Vehicular Platoons
  288. Sparse Bayesian Learning-Based Direct Localization for Distributed Sensor Arrays with Unknown Gain and Phase Errors
  289. Sparse Bayesian Synthetic Aperture Processing Based DOA Estimation with Deformed Towed Arrays
  290. Sparse Channel Representation and Estimation in Near Field Communications
  291. Sparse PCA with False Discovery Rate Controlled Variable Selection
  292. Sparse Regularization Based on Reverse Ordered Weighted L1-Norm and Its Application to Edge-Preserving Smoothing
  293. Sparse Sound Field Representation Using Complex Orthogonal Matching Pursuit
  294. Sparse, Weight-Constrained Arrays With O(N) Aperture for Reduced Mutual Coupling
  295. Sparsely Shared Lora on Whisper for Child Speech Recognition
  296. Sparsespikformer: A Co-Design Framework for Token and Weight Pruning in Spiking Transformer
  297. Spatial Formation-Guided Network for Group Activity Recognition
  298. Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
  299. Spatial-Temporal Interaction Decoding Transformer for Unsupervised Multivariate Time Series Anomaly Detection
  300. Spatio-Temporal Action Detection with a Motion Sense and Semantic Correction Framework
  301. Spatio-Temporal Correlation Learning for Multiple Object Tracking
  302. Spatio-Temporal Data Mining with Information Integrity Protection: Graph Signal Based Air Quality Prediction
  303. Spatiotemporal Group Anomaly Detection via Graph Total Variation on Tensors
  304. Speak While You Think: Streaming Speech Synthesis During Text Generation
  305. Speaker Adaptation For Enhancement Of Bone-Conducted Speech
  306. Speaker Anonymization Using Neural Audio Codec Language Models
  307. Speaker-Adaptive Lipreading Via Spatio-Temporal Information Learning
  308. Speaker-Centric Multimodal Fusion Networks for Emotion Recognition in Conversations
  309. SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
  310. Spectral Analysis of Vowels and Fricatives at Varied Levels of Dysarthria Severity for Amyotrophic Lateral Sclerosis
  311. Spectral Graph Neural Networks with Generalized Laguerre Approximation
  312. Spectro-Spatial Hyperspectral Image Reconstruction From Interferometric Acquisitions
  313. Spectrogram Smoothing for Estimation of the Evolutionary Spectra of Uniformly Modulated Processes
  314. SpectrumNet: Spectrum-Based Trajectory Encode Neural Network for Pedestrian Trajectory Prediction
  315. Speech Collage: Code-Switched Audio Generation by Collaging Monolingual Corpora
  316. Speech Emotion Recognition with Distilled Prosodic and Linguistic Affect Representations
  317. Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone Signal
  318. Speech Foundation Models on Intelligibility Prediction for Hearing-Impaired Listeners
  319. Speech Guided Masked Image Modeling for Visually Grounded Speech
  320. Speech Relationship Learning for Cross-Corpus Speech Emotion Recognition
  321. Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition
  322. Speech-Driven Emotional 3d Talking Face Animation Using Emotional Embeddings
  323. SpeechDPR: End-To-End Spoken Passage Retrieval For Open-Domain Spoken Question Answering
  324. Spiking Structured State Space Model for Monaural Speech Enhancement
  325. Spiking-Leaf: A Learnable Auditory Front-End for Spiking Neural Networks
  326. Spiral Shape Matters: Novel Bio-Inspired Cochlear Cepstrum
  327. Spontts: Modeling and Transferring Spontaneous Style for TTS
  328. Spoofing Attack Augmentation: Can Differently-Trained Attack Models Improve Generalisation?
  329. Srcodec: Split-Residual Vector Quantization for Neural Speech Codec
  330. Stability of Graph Convolutional Neural Networks Through The Lens of Small Perturbation Analysis
  331. Stable Distillation: Regularizing Continued Pre-Training for Low-Resource Automatic Speech Recognition
  332. Stable Knowledge Transfer for Contrastive Distillation
  333. Stable Optimization for Large Vision Model Based Deep Image Prior in Cone-Beam CT Reconstruction
  334. StableMiss+: Prediction with Incomplete Data Under Agnostic Mask Distribution Shift
  335. Stack-and-Delay: A New Codebook Pattern for Music Generation
  336. Stage-Regularized Neural Stein Critics For Testing Goodness-Of-Fit Of Generative Models
  337. State-Augmented Information Routing In Communication Systems With Graph Neural Networks
  338. Stateful Conformer with Cache-Based Inference for Streaming Automatic Speech Recognition
  339. Statistical and Computational Limits of Detecting and Recovering Hidden Submatrices
  340. Stealthy Backdoor Attack Towards Federated Automatic Speaker Verification
  341. Stein Variational Gradient Descent-Based Detection for Random Access with Preambles in MTC
  342. Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity Consistency
  343. Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split Network
  344. Stethoscope-Guided Supervised Contrastive Learning for Cross-Domain Adaptation on Respiratory Sound Classification
  345. Stochastic Configuration Networks for Laboratory Seismic Time-to-Failure Prediction
  346. StofNet: Super-Resolution Time of Flight Network
  347. StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
  348. Straightforward Adaptation of Particle Filter to Fish Eye Images for Top View Pedestrian Tracking
  349. Strategic Arms with Side Communication Prevail Over Low-Regret MAB Algorithms
  350. Streaming Active Learning for Regression Problems Using Regression via Classification
  351. Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
  352. String Sound Synthesizer On Gpu-Accelerated Finite Difference Scheme
  353. Structure Matters: Analyzing Videos Via Graph Neural Networks for Social Media Platform Attribution
  354. Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator
  355. Structure-Informed Positional Encoding for Music Generation
  356. Study of Abuse Detection in Continuous Speech for Indian Languages
  357. Style Adaptation for Domain-Adaptive Semantic Segmentation
  358. Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis
  359. Subgroup Identification Through Multiplex Community Structure Within Functional Connectivity Networks
  360. Subnetwork-To-Go: Elastic Neural Network with Dynamic Training and Customizable Inference
  361. Subspace-Based Co-Array Processing For Nested Arrays without Eigendecomposition
  362. Subspace-Based Detection in OFDM ISAC Systems Under Different Constellations
  363. Subtype-Specific Biomarkers of Alzheimer's Disease from Anatomical and Functional Connectomes via Graph Neural Networks
  364. Summarizing Community-Based Question-Answer Pairs with Focus Rectification
  365. Sunflower Strategy for Bayesian Relational Data Analysis
  366. SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
  367. Supplementing Missing Visions Via Dialog for Scene Graph Generations
  368. Surface-Constrained Progressive Feature Preserving Point Cloud Compression
  369. SweepMM: A High-Quality Multimodal Dataset for Sweeping Robots in Home Scenarios for Vision-Language Model
  370. Syllable Level Features for Parkinson's Disease Detection from Speech
  371. Symmetric Consistency with Cross-Domain Mixup for Cross-Modality Cardiac Segmentation
  372. Symmetric VAR(1) Modelling with Guaranteed Stability
  373. Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley Synthesis
  374. Synchformer: Efficient Synchronization From Sparse Cues
  375. Synonym Replacement and Generation Enhancement for Document Augmentation
  376. SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
  377. Synthesizing Aβ-Pet Via An Image And Label Conditioning Latent Diffusion Model For Detecting Amyloid Status
  378. Synthesizing Black-Box Anti-Forensics Deepfakes With High Visual Quality
  379. Synthetic Conversations Improve Multi-Talker ASR
  380. Synthia's Melody: A Benchmark Framework for Unsupervised Domain Adaptation in Audio
  381. Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset
  382. T-EnFP: An Efficient Transformer Encoder-Based System for Driving Behavior Classification
  383. T-Foley: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
  384. T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
  385. T-SOT FNT: Streaming Multi-Talker ASR with Text-Only Domain Adaptation Capability
  386. TA2P: Task-Aware Adaptive Pruning Method for Image Classification on Edge Devices
  387. TACos: Learning Temporally Structured Embeddings for Few-Shot Keyword Spotting with Dynamic Time Warping
  388. TALDS-Net: Task-Aware Adaptive Local Descriptors Selection for Few-Shot Image Classification
  389. TAROT: A Hierarchical Framework with Multitask co-pretraining on Semi-Structured Data Towards Effective Person-Job fit
  390. TB-ResNet: Bridging the Gap from TDNN to ResNet in Automatic Speaker Verification with Temporal-Bottleneck Enhancement
  391. TCMP: End-to-End Topologically Consistent Magnitude Pruning for Miniaturized Graph Convolutional Networks
  392. TCNAS: Transformer Architecture Evolving in Code Clone Detection
  393. TD-GPT: Target Protein-Specific Drug Molecule Generation GPT
  394. TDT-KWS: Fast and Accurate Keyword Spotting Using Token-and-Duration Transducer
  395. TF-SepNet: An Efficient 1D Kernel Design in Cnns for Low-Complexity Acoustic Scene Classification
  396. TIA: A Teaching Intonation Assessment Dataset in Real Teaching Situations
  397. TNFormer: Single-Pass Multilingual Text Normalization with a Transformer Decoder Model
  398. TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-Device ASR Models
  399. TRET: Two Stream-Based Regionally Enhanced Transformers for Person Re-Identification
  400. TRLS: A Time Series Representation Learning Framework Via Spectrogram for Medical Signal Processing
  401. TRUST-SER: On The Trustworthiness Of Fine-Tuning Pre-Trained Speech Embeddings For Speech Emotion Recognition
  402. Tackling Electrode Shift in Gesture Recognition with HD-EMG Electrode Subsets
  403. Tag Antenna Structure Calibrated Backscattering Signal Detection
  404. Tail Classes Matter: Long-Tailed Object Detection Revisited
  405. TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning
  406. Talking Face Generation for Impression Conversion Considering Speech Semantics
  407. Taming Prompt-Based Data Augmentation for Long-Tailed Extreme Multi-Label Text Classification
  408. Target Localization Based on Multistatic Mimo Radar via Double Coupled Canonical Polyadic Decomposition
  409. Target Optimization Direction Guided Transfer Learning for Image Classification
  410. Target Signal Power Improvement and Clutter Suppression via Beamforming for Integrated Sensing and Communication Systems
  411. Target Speaker Extraction by Directly Exploiting Contextual Information in the Time-Frequency Domain
  412. Target Speech Extraction with Pre-Trained Self-Supervised Learning Models
  413. Task Indicating Transformer for Task-Conditional Dense Predictions
  414. Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
  415. Task Selection and Assignment for Multi-Modal Multi-Task Dialogue Act Classification with Non-Stationary Multi-Armed Bandits
  416. Task Vector Algebra for ASR Models
  417. Task-Wise Prompt Query Function for Rehearsal-Free Continual Learning
  418. Template-Guided Data Augmentation for Unbiased Scene Graph Generation
  419. Tempo Estimation as Fully Self-Supervised Binary Classification
  420. Temporal Conditional Coding for Dynamic Point Cloud Geometry Compression
  421. Temporal Convolution Shrinkage Network for Keyword Spotting
  422. Temporal Inconsistency-Based Active Learning
  423. Temporal Knowledge Graph Embedding using Householder Transformations
  424. Temporal Relational Context Learning for Extrapolation Reasoning on Temporal Knowledge Graphs
  425. Temporal-Spatial Prediction: Pre-Training on Diverse Datasets for EEG Classification
  426. Temporally-Guided Total Variation For Robust Spatiotemporal Fusion Of Satellite Images
  427. Ten-Guard: Tensor Decomposition for Backdoor Attack Detection in Deep Neural Networks
  428. Tensor Decomposition-Based Data Fusion for Biomarker Extraction from Multiple EEG Experiments
  429. Tensor Graph Decomposition for Temporal Networks
  430. Tensor Low-Rank Approximation of Finite-Horizon Value Functions
  431. Tensor Reconstruction-Based Sparse Array 2-D DOA Estimation of Mixed Coherent and Uncorrelated Signals
  432. Tensor-Guided Interpolation For Off-Grid Power Spectrum Map Construction
  433. Tensorial Convolutive Blind Source Separation
  434. Test-Time Distribution Learning Adapter for Cross-Modal Visual Reasoning
  435. Text Region Multiple Information Perception Network for Scene Text Detection
  436. Text-Driven Talking Face Synthesis by Reprogramming Audio-Driven Models
  437. Text-Only Unsupervised Domain Adaptation for Neural Transducer-Based ASR Personalization Using Synthesized Data
  438. Text-Video Completion Networks With Motion Compensation And Attention Aggregation
  439. Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable Attribute
  440. TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech Models
  441. Textual Tokens Classification for Multi-Modal Alignment in Vision-Language Tracking
  442. Texture and Normal Map Estimation for 3D Face Reconstruction
  443. Texture-Unet: A Texture-Aware Network for Bone Marrow Smear Whole-Slide Image Region of Interest Segmentation
  444. The 2nd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing Aid Intelligibility Prediction
  445. The Collaboration of 3D Convolutions and CRO-TSM in Lipreading
  446. The Devil is in Details: Delving Into Lite FFN Design for Vision Transformers
  447. The Double-Edged Sword Of Ai Safety: Balancing Anomaly Detection and OOD Generalization Via Model Anchoring
  448. The Effects of Loudness and Smiling on Timbre Features: Implications for Charismatic Voices in Mandarin, German and Danish
  449. The Joint Grid-Free DOA and Polarization Estimation Algorithm based on Atomic Norm Minimization
  450. The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction
  451. The Power of Few: Accelerating and Enhancing Data Reweighting with Coreset Selection
  452. The Rao, Wald, And Likelihood-Ratio Tests under Generalized Self-Concordance
  453. The Selectivity and Competition of the Mind's Eye in Visual Perception
  454. Theme-Enhanced Hard Negative Sample Mining for Open-Domain Question Answering
  455. Think as People: Context-Driven Multi-Image News Captioning with Adaptive Dual Attention
  456. Three-Dimensional Decoupled Atomic Norm Minimization
  457. Three-Dimensional Sound Wave Propagation Reproduction by CE-FDTD Simulation Applying Actual Radiation Characteristics
  458. Three-Dimensional Spatial-Temporal Near-Field Passive Localization Based on an Exact Spatial Propagation Model
  459. Through-The-Wall Radar Imaging With Wall Clutter Removal Via Riemannian Optimization On The Fixed-Rank Manifold
  460. Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
  461. Time Changed Normalizing Flows for Accurate SDE Modeling
  462. Time-Interval Visual Saliency Prediction in Mammogram Reading
  463. Time-Modulated Intelligent Reflecting Surface for Waveform Security
  464. Titan: Bringing the Deep Image Prior to Implicit Representations
  465. Token-Based Spatiotemporal Representation of the Events
  466. Tokenmotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection VIA Learnable Token Selection
  467. Topological Neural Networks over the Air
  468. Topology-Dependent Privacy Bound for Decentralized Federated Learning
  469. Topology-Regularized Self-Knowledge Distillation for Transductive-Inductive Learning of Brain Disorder Diagnosis
  470. Touring Sampling With Pushforward Maps
  471. Toward Quantifiable Face age Transformation
  472. Toward Sufficient Spatial-Frequency Interaction for Gradient-Aware Underwater Image Enhancement
  473. Towards 3D Computational Persicopy with an Ordinary Camera: a Separable Non-Linear Least Squares Formulation
  474. Towards A World-English Language Model for on-Device Virtual Assistants
  475. Towards ASR Robust Spoken Language Understanding Through in-Context Learning with Word Confusion Networks
  476. Towards Automatic Data Augmentation for Disordered Speech Recognition
  477. Towards Building The Federatedgpt: Federated Instruction Tuning
  478. Towards Controlled Table-to-Text Generation with Scientific Reasoning
  479. Towards Disease-Aware Self-Supervised Dynamic Brain Network Learning For Mental Diagnosis
  480. Towards Efficient Modeling and Inference in Multi-Dimensional Gaussian Process State-Space Models
  481. Towards Enabling DPOAE Estimation on Single-Speaker Earbuds
  482. Towards End-to-End Spoken Grammatical Error Correction
  483. Towards Faster End-to-End Data Transmission Over Voice Channels
  484. Towards Generic Deepfake Detection with Dynamic Curriculum
  485. Towards High Resolution Weather Monitoring With Sound Data
  486. Towards High-Performance and Low-Latency Feature-Based Speaker Adaptation of Conformer Speech Recognition Systems
  487. Towards Improving Speech Emotion Recognition Using Synthetic Data Augmentation from Emotion Conversion
  488. Towards Intelligent Design: A Self-Driven Framework for Collocated Clothing Synthesis Leveraging Fashion Styles and Textures
  489. Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate Speech
  490. Towards Multi-Domain Face Landmark Detection with Synthetic Data from Diffusion Model
  491. Towards Omniscient Feature Alignment for Video Rescaling
  492. Towards Optimal Voice Disentanglement with Weak Supervision
  493. Towards Optimized Multi-Channel Modulo-ADCs: Moduli Selection Strategies and Bit Depth Analysis
  494. Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-Training and Multi-Modal Tokens
  495. Towards Resource-Efficient and Secure Federated Multimedia Recommendation
  496. Towards Robust Multimodal Prompting with Missing Modalities
  497. Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
  498. Towards Video-Text Retrieval Adversarial Attack
  499. Towards a Unified View of Adversarial Training: A Contrastive Perspective
  500. Towards an Interpretable Representation of Speaker Identity via Perceptual Voice Qualities
  501. Towards an Objective Quality Metric for Interpolated Directional Room Impulse Responses
  502. Tracking Beyond the Unambiguous Range with Modulo Single-Photon Lidar
  503. Tracking of Multiple Spawning Targets with Heterogeneous Sensors for Seabed-To-Space Situational Awareness
  504. Trades++: Enhancing Multi-Object Tracking of Real Low Confidence Targets Using a Pyramid-Like Self-Attention Model
  505. Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing
  506. Training Audio Captioning Models without Audio
  507. Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
  508. Training Ultra-Low-Latency Spiking Neural Networks from Scratch
  509. Trajectory set Empowered Hypergraph Transformer for Mobile Sensor Based Traffic Prediction
  510. TranSentence: speech-to-speech Translation via Language-Agnostic Sentence-Level Speech Encoding without Language-Parallel Data
  511. TransAVS: End-to-End Audio-Visual Segmentation with Transformer
  512. TransCycle: A Data Augmentation Method for 3D Human Pose Estimation
  513. TransMUSIC: A Transformer-Aided Subspace Method for DOA Estimation with Low-Resolution ADCS
  514. Transducers with Pronunciation-Aware Embeddings for Automatic Speech Recognition
  515. Transfer the Linguistic Representations from TTS to Accent Conversion with Non-Parallel Data
  516. Transferable Models for Bioacoustics with Human Language Supervision
  517. Transferring Structure Knowledge: A New Task to Fake News Detection towards Cold-Start Propagation
  518. Transformer Model with Multi-Type Classification Decisions for Intrusion Attack Detection of Track Traffic and Vehicle
  519. Transformer-Inspired Lightweight Model for Efficient Time Series Forecasting
  520. Transforming Cardiovascular Health: a Transformer-Based Approach to Continuous, Non-Invasive Blood Pressure Estimation via Radar Sensing
  521. Translatotron 3: Speech to Speech Translation with Monolingual Data
  522. Transmit Beampattern Optimization for MIMO-ISAC Systems with Hybrid Beamforming
  523. Transmitting Data Through Reconfigurable Intelligent Surface: A Spatial Sigma-Delta Modulation Approach
  524. Tree Network Design for Faster Distributed Machine Learning Process with Distributed Dual Coordinate Ascent
  525. Tree of Uncertain Thoughts Reasoning for Large Language Models
  526. Treemil: A Multi-Instance Learning Framework for Time Series Anomaly Detection with Inexact Supervision
  527. Trend-Heuristic Reinforcement Learning Framework for News-Oriented Stock Portfolio Management
  528. Trusted Deep Domain Adaptation with Uncertainty Measure Based on Evidence Theory
  529. Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
  530. Two-Edge-Resolved 3d Non-Line-of-Sight Imaging: A Fisher Information Equalized Discretization
  531. Two-Stage Acoustic Echo Cancellation Network with Dual-Path Alignment
  532. Two-Stage Transfer Learning for Fusion and Classification of Airborne Hyperspectral Imagery
  533. Two-Step Knowledge Distillation for Tiny Speech Enhancement
  534. Type-Aware Decoding Via Explicitly Aggregating Event Information for Document-Level Event Extraction
  535. U2R: Underwater Ultrasonic Reflection Wave Dataset Toward Pose-Invariant Material Recognition
  536. UAV Operation Time Minimization for Wireless-Powered Data Collection
  537. UAV-Based Dynamic Object Tracking with Radio Map
  538. UNAD: Universal Anatomy-Initialized Noise Distribution Learning Framework Towards Low-Dose CT Denoising
  539. UNIDEAL: Curriculum Knowledge Distillation Federated Learning
  540. UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
  541. UNeC: Unsupervised Exploring In Controllable Space
  542. USM-Lite: Quantization and Sparsity Aware Fine-Tuning for Speech Recognition with Universal Speech Models
  543. USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
  544. Ultra Low Complexity Deep Learning Based Noise Suppression
  545. Ultra-Lightweight Neural Differential DSP Vocoder for High Quality Speech Synthesis
  546. Ultra-Low Delay Lossless Compression of Higher Order Ambisonics
  547. Uncertainty Quantification in Deep Learning Based Kalman Filters
  548. Uncertainty-Guided Contrastive Learning For Single Source Domain Generalisation
  549. Uncertainty-Guided Person Search Model with Auxiliary Shallow Feature Exploration
  550. Uncertainty-Guided Physics-Driven Deep Learning Reconstruction via Cyclic Measurement Consistency
  551. Uncovering Strong Ties: A Study of Indirect Sybil Attack on Signed Social Network
  552. Underlying-Complementarity and Surrounding-Correspondence for Multi-View Clustering
  553. Understanding Data Augmentation From A Robustness Perspective
  554. Understanding Gaussian Noise Mismatch: A Hellinger Distance Approach
  555. Understanding Probe Behaviors Through Variational Bounds of Mutual Information
  556. UniX-Encoder: A Universal X-Channel Speech Encoder for AD-HOC Microphone Array Speech Processing
  557. Unidirectional Brain-Computer Interface: Artificial Neural Network Encoding Natural Images to FMRI Response in the Visual Cortex
  558. Unified Analysis of Correlation-Aware Joint Sparse Support Recovery with ℓ0-Norm Constraint
  559. Unified Pretraining Target Based Video-Music Retrieval with Music Rhythm and Video Optical Flow Information
  560. Unified Probability Distributions of Generalized Composite Fading with Inverse-Type Distributions of Large-Scale Shadowing/Fluctuations
  561. Unified Speech and Gesture Synthesis Using Flow Matching
  562. Unified Srgb Real Noise Synthesizing with Adaptive Feature Modulation
  563. Unifying One-Shot Voice Conversion and Cloning with Disentangled Speech Representations
  564. Unimodal Aggregation for CTC-Based Speech Recognition
  565. Unintended Memorization in Large ASR Models, and How to Mitigate It
  566. Unitary Approximate Message Passing for Matrix Factorization
  567. Universal Adversarial Attack Against Speaker Recognition Models
  568. Unlabelled Sensing with Priors: Algorithm and Bounds
  569. Unleashing Trigger-Free Event Detection: Revealing Event Correlations Via a Contrastive Derangement Framework
  570. Unlocking Deep Learning: A BP-Free Approach for Parallel Block-Wise Training of Neural Networks
  571. Unravel Anomalies: an End-to-End Seasonal-Trend Decomposition Approach for Time Series Anomaly Detection
  572. Unraveling Explainable Reinforcement Learning Using Behavior Tree Structures
  573. Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric Gan
  574. Unrolled Proximal Gradient Descent Method for Non-Negative Least Squares Problem
  575. Unsupervised Accent Adaptation Through Masked Language Model Correction of Discrete Self-Supervised Speech Units
  576. Unsupervised Acoustic Scene Mapping Based on Acoustic Features and Dimensionality Reduction
  577. Unsupervised Anomaly Detection for Multivariate Time Series Using Diffusion Model
  578. Unsupervised Continual Learning of Image Representation Via Rememory-Based Simsiam
  579. Unsupervised Disparity Estimation for Light Field Videos
  580. Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
  581. Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport
  582. Unsupervised Human Activity Recognition Via Large Language Models and Iterative Evolution
  583. Unsupervised Learning Based End-to-End Delayless Generative Fixed-Filter Active Noise Control
  584. Unsupervised Learning of Facial Optical Flow via Occlusion-Aware Global-Local Matching
  585. Unsupervised Learning of Neural Semantic Mappings with the Hungarian Algorithm for Compositional Semantics
  586. Unsupervised Multi-Channel Separation And Adaptation
  587. Unsupervised Multi-Domain Data Selection for Asr Fine-Tuning
  588. Unsupervised Multiple Choices Question Answering Via Universal Corpus
  589. Unsupervised Optimal Power Flow Using Graph Neural Networks
  590. Unsupervised Pitch-Timbre Disentanglement of Musical Instruments Using a Jacobian Disentangled Sequential Autoencoder
  591. Unsupervised Remote Sensing Haze Removal Based on Saliency-Guided Transmission Refinement
  592. Unsupervised Speech Enhancement with Diffusion-Based Generative Models
  593. Unsupervised Speech Recognition with N-skipgram and Positional Unigram Matching
  594. Unsupervised Topic-Conditional Extractive Summarization
  595. Unsupervised multiple domain translation through controlled Disentanglement in variational autoencoder
  596. Updated Corpora and Benchmarks for Long-Form Speech Recognition
  597. Uplink Symbol Detection in Dynamic TDD Mimo Systems with AP-AP Interference
  598. Urban Traffic Flow Forecasting Based on Spatial-Temporal Graph Contrastive Learning
  599. User-Assisted Networked Sensing in OFDM Cellular Network with Erroneous Anchor Position Information
  600. Using Clustering to Improve the Performance of few-shot Learning
  601. Using Temporal Consistency for Compressed Sensing in High-Resolution mmWave Sounding
  602. Utilizing Second-Order Information in Noisy Information-Sharing Environments for Distributed Optimization
  603. V-DDPM: MRI Rician Noise Removal Model Based on VST and DDPM
  604. VCD: A Video Conferencing Dataset for Video Compression
  605. VFD-Net: Vocoder Fingerprints Detection for Fake Audio
  606. VGDIFFZERO: Text-To-Image Diffusion Models Can Be Zero-Shot Visual Grounders
  607. VIC-KD: Variance-Invariance-Covariance Knowledge Distillation to Make Keyword Spotting More Robust Against Adversarial Attacks
  608. VK-G2T: Vision and Context Knowledge Enhanced Gloss2text
  609. VL-FAS: Domain Generalization via Vision-Language Model For Face Anti-Spoofing
  610. VMCC-NET: Uncovering Challenging Regions in Semi-Supervised Medical Image Segmentation with Voxel Mask Based Cyclic-Consistency Network
  611. VRDMG: Vocal Restoration via Diffusion Posterior Sampling with Multiple Guidance
  612. VT-ReID: Learning Discriminative Visual-Text Representation for Polyp Re-Identification
  613. Variance Reduction Can Improve Trade-Off in Multi-Objective Learning
  614. Variational Analysis of Adversarial Regularization for Solving Inverse Problems
  615. Variational Connectionist Temporal Classification for Order-Preserving Sequence Modeling
  616. Vector Approximate Message Passing for Not So Large N.I.I.D. Generalized I/O Linear Models
  617. Vector Approximate message Passing with Arbitrary I.I.D. Noise Priors
  618. Vector Nonlinear Hawkes Model with Inhibition
  619. Vector Quantization Knowledge Transfer for End-to-End Text Image Machine Translation
  620. ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition
  621. Video Anomaly Prediction: Problem, Dataset and Method
  622. Video-Language Graph Convolutional Network for Human Action Recognition
  623. View Crafting For Instance-Level Representation from Scene Images
  624. Viewing Writing as Video: Optical Flow based Multi-Modal Handwritten Mathematical Expression Recognition
  625. Vision Transformer with 2D Explicit Position Encoding
  626. Vision-Sensor Attention Based Continual Multimodal Egocentric Activity Recognition
  627. Visual Adapt for RGBD Tracking
  628. Visual Prompt Tuning for Weakly Supervised Phrase Grounding
  629. Visual Speech Recognition for Languages with Limited Labeled Data Using Automatic Labels from Whisper
  630. Visual-Linguistic Representation Learning with Deep Cross-Modality Fusion for Referring Multi-Object Tracking
  631. Visually Dehallucinative Instruction Generation
  632. Visually Guided Binaural Audio Generation with Cross-Modal Consistency
  633. Vocal Fold Dynamics for Automatic Detection of Amyotrophic Lateral Sclerosis from Voice
  634. Voice Anonymization for All-Bias Evaluation of the Voice Privacy Challenge Baseline Systems
  635. Voice Toxicity Detection Using Multi-Task Learning
  636. VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching
  637. VoiceLDM: Text-to-Speech with Environmental Context
  638. Volumetric 3d Point Cloud Attribute Compression: Learned Polynomial Bilateral Filter for Prediction
  639. VoxMM: Rich Transcription of Conversations in the Wild
  640. Voxblink: A Large Scale Speaker Verification Dataset on Camera
  641. VoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation Tasks
  642. Vulnerability of Face age Verification to Replay Attacks
  643. WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
  644. WFTNet: Exploiting Global and Local Periodicity in Long-Term Time Series Forecasting
  645. WI-FI based Indoor Monitoring Enhanced by Multimodal Fusion
  646. WIFIACT: Enhancing Human Sensing Through Environment Robust Preprocessing And Bayesian Self-Supervised Learning
  647. Water Leak Detection via Domain Adaptation
  648. WaterDiff: Perceptual Image Watermarks Via Diffusion Model
  649. Wav2vec-VC: Voice Conversion via Hidden Representations of Wav2vec 2.0
  650. Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition
  651. Wavelet-Guided Acceleration of Text Inversion in Diffusion-Based Image Editing
  652. Wavelet-Inspired Multiscale Graph Convolutional Recurrent Network for Traffic Forecasting
  653. Weakly Semi-Supervised Tool Detection in Minimally Invasive Surgery Videos
  654. Weakly Supervised Few-Shot Segmentation Through Textual Prompt
  655. Weakly-Supervised Crowd Counting with Token Attention and Fusion: A Simple and Effective Baseline
  656. What Do Neural Networks Listen to? Exploring the Crucial Bands in Speech Enhancement Using SINC-Convolution
  657. What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise Analysis
  658. When Green Learning Meets Federated Learning: Toward Distributed Learning with Low Complexity and Model Heterogeneity
  659. When Training-Free Nas Meets Vision Transformers: A Neural Tangent Kernel Perspective
  660. Which is the Better Teacher Action? A New Ranking Model and Dataset
  661. Whisper-Based Transfer Learning for Alzheimer Disease Classification: Leveraging Speech Segments with Full Transcripts as Prompts
  662. Widrow-Hoff LMS Adaline Demonstrator for Schools and Colleges
  663. Window-Based Convolutional Sparse Coding: Towards A Unified Framework
  664. X-CAUNET: Cross-Color Channel Attention with Underwater Image-Enhancing Transformer
  665. XMP: A Cross-Attention Multi-Scale Performer for File Fragment Classification
  666. YOLO-Med : Multi-Task Interaction Network for Biomedical Images
  667. ZE-FESG: A Zero-Shot Feature Extraction Method Based on Semantic Guidance for No-Reference Video Quality Assessment
  668. ZIV-Zakai Bound for DOA Estimation with Gain-Phase Error
  669. Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages
  670. Zero Shot Audio To Audio Emotion Transfer With Speaker Disentanglement
  671. Zero- and Few-Shot Sound Event Localization and Detection
  672. Zero-Shot Co-Salient Object Detection Framework
  673. Zero-Shot Imitation Policy Via Search In Demonstration Dataset
  674. Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language Model
  675. Zero-Shot Object Detection with Partitioned Contrastive Feature Alignment
  676. Zigzag Attention: A Structural Aware Module For Lane Detection
  677. mmBaT: A Multi-Task Framework for Mmwave-Based Human Body Reconstruction and Translation Prediction
  678. uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models
  679. uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures

Looking for submission deadlines instead? See the conference deadline calendar.