← All conferences

ICASSP 2025 Accepted Papers

The full list of 3,298 papers accepted at ICASSP 2025 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. "I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities
  2. 2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images
  3. 30+ Years of Source Separation Research: Achievements and Future Challenges
  4. 3D Gaussian Splatting with Grouped Uncertainty for Unconstrained Images
  5. 3D Mesh Saliency Based on Dictionary Learning with Multi-Level Laplacian-Beltrami Operator
  6. 3D Shape Classification by Registration: Neural-Network-Free and Training-Free
  7. 3D TDOA-AOA Quaternion Based Acoustic SLAM for Drone Localization and Source Mapping
  8. 3D View Optimization for Improving Image Aesthetics
  9. 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
  10. 3DGCQA: A Quality Assessment Database for 3D AI-Generated Contents
  11. 3DSignDiff: Towards 3D Sign Language Gesture Generation
  12. 3GPP IVAS Codec - Perspectives on Development, Testing and Standardization
  13. A 3D Attenuation Coefficient based Degradation Estimation for Real Non-Homogeneous Dehazing
  14. A Bandwidth Efficient Dual Function Radar Communication System Based on a MIMO Radar Using OTFS Waveforms
  15. A Bayesian Interpretation of Adaptive Low-Rank Adaptation
  16. A Bayesian Perspective on Uncertainty Quantification for Estimated Graph Signals
  17. A Bilinear Source Separation, Dereverberation, and Background Noise Suppression Algorithm for Augmented Reality Applications
  18. A Blind Super-Resolution Method for Near-Field Channel Estimation with Angle-Range Recovery
  19. A Block Term Decomposition Model Based Algorithm for Tensor Completion of Multidimensional Harmonic Signals
  20. A CT-based Prediction System for Determining Respiratory Support Level in COVID-19 Patients
  21. A Chinese Expressive Long-dialogue Speech Dataset with Scripts
  22. A Clinical Knowledge-Driven Fine-Tuning Strategy for Applying Foundation Model to Fully Automatic Acute Ischemic Stroke Lesion Segmentation on Non-Contrast CT Scans
  23. A Comparative Analysis of Generalised Echo and Interference Cancelling and Extended Multichannel Wiener Filtering for Combined Noise Reduction and Acoustic Echo Cancellation
  24. A Comparative Study of Invariance-Aware Loss Functions for Deep Learning-based Gridless Direction-of-Arrival Estimation
  25. A Conditional KAN Diffusion Network for Human Activity Recognition with Missing Sensor Signal Series
  26. A Continual Learning Approach for Embodied Question Answering with Generative Adversarial Imitation Learning
  27. A Convolutional Recurrent Mixer Network For Radar Meteorological Image Super-Resolution
  28. A Cost-effective Solution for Remote Sensing Image Segmentation via Train/Test-Time Adaptation
  29. A Counterfactual Ultrasound Anti-Interference Self-Supervised Network for B-mode Ultrasound Tongue Extraction
  30. A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
  31. A Cross-Modal Multi-Attitude Framework for the Generation of Space Target ISAR Images
  32. A Deformable-Based Source-Free Unsupervised Domain Adaptation Method for Cervical Cell Detection
  33. A Diffusion Model over Directed Acyclic Graphs for Event Schema Generation
  34. A Distillation-based Future-aware Graph Neural Network for Stock Trend Prediction
  35. A Divide-and-conquer Approach for Sparse Recovery in High Dimensions
  36. A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data
  37. A Domain Adversarial Learning Framework for Major Depression Disorder Diagnosis
  38. A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous Interpretation
  39. A Dual-Perspective Metaphor Detection Framework Using Large Language Models
  40. A Dual-Stream Network with Non-Stationary Characteristics-Enhanced for SST Image Prediction
  41. A Dynamic Edge-Selection Mechanism in HRV Hypergraph Learning for Improved Stress Detection
  42. A Dynamical Equation Approach For Quasi-Periodic Gaussian Processes
  43. A Fast Saturation Based Dehazing Framework with Accelerated Convolution and Attention Block
  44. A Federated Learning Network Intrusion Detection System for Multiple Imbalances
  45. A Federated Learning-Based Intrusion Detection System for Satellite-Terrestrial Integrated Networks
  46. A Framework Based on Data Augmentation for Knowledge Graph Entity Typing
  47. A Frequency-aware Augmentation Network for Mental Disorders Assessment from Audio
  48. A Fuzzy C-Means Clustering Algorithm for Real Medical Image Segmentation
  49. A GNSS-IR Aided Multispectral Satellite Data Fusion for Meter-Level Wide-Area Volumetric Soil Moisture Estimation
  50. A Generalized Graph Signal Processing Framework for Multiple Hypothesis Testing over Networks
  51. A Generative-Augmented Deep Matrix Factorization Model for POI Recommendations
  52. A Geometry-Based Node Activation Method for Relative Localization
  53. A Graph-Based Generative Adversarial Network Model for Inferring Task-State from Resting-State Functional Connectivity Networks
  54. A Grouping Strategy-Based Progressive Fusion Network for Hyperspectral Image Super-Resolution
  55. A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
  56. A Hierarchical Flow for Few-shot Anomaly Detection via Global-local Aggregation Strategy
  57. A Hierarchical Reasoning Framework for Complex Question Answering over Knowledge Graph with Reinforcement Learning
  58. A Hierarchical Taxonomy For Deep State Space Models
  59. A High-Precision Character Cartoon Style Transfer Method Based on VToonify and Diffusion Models
  60. A Hybrid Model for Weakly-Supervised Speech Dereverberation
  61. A Hybrid Probabilistic-Deterministic Model Recursively Enhancing Speech
  62. A Joint Time-Frequency Attention for Leakage Detection in Water Distribution Networks Using Time Series Decomposition
  63. A Key to Effective Multi-task Learning: Separate Query Selection for Task-Synergized Handling and Node Utilization
  64. A Label Co-occurrence Transformation Network for Joint Empathy Detection and Empathy Intent Classification
  65. A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
  66. A Lowrate Variable-Bias Integrate-and-Fire Time Encoding Machine
  67. A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
  68. A Margin-Maximizing Fine-Grained Ensemble Method
  69. A Method for Removing Reflections from Water Surface Images Based on Pre-trained Image Restoration
  70. A Metric for Predicting the Quality of Ambisonic Spatial Audio Reproduced Using Spatially Interpolated or Extrapolated Room Impulse Responses
  71. A MoE Multimodal Graph Attention Network Framework for Multimodal Emotion Recognition
  72. A Model Stealing Attack Against Multi-Exit Networks
  73. A Modified Gain Normalized Step Size Adaptive Algorithm for Improved Online Secondary Path Modelling in Active Noise Control
  74. A Modified Nonlinear Matched Filter for Skewed Noise Based on the Gram-Charlier Expansion
  75. A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
  76. A Multi-Agent Multi-Environment Mixed Q-Learning for Partially Decentralized Wireless Network Optimization
  77. A Multi-Label EEG Dataset for Mental Attention State Classification in Online Learning
  78. A Multi-Modal Information Fusion Model for Automatic Sleep Staging
  79. A Multi-Prior Fusion Network for Video-based Micro-Expression Recognition
  80. A Multi-Stage Feature Pipeline on Timestamped Speech Transcriptions for Dementia Assessment
  81. A Multi-Wavelength Optical Sensing Framework for Calibration-Free Wearable Blood Pressure Monitoring
  82. A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment
  83. A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information
  84. A Multi-scenario Attention-based Generative Model for Personalized Blood Pressure Time Series Forecasting
  85. A Near-Field 3D Parameter Estimation Method Based on a Symmetric Enhanced Nested Array
  86. A New Model for Prototype-based Continual Learning in Hyperspherical Space
  87. A Noisy Label Filter based on GMM Binary Classification for Speaker Verification
  88. A Non-autoregressive Model for Joint STT and TTS
  89. A Novel Audio-Visual Multimodal Semi-Supervised Model Based on Graph Neural Networks for Depression Detection
  90. A Novel Compressive Compound Word Encoding and Independent Word Attention for Symbolic Music Generation
  91. A Novel Decision-Making Model for Playing Board Game Combining Planning and Opponent Behaviors
  92. A Novel Multimodal Method for Decoding Speech Perception from Brain Activities
  93. A Novel Network for Short-Term Wind Speed Prediction: Mitigating Distribution Shift and Feature Loss
  94. A Novel Self-Supervised Contrastive Learning Framework for Masked EEG Motor Imagery Modeling
  95. A Novel Single Continuous Shot Multiple Lesions Endoscopy Report Generation
  96. A Novel Split Deep Unfolding Transformer for Pan-Sharpening
  97. A Novel Underwater Acoustic Signal Denoising Model Based on Complex Convolution Dual-branch Multi-scale Attention Network
  98. A Novel Weighted Sparse Component Analysis for Underdetermined Blind Speech Separation
  99. A Parametric Non-Negative Coupled Canonical Polyadic Decomposition Algorithm for Hyperspectral Super-Resolution
  100. A Plug-and-Play Diffusion-Styled Conversion Model for Domain Discrepancies in Medical Image Segmentation
  101. A Practical Gated Recurrent Transformer Network Incorporating Multiple Fusions for Video Denoising
  102. A Pre-trained Plug-in Mixture-of-LoRAs Model for Transferable Sequential Recommendation
  103. A Pre-training Framework that Encodes Noise Information for Speech Quality Assessment
  104. A Privacy-Preserving Cross-Modal Retrieval Scheme Based on CLIP and Deep Hashing
  105. A Progressive Local Variance-guided Strategy for Improving Data Augmentation Reliability
  106. A Prompt Learning Framework with Large Language Model Augmentation for Few-shot Multi-label Intent Detection
  107. A Proximal Variable Smoothing for Nonsmooth Minimization Involving Weakly Convex Composite with MIMO Application
  108. A Quality-Aware Sampling Framework for Efficient 3D Point Cloud Transmission
  109. A Quantitative Metric Selection Approach for Time-series Forecasting Foundation Models
  110. A Ranking Scheme for Trust Region Multi-agent Reinforcement Learning
  111. A Reinforcement Learning Agent Controlled Multi-branch Small Object Detection Framework
  112. A Riemannian Approach to Ground Metric Learning for Optimal Transport
  113. A Risk Prediction Model for Real Estate Corporations Using High-Target Semantic BERT and Improved GRU
  114. A Robust Distributed Recurrent Neural Network for Multi-Agent Consensus Control
  115. A Robust Online Miscalibration Detection and Correction Method for LiDAR-Camera
  116. A Robust Quality Evaluator for Panoramic Videos
  117. A Scale-Adaptive and Background-Robust Method for Surface Defect Detection
  118. A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language Models
  119. A Self-supervised UAV Detection Method Based on Channel State Information
  120. A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision
  121. A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
  122. A Spherical-Harmonic Domain Selective Spatial Active Noise Control System Based on Sound Field Reproduction
  123. A Structured Neural Network Approach for Learning Improved Iterative Algorithms for SBL
  124. A Study of Improving The Privacy-Utility Trade-off of Task-specific Models with Learnable Privacy
  125. A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification
  126. A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
  127. A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids
  128. A Task-Oriented Real-Time and Robust Feature Compression and Selection Method in Collaborative Intelligence System
  129. A Teacher Action Quality Assessment Method Based on Label Constraint Strategy
  130. A Training-Free Correlation-Weighted Model for Zero-/Few-Shot Industrial Anomaly Detection with Retrieval Augmentation
  131. A Transmitter-Model Unaware Generative Image Compression Framework for Semantic Communication
  132. A Triangular Stable Node Network based on Self-supervised Learning for personalized prediction
  133. A Two-Stage AIGC Image Quality Assessment with T2I Correspondence and Visual Perception
  134. A Two-timescale Primal-dual Algorithm for Decentralized Optimization with Compression
  135. A Unified Hardware Accelerator for Fast Fourier Transform and Number Theoretic Transform
  136. A Unified Joint Contrastive Triplet Loss with Temporal and Frequency Signal Fusion for Diagnosing Heart Murmurs
  137. A Unified Metric for Simultaneous Evaluation of Error Rate and Annotation Cost
  138. A Unified Model for Oral Reading Fluency and Student Prosody
  139. A Unified Spatiotemporal Frequency Graph Neural Network for fMRI-based Brain Functional Connectivity Analysis
  140. A Weakly Supervised Semantic Segmentation Model with Enhanced CLIP Feature Extraction
  141. A Weighted Cross-entropy Loss for Mitigating LLM Hallucinations in Cross-lingual Continual Pretraining
  142. A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction
  143. A decade of DCASE: Achievements, practices, evaluations and future challenges
  144. A first-order DirAC-based parametric Ambisonic coder for immersive communications
  145. A novel multimodal personality prediction method based on pretrained models and graph relational transformer network
  146. A spectrum-enhanced attention model for semantic segmentation of remote sensing images
  147. A-PeARCNN: a Physics-encoded AutoRegressive Convolutional Neural Network with AttentionNet for Solving Partial Differential Equations
  148. A2B: Neural Rendering of Ambisonic Recordings to Binaural
  149. A2GP-SF: Enhancing Few-shot Class Incremental Learning via Attribute Generative Prompting and Adaptive Sharpness Flattening
  150. AAD-DCE: An Aggregated Multimodal Attention Mechanism for Early and Late Dynamic Contrast Enhanced Prostate MRI Synthesis
  151. AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
  152. ACRL-10K: A Dataset for Air Conditioner Refrigerant Leak Smoke Detection
  153. AD2T: Adversarial Distortion Domain Translation for Robust Watermarking against Non-differentiable Distortions
  154. ADC-GS: Pose-Free 3D Gaussian Splatting with Adaptive Depth Consistency
  155. ADC: Enhancing Function Calling Via Adversarial Datasets and Code Line-Level Feedback
  156. ADD: A Detection Method for Image-Processing Adversarial Defenses
  157. AER-LLM: Ambiguity-aware Emotion Recognition Leveraging Large Language Models
  158. AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated Images
  159. AGR: Age Group fairness Reward for Bias Mitigation in LLMs
  160. AI-Generated Music Detection and its Challenges
  161. AIDC: Benchmark for Analytical Learning in Incremental Disease Classification
  162. AKI360: Enabling Highly Interactive 360-degree Video Streaming by Adaptive Keyframe Interval
  163. ALIC: Adaptive Fusion Entropy Model for Learned Image Compression
  164. AMNS: Attention-Weighted Selective Mask and Noise Label Suppression for Text-to-Image Person Retrieval
  165. AMSER: Accelerate Mobile Speech Emotion Recognition with Signal Compression
  166. AMuSE: Attentive Multilingual Speech Encoding for Zero-Prior ASR
  167. ANASETC: Automatic Neural Architecture Search for Encrypted Traffic Classification
  168. AP-Net: Semi-Supervised Ultrasound Cardiac Segmentation Using Enhanced Anatomical Prior
  169. APLASE: Compression using Adaptive Piecewise Linear Approximation and Sparse Encoding
  170. APTSniffer: Detecting APT Attack Traffic Using Retrieval-Augmented Large Language Models
  171. ARIG-GCN: Anatomical Relationship and Isomorphic Graph Approximation Guided Graph Convolutional Network for Automated ASPECTS Scoring on Non-Contrast CT
  172. ARM : nnU-Net with Arena Mechanism for Medical Image Segmentation
  173. AS-Net: Adaptive Style-aware Network for Handwritten Text Generation
  174. ASANet: Scene Text Recognition With Alternate Self-Attention
  175. ASCDomain: Domain Invariant Device-Adversarial Isotropic Knowledge Distillation Convolutional Neural Architecture
  176. ASFC-NeRF: Large-Scale Scene Rendering with Adaptive Sampling and Feature-aware Compression
  177. ASR Benchmarking: Need for a More Representative Conversational Dataset
  178. ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party Scenarios
  179. ATP-TTS: Adaptive Thresholding Pseudo-Labeling for Low-Resource Multi-Speaker Text-to-Speech
  180. AUIED3K: A New Andaman Underwater Image Enhancement Dataset for Deep Learning-Driven Image Enhancement with Minimum Loss Dehazing
  181. AVS3P10 Standard for Real-time Speech Coding
  182. Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
  183. Accelerating Computation for Large-Scale Wide-Band RF Imaging
  184. Accelerating Convergence in Bounding Box Regression with a Refined IoU Loss Function
  185. Accelerometer-Based Person-in-Bed Detection Challenge
  186. AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
  187. Accompaniment Prompt Adherence: A measure for evaluating music accompaniment systems
  188. Accurate 3D Facial Paralysis Analysis Using Multi-View Infrared Structured Light System
  189. Accurate Hardware Trojan Detection for SGIN Device: A Prompt-Tuning and LangChain Approach
  190. AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing
  191. Achieving Robustness in Blind Modulo Analog-to-Digital Conversion
  192. Acoustic Identification of Individual Animals with Hierarchical Contrastive Learning
  193. Acoustic Position Estimation of a Silent Listener
  194. Active Learning for Long-Tailed Annotation
  195. Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
  196. Active Visual Learning for Robots with Dueling Deep Q-Networks and Transformer Encoders
  197. AdaBoost-Based Channel Estimation in One-Bit Millimeter-Wave MIMO
  198. AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
  199. AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
  200. AdapFed: Adaptive Devices Training Strategy for Heterogeneous Federated Learning
  201. Adapt and Feature Translation for Class-Incremental Learning with Pre-Trained Models
  202. AdaptVC: High Quality Voice Conversion with Adaptive Learning
  203. Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
  204. Adapting Large Language Model for Spatio-Temporal Understanding in Next Point-of-Interest Prediction
  205. Adapting Large Language Models to Forecast in Frequency Domain
  206. Adapting Single-Channel Pre-trained Transformer Models for Multi-Channel Sound Event Localization and Detection
  207. Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
  208. Adapting Without Seeing: Text-Aided Domain Adaptation for Adapting CLIP-like Models to Novel Domains
  209. Adaptive Acquisition in Bayesian Optimization with Agnostic Ensembles
  210. Adaptive Aspect Ratios with Patch-Mixup-ViT-based Vehicle ReID
  211. Adaptive Canonical Correlation Analysis With Application to Time Synchronization for Signal Alignment
  212. Adaptive Central Frequencies Locally Competitive Algorithm for Speech
  213. Adaptive Compression of Supervised and Self-Supervised Models for Green Speech Recognition
  214. Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding
  215. Adaptive Decoding for Efficient Automatic Speech Recognition
  216. Adaptive Feature Aggregation for In-Air Handwritten Trajectory
  217. Adaptive Fine-Grained Feature Mining and RoI Feature Interaction Network for Small Object Detection in Aerial Images
  218. Adaptive Gradient-Based Timesurface for Event-based Detection
  219. Adaptive Large Language Models via Attention Shortcuts
  220. Adaptive Layered-Trust Robust Defense Mechanism for Personalized Federated Learning
  221. Adaptive Lossless Compression for Genomics Data by Multiple (s, k)-mer Encoding and XLSTM
  222. Adaptive Multi-Scale Local Correction for Semi-Supervised 3D Medical Image Segmentation
  223. Adaptive Password Guessing Framework Using Various Datasets
  224. Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known Attributes
  225. Adaptive Receptive Field Convolution for Top-view Fisheye Images Segmentation
  226. Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose Estimation
  227. Adaptive Sparse Feature Location Activation Strategy for Sparse Detectors on Drone Images
  228. Adaptive Spatiotemporal Augmentation for Improving Dynamic Graph Learning
  229. Adaptive Time-Frequency Attention Network for Sleep Stage Classification Using Respiratory Signals
  230. Adaptive-Similarity-Based Brain Dynamic Functional Connectivity with Spatial-Temporal Attention and Domain Adaptation for Schizophrenia Diagnosis
  231. AdaptiveDrop: A Simple Adaptive Label Noise Filtering Scheme for Enhanced Self-supervised Speaker Verification
  232. Addressing Emotion Ambiguity and Annotator Subjectivity for Enhanced Speech Emotion Labeling
  233. Addressing Pilot Contamination in Channel Estimation with Variational Autoencoders
  234. Addressing Speed-Induced Dispersion in Stepped-Frequency PMCW Radar Systems
  235. Adopting Whisper for Confidence Estimation
  236. Advanced Graph-MLPs Distillation based on Global and Local Hyperbolic Geometry Learning
  237. Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
  238. Advances in Microphone Array Processing and Multichannel Speech Enhancement
  239. Advancing Active Speaker Detection for Egocentric Videos
  240. Advancing Dark Action Recognition via Modality Fusion and Dark-to-Light Diffusion Model
  241. Advancing Few-Shot Class-Incremental Learning with Virtual Prototype Guidance Prompting
  242. Advancing High-Resolution and Efficient Automotive Radar Imaging through Domain-Informed 1D Deep Learning
  243. Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset
  244. Advancing Non-intrusive Suppression on Enhancement Distortion for Noise Robust ASR
  245. Advancing Paired Image-Mask Synthesis for Automated Nanoparticle Phenotyping
  246. Advancing SAR Image Robustness: Integrating Diffusion Models for Adversarial Purification and Speckle Noise Suppression
  247. Advancing Single-Snapshot DOA Estimation with Siamese Neural Networks for Sparse Linear Arrays
  248. Advancing Streaming ASR with Chunk-wise Attention and Trans-chunk Selective State Spaces
  249. Adversarial Feature Disentanglement Framework for Voice Pathology Detection
  250. Adversarial Knowledge Transfer for Black-Box Model Inversion Attack
  251. Adversarial Learning For End-To-End Cochlear Speech Denoising Using Lightweight Deep Learning Models
  252. Adversarial Speech-Text Pre-Training for Speech Translation
  253. Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment Analysis
  254. Adversarial Training and Gradient Optimization for Partially Deepfake Audio Localization
  255. Aesthetic Perception Prompting for Interpretable Image Aesthetics Assessment with MLLMs
  256. Age of Gossip with the Push-Pull Protocol
  257. AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation
  258. Agentic Copyright Watermarking against Adversarial Evidence Forgery with Purification-Agnostic Curriculum Proxy Learning
  259. Algorithm Design for Continual Learning in IoT Networks
  260. Aligned Contrastive Learning for Text-to-Music Retrieval
  261. Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations
  262. Aligning Text-to-Image Diffusion Models without Human Feedback
  263. Alignment-Free Training for Transducer-based Multi-Talker ASR
  264. Ambisonics Binaural Rendering via Masked Magnitude Least Squares
  265. Ambisonics Coding in IVAS: A Hybrid SPAR and DirAC System
  266. Amplitude-Guidance Low-Light Image Enhancement with Frequency-based Channel Attention
  267. An Abnormal Audio Generation Method for Fault Diagnosis of Power Transformers
  268. An Adaptive Framework for Multi-View Clustering Leveraging Conditional Entropy Optimization
  269. An Adversarial Perturbation Generation Method for Image Anti-Forensics Based on Dual-Path Spatial Attention GAN
  270. An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS Diagnosis
  271. An Attribute-Enriched Dataset and Auto-Annotated Pipeline for Open Detection
  272. An Automatic Extrinsic Calibration Method for LiDAR-Camera Fusion via Combining Semantic and Geometric Features
  273. An Efficient Hybrid Quantum Variational Classifier With Matrix Product State
  274. An Efficient Pore Annotation Framework for Tight Sandstone Images with Segment Anything Model
  275. An Efficient Residual-based Low-dose PET Reconstruction with Spatial-Frequency Integration
  276. An Efficient Sample Utilization Method for Deep Learning Based on Class Uncertainty
  277. An Efficient and Streaming Audio Visual Active Speaker Detection System
  278. An End-to-End Graph-Guided Spatiotemporal Model for Adaptive Frame-Level Facial Affect Analysis in the Wild
  279. An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM
  280. An Exceptional Dataset For Rare Pancreatic Tumor Segmentation
  281. An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation
  282. An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
  283. An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
  284. An Improved Planar Approximation Localization Method in Distributed Airborne Radars
  285. An Information-Theoretic Analysis of Thompson Sampling with Infinite Action Spaces
  286. An Interactive Evaluation Framework for Empathetic Response Generation
  287. An Intra- and Cross-frame Topological Consistency Scheme for Semi-supervised Atherosclerotic Coronary Plaque Segmentation
  288. An LSTM Feature Imitation Network for Hand Movement Recognition from sEMG Signals
  289. An Optimized GPU-based Acceleration of CRYSTALS-Dilithium
  290. An Underwater Image Quality Dataset with Renewed Pairwise Voting
  291. AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
  292. Analysis and Calibration of Nonlinear Power Amplifiers in Wideband OFDM-Based LEO Satellite Communication System
  293. Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
  294. Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
  295. Anchor-Prompt-based Segmentation and Embedding Model
  296. Anchored Monotonic Alignment and Representation Substitution for Rare Spontaneous Behaviors in Spontaneous Speech Synthesis
  297. Anima2: Cross-Species Animal Animation through Image-to-Video Synthesis with Subject Alignment
  298. AnimateSketches: Animate Sketches with Instance-Aware Mask
  299. Animation Anycolor: Enhancing Line Drawing Colorization with Keypoint Matching
  300. Annealing Distillation Algorithm for Transferring Unsupervised Clustering Knowledge to Supervised Student Models
  301. Apollo: Band-sequence Modeling for High-Quality Audio Restoration
  302. Appearance- and Orientation-aware Fine-grained Rotated Ship Detection in High-Resolution Satellite Imagery
  303. Appearance-adapter: A Self-supervised Pose-guided Human Image Synthesis Approach
  304. Approximation and Analysis of the One-Bit Hermite Law
  305. Archetypal Analysis for Binary Data
  306. ArtTwin: A Novel Concept of Developing Digital Twin of Human Arterial System
  307. Artistic Image Aesthetics Assessment Assisted by Photographic Visual Attributes
  308. Assessing Robustness of Multi-Modal Large Language Models in Image Classification through Hierarchical WordNet-Based Evaluation
  309. Asymptotic Behavior Analysis of Antenna Selection via Sparsity-Induced Precoder
  310. Atom-Constrained Maximum Likelihood Gridless DOA with Wirtinger Gradients
  311. Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
  312. Attention Augmented Structure-centric Bias Mitigation with Feature Disentanglement
  313. Attention Disentanglement for Semantic Diffusion Modeling in Text-to-Image Generation
  314. Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio Codecs
  315. Attention-Based Beamformer For Multi-Channel Speech Enhancement
  316. Attention-Driven Causal Discovery: From Transformer Matrices to Granger Causal Graphs for Non-Stationary Time-series Data
  317. Attention-Enhanced Feature Fusion Network for No-Reference Image Quality Assessment
  318. Attention-Enhanced Short-Time Wiener Solution for Acoustic Echo Cancellation
  319. Attribute Conditional Diffusion-Augmented Person Re-Identification
  320. Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling
  321. Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
  322. Audio Decoding by Inverse Problem Solving
  323. Audio Diffusion with Large Language Models
  324. Audio Explanation Synthesis with Generative Foundation Models
  325. Audio Features Investigation for Singing Voice Deepfake Detection
  326. Audio Sparse-Transformer for Speech Classification
  327. Audio Texture Manipulation by Exemplar-Based Analogy
  328. Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
  329. Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker Detection
  330. Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
  331. Audio-Visual Representation Learning For Lip-Sync Estimation Through Ranking Augmented Contrastive Training
  332. AudioBERT: Audio Knowledge Augmented Language Model
  333. AudioCache: Accelerate Audio Generation With Training-Free Layer Caching
  334. AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
  335. AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
  336. AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
  337. Audiogram-Informed End-to-End Noise Reduction and Wide Dynamic Range Compression for Hearing Aids
  338. Audiopedia: Audio QA with Knowledge
  339. Augmenting Short Enrollment Speech via Synthesis for Target Speaker Extraction
  340. AuscMLLM: Bridging Classification and Reasoning in Heart Sound Analysis with a Multimodal Large Language Model
  341. Automated Exposure Mapping for Networked Interference
  342. Automated Extraction of Spatio-Semantic Graphs for Identifying Cognitive Impairment
  343. Automated Graph Attention Network for Heterogeneous Entity Resolution
  344. Automatic Adaption of the Step Size in Gradient Descent Training
  345. Automatic Detection of Domain Shifts in Speech Enhancement Systems Using Confidence-Based Metrics
  346. Automatic Geometric Quantification and Rupture Risk Evaluation of 3D Intracranial Aneurysms
  347. Automatic Labelling & Semantic Segmentation with 4D Radar Tensors
  348. Automatic Numbering and Pathological Recognition of Pediatric Teeth Using CNN and Attention Mechanisms
  349. Automatic Parkinson's disease detection from speech: Layer selection vs adaptation of foundation models
  350. Automatic Speech Recognition and Spoken Language Understanding of Maritime Radio Communications: A case study with Singapore data
  351. Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing
  352. Automatic recognition of rodent call types using deep supervectors
  353. Automotive Radar Target Detection in Widely Separated and Distributed Aperture Radar Systems
  354. Autoregressive Density Estimation Transformers for Multivariate Time Series Anomaly Detection
  355. Autoregressive Language Model with Historical Context Re-encoding
  356. Auxiliary Tasks Benefit Skeleton-based Action Recognition
  357. Avoiding Domain Drift and Constant Predictions with Diffusion Enhanced Vector-Quantized Autoencoders for Temperature Predictions
  358. BAD: Bidirectional Auto-Regressive Diffusion for Text-to-Motion Generation
  359. BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
  360. BCG data imputation via multimodal feature alignment and semantic sequence prediction
  361. BCS-Net: Multi-Task Breast Cancer Screening Network Enhanced by Multi-Modality Attention
  362. BDCKD: Unlocking the Power of Brownian Distance Covariance in Knowledge Distillation
  363. BDGAN: Boundary and Diversity-aware Generative Adversarial Network for Imbalanced Medical Image Augmentation
  364. BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
  365. BIAWDiff: Enhancing Low-Light Images with Bio-Inspired Attention and Wavelet Diffusion
  366. BID-Net: Balanced Incremental Distillation Network for Fair Dermatological Disease Diagnosis
  367. BIF: A Biosignature Identification Framework for Model-agnostic Interpretation of MVI Diagnosis Models in HCC
  368. BIGFR: Bridging Individual and Group Fairness in Recommendation Systems
  369. BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
  370. BP-GPT: Auditory Neural Decoding Using fMRI-prompted LLM
  371. BRDIA: Bidirectional Reasoning with Dynamic Instruction Adjustment for Multi-hop KGQA
  372. BS-Breath: Respiration Sensing with Cell-free Massive MIMO
  373. BadRefSR: Backdoor Attacks Against Reference-based Image Super Resolution
  374. Band Prompting Aided SAR and Multi-Spectral Data Fusion Framework for Local Climate Zone Classification
  375. Basis Function Learning for Variable-Length and Continuous-Indexed Signals
  376. Basket-Enhanced Heterogenous Hypergraph for Price-Sensitive Next Basket Recommendation
  377. Bayesian Filtering on Graphs
  378. Bayesian Nonparametric Clustering for Source Counting with a Small Aperture Microphone Array
  379. BeatKAN: An Efficient and Drum-Attuned Beat Tracking Method Using Kolmogorov-Arnold Networks
  380. Benchmarking Music Generation Models and Metrics via Human Preference Studies
  381. Bernoulli-Gaussian Scale Mixture Model and BP Method for Multi-Snapshot Sparse Signal Recovery
  382. Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter Network
  383. Beyond Jensen's Inequality: Speeding Up ML Estimation of Generalized Hyperbolic Distributions
  384. Beyond Point Annotation: A Weakly Supervised Network Guided by Multi-Level Labels Generated from Four-Point Annotation for Thyroid Nodule Segmentation in Ultrasound Image
  385. Beyond Speaker Identity: Text Guided Target Speech Extraction
  386. Beyond Uniformity: Deblurring Images With Complex Noise Patterns Using Half Quadratic Splitting
  387. Bi-attention pyramid network for small defect with complex background in industrial detection
  388. BiCG: Binaural Cue Generation from Unified HRTF Datasets
  389. BiMA: Bidimensional multi-level attention embedded network for single-frame infrared small target detection
  390. Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling
  391. Big-Moe: Bypassing Isolated Gating For Generalized Multimodal Face Anti-Spoofing
  392. Bilevel Learning for Low-Light Image Enhancement and Detection
  393. Bilingual Dual-Head Deep Model for Parkinson's Disease Detection from Speech
  394. Binary Representation Learning for Discriminative Acoustic Unit Discovery
  395. Binary Stochastic Flip Optimization for Training Binary Neural Networks
  396. Biodenoising: Animal Vocalization Denoising without Access to Clean Data
  397. Birds of a Feather: Learning to Retrieve Dance Poses From Music Via Ground-Truth Annotation Lifting
  398. Black-Box Adversarial Defense Against Voice Conversion Using Latent Space Perturbation
  399. Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features
  400. Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
  401. BloomCoreset: Fast Coreset Sampling using Bloom Filters for Fine-Grained Self-Supervised Learning
  402. BlurPaint: Image Inpainting using Blurring Diffusion Models
  403. Boli: A dataset for understanding stuttering experience and analyzing stuttered speech
  404. Bone Conducted Signal Guided Speech Enhancement For Voice Assistant on Earbuds
  405. Boolean Matrix Tri-Factorization
  406. Boolean matrix compressed sensing
  407. Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
  408. Boosting Jailbreak Attack with Momentum
  409. Boosting Large Language Model for Speech Synthesis: An Empirical Study
  410. Boosting Lightweight Camouflaged Object Detection with Multi-Scale Context and Boundary Awareness
  411. Boosting Movie and TV Tag Accuracy with Knowledge Graphs
  412. Boosting Open-Vocabulary Object Detection Performance via Class-Agnostic Pseudo-Labels and MultiModal Hybrid Knowledge
  413. Boosting Stereo Image Noise Removal by Learning Uncertainty and Enriched Features
  414. Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
  415. Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size
  416. Bootstrapping LLM-based Fact-checking via Iterative Rationalization Finetuning
  417. Bootstrapping Language-Audio Pre-training for Music Captioning
  418. Bottleneck-Constrained Contrastive Decoupled Network for Multimodal Aspect-based Sentiment Classification
  419. Boundary-Driven Table-Filling with Cross-Granularity Contrastive Learning for Aspect Sentiment Triplet Extraction
  420. Brain MRI Segmentation with Language-Driven Detection and Context-Aware Descriptions
  421. BrainChat: Interactive Semantic Information Decoding from fMRI Using Large-Scale Vision-Language Pretrained Models
  422. BrainVis: Exploring the Bridge between Brain and Visual Signals via Image Reconstruction
  423. Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
  424. Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
  425. Bridge-SR: Schrödinger Bridge for Efficient SR
  426. Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation
  427. Bridging Neural and Symbolic Reasoning: A Dual-System Framework for Interpretable Question Answering
  428. Bridging Speech and Text Foundation Models with ReShape Attention
  429. Bridging Task Boundaries: Remote Sensing Image-Text Retrieval via Dictionary-Driven Adaptation
  430. Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences
  431. Bridging the Modality Gap for Speech-image Retrieval with Text Supervision
  432. Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice
  433. C2AD: Dual Consistency Learning for Zero-Shot Anomaly Detection
  434. C3D-VIT: Consistency-Aware 3D Vision Transformer for Face Forgery Detection
  435. CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
  436. CA-UAP: Content-Agnostic Universal Adversarial Perturbation for Enhanced Generalization
  437. CAAL-Unet: a Confusion Area Attention Lightweight Network for Brain Tumor Segmentation
  438. CAF-YOLO: A Robust Framework for Multi-Scale Lesion Detection in Biomedical Imagery
  439. CAMDet: Condition-Adaptive Multispectral Object Detection Using a Visible-Thermal Translation Model
  440. CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
  441. CAPAST: Content Affinity Preserved Arbitrary Style Transfer
  442. CASC-XVC: Zero-Shot Cross-Lingual Voice Conversion with Content Accordant and Speaker Contrastive Losses
  443. CASleepNet: A Cross Attention-based multimodal fusion approach for sleep staging with EEG and EOG
  444. CAT-Net: A Co-Adaptive Transfer Learning Network for BCI-Assisted Neurorehabilitation
  445. CAW-CL: Cascaded Adaptive Weighted Contrastive Learning for Unsupervised Ultrasound Plane-Wave Image Reconstruction
  446. CE-FFT: Communication-Efficient Federated Fine-Tuning for Large Language Models via Quantization and In-Context Learning
  447. CEMSSL: Conditional Embodied Self-Supervised Learning is All You Need for High-precision Multi-solution Inverse Kinematics of Robot Arms
  448. CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion
  449. CGDD: Contrastive Gaussian-Dirac Diffusion Model
  450. CGEDN: Approximation of Graph Edit Distance with Path Generation via Learning Node Matching
  451. CGNet: Classification-Guided Multi-Task Interactive Network for Hyperspectral and Multispectral Image Fusion
  452. CHASE: Channel-Wise and Spatial Attention for Early Exiting in Image Classification
  453. CIEGCL: Counterfactual Intervention Enhancing Graph Contrastive Learning in Implicit Feedback
  454. CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
  455. CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition
  456. CLHi-MTS: A Contrastive Learning-Based Hierarchical Framework for Masked Medical Time-Series Modeling
  457. CLIPGaze: Zero-Shot Goal-Directed Scanpath Prediction Using CLIP
  458. CMFNThinker: A Novel Cross-source Multi-modal Fake News Detection Model
  459. CMGait: Enhancing Cross-Modality Gait Recognition between LiDAR and RGB through Contrastive Identity-consistent Feature Aggregation
  460. CMoS: Customizing Model Structures for Personalized Federated Learning
  461. COAST: Contrastive Learning with Augmented Spatio-Temporal Encoding for Next POI Recommendation
  462. COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
  463. COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
  464. COREMIL: Contextual Position Encoding-based Retrievable Multiple Instance Learning for Slide-level Classification
  465. COSMIC waveforms for Integrated Communication and Imaging
  466. CPA-Enhancer: Chain-of-Thought Prompted Adaptive Enhancer for Downstream Vision Tasks Under Unknown Degradations
  467. CPL: Curriculum Pseudo Labeling for Weakly Supervised Temporal Forgery Localization
  468. CPSNet: Comprehensive Enhancement Representation for Polyp Segmentation Task
  469. CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
  470. CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze Estimation
  471. CSD: Weather forecasting with graph neural network based on cross-scale diffusivity
  472. CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity Classification
  473. CSS: Overcoming Pose and Scene Challenges in Crowd-Sourced 3D Gaussian Splatting
  474. CT Image Prediction Of PD-1 Gastric Cancer Patients Based On The PLSG Framework
  475. CTGDiff: A Conditional Diffusion Model for Cardiotocography Signal Synthesis
  476. CabiNet: A Deep Learning Framework for Multiclass Medical Image Segmentation from Multiple Single Class Datasets
  477. Calibration of Multiple Asynchronous Microphone Arrays using Hybrid TDOA
  478. Camouflaged Object Detection via Neural Architecture Search
  479. Camouflaged Object Detection with CNN-Transformer Harmonization and Calibration
  480. Can AI See What We Can't? Leveraging Deep Learning and Multi-Temporal Satellite Data to Revolutionize Crop Type Mapping and Yield Prediction
  481. Can Automated Speech Recognition Errors Provide Valuable Clues for Alzheimer's Disease Detection?
  482. Can Fairness and Robustness Be Simultaneously Achieved Under Byzantine Attacks?
  483. Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
  484. Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
  485. Can Quality Survive Scale? Toward an Equal-Quality Instance-Dependent Label Noise Model
  486. Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?
  487. Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis Model
  488. CapT: A Hierarchical Capsule Representation Learning Approach for Class Continual Learning
  489. Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning
  490. CardioFlow: Learning to Generate ECG from PPG with Rectified Flow
  491. CardioRiskNet: Attention-based CVAE-enabled GCN for Risk Prediction in STEMI
  492. Carver: Learning to Reconstruct Right Ventricle from Sparse Multi-View 2D Echocardiograms
  493. CascadePAIE: Reallocating Relevance for Event Roles and Event Text in Event Argument Extraction
  494. Catch Causal Signals from Edges for Label Imbalance in Graph Classification
  495. Cauchy-Schwarz Divergence Transfer Entropy
  496. Causal Debiasing for Visual Commonsense Reasoning
  497. Causal Feature Supervision Decoupling: A Novel Method for Clothes-Changing Person Re-identification Algorithm
  498. Causal Speech Enhancement Based on a Two-Branch Nested U-Net Architecture Using Self-Supervised Speech Embeddings
  499. Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
  500. Causal fMRI-Mamba: Causal State Space Model for Neural Decoding and Brain Task States Recognition
  501. Causality-Guided Context-Aware Multimodal Public Speaking Anxiety Detection for Out-of-Distribution Generalization
  502. Certainty-guided Reasoning and Refinement Network for Camouflaged Object Detection
  503. Chain-of-Thought Prompting for Speech Translation
  504. Chained Motion Vector Prediction for Video Coding
  505. ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning
  506. Channel and space-based joint rate allocation algorithm
  507. Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
  508. ChannelMixer: A Hybrid CNN-Transformer Framework for Enhanced Multivariate Long-Term Time Series Forecasting
  509. Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts
  510. Chat-Driven 3D Human Pose and Shape Editing with Large Language Models
  511. ChatCAD: An MLLM-Guided Framework for Zero-shot CAD Drawing Restoration
  512. CheapNVS: Real-Time On-Device Narrow-Baseline Novel View Synthesis
  513. Chinese Speech Processing via Chinese Character Feature
  514. Chitrarth: Bridging Vision and Language for a Billion People
  515. ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
  516. CiGA: A Cross-Layer Fine-Grained Attention Correction Method for Large Language Model
  517. Class Relevance Learning for Out-of-Distribution Detection
  518. Class Semantic Prompts Enhanced Prototypical Fusion Method for Few-shot Named Entity Recognition
  519. Class-Difficulty Aware Hybrid Active Learning
  520. Class-wise Adaptive Logits Distillation with Meta-Learning
  521. Classification Error Bound for Low Bayes Error Conditions in Machine Learning
  522. Classification Inconsistency Alignment Network for Cross-corpus Speech Emotion Recognition
  523. Classification of Eye-Tracking Data Based on Spatiotemporal Attention Encoding
  524. Classification of Zhuang Dialect combined with Bert and SimAM
  525. Classifier-Guided Captioning Across Modalities
  526. Classifying Music-Induced Emotion Using Multi-Modal Ensembles of EEG and Audio Feature Models
  527. Climate Downscaling Using Neural Operator: Spatiotemporal Multimodal Fusion Operator with State-Query Coupled Kernel
  528. ClingTP: Curriculum Learning based Multi-style Title Prefix Generation
  529. Clinically Robust Polyp Segmentation: Enhanced Generalization and Perturbation Resistance
  530. Cloth-debiasing with Stable Diffusion in Cloth-changing Person Re-identification
  531. Cluster-Perceptive Graph Contrastive Learning for Community Detection
  532. Cluster-Refined Optimal Transport for Unsupervised Action Segmentation
  533. Clutter Resilient Occlusion Avoidance for Tightly-Coupled Motion-Assisted Detection
  534. Co-Attention Based Multi-Channel TF-GridNet for Speech Separation with Ad-Hoc Microphone Arrays
  535. Co-training with Progressive Distribution Alignment and Uncertainty-Interactive Relabeling for Semi-Supervised Domain Adaptive Semantic Segmentation
  536. CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
  537. CoGAP: A Personalized Federated Learning Method Using Collaborative Optimization for Medical Image Classification
  538. CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
  539. CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
  540. Coarse-to-Fine Text-to-Music Latent Diffusion
  541. Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFi
  542. Code Drift: Towards Idempotent Neural Audio Codecs
  543. Codec-ASV: Exploring Neural Audio Codec For Speaker Representation Learning
  544. CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks
  545. Cognitive Decline Detection using DLB Extraction Pipelines
  546. Cognitive Load Monitoring via Earable Acoustic Sensing
  547. Cognitive MIMO Radar Beamforming for Target Tracking Using a BCRB-based Criterion
  548. Cohort-Sensitive Labeling: An Effective Approach for Enhancing ASR Performance
  549. Col-OLHTR: A Novel Framework for Multimodal Online Handwritten Text Recognition
  550. Collaborative Association Network for Multi-view Multi-Human Association and Tracking using Constraint Optimization and Object Search
  551. Collaborative Automotive Radar Sensing via Mixed-Precision Distributed Array Completion
  552. Collaborative Dual-Branch Spatial-Frequency Enhancement Network for Low-Light Images
  553. Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
  554. Collaborative Personalized Federated Learning via Exponential Moving Average Optimization
  555. Collaborative Semantics-Assisted Large Language Models for Next POI Recommendation
  556. Collision-less and Balanced Sampling for Language-Queried Audio Source Separation
  557. Collusion-resistant Black-box Watermarking in Federated Learning through Weight Relevance Analysis
  558. Colored Point Cloud-based Mesh Registration for Enhancing Inter-frame Coding of Texture Video in V-DMC
  559. Colorization Network Watermarking in the CIE-Lab Domain
  560. Combined object-based audio and MASA format for enhanced spatial mobile communication
  561. Combining Loss-aware Curriculum Learning with Incomplete Graph Neural Networks
  562. Combining Spatio-Temporal Networks and Graph Attention Architectures for EEG-Based Workload Classification
  563. Commonality Augmented Disentanglement for Multimodal Crowdfunding Success Prediction
  564. Communication-efficient Exact Diffusion for Decentralized Learning
  565. Communication-efficient Verifiable and Oblivious Aggregation with Client Dropouts
  566. Community-entropy Based Graph Structure Learning for Topology-imbalance
  567. CompMTL: Layer-Wise Competitive Multi-Task Learning
  568. Compact Neural TTS Voices for Accessibility
  569. Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
  570. Compgen: Synthesis and Generation of Faces From Edgemaps
  571. Complementary Graph Learning and Prompt-based Cross-modal Generation for Missing-modality Fake News Detection
  572. Complementary Learning System Theory-based Active Learning for Audio Classification
  573. Complete Reconstruction of the Tongue Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI Data
  574. Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
  575. Complex Coprime Frequency Sum Based Signal Representation for Period Estimation
  576. Complex Open Information Extraction with Heterogeneous Syntax Forests
  577. ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling
  578. Component-wise Self-Correction Network for Human Motion Prediction
  579. Compositional Audio Representation Learning
  580. Comprehensive Feature Processing Based on Attention Mechanism for Co-Salient Object Detection
  581. Comprehensive Perturbation Consistency for Semi-Supervised Change Detection in Remote Sensing Images
  582. Compressing a Flow-Based Privacy Protection Model via a Novel Joint Distilling and Pruning Method
  583. Compressive Imaging Reconstruction via Conditional Diffusion Model With Augmented Measurements
  584. ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
  585. ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
  586. ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting
  587. Concentrating Harder for Faster Audio Transformer
  588. Conditional Convolutions for End-to-End Single-Stage Video Text Detection
  589. Conditional Deep Canonical Time Warping
  590. Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning
  591. Conditional-Balanced Adversarial Delta Tuning for Cross-Domain Implicit Discourse Relation Recognition
  592. Conformal Prediction for Manifold-based Source Localization with Gaussian Processes
  593. Confusion-Aware Prototypical Contrastive Learning for Open-Vocabulary Object Detection
  594. Consensus Graph Filter Learning for Multiple Graph Clustering
  595. Consensus Graph-Based Spectral Ensemble Clustering via Low-Rank Tensor Learning
  596. Conservative Offline Meta-Reinforcement Learning with Task Similarity Measurement
  597. Constraint-Awareness and Graph Reasoning for Temporal Question Answering
  598. Constructing Datasets From Public Police Body Camera Footage
  599. Contactless Nighttime Stress Monitoring with mmWave Radar
  600. Contactless Vital Sign Monitoring for Multiple People Using a Millimeter-wave MIMO Radar
  601. Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification
  602. Content-Aware Dynamic Superpixel Segmentation
  603. Context-Aware Multi-Scale Polyp Segmentation Network
  604. Context-Guided Active Domain Adaptation for Blended Target Domain
  605. Contextual ASR with Retrieval Augmented Large Language Model
  606. Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
  607. Contextual Value Alignment
  608. Contextualization of ASR with LLM using phonetic retrieval-based augmentation
  609. Continual Self-supervised Learning Considering Medical Domain Knowledge in Chest CT Images
  610. Continual Unsupervised Domain Adaptation for Audio Deepfake Detection
  611. Continuous-Discrete Differentiable Particle Filters for Irregular Time Series
  612. Continuously Learning New Words in Automatic Speech Recognition
  613. Continuously Learning Video-level Object Tokens for Robust UAV tracking
  614. Contrast Memory for Unsupervised Anomaly Detection
  615. Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
  616. Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
  617. Contrastive Learning via Randomly Generated Deep Supervision
  618. Contrastive Lyrics Alignment with a Timestamp-Informed Loss
  619. Contrastive Pre-Training and Post-Tuning for Heterogeneous Graph Learning
  620. ControlMol: Adding Substructure Control To Molecule Diffusion Models
  621. Controllable Forgetting Mechanism for Few-Shot Class-Incremental Learning
  622. Controllable Generative Model for Brain Evolution
  623. Controlling the Number of Sample-Contributive Vertices in Generalized Sampling of Graph Signals
  624. Convergence Analysis of alpha-SVRG under Strong Convexity
  625. ConvexECG: Lightweight and Explainable Neural Networks for Personalized, Continuous Cardiac Monitoring
  626. Convolutional Retentive Network for EEG Decoding
  627. Convolutional Sparse Coding with Multipath Orthogonal Matching Pursuit
  628. Cooperative ISAC for Localization and Velocity Estimation Using OFDM Waveforms in Cell-Free MIMO Systems
  629. Cooperative Multi-Target Tracking Based on Multi-Detection TPHD in MIMO-OFDM Systems
  630. Cooperative Neural Radiance Field for Dynamic Scene Deblurring
  631. Cooperative and Competitive Functional Connectivity Based on Improved Ising Model
  632. CorrGAN: Simultaneous Learning of Speech Enhancement and Perceptual Quality Loss Functions
  633. Correlated Attention in Transformers for Multivariate Time Series
  634. Correlated Multiple IHC Virtual Staining for Breast Histopathological Images
  635. Correlative3D: Inter-Object Correlation-Aware 3D Scene Understanding
  636. Covariance Change Point Detection for Graph Signals
  637. Covert and Potent: A Weather-Camouflaged Backdoor Attacks on Self-Supervised Learning
  638. Cramér-Rao Bounds for Wideband Near-Field Sensing
  639. Credible and Detailed 3D Face Reconstruction in Large Pose
  640. CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
  641. Critically-Damped Third-Order Langevin Dynamics
  642. CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language Understanding
  643. Cross-Channel Unlabeled Sensing over a Union of Signal Subspaces
  644. Cross-Component Residual Prediction for Geometry-Based Point Cloud Compression
  645. Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype Reprojection
  646. Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image Recognition
  647. Cross-Layer Graph Knowledge Distillation for Image Recognition
  648. Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
  649. Cross-Modality Fusion Mamba for All-in-One Extreme Weather-Degraded Image Restoration
  650. Cross-Talk Detection in the IVAS Stereo Codec Based on GCC-PHAT
  651. Cross-Template-Based Hypergraph Transformer
  652. Cross-attention Inspired Selective State Space Models for Target Sound Extraction
  653. Cross-lingual Evaluation Of Hypernasality Using Wav2Vec2 Features
  654. Cross-modal Gaussian Localization Distillation for Optical Information guided SAR Object Detection
  655. CrossHash: Cross-scale Vision Transformer Hashing for Image Retrieval
  656. CrossSleep: Multi-Scale Attention with Cross-Time Learning for Single Channel EEG-Based Sleep Staging
  657. Crowdsourced Homophily Ties Based Graph Annotation Via Large Language Model
  658. Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
  659. CurMIM: Curriculum Masked Image Modeling
  660. Curriculum Contrastive Learning for Aspect-based Sentiment Analysis
  661. Curriculum Learning aided Audio-Visual Speech Recognition with Arbitrary Speaker Number
  662. CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation
  663. D2-MLP: Dynamic Decomposed MLP Mixer for Medical Image Segmentation
  664. D2S: Towards Efficient Sparse 3D Object Detection via Dense to Sparse Knowledge Distillation
  665. D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription
  666. DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks
  667. DACAT: Dual-stream Adaptive Clip-aware Time Modeling for Robust Online Surgical Phase Recognition
  668. DAEF-VS: An Efficient Universal VoIP Steganalysis Framework Based on Domain-Aware Knowledge
  669. DAREK - Distance Aware Error for Kolmogorov Networks
  670. DARN: An Attention-Based Neural Network Using Residual Blocks for Sleep Micro-Events Detection
  671. DARNet: A Dual Attention Residual Network for Medical Image Classification
  672. DASSL: Domain Agnostic Self-Supervised Learning with Multiple Missing Information Reconstruction Branches
  673. DATA-VSR: Dynamic Trajectory Attention and Texture Adaptive Rooter for Video Super-Resolution
  674. DBCR: Exploiting Both Intra-cluster and Extra-cluster Relations for Compositional Reasoning
  675. DCASI: A Sequence-based Attack Investigation Method Using DTW Contrastive Learning
  676. DCCMamba: A Dual-stream Cross-time and Cross-feature with Mamba for Multivariate Time Series Forecasting
  677. DCCT-Net: A Network Combined Dynamic CNN and Transformer for Image Compressive Sensing
  678. DCD-MUSIC: Deep-Learning-Aided Cascaded Differentiable MUSIC Algorithm for Near-Field Localization of Multiple Sources
  679. DCFormer: Divide-and-Conquer in 3D Human Pose Estimation Tasks
  680. DCIM-AVSR: Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
  681. DDA: Distillation-Driven Acceleration of the Reverse Diffusion Process for Stochastic Multi-Ship Trajectory Prediction
  682. DDNet: Deformable Convolution and Dense FPN for Surface Defect Detection in Recycled Books
  683. DDNet: Exploring Dual Dependencies for Long-Term Time Series Forecasting
  684. DDSP Guitar Amp: Interpretable Guitar Amplifier Modeling
  685. DEBT: Enhancing Entity Alignment in Knowledge Graphs through Description Enrichment and Bootstrap Training
  686. DEFormer: DCT-driven Enhancement Transformer for Low-light Image and Dark Vision
  687. DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis
  688. DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models
  689. DETCP: Self-Detoxifying Language Models With Contrastive Pairs
  690. DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
  691. DFMA: Adaptive Dual Fusion for Multimodal Relation Extraction with Mutual Attention
  692. DFNeRF: Disentangled Facial Neural Radiance Fields for Text-based Editing of Free-view Talking Head
  693. DFT-Spread-Based OTFS Waveform Design With Good Peak-to-Average Power Ratio for Joint Sensing and Communications
  694. DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
  695. DGJA: Dependency Graph-enhanced Joint Attention Structure for Multimodal Sarcasm Detection
  696. DH-VTON: Deep Text-Driven Virtual Try-On via Hybrid Attention Learning
  697. DICS: Find Domain-Invariant and Class-Specific Features for Out-of-Distribution Generalization
  698. DKD2L: Dual Knowledge Distillation Dynamic Learning for sketch-based 3D shape retrieval
  699. DLM-VMTL: A Double LayerMapper For Heterogeneous Data Video Multi-Task Prompt Learning
  700. DMIBot: Dynamic Multimodal Interaction for Twitter Bot Detection
  701. DMKPN: Image Deblurring Under Multi-Factor Aliasing Diffusion Degradation
  702. DN-DR: Discriminative Network with Dual Reconstruction for Image Anomaly Detection
  703. DOA Estimation Based on Enhanced SRP-MVDR Using Kronecker Product Decomposition for Large Rectangular Microphone Arrays
  704. DOA Estimation of Coherent Sources Using Residual Network-based Subspace Reconstruction
  705. DOSE: Drum One-Shot Extraction from Music Mixture
  706. DPC: Large Model Alignment Method based on Decoding Probability Correction
  707. DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
  708. DPM-LVSN: A Diffusion Probabilistic Model-based Left Ventricular Segmentation Network
  709. DRANet: Dual-threshold Guided Reliability Aware Network for Semi-Supervised Image Semantic Segmentation
  710. DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
  711. DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images
  712. DRSFANet: Dual-Path CNN with Residual and Frequency Attention for Image Denoising
  713. DS-BTIAN: A Novel Deep-Shallow Bidirectional Transformer Interactive Attention Network for Multimodal Emotion Recognition
  714. DSDIR: A Two-Stage Method for Addressing Noisy Long-Tailed Problems in Malicious Traffic Detection
  715. DSDN-Net: An Effective Network for Semantic Segmentation in Open-Pit Coal Mining Areas for Land Cover Recognition
  716. DSFormer: Deformable Pointformer for 3D Salient Object Detection
  717. DSINet: Towards Real-Time Target Speaker Extraction with Dynamic Speaker Information Fusion
  718. DSSM: Dual State Space Model For Human Motions Generation
  719. DTR: Dynamic Tree-Ring Watermarking Framework for Diffusion-Based Video Generation
  720. DU-PMVS: Learned Patchmatch Multi-View Stereo Based on Deformable Feature Pyramid and Uncertainty Awareness Modeling
  721. DULRTC-RME: A Deep Unrolled Low-rank Tensor Completion Network for Radio Map Estimation
  722. DUNE: Sim2Real Transfer for Depth-based Navigation in Unstructured Dynamic Indoor Environments
  723. DVM: Towards Controllable LLM Agents in Social Deduction Games
  724. DX2CT: Diffusion Model for 3D CT Reconstruction from Bi or Mono-planar 2D X-ray(s)
  725. DapPep: Domain Adaptive Peptide-agnostic Learning for Universal T-cell Receptor-antigen Binding Affinity Prediction
  726. Dark Experience for Incremental Keyword Spotting
  727. Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
  728. Data Glove-based Personalized Continuous Gesture Segmentation
  729. Data-Aided Regularization of Direct-Estimate Combiner in Distributed MIMO Systems
  730. Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
  731. Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech Enhancement
  732. Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
  733. Data-Free Post-Training Quantization with Block-wise Enhanced Sample Generation
  734. Data-driven Processing using Parametric Neural Network for Improved Bluetooth Channel Sounding Distance Estimation
  735. De-confusing Hard Samples for Text Semantic Hashing
  736. DeBeauty: A Joint Framework for Facial Beautification Removal Based on Spatial Collaborative Adaptation and Hyperplane Relocation
  737. DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
  738. Debiased Estimation for Cross-Domain Cold Start Recommendation
  739. Debiased Prototype Evolving for Point Cloud Domain Adaptation via 3D Foundation Models
  740. Debiased Training For Semi-supervised Sound Event Detection
  741. Decentralized Federated Dataset Dictionary Learning for Multi-Source Domain Adaptation
  742. Decentralized Online Ensembles of Gaussian Processes for Multi-Agent Systems
  743. Decentralized Stochastic Successive Convex Approximation for composite non-convex problems with non-linear functional constraints
  744. Decision-Aided Progressive Symbol Phase Equalizer in Sweep Spread Carrier Underwater Acoustic Communications
  745. Decoding Brain Structure and Gene Expression Interactions in Alzheimer's Disease Pathology
  746. Decoding the Unintelligible: Neural Speech Tracking in Low Signal-to-Noise Ratios
  747. Decoupled Feature Matching for Few-shot Counting and Localization
  748. DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors Decoupling
  749. Decoupling While Coupling: Towards More Accurate Stereo Image Sand Removal Beyond Certainty
  750. Decreasing Word Error Rates in Paragraph Handwritten Text Recognition with Synthetic Data
  751. Deep Diffusion Gradients Leakage in Federated Learning
  752. Deep Dynamic Probabilistic Canonical Correlation Analysis
  753. Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments
  754. Deep Feedback Cancellation for Hearing Aids with Improved System Stability and Sound Quality
  755. Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
  756. Deep Joint Source-Channel Coding for Wireless Point Cloud Transmission
  757. Deep Learning Amplified Early Stopping Bias: Overestimating Performance on Small Datasets
  758. Deep Learning for Modulo Sampling of FRI Signals
  759. Deep Learning-Based Perceptual Vibrotactile Codec with Rate Scalability
  760. Deep Metamorphic Registration for Tumor-Affected Medical Image Alignment
  761. Deep Model Pruning without Finetuning for Few Category Datasets
  762. Deep Receiver for Multi-Layer Data Transmission with Superimposed Pilots
  763. Deep Support Vein Machine for Lung Parcellation
  764. Deep Sylvester Posterior Inference for Adaptive Compressed Sensing in Ultrasound Imaging
  765. Deep Time Series Anomaly Detection with Local Temporal Pattern Learning
  766. Deep Transfer Regression for EEG-based Driving Fatigue Detection
  767. Deep Unfolded Approximate Message Passing for Quantitative Acoustic Microscopy Image Reconstruction
  768. Deep Unfolding Using Score-based Generative Networks for Automotive Radar Interference Mitigation
  769. Deep Unfolding of Full Waveform Inversion for Quantitative Ultrasound Imaging
  770. Deep Variational Sequential Monte Carlo for High-Dimensional Observations
  771. Deep-Relative-Trust-Based Diffusion for Decentralized Deep Learning
  772. DeepMatch: Navigating the Complexities of Underwater Textures for Enhanced Keypoint Matching
  773. DeepPEM-AFC: An Improved Prediction-Error-Method-based Adaptive Feedback Cancellation with Deep Learning for Hearing Aids
  774. DeepPreNet: A Deep Learning Pre-Processing Method for Speech Distortion Correction in Parametric Array Loudspeaker
  775. Deepfake Detection of Singing Voices With Whisper Encodings
  776. Deeply Coupling EEG Signals and Eye Movements for Multi-Modal and Region-Aware Emotion Recognition
  777. DeformAvatar: Point-Based Human Avatar Re-targeting and Rendering
  778. Deformable Attention-Based Edge-Aware Network for Single Image Super-Resolution
  779. Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
  780. Delving Into Coarse-Fine Feature Interaction Alignment for UAV Object Detection
  781. Delving into Transformer-based Network Architecture for Guided Depth Super-Resolution
  782. Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification
  783. Denoising and Restoring Channel State Information for 5G Indoor Positioning in Low-SNR Scenarios
  784. Dense Point Clouds Matter: Dust-GS for Scene Reconstruction from Sparse Viewpoints
  785. Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto Accompaniments
  786. Density-Adaptive Fuzzy Clustering with Isolation Kernel
  787. Density-aware and Depth-aware Visual Representation for Zero-Shot Object Counting
  788. DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection
  789. Description-Based Controllable Text-to-Speech With Cross-Lingual Voice Control
  790. Design and Optimization of Superdirective Beamforming and Post-Filtering for Speech Enhancement
  791. Design of Multiple Binary Waveforms for the Joint MIMO Radar and Communications
  792. Design of Robust Differential Beamformers with Microphone Arrays of Arbitrary Planar Geometry
  793. DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speech
  794. Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
  795. Detecting OOD Samples via Optimal Transport Scoring Function
  796. Detecting and Defending Against Adversarial Attacks on Automatic Speech Recognition via Diffusion Models
  797. Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
  798. Developing a Multilingual Dataset and Evaluation Metrics for Code-Switching: A Focus on Hong Kong's Polylingual Dynamics
  799. Device Selection for Resource-Efficient Edge Caching in a Federated Learning Framework
  800. Device-aware Optical Adversarial Attack for a Portable Projector-camera System
  801. DiGradPatch: Black-Box Patch Attacks via Diffusion-Based Double Gradient and Sensitive Distribution Guidance
  802. Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver
  803. Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
  804. Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance
  805. DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
  806. DiffAttack: Imperceptible and Transferable Audio Adversarial Attack via Diffusion Model
  807. DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
  808. DiffDesign: A diffusion model using garment Knowledge-Enhanced for Fashion Design Synthesis
  809. DiffETM: Diffusion Process Enhanced Embedded Topic Model
  810. DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
  811. DiffKillR: Killing and Recreating Diffeomorphisms for Cell Annotation in Dense Microscopy Images
  812. DiffListener: Discrete Diffusion Model for Listener Generation
  813. DiffMEL: A large-scale difficulty-graded dataset for Multimodal Entity Linking
  814. DiffRS: An Extensible Diffusion Model for Remote Sensing Image Generation
  815. DiffSR: Learning Radar Reflectivity Synthesis via Diffusion Model from Satellite Observations
  816. DiffSSD: A Diffusion-Based Dataset For Speech Forensics
  817. Difference Bonds Consistency and Complementarity to Enhance Multimodal Representation Learning
  818. Differentially Private Distribution Estimation Using Functional Approximation
  819. Differentially Private and Communication-efficient Decentralized Learning Using Deep Quantizers
  820. DiffuseFIST: A Fast Image-guided Style Transfer Method for Adapting Large-scale Diffusion Models
  821. Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face Generation
  822. Diffusion Augmentation Sub-center Modeling for Unsupervised Anomalous Sound Detection with Partially Attribute-Unavailable Conditions
  823. Diffusion Counterfactual-Based Anomaly Detection in Class-Imbalanced Data
  824. Diffusion Features to Bridge Domain Gap for Semantic Segmentation
  825. Diffusion Learning Over Adaptive Competing Networks
  826. Diffusion Model Based Image Reconstruction in Lensless Imaging
  827. Diffusion Model with Multi-layer Wavelet Transform for Low-Light Image Enhancement
  828. Diffusion Models are Good Unsupervised Class-agnostic Shape Part Segmentators
  829. Diffusion Models are Zero-Shot Generative Text-Vision Retrievers
  830. Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
  831. Diffusion-based Data Augmentation for Object Counting Problems
  832. Diffusion-based Identity-Preserving Facial Privacy Protection
  833. Diffusion-based Target Device Style Transfer for Robust Acoustic Scene Classification
  834. Diffusion-based Unsupervised Audio-visual Speech Enhancement
  835. Digital Operating Mode Classification of Real-World Amateur Radio Transmissions
  836. Digital Twin-Driven Bearing-Fault Detection in Induction Motor and Drives using Graph Sampling and Aggregation Network
  837. Dike: Enhancing Fairness and Efficiency in GPU Clusters for Deep Learning
  838. Dilated Convolution for Time Series Learning
  839. Dimensionality-Reduced Spatial Bipartite Graph Clustering for Hyperspectral and LiDAR Data
  840. Directional Source Separation for Robust Speech Recognition on Smart Glasses
  841. DirichNet Model for Detection of TMS-Induced Speech Errors in Patients Undergoing Epilepsy Surgery
  842. Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge
  843. Discriminating Mizo Hunting and War Chants using Acoustic Features
  844. Disentangle Heart Rate Signals for Improved Stress Detection
  845. Disentangled Representation Learning for Chinese Handwriting Recognition
  846. Disentanglement Analysis in Deep Latent Variable Models Matching Aggregate Posterior Distributions
  847. Disentangling Hierarchical Features for Anomalous Sound Detection Under Domain Shift
  848. Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
  849. Disparity-Guided Cross-View Transformer For Stereo Image Super-Resolution
  850. Distance Based Single-Channel Target Speech Extraction
  851. Distill To Detect: Amplifying Anomalies in Backdoor Models through Knowledge Distillation
  852. DistillW2N: A Lightweight One-Shot Whisper to Normal Voice Conversion Model Using Distillation of Self-Supervised Features
  853. Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
  854. Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
  855. Distilling Knowledge from Large Video Models for Driver Visual Attention Prediction
  856. Distributed ATC Particle Filters for Cooperative Quaternion Tracking
  857. Distributed IRSs Mitigate Spatial Wideband & Beam Split Effects
  858. Distributed Interference Alignment Precoding and Detection for MU-MIMO OTSM Downlink in Time-Varying Channels
  859. Distributed Navigation with Dynamic Obstacles
  860. Distributed-Robust Source Localization in Wireless Acoustic Sensor Networks
  861. Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure Segmentation
  862. Distributionally Robust Kalman Filtering over an Infinite-Horizon
  863. Diverse Collaboration in Multi-Agent Reinforcement Learning via Self-Adaptive Method
  864. Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding
  865. Diversity Matters: Co-training for Semi-Supervised Change Detection in Remote Sensing Images
  866. Diversity Seeking Techniques for Red-Teaming Large Language Models
  867. Divide-and-Conquer Variational Bayesian Inference for Multi-task Learning of High-resolution SAR Imagery
  868. Do Less and Achieve More: Free Condition Video Outpainting with Diffusion Model
  869. Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning
  870. DoA-Aided MMSE Channel Estimation for Wireless Communication Systems
  871. DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
  872. Domain Connection based Unsupervised Domain Adaptation for Semantic Segmentation
  873. Domain Obfuscation for Efficient Secure Aggregation of Sparse Feature Vectors
  874. Domain-Aware Knowledge Debiasing for Generalizable Video Understanding in CLIP
  875. Domain-Incremental Learning for Audio Classification
  876. Domain-Independent Automatic Generation of Descriptive Texts for Time-Series Data
  877. Domain-Specific Adaptation in Speech Emotion Recognition Using Emotional Distribution Alignment
  878. Domain-aware Node Representation Learning for Graph Out-of-Distribution Generalization
  879. Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network
  880. Doppler Single-Photon Lidar
  881. Double Domain Converter Transformer For Improving EEG-Based Emotion Recognition from Video to Game Scenarios
  882. DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models
  883. DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
  884. DreamHA: Towards High-Quality Human Animation with Image-to-Video Diffusion Models
  885. DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
  886. Driver Reaction Time Prediction Through Adaptive Evolutionary Synchrony Window and Convolutional-LSTM
  887. DuCol: Text-Tag Adaptive Colorization of Dual-Character Line Art
  888. DuPI: Dual-resolution Pseudo-label Integration for Semi-supervised Instance Segmentation
  889. Dual Attention for Space-Time Video Super-Resolution
  890. Dual Decoder for Fast Inference in Natural Language Generation
  891. Dual Encoders for Diffusion-based Image Inpainting
  892. Dual Multi-Scale GCN with Deformable Temporal Kernel for Skeleton-based Action Recognition
  893. Dual Path Unsupervised Real Image Denoising
  894. Dual Position Attention Time-Frequency Network for Binaural Audio Synthesis
  895. Dual Trajectory Revised Diffusion Model for Time Series Forecasting
  896. Dual-Domain Feature-Guided Task Alignment for Enhanced Small Object Detection
  897. Dual-Frequency Spatio-Temporal Phase Unwrapping
  898. Dual-Function Waveform Design in Wireless Sensor Networks via SoS Optimization
  899. Dual-Modality Guided Artistic Style Transfer with Pre-trained Diffusion Models
  900. Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection
  901. Dual-Path Consistency Unsupervised Domain Adaptation for Nighttime Semantic Segmentation
  902. Dual-Path Contrastive Short Text Clustering with High-order Random Walk
  903. Dual-Path Model for Pulmonary Artery Segmentation
  904. Dual-Population Watermark Vaccine: Efficient and Imperceptible Adversarial Attack for Watermarked Image Protection
  905. Dual-Process Watermarked Diffusion: Integrating Watermarking With Denoising in Point Clouds
  906. Dual-Pyramid Attention Collaborative Network for Oracle Bone Inscription Classification
  907. Dual-Space Augmented Intrinsic-LoRA for Wind Turbine Segmentation
  908. Dual-Triple Transformer Networks for Accurate CT Pleural Effusion Segmentation
  909. Dual-energy CT metal artifact reduction by combined material decomposition and projection domain threshold segmentation
  910. Dual-level AMR Injection for Prompt-based Event Argument Extraction
  911. Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation
  912. Dynamic Category Queries Transformer for Generalized Few-shot Semantic Segmentation
  913. Dynamic Dictionary Design for Localization in Automotive Radar Systems
  914. Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
  915. Dynamic Graph Convolutional Networks with Spatiotemporal Missing Pattern Awareness
  916. Dynamic Graph Multi-granularity Attribute Scene Evolution Sequence Recommendation
  917. Dynamic Graph Recommendation via Sparse Augmentation and Singular Adaptation
  918. Dynamic Incentive Model for Federated Learning Model Trading via Evolutionary Game Theory
  919. Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing
  920. Dynamic Object Queries for Transformer-based Incremental Object Detection
  921. Dynamic Prototype Rehearsal for Continual ECG Arrhythmia Detection
  922. Dynamic ROI Adaptation for Accurate Non-Contact Heart Rate Estimation Using VGG-13 based Encoder-Decoder Model and Facial Landmarks
  923. Dynamic Routing and Calibration for Few-Shot Object Detection
  924. Dynamic SRM Curriculum for Trustworthy Multi-modal Classification
  925. Dynamic Soft Contrastive Learning for Time Series Anomaly Detection
  926. Dynamic Sparse Encoding and Cross-Temporal Attention for Remote Sensing Image Change Detection
  927. Dynamic Speech Generation to Enhance Intelligibility in Noisy Environments
  928. Dynamic SpikFormer: Low-Latency & Energy-Efficient Spiking Neural Networks with Dynamic Time Steps for Vision Transformers
  929. Dynamic Structure Hypergraph for Document-level Event Extraction
  930. Dynamic-static Feature Fusion with Multi-scale Attention for Continuous Blood Glucose Prediction
  931. DynamicAttention: Dynamic KV Cache for Disaggregate LLM Inference
  932. Dynamically Causal-Enhanced Exercise Representations for Adaptive Knowledge Tracing
  933. Dynamically Optimize MTD Strategy in Satellite Computing Systems Using A2C Reinforcement Learning
  934. Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech Recognition
  935. E-RNS : Enhancing Negative Sample Quality from Gradient Perspective for Graph Recommendation
  936. E-URES 2.0: Efficient User-Centric Residual-Echo Suppression with a Lightweight Neural Network
  937. E1 TTS: Simple and Fast Non-Autoregressive TTS
  938. ECBANet: Exploiting Complementary Information for Efficient Burst Super-Resolution
  939. ECG-guided individual identification via PPG
  940. ECSNN: Spiking Neural Networks for Efficient Exposure Correction in Endoscopy Imaging
  941. EDSep: An Effective Diffusion-Based Method for Speech Source Separation
  942. EEG Correlation Analysis-guided Graph Local Enhanced Feature Learning For Emotion Recognition
  943. EEG Decoding and Visual Reconstruction via 3D Geometric with Nonstationarity Modelling
  944. EEG-Music Emotion Recognition: Challenge Overview
  945. EEG-ReMinD: Enhancing Neurodegenerative EEG Decoding through Self-Supervised State Reconstruction-Primed Riemannian Dynamics
  946. EFL-PEFT: A communication Efficient Federated Learning framework using PEFT sparsification for ASR
  947. EGAS: Enhanced Geometry-aware 3D Asset Generation Using Gaussian Splatting
  948. EGENN: An Efficient Graph-Enhanced Neural Network for Multivariate Time Series Forecasting
  949. EM-MIAs: Enhancing Membership Inference Attacks in Large Language Models through Ensemble Modeling
  950. EMMeTT: Efficient Multimodal Machine Translation Training
  951. EP-SAM: An Edge-Detection Prompt SAM Based Efficient Framework for Ultra-Low Light Video Segmentation
  952. EPCPE: A Real-time End-to-End Pipeline for RGB-based Category-level 6D Pose Estimation
  953. EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities
  954. EPI-Mamba: State Space Model for Semantic Segmentation from Light Fields
  955. EPIC: Error Pattern Informed Correction for Classroom ASR with Limited Labeled Data
  956. ERGNN: Spectral Graph Neural Network With Explicitly-Optimized Rational Graph Filters
  957. ES-NeRF: Enhancing Segmentation in NeRF with CLIP
  958. ETDE-Net: An End-to-End Time-Domain Enhancement Network for LPI Radar Signals
  959. EagerLog: Active Learning Enhanced Retrieval Augmented Generation for Log-based Anomaly Detection
  960. Earbuds Orientation Alignment Based on Markov Chain Monte Carlo Sampling
  961. Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
  962. Easing Optimization Paths: a Circuit Perspective
  963. Easy, Interpretable, Effective: openSMILE for voice deepfake detection
  964. Easy-to-hard Instance-level Feature Fusion for Co-saliency Detection
  965. EasyControl: Adding Control to Video Diffusion for Controllable Video Generation and Interpolation
  966. Edge First: Edge-Guided Geometry for Superior 3D Roof Wireframe Reconstruction
  967. Edge-aware Laplacian Pyramid Network for Efficient Image Deblurring
  968. Edge-interaction Mamba Network for MRI Brain Tumor Segmentation
  969. Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
  970. Effective Context Modeling Framework for Emotion Recognition in Conversations
  971. Effective Integration of KAN for Keyword Spotting
  972. Effective Pre-Training of Audio Transformers for Sound Event Detection
  973. Effective Techniques for Scaling Audio Encoder Pretraining
  974. Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
  975. EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference Speed
  976. Efficient Anchor Graph Clustering Through Enhanced Within-Cluster Homogeneity
  977. Efficient Co-Approximate Parallel Compressive Depth Reconstruction on FPGA
  978. Efficient Co-clustering via Anchor-refined Label Spreading
  979. Efficient Data-Dependent Random Projection for Least Square Regressions
  980. Efficient Dataset Distillation through Low-Rank Space Sampling
  981. Efficient Defocus Deblurring Networks based on Diffusion Models
  982. Efficient Estimation of Kernel Matrix Spectral Norm using Random Features
  983. Efficient Extreme Large-Scale Speaker Verification: Dynamic Active Sub Fully-Connected Layers for Faster Training and Memory Optimization
  984. Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging Scenarios
  985. Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
  986. Efficient Fusion of Computationally Diverse Modalities Using Chunking and Cross-Attention
  987. Efficient Global Attention and Correlation-Aware Fusion for Hyperspectral Image Classification
  988. Efficient Gridless Wideband Direction-of-Arrival Estimation From Many Frequencies
  989. Efficient Hierarchical Domain Adaptive Thermal Infrared Tracking
  990. Efficient Infrared Image Super-Resolution Reconstruction via Guided Filter Coefficients Estimation with Parallax Attention Mechanism
  991. Efficient Large-Scale Scene Point Cloud Upsampling with Implicit Neural Networks and Spatial Hashing
  992. Efficient Learning of Balanced Signed Graphs via Iterative Linear Programming
  993. Efficient Localized Perception for Resource-Constrained Vision Systems
  994. Efficient Long Document Ranking via Adaptive Token Pruning with Query-Document Alignment
  995. Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation
  996. Efficient Long-Form Speech Recognition for General Speech In-Context Learning
  997. Efficient MDCT-Based Multi-Channel Coding with Perceptual Whitening and Broadband ILD Compensation
  998. Efficient Modeling and Low Complexity Implementation of Rate Estimation in Versatile Video Coding
  999. Efficient Multi-branch Black-box Semantic-aware Targeted Attack Against Deep Hashing Retrieval
  1000. Efficient Non-Sequential Relational Modeling for Temporal Knowledge Graph Link Predictions

Looking for submission deadlines instead? See the conference deadline calendar.