← All conferences

ICASSP 2025 Accepted Papers

The full list of 3,298 papers accepted at ICASSP 2025 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

  1. Efficient Object Placement Via LLM and Diffusion Model
  2. Efficient Prototypical Classifier for Class-Incremental Learning
  3. Efficient Pruning for Large-Scale Seq2Seq Speech Models without Back-Propagation
  4. Efficient Quality Controllable Neural Image Compression based on QD-Model
  5. Efficient Quantization and Denoising Using Local Graph Fourier Frames
  6. Efficient Spatial Audio Rendering Via Differentiable FIR To IIR Estimation
  7. Efficient Streaming LLM for Speech Recognition
  8. Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression
  9. Efficient Visual Storytelling through Descriptive Words Distillation and Dynamic Decoding
  10. Efficient and Effective Model Extraction
  11. Efficient and Expandable Token-Level Approach for Multi-Domain Sensitive Information Classification
  12. Efficient-USR: Prompt Guided Dual-Domain Feature Information for Efficient Underwater Image Super-Resolution
  13. EfficientNet-Gaze: Integrating Multi-Scale Feature Extraction with Frequency Domain Analysis for Efficient Gaze Estimation
  14. EfficientSleepNet: A Novel Lightweight End-to-End Model for Automated Sleep Staging on Single-Channel EEG
  15. EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates
  16. Egocentric Speaker Diarization with Vision-Guided Clustering and Adaptive Speech Re-detection
  17. Electrocardiogram Report Generation and Question Answering via Retrieval-Augmented Self-Supervised Modeling
  18. Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition
  19. Embedding Enhanced MLP Enables Simple and Extensible Spatiotemporal Forecasting
  20. Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
  21. Emotion information recovery potential of wav2vec2 network fine-tuned for speech recognition task
  22. Emotion-Preserving Prosody Anonymization Network for Voice Privacy Protection
  23. Emotion-aware Structural Enhancement Graph Auto-Encoder for Rumor Detection
  24. Emotional Knowledge Self-Distillation in Dialogue
  25. Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
  26. Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
  27. Enabling DMG Wi-Fi Sensing in Data Transmission Intervals by Exploiting Beam Training Codebook
  28. End-to-end acoustic-articulatory dysarthric speech recognition leveraging large-scale pretrained acoustic features
  29. Energy Backdoor Attack to Deep Neural Networks
  30. Energy Consumption Trends in Sound Event Detection Systems
  31. Energy-based Model Guided Self-Supervised Learning for Speaker Verification
  32. Enhanced Breast Cancer Molecular Biomarker Classification: A Novel Two-Stage Machine Learning Pipeline for Accurate Histological Analysis of Whole Slide Images
  33. Enhanced Control for Diffusion Bridge in Image Restoration
  34. Enhanced Corneal Endothelial Cell Segmentation via Frequency-Selected Residual Fourier Diffusion Models
  35. Enhanced Loudspeaker Membrane Excursion Control Method Using Low Latency Distortion Prediction and Efficient LSTM Networks
  36. Enhanced Multimodal Depression Detection With Emotion Prompts
  37. Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph Propagation
  38. Enhanced Sparse Bayesian Learning Methods with Application to Massive MIMO Channel Estimation
  39. Enhanced Weakly Supervised Few-shot Classification & Segmentation
  40. Enhancing 3D Medical Image Understanding with 2D Multimodal Large Language Models
  41. Enhancing 6D Pose Estimation with Cross-modal Fusion Network and Density-peak Keypoint Localization
  42. Enhancing Age-Related Robustness in Children Speaker Verification
  43. Enhancing Autonomous Driving through Dual-Process Learning with Behavior and Reflection Integration
  44. Enhancing Autonomous Vehicle Planning With a Robust Fault-Tolerant Mechanism for Action-Induced Agent Detection
  45. Enhancing Boundary-Handling Strategies for Convolutional Sparse Representation Model with 46 × 46 Convolution-Multiplication Properties
  46. Enhancing Change Detection in Remote Sensing: Integrating Synthetic Data with Semi-Supervised Learning
  47. Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality Learning
  48. Enhancing Complex Formula Recognition with Hierarchical Detail-Focused Network
  49. Enhancing Continual Learning for Medical Imaging: Efficient Knowledge Transfer and Multi-Disease Prediction
  50. Enhancing Convolutional Models for Indoor Radio Mapping via Ray Marching
  51. Enhancing Cross-Domain Slot Filling with Joint LLM Data Generation and Data Curation
  52. Enhancing DETR Efficiency with Inter-Object Relationship and Semantic Spectral Decomposition-Based Distillation
  53. Enhancing Data-Free Class-Incremental Learning via Image-Centric Dual Distillation
  54. Enhancing Document-Level Relation Extraction through Entity-Pair-Level Interaction Modeling
  55. Enhancing EEG-based Covert Speech Decoding through Knowledge Transfer
  56. Enhancing Emotion Reasoning for Image Multi-Emotion Prediction
  57. Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
  58. Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
  59. Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
  60. Enhancing Extrapolation Reasoning on Temporal Knowledge Graphs with Logic Rules and Queries
  61. Enhancing Fairness in Gaussian Mixture Clustering through Impact Factor
  62. Enhancing Federated Domain Adaptation via Multi-Granular Fine-Grained Alignment
  63. Enhancing Federated Knowledge Distillation in Heterogeneous and Non-IID Scenarios
  64. Enhancing Few-Shot Out-of-Distribution Detection with Gradient Aligned Context Optimization
  65. Enhancing Generalized EEG Classification with Decomposed Statistics-diverse Feature Augmentation
  66. Enhancing Graph-based Fraud Detection by Adversarial Confidence Reweighting
  67. Enhancing Image Editing with Chain-of-Thought Reasoning and Multimodal Large Language Models
  68. Enhancing Image Generation Fidelity via Progressive Prompts
  69. Enhancing Imaging Generation through Implicit Neural Representations and HyperNetwork for Spatial Variability
  70. Enhancing Incomplete Multimodal Learning via Modal Complementary Recovering
  71. Enhancing Information Extraction with METORIE: A Metaphor and Trap-Based Dataset for Cross-Domain Fine-Tuning
  72. Enhancing Large Language Model Inference Efficiency via Lookahead Cache Filtering
  73. Enhancing Large Language Models on Domain-specific Tasks: A Novel Training Strategy via Domain Adaptation and Preference Alignment
  74. Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
  75. Enhancing Long-Term Capabilities of Large Language Models via Discourse Sub-graph Analysis
  76. Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
  77. Enhancing Multi-Channel Speech with Limited Microphones via Spherical Harmonic Transform
  78. Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
  79. Enhancing Multimodal Analogical Reasoning Through Triplet Interaction
  80. Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
  81. Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention
  82. Enhancing Multivariate Time Series Forecasting with Multi-scale Moving Transformation
  83. Enhancing Network Calibration for Low-Cost Gas Sensor Networks Through Adaptive Similarity Search
  84. Enhancing Out-of-Distribution Detection through Dynamic Activation Function
  85. Enhancing Precision in Image-Guided Spine Surgery through the Prediction of Occluded Fiducials Utilizing ResNet Architecture
  86. Enhancing Privacy in Radar-Based Vital Sign Monitoring Via Non-Linear FMCW Waveforms
  87. Enhancing Prosody Transfer in Speech Synthesis by Using Prosodically-Aligned References
  88. Enhancing Robustness of Implicit Neural Representations Against Weight Perturbations
  89. Enhancing Session-Based Recommendation with Hypergraph Motifs and Contrastive Learning
  90. Enhancing Small Model Performance in Educational Classification Tasks through Knowledge Distillation
  91. Enhancing Speech Emotion Recognition with Speech Dynamic Modeling and Multi-Modal Knowledge Distillation
  92. Enhancing Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
  93. Enhancing Stutter Detection using Long-Term Average Spectrum Values
  94. Enhancing TTS Stability in Hebrew using Discrete Semantic Units
  95. Enhancing Task-Specific Feature Learning with LLMs for Multimodal Emotion and Intent Joint Understanding
  96. Enhancing Teacher Classroom Behavior Descriptions: A Spatio-Temporal Graph-Based Method for Video Captioning
  97. Enhancing Text Annotation Through Rationale-Driven Collaborative Few-Shot Prompting
  98. Enhancing Time Series Prediction with Evolutionary Algorithm-based Optimization of LSTM
  99. Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes
  100. Enhancing Video-Text Matching via Sparse Stratified Sampling
  101. Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
  102. Enhancing Vision: Harmonizing Frequency for Imaging Quality and Perception Accuracy
  103. Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning
  104. Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
  105. Enhancing Zero-Shot Cross-Lingual Event Argument Extraction with Language-Independent Information
  106. Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration Prediction
  107. Enhancing Zero-Shot Relation Extraction through Staged Interaction with Large Language Models
  108. Enhancing the Robustness of LiDAR-based Object Detection under Disappearing Attacks
  109. Epigraph Based Multilevel Optimization (EMO) for Enhancing Chain-of-Thought Reasoning Capabilities
  110. EqGAN: Reformation-based Feature Equalization Fusion for Few-shot Image Generation
  111. Error Bounds Revisited, and How to Use Bayesian Statistics While Remaining a Frequentist
  112. Error Feedback Approach for Quantization Noise Reduction of Distributed Graph Filters
  113. Essentia: Boosting Artifact Removal from EEG through Semantic Guidance Utilizing Diffusion Model
  114. Estimating Instrument Spectral Response Functions Using Sparse Representations and Quadratic Envelopes
  115. Estimating Multi-chirp Parameters using Curvature-guided Langevin Monte Carlo
  116. Estimating Musical Surprisal in Audio
  117. Estimating the Number and Locations of Boundaries in Reverberant Environments with Deep Learning
  118. Estimation of Doppler, Range, and Direction of Targets in Wideband Bistatic Automotive Radar
  119. Estimation of Multi-Attribute Differential Graphs with Non-Convex Penalties
  120. EvaSR: Rethinking Efficient Visual Attention Design for Image Super-Resolution
  121. Evaluating Contrastive Methodologies for Music Representation Learning Using Playlist Data
  122. Evaluating Snippet Significance: A Framework for Audio and Text-Based Dialogue Summarization
  123. Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation
  124. Evaluating the Posterior Sampling Ability of Plug&Play Diffusion Methods in Sparse-View CT
  125. Evaluating the security of public surrogate watermark detectors
  126. Evaluation of Deep Audio Representations for Hearables
  127. Evaluation of Wearable Head BCG for PTT Measurement in Blood Pressure Intervention
  128. Event Masked Autoencoder: Point-wise Action Recognition with Event-Based Cameras
  129. Event-Driven Prony: Towards Asynchronous Spectral Estimation
  130. Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration
  131. EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking
  132. Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
  133. Evidential Deep Learning with Reweighted Margin Adjustment for Uncertainty-Driven Cervical OCT Image Diagnosis
  134. Evidential Neural GPLDA: A Novel Approach to Quantify Prediction Uncertainty in Speaker Verification Systems
  135. Evidential-TTS: High Fidelity Zero-Shot Text-to-Speech Using Evidential Deep Learning
  136. ExVC: Leveraging Mixture of Experts Models for Efficient Zero-shot Voice Conversion
  137. Exact Rotation Invariant Robust PCA
  138. Exact Solutions of the Inner Optimization Problem of Adversarial Robustness
  139. Explainable Adversarial Attacks on Coarse-to-Fine Classifiers
  140. Explainable Detection of Alzheimer's Disease Through Analysis of Human Behavior in Video
  141. Explainable Orthogonal Attention Networks for EEG-based Analysis: Leveraging Disentangled Representations to Enhance Diagnosis
  142. Explainable Reinforcement Learning for Trajectory Design in UAV-assisted Wireless Networks
  143. Explaining Representations in Correlation-based Deep Multiview Representation Learning
  144. Explaining Speaker and Spoof Embeddings via Probing
  145. Explicit Mutual Information Maximization for Self-Supervised Learning
  146. Explicit Spatial Hint and Implicit Logits Relation: Distilling Heterogeneous Knowledge From Vision Transformer to CNN
  147. Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
  148. Exploiting Attention-to-Motion via Transformer for Versatile Video Frame Interpolation
  149. Exploiting Beam-Split in IRS-aided Systems via OFDMA
  150. Exploiting Foundation Models for Label-Efficient Few-Shot Learning via Feature Coupling: A Case Study of cardiac CT Segmentation
  151. Exploiting Robust Model Watermarking Against the Model Fine-Tuning Attack via Flat Minima Aware Optimizers
  152. Exploiting Wavelet Scattering Transform & Squeeze-Excitation Blocks with Cross-Modal Attention for Multi-modal Emotion Recognition
  153. Exploiting the Relationship within the Unlabelled Samples by Set Matching for Generalized Category Discovery
  154. Exploration of Sequence-wise Optimized Parameters for Low Complexity Enhancement Video Coding (LCEVC) on 4K Content
  155. Explore the Hallucination on Low-level Perception for MLLMs
  156. Exploring Acoustic Foundations in Speech Production Assessment Models for Children with Cochlear Implants
  157. Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
  158. Exploring Acted Sleepy Speech to Advance Real-World Sleepiness Estimation and Cognitive Degradation Detection
  159. Exploring Antenna Placement Configurations with a Radar-based Silent Speech Interface
  160. Exploring Facial Kinship Verification through Contactless Heart Activity Analysis
  161. Exploring Generalization Boundaries of Unsupervised Industrial Anomaly Detection Models through Attribute Perturbation
  162. Exploring Graph-aware Reasoning and Bidirectional Selection for Vision-Language Navigation
  163. Exploring Group Theory for Optimal Cognitive Radar Waveform Design
  164. Exploring Inter-Variate and Long-Term Dependencies to Boost Multivariate Time Series Forecasting
  165. Exploring Kolmogorov-Arnold networks for realistic image sharpness assessment
  166. Exploring Large Language Models for Knowledge Graph Completion
  167. Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
  168. Exploring Meta Evidence for Prompt Optimization
  169. Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
  170. Exploring Simple Siamese Network for High-Resolution Video Quality Assessment
  171. Exploring Spectral Signatures of Chinese liquor using Machine Learning and SHapley Additive exPlanations
  172. Exploring Temporal Constraints for Unsupervised Iris Motion Tracking in AS-OCT Videos
  173. Exploring Text-Queried Sound Event Detection with Audio Source Separation
  174. Exploring Triple Knowledge Cues for Zero-Shot Human-Object Interaction Detection
  175. Exploring the Differences between Deaf and Hearing Infant Cries
  176. Exploring the Distribution of Cell Subpopulations in Pancreatic Ductal Adenocarcinoma Slides by Joint Spatial Transcriptomics and Pathology Data
  177. Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion
  178. Exploring the Interpretability of EEG-Inception Convolutional Neural Networks for Epilepsy Prediction
  179. Exploring the Robustness of In-Context Learning with Noisy Labels
  180. Exploring the Role of CLIP Global Visual Features in Multimodal Large Language Models
  181. Exponential Convergence of Stochastic Mirror Descent in Over-parameterized Linear Models
  182. Extending MPR for Locating a Moving Object Based on TDOA and FDOA
  183. Extending Whisper for Emotion Prediction Using Word-level Pseudo Labels
  184. Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
  185. Extracting Sparse Specialist Models from Generalist Models
  186. Extremum Encoding for Joint Baseband Signal Compression and Time-Delay Estimation for Distributed Systems
  187. Eye Movements as Images: A Multimodal Framework for Eye Movements Representation
  188. F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music Generation
  189. FA-GAN: Defense Against Adversarial Attacks in Automatic Modulation Recognition
  190. FABLE: A Bundle Method For Federated Learning In Wireless Systems
  191. FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
  192. FAF-Filt: Frequency-aware Fourier Filter for Sound Event Detection
  193. FARE: A Deep Learning-Based Framework for Radar-Based Face Recognition and Out-of-Distribution Detection
  194. FAST: Fast Audio Spectrogram Transformer
  195. FASTER: Face Attribute Sliders with Semantic Rewards
  196. FAWL: Weakly-Supervised Video Corpus Moment Retrieval with Frame-Wise Auxiliary Alignment and Weighted Contrastive Learning
  197. FBI-Net: Frequency Band Integration Network for Infrared Small Target Segmentation
  198. FBSE-FTFCWT-Based Novel Automated Framework for Dysarthric Speech Detection
  199. FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining
  200. FCoDT-Net: A Novel Framework for High-Precision Medical Image Segmentation Using Contextual Distillation Transformer
  201. FDDSGCN: Fractional Decoupling Dynamic Spatiotemporal Graph Convolutional Network for Traffic Forecasting
  202. FDR Control for Complex-Valued Data with Application in Single Snapshot Multi-Source Detection and DOA Estimation
  203. FDR-Controlled Portfolio Optimization for Sparse Financial Index Tracking
  204. FEA-DETR: An Enhanced ConvNet for Detecting Prohibited Objects in X-Ray Images Using Frequency and Edge Aware Information
  205. FG3DFormer: Fine-Grained 3D Shape Classification Based on Vision Transformer
  206. FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection
  207. FKAN-GMFNet: Fourier Kolmogorov-Arnold-based Group Multi-scale Fusion Network for Aneurysm Image Segmentation
  208. FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing
  209. FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
  210. FR2ViT: Finetuning-free Token Reduction for Dense Prediction Through a Refinement-Reactivation Architecture
  211. FSENet: Frequency Separation Enhancement Network for Super-Resolution
  212. FUTGA-MIR: Enhancing Fine-grained and Temporally-aware Music Understanding with Music Information Retrieval
  213. FUVAS: Few-shot Unsupervised Video Anomaly Segmentation via Low-Rank Factorization of Spatio-Temporal Features
  214. Face Relighting with Ratio Function for Explicit Geometric Representation
  215. Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping
  216. Facial Expression Recognition with DToF Sensing
  217. Facilitating Semi-Supervised Pedestrian Detection with Structurally Controllable Instance Synthesis
  218. Factorized-VITS: Decoupling Prosody and Text in End-to-End Speech Synthesis without External or Secondary Aligner
  219. Fading-Invariant Adversarial Attacks on Neural Modulation Recognition
  220. Fair CoVariance Neural Networks
  221. Fair MP-BOOST: Fair and Interpretable Minipatch Boosting
  222. FairAdapter: Detecting AI-generated Images with Improved Fairness
  223. Faithful Self-Refinement in Mathematical Reasoning via Progressive Back-Translation
  224. FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
  225. Fast Adaptation of Pretrained Speaker Verification System for Source Speaker Tracking
  226. Fast DCT+: A Family of Fast Transforms Based on Rank-One Updates of the Path Graph
  227. Fast DPCNs for Feature Extraction without Labels
  228. Fast Sparse DFT Computation for Arbitrary Length by Circular Convolution
  229. Fast Sparse Learning from Streaming Data with LASSO
  230. Fast Structured Orthogonal Dictionary Learning using Householder Reflections
  231. Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
  232. Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
  233. Fast and Robust High Resolution Frequency Estimation of Damped Signals
  234. Fast inter-frame coding for dynamic meshes via supervoxel-based shape matching
  235. Faster Speech-LLaMA Inference with Multi-token Prediction
  236. FasterGold-DETR: An Efficient End-to-End Fire Detection Model via Gather-and-Distribute Mechanism
  237. Feature Disentangling Dual-stream Network for User Bias Alleviation in Social Media Prediction
  238. Feature Refinement Decomposition and Relation Preference Enhancement for Remote Sensing Change Detection
  239. FedCAda: Adaptive Client-Side Optimization for Accelerated and Stable Federated Learning
  240. FedDiT: Federated Learning by Distillation Token Enhanced Vision Transformer
  241. FedDiffRec: A Module-wise Training Approach for Diffusion-Based Recommendation in Federated Learning
  242. FedFLD: Heterogeneous Federated Learning via Forget-Less Distillation
  243. FedImp: Federated Learning Using Important Layers of Client Models for the Diagnosis of Breast Cancer Histopathology Images
  244. FedRPN: An Efficient Framework for Optimizing System Heterogeneity in Federated Learning
  245. FedSe: Group-Based Sequential Training Strategies for Mitigating Label Skew in Federated Learning
  246. FedTG: Text-guided Federated Domain Generalization
  247. FedTLU: Federated Learning with Targeted Layer Updates
  248. Federated Cross-Client Collaborative Filtering with Tensor Compressive Learning
  249. Federated Domain Generalization with Label Smoothing and Balanced Decentralized Training
  250. Federated Hybrid-Supervised Learning for Universal Medical Image Segmentation
  251. Federated Learning with Heterogeneous Feature Adaptation for Human Activity Recognition
  252. Federated Prototype Guided Adaption for Vision-Language Models
  253. Federated Smoothing ADMM for Quantile Regression with Non-Convex Sparse Penalties
  254. FeedbackFuzz: Fuzzing Processors via Intricate Program Generation with Feedback Engine
  255. Few-Shot Object Detection in Satellite Imagery with Feature Fusion Pyramid and Adaptive Region Proposal Networks
  256. Few-shot Image Classification based on Attribute Prediction and Selection
  257. Few-shot Keyword-incremental Learning Using Compositional Information
  258. FiTGAN: Content Fusion with Style Transformation for Few-shot Image Generation
  259. Filtering Resistant Large Language Model Watermarking via Style Injection
  260. Find Details in Long Videos: Tower-of-Thoughts and Self-Retrieval Augmented Generation for Video Understanding
  261. Fine-Grained Global Modeling Learning for Personalized Federated Sequential Recommender
  262. Fine-grained Vital Sign Reconstruction through Machine Learning on Multi-channel Radar Signals
  263. Fine-portraitist: Visualizing the Speaker's Face Portrait during Speech Listening
  264. Fine-tuning TitaNet-Large Model for Speaker Anonymization Attacker Systems
  265. Fioma: Towards Open-Set Semi-Supervised Specific Emitter Identification
  266. First-frame Supervised Video Polyp Segmentation via Propagative and Semantic Dual-teacher Network
  267. First-order State Space Model for Lightweight Image Super-resolution
  268. Flare-Aware RWKV for Flare Removal
  269. FlashSR: One-step Versatile Audio Super-resolution via Diffusion Distillation
  270. Flexible Event-Driven Biological Imaging via Bayesian Inference
  271. Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching for Speaker Diarization
  272. FlowMAC: Conditional Flow Matching for Audio Coding at Low Bit Rates
  273. FlowSE: Flow Matching-based Speech Enhancement
  274. FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
  275. Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
  276. Fooling the Forgers: A Multi-Stage Framework for Audio Deepfake Detection
  277. Foreground-aware Prototypical Network for Prohibited Item Detection from X-ray Scans
  278. ForensiCam-215K: A Large Scale Image and Video Dataset for Forensic Analysis
  279. Forensics Analysis of Residual Noise Texture in digital Images for Detection of Deepfake
  280. Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
  281. Found In The Distribution: Utilizing Latent Dirichlet Allocation Improves Long Context Comprehension of Large Language Models
  282. Foundation Model and Temporal Priors-guided Transductive Few-shot Action Recognition
  283. Foundation Models Boost Low-Level Perceptual Similarity Metrics
  284. Fourth-Order Cumulant Based 3-D Near-Field Underdetermined Parameter Estimation With Exact Spatial Propagation Model
  285. Fractional-Order Hyperbolic Tangent Based Adaptive Algorithm for Feedback Control in Hearing Aids
  286. Frank-Wolfe Method with Proximal Regularization for Constrained Federated Learning with Non-iid Data
  287. FreeAlign: Superior Text-Image Alignment by Modulating Prompt Attention
  288. FreeLesion: Synthetic Image-Mask Pairs for Fundus Lesion Segmentation via Curriculum Learning and Feature-Loss Guided Filtering
  289. FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
  290. FreeSegDiff: Annotation-free Saliency Segmentation with Diffusion Models
  291. Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection
  292. FreqSense: Universal and Low-Latency Adversarial Example Detection for Speaker Recognition with Interpretability in Frequency Domain
  293. Frequency Agnostic Tissue Characterization in Ultrasound Imaging using Backscattered Signal Statistics
  294. Frequency Domain Information Integrated Network for Low-Light Image Enhancement
  295. Frequency-Based Federated Domain Generalization for Polyp Segmentation
  296. Frequency-Domain Guided Multiple Parallel Kernels Network for Low-Light Remote Sensing Image Enhancement
  297. Frequency-Domain Popularity Forecasting with Shape-Based Retrieval
  298. Frequency-Space Margin Perception for Open Set Knowledge Distillation
  299. Frequency-enhanced Comprehensive Dependency Attention for Time Series Anomaly Detection
  300. Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
  301. From Characters to Subwords: Modeling Unit Conversion for Low-resource Speech Recognition
  302. From Pixels to Voice: A Simple and Efficient End-to-End Spoken Image Description Approach via Vision Codec Language Models
  303. From Voices to Beats: Enhancing Music Deepfake Detection by Identifying Forgeries in Background
  304. FruitMMBench: A Multi-modal Benchmark for Fruit Quality Assessment
  305. Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
  306. Full-Reference Point Cloud Quality Assessment with Multimodal Large Language Models
  307. Full-text Error Correction for Chinese Speech Recognition with Large Language Model
  308. Fully Connected Tensor Network based Brain Structural Feature Extraction for Early Alzheimer's Disease Detection
  309. Fully Spiking Neural Network for Legged Robots
  310. Functional Near-Infrared Spectroscopy Feature Extraction with Application in Workload Estimation
  311. Fundamental Social Learning Scaling Law for Tracking Hidden Markov Models
  312. Fusing Multimodality of Large Language Models and Satellite Imagery via Simplicial Contrastive Learning for Latent Urban Feature Identification and Environmental Application
  313. Fusion of Information in Multiple Particle Filtering in the Presence of Unknown Static Parameters
  314. Fusion-OSR: Cross-Domain Contrastive Learning with Weibull Calibration for Time Series Open Set Recognition
  315. FusionClassNet: A Multi-Scale Feature Fusion Network with Contrastive Loss-Driven Classification for Enhanced Lung Tumor Image Representations
  316. FuzzyMIL: Decoupling Pathological Phenotypes through Deep Fuzzy Clustering for Efficient Whole Slide Image Analysis
  317. G-Depth: An Efficient Graph Method for Robust Depth Completion
  318. GADACE: Graph Anomaly Detection Combining Attribute Contrast and Structure Reconstruction
  319. GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
  320. GATOmics: A Novel Multi-Omics Graph Attention Network Model for Cancer Driver Gene Detection
  321. GBA-Net: A Method for 3D Brain Tumor Segmentation Based on Multi-scale Gaussian Boundary Attention
  322. GCAT: Gated Convolutional Attention Transformer for Efficient Image Super-Resolution
  323. GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning
  324. GDDA: Semantic OOD Detection on Graphs under Covariate Shift via Score-Based Diffusion Models
  325. GDFDNet: A Novel Graph-Based Dynamically Fused Dual-Stream Network for Accuracy Prohibited Items Detection
  326. GDRIVE: Adaptive Object Detection in Autonomous Vehicles via Graph-Based Feature Learning
  327. GEE Maximization in UAV-Aided Mobile IoT Networks Using Deep Reinforcement Learning
  328. GEGA: Graph Convolutional Networks and Evidence Retrieval Guided Attention for Enhanced Document-level Relation Extraction
  329. GEMD-UNet: Graph Structure Enhanced Multi-dimensional Learning Unet for Cloud Detection
  330. GENIE: Socially Unbiased Generative Text-to-Image Editing
  331. GIST: Guided Interpretable Large Language Model Strategy Transfer for Multi-Task Reinforcement Learning
  332. GLST-GCN: Global-Local Spatio-Temporal Graph Convolutional Network for Skeleton-based Hand Motion Prediction
  333. GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation
  334. GMCL: Graph-Enhanced Multimodal Contrastive Learning for Rumor Detection
  335. GMM-Based Bootstrap Prototype-Aware Learning For Weakly Supervised Semantic Segmentation
  336. GMMCL: Adaptive Concept Drift in Data Streams with Gaussian Mixture Models based on Contrastive Learning
  337. GNCL: A Graph Neural Network with Consistency Loss for Segment-Level Spoofed Speech Detection
  338. GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision Tasks
  339. GPPT: Gaussian Process-infused Prompt Tuning for Vision-language Models
  340. GPT-C: Generative PrompT Compression
  341. GPT-LAD: Leveraging Large Multimodal Models for Logical Anomaly Detection
  342. GRACED: A Plug-and-Play Solution for Certifiable Graph Classification
  343. GREST: Ghost Targets Removal Algorithm Using Multipath Angle Estimation
  344. GS-PT: Exploiting 3D Gaussian Splatting for Comprehensive Point Cloud Understanding via Self-supervised Learning
  345. GSMM: Efficient Global Sparsification for Resource-Conscious Multimodal Models
  346. GateM2Former: Gated Feature Selection and Expert Modeling in Multimodal Emotion Recognition
  347. Gated Cross-Attention Network for Depth Completion
  348. Gaussian Constrained Diffeomorphic Deformation Network for Panoramic Semantic Segmentation
  349. Gaussian Difference: Find Any Change Instance in 3D Scenes
  350. Gaussian-Face: Talking Head Generation with Hybrid Density via 3D Gaussian Splatting
  351. GaussianEnhancer: A General Rendering Enhancer for Gaussian Splatting
  352. GaussianSlicer: Efficient Surface Reconstruction from Cross-sectional Slices with Gaussian Splatting
  353. Gaze-Assisted Human-Centric Domain Adaptation for Cardiac Ultrasound Image Segmentation
  354. Gaze-GZ: Generalized Gaze Estimation with Multi-scale Gaze Zone Prediction
  355. GeMIMO: Searching the Cores of X-formers for Time Series Forecasting
  356. Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
  357. General Dynamic Regularization Federated Learning with Hybrid Sharpness-Aware Minimization
  358. Generalizable Articulated Object Perception with Superpoints
  359. Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
  360. Generalizable Indoor Path Loss Prediction
  361. Generalizable Real-time Accelerated Dynamic MRI
  362. Generalization Guarantee of Decentralized Learning with Heterogeneous Data
  363. Generalize Audio Deepfake Algorithm Recognition via Attribution Enhancement
  364. Generalized Approximate Message-Passing for Compressed Sensing with Sublinear Sparsity
  365. Generalized Graph Signal Reconstruction via the Uncertainty Principle
  366. Generalized Linear Models with 1-Bit Measurements: Asymptotics of the Maximum Likelihood Estimator
  367. Generate E-commerce Product Background by Integrating Category Commonality and Personalized Style
  368. Generating Apoptosis-Inducing Anticancer Peptides Targeting BCL-xL Using Latent Diffusion Models on Small Datasets
  369. Generating Customized 4D Motions from Text Inputs Using Spatial-Temporal Slicing Approaches
  370. Generating Editable Head Avatars with 3D Gaussian GANs
  371. Generating Gezi Opera Scores with a Large Language Model and a High-Quality Dataset
  372. Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
  373. Generating Targeted Universal Adversarial Perturbation against Automatic Speech Recognition via Phoneme Tailoring
  374. Generating Vocals from Lyrics and Musical Accompaniment
  375. Generative Adversarial Network with Adaptive Synthesis for Brain-Computer Interfaces in Motor Imagery Classification
  376. Generative Adversarial Network with Structured Semantic Prompts Constrainting Clip for Text-to-Image
  377. Generative Dataset Distillation Based on Self-knowledge Distillation
  378. Generative Diffusion Model-based Energy Management in Networked Energy Systems
  379. Generative Expansion of Small Datasets: An Expansive Graph Approach
  380. Generative Model based Optical Response Prediction for Plasmonic Sensing
  381. Generative Sensing: Pre-training LiDAR with Masked Autoencoders for Ultra-Frugal Perception
  382. Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
  383. Geodesic Mean Threshold Scheme on Riemannian Manifold for EEG Signal Classification
  384. Geometric Feature-Driven Metric Learning for 3D Craniofacial Superimposition
  385. Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
  386. Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
  387. Global Context MambaVision for EEG-based Emotion Recognition
  388. Global Enhanced Frame Prompt Tuning for Sound Event Detection
  389. Global Static Pruning via Adaptive Sample Complexity Awareness
  390. Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model
  391. Globally Normalizing the Transducer for Streaming Speech Recognition
  392. GoLoColor: Towards Global-Local Semantic Aware Image Colorization
  393. GraFPrint: A GNN-Based Approach for Audio Identification
  394. GradPFL: Gradient-Driven Adaptive Clustering in Personalized Federated Learning
  395. Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
  396. Gradient-Oriented Clustered Federated Learning With Efficient Knowledge Sharing in Non-IID Settings
  397. Gram: A Large-Scale General EEG Model for Raw Data Classification and Restoration Tasks
  398. Granularity-Aware Contrastive Learning for Fine-Grained Action Recognition
  399. Graph Anomaly Detection via Multi-Scale Reconstruction of Graph Encoder-Decoder Networks
  400. Graph Contrastive Learning with Decoupled Augmentation
  401. Graph Embedded Stochastic Configuration Networks for Imbalanced Data Classification
  402. Graph Learning with Low-rank and Diagonal Structures: A Riemannian Geometric Approach
  403. Graph Neural Networks Meet Probabilistic Graphical Models: A Survey
  404. Graph Neural Networks for Parkinson's Disease Detection
  405. Graph Pooling via Dropping Task-Irrelevant Nodes
  406. Graph Refinement in Latent Space: A Hypergraph Convolution for Underwater Object Detection
  407. Graph Signal Reconstruction via Koopman Autoencoder
  408. Graph Structure Learning via Transfer Entropy for Multivariate Time Series Anomaly Detection
  409. Graph Topology Identification Based on Covariance Matching
  410. Graph-Driven Insights: Enhancing Stock Market Prediction with Relational Temporal Dynamics
  411. Graph-Enhanced Dual-Stream Feature Fusion with Pre-Trained Model for Acoustic Traffic Monitoring
  412. Graph-based Signal Sampling with Adaptive Subspace Reconstruction for Spatially-irregular Sensor Data
  413. GraphDAE-PU: Graph Denosing Auto-Encoder for Arbitrary-Scale Point Cloud Upsampling
  414. GraphVCM: Virtual Center Mixing with Distance-Aware Regulation for Class Imbalanced Node Classification
  415. Grey Wolf Optimizer Algorithm Based Active Noise Control Without Secondary Path Identification
  416. Group Modeling and Recommendation Based on Multi-Behavior Interactions in Live Streaming E-Commerce
  417. Group-CLIP Uncertainty Modeling for Group Re-Identification
  418. Group-wise Semantic-enhanced Interaction Network for Remote Sensing Spatio-Temporal Fusion
  419. Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker Recognition
  420. Grouping-Based Crowding Differential Evolution Approaches for Multimodal Feature Selection
  421. Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion Models
  422. Guided Speaker Embedding
  423. Guiding Inter-domain Class Balancing With Salient Features For Domain Adaptive Object Detection
  424. Guitar-TECHS: An Electric Guitar Dataset Covering Techniques, Musical Excerpts, Chords and Scales Using a Diverse Array of Hardware
  425. HANet: A Harmonic Attention-Based Network for Singing Melody Extraction from Polyphonic Music
  426. HAPG-SAQAM: Human Auditory Perception Guided Spatial Audio Quality Assessment Metric
  427. HATTM: A Novel Hybrid Attention Model for Ethereum Phishing Scams Detection
  428. HBRW: A Hardness-Based Re-Weighting Approach for Long-tailed Medical Image Classification
  429. HCLTS: Mining Customers' Consumption Patterns in Natural Gas Time Series with Hierarchical Contrastive Learning
  430. HCoTT: Hierarchical Chain-of-Thought Distillation
  431. HDA-GS: Hierarchical Density-Controlled for Anisotropic 3D Gaussian Splatting
  432. HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
  433. HDRec: Hierarchical Distillation for Enhanced LLM-based Recommendation Systems
  434. HFE-RWKV: High-Frequency Enhanced RWKV Model for Efficient Left Ventricle Segmentation in Pediatric Echocardiograms
  435. HFLR: Optimizing GNN Training via High-Fixed-Low-Resampling
  436. HFedPFS: Heterogeneous Federated Learning with Personalized Data Feature Sharing
  437. HGNet: Hash Generation Network Guided by High Frequency Information for Fine-Grained Image Retrieval
  438. HID-NAS: A Novel Neural Architecture Search Pipeline for High Information Density Data
  439. HLTCOE Submission to the VoicePrivacy Attacker Challenge
  440. HR-SKGs: Hyper-Relational Semantic Knowledge Graphs for Multi-hop Reading Comprehension
  441. HRTF Estimation using a Score-based Prior
  442. HYB-VITON: A Hybrid Approach to Virtual Try-On Combining Explicit and Implicit Warping
  443. HYMAN: Hybrid Memory and Attention Network for Unsupervised Anomaly Detection
  444. HamaraAwaz: Advancing Low-Latency Streaming TTS for Multilingual Speech in Indian Languages
  445. HandS3C: 3D Hand Mesh Reconstruction with State Space Spatial Channel Attention from RGB images
  446. Hard Sample Aware Robust Contrastive Learning for Multi-View Clustering
  447. Harmonizing for defect visibility with Fine-Grained Hierarchical Interaction Learning
  448. Harnessing Content and Structure in ID for Multimodal Recommendation
  449. Harnessing Contrastive Learning and Neural Transformation for Time Series Anomaly Detection
  450. Harnessing Dimensional Contrast and Information Compensation for Sentence Embedding Enhancement
  451. Harnessing Light Field Angular Cues and Spatial Geometries for Semantic Segmentation
  452. Harnessing the Potential of Omnidirectional UAVs in RIS-Enabled Wireless Networks
  453. Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model for Guiding End-to-End Speech Recognition
  454. HazeCLIP: Towards Language Guided Real-World Image Dehazing
  455. Hazy Remote Sensing Image Semantic Segmentation with Weak Annotations via Pre-training Optimization and Co-training
  456. Heart Sounds for High Blood Pressure Prediction
  457. Hedging Is Not All You Need: A Simple Baseline for Online Learning Under Haphazard Inputs
  458. Heterogeneous Data-based Cross-domain Few-shot Classification Method of Hyperspectral Image
  459. Heterogeneous Graph Convolutional Neural Networks for EEG-fNIRS Bimodal Emotion Recognition
  460. Heterogeneous Graph Dual-structure Optimization Based Attribute-aware for Recommendation
  461. Heterogeneous Packet Translation for Cross-Technology Communication
  462. HiE-VL: A Large Vision-Language Model with Hierarchical Adapter for Handwritten Mathematical Expression Recognition
  463. HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
  464. HiLiteMamba: A Lightweight and High-Frequency Aware Network for Single Image Super-Resolution
  465. HiRes: Hierarchical Feature Optimization and Rescorer for Automatic ICD Coding
  466. HieClip: Hierarchical CLIP with Explicit Alignment for Zero-Shot Anomaly Detection
  467. Hierarchical Bayesian Estimation of COVID-19 Reproduction Number
  468. Hierarchical Context Interaction and Reasoning with Transformer for Emotion Recognition
  469. Hierarchical Expectation Propagation for Semi-Blind Channel Estimation in Cell-Free Networks
  470. Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
  471. Hierarchical Loss for Bi-Level Classification of Speech into Language and Dialects
  472. Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction Generation
  473. Hierarchical Nash Equilibrium over Variational Equilibria via Fixed-point Set Expression of Quasi-nonexpansive Operator
  474. Hierarchical Perceptual Distillation Network for Lightweight Image Super-Resolution Reconstruction
  475. Hierarchical Prompt Tuning for System-Incremental Log Analysis
  476. Hierarchical Proxy Learning for Cloth-Changing Person Re-Identification
  477. Hierarchical Relation Distillation for Efficient 3D Visual Grounding
  478. Hierarchical Similarity Loss Enhanced Depth and Structural Fidelity in Monocular RGB-to-Depth Mapping with Adversarial Training
  479. Hierarchical Spatial-Temporal Enhancement Network For Continuous Sign Language Recognition
  480. Hierarchical Spatiotemporal Attention Network for Fine-grained Brain Cognitive State Recognition
  481. High-Efficiency Modulation Classification With Temporal-Frequency Analysis Based on Multi-channel Filter Bank
  482. High-Fidelity Editable Portrait Synthesis with 3D GAN Inversion
  483. High-Fidelity Music Vocoder using Neural Audio Codecs
  484. High-Fidelity Single-View Reconstruction of Indoor Scenes using 3D Shape Prior Template and Pixel-Aligned Deformation
  485. High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity Consistency
  486. High-Resolution Gait Micro-Doppler Synthesis from Videos Over Diverse Trajectories
  487. High-Resolution Speech Restoration with Latent Diffusion Model
  488. Higher-Order Topological Directionality and Directed Simplicial Neural Networks
  489. Homogeneous Graph Extraction: An Approach to Learning Heterogeneous Graph Embedding
  490. Hop-level Direct Preference Optimization for Knowledge Graph Reasoning with Trees
  491. How Machines Perceive Rooms - Regions of Relevance in Room Impulse Responses
  492. How Redundant Is the Transformer Stack in Speech Representation Models?
  493. How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios
  494. Human Action Recognition in Multi-Level Convolutional Temporal Attention Network
  495. Hybrid Coding and Weakly-Supervised Approach for Depth Estimation from Wrapped Phase
  496. Hybrid Content Caching Empowered By AIGC in Wireless Networks
  497. Hybrid Contrastive Learning Decoupling Speech Emotion Recognition
  498. Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification
  499. Hybrid Feature Fusion for Enhancing Medical Document Embedding
  500. Hybrid Feature Global Attention Network for Noisy-reverberant Speech Enhancement
  501. Hybrid Losses for Hierarchical Embedding Learning
  502. Hybrid Offline Passive Grammatical Inference and Online Planning for Non-Markovian Tasks
  503. Hybrid Precoding in mmWave Multiuser MIMO Systems with Delay Alignment Modulation (DAM)
  504. Hybrid Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
  505. Hybrid Spatial-Frequency Attention Network For Fine-Grained Skeleton-Based Action Recognition
  506. Hybrid predictive and parametric stereo coding for voice and audio communications
  507. HypCAD: Geometry-Enhanced Hyperbolic Contrastive Learning for CAD Model Retrieval
  508. Hyper-Refinement for Low-Rank Adaptation
  509. Hyper-adapter for Parameter-Efficient Multilingual ASR Adaptation
  510. HyperDiff: Masked Diffusion Model with High-efficient Transformer for Hyperspectral Image Cross-Scene Classification
  511. HyperKAN: Hypergraph Representation Learning with Kolmogorov-Arnold Networks
  512. HyperMST: Multi-scale Spatio-Temporal Hypercorrelation Network for POI Recommendation
  513. HyperSDT: HyperNetwork Slide Decision Tree for Interpretable Tabular Learning
  514. HyperSF: A Hypergraph Representation Learning Method Based on Structural Fusion
  515. HyperSMOTE: A Hypergraph-based Oversampling Approach for Imbalanced Node Classifications
  516. Hyperbolic Distance Based on EMD and Diffusion for Hyperspectral Imaging
  517. Hyperbolic Multimodal Knowledge Graph Embedding
  518. Hyperbolic PHATE: Visualizing Continuous Hierarchy of Latent Differentiation Structures
  519. Hyperedge Representations with Hypergraph Wavelets: Applications to Spatial Transcriptomics
  520. Hypergradient-free Training for Deep Equilibrium Models
  521. Hypergraph-Based Dynamic Graph Node Classification
  522. Hyperspectral Image Reconstruction with Unseen Material Detection
  523. Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
  524. I-KAN: Reconstructing Over-Range Inertial Signals
  525. ICAA-Mamba: Vision Mamba for Image Color Aesthetics Assessment
  526. ICIMG-Net: Inject Context Information to Motion Generation for Optical Flow Estimation
  527. ID-RWKV: Image Deraining RWKV
  528. IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs
  529. IEEE 802.11ad-Aided 5-D Sensing With a UAV Swarm in Urban Environment
  530. ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation
  531. INFR-GC: Interpretable Feature Representations for Granger Causality in Cortico-muscular Interactions
  532. INN-PAR: Invertible Neural Network for PPG to ABP Reconstruction
  533. INN-based Secure Steganography Using Lost Information as Adversarial Perturbations
  534. IOR: Inversed Objects Replay for Incremental Object Detection
  535. IOVS4NeRF: Incremental Optimal View Selection for Large-Scale NeRFs
  536. IPNet: Interpretable Prototype Network for Multi-Source Domain Adaptation
  537. IPP-Net: A Generalizable Deep Neural Network Model for Indoor Pathloss Radio Map Prediction
  538. ITMO language diarization and identification systems for the DISPLACE 2024 challenge
  539. ITW-DehazeFormer: Imaging through Turbid Water Using Improved DehazeFormer
  540. Identical Human Preference Alignment Paradigm for Text-to-Image Models
  541. Identical-Delay Based 2-D DOA and Frequency Joint Estimation With Sub-Nyquist Sampling for URA
  542. Identification and Correction of Permutation Errors in Compressed Sensing-Based Group Testing
  543. Identifying Adversarial Attacks in Crowdsourcing via Dense Subgraph Detection
  544. Identifying Bots on Social Media through Coordinated Group Perception
  545. Identifying and Mitigating Mismatched Language Code in Multilingual ASR
  546. Identity-Agnostic Learning for Deepfake Face Detection
  547. Identity-Preserving Audio-Driven Holistic Human Motion Video Generation
  548. Identity-Preserving Diffusion for Face Restoration
  549. Identity-aware Feature Decoupling Learning for Clothing-change Person Re-identification
  550. IdentityLock: An Identity-aware Backdoor strategy for Face Swapping Defense
  551. Image Compressive Sensing With Adaptive Sampling by Median Filtering
  552. Image-assisted Label Connective Completion for Vessel Segmentation with Insufficient Annotations
  553. ImageFlowNet: Forecasting Multiscale Image-Level Trajectories of Disease Progression with Irregularly-Sampled Longitudinal Medical Images
  554. Imitating Human Selective Attention Using Dual Policy Network for Scanpath Prediction
  555. ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model
  556. Impact of Glyph Information on Latent Space Diffusion Models for Accurate Handwritten Text Generation
  557. Impact of Temporal Precision on Speech Synthesis Accuracy From Electrocorticographic Brain Signals
  558. Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
  559. Imperceptible Adversarial Attacks on Point Clouds Guided by Point-to-Surface Field
  560. Imperceptible Transfer Attack on Large Vision-Language Models
  561. Implanting Robust Watermarks in Latent Diffusion Models for Video Generation
  562. Implementing Finite Impulse Response Filters on Quantum Computers
  563. Implicit Neural Representations with Fourier Kolmogorov-Arnold Networks
  564. Implicit and Explicit Rule Injection for Complex Query Answering over Knowledge Graphs
  565. Importance-Awareness Masking Network for Robust Document Retrieval
  566. Improved Bounds For Online Convex Optimization
  567. Improved Cross-Lingual Speaker Verification Using Speaker Sensitive Feature Guidance and Fine-grained Phonetic Information
  568. Improved Extrinsic Calibration of Acoustic Cameras via Batch Optimization
  569. Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
  570. Improved Image Classification with Manifold Neural Networks
  571. Improved Motion Plane Adaptive 360-Degree Video Compression Using Affine Motion Models
  572. Improved Out-of-domain Detection in VAE Latent Spaces with Boundary-driven Regularisation
  573. Improved Pitch and Voicing Determination Using the Reflected Root Chirp Group Delay Spectrum
  574. Improved Recognition of the Speech of People with Parkinson's Who Stutter
  575. Improved Techniques for Offline Reinforcement Learning: Advantage Value Estimation and Layernorm
  576. Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions
  577. Improving 5G Positioning Through Signal-to-Noise Ratio Recognition Training
  578. Improving Acoustic Scene Classification in Low-Resource Conditions
  579. Improving Adversarial Transferability through Channel-wise Scaling and Frequency-random Dropping
  580. Improving Compressive Imaging Recovery via Measurement Augmentation
  581. Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC Loss
  582. Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces
  583. Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
  584. Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASR
  585. Improving Embeddings by Refining Meanings for Temporal Knowledge Graph Link Predictions
  586. Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMs
  587. Improving GAN Performance Using Confidence-Aware Discrimination
  588. Improving Generated and Retrieved Knowledge Combination Through Zero-shot Generation
  589. Improving Height Prediction for Vision-Based Roadside 3D Object Detection
  590. Improving Irregular Text Recognition with Adaptive Feature Compression
  591. Improving Knowledge Base Question Answering via Retrieval Enhancement and Stepwise Reasoning
  592. Improving Knowledge Distillation via Cross-Modal Insights from CLIP
  593. Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
  594. Improving Micro-expression Recognition using Multi-sequence Driven Face Generation
  595. Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
  596. Improving Multimodal Human Pose Estimation by Adversarial Modality Enhancement†
  597. Improving Multimodal Large Language Models through Combining Resampler and MLP Projections
  598. Improving Open-Ended Referring Expression Comprehension via Dual-Language Constraints
  599. Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation Adjustment
  600. Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
  601. Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
  602. Improving Robustness of Post-hoc Calibration Against Common Corruptions By Learnable Augmentation
  603. Improving Sidescan Sonar Performance Using Array Upsampling Beamforming Synthetic Aperture
  604. Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
  605. Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
  606. Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
  607. Improvised Performance Following in Real Time for Automatic Accompaniment
  608. In Search of Optimal Pretraining Strategy for Robust Speaker Recognition
  609. In-Context Multitask Learning for Few-shot Fine-tuning of Large Language Models in Traditional Chinese Medicine Tongue Diagnosis
  610. Incorporate Global Information from Entire Datasets for Knowledge Tracing via Mini-Batch Input
  611. Incorporating Improved Sinusoidal Threshold-based Semi-supervised Method and Diffusion Models for Osteoporosis Diagnosis
  612. Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
  613. Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
  614. Individual Fairness for Fuzzy C-Means Clustering
  615. Indoor Airflow Imaging Using Physics-Informed Schlieren Tomography
  616. Indoor Sensing with Measurements
  617. Infant Cry Detection Using Causal Temporal Representation
  618. Inference Retrieval-Augmented Multi-Modal Chain-of-Thoughts Reasoning for Language Models
  619. Influence of Oropharyngeal Esophageal Cavity Geometry and Beak Angle on Vocal Tract Resonance of Birds using Computational Modeling
  620. Influence-Based Channel Reweighting for Multivariate Time Series Forecasting
  621. InfoHarmonizer Graph Contrastive Clustering
  622. InfoMin-based Query Embedding Optimization For Query-based Universal Sound Separation
  623. Information-Theoretic Minimax Regret Bounds for Reinforcement Learning based on Duality
  624. Infrared and Visible Image Fusion with Hierarchical Human Perception
  625. InjectTST: Injecting Global Information into Independent Channels for Long Time Series Forecasting
  626. Injecting Global Context for Multivariate Time Series Forecasting on Variable Subsets
  627. Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech Recognition
  628. Input Uncertainty Attribution by Uncertainty Propagation
  629. InsectMamba: State Space Model with Adaptive Composite Features for Insect Recognition
  630. Inside and Inside: Efficient Anomaly Detection by Fully Capturing the Detailed Dynamics
  631. InstAD: Instance-aware Segmentation Framework for Zero-shot Multi-instance Anomaly Detection
  632. Instance Segmentation of Airway Anatomies Using Mask R-CNN Prompt Adaptation-SAM
  633. Instance-wise Feature Acquisition with Classifier Selection Option for Structured Data Instances
  634. InstantSpeech: Instant Synchronous Text-to-Speech Synthesis for LLM-driven Voice Chatbots
  635. Instantaneous Trajectory Prediction via Latent Bidirectional Cooperative Diffusion
  636. Integrated Global-Local Gaussian Attention for Image Compression
  637. Integrated Interpolation and Matrix Completion for Radio Map Estimation: A Convex Optimization Approach
  638. Integrating Adaptive Sampling for Optimal Learned Video Compression
  639. Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
  640. Integrating Concept Associations for Query Focused Knowledge Summarization
  641. Integrating Failures in Robot Skill Acquisition with Offline Action-Sequence Diffusion RL
  642. Integrating Multi-Scale Compression Attention with Edge Detection for Ultrasound Tumor Segmentation
  643. Integrating Pause Information with Word Embeddings in Language Models for Alzheimer's Disease Detection from Spontaneous Speech
  644. Integrating Potential Pronunciations for Enhanced Mispronunciation Detection and Diagnosis Ability in LLMs
  645. Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection
  646. Intelligent Target Maneuverability in Presence of Tracking with Multiple Radars
  647. Intent-driven In-context Learning for Few-shot Dialogue State Tracking
  648. Inter- and Intra-Sentence Cuer-Invariant Representation Learning for Generalizable Cued Speech Recognition
  649. Inter-Frame Skip Coding Mode For Point Cloud Geometry Compression in Solid G-PCC
  650. Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos
  651. Interactive and Balanced Multimodal Learning via Cross Attention and Gradient Modulation for Compressed Video Action Recognition
  652. Interference-Resilient Hybrid Multi-Antenna ARQ
  653. Interpolation Filter Design for Sample Rate Independent Audio Effect RNNs
  654. Interpolation for Weight-Constrained Nested Arrays Having Non-Central ULA Segments in the Coarray
  655. Interpreting Deep Neural Network-Based Receiver Under Varying Signal-To-Noise Ratios
  656. Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
  657. Intra-modal Relation and Emotional Incongruity Learning using Graph Attention Networks for Multimodal Sarcasm Detection
  658. Intrusion Detection for Intelligent Transportation Systems: A lightweight interpretable model
  659. InvGS: a Novel Real-Time Inverse Rendering Framework Utilizing 3D Gaussian Splatting
  660. Invariant Model Learning on Local-Aware Wasserstein Geodesic for Domain Adaptation
  661. Investigating F0 Estimation in Speech Synthesis from Real-time MRI Articulatory Data
  662. Investigating Factors Related to the Naturalness of Synthesized Unison Singing
  663. Investigating Numerical Translation with Large Language Models
  664. Investigating Training Objectives for Generative Speech Enhancement
  665. Investigating the Sensitivity of Pre-trained Audio Embeddings to Common Effects
  666. Investigating voiced and unvoiced regions of speech for audio deepfake detection
  667. Investigation of Spatial Self-Supervised Learning and Its Application to Target Speaker Speech Recognition
  668. Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
  669. Investigation of perceptual music similarity focusing on each instrumental part
  670. Is It Still Fair? Investigating Gender Fairness in Cross-Corpus Speech Emotion Recognition
  671. Iterative Operator Sketching Framework for Large-Scale Imaging Inverse Problems
  672. JANE: Joint Angle Networks Assisting 3D Human Pose Estimation
  673. JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
  674. JSUnet: A New Hybrid U-shaped Network for Jamming Suppression
  675. Jack of All Trades, Master of None: PMP-Guided Adaptive Multi-Teacher Distillation with Meta-Learning
  676. Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
  677. Joint Beamforming Design for Multi-Functional RIS-Aided Over-the-Air Computation
  678. Joint Edge and Regional Depth Enhancement Network for Camouflaged Object Detection
  679. Joint Energy-Based Optimization of Binary Offloading Decisions and Communication Resources in TDMA Systems, via Dynamic Programming
  680. Joint Feature and Kernel Fusion for Improved Depth-Aware Panoptic Segmentation
  681. Joint Multi-Scale Contextual and Noise Suppression for Group Emotion Recognition
  682. Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration With Improved Intelligibility
  683. Joint Semantic Segmentation of Optical and SAR Image in Hazy Environments via Cross-modal Information Rectification and Cross-attention Fusion
  684. Joint Space-Time Adaptive Processing and Beamforming Design for Cell-Free ISAC Systems
  685. Joint Task Offloading and Routing in Wireless Multi-hop Networks Using Biased Backpressure Algorithm
  686. Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank Adaptation
  687. Joint-Wise Distributed Perception Graph Convolutional Network for Skeleton-Based Action Recognition
  688. JointSwinUNETR: an Efficient Feature-enhanced Architecture for Small Intestine Cine MRI Segmentation
  689. Jointly Optimal Array Geometries and Waveforms in Active Sensing: New Insights Into Array Design via the Cramér-Rao Bound
  690. Jointly Optimizing Data Discretization and Naive Bayes Classifier via Multi-Objective Optimization
  691. K-HashFed: Communication Efficient Federated Learning through Gradient Clustering and Hashing
  692. KABON: Knowledge Aggregation with Vision-Language Model for Black-Box Open-Set Domain Adaptation
  693. KAFQN: Kolmogorov-Arnold Fuzzy-guided Q-Network in Reinforcement Learning
  694. KAN v.s. MLP for Offline Reinforcement Learning
  695. KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head Generation
  696. KAN-HyperpointNet for Point Cloud Sequence-Based 3D Human Action Recognition
  697. KANGAN-AVSS: Kolmogorov-Arnold Network Based Generative Adversarial Networks for Audio-Visual Speech Synthesis
  698. KARLM: Enhancing LLM-based Recommendation Systems with Knowledge Bases
  699. KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification
  700. KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration
  701. KCE-Unet: A novel music denoising method with KANConv ECA Unet
  702. KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation
  703. KGD-GNN: A Knowledge-Guided Graph Neural Network for Myocardial Infarction Localization via 12-lead ECG
  704. KIKE: Linguistic Steganalysis Based on Knowledge Infusion and Knowledge Encoding
  705. KLFormer: Karhunen-Loève Transform for Robust 3D Human Pose Estimation
  706. KLMN: Knowledge distillation based lightweight multi-clue image forgery detection and localization
  707. KMG-LL: Knowledge-enhanced Multimodal Graph for Dialogue Generation
  708. KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models
  709. Keep what you need : extracting efficient subnetworks from large audio representation models
  710. Keeping Your Eyes on the Fingertip: A Two-Stage In-Air Handwritten Recognition Method
  711. Keeping the Balance: Anomaly Score Calculation for Domain Generalization
  712. Keeping the Best: The K-Best rule for Efficient Quickest Change Detection with Unknown Post-Change Distribution
  713. Kernel-Based Anomaly Detection Using Generalized Hyperbolic Processes
  714. Key Clues Guided Video Character Social Relationship Recognition Enhanced by LLM
  715. Keypoint Aware Masked Image Modelling
  716. Knocking on IP: Unveiling Websites through Cache-Aware Fingerprinting
  717. Know Your Heart Better: Multimodal Cardiac Output Monitoring using Earbuds
  718. Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR
  719. Knowledge Distillation From Ensemble for Spoken Language Identification
  720. Knowledge Distillation for Image Restoration : Simultaneous Learning from Degraded and Clean Images
  721. Knowledge Enhanced Multi-Domain Recommendations in an AI Assistant Application
  722. Knowledge Is Powerful: Art Knowledge-Driven Framework for Painting Style Classification Integrating Multimodal Knowledge
  723. Knowledge Transfer Across Modalities for Weakly Supervised Point Cloud Semantic Segmentation
  724. Knowledge-Enhanced Poetry-Image Synthesis with Large Language Model
  725. Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection
  726. Known-Plaintext Attacks to Thumbnail-Preservation Encryption Using Pix2pix Generative Adversarial Network
  727. Kronecker-structured Sparse Vector Recovery with Application to IRS-MIMO Channel Estimation
  728. L2 · M = C2 Large Language Models Are Covert Channels
  729. L2G: Head Gesture Animation Using an Emotion Guided Language Model
  730. L3D-Pose: Lifting Pose for 3D Avatars from a Single Camera in the Wild
  731. LABEL-SAM: A Semi-Automatic Interactive Annotation Model for Aortic Dissection Segmentation in 3D CTA Image
  732. LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
  733. LAVViT: Latent Audio-Visual Vision Transformers for Speaker Verification
  734. LBPE: Long-token-first Tokenization to Improve Large Language Models
  735. LCE: A Framework for Explainability of Ultrasound Image Based on Concept Discovery
  736. LCFed: An Efficient Clustered Federated Learning Framework for Heterogeneous Data
  737. LDG: Lightweight Deformable 3D Gaussians for Single-View Dynamic Scene Reconstruction
  738. LDGNet: LLMs Debate-Guided Network for Multimodal Sarcasm Detection
  739. LEF-TTS: Lightweight and Efficient End-to-End Text-to-Speech Synthesis With Multi-Stream Generator
  740. LEP: Leveraging Local Entropy Pruning for Sparsity in Large Language Models
  741. LFSRDiff: Light Field Image Super-Resolution via Diffusion Models
  742. LGNet: Linear Graph Representation for Efficient Cold-Start Recommendations
  743. LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
  744. LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
  745. LIMMITS'25: Multilingual Streaming TTS With Neural Codecs for Indian Languages
  746. LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
  747. LKA-ReID: Vehicle Re-Identification with Large Kernel Attention
  748. LKConvPose: A Pose Estimation Model with Large Receptive Field
  749. LKSNeXt: An Efficient Medical Image Segmentation Network with Large Kernels and Lightweight Structure
  750. LLDB: Efficient Low-Light Image Enhancement with Difffusion Bridge
  751. LLFA: Fusing Global Illumination and Local Priors for Low-Light Face Image Enhancement with Adaptor
  752. LLGS: Illuminating Gaussian Splatting via absorptance Modulation
  753. LLM based Text Generation for Improved Low-resource Speech Recognition Models
  754. LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
  755. LLM-Augmented Symbolic RL with Landmark-Based Task Decomposition
  756. LLM-GAN: Constructing Generative Adversarial Network Through Large Language Models for Explainable Fake News Detection
  757. LLM-Guided Dual-Branch Diffusion Model for Fine-Grained Motion Synthesis
  758. LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study
  759. LLMProto: A Hardware-Efficient Finetuning Model for Few-Shot Relation Extraction with Large Language Model
  760. LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
  761. LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
  762. LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
  763. LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with Efficient Narrow-Band and Cross-Band Attention
  764. LMTalker: Sparse Landmark-guided Gaussian Splatting for High-fidelity Talking Head Synthesis
  765. LNLFace: Enhanced Blind Face Restoration With Local and Non-local Lookups
  766. LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video Codec
  767. LOFI: Harnessing Attention Dynamics for Facial Expression Recognition with Noisy Labels
  768. LP-Gaussians: Learnable Parametric Gaussian Splatting for Efficient Dynamic Reconstruction of Single-View Scenes
  769. LPBS: A RL-PPO Driven K8S Batch Processing Task Scheduler
  770. LSTM-QGAN: Scalable NISQ Generative Adversarial Network
  771. LSU-NET: Lightweight Automatic Organs Segmentation Network for Medical Images
  772. LTOS: Layout-controllable Text-object Synthesis via Adaptive Cross-attention Fusions
  773. LV-ReID: Large Language-Vision Alignment Model for Text-based Person Re-identification
  774. LaTeXNet: A Specialized Model for Converting Visual Tables and Equations to LaTeX Code
  775. Label Dependency Aware Loss for Reliable Multi-Label Medical Image Classification
  776. Label Relationship Graph-Enhanced Class Hierarchy for Incremental Classification of Remote Sensing Images
  777. Label-constrained Unsupervised Domain Adaptation for Semantic Segmentation with Diffusion Models
  778. LagTS: Toward Adaptive Lag Relationship Modeling for Multivariate Time Series Forecasting
  779. Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
  780. Language-Queried Target Sound Extraction Without Parallel Training Data
  781. Language-based Audio Moment Retrieval
  782. Large Covariance Matrix Estimation for Groups of Highly Correlated Variables via Nonconvex Optimization
  783. Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
  784. Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
  785. Large Language Model-Empowered Adversarial Fusion for Typhoon Track Prediction
  786. Large Language Models Are Efficient Learners as Zero-Shot Speech Translators
  787. Large Language Models are Strong Audio-Visual Speech Recognition Learners
  788. Large Multimodal Model is a Better Comparator on Facial Beauty Prediction
  789. Large-Scale Recurrent Neural Networks with Fully Homomorphic Encryption for Privacy-Enhanced Speaker Identification
  790. Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
  791. Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
  792. Latent Representation Learning for Multimodal Brain Activity Translation
  793. Latent Space Score-based Diffusion Model for Probabilistic Multivariate Time Series Imputation
  794. Latent Watermarking of Audio Generative Models
  795. LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation
  796. Layer-Animate for Transparent Video Generation
  797. Lead Instrument Detection from Multitrack Music
  798. Learn from Balance: Rectifying Knowledge Transfer for Long-Tailed Scenarios
  799. Learned Approximated Optimization for Rapid Low-Complexity Hybrid Beamforming Design
  800. Learned ReLU-Based Soft Thresholding: A Data-Driven Method for Non-Negative Sparse Signal Recovery
  801. Learned Video Compression With Refined Adaptive Flow Pyramid And Coordinate-Aware Attention
  802. Learning Adaptive Spatial-temporal Structured Correlation Filters for UAV Object Tracking
  803. Learning Binary-Antithetical Information Bottleneck for Generalizable Face Anti-Spoofing
  804. Learning Class Prototypes for Visual Emotion Recognition
  805. Learning Class Unique Features in Fine-Grained Visual Classification
  806. Learning Control of Neural Sound Effects Synthesis from Physically Inspired Models
  807. Learning Deep Frequency Degradation Prior for Remote Sensing Spatio-temporal Fusion
  808. Learning Diffusion Model from Noisy Measurement using Principled Expectation-Maximization Method
  809. Learning Hierarchical Attribute Prompt for Vision-Language Models
  810. Learning Joint Appearance and Shape Co-Representations for Co-Saliency Detection
  811. Learning Markup Language Model for Composite Relationships Extraction
  812. Learning Music Audio Representations With Limited Data
  813. Learning Permutations in Monarch Factorization
  814. Learning Preconditioners in Gates-controlled Deep Unfolding Networks based on Quasi-Newton Methods For Accelerated MRI Reconstruction
  815. Learning Primitive Relations for Compositional Zero-Shot Learning
  816. Learning Rank Constrained Exposure Correction from Unpaired Data
  817. Learning Rate Optimization for Deep Neural Networks Using Lipschitz Bandits
  818. Learning Rich Speech Representations with Acoustic-Semantic Factorization
  819. Learning Semantic Facial Descriptors for Accurate Face Animation
  820. Learning Simultaneous Facial Canonical Correlation Representation for Face Hallucination
  821. Learning Source Disentanglement in Neural Audio Codec
  822. Learning Statistical and Physical Modeling for Consistency Human Motion Prediction
  823. Learning Strategy with Barlow Twins Objective for Emotion-Robust Speaker Verification System
  824. Learning Stroke-Order Dynamics in Few-Shot Font Generation via Sequential Awareness
  825. Learning Structured Compressed Sensing with Automatic Resource Allocation
  826. Learning Time-Varying Graphs from Data with Few Causes
  827. Learning Two-factor Representation for Magnetic Resonance Image Super-resolution
  828. Learning Weighted Least Squares Data Term for Poisson Image Deconvolution
  829. Learning a Sparse Polynomial Approximation to the Transition Function of General State-Space Models
  830. Learning from Ambiguous Data with Hard Labels
  831. Learning from Reconstruction: A Two-Stage Global-to-Local Framework for Temporal Knowledge Graph Completion
  832. Learning in the Model Space: Fault Diagnosis by Co-objective Learning in DynInt Model Space
  833. Learning to Follow Infrared Prior Repersentation for Image Dehazing
  834. Learning to Optimally Sample in MRI for Denoising-Driven Regularization
  835. Learning to Reconstruct Signals With Inexact Sensing Operator via Knowledge Distillation
  836. Learning with Coupled Noisy Labels for Visible-Infrared Person Re-identification via Graph Consistency
  837. Learning with Partial Labels from Conflict-Free and Semi-Supervised Perspective
  838. Learning-Aided Kalman Tracking in Biased Dynamic Systems: The Case of Cable-Driven Robots for Surgery
  839. Learning-Based Utility Estimation with Application to Speech Enhancement of a Moving Speaker
  840. Leave No Stone Unturned: Optimizing Subpattern Information Entropy for Coreset Selection
  841. Leave-One-EquiVariant: Alleviating Invariance-Related Information Loss in Contrastive Music Representations
  842. Lenna: Language Enhanced Reasoning Detection Assistant
  843. Less Is More: Embracing Sparsity and Interpolation with Esiformer for Time Series Forecasting
  844. Less Over More: Interference Sample Gradient Purification For Parallel Continual Learning
  845. Less Yet Robust: Crucial Region Selection for Scene Recognition
  846. Less is Enough: Relation Graph Guided Few-shot Learning for Multi-label Aspect Category Detection
  847. Less is more: Efficient Scene Graph Generation with reparameterization
  848. Let There Be Light: Robust Lensless Imaging Under External Illumination With Deep Learning
  849. Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
  850. Leveraging Boolean Directivity Embedding for Binaural Target Speaker Extraction
  851. Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
  852. Leveraging Heterophily in Spatial-Temporal Graphs for Multivariate Time-Series Forecasting
  853. Leveraging IPA and Articulatory Features as Effective Inductive Biases for Multilingual ASR Training
  854. Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
  855. Leveraging Mixture of Experts for Improved Speech Deepfake Detection
  856. Leveraging Multimodal Diffusion Models to Accelerate Imaging with Side Information
  857. Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification
  858. Leveraging Out-of-Domain Noise for Unsupervised Domain Adaptation in Speech Enhancement
  859. Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
  860. Leveraging Registers in Vision Transformers for Robust Adaptation
  861. Leveraging Self-Supervised Learning for Speaker Diarization
  862. Leveraging Visual Captions for Enhanced Zero-Shot HOI Detection
  863. LiDAR Light Scattering Augmentation (LISA): Physics-based Simulation of Adverse Weather Conditions for 3D Object Detection
  864. LiDAR-SPD: Improving Adversarial Robustness of 3D Object Detection via Spherical Projection and Diffusion
  865. LiRCDepth: Lightweight Radar-Camera Depth Estimation via Knowledge Distillation and Uncertainty Guidance
  866. LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
  867. Lightweight Clustered Federated Learning via Feature Extraction
  868. Lightweight Image Quality Prediction Guided by Perceptual Ranking Feedback
  869. Lightweight Multi-Frequency Enhancement Network for RGB-D Video Salient Object Detection
  870. Lightweight Self-Supervised Monocular Depth Estimation for All-Day Scenes Using Generative Adversarial Network
  871. Lightweight neural front-ends for low-resource on-device Text-to-Speech
  872. Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
  873. Linguistics-Vision Monotonic Consistent Network for Sign Language Production
  874. Linking Known and Unknown: Generalized Cross-Instance Feature Helps Category Discovery
  875. LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
  876. LipReading for Low-resource Languages by Language Dynamic LoRA
  877. LitePest: Real-Time and Efficient Detection of Agricultural Pests Using an Advanced Lightweight Deep Learning Network
  878. LkSFocalNets: Video Action Recognition With Large Kernel Selective Focal Networks
  879. LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
  880. LoRATEE: A Secure and Efficient Inference Framework for Multi-Tenant LoRA LLMs Based on TEE
  881. LoVA: Long-form Video-to-Audio Generation
  882. LocRef-Diffusion: Tuning-Free Layout and Appearance-Guided Generation
  883. Local Adaptive Time-Frequency Bidirectional Synchrosqueezing Transform
  884. Local Feature Alignment Prompt-Tuning for Few-shot Multimodal Aspect Sentiment Analysis
  885. Local Statistics for Generative Image Detection
  886. Localised Frequency Latent Domain Watermarking of DDIM Generated Images
  887. Locally Correctable Lattices
  888. LogSI: A Benchmark for System-Incremental Log Analysis
  889. Long-Range Multi-Scale Fusion for Efficient Single Image Super-Resolution
  890. Long-tailed Oracle Character Recognition Based on Convolutional Neural Networks and Vision Transformers
  891. Longitudinal Wrist PPG Analysis for Reliable Hypertension Risk Screening Using Deep Learning
  892. Look Before You Leap: Problem Elaboration Prompting Improves Mathematical Reasoning in Large Language Models
  893. Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction
  894. LossControl: Defending Membership Inference Attacks by Controlling the Loss
  895. Lossless Phase Conversion Method for Object Wave-based Hologram Compression
  896. Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications
  897. Low Complexity DoA-ToA Signature Estimation for Multi-Antenna Multi-Carrier Systems
  898. Low Complexity Rate Splitting Approach in RIS-Aided Systems Based on Channel Statistics
  899. Low Complexity Riemannian Coordinate-Descent over Symmetric Positive Definite Matrices
  900. Low Complexity Super Resolution for Resampling-based Video Coding
  901. Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
  902. Low Rank and Sparse Fourier Structure in Recurrent Networks Trained on Modular Addition
  903. Low-Complexity Cramér-Rao Lower Bound and Sum Rate Optimization in ISAC Systems
  904. Low-Complexity Neural Speech Dereverberation With Adaptive Target Control
  905. Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear Microphone
  906. Low-Correlation OFDM Waveform Design With Optimally Coded Sub-Carriers for the Joint Sensing and Communications
  907. Low-Light Detector Based on Feature Filtering and Enhancement
  908. Low-Rank Tensors for Multi-Dimensional Markov Models
  909. Low-Rank Transformer Adaptation for Arbitrary Style Transfer
  910. Low-Rank Tucker Decomposition of Multi-Subject Complex-Valued fMRI Data
  911. Low-Rate Modulo Folded ADC for Detecting Linearly Modulated Communication Symbols
  912. Low-Resolution Hierarchical Training for Efficient 3D Gaussian Splatting
  913. Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
  914. Low-rank Adaptation Method for Respiratory Sound Classification: A necessary road towards Large Models
  915. Low-shot Image Classification Using Mixture of Experts
  916. Lunar Tracking: A New Benchmark For Nighttime Tiny Object Tracking
  917. Lungmix: A Mixup-Based Strategy for Generalization in Respiratory Sound Classification
  918. M*: On-Chip Microfluidic Operations With A-Star for Portable Diagnostics
  919. M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
  920. M-MoE: Mixture of Mixture-of-Expert Model for CTC-based Streaming Multilingual ASR
  921. M2F2Net: Multi-stage Mixed Feature Fusion Network For Remote Sensing Change Detection
  922. M2PAIR: A High-Quality Acoustic Impulse Response Computation Model
  923. M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
  924. M3-CVC: Controllable Video Compression with Multimodal Generative Models
  925. M3ADD: A Novel Benchmark for Physiology Signal-based Automatic Depression Detection with Multimodal Multitask Multievent Framework
  926. M3Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential Recommendation
  927. MA-Det: A Discriminative Morphology-Aware Detector for Cervical Lesion Cell Clumps
  928. MACA: Multi-Anchor Classification Approach for Unsupervised Domain Adaptation
  929. MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion
  930. MAEM: A Multi-Aspect Extraction Model for Enhanced Embedding in RAG
  931. MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal Fusion
  932. MAID: Model Attribution via Inverse Diffusion
  933. MAITFuse: Multi-Dimension Adaptive Interaction Transform Network For Infrared-visible Image Fusion
  934. MAJoR: Visual Emotion Analysis via Multi-Attribute Joint Reasoning
  935. MAP Image Recovery with Guarantees using Locally Convex Multi-Scale Energy (LC-MUSE) Model
  936. MAP: Supporting Multimodal Knowledge Graph Completion via Augmented Modality Alignment and Instance Preserving
  937. MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
  938. MDDNet: Multilevel Difference-Enhanced Denoise Network for Unsupervised Change Detection in SAR Images
  939. MDN: Mamba-Driven Dualstream Network For Medical Hyperspectral Image Segmentation
  940. MDNet: Multi-Decoder Network for Abdominal CT Organs Segmentation
  941. MDRNet: Multi-Branch with Different Feature Representations Network for Motor Imagery Classification
  942. MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image Retrieval
  943. MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding Challenge
  944. META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
  945. MF-BERT: A Siamese Pre-training Framework for Motion Forecasting
  946. MFANet: Multi-Feature Aggregation Network for Multi-focus Image Fusion
  947. MFDPonzi: Detecting Ethereum Ponzi Schemes Using Static Features from Novel Opcode Sequences
  948. MFMamba: A Multimodal Fusion State Space Model for Depression Recognition
  949. MFT: Modal Fusion Transformer for Cross-Modal Fusion in 3D Object Detection
  950. MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
  951. MHGNet: Multi-Heterogeneous Graph Neural Network for Traffic Prediction
  952. MHSDB: A Comprehensive Benchmark for Multimodal Humor and Sarcasm Detection Leveraging Foundation Models
  953. MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword Spotting
  954. MIFAE-Forensics: Masked Image-Frequency AutoEncoder for DeepFake Detection
  955. MILE: Multi-Instance Learning for Document Event Argument Extraction
  956. MIMO Channel as a Neural Function: Implicit Neural Representations for Extreme CSI Compression
  957. MINR: Efficient Implicit Neural Representations for Multi-Image Encoding
  958. MKD-YOLO: Multi-Scale and Knowledge-Distilling YOLO for Efficient PPE Compliance Detection
  959. MLNet: Mutual Learning Network to Improve Self-Supervised Representation for Fine-Grained Visual Recognition
  960. MLSDET: Multi-LLM Statistical Deep Ensemble for Chinese AI-Generated Text Detection
  961. MLSwinTNet: A Multi-Level Feature Interaction Network for Low-Light Image Enhancement
  962. MM-LogVec: System Log Anomaly Detection Method Based on Multimodal Representation Learning
  963. MMA-Net: Multi-Modal Attention Network for 2-D Object Detection in Autonomous Driving
  964. MMCD: Memory-Based Multimodal Change Detection
  965. MMEditor: Multimodal Prompt-Driven 3D Gaussian Splatting Editing
  966. MMFN: Multi-Feature Multi-Modal Fusion Network for Diagnosis of Superficial Lymph Node Disease
  967. MMTP: Meta-learning-based Multi-Textual Prompt Tuning for Visual-Language Models
  968. MP-DPCC: A Motion Proxy-Based Dynamic Point Cloud Compression Framework
  969. MPAM-3DGS: Multi-Parametric Adversarial Manipulation for 3D Gaussian Splatting
  970. MPFL: A Decentralised Federated Learning Framework Based on Multi-Population Genetic Algorithm
  971. MPNAS: Multimodal Sentiment Analysis Pruning via Neural Architecture Search
  972. MPOT: Manifold Preserving Optimal Transport for Visual Recognition Under Severe Distribution Shift
  973. MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
  974. MQVAE: Capturing Metastable Dynamics from EEG for Brain-computer Interfaces
  975. MRANet: An Encoder-Decoder Network with Multi-Scale Residual Atrous-Spatial Pyramid Pooling for Seismic Phase Picking
  976. MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
  977. MS-RainMamba: Learning Multi-Scale State Space Models for Single Image Deraining
  978. MS-SCANet: A Multiscale Transformer-Based Architecture with Dual Attention for No-Reference Image Quality Assessment
  979. MS-UFAD: A Large-Scale Dataset for Real-world Unified Face Attack Detection with Text Descriptions
  980. MSA-ITEI: A Novel Method for Multimodal Analysis of Social Media Stickers
  981. MSACC: A Unified Multimodal Sentiment Analysis Framework for High Interpretability and Zero-shot Performance
  982. MSANet: Mixed Spectral and Attention Network for Robust 3D Human Pose Estimation
  983. MSE-based Sampling of Bandlimited Product Graph Signals via Joint Low-pass Impulse Responses
  984. MSECG: Incorporating Mamba for Robust and Efficient ECG Super-Resolution
  985. MSEMG: Surface Electromyography Denoising with a Mamba-based Efficient Network
  986. MSRFormer: Hybrid Scale Self-Attention and Local Fast Convolution Transformer for Facial Expression Recognition
  987. MST-HA: Multi-Modal Signal Fusion with Bayesian Optimization for Robust Industrial Robot Joint Health Assessment
  988. MSTBI: Head CT Detection and Prognostic Assessment of Traumatic Brain Injury Dataset
  989. MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
  990. MTE: Multi Transformation of Entities in Quaternion Vector Space for Temporal Knowledge Graph Completion
  991. MTMDC-GAN: Self-Attention Driven Multi-Scale Temporal Synthesis with Multi-Domain Analysis and Contrastive Learning
  992. MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
  993. MTTM: Memory-Augmented with Mamba for 3D Medical Images Analysis
  994. MULiving: Towards Real-time Multi-User Survival State Monitoring Using Wearable RFID Tags
  995. MUPO-Net: A Multilevel Dual-domain Progressive Enhancement Network with Embedded Attention for CT Metal Artifact Reduction
  996. MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation
  997. MVCBRec: Multi-View Contrastive Learning for Bundle Recommendation
  998. MVDC : A Multi-view Dental Completion Model Based on Contrastive Learning
  999. MX-Font++: Mixture of Heterogeneous Aggregation Experts for Few-shot Font Generation
  1000. MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion

Looking for submission deadlines instead? See the conference deadline calendar.