ICASSP 2025 Accepted Papers
The full list of 3,298 papers accepted at ICASSP 2025 (IEEE International Conference on Acoustics, Speech and Signal Processing). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
- Efficient Object Placement Via LLM and Diffusion Model
- Efficient Prototypical Classifier for Class-Incremental Learning
- Efficient Pruning for Large-Scale Seq2Seq Speech Models without Back-Propagation
- Efficient Quality Controllable Neural Image Compression based on QD-Model
- Efficient Quantization and Denoising Using Local Graph Fourier Frames
- Efficient Spatial Audio Rendering Via Differentiable FIR To IIR Estimation
- Efficient Streaming LLM for Speech Recognition
- Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression
- Efficient Visual Storytelling through Descriptive Words Distillation and Dynamic Decoding
- Efficient and Effective Model Extraction
- Efficient and Expandable Token-Level Approach for Multi-Domain Sensitive Information Classification
- Efficient-USR: Prompt Guided Dual-Domain Feature Information for Efficient Underwater Image Super-Resolution
- EfficientNet-Gaze: Integrating Multi-Scale Feature Extraction with Frequency Domain Analysis for Efficient Gaze Estimation
- EfficientSleepNet: A Novel Lightweight End-to-End Model for Automated Sleep Staging on Single-Channel EEG
- EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates
- Egocentric Speaker Diarization with Vision-Guided Clustering and Adaptive Speech Re-detection
- Electrocardiogram Report Generation and Question Answering via Retrieval-Augmented Self-Supervised Modeling
- Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition
- Embedding Enhanced MLP Enables Simple and Extensible Spatiotemporal Forecasting
- Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
- Emotion information recovery potential of wav2vec2 network fine-tuned for speech recognition task
- Emotion-Preserving Prosody Anonymization Network for Voice Privacy Protection
- Emotion-aware Structural Enhancement Graph Auto-Encoder for Rumor Detection
- Emotional Knowledge Self-Distillation in Dialogue
- Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
- Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
- Enabling DMG Wi-Fi Sensing in Data Transmission Intervals by Exploiting Beam Training Codebook
- End-to-end acoustic-articulatory dysarthric speech recognition leveraging large-scale pretrained acoustic features
- Energy Backdoor Attack to Deep Neural Networks
- Energy Consumption Trends in Sound Event Detection Systems
- Energy-based Model Guided Self-Supervised Learning for Speaker Verification
- Enhanced Breast Cancer Molecular Biomarker Classification: A Novel Two-Stage Machine Learning Pipeline for Accurate Histological Analysis of Whole Slide Images
- Enhanced Control for Diffusion Bridge in Image Restoration
- Enhanced Corneal Endothelial Cell Segmentation via Frequency-Selected Residual Fourier Diffusion Models
- Enhanced Loudspeaker Membrane Excursion Control Method Using Low Latency Distortion Prediction and Efficient LSTM Networks
- Enhanced Multimodal Depression Detection With Emotion Prompts
- Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph Propagation
- Enhanced Sparse Bayesian Learning Methods with Application to Massive MIMO Channel Estimation
- Enhanced Weakly Supervised Few-shot Classification & Segmentation
- Enhancing 3D Medical Image Understanding with 2D Multimodal Large Language Models
- Enhancing 6D Pose Estimation with Cross-modal Fusion Network and Density-peak Keypoint Localization
- Enhancing Age-Related Robustness in Children Speaker Verification
- Enhancing Autonomous Driving through Dual-Process Learning with Behavior and Reflection Integration
- Enhancing Autonomous Vehicle Planning With a Robust Fault-Tolerant Mechanism for Action-Induced Agent Detection
- Enhancing Boundary-Handling Strategies for Convolutional Sparse Representation Model with 46 × 46 Convolution-Multiplication Properties
- Enhancing Change Detection in Remote Sensing: Integrating Synthetic Data with Semi-Supervised Learning
- Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality Learning
- Enhancing Complex Formula Recognition with Hierarchical Detail-Focused Network
- Enhancing Continual Learning for Medical Imaging: Efficient Knowledge Transfer and Multi-Disease Prediction
- Enhancing Convolutional Models for Indoor Radio Mapping via Ray Marching
- Enhancing Cross-Domain Slot Filling with Joint LLM Data Generation and Data Curation
- Enhancing DETR Efficiency with Inter-Object Relationship and Semantic Spectral Decomposition-Based Distillation
- Enhancing Data-Free Class-Incremental Learning via Image-Centric Dual Distillation
- Enhancing Document-Level Relation Extraction through Entity-Pair-Level Interaction Modeling
- Enhancing EEG-based Covert Speech Decoding through Knowledge Transfer
- Enhancing Emotion Reasoning for Image Multi-Emotion Prediction
- Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
- Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
- Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
- Enhancing Extrapolation Reasoning on Temporal Knowledge Graphs with Logic Rules and Queries
- Enhancing Fairness in Gaussian Mixture Clustering through Impact Factor
- Enhancing Federated Domain Adaptation via Multi-Granular Fine-Grained Alignment
- Enhancing Federated Knowledge Distillation in Heterogeneous and Non-IID Scenarios
- Enhancing Few-Shot Out-of-Distribution Detection with Gradient Aligned Context Optimization
- Enhancing Generalized EEG Classification with Decomposed Statistics-diverse Feature Augmentation
- Enhancing Graph-based Fraud Detection by Adversarial Confidence Reweighting
- Enhancing Image Editing with Chain-of-Thought Reasoning and Multimodal Large Language Models
- Enhancing Image Generation Fidelity via Progressive Prompts
- Enhancing Imaging Generation through Implicit Neural Representations and HyperNetwork for Spatial Variability
- Enhancing Incomplete Multimodal Learning via Modal Complementary Recovering
- Enhancing Information Extraction with METORIE: A Metaphor and Trap-Based Dataset for Cross-Domain Fine-Tuning
- Enhancing Large Language Model Inference Efficiency via Lookahead Cache Filtering
- Enhancing Large Language Models on Domain-specific Tasks: A Novel Training Strategy via Domain Adaptation and Preference Alignment
- Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
- Enhancing Long-Term Capabilities of Large Language Models via Discourse Sub-graph Analysis
- Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
- Enhancing Multi-Channel Speech with Limited Microphones via Spherical Harmonic Transform
- Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
- Enhancing Multimodal Analogical Reasoning Through Triplet Interaction
- Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
- Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention
- Enhancing Multivariate Time Series Forecasting with Multi-scale Moving Transformation
- Enhancing Network Calibration for Low-Cost Gas Sensor Networks Through Adaptive Similarity Search
- Enhancing Out-of-Distribution Detection through Dynamic Activation Function
- Enhancing Precision in Image-Guided Spine Surgery through the Prediction of Occluded Fiducials Utilizing ResNet Architecture
- Enhancing Privacy in Radar-Based Vital Sign Monitoring Via Non-Linear FMCW Waveforms
- Enhancing Prosody Transfer in Speech Synthesis by Using Prosodically-Aligned References
- Enhancing Robustness of Implicit Neural Representations Against Weight Perturbations
- Enhancing Session-Based Recommendation with Hypergraph Motifs and Contrastive Learning
- Enhancing Small Model Performance in Educational Classification Tasks through Knowledge Distillation
- Enhancing Speech Emotion Recognition with Speech Dynamic Modeling and Multi-Modal Knowledge Distillation
- Enhancing Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
- Enhancing Stutter Detection using Long-Term Average Spectrum Values
- Enhancing TTS Stability in Hebrew using Discrete Semantic Units
- Enhancing Task-Specific Feature Learning with LLMs for Multimodal Emotion and Intent Joint Understanding
- Enhancing Teacher Classroom Behavior Descriptions: A Spatio-Temporal Graph-Based Method for Video Captioning
- Enhancing Text Annotation Through Rationale-Driven Collaborative Few-Shot Prompting
- Enhancing Time Series Prediction with Evolutionary Algorithm-based Optimization of LSTM
- Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes
- Enhancing Video-Text Matching via Sparse Stratified Sampling
- Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
- Enhancing Vision: Harmonizing Frequency for Imaging Quality and Perception Accuracy
- Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning
- Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
- Enhancing Zero-Shot Cross-Lingual Event Argument Extraction with Language-Independent Information
- Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration Prediction
- Enhancing Zero-Shot Relation Extraction through Staged Interaction with Large Language Models
- Enhancing the Robustness of LiDAR-based Object Detection under Disappearing Attacks
- Epigraph Based Multilevel Optimization (EMO) for Enhancing Chain-of-Thought Reasoning Capabilities
- EqGAN: Reformation-based Feature Equalization Fusion for Few-shot Image Generation
- Error Bounds Revisited, and How to Use Bayesian Statistics While Remaining a Frequentist
- Error Feedback Approach for Quantization Noise Reduction of Distributed Graph Filters
- Essentia: Boosting Artifact Removal from EEG through Semantic Guidance Utilizing Diffusion Model
- Estimating Instrument Spectral Response Functions Using Sparse Representations and Quadratic Envelopes
- Estimating Multi-chirp Parameters using Curvature-guided Langevin Monte Carlo
- Estimating Musical Surprisal in Audio
- Estimating the Number and Locations of Boundaries in Reverberant Environments with Deep Learning
- Estimation of Doppler, Range, and Direction of Targets in Wideband Bistatic Automotive Radar
- Estimation of Multi-Attribute Differential Graphs with Non-Convex Penalties
- EvaSR: Rethinking Efficient Visual Attention Design for Image Super-Resolution
- Evaluating Contrastive Methodologies for Music Representation Learning Using Playlist Data
- Evaluating Snippet Significance: A Framework for Audio and Text-Based Dialogue Summarization
- Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation
- Evaluating the Posterior Sampling Ability of Plug&Play Diffusion Methods in Sparse-View CT
- Evaluating the security of public surrogate watermark detectors
- Evaluation of Deep Audio Representations for Hearables
- Evaluation of Wearable Head BCG for PTT Measurement in Blood Pressure Intervention
- Event Masked Autoencoder: Point-wise Action Recognition with Event-Based Cameras
- Event-Driven Prony: Towards Asynchronous Spectral Estimation
- Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration
- EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking
- Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
- Evidential Deep Learning with Reweighted Margin Adjustment for Uncertainty-Driven Cervical OCT Image Diagnosis
- Evidential Neural GPLDA: A Novel Approach to Quantify Prediction Uncertainty in Speaker Verification Systems
- Evidential-TTS: High Fidelity Zero-Shot Text-to-Speech Using Evidential Deep Learning
- ExVC: Leveraging Mixture of Experts Models for Efficient Zero-shot Voice Conversion
- Exact Rotation Invariant Robust PCA
- Exact Solutions of the Inner Optimization Problem of Adversarial Robustness
- Explainable Adversarial Attacks on Coarse-to-Fine Classifiers
- Explainable Detection of Alzheimer's Disease Through Analysis of Human Behavior in Video
- Explainable Orthogonal Attention Networks for EEG-based Analysis: Leveraging Disentangled Representations to Enhance Diagnosis
- Explainable Reinforcement Learning for Trajectory Design in UAV-assisted Wireless Networks
- Explaining Representations in Correlation-based Deep Multiview Representation Learning
- Explaining Speaker and Spoof Embeddings via Probing
- Explicit Mutual Information Maximization for Self-Supervised Learning
- Explicit Spatial Hint and Implicit Logits Relation: Distilling Heterogeneous Knowledge From Vision Transformer to CNN
- Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
- Exploiting Attention-to-Motion via Transformer for Versatile Video Frame Interpolation
- Exploiting Beam-Split in IRS-aided Systems via OFDMA
- Exploiting Foundation Models for Label-Efficient Few-Shot Learning via Feature Coupling: A Case Study of cardiac CT Segmentation
- Exploiting Robust Model Watermarking Against the Model Fine-Tuning Attack via Flat Minima Aware Optimizers
- Exploiting Wavelet Scattering Transform & Squeeze-Excitation Blocks with Cross-Modal Attention for Multi-modal Emotion Recognition
- Exploiting the Relationship within the Unlabelled Samples by Set Matching for Generalized Category Discovery
- Exploration of Sequence-wise Optimized Parameters for Low Complexity Enhancement Video Coding (LCEVC) on 4K Content
- Explore the Hallucination on Low-level Perception for MLLMs
- Exploring Acoustic Foundations in Speech Production Assessment Models for Children with Cochlear Implants
- Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
- Exploring Acted Sleepy Speech to Advance Real-World Sleepiness Estimation and Cognitive Degradation Detection
- Exploring Antenna Placement Configurations with a Radar-based Silent Speech Interface
- Exploring Facial Kinship Verification through Contactless Heart Activity Analysis
- Exploring Generalization Boundaries of Unsupervised Industrial Anomaly Detection Models through Attribute Perturbation
- Exploring Graph-aware Reasoning and Bidirectional Selection for Vision-Language Navigation
- Exploring Group Theory for Optimal Cognitive Radar Waveform Design
- Exploring Inter-Variate and Long-Term Dependencies to Boost Multivariate Time Series Forecasting
- Exploring Kolmogorov-Arnold networks for realistic image sharpness assessment
- Exploring Large Language Models for Knowledge Graph Completion
- Exploring Local Interpretable Model-Agnostic Explanations for Speech Emotion Recognition with Distribution-Shift
- Exploring Meta Evidence for Prompt Optimization
- Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
- Exploring Simple Siamese Network for High-Resolution Video Quality Assessment
- Exploring Spectral Signatures of Chinese liquor using Machine Learning and SHapley Additive exPlanations
- Exploring Temporal Constraints for Unsupervised Iris Motion Tracking in AS-OCT Videos
- Exploring Text-Queried Sound Event Detection with Audio Source Separation
- Exploring Triple Knowledge Cues for Zero-Shot Human-Object Interaction Detection
- Exploring the Differences between Deaf and Hearing Infant Cries
- Exploring the Distribution of Cell Subpopulations in Pancreatic Ductal Adenocarcinoma Slides by Joint Spatial Transcriptomics and Pathology Data
- Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion
- Exploring the Interpretability of EEG-Inception Convolutional Neural Networks for Epilepsy Prediction
- Exploring the Robustness of In-Context Learning with Noisy Labels
- Exploring the Role of CLIP Global Visual Features in Multimodal Large Language Models
- Exponential Convergence of Stochastic Mirror Descent in Over-parameterized Linear Models
- Extending MPR for Locating a Moving Object Based on TDOA and FDOA
- Extending Whisper for Emotion Prediction Using Word-level Pseudo Labels
- Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
- Extracting Sparse Specialist Models from Generalist Models
- Extremum Encoding for Joint Baseband Signal Compression and Time-Delay Estimation for Distributed Systems
- Eye Movements as Images: A Multimodal Framework for Eye Movements Representation
- F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music Generation
- FA-GAN: Defense Against Adversarial Attacks in Automatic Modulation Recognition
- FABLE: A Bundle Method For Federated Learning In Wireless Systems
- FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
- FAF-Filt: Frequency-aware Fourier Filter for Sound Event Detection
- FARE: A Deep Learning-Based Framework for Radar-Based Face Recognition and Out-of-Distribution Detection
- FAST: Fast Audio Spectrogram Transformer
- FASTER: Face Attribute Sliders with Semantic Rewards
- FAWL: Weakly-Supervised Video Corpus Moment Retrieval with Frame-Wise Auxiliary Alignment and Weighted Contrastive Learning
- FBI-Net: Frequency Band Integration Network for Infrared Small Target Segmentation
- FBSE-FTFCWT-Based Novel Automated Framework for Dysarthric Speech Detection
- FCConDubber: Fine And Coarse Grained Prosody Alignment For Expressive Video Dubbing via Contrastive Audio-Motion Pretraining
- FCoDT-Net: A Novel Framework for High-Precision Medical Image Segmentation Using Contextual Distillation Transformer
- FDDSGCN: Fractional Decoupling Dynamic Spatiotemporal Graph Convolutional Network for Traffic Forecasting
- FDR Control for Complex-Valued Data with Application in Single Snapshot Multi-Source Detection and DOA Estimation
- FDR-Controlled Portfolio Optimization for Sparse Financial Index Tracking
- FEA-DETR: An Enhanced ConvNet for Detecting Prohibited Objects in X-Ray Images Using Frequency and Edge Aware Information
- FG3DFormer: Fine-Grained 3D Shape Classification Based on Vision Transformer
- FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection
- FKAN-GMFNet: Fourier Kolmogorov-Arnold-based Group Multi-scale Fusion Network for Aneurysm Image Segmentation
- FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing
- FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
- FR2ViT: Finetuning-free Token Reduction for Dense Prediction Through a Refinement-Reactivation Architecture
- FSENet: Frequency Separation Enhancement Network for Super-Resolution
- FUTGA-MIR: Enhancing Fine-grained and Temporally-aware Music Understanding with Music Information Retrieval
- FUVAS: Few-shot Unsupervised Video Anomaly Segmentation via Low-Rank Factorization of Spatio-Temporal Features
- Face Relighting with Ratio Function for Explicit Geometric Representation
- Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping
- Facial Expression Recognition with DToF Sensing
- Facilitating Semi-Supervised Pedestrian Detection with Structurally Controllable Instance Synthesis
- Factorized-VITS: Decoupling Prosody and Text in End-to-End Speech Synthesis without External or Secondary Aligner
- Fading-Invariant Adversarial Attacks on Neural Modulation Recognition
- Fair CoVariance Neural Networks
- Fair MP-BOOST: Fair and Interpretable Minipatch Boosting
- FairAdapter: Detecting AI-generated Images with Improved Fairness
- Faithful Self-Refinement in Mathematical Reasoning via Progressive Back-Translation
- FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
- Fast Adaptation of Pretrained Speaker Verification System for Source Speaker Tracking
- Fast DCT+: A Family of Fast Transforms Based on Rank-One Updates of the Path Graph
- Fast DPCNs for Feature Extraction without Labels
- Fast Sparse DFT Computation for Arbitrary Length by Circular Convolution
- Fast Sparse Learning from Streaming Data with LASSO
- Fast Structured Orthogonal Dictionary Learning using Householder Reflections
- Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
- Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
- Fast and Robust High Resolution Frequency Estimation of Damped Signals
- Fast inter-frame coding for dynamic meshes via supervoxel-based shape matching
- Faster Speech-LLaMA Inference with Multi-token Prediction
- FasterGold-DETR: An Efficient End-to-End Fire Detection Model via Gather-and-Distribute Mechanism
- Feature Disentangling Dual-stream Network for User Bias Alleviation in Social Media Prediction
- Feature Refinement Decomposition and Relation Preference Enhancement for Remote Sensing Change Detection
- FedCAda: Adaptive Client-Side Optimization for Accelerated and Stable Federated Learning
- FedDiT: Federated Learning by Distillation Token Enhanced Vision Transformer
- FedDiffRec: A Module-wise Training Approach for Diffusion-Based Recommendation in Federated Learning
- FedFLD: Heterogeneous Federated Learning via Forget-Less Distillation
- FedImp: Federated Learning Using Important Layers of Client Models for the Diagnosis of Breast Cancer Histopathology Images
- FedRPN: An Efficient Framework for Optimizing System Heterogeneity in Federated Learning
- FedSe: Group-Based Sequential Training Strategies for Mitigating Label Skew in Federated Learning
- FedTG: Text-guided Federated Domain Generalization
- FedTLU: Federated Learning with Targeted Layer Updates
- Federated Cross-Client Collaborative Filtering with Tensor Compressive Learning
- Federated Domain Generalization with Label Smoothing and Balanced Decentralized Training
- Federated Hybrid-Supervised Learning for Universal Medical Image Segmentation
- Federated Learning with Heterogeneous Feature Adaptation for Human Activity Recognition
- Federated Prototype Guided Adaption for Vision-Language Models
- Federated Smoothing ADMM for Quantile Regression with Non-Convex Sparse Penalties
- FeedbackFuzz: Fuzzing Processors via Intricate Program Generation with Feedback Engine
- Few-Shot Object Detection in Satellite Imagery with Feature Fusion Pyramid and Adaptive Region Proposal Networks
- Few-shot Image Classification based on Attribute Prediction and Selection
- Few-shot Keyword-incremental Learning Using Compositional Information
- FiTGAN: Content Fusion with Style Transformation for Few-shot Image Generation
- Filtering Resistant Large Language Model Watermarking via Style Injection
- Find Details in Long Videos: Tower-of-Thoughts and Self-Retrieval Augmented Generation for Video Understanding
- Fine-Grained Global Modeling Learning for Personalized Federated Sequential Recommender
- Fine-grained Vital Sign Reconstruction through Machine Learning on Multi-channel Radar Signals
- Fine-portraitist: Visualizing the Speaker's Face Portrait during Speech Listening
- Fine-tuning TitaNet-Large Model for Speaker Anonymization Attacker Systems
- Fioma: Towards Open-Set Semi-Supervised Specific Emitter Identification
- First-frame Supervised Video Polyp Segmentation via Propagative and Semantic Dual-teacher Network
- First-order State Space Model for Lightweight Image Super-resolution
- Flare-Aware RWKV for Flare Removal
- FlashSR: One-step Versatile Audio Super-resolution via Diffusion Distillation
- Flexible Event-Driven Biological Imaging via Bayesian Inference
- Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching for Speaker Diarization
- FlowMAC: Conditional Flow Matching for Audio Coding at Low Bit Rates
- FlowSE: Flow Matching-based Speech Enhancement
- FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
- Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
- Fooling the Forgers: A Multi-Stage Framework for Audio Deepfake Detection
- Foreground-aware Prototypical Network for Prohibited Item Detection from X-ray Scans
- ForensiCam-215K: A Large Scale Image and Video Dataset for Forensic Analysis
- Forensics Analysis of Residual Noise Texture in digital Images for Detection of Deepfake
- Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
- Found In The Distribution: Utilizing Latent Dirichlet Allocation Improves Long Context Comprehension of Large Language Models
- Foundation Model and Temporal Priors-guided Transductive Few-shot Action Recognition
- Foundation Models Boost Low-Level Perceptual Similarity Metrics
- Fourth-Order Cumulant Based 3-D Near-Field Underdetermined Parameter Estimation With Exact Spatial Propagation Model
- Fractional-Order Hyperbolic Tangent Based Adaptive Algorithm for Feedback Control in Hearing Aids
- Frank-Wolfe Method with Proximal Regularization for Constrained Federated Learning with Non-iid Data
- FreeAlign: Superior Text-Image Alignment by Modulating Prompt Attention
- FreeLesion: Synthetic Image-Mask Pairs for Fundus Lesion Segmentation via Curriculum Learning and Feature-Loss Guided Filtering
- FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
- FreeSegDiff: Annotation-free Saliency Segmentation with Diffusion Models
- Freeze and Learn: Continual Learning with Selective Freezing for Speech Deepfake Detection
- FreqSense: Universal and Low-Latency Adversarial Example Detection for Speaker Recognition with Interpretability in Frequency Domain
- Frequency Agnostic Tissue Characterization in Ultrasound Imaging using Backscattered Signal Statistics
- Frequency Domain Information Integrated Network for Low-Light Image Enhancement
- Frequency-Based Federated Domain Generalization for Polyp Segmentation
- Frequency-Domain Guided Multiple Parallel Kernels Network for Low-Light Remote Sensing Image Enhancement
- Frequency-Domain Popularity Forecasting with Shape-Based Retrieval
- Frequency-Space Margin Perception for Open Set Knowledge Distillation
- Frequency-enhanced Comprehensive Dependency Attention for Time Series Anomaly Detection
- Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
- From Characters to Subwords: Modeling Unit Conversion for Low-resource Speech Recognition
- From Pixels to Voice: A Simple and Efficient End-to-End Spoken Image Description Approach via Vision Codec Language Models
- From Voices to Beats: Enhancing Music Deepfake Detection by Identifying Forgeries in Background
- FruitMMBench: A Multi-modal Benchmark for Fruit Quality Assessment
- Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
- Full-Reference Point Cloud Quality Assessment with Multimodal Large Language Models
- Full-text Error Correction for Chinese Speech Recognition with Large Language Model
- Fully Connected Tensor Network based Brain Structural Feature Extraction for Early Alzheimer's Disease Detection
- Fully Spiking Neural Network for Legged Robots
- Functional Near-Infrared Spectroscopy Feature Extraction with Application in Workload Estimation
- Fundamental Social Learning Scaling Law for Tracking Hidden Markov Models
- Fusing Multimodality of Large Language Models and Satellite Imagery via Simplicial Contrastive Learning for Latent Urban Feature Identification and Environmental Application
- Fusion of Information in Multiple Particle Filtering in the Presence of Unknown Static Parameters
- Fusion-OSR: Cross-Domain Contrastive Learning with Weibull Calibration for Time Series Open Set Recognition
- FusionClassNet: A Multi-Scale Feature Fusion Network with Contrastive Loss-Driven Classification for Enhanced Lung Tumor Image Representations
- FuzzyMIL: Decoupling Pathological Phenotypes through Deep Fuzzy Clustering for Efficient Whole Slide Image Analysis
- G-Depth: An Efficient Graph Method for Robust Depth Completion
- GADACE: Graph Anomaly Detection Combining Attribute Contrast and Structure Reconstruction
- GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
- GATOmics: A Novel Multi-Omics Graph Attention Network Model for Cancer Driver Gene Detection
- GBA-Net: A Method for 3D Brain Tumor Segmentation Based on Multi-scale Gaussian Boundary Attention
- GCAT: Gated Convolutional Attention Transformer for Efficient Image Super-Resolution
- GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning
- GDDA: Semantic OOD Detection on Graphs under Covariate Shift via Score-Based Diffusion Models
- GDFDNet: A Novel Graph-Based Dynamically Fused Dual-Stream Network for Accuracy Prohibited Items Detection
- GDRIVE: Adaptive Object Detection in Autonomous Vehicles via Graph-Based Feature Learning
- GEE Maximization in UAV-Aided Mobile IoT Networks Using Deep Reinforcement Learning
- GEGA: Graph Convolutional Networks and Evidence Retrieval Guided Attention for Enhanced Document-level Relation Extraction
- GEMD-UNet: Graph Structure Enhanced Multi-dimensional Learning Unet for Cloud Detection
- GENIE: Socially Unbiased Generative Text-to-Image Editing
- GIST: Guided Interpretable Large Language Model Strategy Transfer for Multi-Task Reinforcement Learning
- GLST-GCN: Global-Local Spatio-Temporal Graph Convolutional Network for Skeleton-based Hand Motion Prediction
- GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation
- GMCL: Graph-Enhanced Multimodal Contrastive Learning for Rumor Detection
- GMM-Based Bootstrap Prototype-Aware Learning For Weakly Supervised Semantic Segmentation
- GMMCL: Adaptive Concept Drift in Data Streams with Gaussian Mixture Models based on Contrastive Learning
- GNCL: A Graph Neural Network with Consistency Loss for Segment-Level Spoofed Speech Detection
- GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision Tasks
- GPPT: Gaussian Process-infused Prompt Tuning for Vision-language Models
- GPT-C: Generative PrompT Compression
- GPT-LAD: Leveraging Large Multimodal Models for Logical Anomaly Detection
- GRACED: A Plug-and-Play Solution for Certifiable Graph Classification
- GREST: Ghost Targets Removal Algorithm Using Multipath Angle Estimation
- GS-PT: Exploiting 3D Gaussian Splatting for Comprehensive Point Cloud Understanding via Self-supervised Learning
- GSMM: Efficient Global Sparsification for Resource-Conscious Multimodal Models
- GateM2Former: Gated Feature Selection and Expert Modeling in Multimodal Emotion Recognition
- Gated Cross-Attention Network for Depth Completion
- Gaussian Constrained Diffeomorphic Deformation Network for Panoramic Semantic Segmentation
- Gaussian Difference: Find Any Change Instance in 3D Scenes
- Gaussian-Face: Talking Head Generation with Hybrid Density via 3D Gaussian Splatting
- GaussianEnhancer: A General Rendering Enhancer for Gaussian Splatting
- GaussianSlicer: Efficient Surface Reconstruction from Cross-sectional Slices with Gaussian Splatting
- Gaze-Assisted Human-Centric Domain Adaptation for Cardiac Ultrasound Image Segmentation
- Gaze-GZ: Generalized Gaze Estimation with Multi-scale Gaze Zone Prediction
- GeMIMO: Searching the Cores of X-formers for Time Series Forecasting
- Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
- General Dynamic Regularization Federated Learning with Hybrid Sharpness-Aware Minimization
- Generalizable Articulated Object Perception with Superpoints
- Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
- Generalizable Indoor Path Loss Prediction
- Generalizable Real-time Accelerated Dynamic MRI
- Generalization Guarantee of Decentralized Learning with Heterogeneous Data
- Generalize Audio Deepfake Algorithm Recognition via Attribution Enhancement
- Generalized Approximate Message-Passing for Compressed Sensing with Sublinear Sparsity
- Generalized Graph Signal Reconstruction via the Uncertainty Principle
- Generalized Linear Models with 1-Bit Measurements: Asymptotics of the Maximum Likelihood Estimator
- Generate E-commerce Product Background by Integrating Category Commonality and Personalized Style
- Generating Apoptosis-Inducing Anticancer Peptides Targeting BCL-xL Using Latent Diffusion Models on Small Datasets
- Generating Customized 4D Motions from Text Inputs Using Spatial-Temporal Slicing Approaches
- Generating Editable Head Avatars with 3D Gaussian GANs
- Generating Gezi Opera Scores with a Large Language Model and a High-Quality Dataset
- Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
- Generating Targeted Universal Adversarial Perturbation against Automatic Speech Recognition via Phoneme Tailoring
- Generating Vocals from Lyrics and Musical Accompaniment
- Generative Adversarial Network with Adaptive Synthesis for Brain-Computer Interfaces in Motor Imagery Classification
- Generative Adversarial Network with Structured Semantic Prompts Constrainting Clip for Text-to-Image
- Generative Dataset Distillation Based on Self-knowledge Distillation
- Generative Diffusion Model-based Energy Management in Networked Energy Systems
- Generative Expansion of Small Datasets: An Expansive Graph Approach
- Generative Model based Optical Response Prediction for Plasmonic Sensing
- Generative Sensing: Pre-training LiDAR with Masked Autoencoders for Ultra-Frugal Perception
- Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
- Geodesic Mean Threshold Scheme on Riemannian Manifold for EEG Signal Classification
- Geometric Feature-Driven Metric Learning for 3D Craniofacial Superimposition
- Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
- Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
- Global Context MambaVision for EEG-based Emotion Recognition
- Global Enhanced Frame Prompt Tuning for Sound Event Detection
- Global Static Pruning via Adaptive Sample Complexity Awareness
- Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model
- Globally Normalizing the Transducer for Streaming Speech Recognition
- GoLoColor: Towards Global-Local Semantic Aware Image Colorization
- GraFPrint: A GNN-Based Approach for Audio Identification
- GradPFL: Gradient-Driven Adaptive Clustering in Personalized Federated Learning
- Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
- Gradient-Oriented Clustered Federated Learning With Efficient Knowledge Sharing in Non-IID Settings
- Gram: A Large-Scale General EEG Model for Raw Data Classification and Restoration Tasks
- Granularity-Aware Contrastive Learning for Fine-Grained Action Recognition
- Graph Anomaly Detection via Multi-Scale Reconstruction of Graph Encoder-Decoder Networks
- Graph Contrastive Learning with Decoupled Augmentation
- Graph Embedded Stochastic Configuration Networks for Imbalanced Data Classification
- Graph Learning with Low-rank and Diagonal Structures: A Riemannian Geometric Approach
- Graph Neural Networks Meet Probabilistic Graphical Models: A Survey
- Graph Neural Networks for Parkinson's Disease Detection
- Graph Pooling via Dropping Task-Irrelevant Nodes
- Graph Refinement in Latent Space: A Hypergraph Convolution for Underwater Object Detection
- Graph Signal Reconstruction via Koopman Autoencoder
- Graph Structure Learning via Transfer Entropy for Multivariate Time Series Anomaly Detection
- Graph Topology Identification Based on Covariance Matching
- Graph-Driven Insights: Enhancing Stock Market Prediction with Relational Temporal Dynamics
- Graph-Enhanced Dual-Stream Feature Fusion with Pre-Trained Model for Acoustic Traffic Monitoring
- Graph-based Signal Sampling with Adaptive Subspace Reconstruction for Spatially-irregular Sensor Data
- GraphDAE-PU: Graph Denosing Auto-Encoder for Arbitrary-Scale Point Cloud Upsampling
- GraphVCM: Virtual Center Mixing with Distance-Aware Regulation for Class Imbalanced Node Classification
- Grey Wolf Optimizer Algorithm Based Active Noise Control Without Secondary Path Identification
- Group Modeling and Recommendation Based on Multi-Behavior Interactions in Live Streaming E-Commerce
- Group-CLIP Uncertainty Modeling for Group Re-Identification
- Group-wise Semantic-enhanced Interaction Network for Remote Sensing Spatio-Temporal Fusion
- Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker Recognition
- Grouping-Based Crowding Differential Evolution Approaches for Multimodal Feature Selection
- Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion Models
- Guided Speaker Embedding
- Guiding Inter-domain Class Balancing With Salient Features For Domain Adaptive Object Detection
- Guitar-TECHS: An Electric Guitar Dataset Covering Techniques, Musical Excerpts, Chords and Scales Using a Diverse Array of Hardware
- HANet: A Harmonic Attention-Based Network for Singing Melody Extraction from Polyphonic Music
- HAPG-SAQAM: Human Auditory Perception Guided Spatial Audio Quality Assessment Metric
- HATTM: A Novel Hybrid Attention Model for Ethereum Phishing Scams Detection
- HBRW: A Hardness-Based Re-Weighting Approach for Long-tailed Medical Image Classification
- HCLTS: Mining Customers' Consumption Patterns in Natural Gas Time Series with Hierarchical Contrastive Learning
- HCoTT: Hierarchical Chain-of-Thought Distillation
- HDA-GS: Hierarchical Density-Controlled for Anisotropic 3D Gaussian Splatting
- HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
- HDRec: Hierarchical Distillation for Enhanced LLM-based Recommendation Systems
- HFE-RWKV: High-Frequency Enhanced RWKV Model for Efficient Left Ventricle Segmentation in Pediatric Echocardiograms
- HFLR: Optimizing GNN Training via High-Fixed-Low-Resampling
- HFedPFS: Heterogeneous Federated Learning with Personalized Data Feature Sharing
- HGNet: Hash Generation Network Guided by High Frequency Information for Fine-Grained Image Retrieval
- HID-NAS: A Novel Neural Architecture Search Pipeline for High Information Density Data
- HLTCOE Submission to the VoicePrivacy Attacker Challenge
- HR-SKGs: Hyper-Relational Semantic Knowledge Graphs for Multi-hop Reading Comprehension
- HRTF Estimation using a Score-based Prior
- HYB-VITON: A Hybrid Approach to Virtual Try-On Combining Explicit and Implicit Warping
- HYMAN: Hybrid Memory and Attention Network for Unsupervised Anomaly Detection
- HamaraAwaz: Advancing Low-Latency Streaming TTS for Multilingual Speech in Indian Languages
- HandS3C: 3D Hand Mesh Reconstruction with State Space Spatial Channel Attention from RGB images
- Hard Sample Aware Robust Contrastive Learning for Multi-View Clustering
- Harmonizing for defect visibility with Fine-Grained Hierarchical Interaction Learning
- Harnessing Content and Structure in ID for Multimodal Recommendation
- Harnessing Contrastive Learning and Neural Transformation for Time Series Anomaly Detection
- Harnessing Dimensional Contrast and Information Compensation for Sentence Embedding Enhancement
- Harnessing Light Field Angular Cues and Spatial Geometries for Semantic Segmentation
- Harnessing the Potential of Omnidirectional UAVs in RIS-Enabled Wireless Networks
- Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model for Guiding End-to-End Speech Recognition
- HazeCLIP: Towards Language Guided Real-World Image Dehazing
- Hazy Remote Sensing Image Semantic Segmentation with Weak Annotations via Pre-training Optimization and Co-training
- Heart Sounds for High Blood Pressure Prediction
- Hedging Is Not All You Need: A Simple Baseline for Online Learning Under Haphazard Inputs
- Heterogeneous Data-based Cross-domain Few-shot Classification Method of Hyperspectral Image
- Heterogeneous Graph Convolutional Neural Networks for EEG-fNIRS Bimodal Emotion Recognition
- Heterogeneous Graph Dual-structure Optimization Based Attribute-aware for Recommendation
- Heterogeneous Packet Translation for Cross-Technology Communication
- HiE-VL: A Large Vision-Language Model with Hierarchical Adapter for Handwritten Mathematical Expression Recognition
- HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
- HiLiteMamba: A Lightweight and High-Frequency Aware Network for Single Image Super-Resolution
- HiRes: Hierarchical Feature Optimization and Rescorer for Automatic ICD Coding
- HieClip: Hierarchical CLIP with Explicit Alignment for Zero-Shot Anomaly Detection
- Hierarchical Bayesian Estimation of COVID-19 Reproduction Number
- Hierarchical Context Interaction and Reasoning with Transformer for Emotion Recognition
- Hierarchical Expectation Propagation for Semi-Blind Channel Estimation in Cell-Free Networks
- Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
- Hierarchical Loss for Bi-Level Classification of Speech into Language and Dialects
- Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction Generation
- Hierarchical Nash Equilibrium over Variational Equilibria via Fixed-point Set Expression of Quasi-nonexpansive Operator
- Hierarchical Perceptual Distillation Network for Lightweight Image Super-Resolution Reconstruction
- Hierarchical Prompt Tuning for System-Incremental Log Analysis
- Hierarchical Proxy Learning for Cloth-Changing Person Re-Identification
- Hierarchical Relation Distillation for Efficient 3D Visual Grounding
- Hierarchical Similarity Loss Enhanced Depth and Structural Fidelity in Monocular RGB-to-Depth Mapping with Adversarial Training
- Hierarchical Spatial-Temporal Enhancement Network For Continuous Sign Language Recognition
- Hierarchical Spatiotemporal Attention Network for Fine-grained Brain Cognitive State Recognition
- High-Efficiency Modulation Classification With Temporal-Frequency Analysis Based on Multi-channel Filter Bank
- High-Fidelity Editable Portrait Synthesis with 3D GAN Inversion
- High-Fidelity Music Vocoder using Neural Audio Codecs
- High-Fidelity Single-View Reconstruction of Indoor Scenes using 3D Shape Prior Template and Pixel-Aligned Deformation
- High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity Consistency
- High-Resolution Gait Micro-Doppler Synthesis from Videos Over Diverse Trajectories
- High-Resolution Speech Restoration with Latent Diffusion Model
- Higher-Order Topological Directionality and Directed Simplicial Neural Networks
- Homogeneous Graph Extraction: An Approach to Learning Heterogeneous Graph Embedding
- Hop-level Direct Preference Optimization for Knowledge Graph Reasoning with Trees
- How Machines Perceive Rooms - Regions of Relevance in Room Impulse Responses
- How Redundant Is the Transformer Stack in Speech Representation Models?
- How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios
- Human Action Recognition in Multi-Level Convolutional Temporal Attention Network
- Hybrid Coding and Weakly-Supervised Approach for Depth Estimation from Wrapped Phase
- Hybrid Content Caching Empowered By AIGC in Wireless Networks
- Hybrid Contrastive Learning Decoupling Speech Emotion Recognition
- Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification
- Hybrid Feature Fusion for Enhancing Medical Document Embedding
- Hybrid Feature Global Attention Network for Noisy-reverberant Speech Enhancement
- Hybrid Losses for Hierarchical Embedding Learning
- Hybrid Offline Passive Grammatical Inference and Online Planning for Non-Markovian Tasks
- Hybrid Precoding in mmWave Multiuser MIMO Systems with Delay Alignment Modulation (DAM)
- Hybrid Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
- Hybrid Spatial-Frequency Attention Network For Fine-Grained Skeleton-Based Action Recognition
- Hybrid predictive and parametric stereo coding for voice and audio communications
- HypCAD: Geometry-Enhanced Hyperbolic Contrastive Learning for CAD Model Retrieval
- Hyper-Refinement for Low-Rank Adaptation
- Hyper-adapter for Parameter-Efficient Multilingual ASR Adaptation
- HyperDiff: Masked Diffusion Model with High-efficient Transformer for Hyperspectral Image Cross-Scene Classification
- HyperKAN: Hypergraph Representation Learning with Kolmogorov-Arnold Networks
- HyperMST: Multi-scale Spatio-Temporal Hypercorrelation Network for POI Recommendation
- HyperSDT: HyperNetwork Slide Decision Tree for Interpretable Tabular Learning
- HyperSF: A Hypergraph Representation Learning Method Based on Structural Fusion
- HyperSMOTE: A Hypergraph-based Oversampling Approach for Imbalanced Node Classifications
- Hyperbolic Distance Based on EMD and Diffusion for Hyperspectral Imaging
- Hyperbolic Multimodal Knowledge Graph Embedding
- Hyperbolic PHATE: Visualizing Continuous Hierarchy of Latent Differentiation Structures
- Hyperedge Representations with Hypergraph Wavelets: Applications to Spatial Transcriptomics
- Hypergradient-free Training for Deep Equilibrium Models
- Hypergraph-Based Dynamic Graph Node Classification
- Hyperspectral Image Reconstruction with Unseen Material Detection
- Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
- I-KAN: Reconstructing Over-Range Inertial Signals
- ICAA-Mamba: Vision Mamba for Image Color Aesthetics Assessment
- ICIMG-Net: Inject Context Information to Motion Generation for Optical Flow Estimation
- ID-RWKV: Image Deraining RWKV
- IDE: A Multi-Agent-Driven Iterative Framework for Dynamic Evaluation of LLMs
- IEEE 802.11ad-Aided 5-D Sensing With a UAV Swarm in Urban Environment
- ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation
- INFR-GC: Interpretable Feature Representations for Granger Causality in Cortico-muscular Interactions
- INN-PAR: Invertible Neural Network for PPG to ABP Reconstruction
- INN-based Secure Steganography Using Lost Information as Adversarial Perturbations
- IOR: Inversed Objects Replay for Incremental Object Detection
- IOVS4NeRF: Incremental Optimal View Selection for Large-Scale NeRFs
- IPNet: Interpretable Prototype Network for Multi-Source Domain Adaptation
- IPP-Net: A Generalizable Deep Neural Network Model for Indoor Pathloss Radio Map Prediction
- ITMO language diarization and identification systems for the DISPLACE 2024 challenge
- ITW-DehazeFormer: Imaging through Turbid Water Using Improved DehazeFormer
- Identical Human Preference Alignment Paradigm for Text-to-Image Models
- Identical-Delay Based 2-D DOA and Frequency Joint Estimation With Sub-Nyquist Sampling for URA
- Identification and Correction of Permutation Errors in Compressed Sensing-Based Group Testing
- Identifying Adversarial Attacks in Crowdsourcing via Dense Subgraph Detection
- Identifying Bots on Social Media through Coordinated Group Perception
- Identifying and Mitigating Mismatched Language Code in Multilingual ASR
- Identity-Agnostic Learning for Deepfake Face Detection
- Identity-Preserving Audio-Driven Holistic Human Motion Video Generation
- Identity-Preserving Diffusion for Face Restoration
- Identity-aware Feature Decoupling Learning for Clothing-change Person Re-identification
- IdentityLock: An Identity-aware Backdoor strategy for Face Swapping Defense
- Image Compressive Sensing With Adaptive Sampling by Median Filtering
- Image-assisted Label Connective Completion for Vessel Segmentation with Insufficient Annotations
- ImageFlowNet: Forecasting Multiscale Image-Level Trajectories of Disease Progression with Irregularly-Sampled Longitudinal Medical Images
- Imitating Human Selective Attention Using Dual Policy Network for Scanpath Prediction
- ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model
- Impact of Glyph Information on Latent Space Diffusion Models for Accurate Handwritten Text Generation
- Impact of Temporal Precision on Speech Synthesis Accuracy From Electrocorticographic Brain Signals
- Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
- Imperceptible Adversarial Attacks on Point Clouds Guided by Point-to-Surface Field
- Imperceptible Transfer Attack on Large Vision-Language Models
- Implanting Robust Watermarks in Latent Diffusion Models for Video Generation
- Implementing Finite Impulse Response Filters on Quantum Computers
- Implicit Neural Representations with Fourier Kolmogorov-Arnold Networks
- Implicit and Explicit Rule Injection for Complex Query Answering over Knowledge Graphs
- Importance-Awareness Masking Network for Robust Document Retrieval
- Improved Bounds For Online Convex Optimization
- Improved Cross-Lingual Speaker Verification Using Speaker Sensitive Feature Guidance and Fine-grained Phonetic Information
- Improved Extrinsic Calibration of Acoustic Cameras via Batch Optimization
- Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
- Improved Image Classification with Manifold Neural Networks
- Improved Motion Plane Adaptive 360-Degree Video Compression Using Affine Motion Models
- Improved Out-of-domain Detection in VAE Latent Spaces with Boundary-driven Regularisation
- Improved Pitch and Voicing Determination Using the Reflected Root Chirp Group Delay Spectrum
- Improved Recognition of the Speech of People with Parkinson's Who Stutter
- Improved Techniques for Offline Reinforcement Learning: Advantage Value Estimation and Layernorm
- Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions
- Improving 5G Positioning Through Signal-to-Noise Ratio Recognition Training
- Improving Acoustic Scene Classification in Low-Resource Conditions
- Improving Adversarial Transferability through Channel-wise Scaling and Frequency-random Dropping
- Improving Compressive Imaging Recovery via Measurement Augmentation
- Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC Loss
- Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces
- Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
- Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASR
- Improving Embeddings by Refining Meanings for Temporal Knowledge Graph Link Predictions
- Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMs
- Improving GAN Performance Using Confidence-Aware Discrimination
- Improving Generated and Retrieved Knowledge Combination Through Zero-shot Generation
- Improving Height Prediction for Vision-Based Roadside 3D Object Detection
- Improving Irregular Text Recognition with Adaptive Feature Compression
- Improving Knowledge Base Question Answering via Retrieval Enhancement and Stepwise Reasoning
- Improving Knowledge Distillation via Cross-Modal Insights from CLIP
- Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
- Improving Micro-expression Recognition using Multi-sequence Driven Face Generation
- Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
- Improving Multimodal Human Pose Estimation by Adversarial Modality Enhancement†
- Improving Multimodal Large Language Models through Combining Resampler and MLP Projections
- Improving Open-Ended Referring Expression Comprehension via Dual-Language Constraints
- Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation Adjustment
- Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
- Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
- Improving Robustness of Post-hoc Calibration Against Common Corruptions By Learnable Augmentation
- Improving Sidescan Sonar Performance Using Array Upsampling Beamforming Synthetic Aperture
- Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
- Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
- Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
- Improvised Performance Following in Real Time for Automatic Accompaniment
- In Search of Optimal Pretraining Strategy for Robust Speaker Recognition
- In-Context Multitask Learning for Few-shot Fine-tuning of Large Language Models in Traditional Chinese Medicine Tongue Diagnosis
- Incorporate Global Information from Entire Datasets for Knowledge Tracing via Mini-Batch Input
- Incorporating Improved Sinusoidal Threshold-based Semi-supervised Method and Diffusion Models for Osteoporosis Diagnosis
- Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
- Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
- Individual Fairness for Fuzzy C-Means Clustering
- Indoor Airflow Imaging Using Physics-Informed Schlieren Tomography
- Indoor Sensing with Measurements
- Infant Cry Detection Using Causal Temporal Representation
- Inference Retrieval-Augmented Multi-Modal Chain-of-Thoughts Reasoning for Language Models
- Influence of Oropharyngeal Esophageal Cavity Geometry and Beak Angle on Vocal Tract Resonance of Birds using Computational Modeling
- Influence-Based Channel Reweighting for Multivariate Time Series Forecasting
- InfoHarmonizer Graph Contrastive Clustering
- InfoMin-based Query Embedding Optimization For Query-based Universal Sound Separation
- Information-Theoretic Minimax Regret Bounds for Reinforcement Learning based on Duality
- Infrared and Visible Image Fusion with Hierarchical Human Perception
- InjectTST: Injecting Global Information into Independent Channels for Long Time Series Forecasting
- Injecting Global Context for Multivariate Time Series Forecasting on Variable Subsets
- Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech Recognition
- Input Uncertainty Attribution by Uncertainty Propagation
- InsectMamba: State Space Model with Adaptive Composite Features for Insect Recognition
- Inside and Inside: Efficient Anomaly Detection by Fully Capturing the Detailed Dynamics
- InstAD: Instance-aware Segmentation Framework for Zero-shot Multi-instance Anomaly Detection
- Instance Segmentation of Airway Anatomies Using Mask R-CNN Prompt Adaptation-SAM
- Instance-wise Feature Acquisition with Classifier Selection Option for Structured Data Instances
- InstantSpeech: Instant Synchronous Text-to-Speech Synthesis for LLM-driven Voice Chatbots
- Instantaneous Trajectory Prediction via Latent Bidirectional Cooperative Diffusion
- Integrated Global-Local Gaussian Attention for Image Compression
- Integrated Interpolation and Matrix Completion for Radio Map Estimation: A Convex Optimization Approach
- Integrating Adaptive Sampling for Optimal Learned Video Compression
- Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
- Integrating Concept Associations for Query Focused Knowledge Summarization
- Integrating Failures in Robot Skill Acquisition with Offline Action-Sequence Diffusion RL
- Integrating Multi-Scale Compression Attention with Edge Detection for Ultrasound Tumor Segmentation
- Integrating Pause Information with Word Embeddings in Language Models for Alzheimer's Disease Detection from Spontaneous Speech
- Integrating Potential Pronunciations for Enhanced Mispronunciation Detection and Diagnosis Ability in LLMs
- Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection
- Intelligent Target Maneuverability in Presence of Tracking with Multiple Radars
- Intent-driven In-context Learning for Few-shot Dialogue State Tracking
- Inter- and Intra-Sentence Cuer-Invariant Representation Learning for Generalizable Cued Speech Recognition
- Inter-Frame Skip Coding Mode For Point Cloud Geometry Compression in Solid G-PCC
- Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos
- Interactive and Balanced Multimodal Learning via Cross Attention and Gradient Modulation for Compressed Video Action Recognition
- Interference-Resilient Hybrid Multi-Antenna ARQ
- Interpolation Filter Design for Sample Rate Independent Audio Effect RNNs
- Interpolation for Weight-Constrained Nested Arrays Having Non-Central ULA Segments in the Coarray
- Interpreting Deep Neural Network-Based Receiver Under Varying Signal-To-Noise Ratios
- Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
- Intra-modal Relation and Emotional Incongruity Learning using Graph Attention Networks for Multimodal Sarcasm Detection
- Intrusion Detection for Intelligent Transportation Systems: A lightweight interpretable model
- InvGS: a Novel Real-Time Inverse Rendering Framework Utilizing 3D Gaussian Splatting
- Invariant Model Learning on Local-Aware Wasserstein Geodesic for Domain Adaptation
- Investigating F0 Estimation in Speech Synthesis from Real-time MRI Articulatory Data
- Investigating Factors Related to the Naturalness of Synthesized Unison Singing
- Investigating Numerical Translation with Large Language Models
- Investigating Training Objectives for Generative Speech Enhancement
- Investigating the Sensitivity of Pre-trained Audio Embeddings to Common Effects
- Investigating voiced and unvoiced regions of speech for audio deepfake detection
- Investigation of Spatial Self-Supervised Learning and Its Application to Target Speaker Speech Recognition
- Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
- Investigation of perceptual music similarity focusing on each instrumental part
- Is It Still Fair? Investigating Gender Fairness in Cross-Corpus Speech Emotion Recognition
- Iterative Operator Sketching Framework for Large-Scale Imaging Inverse Problems
- JANE: Joint Angle Networks Assisting 3D Human Pose Estimation
- JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
- JSUnet: A New Hybrid U-shaped Network for Jamming Suppression
- Jack of All Trades, Master of None: PMP-Guided Adaptive Multi-Teacher Distillation with Meta-Learning
- Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
- Joint Beamforming Design for Multi-Functional RIS-Aided Over-the-Air Computation
- Joint Edge and Regional Depth Enhancement Network for Camouflaged Object Detection
- Joint Energy-Based Optimization of Binary Offloading Decisions and Communication Resources in TDMA Systems, via Dynamic Programming
- Joint Feature and Kernel Fusion for Improved Depth-Aware Panoptic Segmentation
- Joint Multi-Scale Contextual and Noise Suppression for Group Emotion Recognition
- Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration With Improved Intelligibility
- Joint Semantic Segmentation of Optical and SAR Image in Hazy Environments via Cross-modal Information Rectification and Cross-attention Fusion
- Joint Space-Time Adaptive Processing and Beamforming Design for Cell-Free ISAC Systems
- Joint Task Offloading and Routing in Wireless Multi-hop Networks Using Biased Backpressure Algorithm
- Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank Adaptation
- Joint-Wise Distributed Perception Graph Convolutional Network for Skeleton-Based Action Recognition
- JointSwinUNETR: an Efficient Feature-enhanced Architecture for Small Intestine Cine MRI Segmentation
- Jointly Optimal Array Geometries and Waveforms in Active Sensing: New Insights Into Array Design via the Cramér-Rao Bound
- Jointly Optimizing Data Discretization and Naive Bayes Classifier via Multi-Objective Optimization
- K-HashFed: Communication Efficient Federated Learning through Gradient Clustering and Hashing
- KABON: Knowledge Aggregation with Vision-Language Model for Black-Box Open-Set Domain Adaptation
- KAFQN: Kolmogorov-Arnold Fuzzy-guided Q-Network in Reinforcement Learning
- KAN v.s. MLP for Offline Reinforcement Learning
- KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head Generation
- KAN-HyperpointNet for Point Cloud Sequence-Based 3D Human Action Recognition
- KANGAN-AVSS: Kolmogorov-Arnold Network Based Generative Adversarial Networks for Audio-Visual Speech Synthesis
- KARLM: Enhancing LLM-based Recommendation Systems with Knowledge Bases
- KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification
- KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration
- KCE-Unet: A novel music denoising method with KANConv ECA Unet
- KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation
- KGD-GNN: A Knowledge-Guided Graph Neural Network for Myocardial Infarction Localization via 12-lead ECG
- KIKE: Linguistic Steganalysis Based on Knowledge Infusion and Knowledge Encoding
- KLFormer: Karhunen-Loève Transform for Robust 3D Human Pose Estimation
- KLMN: Knowledge distillation based lightweight multi-clue image forgery detection and localization
- KMG-LL: Knowledge-enhanced Multimodal Graph for Dialogue Generation
- KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models
- Keep what you need : extracting efficient subnetworks from large audio representation models
- Keeping Your Eyes on the Fingertip: A Two-Stage In-Air Handwritten Recognition Method
- Keeping the Balance: Anomaly Score Calculation for Domain Generalization
- Keeping the Best: The K-Best rule for Efficient Quickest Change Detection with Unknown Post-Change Distribution
- Kernel-Based Anomaly Detection Using Generalized Hyperbolic Processes
- Key Clues Guided Video Character Social Relationship Recognition Enhanced by LLM
- Keypoint Aware Masked Image Modelling
- Knocking on IP: Unveiling Websites through Cache-Aware Fingerprinting
- Know Your Heart Better: Multimodal Cardiac Output Monitoring using Earbuds
- Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR
- Knowledge Distillation From Ensemble for Spoken Language Identification
- Knowledge Distillation for Image Restoration : Simultaneous Learning from Degraded and Clean Images
- Knowledge Enhanced Multi-Domain Recommendations in an AI Assistant Application
- Knowledge Is Powerful: Art Knowledge-Driven Framework for Painting Style Classification Integrating Multimodal Knowledge
- Knowledge Transfer Across Modalities for Weakly Supervised Point Cloud Semantic Segmentation
- Knowledge-Enhanced Poetry-Image Synthesis with Large Language Model
- Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection
- Known-Plaintext Attacks to Thumbnail-Preservation Encryption Using Pix2pix Generative Adversarial Network
- Kronecker-structured Sparse Vector Recovery with Application to IRS-MIMO Channel Estimation
- L2 · M = C2 Large Language Models Are Covert Channels
- L2G: Head Gesture Animation Using an Emotion Guided Language Model
- L3D-Pose: Lifting Pose for 3D Avatars from a Single Camera in the Wild
- LABEL-SAM: A Semi-Automatic Interactive Annotation Model for Aortic Dissection Segmentation in 3D CTA Image
- LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
- LAVViT: Latent Audio-Visual Vision Transformers for Speaker Verification
- LBPE: Long-token-first Tokenization to Improve Large Language Models
- LCE: A Framework for Explainability of Ultrasound Image Based on Concept Discovery
- LCFed: An Efficient Clustered Federated Learning Framework for Heterogeneous Data
- LDG: Lightweight Deformable 3D Gaussians for Single-View Dynamic Scene Reconstruction
- LDGNet: LLMs Debate-Guided Network for Multimodal Sarcasm Detection
- LEF-TTS: Lightweight and Efficient End-to-End Text-to-Speech Synthesis With Multi-Stream Generator
- LEP: Leveraging Local Entropy Pruning for Sparsity in Large Language Models
- LFSRDiff: Light Field Image Super-Resolution via Diffusion Models
- LGNet: Linear Graph Representation for Efficient Cold-Start Recommendations
- LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
- LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
- LIMMITS'25: Multilingual Streaming TTS With Neural Codecs for Indian Languages
- LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
- LKA-ReID: Vehicle Re-Identification with Large Kernel Attention
- LKConvPose: A Pose Estimation Model with Large Receptive Field
- LKSNeXt: An Efficient Medical Image Segmentation Network with Large Kernels and Lightweight Structure
- LLDB: Efficient Low-Light Image Enhancement with Difffusion Bridge
- LLFA: Fusing Global Illumination and Local Priors for Low-Light Face Image Enhancement with Adaptor
- LLGS: Illuminating Gaussian Splatting via absorptance Modulation
- LLM based Text Generation for Improved Low-resource Speech Recognition Models
- LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
- LLM-Augmented Symbolic RL with Landmark-Based Task Decomposition
- LLM-GAN: Constructing Generative Adversarial Network Through Large Language Models for Explainable Fake News Detection
- LLM-Guided Dual-Branch Diffusion Model for Fine-Grained Motion Synthesis
- LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study
- LLMProto: A Hardware-Efficient Finetuning Model for Few-Shot Relation Extraction with Large Language Model
- LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
- LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
- LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
- LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with Efficient Narrow-Band and Cross-Band Attention
- LMTalker: Sparse Landmark-guided Gaussian Splatting for High-fidelity Talking Head Synthesis
- LNLFace: Enhanced Blind Face Restoration With Local and Non-local Lookups
- LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video Codec
- LOFI: Harnessing Attention Dynamics for Facial Expression Recognition with Noisy Labels
- LP-Gaussians: Learnable Parametric Gaussian Splatting for Efficient Dynamic Reconstruction of Single-View Scenes
- LPBS: A RL-PPO Driven K8S Batch Processing Task Scheduler
- LSTM-QGAN: Scalable NISQ Generative Adversarial Network
- LSU-NET: Lightweight Automatic Organs Segmentation Network for Medical Images
- LTOS: Layout-controllable Text-object Synthesis via Adaptive Cross-attention Fusions
- LV-ReID: Large Language-Vision Alignment Model for Text-based Person Re-identification
- LaTeXNet: A Specialized Model for Converting Visual Tables and Equations to LaTeX Code
- Label Dependency Aware Loss for Reliable Multi-Label Medical Image Classification
- Label Relationship Graph-Enhanced Class Hierarchy for Incremental Classification of Remote Sensing Images
- Label-constrained Unsupervised Domain Adaptation for Semantic Segmentation with Diffusion Models
- LagTS: Toward Adaptive Lag Relationship Modeling for Multivariate Time Series Forecasting
- Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
- Language-Queried Target Sound Extraction Without Parallel Training Data
- Language-based Audio Moment Retrieval
- Large Covariance Matrix Estimation for Groups of Highly Correlated Variables via Nonconvex Optimization
- Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
- Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
- Large Language Model-Empowered Adversarial Fusion for Typhoon Track Prediction
- Large Language Models Are Efficient Learners as Zero-Shot Speech Translators
- Large Language Models are Strong Audio-Visual Speech Recognition Learners
- Large Multimodal Model is a Better Comparator on Facial Beauty Prediction
- Large-Scale Recurrent Neural Networks with Fully Homomorphic Encryption for Privacy-Enhanced Speaker Identification
- Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
- Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
- Latent Representation Learning for Multimodal Brain Activity Translation
- Latent Space Score-based Diffusion Model for Probabilistic Multivariate Time Series Imputation
- Latent Watermarking of Audio Generative Models
- LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation
- Layer-Animate for Transparent Video Generation
- Lead Instrument Detection from Multitrack Music
- Learn from Balance: Rectifying Knowledge Transfer for Long-Tailed Scenarios
- Learned Approximated Optimization for Rapid Low-Complexity Hybrid Beamforming Design
- Learned ReLU-Based Soft Thresholding: A Data-Driven Method for Non-Negative Sparse Signal Recovery
- Learned Video Compression With Refined Adaptive Flow Pyramid And Coordinate-Aware Attention
- Learning Adaptive Spatial-temporal Structured Correlation Filters for UAV Object Tracking
- Learning Binary-Antithetical Information Bottleneck for Generalizable Face Anti-Spoofing
- Learning Class Prototypes for Visual Emotion Recognition
- Learning Class Unique Features in Fine-Grained Visual Classification
- Learning Control of Neural Sound Effects Synthesis from Physically Inspired Models
- Learning Deep Frequency Degradation Prior for Remote Sensing Spatio-temporal Fusion
- Learning Diffusion Model from Noisy Measurement using Principled Expectation-Maximization Method
- Learning Hierarchical Attribute Prompt for Vision-Language Models
- Learning Joint Appearance and Shape Co-Representations for Co-Saliency Detection
- Learning Markup Language Model for Composite Relationships Extraction
- Learning Music Audio Representations With Limited Data
- Learning Permutations in Monarch Factorization
- Learning Preconditioners in Gates-controlled Deep Unfolding Networks based on Quasi-Newton Methods For Accelerated MRI Reconstruction
- Learning Primitive Relations for Compositional Zero-Shot Learning
- Learning Rank Constrained Exposure Correction from Unpaired Data
- Learning Rate Optimization for Deep Neural Networks Using Lipschitz Bandits
- Learning Rich Speech Representations with Acoustic-Semantic Factorization
- Learning Semantic Facial Descriptors for Accurate Face Animation
- Learning Simultaneous Facial Canonical Correlation Representation for Face Hallucination
- Learning Source Disentanglement in Neural Audio Codec
- Learning Statistical and Physical Modeling for Consistency Human Motion Prediction
- Learning Strategy with Barlow Twins Objective for Emotion-Robust Speaker Verification System
- Learning Stroke-Order Dynamics in Few-Shot Font Generation via Sequential Awareness
- Learning Structured Compressed Sensing with Automatic Resource Allocation
- Learning Time-Varying Graphs from Data with Few Causes
- Learning Two-factor Representation for Magnetic Resonance Image Super-resolution
- Learning Weighted Least Squares Data Term for Poisson Image Deconvolution
- Learning a Sparse Polynomial Approximation to the Transition Function of General State-Space Models
- Learning from Ambiguous Data with Hard Labels
- Learning from Reconstruction: A Two-Stage Global-to-Local Framework for Temporal Knowledge Graph Completion
- Learning in the Model Space: Fault Diagnosis by Co-objective Learning in DynInt Model Space
- Learning to Follow Infrared Prior Repersentation for Image Dehazing
- Learning to Optimally Sample in MRI for Denoising-Driven Regularization
- Learning to Reconstruct Signals With Inexact Sensing Operator via Knowledge Distillation
- Learning with Coupled Noisy Labels for Visible-Infrared Person Re-identification via Graph Consistency
- Learning with Partial Labels from Conflict-Free and Semi-Supervised Perspective
- Learning-Aided Kalman Tracking in Biased Dynamic Systems: The Case of Cable-Driven Robots for Surgery
- Learning-Based Utility Estimation with Application to Speech Enhancement of a Moving Speaker
- Leave No Stone Unturned: Optimizing Subpattern Information Entropy for Coreset Selection
- Leave-One-EquiVariant: Alleviating Invariance-Related Information Loss in Contrastive Music Representations
- Lenna: Language Enhanced Reasoning Detection Assistant
- Less Is More: Embracing Sparsity and Interpolation with Esiformer for Time Series Forecasting
- Less Over More: Interference Sample Gradient Purification For Parallel Continual Learning
- Less Yet Robust: Crucial Region Selection for Scene Recognition
- Less is Enough: Relation Graph Guided Few-shot Learning for Multi-label Aspect Category Detection
- Less is more: Efficient Scene Graph Generation with reparameterization
- Let There Be Light: Robust Lensless Imaging Under External Illumination With Deep Learning
- Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
- Leveraging Boolean Directivity Embedding for Binaural Target Speaker Extraction
- Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
- Leveraging Heterophily in Spatial-Temporal Graphs for Multivariate Time-Series Forecasting
- Leveraging IPA and Articulatory Features as Effective Inductive Biases for Multilingual ASR Training
- Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
- Leveraging Mixture of Experts for Improved Speech Deepfake Detection
- Leveraging Multimodal Diffusion Models to Accelerate Imaging with Side Information
- Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification
- Leveraging Out-of-Domain Noise for Unsupervised Domain Adaptation in Speech Enhancement
- Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
- Leveraging Registers in Vision Transformers for Robust Adaptation
- Leveraging Self-Supervised Learning for Speaker Diarization
- Leveraging Visual Captions for Enhanced Zero-Shot HOI Detection
- LiDAR Light Scattering Augmentation (LISA): Physics-based Simulation of Adverse Weather Conditions for 3D Object Detection
- LiDAR-SPD: Improving Adversarial Robustness of 3D Object Detection via Spherical Projection and Diffusion
- LiRCDepth: Lightweight Radar-Camera Depth Estimation via Knowledge Distillation and Uncertainty Guidance
- LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
- Lightweight Clustered Federated Learning via Feature Extraction
- Lightweight Image Quality Prediction Guided by Perceptual Ranking Feedback
- Lightweight Multi-Frequency Enhancement Network for RGB-D Video Salient Object Detection
- Lightweight Self-Supervised Monocular Depth Estimation for All-Day Scenes Using Generative Adversarial Network
- Lightweight neural front-ends for low-resource on-device Text-to-Speech
- Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
- Linguistics-Vision Monotonic Consistent Network for Sign Language Production
- Linking Known and Unknown: Generalized Cross-Instance Feature Helps Category Discovery
- LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
- LipReading for Low-resource Languages by Language Dynamic LoRA
- LitePest: Real-Time and Efficient Detection of Agricultural Pests Using an Advanced Lightweight Deep Learning Network
- LkSFocalNets: Video Action Recognition With Large Kernel Selective Focal Networks
- LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
- LoRATEE: A Secure and Efficient Inference Framework for Multi-Tenant LoRA LLMs Based on TEE
- LoVA: Long-form Video-to-Audio Generation
- LocRef-Diffusion: Tuning-Free Layout and Appearance-Guided Generation
- Local Adaptive Time-Frequency Bidirectional Synchrosqueezing Transform
- Local Feature Alignment Prompt-Tuning for Few-shot Multimodal Aspect Sentiment Analysis
- Local Statistics for Generative Image Detection
- Localised Frequency Latent Domain Watermarking of DDIM Generated Images
- Locally Correctable Lattices
- LogSI: A Benchmark for System-Incremental Log Analysis
- Long-Range Multi-Scale Fusion for Efficient Single Image Super-Resolution
- Long-tailed Oracle Character Recognition Based on Convolutional Neural Networks and Vision Transformers
- Longitudinal Wrist PPG Analysis for Reliable Hypertension Risk Screening Using Deep Learning
- Look Before You Leap: Problem Elaboration Prompting Improves Mathematical Reasoning in Large Language Models
- Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction
- LossControl: Defending Membership Inference Attacks by Controlling the Loss
- Lossless Phase Conversion Method for Object Wave-based Hologram Compression
- Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications
- Low Complexity DoA-ToA Signature Estimation for Multi-Antenna Multi-Carrier Systems
- Low Complexity Rate Splitting Approach in RIS-Aided Systems Based on Channel Statistics
- Low Complexity Riemannian Coordinate-Descent over Symmetric Positive Definite Matrices
- Low Complexity Super Resolution for Resampling-based Video Coding
- Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
- Low Rank and Sparse Fourier Structure in Recurrent Networks Trained on Modular Addition
- Low-Complexity Cramér-Rao Lower Bound and Sum Rate Optimization in ISAC Systems
- Low-Complexity Neural Speech Dereverberation With Adaptive Target Control
- Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear Microphone
- Low-Correlation OFDM Waveform Design With Optimally Coded Sub-Carriers for the Joint Sensing and Communications
- Low-Light Detector Based on Feature Filtering and Enhancement
- Low-Rank Tensors for Multi-Dimensional Markov Models
- Low-Rank Transformer Adaptation for Arbitrary Style Transfer
- Low-Rank Tucker Decomposition of Multi-Subject Complex-Valued fMRI Data
- Low-Rate Modulo Folded ADC for Detecting Linearly Modulated Communication Symbols
- Low-Resolution Hierarchical Training for Efficient 3D Gaussian Splatting
- Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
- Low-rank Adaptation Method for Respiratory Sound Classification: A necessary road towards Large Models
- Low-shot Image Classification Using Mixture of Experts
- Lunar Tracking: A New Benchmark For Nighttime Tiny Object Tracking
- Lungmix: A Mixup-Based Strategy for Generalization in Respiratory Sound Classification
- M*: On-Chip Microfluidic Operations With A-Star for Portable Diagnostics
- M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
- M-MoE: Mixture of Mixture-of-Expert Model for CTC-based Streaming Multilingual ASR
- M2F2Net: Multi-stage Mixed Feature Fusion Network For Remote Sensing Change Detection
- M2PAIR: A High-Quality Acoustic Impulse Response Computation Model
- M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
- M3-CVC: Controllable Video Compression with Multimodal Generative Models
- M3ADD: A Novel Benchmark for Physiology Signal-based Automatic Depression Detection with Multimodal Multitask Multievent Framework
- M3Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential Recommendation
- MA-Det: A Discriminative Morphology-Aware Detector for Cervical Lesion Cell Clumps
- MACA: Multi-Anchor Classification Approach for Unsupervised Domain Adaptation
- MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion
- MAEM: A Multi-Aspect Extraction Model for Enhanced Embedding in RAG
- MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal Fusion
- MAID: Model Attribution via Inverse Diffusion
- MAITFuse: Multi-Dimension Adaptive Interaction Transform Network For Infrared-visible Image Fusion
- MAJoR: Visual Emotion Analysis via Multi-Attribute Joint Reasoning
- MAP Image Recovery with Guarantees using Locally Convex Multi-Scale Energy (LC-MUSE) Model
- MAP: Supporting Multimodal Knowledge Graph Completion via Augmented Modality Alignment and Instance Preserving
- MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
- MDDNet: Multilevel Difference-Enhanced Denoise Network for Unsupervised Change Detection in SAR Images
- MDN: Mamba-Driven Dualstream Network For Medical Hyperspectral Image Segmentation
- MDNet: Multi-Decoder Network for Abdominal CT Organs Segmentation
- MDRNet: Multi-Branch with Different Feature Representations Network for Motor Imagery Classification
- MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image Retrieval
- MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding Challenge
- META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
- MF-BERT: A Siamese Pre-training Framework for Motion Forecasting
- MFANet: Multi-Feature Aggregation Network for Multi-focus Image Fusion
- MFDPonzi: Detecting Ethereum Ponzi Schemes Using Static Features from Novel Opcode Sequences
- MFMamba: A Multimodal Fusion State Space Model for Depression Recognition
- MFT: Modal Fusion Transformer for Cross-Modal Fusion in 3D Object Detection
- MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
- MHGNet: Multi-Heterogeneous Graph Neural Network for Traffic Prediction
- MHSDB: A Comprehensive Benchmark for Multimodal Humor and Sarcasm Detection Leveraging Foundation Models
- MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword Spotting
- MIFAE-Forensics: Masked Image-Frequency AutoEncoder for DeepFake Detection
- MILE: Multi-Instance Learning for Document Event Argument Extraction
- MIMO Channel as a Neural Function: Implicit Neural Representations for Extreme CSI Compression
- MINR: Efficient Implicit Neural Representations for Multi-Image Encoding
- MKD-YOLO: Multi-Scale and Knowledge-Distilling YOLO for Efficient PPE Compliance Detection
- MLNet: Mutual Learning Network to Improve Self-Supervised Representation for Fine-Grained Visual Recognition
- MLSDET: Multi-LLM Statistical Deep Ensemble for Chinese AI-Generated Text Detection
- MLSwinTNet: A Multi-Level Feature Interaction Network for Low-Light Image Enhancement
- MM-LogVec: System Log Anomaly Detection Method Based on Multimodal Representation Learning
- MMA-Net: Multi-Modal Attention Network for 2-D Object Detection in Autonomous Driving
- MMCD: Memory-Based Multimodal Change Detection
- MMEditor: Multimodal Prompt-Driven 3D Gaussian Splatting Editing
- MMFN: Multi-Feature Multi-Modal Fusion Network for Diagnosis of Superficial Lymph Node Disease
- MMTP: Meta-learning-based Multi-Textual Prompt Tuning for Visual-Language Models
- MP-DPCC: A Motion Proxy-Based Dynamic Point Cloud Compression Framework
- MPAM-3DGS: Multi-Parametric Adversarial Manipulation for 3D Gaussian Splatting
- MPFL: A Decentralised Federated Learning Framework Based on Multi-Population Genetic Algorithm
- MPNAS: Multimodal Sentiment Analysis Pruning via Neural Architecture Search
- MPOT: Manifold Preserving Optimal Transport for Visual Recognition Under Severe Distribution Shift
- MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
- MQVAE: Capturing Metastable Dynamics from EEG for Brain-computer Interfaces
- MRANet: An Encoder-Decoder Network with Multi-Scale Residual Atrous-Spatial Pyramid Pooling for Seismic Phase Picking
- MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
- MS-RainMamba: Learning Multi-Scale State Space Models for Single Image Deraining
- MS-SCANet: A Multiscale Transformer-Based Architecture with Dual Attention for No-Reference Image Quality Assessment
- MS-UFAD: A Large-Scale Dataset for Real-world Unified Face Attack Detection with Text Descriptions
- MSA-ITEI: A Novel Method for Multimodal Analysis of Social Media Stickers
- MSACC: A Unified Multimodal Sentiment Analysis Framework for High Interpretability and Zero-shot Performance
- MSANet: Mixed Spectral and Attention Network for Robust 3D Human Pose Estimation
- MSE-based Sampling of Bandlimited Product Graph Signals via Joint Low-pass Impulse Responses
- MSECG: Incorporating Mamba for Robust and Efficient ECG Super-Resolution
- MSEMG: Surface Electromyography Denoising with a Mamba-based Efficient Network
- MSRFormer: Hybrid Scale Self-Attention and Local Fast Convolution Transformer for Facial Expression Recognition
- MST-HA: Multi-Modal Signal Fusion with Bayesian Optimization for Robust Industrial Robot Joint Health Assessment
- MSTBI: Head CT Detection and Prognostic Assessment of Traumatic Brain Injury Dataset
- MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
- MTE: Multi Transformation of Entities in Quaternion Vector Space for Temporal Knowledge Graph Completion
- MTMDC-GAN: Self-Attention Driven Multi-Scale Temporal Synthesis with Multi-Domain Analysis and Contrastive Learning
- MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
- MTTM: Memory-Augmented with Mamba for 3D Medical Images Analysis
- MULiving: Towards Real-time Multi-User Survival State Monitoring Using Wearable RFID Tags
- MUPO-Net: A Multilevel Dual-domain Progressive Enhancement Network with Embedded Attention for CT Metal Artifact Reduction
- MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation
- MVCBRec: Multi-View Contrastive Learning for Bundle Recommendation
- MVDC : A Multi-view Dental Completion Model Based on Contrastive Learning
- MX-Font++: Mixture of Heterogeneous Aggregation Experts for Few-shot Font Generation
- MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
ICASSP accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.