← Search

Long Ma

59 accepted papers

2026

BEYOND FACE SWAPPING: A DIFFUSION-BASED DIGITAL HUMAN BENCHMARK FOR MULTIMODAL DEEPFAKE DETECTION

ICASSP 2026poster

In recent years, the explosive advancement of deepfake technology has posed a critical and escalating threat to public security: diffusion-based digital human generation. Unlike traditional face manipulation methods, such models can generate highly realistic videos with consistency via multimodal co…

Cited by 0SourcePDFScholar
2026

Benchmarking Endoscopic Surgical Image Restoration and Beyond

CVPR 2026

In endoscopic surgery, a clear and high-quality visual field is critical for surgeons to make accurate intraoperative decisions. However, persistent visual degradation, including smoke generated by energy devices, lens fogging from thermal gradients, and lens contamination due to blood or tissue flu

Cited by 0SourcecodeScholar
2026

BiPA: Bilevel Prompt Adaptation for Underwater Instance Segmentation

CVPR 2026

Underwater instance segmentation is essential for fine-grained scene understanding. However, underwater imagery exhibits a strong domain gap from in-air vision due to severe degradation (e.g., turbidity). Consequently, despite its general segmentation ability, SAM degrades sharply underwater. In thi

Cited by 0SourcecodeScholar
2026

Bridging Human Evaluation to Infrared and Visible Image Fusion

CVPR 2026

Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and objective metrics, often resulting in fusion outcomes that do not align with human visual preferences. This challenge is

Cited by 0SourcecodeScholar
2026

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

AAAI 2026technical

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their

Cited by 0SourcePDFScholar
2026

Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing

CVPR 2026

Face Anti-Spoofing (FAS) typically depends on a single visual modality when defending against presentation attacks such as print attacks, screen replays, and 3D masks, resulting in limited generalization across devices, environments, and attack types. Meanwhile, Multimodal Large Language Models (MLL

Cited by 0SourcecodeScholar
2026

LEVERAGING LARGE MULTIMODAL MODELS FOR AUDIO-VIDEO DEEPFAKE DETECTION: A PILOT STUDY

ICASSP 2026oral

Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-specific models: they work well on curated tests but scale poorly and generalize weakly across domains. We introduce AV-LM…

Cited by 0SourcePDFScholar
2026

Learning 3D Occupancy from Beam Overlap in 2D Rotating mmWave Radar

AAAI 2026technical

Robust 3D perception under adverse weather is critical for autonomous systems. While mmWave Radars are inherently weather-resistant, conventional 2D rotating Radar sensors lack direct elevation resolution, limiting their 3D perception ability. Although 4D imaging radars can provide elevation informa

Cited by 0SourcePDFScholar
2026

Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation

IJCAI 2026

Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are expensive. Point supervision with online label evolution reduces annotation cost; however, lightweight CNN detectors often lack sufficient semantics,

Cited by 0Scholar
2026

Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management

ICLR 2026poster

Large Language Models (LLMs) suffer from significant performance degradation when processing long contexts due to proactive interference, where irrelevant information in earlier parts of the context disrupts reasoning and memory recall. While most research focuses on external memory systems to augme…

Cited by 0SourceScholar
2026

Streaming Diffusion Model for Fast Infrared and Visible Video Fusion

CVPR 2026

Infrared and visible video fusion is pivotal for robust perceptual systems, aiming to synthesize a comprehensive video stream that leverages both thermal resilience and textured details. However, prevailing methods, by treating videos as sequences of independent frames, inherently introduce temporal

Cited by 0SourcecodeScholar
2026

Taming Generative Diffusion Model for Task-Oriented Infrared Imaging

CVPR 2026

Infrared imaging is essential for perception in harsh environments. However, dynamically coupled degradation factors severely impair visual quality and downstream semantic accuracy. Although generative diffusion models provide strong image restoration priors, high computational cost and physical inc

Cited by 0SourcecodeScholar
2026

Your One-Stop Solution for AI-Generated Video Detection

CVPR 2026

Recent advances in generative modeling can create remarkably realistic synthetic videos, making it increasingly difficult for humans to distinguish them from real ones and necessitating reliable detection methods. However, two key limitations hinder the development of this field.**From the dataset p

Cited by 0SourcecodeScholar
2025

Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement Learning

ICML 2025poster

The ability to adapt to new environments with noisy dynamics and unseen objectives is crucial for AI agents. In-context reinforcement learning (ICRL) has emerged as a paradigm to build adaptive policies, employing a **context** trajectory of the test-time interactions to infer the true task and the…

Cited by 0SourcePDFScholar
2025

Binary Representation Learning for Discriminative Acoustic Unit Discovery

ICASSP 2025accepted

Acoustic Unit Discovery (AUD) aims to obtain phoneme-like units that preserve linguistically significant information while removing paralinguistic details. Although Contrastive Predictive Coding (CPC) has emerged as a leading self-supervised representation learning method for this task, CPC-based me…

Cited by 0SourceScholar
2025

CoA: Towards Real Image Dehazing via Compression-and-Adaptation

CVPR 2025poster

Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both eff…

2025

DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion

CVPR 2025poster

Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resultin…

2025

DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared Imaging

CVPR 2025poster

Thermal imaging is often compromised by dynamic, complex degradations caused by hardware limitations and unpredictable environmental factors. The scarcity of high-quality infrared data, coupled with the challenges of dynamic, intricate degradations, makes it difficult to recover details using exis…

2025

DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution

CVPR 2025poster

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challen…

2025

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

ICASSP 2025accepted

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversi…

Cited by 0SourceScholar
2025

EchoDiffusion: Waveform Conditioned Diffusion Models for Echo-Based Depth Estimation

AAAI 2025technical

To extract spatial information, depth estimation using conventional echo-based methods typically employs models with encoder-decoder architectures, such as UNet. However, these methods may face challenges in extracting fine details from echo waveforms and handling multi-scale feature extraction with…

2025

Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark

NeurIPS 2025poster

We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one e…

Cited by 0SourceScholar
2025

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

ICML 2025poster

The GPT-4o's excellent duplex speech interaction ability has given users an impressive experience. Researchers have recently proposed several multimodal LLMs to achieve user-agent speech-to-speech conversations. In this paper, we propose a novel speech-text multimodal LLM architecture called Freeze-…

Cited by 32SourcePDFScholar
2025

From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face Deepfakes

NeurIPS 2025poster

Detecting deepfakes has been an increasingly important topic, especially given the rapid development of AI generation techniques. In this paper, we ask: How can we build a universal detection framework that is effective for most facial deepfakes? One significant challenge is the wide variety of deep…

Cited by 0SourceScholar
2025

M-MoE: Mixture of Mixture-of-Expert Model for CTC-based Streaming Multilingual ASR

ICASSP 2025accepted

The Mixture-of-Expert (MoE) structure has been effectively utilized in multilingual ASR tasks. However, the potential of external language information remains underutilized. In this paper, we introduce the Mixture of MoE (M-MoE) structure, featuring multiple language-specific MoEs and a language-unk…

Cited by 0SourceScholar
2025

PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors

CVPR 2025poster

This paper presents PCDreamer, a novel method for point cloud completion. Traditional methods typically extract features from partial point clouds to predict missing regions, but the large solution space often leads to unsatisfactory results. More recent approaches have started to use images as extr…

Cited by 1SourcePDFScholar
2025

Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyond

CVPR 2025poster

Recently, enhancing image quality in the original RAW domain has garnered significant attention, with denoising and reconstruction emerging as fundamental tasks. Although some works attempt to couple these tasks, they primarily focus on cascade learning while neglecting task associativity within a b…

2025

Simulating Human-like Daily Activities with Desire-driven Autonomy

ICLR 2025poster

Desires motivate humans to interact autonomously with the complex world. In contrast, current AI agents require explicit task specifications, such as instructions or reward functions, which constrain their autonomy and behavioral diversity. In this paper, we introduce a Desire-driven Autonomous Agen…

Cited by 2SourcePDFScholar
2025

Social World Model-Augmented Mechanism Design Policy Learning

NeurIPS 2025poster

Designing adaptive mechanisms to align individual and collective interests remains a central challenge in artificial social intelligence. Existing methods often struggle with modeling heterogeneous agents possessing persistent latent traits (e.g., skills, preferences) and dealing with complex multi-…

Cited by 0SourceScholar
2025

TextMEF: Text-guided Prompt Learning for Multi-exposure Image Fusion

IJCAI 2025

Multi-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in uns

Cited by 0SourcePDFScholar
2025

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

NeurIPS 2025spotlight

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in bot…

Cited by 0SourcecodeScholar
2024

3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editing

ECCV 2024poster

"The current GAN inversion methods typically can only edit the appearance and shape of a single object and background while overlooking spatial information. In this work, we propose a 3D editing framework, to enable multifaceted editing of affine information (scale, translation, and rotation) on mul…

2024

Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution

ECCV 2024poster

"Image super-resolution (SR) is a critical technique for enhancing image quality, playing a vital role in image enhancement. While recent advancements, notably transformer-based methods, have advanced the field, infrared image SR remains a formidable challenge. Due to the inherent characteristics of…

2024

Fast Peer Adaptation with Context-aware Exploration

ICML 2024poster

Fast adapting to unknown peers (partners or opponents) with different strategies is a key challenge in multi-agent games. To do so, it is crucial for the agent to probe and identify the peer’s strategy efficiently, as this is the prerequisite for carrying out the best response in adaptation. However…

Cited by 2SourcePDFScholar
2024

Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-Free Multi-Exposure Image Fusion

AAAI 2024technical

Multi-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements, the field grapples with challenges, notably the reliance on manual designs for network structures and loss functions,…

2024

Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation

AAAI 2024technical

Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for regular detectors. However, because of the disparity in task objectives between the enhancer and detector, this paradigm c…

Cited by 12SourcePDFScholar
2024

Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications

IJCAI 2024poster

Multi-modality image fusion aims to integrate images from multiple sensors, producing an image that is visually appealing and offers more comprehensive information than any single one. To ensure high visual quality and facilitate accurate subsequent perception tasks, previous methods have often casc…

2023

Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond

IJCAI 2023poster

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying co…

2023

Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation

ICCV 2023oral

Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach `Best of Both Worlds'. To overcome this issue, in this paper, we propos…

Cited by 172PDFcodeScholar
2022

Conversational Speech Recognition by Learning Conversation-Level Characteristics

ICASSP 2022accepted

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can naturally take advantages from specific characteristics of conversation, such as role preference and topical coherence. This p…

Cited by 0SourceScholar
2022

Hierarchical Bilevel Learning with Architecture and Loss Search for Hadamard-based Image Restoration

IJCAI 2022poster

In the past few decades, Hadamard-based image restoration problems (e.g., low-light image enhancement) attract wide concerns in multiple areas related to artificial intelligence. However, existing works mostly focus on heuristically defining architecture and loss by the engineering experiences that…

Cited by 3SourcePDFScholar
2022

Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models

ICASSP 2022accepted

Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models. Due to the conditional independence assumption, CTC-based models are always weaker than attention-based e…

Cited by 35SourceScholar
2022

Semantic-aware Texture-Structure Feature Collaboration for Underwater Image Enhancement

ICRA 2022poster

Underwater image enhancement has become an attractive topic as a significant technology in marine engi-neering and aquatic robotics. However, the limited number of datasets and imperfect hand-crafted ground truth weaken its robustness to unseen scenarios, and hamper the application to high-level vis…

Cited by 34SourcecodeScholar
2022

Toward Fast, Flexible, and Robust Low-Light Image Enhancement

CVPR 2022oral

Existing low-light image enhancement techniques are mostly not only difficult to deal with both visual quality and computational efficiency but also commonly invalid in unknown complex scenarios. In this paper, we develop a new Self-Calibrated Illumination (SCI) learning framework for fast, flexible…

Cited by 810PDFcodeScholar
2021

GTA-Net: Gradual Temporal Aggregation Network for Fast Video Deraining

ICASSP 2021accepted

Recently, the development of intelligent technology arouses the requirements of high-quality videos. Rain streak is a frequent and inevitable factor to degrade the video. Many researchers have put their energies into eliminating the adverse effects of rainy video. Unfortunately, how to fully utilize…

Cited by 0SourceScholar
2021

NASA: A Noise-Adaptive and Structure-Aware Learning Framework for Image Deblurring

ICASSP 2021accepted

Image deblurring is a classical low-level visual processing task, which aims to recover a potentially noise-free sharp image from the blurred image. Existing prior-based and learning-based methods usually need to manually set some vital auxiliary components (e.g., noise level). It brings about extre…

Cited by 0SourceScholar
2021

Retinex-Inspired Unrolling With Cooperative Prior Architecture Search for Low-Light Image Enhancement

CVPR 2021poster

Low-light image enhancement plays very important roles in low-level vision areas. Recent works have built a great deal of deep learning models to address this task. However, these approaches mostly rely on significant architecture engineering and suffer from high computational burden. In this paper,…

Cited by 884PDFcodeScholar
2021

Temporal Rain Decomposition with Spatial Structure Guidance for Video Deraining

ICASSP 2021accepted

Recently, removing rain streaks from videos has drawn wide concerns in vision and multimedia communities. But existing works ignore the depicts of image inherent structure and rain location to cause details loss, and their adopted manners of exploiting temporal information are still insufficient. In…

Cited by 0SourceScholar
2020

Principle-Inspired Multi-Scale Aggregation Network for Extremely Low-Light Image Enhancement

ICASSP 2020accepted

The under-exposure and low-light environments are common to degrade the image-quality with invisible information. To ameliorate this case, a copious of low-light image enhancement methods are developed. However, these existing works are hard to handle extremely low-light conditions with noises, even…

Cited by 0SourceScholar
2020

Sequential Deep Unrolling With Flow Priors For Robust Video Deraining

ICASSP 2020accepted

Video deraining has attracted wide attention since the urgent demand of high-quality video in recent years. The indistinct details and nonideal deraining effects are the most common defects in existing techniques, whose cause lies in the insufficient usage of single-frame image and temporal informat…

Cited by 0SourceScholar
2018

A Bridging Framework for Model Optimization and Deep Propagation

NeurIPS 2018poster

Optimizing task-related mathematical model is one of the most fundamental methodologies in statistic and learning areas. However, generally designed schematic iterations may hard to investigate complex data distributions in real-world applications. Recently, training deep propagations (i.e., network…

Cited by 20SourcePDFScholar
2018

Deep Layer Prior Optimization for Single Image Rain Streaks Removal

ICASSP 2018accepted

Visible distortions caused by rain streaks have significant negative effects on the performance of many vision and learning algorithms. Most of the existing deraining approaches propose to build complex prior models to formulate the appearance of rain streaks. Unfortunately, these human-designed pri…

Cited by 0SourceScholar
2018

Robust Haze Removal Via Joint Deep Transmission and Scene Propagation

ICASSP 2018accepted

Haze is one of the most important factors which reduce the outdoor image quality. Existing approaches often aim to design their models based on principles of hazes. However, even with exactly modeled haze distribution, it is still a challenging task due to factors in real scenario, such as noises, h…

Cited by 0SourceScholar
2017

Adaptation of PLDA for multi-source text-independent speaker verification

ICASSP 2017accepted

Probabilistic linear discriminant analysis (PLDA) is widely described as an effective model for text-independent speaker verification in the i-vector space. The PLDA scoring function is typically formulated as the likelihood ratio between the speaker-adapted and the universal PLDAs. In this case, th…

Cited by 0SourceScholar