← Search

Lei Ma

56 accepted papers

2026

Beyond Visual Reconstruction Quality: Object Perception-aware 3D Gaussian Splatting for Autonomous Driving

ICLR 2026poster

Reconstruction techniques, such as 3D Gaussian Splatting (3DGS), are increasingly used for generating scenarios in autonomous driving system (ADS) research. Existing 3DGS-based works for autonomous driving scenario generation have, through various optimizations, achieved high visual similarity in re…

Cited by 0SourcecodeScholar
2026

CACR: Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

ICML 2026poster

The task of temporal answer grounding in instructional videos (TAGV), which aims to locate precise video segments that respond to natural language queries, is increasingly important for direct video answer retrieval. This task remains challenging due to the need to comprehend semantically complex qu…

Cited by 0SourceScholar
2026

Can Protective Watermarking Safeguard the Copyright of 3D Gaussian Splatting?

AAAI 2026technical

3D Gaussian Splatting (3DGS) has emerged as a powerful representation for 3D scenes, widely adopted due to its exceptional efficiency and high-fidelity visual quality. Given the significant value of 3DGS assets, recent works have introduced specialized watermarking schemes to ensure copyright protec

Cited by 0SourcePDFScholar
2026

EcoDiffusion: Uncertainty-Aware Emulation of Ecosystem Processes with Conditional Diffusion for Long Sequences with Single-Step Initialization

AAAI 2026technical

Terrestrial ecosystems constitute a major component of the global carbon sink and play a critical role in regulating the global carbon cycle. Although process-based models such as the Ecosystem Demography (ED) model are widely used to simulate these dynamics and widely adopted in research and applic

Cited by 0SourcePDFScholar
2026

FiSeR: Fine-Grained Source Representations for Cross-Domain AI Image Detection

ICML 2026poster

Real-world synthetic image detectors often generalize poorly under domain shift despite strong in-domain performance. Using unsupervised UMAP projections, we find that natural and synthetic features remain partially separable on unseen datasets, yet performance still drops, suggesting that the class…

Cited by 0SourceScholar
2026

From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning

AAAI 2026technical

LiDAR-based 3D object detection models often struggle to generalize to real-world environments due to limited object diversity in existing datasets. To tackle it, we introduce the first generalized cross-domain few-shot (GCFS) task in 3D object detection, aiming to adapt a source-pretrained model to

Cited by 0SourcePDFScholar
2026

MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM Agents

AAAI 2026technical

Physical adversarial attacks in driving scenarios can expose critical vulnerabilities in visual perception models. However, developing such attacks remains non-trivial due to diverse real-world environmental influences. Existing approaches either struggle to generalize to dynamic environments or fai

Cited by 0SourcePDFScholar
2026

ModalPatch: A Plug-And-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop

ICRA 2026poster

Multi-modal 3D object detection is pivotal for autonomous driving, integrating complementary sensors like LiDAR and cameras. However, its real-world reliability is challenged by transient data interruptions and missing, where modalities can momentarily drop due to hardware glitches, adverse weather,…

2026

Motus: A Unified Latent Action World Model

CVPR 2026

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this pa

Cited by 0SourcecodeScholar
2026

Nano3D: A Training-Free Approach for Efficient 3D Editing Without Masks

ICLR 2026poster

3D object editing is essential for interactive content creation in gaming, animation, and robotics, yet current approaches remain inefficient, inconsistent, and often fail to preserve unedited regions. Most methods rely on editing multi-view renderings followed by reconstruction, which introduces ar…

Cited by 0SourcecodeScholar
2026

Sample-Efficient Learning with Online Expert Correction for Autonomous Catheter Steering in Endovascular Bifurcation Navigation

ICRA 2026poster

Robot-assisted endovascular intervention offers a safe and effective solution for remote catheter manipulation, reducing radiation exposure while enabling precise navigation. Reinforcement learning (RL) has recently emerged as a promising approach for autonomous catheter steering; however, conventio…

2026

Splats in Splats: Robust and Effective 3D Steganography Towards Gaussian Splatting

AAAI 2026technical

3D Gaussian splatting (3DGS) has demonstrated impressive 3D reconstruction performance with explicit scene representations. Given the widespread application of 3DGS in 3D reconstruction and generation tasks, there is an urgent need to protect the copyright of 3DGS assets. However, existing copyright

Cited by 0SourcePDFScholar
2026

TAlignDiff: Automatic Tooth Alignment assisted by Diffusion-based Transformation Learning

CVPR 2026

Orthodontic treatment hinges on tooth alignment, which significantly affects occlusal function, facial aesthetics, and patients' quality of life. Current deep learning approaches often predict transformation matrices for the misaligned tooth point cloud via point-to-point geometric constraints to ac

Cited by 0SourceScholar
2026

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

ICML 2026poster

While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks …

Cited by 0SourceScholar
2025

CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest Ecosystems

NeurIPS 2025poster

Forest ecosystems play a critical role in the Earth system as major carbon sinks that are essential for carbon neutralization and climate change mitigation. However, the Earth has undergone significant deforestation and forest degradation, and the remaining forested areas are also facing increasing…

Cited by 0SourcecodeScholar
2025

CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification

AAAI 2025technical

Large Language Models (LLMs) have made significant progress in code generation, offering developers groundbreaking automated programming support. However, LLMs often generate code that is syntactically correct and even semantically plausible, but may not execute as expected or fulfill specified requ…

2025

DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning

NeurIPS 2025poster

Detecting AI-involved text is essential for combating misinformation, plagiarism, and academic misconduct. However, AI text generation includes diverse collaborative processes (AI-written text edited by humans, human-written text edited by AI, and AI-generated text refined by other AI), where vario…

Cited by 0SourcecodeScholar
2025

DUNE: Sim2Real Transfer for Depth-based Navigation in Unstructured Dynamic Indoor Environments

ICASSP 2025accepted

Collision-free navigation in dynamic environments, especially with moving pedestrians, is crucial for mobile robots. This paper introduces DUNE, a depth-based policy trained in simulation for collision-free navigation of Ackermann mobile robots in unstructured indoor environments. DUNE uses a CNN-LS…

Cited by 0SourceScholar
2025

DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches

NeurIPS 2025poster

Stereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help revealing vulnerabilities before deployment. Previous wo…

Cited by 0SourcecodeScholar
2025

Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update

AAAI 2025technical

Warning: This paper contains offensive content that may disturb some readers. Vision-language models (VLMs) demonstrate strong multimodal capabilities but have been found to be more susceptible to generating harmful content compared to their backbone large language models (LLMs). Our investigation r…

2025

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

EMNLP 2025

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting

Cited by 0SourcePDFScholar
2025

MindCustomer: Multi-Context Image Generation Blended with Brain Signal

ICML 2025poster

Advancements in generative models have promoted text- and image-based multi-context image generation. Brain signals, offering a direct representation of user intent, present new opportunities for image customization. However, it faces challenges in brain interpretation, cross-modal context fusion an…

Cited by 0SourcePDFScholar
2025

Multilingual Blending: Large Language Model Safety Alignment Evaluation with Language Mixture

NAACL 2025findings

As safety remains a crucial concern throughout the development lifecycle of Large Language Models (LLMs), researchers and industrial practitioners have increasingly focused on safeguarding and aligning LLM behaviors with human preferences and ethical standards. LLMs, trained on extensive multilingua…

Cited by 0SourcePDFScholar
2025

SpikeGS: Reconstruct 3D Scene Captured by a Fast-Moving Bio-Inspired Camera

AAAI 2025technical

3D Gaussian Splatting (3DGS) has been proven to exhibit exceptional performance in reconstructing 3D scenes. However, the effectiveness of 3DGS heavily relies on sharp images, and fulfilling this requirement presents challenges in real-world scenarios particularly when utilizing fast-moving cameras.…

Cited by 0SourcePDFScholar
2025

TESTEVAL: Benchmarking Large Language Models for Test Case Generation

NAACL 2025findings

For program languages, testing plays a crucial role in the software development cycle, enabling the detection of bugs, vulnerabilities, and other undesirable behaviors. To perform software testing, testers need to write code snippets that execute the program under test. Recently, researchers have re…

2025

Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling

ACL 2025long

Expressive zero-shot voice conversion (VC) is a critical and challenging task that aims to transform the source timbre into an arbitrary unseen speaker while preserving the original content and expressive qualities. Despite recent progress in zero-shot VC, there remains considerable potential for im…

2025

TreeFinder: A US-Scale Benchmark Dataset for Individual Tree Mortality Monitoring Using High-Resolution Aerial Imagery

NeurIPS 2025poster

Monitoring individual tree mortality at scale has been found to be crucial for understanding forest loss, ecosystem resilience, carbon fluxes, and climate-induced impacts. However, the fine-granularity monitoring faces major challenges on both the data and methodology sides because: (1) finding isol…

Cited by 0SourcecodeScholar
2024

CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models

ACL 2024findings

The advancement of large language models (LLMs) has enhanced the ability to generalize across a wide range of unseen natural language processing (NLP) tasks through instruction-following.Yet, their effectiveness often diminishes in low-resource languages like Chinese, exacerbated by biased evaluatio…

2024

Correspondence-Free Non-Rigid Point Set Registration Using Unsupervised Clustering Analysis

CVPR 2024highlight

This paper presents a novel non-rigid point set registration method that is inspired by unsupervised clustering analysis. Unlike previous approaches that treat the source and target point sets as separate entities we develop a holistic framework where they are formulated as clustering centroids and…

2024

GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition

ICASSP 2024accepted

Contrastive cross-modality pretraining has recently exhibited impressive success in diverse fields, whereas there is limited research on their merits in speech emotion recognition (SER). In this paper, we propose GEmo-CLAP, a kind of gender-attribute-enhanced contrastive language-audio pretraining (…

Cited by 0SourceScholar
2024

ISR-LLM: Iterative Self-Refined Large Language Model for Long-Horizon Sequential Task Planning

ICRA 2024poster

Motivated by the substantial achievements of Large Language Models (LLMs) in the field of natural language processing, recent research has commenced investigations into the application of LLMs for complex, long-horizon sequential task planning challenges in robotics. LLMs are advantageous in offerin…

Cited by 73SourcecodeScholar
2024

LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking Attacks

ICLR 2024poster

Visual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art (SOTA) trackers often fail when faced with adversarial pertur…

2024

Learning from Pattern Completion: Self-supervised Controllable Generation

NeurIPS 2024poster

The human brain exhibits a strong ability to spontaneously associate different visual attributes of the same or similar visual scene, such as associating sketches and graffiti with real-world visual objects, usually without supervising information. In contrast, in the field of artificial intelligenc…

2024

Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization

ICLR 2024spotlight

The out-of-distribution (OOD) problem generally arises when neural networks encounter data that significantly deviates from the training data distribution, i.e., in-distribution (InD). In this paper, we study the OOD problem from a neuron activation view. We first formulate neuron activation states…

2024

PrivAuditor: Benchmarking Data Protection Vulnerabilities in LLM Adaptation Techniques

NeurIPS 2024spotlight

Large Language Models (LLMs) are recognized for their potential to be an important building block toward achieving artificial general intelligence due to their unprecedented capability for solving diverse tasks. Despite these achievements, LLMs often underperform in domain-specific tasks without tra…

Cited by 2SourcePDFScholar
2024

Retrospective for the Dynamic Sensorium Competition for predicting large-scale mouse primary visual cortex activity from videos

NeurIPS 2024poster

Understanding how biological visual systems process information is challenging because of the nonlinear relationship between visual input and neuronal responses. Artificial neural networks allow computational neuroscientists to create predictive models that connect biological and machine vision. Ma…

Cited by 3SourcePDFScholar
2023

DeepGemini: Verifying Dependency Fairness for Deep Neural Network

AAAI 2023technical

Deep neural networks (DNNs) have been widely adopted in many decision-making industrial applications. Their fairness issues, i.e., whether there exist unintended biases in the DNN, receive much attention and become critical concerns, which can directly cause negative impacts in our daily life and po…

2023

Evading DeepFake Detectors via Adversarial Statistical Consistency

CVPR 2023poster

In recent years, as various realistic face forgery techniques known as DeepFake improves by leaps and bounds, more and more DeepFake detection techniques have been proposed. These methods typically rely on detecting statistical differences between natural (i.e., real) and DeepFake-generated images i…

Cited by 59SourcePDFScholar
2023

Hybridformer: Improving Squeezeformer with Hybrid Attention and NSR Mechanism

ICASSP 2023accepted

SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attention (SA). In addition, limited by the large convolution kernel size, the local modeling ability of SqueezeFormer is insuf…

Cited by 0SourceScholar
2023

Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias

NeurIPS 2023poster

The scarcity of data presents a critical obstacle to the efficacy of medical vision-language pre-training (VLP). A potential solution lies in the combination of datasets from various language communities. Nevertheless, the main challenge stems from the complexity of integrating diverse syntax and se…

2023

Neural Episodic Control with State Abstraction

ICLR 2023top-25%

Existing Deep Reinforcement Learning (DRL) algorithms suffer from sample inefficiency. Generally, episodic control-based approaches are solutions that leverage highly rewarded past experiences to improve sample efficiency of DRL algorithms. However, previous episodic control-based approaches fail to…

Cited by 14SourcePDFScholar
2023

Unsupervised Optical Flow Estimation with Dynamic Timing Representation for Spike Camera

NeurIPS 2023poster

Efficiently selecting an appropriate spike stream data length to extract precise information is the key to the spike vision tasks. To address this issue, we propose a dynamic timing representation for spike streams. Based on multi-layers architecture, it applies dilated convolutions on temporal dime…

2022

Modeling The Detection Capability Of High-Speed Spiking Cameras

ICASSP 2022accepted

The novel working principle enables spiking cameras to capture high-speed moving objects. However, the applications of spiking cameras can be affected by many factors, such as brightness intensity, detectable distance, and the maximum speed of moving targets. Improper settings such as weak ambient b…

Cited by 0SourceScholar
2022

Optimal Time Trajectory Generation and Tracking Control for Over-Actuated Multirotors With Large-Angle Maneuvering Capability

RA-L 2022

This paper presents an optimal time trajectory generation method for over-actuated multirotors. Different from underactuated multi-rotors that can only track a 4-D trajectory, over-actuated multi-rotors have the ability to track a 6-D trajectory. The proposed method can generate a 3-degree of freedo

Cited by 6SourceScholar
2021

Decision-Guided Weighted Automata Extraction from Recurrent Neural Networks

AAAI 2021technical

Recurrent Neural Networks (RNNs) have demonstrated their effectiveness in learning and processing sequential data (e.g., speech and natural language). However, due to the black-box nature of neural networks, understanding the decision logic of RNNs is quite challenging. Some recent progress has been…

Cited by 25SourcePDFScholar
2021

EfficientDeRain: Learning Pixel-wise Dilation Filtering for High-Efficiency Single-Image Deraining

AAAI 2021technical

Single-image deraining is rather challenging due to the unknown rain model. Existing methods often make specific assumptions of the rain model, which can hardly cover many diverse circumstances in the real world, compelling them to employ complex optimization or progressive refinement. This, however…

2021

Learning To Adversarially Blur Visual Object Tracking

ICCV 2021poster

Motion blur caused by the moving of the object or camera during the exposure can be a key challenge for visual object tracking, affecting tracking accuracy significantly. In this work, we explore the robustness of visual object trackers against motion blur from a new angle, i.e., adversarial blur at…

Cited by 60PDFcodeScholar
2021

RNNRepair: Automatic RNN Repair via Model-based Analysis

ICML 2021spotlight

Deep neural networks are vulnerable to adversarial attacks. Due to their black-box nature, it is rather challenging to interpret and properly repair these incorrect behaviors. This paper focuses on interpreting and repairing the incorrect behaviors of Recurrent Neural Networks (RNNs). We propose a l…

Cited by 25SourcePDFScholar
2021

Waypoints updating based on Adam and ILC for path learning in physical human-robot interaction

ICRA 2021poster

This paper presents a novel method for learning and tracking of the desired path of the human partner in physical human-robot interaction. Combining the Adam optimization algorithm with iteration learning control (ILC), a path learning method is designed to generate and update reference waypoints ac…

Cited by 10SourceScholar
2020

FakeSpotter: A Simple yet Robust Baseline for Spotting AI-Synthesized Fake Faces

IJCAI 2020poster

In recent years, generative adversarial networks (GANs) and its variants have achieved unprecedented success in image synthesis. They are widely adopted in synthesizing facial images which brings potential security concerns to humans as the fakes spread and fuel the misinformation. However, robust d…

2020

SPARK: Spatial-aware Online Incremental Attack Against Visual Tracking

ECCV 2020poster

Adversarial attacks of deep neural networks have been intensively studied on image, audio, natural language, patch, and pixel classification tasks. Nevertheless, as a typical, while important real-world application, the adversarial attacks of online video object tracking that traces an object's movi…

Cited by 112SourcePDFScholar
2020

Watch out! Motion is Blurring the Vision of Your Deep Neural Networks

NeurIPS 2020poster

The state-of-the-art deep neural networks (DNNs) are vulnerable against adversarial examples with additive random-like noise perturbations. While such examples are hardly found in the physical world, the image blurring effect caused by object motion, on the other hand, commonly occurs in practice, m…