← Search

Huimin Ma

39 accepted papers

2026

MEDUSA: Motion Elimination in Diffusion Using Spectral Attack

ICML 2026poster

With the widespread application of Video Diffusion Models (VDMs), video synthesis has achieved remarkable temporal dynamics. Image-to-Video (I2V) generation allows users to provide reference images, which enables attackers to inject adversarial noise into these conditions. Due to the robust spatio-t…

Cited by 0SourceScholar
2026

Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images

CVPR 2026

Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D vehicle generation methods are often trained on synthetic data with significant domain gaps from real-world distributions. The generated models ofte

Cited by 0SourcecodeScholar
2026

Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models

CVPR 2026

As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human-like Theory of Mind (ToM). Most existing ToM evaluations, however, are centered on text-based inputs, while scenarios relying solely on visual infor

Cited by 0SourceScholar
2025

A novel multimodal personality prediction method based on pretrained models and graph relational transformer network

ICASSP 2025accepted

Multimodal personality analysis aims to identify and express human personality traits in videos. However, RNN and its variants have a limited ability to learn long-term temporal dependencies and existing methods neglect bimodal association features. Based on the fact that visual modalities play a do…

Cited by 0SourceScholar
2025

AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios

NeurIPS 2025poster

By sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous work focus on vehicle-to-vehicle and vehicle-to-infrastructure collaboration, with limited attention to aerial perspectiv…

Cited by 0SourcecodeScholar
2025

A²RNet: Adversarial Attack Resilient Network for Robust Infrared and Visible Image Fusion

AAAI 2025technical

Infrared and visible image fusion (IVIF) is a crucial technique for enhancing visual performance by integrating unique information from different modalities into one fused image. Exiting methods pay more attention to conducting fusion with undisturbed data, while overlooking the impact of deliberate…

2025

DADet: Safeguarding Image Conditional Diffusion Models against Adversarial and Backdoor Attacks via Diffusion Anomaly Detection

ICCV 2025poster

While image conditional diffusion models demonstrate impressive generation capabilities, they exhibit high vulnerability when facing backdoor and adversarial attacks. In this paper, we define a scenario named diffusion anomaly where the generated results of a reverse process under attack deviate sig…

Cited by 0SourcePDFScholar
2025

Enhancing Autonomous Driving through Dual-Process Learning with Behavior and Reflection Integration

ICASSP 2025accepted

Contemporary autonomous driving (AD) methodologies, which predominantly convert visual features into control directives, face long-tail challenges due to constraints imposed by limited data distribution. Conversely, human drivers exhibit proficiency in such conditions, underscoring the significance…

Cited by 0SourceScholar
2025

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models

ICML 2025poster

As large language models evolve, there is growing anticipation that they will emulate human-like Theory of Mind (ToM) to assist with routine tasks. However, existing methods for evaluating machine ToM focus primarily on unimodal models and largely treat these models as black boxes, lacking an interp…

Cited by 0SourcePDFScholar
2025

Image-to-video Adaptation with Outlier Modeling and Robust Self-learning

AAAI 2025technical

The image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Exis…

2025

Kaleidoscopic Background Attack: Disrupting Pose Estimation with Multi-Fold Radial Symmetry Textures

ICCV 2025poster

Camera pose estimation is a fundamental computer vision task that is essential for applications like visual localization and multi-view stereo reconstruction. In the object-centric scenarios with sparse inputs, the accuracy of pose estimation can be significantly influenced by background textures th…

Cited by 0SourcePDFScholar
2025

MVSMamba: Multi-View Stereo with State Space Model

NeurIPS 2025poster

Robust feature representations are essential for learning-based Multi-View Stereo (MVS), which relies on accurate feature matching. Recent MVS methods leverage Transformers to capture long-range dependencies based on local features extracted by conventional feature pyramid networks. However, the qua…

Cited by 0SourcecodeScholar
2025

MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network

ICCV 2025poster

Learning-based Multi-View Stereo (MVS) methods aim to predict depth maps for a sequence of calibrated images to recover dense point clouds. However, existing MVS methods often struggle with challenging regions, such as textureless regions and reflective surfaces, where feature matching fails. In con…

2025

ProtoCar: Learning 3D Vehicle Prototypes from Single-View and Unconstrained Driving Scene Images

AAAI 2025technical

Reconstructing 3D models from sensor data is a valuable and promising direction for developing testing and validation environments in applications like autonomous driving. However, existing methods for 3D modeling often rely on extensive multi-view data or controlled conditions, making them difficul…

Cited by 0SourcePDFScholar
2025

RRT-MVS: Recurrent Regularization Transformer for Multi-View Stereo

AAAI 2025technical

Learning-based multi-view stereo methods aim to predict depth maps for reconstructing dense point clouds. These methods rely on regularization to reduce redundancy in the cost volume. However, existing methods have limitations: CNN-based regularization is restricted to local receptive fields, while…

Cited by 0SourcePDFScholar
2025

RhythmMamba: Fast, Lightweight, and Accurate Remote Physiological Measurement

AAAI 2025technical

Remote photoplethysmography (rPPG) is a method for non-contact measurement of physiological signals from facial videos, holding great potential in various applications such as healthcare, affective computing, and anti-spoofing. Existing deep learning methods struggle to address two core issues of rP…

2025

SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance Segmentation

CVPR 2025poster

In the field of zero-shot 3D instance segmentation, existing 2D-to-3D lifting methods typically obtain 2D segmentation across multiple RGB frames using vision foundation models, which are then projected and merged into 3D space. However, since the inference of vision foundation models on a single fr…

2025

Synergistic Spotting and Recognition of Micro-Expression via Temporal State Transition

ICASSP 2025accepted

Micro-expressions are involuntary facial movements that cannot be consciously controlled, conveying subtle cues with substantial real-world applications. The analysis of micro-expressions generally involves two main tasks: spotting micro-expression intervals in long videos and recognizing the emotio…

Cited by 0SourceScholar
2024

Center of Pressure Estimation by Analyzing Walking Videos

ICASSP 2024accepted

Center of pressure (COP) serves as a widely utilized indicator for evaluating balance-related issues, e.g., gait quality of neurological disorders, fall risk of the elderly, and recovery of the injured. Existing methods for acquiring COP mostly rely on expensive force platforms or wearable force-sen…

Cited by 0SourceScholar
2024

Enhancing Adversarial Transferability in Object Detection with Bidirectional Feature Distortion

ICASSP 2024accepted

Previous works have shown that perturbing internal-layer features can significantly enhance the transferability of black-box attacks in classifiers. However, these methods have not achieved satisfactory performance when applied to detectors due to the inherent differences in features between detecto…

Cited by 0SourceScholar
2024

ODGEN: Domain-specific Object Detection Data Generation with Diffusion Models

NeurIPS 2024poster

Modern diffusion-based image generative models have made significant progress and become promising to enrich training data for the object detection task. However, the generation quality and the controllability for complex scenes containing multi-class objects and dense objects with occlusions remain…

Cited by 5SourcePDFScholar
2024

PLS: Unsupervised Domain Adaptation for 3d Object Detection Via Pseudo-Label Sizes

ICASSP 2024accepted

3D object detection has gained increasing attention in modern autonomous driving systems. However, the performance of the detector significantly degrades during cross-domain deployment due to domain shift. The detector is inevitably biased towards its training dataset when employed on a target datas…

Cited by 0SourceScholar
2024

Step Vulnerability Guided Mean Fluctuation Adversarial Attack against Conditional Diffusion Models

AAAI 2024technical

The high-quality generation results of conditional diffusion models have brought about concerns regarding privacy and copyright issues. As a possible technique for preventing the abuse of diffusion models, the adversarial attack against diffusion models has attracted academic attention recently. In…

2024

Transferable Adversarial Attacks for Object Detection Using Object-Aware Significant Feature Distortion

AAAI 2024technical

Transferable black-box adversarial attacks against classifiers by disturbing the intermediate-layer features have been extensively studied in recent years. However, these methods have not yet achieved satisfactory performances when directly applied to object detectors. This is largely because the fe…

2024

Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and Visual Analysis Strategy

ECCV 2024poster

"In the domain of Camouflaged Object Segmentation (COS), despite continuous improvements in segmentation performance, the underlying mechanisms of effective camouflage remain poorly understood, akin to a black box. To address this gap, we present the first comprehensive study to examine the impact o…

2024

Unveiling the Dynamics of Information Interplay in Supervised Learning

ICML 2024poster

In this paper, we use matrix information theory as an analytical tool to analyze the dynamics of the information interplay between data representations and classification head vectors in the supervised learning process. Specifically, inspired by the theory of Neural Collapse, we introduce matrix mut…

Cited by 3SourcePDFScholar
2024

Upper-body Hierarchical Graph for Skeleton Based Emotion Recognition in Assistive Driving

ECCV 2024poster

"Emotion recognition plays a crucial role in enhancing the safety and enjoyment of assistive driving experiences. By enabling intelligent systems to perceive and understand human emotions, we can significantly improve human-machine interactions. Current research in emotion recognition emphasizes fac…

2023

Defending Against Universal Patch Attacks by Restricting Token Attention in Vision Transformers

ICASSP 2023accepted

Previous works reveal that similar to CNNs, vision transformers (ViT) are also vulnerable to universal adversarial patch attacks. In this paper, we empirically reveal and mathematically explain that the shallow tokens in the transformer and the attention of the network can largely influence the clas…

Cited by 0SourceScholar
2023

FD-Align: Feature Discrimination Alignment for Fine-tuning Pre-Trained Models in Few-Shot Learning

NeurIPS 2023poster

Due to the limited availability of data, existing few-shot learning methods trained from scratch fail to achieve satisfactory performance. In contrast, large-scale pre-trained models such as CLIP demonstrate remarkable few-shot and zero-shot capabilities. To enhance the performance of pre-trained mo…

2021

Defending Against Universal Adversarial Patches by Clipping Feature Norms

ICCV 2021poster

Physical-world adversarial attacks based on universal adversarial patches have been proved to be able to mislead deep convolutional neural networks (CNNs), exposing the vulnerability of real-world visual classification systems based on CNNs. In this paper, we empirically reveal and mathematically ex…

Cited by 37PDFScholar
2021

Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent Problems

NeurIPS 2021poster

We introduce an automatic curriculum algorithm, Variational Automatic Curriculum Learning (VACL), for solving challenging goal-conditioned cooperative multi-agent reinforcement learning problems. We motivate our curriculum learning paradigm through a variational perspective, where the learning objec…

Cited by 45SourcePDFScholar
2018

Weakly-Supervised Semantic Segmentation by Iteratively Mining Common Object Features

CVPR 2018poster

Weakly-supervised semantic segmentation under image tags supervision is a challenging task as it directly associates high-level semantic to low-level appearance. To bridge this gap, in this paper, we propose an iterative bottom-up and top-down framework which alternatively expands object regions and…

Cited by 375SourcePDFScholar
2016

Monocular 3D Object Detection for Autonomous Driving

CVPR 2016poster

The goal of this paper is to perform 3D object detection in single monocular images in the domain of autonomous driving. Our method first aims to generate a set of candidate class-specific object proposals, which are then run through a standard CNN pipeline to obtain high-quality object detections.…

Cited by 1263PDFScholar
2015

3D Object Proposals for Accurate Object Class Detection

NeurIPS 2015poster

The goal of this paper is to generate high-quality 3D object proposals in the context of autonomous driving. Our method exploits stereo imagery to place proposals in the form of 3D bounding boxes. We formulate the problem as minimizing an energy function encoding object size priors, ground plane a…

Cited by 1092SourcePDFScholar
2015

Improving Object Proposals With Multi-Thresholding Straddling Expansion

CVPR 2015poster

Recent advances in object detection have exploited object proposals to speed up object searching. However, many of existing object proposal generators have strong localization bias or require computationally expensive diversification strategies. In this paper, we present an effective approach to add…

Cited by 96SourcePDFScholar