← Search

Wei Ma

35 accepted papers

2026

CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics

ICLR 2026poster

We introduce CMPhysBench, designed to assess the proficiency of Large Language Models (LLMs) in Condensed Matter Physics, as a novel Benchmark. CMPhysBench is composed of more than 520 graduate-level meticulously curated questions covering both representative subfields and foundational theoretical f…

Cited by 0SourcecodeScholar
2026

DualScope: Capturing Critical Spatial and Temporal Cues for Distracted Driving Activity Recognition

AAAI 2026technical

Accurately recognizing distracted driving activities in real-world scenarios is essential for improving road and pedestrian safety. However, existing approaches are prone to attending to irrelevant scene context and are susceptible to interference from redundant frames, compromising their robustness

Cited by 0SourcePDFScholar
2026

E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving

CVPR 2026

End-to-end autonomous driving (AD) systems increasingly adopt vision-language-action (VLA) models, yet they ignore the passenger's emotional state, which is central to comfort and AD acceptance. We introduce Open-Domain End-to-End (OD-E2E) AD, where an autonomous vehicle must interpret free-form nat

Cited by 0SourceScholar
2026

From Sequential to Recursive: Enhancing Decision-Focused Learning with Bidirectional Feedback

AAAI 2026technical

Decision-focused learning (DFL) has emerged as a powerful end-to-end alternative to conventional predict-then-optimize (PTO) pipelines by directly optimizing predictive models through downstream decision losses. Existing DFL frameworks are limited by their strictly sequential structure, referred to

Cited by 0SourcePDFScholar
2026

ProCURE: Addressing the Programming Concept Understanding Gap for Code Generation in LLMs via Concept-Aware Consistency Learning

IJCAI 2026

Although Large Language Models (LLMs) excel at code generation, recent research reveals that they exhibit an insufficient grasp of core programming concepts, such as data flow and control flow. This limitation undermines their robustness when encountering variations in these concepts in practice; ho

Cited by 0Scholar
2026

ROI-GSurFisher: Next Best View Selection for Active Gaussian Splatting Via Fisher Information of ROI-Selected Gaussian Surfels

ICRA 2026poster

Next Best View (NBV) selection is critical for achieving high-quality 3D reconstruction in unknown environments. This paper presents an active NBV selection approach tailored for Gaussian Splatting (GS), a widely adopted 3D reconstruction technique that has recently gained significant attention and …

Cited by 0codeScholar
2026

RealNet: Efficient and Unsupervised Detection of AI-Generated Images via Real-Only Representation Learning

AAAI 2026technical

Detecting AI-generated images remains a persistent challenge, as existing detectors often struggle to generalize to forgeries produced by previously unseen generative models. This generalization gap mainly stems from entanglement with semantic content and overfitting to model-specific artifacts. Mor

Cited by 0SourcePDFScholar
2026

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

ICML 2026poster

Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high training efficiency. Although these EDistilled LLMs achieve state-of-the-art (SOTA) performance on general ability benchmarks relative to similarly size…

Cited by 0SourceScholar
2026

SUBTA: A Framework for Supported User-Guided Bimanual Teleoperation in Structured Assembly

ICRA 2026poster

In human-robot collaboration, shared autonomy enhances human performance through precise, intuitive support. Effective robotic assistance requires accurately inferring human intentions and understanding task structures to determine optimal support timing and methods. In this paper, we present SUBTA,…

2026

Steerable Adversarial Scenario Generation through Test-Time Preference Alignment

ICLR 2026poster

Adversarial scenario generation is a cost-effective approach for safety assessment of autonomous driving systems. However, existing methods are often constrained to a single, fixed trade-off between competing objectives such as adversariality and realism. This yields behavior-specific models that c…

Cited by 0SourcecodeScholar
2025

3SAT: A Simple Self-Supervised Adversarial Training Framework

AAAI 2025technical

The combination of self-supervised learning and adversarial training (AT) can significantly improve the adversarial robustness of self-supervised models. However, the robustness of self-supervised adversarial training (self-AT) still lags behind that of state-of-the-art (SOTA) supervised AT (sup-AT)…

2025

A Federated Learning-Based Intrusion Detection System for Satellite-Terrestrial Integrated Networks

ICASSP 2025accepted

The emergence of Satellite-Terrestrial Integrated Networks (STIN) has significantly expanded terrestrial network coverage but introduced new security threats. Current Intrusion Detection Systems (IDSs) for STIN mostly consider the distributed nature of satellites, overlooking the computational limit…

Cited by 0SourceScholar
2025

DASSL: Domain Agnostic Self-Supervised Learning with Multiple Missing Information Reconstruction Branches

ICASSP 2025accepted

Self-supervised learning (SSL) is a technique used to learn feature representations from unlabeled data. However, existing SSL frameworks either rely too heavily on domain knowledge due to their design based on feature invariance, leading to a lack of domain transferability, or they are based on aut…

Cited by 0SourceScholar
2025

Geolocation Representation from Large Language Models Are Generic Enhancers for Spatio-Temporal Learning

AAAI 2025technical

In the geospatial domain, universal representation models are significantly less prevalent than their extensive use in natural language processing and computer vision. This discrepancy arises primarily from the high costs associated with the input of existing representation models, which often requi…

Cited by 7SourcePDFScholar
2025

Online Identification of Equivalent Roll Center for Underwater Gliders to Weaken the Repeated Yawing

RA-L 2025

Underwater gliders (UG) experience unexpected yaw during each dormancy stage of the controller, necessitating additional adjustments, which manifests as repeated yawing throughout the entire profiling time. Repeated yawing is frequently observed during the long-term deployment of the UGs, which can

Cited by 1SourceScholar
2025

Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning

EMNLP 2025

Vision-language models (VLMs) excel in many downstream tasks but struggle with spatial reasoning, which is crucial for navigation and interaction with physical environments. Specifically, many spatial reasoning tasks rely on fundamental two-dimensional (2D) capabilities, yet our evaluation shows tha

2024

DEIE: Benchmarking Document-level Event Information Extraction with a Large-scale Chinese News Dataset

COLING 2024main

A text corpus centered on events is foundational to research concerning the detection, representation, reasoning, and harnessing of online events. The majority of current event-based datasets mainly target sentence-level tasks, thus to advance event-related research spanning from sentence to documen…

2024

Fast and Accurate Root Cause Analysis Based on Signalling Messages for 5G Networks

ICASSP 2024accepted

The ever-increasing complexity and scale of 5G communication networks pose huge challenges to network operations. Root cause analysis is considered as a promising method for fault detection. However, it still suffers challenges of severely uneven distribution of fault data, low accuracy in root caus…

Cited by 0SourceScholar
2024

ItiNera: Integrating Spatial Optimization with Large Language Models for Open-domain Urban Itinerary Planning

EMNLP 2024industry

Citywalk, a recently popular form of urban travel, requires genuine personalization and understanding of fine-grained requests compared to traditional itinerary planning. In this paper, we introduce the novel task of Open-domain Urban Itinerary Planning (OUIP), which generates personalized urban iti…

2024

Manticore: An Unsupervised Intrusion Detection System Based on Contrastive Learning in 5G Networks

ICASSP 2024accepted

The increasing complexity and openness of 5G networks naturally enlarge the attack surface and introduce new vulnerabilities, thereby posing challenges to the performance of existing intrusion detection systems (IDSs). Current IDSs solely rely on statistical features, which may suffer from low accur…

Cited by 0SourceScholar
2024

Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization

NeurIPS 2024poster

Self-supervised learning (SSL) has rapidly advanced in recent years, approaching the performance of its supervised counterparts through the extraction of representations from unlabeled data. However, dimensional collapse, where a few large eigenvalues dominate the eigenspace, poses a significant obs…

Cited by 2SourcePDFScholar
2024

Preventing Model Collapse in Deep Canonical Correlation Analysis by Noise Regularization

NeurIPS 2024poster

Multi-View Representation Learning (MVRL) aims to learn a unified representation of an object from multi-view data. Deep Canonical Correlation Analysis (DCCA) and its variants share simple formulations and demonstrate state-of-the-art performance. However, with extensive experiments, we observe the…

Cited by 0SourcePDFScholar
2024

Subtle Signatures, Strong Shields: Advancing Robust and Imperceptible Watermarking in Large Language Models

ACL 2024findings

The widespread adoption of Large Language Models (LLMs) has led to an increase in AI-generated text on the Internet, presenting a crucial challenge to differentiate AI-created content from human-written text. This challenge is critical to prevent issues of authenticity, trust, and potential copyrigh…

Cited by 3SourcePDFScholar
2024

UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and Localization

CVPR 2024poster

We present UnionFormer a novel framework that integrates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically we construct a BSFI-Net to extract tampering features from RGB and noise views achieving enhanced responsiveness to boundary…

Cited by 10SourcePDFScholar
2023

A Black-Box Attack on Code Models via Representation Nearest Neighbor Search

EMNLP 2023long findings

Existing methods for generating adversarial code examples face several challenges: limted availability of substitute variables, high verification costs for these substitutes, and the creation of adversarial samples with noticeable perturbations. To address these concerns, our proposed approach, RNNS…

Cited by 0SourceScholar
2023

Multi-Local Attention for Speech-Based Depression Detection

ICASSP 2023accepted

This article shows that an attention mechanism, the Multi-Local Attention, can improve a depression detection approach based on Long Short-Term Memory Networks. Besides leading to higher performance metrics (e.g., Accuracy and F1 Score), Multi-Local Attention improves two other aspects of the approa…

Cited by 0SourceScholar
2023

Retrieve-and-Sample: Document-level Event Argument Extraction via Hybrid Retrieval Augmentation

ACL 2023long

Recent studies have shown the effectiveness of retrieval augmentation in many generative NLP tasks. These retrieval-augmented methods allow models to explicitly acquire prior external knowledge in a non-parametric manner and regard the retrieved reference instances as cues to augment text generation…

2022

CLIO: Role-interactive Multi-event Head Attention Network for Document-level Event Extraction

COLING 2022main

Transforming the large amounts of unstructured text on the Internet into structured event knowledge is a critical, yet unsolved goal of NLP, especially when addressing document-level text. Existing methods struggle in Document-level Event Extraction (DEE) due to its two intrinsic challenges: (a) Nes…

Cited by 11SourcePDFScholar
2019

Missing Not at Random in Matrix Completion: The Effectiveness of Estimating Missingness Probabilities Under a Low Nuclear Norm Assumption

NeurIPS 2019poster

Matrix completion is often applied to data with entries missing not at random (MNAR). For example, consider a recommendation system where users tend to only reveal ratings for items they like. In this case, a matrix completion method that relies on entries being revealed at uniformly sampled row and…

2017

Double-bit quantization and weighting for nearest neighbor search

ICASSP 2017accepted

Binary embedding is an effective way for nearest neighbor (NN) search as binary code is storage efficient and fast to compute. It tries to convert real-value signatures into binary codes while preserving similarity of the original data. However, it greatly decreases the discriminability of original…

Cited by 0SourceScholar