← Search

Wentao Mo

6 accepted papers

2026

FROM KNOWING TO DOING PRECISELY: A GENERAL SELF-CORRECTION AND TERMINATION FRAMEWORK FOR VLA MODELS

ICASSP 2026poster

While vision-language-action (VLA) models for embodied agents integrate perception, reasoning, and control, they remain constrained by two critical weaknesses: first, during grasping tasks, the action tokens generated by the language model often exhibit subtle spatial deviations from the target obje…

Cited by 0SourcePDFScholar
2026

TWINFUZZ: Dual-Model Fuzzing for Robustness Generalization in Deep Learning

AAAI 2026technical

Deep learning (DL) models are increasingly deployed in safety-critical applications such as face recognition, autonomous driving, and medical diagnosis. Despite their impressive accuracy, they remain vulnerable to adversarial examples - subtle perturbations that can cause incorrect predictions, i.e.

Cited by 0SourcePDFScholar
2026

Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach

AAAI 2026technical

Video-based multimodal large language models (V-MLLMs) have shown vulnerability to adversarial examples in video-text multimodal tasks. However, the transferability of adversarial videos to unseen models—a common and practical real-world scenario—remains unexplored. In this paper, we pioneer an in

Cited by 0SourcePDFScholar
2024

3D Vision and Language Pretraining with Large-Scale Synthetic Data

IJCAI 2024poster

3D Vision-Language Pre-training (3D-VLP) aims to provide a pre-train model which can bridge 3D scenes with natural language, which is an important technique for embodied intelligence. However, current 3D-VLP datasets are hindered by limited scene-level diversity and insufficient fine-grained annot…

2024

Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA

AAAI 2024technical

In 3D Visual Question Answering (3D VQA), the scarcity of fully annotated data and limited visual content diversity hampers the generalization to novel scenes and 3D concepts (e.g., only around 800 scenes are utilized in ScanQA and SQA dataset). Current approaches resort supplement 3D reasoning with…