← Search

Zhiqiang Ma

12 accepted papers

2026

LAOF: Robust Latent Action Learning with Optical Flow Constraints

CVPR 2026

Learning latent actions from large-scale videos is crucial for the pre-training of scalable embodied foundation models, yet existing methods often struggle with action-irrelevant distractors. Although incorporating action supervision can alleviate these distractions, its effectiveness is restricted

Cited by 0SourcecodeScholar
2026

Perturb Your Data: Paraphrase-Guided Training Data Watermarking

AAAI 2026technical

Training data detection is critical for enforcing copyright and data licensing, as Large Language Models (LLM) are trained on massive text corpora scraped from the internet. We present SPECTRA, a watermarking approach that makes training data reliably detectable even when it comprises less than 0.00

Cited by 0SourcePDFScholar
2025

CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation

ACL 2025long

Due to their ability to process long and complex contexts, LLMs can offer key benefits to the Legal domain, but their adoption has been hindered by their tendency to generate unfaithful, ungrounded, or hallucinatory outputs. While Retrieval-Augmented Generation offers a promising solution by groundi…

Cited by 0SourcePDFScholar
2025

VDTF-ACT: ACT-based Multimodal Space Fine Manipulation Method with Visual Depth Tactile Fusion

IROS 2025

Autonomous fine manipulation in space for orbital assembly continues to present a critical challenge in the field of aerospace engineering. Under low-gravity conditions, during satellite manipulator operations on free-floating objects, the absence of significant gravitational forces and friction con

Cited by 0SourcecodeScholar
2024

Aligning Human Intent From Imperfect Demonstrations With Confidence-Based Inverse Soft-Q Learning

RA-L 2024

Imitation learning attracts much attention for its ability to allow robots to quickly learn human manipulation skills through demonstrations. However, in the real world, human demonstrations often exhibit random behavior that is not intended by humans. Collecting high-quality human datasets is both

Cited by 4SourcecodeScholar
2024

DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding

ACL 2024long

Enterprise documents such as forms, receipts, reports, and other such records, often carry rich semantics at the intersection of textual and spatial modalities. The visual cues offered by their complex layouts play a crucial role in comprehending these documents effectively. In this paper, we presen…

2024

Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation

EMNLP 2024finding

Language models are capable of memorizing detailed patterns and information, leading to a double-edged effect: they achieve impressive modeling performance on downstream tasks with the stored knowledge but also raise significant privacy concerns. Traditional differential privacy based training appro…

Cited by 1SourcePDFScholar
2024

The State of the Art of Large Language Models on Chartered Financial Analyst Exams

EMNLP 2024industry

The Chartered Financial Analyst (CFA) program is one of the most widely recognized financial certifications globally. In this work, we test a variety of state-of-the-art large language models (LLMs) on mock CFA exams to provide an overview of their financial analysis capabilities using the same eval…

Cited by 2SourcePDFScholar
2024

“What is the value of templates?” Rethinking Document Information Extraction Datasets for LLMs

EMNLP 2024finding

The rise of large language models (LLMs) for visually rich document understanding (VRDU) has kindled a need for prompt-response, document-based datasets. As annotating new datasets from scratch is labor-intensive, the existing literature has generated prompt-response datasets from available resource…

Cited by 0SourcePDFScholar
2023

Reducing the GAP Between Streaming and Non-Streaming Transducer-Based ASR by Adaptive Two-Stage Knowledge Distillation

ICASSP 2023accepted

Transducer is one of the mainstream frameworks for streaming speech recognition. There is a performance gap between the streaming and non-streaming transducer models due to limited context. To reduce this gap, an effective way is to ensure that their hidden and output distributions are consistent, w…

Cited by 0SourceScholar
2022

ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering

EMNLP 2022main

With the recent advance in large pre-trained language models, researchers have achieved record performances in NLP tasks that mostly focus on language pattern matching. The community is experiencing the shift of the challenge from how to model language to the imitation of complex reasoning abilities…

2022

Maximized Hydrodynamic Stimulation Strategy for Placement of Differential Pressure and Velocity Sensors in Artificial Lateral Line Systems

RA-L 2022

Fish can perceive the surrounding flow field using their lateral line systems, consisting of canal neuromasts (CNs) for flow pressure gradient perception and superficial neuromasts (SNs) for flow velocity detection. Although various artificial lateral line (ALL) systems have been developed inspired

Cited by 21SourceScholar