← Search

Ziming Li

10 accepted papers

2025

DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation

ICCV 2025poster

Non-prehensile manipulation is crucial for handling objects that are too thin, large, or otherwise ungraspable in unstructured environments. While conventional planning-based approaches struggle with complex contact modeling, learning-based methods have recently emerged as a promising alternative. H…

Cited by 0SourcePDFScholar
2025

Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests

COLING 2025main

Indirect User Requests (IURs), such as “It’s cold in here” instead of “Could you please increase the temperature?” are common in human-human task-oriented dialogue and require world knowledge and pragmatic reasoning from the listener. While large language models (LLMs) can handle these requests effe…

Cited by 1SourcePDFScholar
2024

Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction

EMNLP 2024main

Automated red teaming is an effective method for identifying misaligned behaviors in large language models (LLMs). Existing approaches, however, often focus primarily on improving attack success rates while overlooking the need for comprehensive test case coverage. Additionally, most of these method…

2024

MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning

ACL 2024findings

Large language models (LLMs), despite their remarkable progress across various general domains, encounter significant barriers in medicine and healthcare. This field faces unique challenges such as domain-specific terminologies and reasoning over specialized knowledge. To address these issues, we pr…

2023

PartManip: Learning Cross-Category Generalizable Part Manipulation Policy From Point Cloud Observations

CVPR 2023poster

Learning a generalizable object manipulation policy is vital for an embodied agent to work in complex real-world scenes. Parts, as the shared components in different object categories, have the potential to increase the generalization ability of the manipulation policy and achieve cross-category obj…

Cited by 40SourcePDFScholar
2023

QAP: A Quantum-Inspired Adaptive-Priority-Learning Model for Multimodal Emotion Recognition

ACL 2023findings

Multimodal emotion recognition for video has gained considerable attention in recent years, in which three modalities (i.e., textual, visual and acoustic) are involved. Due to the diverse levels of informational content related to emotion, three modalities typically possess varying degrees of contri…

Cited by 16SourcePDFScholar
2023

TrojanSQL: SQL Injection against Natural Language Interface to Database

EMNLP 2023long main

The technology of text-to-SQL has significantly enhanced the efficiency of accessing and manipulating databases. However, limited research has been conducted to study its vulnerabilities emerging from malicious user interaction. By proposing TrojanSQL, a backdoor-based SQL injection framework for t…

Cited by 0SourceScholar
2022

AMOA: Global Acoustic Feature Enhanced Modal-Order-Aware Network for Multimodal Sentiment Analysis

COLING 2022main

In recent years, multimodal sentiment analysis (MSA) has attracted more and more interest, which aims to predict the sentiment polarity expressed in a video. Existing methods typically 1) treat three modal features (textual, acoustic, visual) equally, without distinguishing the importance of differe…

Cited by 25SourcePDFScholar