← Search

Qingyuan Hu

6 accepted papers

2024

PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs

ACL 2024findings

While Large language models (LLMs) have demonstrated considerable capabilities across various natural language tasks, they often fall short of the performance achieved by domain-specific state-of-the-art models. One potential approach to enhance domain-specific capabilities of LLMs involves fine-tun…

2024

Position: Towards Unified Alignment Between Agents, Humans, and Environment

ICML 2024poster

The rapid progress of foundation models has led to the prosperity of autonomous agents, which leverage the universal capabilities of foundation models to conduct reasoning, decision-making, and environmental interaction. However, the efficacy of agents remains limited when operating in intricate, re…

Cited by 4SourcePDFScholar
2023

ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos

EMNLP 2023long main

Multimodal counterfactual reasoning is a vital yet challenging ability for AI systems. It involves predicting the outcomes of hypothetical circumstances based on vision and language inputs, which enables AI models to learn from failures and explore hypothetical scenarios. Despite its importance, the…

Cited by 0SourcecodeScholar
2023

Learning Action Conditions from Instructional Manuals for Instruction Understanding

ACL 2023long

The ability to infer pre- and postconditions of an action is vital for comprehending complex instructions, and is essential for applications such as autonomous instruction-guided agents and assistive AI that supports humans to perform physical tasks. In this work, we propose a task dubbed action con…

2020

Stochastic Batch Augmentation with An Effective Distilled Dynamic Soft Label Regularizer

IJCAI 2020poster

Data augmentation have been intensively used in training deep neural network to improve the generalization, whether in original space (e.g., image space) or representation space. Although being successful, the connection between the synthesized data and the original data is largely ignored in traini…

Cited by 0SourcePDFScholar