← Search

Zhiwei Liu

30 accepted papers

2026

Position: Vector Prompt Interfaces Should Be Exposed to Enable Customization of Large Language Models

ICML 2026poster

As large language models (LLMs) transition from research prototypes to real-world systems, customization has emerged as a central bottleneck. While text prompts can already customize LLM behavior, we argue that text-only prompting does not constitute a suitable control interface for scalable, stable…

Cited by 0SourceScholar
2026

Test-Time Adaptation for LLM Agents via Environment Interaction

ICLR 2026poster

Large language model (LLM)-based agents struggle to generalize to novel and complex environments, such as unseen websites or new sets of functions, due to a fundamental mismatch between their pre-training and test-time conditions. This challenge stems from two distinct failure modes: a syntactic mis…

Cited by 0SourcecodeScholar
2026

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels

ICLR 2026poster

Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more data-efficient solution capable of bridging this gap, yet its appl…

Cited by 0SourcecodeScholar
2025

APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

NeurIPS 2025poster

Training effective AI agents for multi-turn interactions requires high-quality data that captures realistic human-agent dynamics, yet such data is scarce and expensive to collect manually. We introduce APIGen-MT, a two-phase framework that generates verifiable and diverse multi-turn agent data. In t…

Cited by 0SourceScholar
2025

ActionStudio: A Lightweight Framework for Data and Training of Large Action Models

EMNLP 2025

Large Action models are essential for enabling autonomous agents to perform complex tasks. However, training such models remains challenging due to the diversity of agent environments and the complexity of noisy agentic data. Existing infrastructure offers limited support for scalable, agent-specifi

2025

Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents

ICLR 2025poster

Large language model (LLM) agents have shown great potential in solving real-world software engineering (SWE) problems. The most advanced open-source SWE agent can resolve over 27% of real GitHub issues in SWE-Bench Lite. However, these sophisticated agent frameworks exhibit varying strengths, excel…

Cited by 10SourcePDFScholar
2025

LATTE: Learning to Think with Vision Specialists

EMNLP 2025

While open-source vision-language models perform well on simple question-answering, they still struggle with complex questions that require both perceptual and reasoning capabilities. We propose LATTE, a family of vision-language models that have LeArned to Think wiTh vision spEcialists. By offloadi

2025

PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data

ACL 2025finding

Personalization is essential for AI assistants, especially in private AI settings where models are expected to interpret users’ personal data (e.g., conversations, app usage) to understand their background, preferences, and social context. However, due to privacy concerns, existing academic research…

Cited by 23SourcePDFScholar
2025

RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional Information

ACL 2025long

Misinformation is prevalent in various fields such as education, politics, health, etc., causing significant harm to society. However, current methods for cross-domain misinformation detection rely on effort- and resource-intensive fine-tuning and complex model structures. With the outstanding perfo…

2025

Selective Preference Optimization via Token-Level Reward Function Estimation

EMNLP 2025

Recent advancements in LLM alignment leverage token-level supervisions to perform fine-grained preference optimization. However, existing token-level alignment methods either optimize on all available tokens, which can be noisy and inefficient, or perform selective training with complex and expensiv

Cited by 0SourcePDFScholar
2025

VLIMNet: A Visible Light And Infrared Image Matching Network Based On Segment Anything Model And SuperPoint

ICASSP 2025accepted

This paper introduces a novel method for matching visible light and infrared images, termed the Visible Light and Infrared Image Matching Network (VLIMNet). In the image encoding stage, we incorporate a generative architecture-based modality transformation network after the SuperPoint encoder, enabl…

Cited by 0SourceScholar
2025

xLAM: A Family of Large Action Models to Empower AI Agent Systems

NAACL 2025long

Autonomous agents powered by large language models (LLMs) have attracted significant research interest. However, the open-source community faces many challenges in developing specialized models for agent tasks, driven by the scarcity of high-quality agent datasets and the absence of standard protoco…

2024

APIGen: Automated PIpeline for Generating Verifiable and Diverse Function-Calling Datasets

NeurIPS 2024poster

The advancement of function-calling agent models requires diverse, reliable, and high-quality datasets. This paper presents APIGen, an automated data generation pipeline designed to synthesize high-quality datasets for function-calling applications. We leverage APIGen and collect 3,673 executable AP…

2024

FinBen: A Holistic Financial Benchmark for Large Language Models

NeurIPS 2024poster

LLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of comprehensive benchmarks, the rapid development of LLMs, and the complexity of financial tasks. In this paper, we introduce FinBen, the first extensive open-source evaluati…

2024

MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models

NeurIPS 2024poster

Recent advancements in large language models (LLMs) focus on aligning to heterogeneous human expectations and values via multi-objective preference alignment. However, existing methods are dependent on the policy model parameters, which require high-cost repetition of their alignment algorithms for…

2024

PFDM: Parser-Free Virtual Try-On via Diffusion Model

ICASSP 2024accepted

Virtual try-on can significantly improve the garment shopping experiences in both online and in-store scenarios, attracting broad interest in computer vision. However, to achieve high-fidelity try-on performance, most state-of-the-art methods still rely on accurate segmentation masks, which are ofte…

Cited by 0SourceScholar
2024

Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization

ICLR 2024spotlight

Recent months have seen the emergence of a powerful new trend in which large language models (LLMs) are augmented to become autonomous language agents capable of performing objective oriented multi-step tasks on their own, rather than merely responding to queries from human users. Most existing lang…

2023

A 3.4-Millimeter Flea-Sized Robot With Powerful Jumping and Fast Crawling Locomotion

RA-L 2023

Fleas in nature generally have potent muscles for jumping and crawling abilities to overcome obstacles in complex environments, while it is challenging for flea-sized robots to have powerful actuators due to the size effect. This work presents a novel high-voltage pulsed actuator for a flea-sized ro

Cited by 15SourceScholar
2023

CHEER: Centrality-aware High-order Event Reasoning Network for Document-level Event Causality Identification

ACL 2023long

Document-level Event Causality Identification (DECI) aims to recognize causal relations between events within a document. Recent studies focus on building a document-level graph for cross-sentence reasoning, but ignore important causal structures — there are one or two “central” events that prevail…

Cited by 19SourcePDFScholar
2023

Hierarchical Hypergraph Recurrent Attention Network for Temporal Knowledge Graph Reasoning

ICASSP 2023accepted

Temporal knowledge graph (TKG) serves as an essential tool in modeling complex event relations among real-world entities. A temporal knowledge graph can be viewed as a collection of knowledge graph snapshots ordered by time. Reasoning over such graphs remains nontrivial as temporal causal dependenci…

Cited by 0SourceScholar
2022

A Centimeter-Scale Electrohydrodynamic Multi-Modal Robot Capable of Rolling, Hopping, and Taking Off

RA-L 2022

Insects and animals in nature generally have various modes of locomotion to adapt to complex environments, such as crawling, running, flying, and jumping. Achieving multi-locomotion in a centimeter-scale robot requires complex structures and mechanisms that are normally difficult to design and fabri

Cited by 7SourceScholar
2022

Untethered Microrobots Driven by kV-Level Capacitive Actuators via Mechanical Electrostatic Inverters

RA-L 2022

Several microrobots driven by capacitive actuators have shown excellent performance because of the high power density and low working current of the actuators. However, these capacitive actuators need high alternating working voltage (hundreds or thousands of volts), which results in heavy and bulky

Cited by 5SourceScholar
2021

Few-Shot Intent Detection via Contrastive Pre-Training and Fine-Tuning

EMNLP 2021main

In this work, we focus on a more challenging few-shot intent detection scenario where many intents are fine-grained and semantically similar. We present a simple yet effective few-shot intent detection schema via contrastive pre-training and fine-tuning. Specifically, we first conduct self-supervise…

2021

Low Voltage Control of Micro-Ionic Thrusters Using the Electrostatic Induced Potential of the Collector

RA-L 2021

This work presents a low voltage (5 V) control method for micro-ionic thrusters with high operating voltage (>1 kV). When the collector of a micro-ionic thruster is connected to the negative electrode through a 5V switch circuit, and a high DC voltage (1-10 kV) is applied to the emitter, we can get

Cited by 7SourceScholar
2021

PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity Recognition

EMNLP 2021main

Cross-domain Named Entity Recognition (NER) transfers the NER knowledge from high-resource domains to the low-resource target domain. Due to limited labeled resources and domain shift, cross-domain NER is a challenging task. To address these challenges, we propose a progressive domain adaptation Kno…

Cited by 26SourcePDFScholar
2020

A Large-Scale Deep Architecture for Personalized Grocery Basket Recommendations

ICASSP 2020accepted

With growing consumer adoption of online grocery shopping through platforms such as Amazon Fresh, Instacart, and Walmart Grocery, there is a pressing business need to provide relevant recommendations throughout the customer journey. In this paper, we introduce a production within-basket grocery reco…

Cited by 0SourceScholar
2020

Adaptive Variance Based Label Distribution Learning For Facial Age Estimation

ECCV 2020poster

Estimating age from a single facial image is a classic and challenging topic in computer vision. One of its most intractable issues is label ambiguity, i.e., face images from adjacent age of the same person are often indistinguishable. Some existing methods adopt distribution learning to tackle this…

Cited by 78SourcePDFScholar
2020

Identity-Guided Human Semantic Parsing for Person Re-Identification

ECCV 2020poster

Existing alignment-based methods have to employ the pre-trained human parsing models to achieve the pixel-level alignment, and cannot identify the personal belongings (e.g., backpacks and reticule) which are crucial to person re-ID. In this paper, we propose the identity-guided human semantic parsin…

2019

Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark Detection

CVPR 2019poster

Recently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do…

Cited by 74PDFScholar