← Search

Zhiyu Chen

32 accepted papers

2026

HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation

ICLR 2026poster

Agentic Retrieval-Augmented Generation (RAG) is a powerful technique for incorporating external information that Large Language Models (LLMs) lack, enabling better problem solving and question answering. However, suboptimal search behaviors exist widely, such as over-search (retrieving information a…

Cited by 0SourceScholar
2026

LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform

ICML 2026poster

Literature reviews are essential to reflect the landscape of research fields. Large language models, especially deep research agents, have recently shown strong capabilities in automated literature review generation. However, it remains a challenging task to rigorously evaluate the scientific value …

Cited by 0SourceScholar
2026

OmniClimb: A Passively Adaptive Rail-Climbing Robot for Multi-Curvature Mobility

RA-L 2026

Stable mobility of rail-climbing robots in complex aerial worksites remains a significant challenge for industrial inspection and maintenance. Existing designs often struggle with reliable adhesion, efficient locomotion, and adaptability to omnidirectional movement on rails with varying curvature an

Cited by 0SourceScholar
2026

Principled RL for Flow Matching Emerges From the Chunk-level Policy Optimization

ICML 2026poster

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive …

Cited by 0SourceScholar
2026

RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue

AAAI 2026technical

Food rescue organizations simultaneously tackle food insecurity and waste by working with volunteers to redistribute food from donors who have excess to recipients who need it. Volunteer feedback allows food rescue organizations to identify issues early and ensure volunteer satisfaction. However, fo

Cited by 0SourcePDFScholar
2025

CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy

NAACL 2025long

There is a significant gap between patient needs and available mental health support today. In this paper, we aim to thoroughly examine the potential of using Large Language Models (LLMs) to assist professional psychotherapy. To this end, we propose a new benchmark, CBT-Bench, for the systematic eva…

Cited by 4SourcePDFScholar
2025

IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction

ACL 2025finding

While large language models (LLMs) have been thoroughly evaluated for deductive and inductive reasoning, their proficiency in holistic rule learning in interactive environments remains less explored. We introduce RULEARN, a novel benchmark to assess the rule-learning abilities of LLM agents in inter…

2025

LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Research

EMNLP 2025

Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reproducing code from research papers, especially in the NLP domain, remains underexplored. This task includes unique complex r

2025

ObfusLM: Privacy-preserving Language Model Service against Embedding Inversion Attacks

ACL 2025long

As the rapid expansion of Machine Learning as a Service (MLaaS) for language models, concerns over the privacy of client inputs during inference or fine-tuning have correspondingly escalated. Recently, solutions have been proposed to safeguard client privacy by obfuscation techniques. However, the s…

2025

Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty

EMNLP 2025

Agentic Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by enabling dynamic, multi-step reasoning and information retrieval. However, these systems often exhibit sub-optimal search behaviors like over-search (retrieving redundant information) and under-search (faili

Cited by 0SourcePDFScholar
2025

Wizard of Shopping: Target-Oriented E-commerce Dialogue Generation with Decision Tree Branching

ACL 2025long

The goal of conversational product search (CPS) is to develop an intelligent, chat-based shopping assistant that can directly interact with customers to understand shopping intents, ask clarification questions, and find relevant products. However, training such assistants is hindered mainly due to t…

Cited by 0SourcePDFScholar
2024

Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data

COLING 2024main

Large Language Models (LLMs) operating in 0-shot or few-shot settings achieve competitive results in Text Classification tasks. In-Context Learning (ICL) typically achieves better accuracy than the 0-shot setting, but it pays in terms of efficiency, due to the longer input prompt. In this paper, we…

Cited by 2SourcePDFScholar
2024

Generative Explore-Exploit: Training-free Optimization of Generative Recommender Systems using LLM Optimizers

ACL 2024long

Recommender systems are widely used to suggest engaging content, and Large Language Models (LLMs) have given rise to generative recommenders. Such systems can directly generate items, including for open-set tasks like question suggestion. While the world knowledge of LLMs enables good recommendation…

Cited by 4SourcePDFScholar
2024

Identifying High Consideration E-Commerce Search Queries

EMNLP 2024industry

In e-commerce, high consideration search missions typically require careful and elaborate decision making, and involve a substantial research investment from customers. We consider the task of identifying High Consideration (HC) queries. Identifying such queries enables e-commerce sites to better se…

Cited by 0SourcePDFScholar
2024

Large Language Models as Zero-shot Dialogue State Tracker through Function Calling

ACL 2024long

Large language models (LLMs) are increasingly prevalent in conversational systems due to their advanced understanding and generative capabilities in general contexts. However, their effectiveness in task-oriented dialogues (TOD), which requires not only response generation but also effective dialogu…

2024

Multimodal Procedural Planning via Dual Text-Image Prompting

EMNLP 2024finding

Embodied agents have achieved prominent performance in following human instructions to complete tasks. However, the potential of providing instructions informed by texts and images to assist humans in completing tasks remains underexplored. To uncover this capability, we present the multimodal proce…

2024

PATIENT-𝜓: Using Large Language Models to Simulate Patients for Training Mental Health Professionals

EMNLP 2024main

Mental illness remains one of the most critical public health issues. Despite its importance, many mental health professionals highlight a disconnect between their training and actual real-world patient practice. To help bridge this gap, we propose PATIENT-𝜓, a novel patient simulation framework for…

2023

Answering Unanswered Questions through Semantic Reformulations in Spoken QA

ACL 2023industry

Spoken Question Answering (QA) is a key feature of voice assistants, usually backed by multiple QA systems. Users ask questions via spontaneous speech that can contain disfluencies, errors, and informal syntax or phrasing. This is a major challenge in QA, causing unanswered questions or irrelevant a…

2023

Empowering Psychotherapy with Large Language Models: Cognitive Distortion Detection through Diagnosis of Thought Prompting

EMNLP 2023short findings

Mental illness remains one of the most critical public health issues of our time, due to the severe scarcity and accessibility limit of professionals. Psychotherapy requires high-level expertise to conduct deep, complex reasoning and analysis on the cognition modeling of the patients. In the era of…

Cited by 0SourceScholar
2023

Generate-then-Retrieve: Intent-Aware FAQ Retrieval in Product Search

ACL 2023industry

Frequently Asked Question (FAQ) retrieval aims at retrieving question-answer pairs for a given a user query. Integrating FAQ retrieval with product search can not only empower users to make more informed purchase decisions, but also enhance user retention through efficient post-purchase support. Pro…

Cited by 5SourcePDFScholar
2023

Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling

ICML 2023poster

Health conditions among patients in intensive care units (ICUs) are monitored via electronic health records (EHRs), composed of numerical time series and lengthy clinical note sequences, both taken at $\textit{irregular}$ time intervals. Dealing with such irregularity in every modality, and integrat…

2023

MultiCoNER v2: a Large Multilingual dataset for Fine-grained and Noisy Named Entity Recognition

EMNLP 2023short findings

We present MULTICONER V2, a dataset for fine-grained Named Entity Recognition covering 33 entity classes across 12 languages, in both monolingual and multilingual settings. This dataset aims to tackle the following practical challenges in NER: (i) effective handling of fine-grained classes that incl…

Cited by 0SourceScholar
2022

ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering

EMNLP 2022main

With the recent advance in large pre-trained language models, researchers have achieved record performances in NLP tasks that mostly focus on language pattern matching. The community is experiencing the shift of the challenge from how to model language to the imitation of complex reasoning abilities…

2022

F8Net: Fixed-Point 8-bit Only Multiplication for Network Quantization

ICLR 2022oral

Neural network quantization is a promising compression technique to reduce memory footprint and save energy consumption, potentially leading to real-time inference. However, there is a performance gap between quantized and full-precision models. To reduce it, existing quantization approaches require…

2022

KETOD: Knowledge-Enriched Task-Oriented Dialogue

NAACL 2022findings

Existing studies in dialogue system research mostly treat task-oriented dialogue and chit-chat as separate domains. Towards building a human-like assistant that can converse naturally and seamlessly with users, it is important to build a dialogue system that conducts both types of conversations effe…

2022

Reinforced Question Rewriting for Conversational Question Answering

EMNLP 2022industry

Conversational Question Answering (CQA) aims to answer questions contained within dialogues, which are not easily interpretable without context. Developing a model to rewrite conversational questions into self-contained ones is an emerging solution in industry settings as it allows using existing si…

Cited by 27SourcePDFScholar
2021

FinQA: A Dataset of Numerical Reasoning over Financial Data

EMNLP 2021main

The sheer volume of financial statements makes it difficult for humans to access and analyze a business’s financials. Robust numerical reasoning likewise faces unique challenges in this domain. In this work, we focus on answering deep questions over financial data, aiming to automate the analysis of…

2021

NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile Acceleration

CVPR 2021poster

With the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Prior methods towards this goal, including model compression and network architecture search (NAS), are largely performed i…

Cited by 34PDFcodeScholar
2021

NUANCED: Natural Utterance Annotation for Nuanced Conversation with Estimated Distributions

EMNLP 2021finding

Existing conversational systems are mostly agent-centric, which assumes the user utterances will closely follow the system ontology. However, in real-world scenarios, it is highly desirable that users can speak freely and naturally. In this work, we attempt to build a user-centric dialogue system fo…

2020

Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text Classification

ICML 2020poster

Extreme multi-label text classification (XMTC) is a task for tagging a given text with the most relevant labels from an extremely large label set. We propose a novel deep learning method called APLC-XLNet. Our approach fine-tunes the recently released generalized autoregressive pretrained model (XLN…