← Search

Fengli Xu

14 accepted papers

2026

AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent

ICML 2026poster

In modern AI research, baseline and dataset selection is a high-stakes decision in experimental design. It operationalizes a research idea into a concrete evaluation protocol and largely determines the validity and comparability of empirical conclusions. However, making appropriate choices is increa…

Cited by 0SourceScholar
2026

AgentSwift: Efficient LLM Agent Design via Value-Guided Hierarchical Search

AAAI 2026technical

Large language model (LLM) agents have demonstrated strong capabilities across diverse domains, yet automated agent design remains a significant challenge. Current automated agent design approaches are often constrained by limited search spaces that primarily optimize workflows but fail to integrate

Cited by 0SourcePDFScholar
2026

FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents

ICLR 2026poster

Mobile GUI agents are becoming critical tools to improve user experience on smart devices, with multimodal large language models (MLLMs) emerging as the dominant paradigms in this domain. Current agents, however, rely on explicit human instructions, overlooking the potential to leverage the contextu…

Cited by 0SourcecodeScholar
2026

LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform

ICML 2026poster

Literature reviews are essential to reflect the landscape of research fields. Large language models, especially deep research agents, have recently shown strong capabilities in automated literature review generation. However, it remains a challenging task to rigorously evaluate the scientific value …

Cited by 0SourceScholar
2026

Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs

ICLR 2026poster

Large language models (LLMs) acquire extensive prior knowledge through large-scale pretraining and can be further enhanced via supervised fine-tuning (SFT) or reinforcement learning (RL)-based post-training. A growing body of evidence has shown that RL fine-tuning improves the capability of LLMs bey…

Cited by 0SourcecodeScholar
2026

ResMAS: Resilience Optimization in LLM-based Multi-agent Systems

AAAI 2026technical

Large Language Model-based Multi-Agent Systems (LLM-based MAS), where multiple LLM agents collaborate to solve complex tasks, have shown impressive performance in many areas. However, MAS are typically distributed across different devices or environments, making them vulnerable to perturbations such

Cited by 0SourcePDFScholar
2025

AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems

NeurIPS 2025spotlight

The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs’ advanced reasoning and role-playing capabilities to enable autonomous, adaptive decision-making. Unlike traditional recommendation approa…

Cited by 0SourcecodeScholar
2025

AgentSquare: Automatic LLM Agent Search in Modular Design Space

ICLR 2025poster

Recent advancements in Large Language Models (LLMs) have led to a rapid growth of agentic systems capable of handling a wide range of complex tasks. However, current research largely relies on manual, task-specific design, limiting their adaptability to novel tasks. In this paper, we introduce a new…

2025

Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models

ICML 2025poster

Chain-of-Thought (CoT) technique has proven effective in improving the performance of large language models (LLMs) on complex reasoning tasks. However, the performance gains are inconsistent across different tasks, and the underlying mechanism remains a long-standing research question. In this work,…

2024

HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction

NeurIPS 2024poster

Citation networks are critical infrastructures of modern science, serving as intricate webs of past literature and enabling researchers to navigate the knowledge production system. To mine information hiding in the link space of such networks, predicting which previous papers (candidates) will a new…

2021

AttnMove: History Enhanced Trajectory Recovery via Attentional Network

AAAI 2021technical

A considerable amount of mobility data has been accumulated due to the proliferation of location-based service. Nevertheless, compared with mobility data from transportation systems like the GPS module in taxis, this kind of data is commonly sparse in terms of individual trajectories in the sense th…