← Search

Haochen Zhang

15 accepted papers

2025

A Variable Admittance Control Strategy for Stable and Compliant Human-Robot Physical Interaction

RA-L 2025

Admittance control is an important method for providing collaborative robots with precise manipulation and flexible contact behavior in industrial settings that often involve physical interaction. However, too rigid or high-frequency interactions by non-specialists will jeopardise the stability of t

Cited by 5SourceScholar
2025

Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining

NeurIPS 2025poster

Low-rank optimization has emerged as a promising approach to enabling memory-efficient training of large language models (LLMs). Existing low-rank optimization methods typically project gradients onto a low-rank subspace, reducing the memory cost of storing optimizer states. A key challenge in these…

Cited by 0SourceScholar
2025

CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems

ACL 2025finding

Recommender systems play a pivotal role in providing relevant content to users. With the rapid development of large language models (LLMs), researchers have begun utilizing LLMs to build more powerful recommender systems. However, existing approaches that focus on aligning LLMs with recommendation t…

2025

Federated $Q$-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost

ICLR 2025poster

In this paper, we consider model-free federated reinforcement learning for tabular episodic Markov decision processes. Under the coordination of a central server, multiple agents collaboratively explore the environment and learn an optimal policy without sharing their raw data. Despite recent advanc…

Cited by 6SourcePDFScholar
2025

Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition

ICLR 2025spotlight

We study the gap-dependent bounds of two important algorithms for on-policy $Q$-learning for finite-horizon episodic tabular Markov Decision Processes (MDPs): UCB-Advantage (Zhang et al. 2020) and Q-EarlySettled-Advantage (Li et al. 2021). UCB-Advantage and Q-EarlySettled-Advantage improve upon the…

Cited by 3SourcePDFScholar
2025

IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes

ICRA 2025

With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can operate in diverse environments given natural language input. One such application is indoor navigation using natural la

Cited by 3SourcecodeScholar
2025

Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning

NeurIPS 2025poster

Motivated by real-world settings where data collection and policy deployment—whether for a single agent or across multiple agents—are costly, we study the problem of on-policy single-agent reinforcement learning (RL) and federated RL (FRL) with a focus on minimizing burn-in costs (the sa…

Cited by 0SourceScholar
2025

SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models

IROS 2025

Interpreting object-referential language and grounding objects in 3D with spatial relations and attributes is essential for robots operating alongside humans. However, this task is often challenging due to the diversity of scenes, large number of fine-grained objects, and complex free-form nature of

Cited by 8SourcecodeScholar
2025

Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory

AISTATS 2025poster

Lifelong reinforcement learning (RL) has been developed as a paradigm for extending single-task RL to more realistic, dynamic settings. In lifelong RL, the "life" of an RL agent is modeled as a stream of tasks drawn from a task distribution. We propose EPIC (Empirical PAC-Bayes that Improves Continu…

Cited by 0SourceScholar
2024

Jellyfish: Instruction-Tuning Local Large Language Models for Data Preprocessing

EMNLP 2024main

This paper explores the utilization of LLMs for data preprocessing (DP), a crucial step in the data mining pipeline that transforms raw data into a clean format. We instruction-tune local LLMs as universal DP task solvers that operate on a local, single, and low-priced GPU, ensuring data security an…

Cited by 8SourcePDFScholar
2021

Unsupervised Real-World Super-Resolution: A Domain Adaptation Perspective

ICCV 2021poster

Most existing convolution neural network (CNN) based super-resolution (SR) methods generate their paired training dataset by artificially synthesizing low-resolution (LR) images from the high-resolution (HR) ones. However, this dataset preparation strategy harms the application of these CNNs in real…

Cited by 61PDFScholar