← Search

fangwei zhong

27 accepted papers

2026

EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation

ICML 2026poster

Uncovering the causal mechanisms of educational social dynamics is critical for designing effective pedagogical policies. However, traditional methods face a fundamental dilemma. vational studies often lack causal power, while controlled experiments are ethically prohibitive. While LLM-powered multi…

Cited by 0SourceScholar
2026

TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking

ICRA 2026poster

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously following moving targets is essential. Recent advances have enabled language-guided tracking in complex and unstructured…

2025

Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective

ACL 2025finding

As large language models (LLMs) become increasingly integrated into critical applications, aligning their behavior with human values presents significant challenges. Current methods, such as Reinforcement Learning from Human Feedback (RLHF), typically focus on a limited set of coarse-grained values…

2025

Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement Learning

ICML 2025poster

The ability to adapt to new environments with noisy dynamics and unseen objectives is crucial for AI agents. In-context reinforcement learning (ICRL) has emerged as a paradigm to build adaptive policies, employing a **context** trajectory of the test-time interactions to infer the true task and the…

Cited by 0SourcePDFScholar
2025

Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia

NeurIPS 2025poster

Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing eval…

Cited by 0SourceScholar
2025

Simulating Human-like Daily Activities with Desire-driven Autonomy

ICLR 2025poster

Desires motivate humans to interact autonomously with the complex world. In contrast, current AI agents require explicit task specifications, such as instructions or reward functions, which constrain their autonomy and behavioral diversity. In this paper, we introduce a Desire-driven Autonomous Agen…

Cited by 2SourcePDFScholar
2025

UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

ICCV 2025poster

We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environments. We also provide a rich variety of playable entities, including humans, animals, robots, and vehicles for embodied…

2025

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

IROS 2025

We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our approach combines the off-the-shelf active tracking methods with

Cited by 3SourceScholar
2024

Fast Peer Adaptation with Context-aware Exploration

ICML 2024poster

Fast adapting to unknown peers (partners or opponents) with different strategies is a key challenge in multi-agent games. To do so, it is crucial for the agent to probe and identify the peer’s strategy efficiently, as this is the prerequisite for carrying out the best response in adaptation. However…

Cited by 2SourcePDFScholar
2024

Fine Tuning Out-of-Vocabulary Item Recommendation with User Sequence Imagination

NeurIPS 2024spotlight

Recommending out-of-vocabulary (OOV) items is a challenging problem since the in-vocabulary (IV) items have well-trained behavioral embeddings but the OOV items only have content features. Current OOV recommendation models often generate 'makeshift' embeddings for OOV items from content features and…

Cited by 3SourcePDFScholar
2024

Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy

NeurIPS 2024poster

Diplomacy is one of the most sophisticated activities in human society, involving complex interactions among multiple parties that require skills in social reasoning, negotiation, and long-term strategic planning. Previous AI agents have demonstrated their ability to handle multi-step games and larg…

2023

GFPose: Learning 3D Human Pose Prior With Gradient Fields

CVPR 2023poster

Learning 3D human pose prior is essential to human-centered AI. Here, we present GFPose, a versatile framework to model plausible 3D human poses for various applications. At the core of GFPose is a time-dependent score network, which estimates the gradient on each body joint and progressively denois…

2023

Learning Semantic-Agnostic and Spatial-Aware Representation for Generalizable Visual-Audio Navigation

RA-L 2023

Visual-audio navigation (VAN) is attracting more and more attention from the robotic community due to its broad applications, e.g., household robots and rescue robots. In this task, an embodied agent must search for and navigate to the sound source with egocentric visual and audio observations. Howe

Cited by 12SourcecodeScholar
2023

Proactive Multi-Camera Collaboration for 3D Human Pose Estimation

ICLR 2023poster

This paper presents a multi-agent reinforcement learning (MARL) scheme for proactive Multi-Camera Collaboration in 3D Human Pose Estimation in dynamic human crowds. Traditional fixed-viewpoint multi-camera solutions for human motion capture (MoCap) are limited in capture space and susceptible to dyn…

Cited by 17SourcePDFScholar
2023

RSPT: Reconstruct Surroundings and Predict Trajectory for Generalizable Active Object Tracking

AAAI 2023technical

Active Object Tracking (AOT) aims to maintain a specific relation between the tracker and object(s) by autonomously controlling the motion system of a tracker given observations. It is widely used in various applications such as mobile robots and autonomous driving. However, Building a generalizable…

2022

Disentangling Disease-related Representation from Obscure for Disease Prediction

ICML 2022spotlight

Disease-related representations play a crucial role in image-based disease prediction such as cancer diagnosis, due to its considerable generalization capacity. However, it is still a challenge to identify lesion characteristics in obscured images, as many lesions are obscured by other tissues. In t…

Cited by 5SourcePDFScholar
2022

MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control

NeurIPS 2022accept

We introduce the Multi-Agent Tracking Environment (MATE), a novel multi-agent environment simulates the target coverage control problems in the real world. MATE hosts an asymmetric cooperative-competitive game consisting of two groups of learning agents--"cameras" and "targets"--with opposing intere…

2022

TarGF: Learning Target Gradient Field to Rearrange Objects without Explicit Goal Specification

NeurIPS 2022accept

Object Rearrangement is to move objects from an initial state to a goal state. Here, we focus on a more practical setting in object rearrangement, i.e., rearranging objects from shuffled layouts to a normative target distribution without explicit goal specification. However, it remains challenging f…

Cited by 36SourcePDFScholar
2022

ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of Mind

ICLR 2022poster

Being able to predict the mental states of others is a key factor to effective social interaction. It is also crucial for distributed multi-agent systems, where agents are required to communicate and cooperate. In this paper, we introduce such an important social-cognitive skill, i.e. Theory of Mind…

2021

Towards Distraction-Robust Active Visual Tracking

ICML 2021spotlight

In active visual tracking, it is notoriously difficult when distracting objects appear, as distractors often mislead the tracker by occluding the target or bringing a confusing appearance. To address this issue, we propose a mixed cooperative-competitive multi-agent game, where a target and multiple…

Cited by 45SourcePDFScholar
2020

Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor Networks

NeurIPS 2020poster

Maximum target coverage by adjusting the orientation of distributed sensors is an important problem in directional sensor networks (DSNs). This problem is challenging as the targets usually move randomly but the coverage range of sensors is limited in angle and distance. Thus, it is required to coor…

2019

AD-VAT: An Asymmetric Dueling mechanism for learning Visual Active Tracking

ICLR 2019poster

Visual Active Tracking (VAT) aims at following a target object by autonomously controlling the motion system of a tracker given visual observations. Previous work has shown that the tracker can be trained in a simulator via reinforcement learning and deployed in real-world scenarios. However, during…

2019

CRAVES: Controlling Robotic Arm With a Vision-Based Economic System

CVPR 2019poster

Training a robotic arm to accomplish real-world tasks has been attracting increasing attention in both academia and industry. This work discusses the role of computer vision algorithms in this field. We focus on low-cost arms on which no sensors are equipped and thus all decisions are made upon visu…

Cited by 71PDFScholar
2018

End-to-end Active Object Tracking via Reinforcement Learning

ICML 2018oral

We study active object tracking, where a tracker takes as input the visual observation (i.e. frame sequence) and produces the camera control signal (e.g., move forward, turn left, etc). Conventional methods tackle the tracking and the camera control separately, which is challenging to tune jointly.…

Cited by 113SourcePDFScholar