← Search

Zheng Tian

16 accepted papers

2026

Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) combine the linguistic strengths of LLMs with the ability to process multimodal data, enabling them to address a broader range of tasks. This progression highlights a shift from language-only reasoning to integrated vision–language reasoning in children's dev…

Cited by 0SourcecodeScholar
2026

Robust High-Precision Trajectory Planning for Payload Transportation in Overhead Cranes: A Disturbance-Aware Approach

RA-L 2026

The coordinated motion of the trolley and hoisting rope improves crane flexibility but poses challenges in precise trajectory conversion and tracking due to disturbances and inaccessible low-level controllers. This letter proposes a disturbance-aware high-precision trajectory planning method integra

Cited by 1SourceScholar
2026

Safety-Critical Steering Control for Rubber-Tired Container Gantry Cranes: A State-Interlocked CBF Approach

RA-L 2026

The rubber-tired container gantry crane (RTG) is a type of heavy-duty lifting equipment commonly used in container yards, which is driven by two-side rubber tires and steered via differential drive. While moving along the desired path, the RTG must remain centered of the lane with restricted heading

Cited by 0SourceScholar
2026

SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs

ICLR 2026poster

Humans can imagine and manipulate visual images mentally, a capability known as \textit{spatial visualization}. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through spatial visualization remains insufficiently evaluated…

Cited by 0SourcecodeScholar
2026

UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots

RA-L 2026

Achieving generalizable whole-body motion control is essential for deploying humanoid robots in real-world environments. However, existing MLP-based policies trained under partial observations often suffer from limited expressiveness and struggle to maintain global consistency. These shortcomings ma

Cited by 41SourceScholar
2025

Multi-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal Properties

CVPR 2025poster

Object anomaly detection is essential for industrial quality inspection, yet traditional single-sensor methods face critical limitations. They fail to capture the wide range of anomaly types, as single sensors are often constrained to either external appearance, geometric structure, or internal prop…

2024

An LLM-driven Framework for Multiple-Vehicle Dispatching and Navigation in Smart City Landscapes

ICRA 2024poster

In the context of smart cities, autonomous vehicles, such as unmanned delivery vehicles and taxis are gradually gaining acceptance. However, their application scenarios remain significantly fragmented. Typically, an Autonomous Multi-Functional Vehicle (AMFV) is not engaged in other scenarios when id…

Cited by 15SourceScholar
2024

Language and Sketching: An LLM-driven Interactive Multimodal Multitask Robot Navigation Framework

ICRA 2024poster

The socially-aware navigation system has evolved to adeptly avoid various obstacles while performing multiple tasks, such as point-to-point navigation, human-following, and -guiding. However, a prominent gap persists: in Human-Robot Interaction (HRI), the procedure of communicating commands to robot…

Cited by 20SourceScholar
2024

Off-Agent Trust Region Policy Optimization

IJCAI 2024poster

Leveraging the experiences of other agents offers a powerful mechanism to enhance policy optimization in multi-agent reinforcement learning (MARL). However, contemporary MARL algorithms often neglect experience sharing possibilities or adopt a simple approach via direct parameter sharing. Our work e…

Cited by 0SourcePDFScholar
2024

Tri-Modal Motion Retrieval by Learning a Joint Embedding Space

CVPR 2024highlight

Text-to-motion tasks have been the focus of recent advancements in the human motion domain. However the performance of text-to-motion tasks have not reached its potential primarily due to the lack of motion datasets and the pronounced gap between the text and motion modalities. To mitigate this chal…

Cited by 5SourcePDFScholar
2023

Multi-embodiment Legged Robot Control as a Sequence Modeling Problem

ICRA 2023poster

Robots are traditionally bounded by a fixed embodiment during their operational lifetime, which limits their ability to adapt to their surroundings. Co-optimizing control and morphology of a robot, however, is often inefficient due to the complex interplay between the controller and morphology. In t…

Cited by 15SourceScholar
2023

Order Matters: Agent-by-agent Policy Optimization

ICLR 2023poster

While multi-agent trust region algorithms have achieved great success empirically in solving coordination tasks, most of them, however, suffer from a non-stationarity problem since agents update their policies simultaneously. In contrast, a sequential scheme that updates policies agent-by-agent pro…

2023

Sim-to-Real Transfer for Quadrupedal Locomotion via Terrain Transformer

ICRA 2023poster

Deep reinforcement learning has recently emerged as an appealing alternative for legged locomotion over multiple terrains by training a policy in physical simulation and then transferring it to the real world (i.e., sim-to-real transfer). Despite considerable progress, the capacity and scalability o…

Cited by 22SourceScholar
2022

M2N: Mesh Movement Networks for PDE Solvers

NeurIPS 2022accept

Numerical Partial Differential Equation (PDE) solvers often require discretizing the physical domain by using a mesh. Mesh movement methods provide the capability to improve the accuracy of the numerical solution without introducing extra computational burden to the PDE solver, by increasing mesh re…

Cited by 19SourcePDFScholar
2020

SMARTS: An Open-Source Scalable Multi-Agent RL Training School for Autonomous Driving

CoRL 2020

Interaction is fundamental in autonomous driving (AD). Despite more than a decade of intensive R&D in AD, how to dynamically interact with diverse road users in various contexts still remains unsolved. Multi-agent learning has recently seen big breakthroughs and has much to offer towards solving rea