← Search

Zhou FANG

11 accepted papers

2026

MoE-Powered Fast VLMs Via Curriculum Learning-Based Knowledge Distillation: Taming Regular and Corner Cases in Autonomous Driving

ICRA 2026poster

Autonomous driving has advanced significantly with the integration of large Vision-Language Models (VLMs), which excel in understanding and analyzing driving data. However, existing VLMs face challenges, particularly in terms of latency, which is crucial for real-time driving tasks. While shrinking …

Cited by 0Scholar
2026

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have garnered significant attention recently and demonstrate outstanding capabilities in various tasks such as OCR, VQA, captioning, etc. However, hallucination remains a persistent issue. While numerous methods have been proposed to mitigate hallucinations,

Cited by 0SourcePDFScholar
2025

Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning

NeurIPS 2025poster

Reinforcement learning (RL) heavily depends on well-designed reward functions, which are often biased and difficult to design for complex behaviors. Preference-based RL (PbRL) addresses this by learning reward models from human feedback, but its practicality is constrained by a critical dilemma: wh…

Cited by 0SourceScholar
2024

General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion Process

NeurIPS 2024poster

Articulated object manipulation in real images is a fundamental step in computer and robotic vision tasks. Recently, several image editing methods based on diffusion models have been proposed to manipulate articulated objects according to text prompts. However, these methods often generate weird art…

Cited by 0SourcePDFScholar
2024

PACE: Pose Annotations in Cluttered Environments

ECCV 2024poster

"We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered scenarios. PACE provides a large-scale real-world benchmark for both instance-level and category-level settings. The ben…

2024

Primitive-Based 3D Human-Object Interaction Modelling and Programming

AAAI 2024technical

Embedding Human and Articulated Object Interaction (HAOI) in 3D is an important direction for a deeper human activity understanding. Different from previous works that use parametric and CAD models to represent humans and objects, in this work, we propose a novel 3D geometric primitive-based languag…

Cited by 3SourcePDFScholar
2024

vMFER: Von Mises-Fisher Experience Resampling Based on Uncertainty of Gradient Directions for Policy Improvement

IJCAI 2024poster

Reinforcement Learning (RL) is a widely employed technique in decision-making problems, encompassing two fundamental operations -- policy evaluation and policy improvement. Enhancing learning efficiency remains a key challenge in RL, with many efforts focused on using ensemble critics to boost polic…

Cited by 1SourcePDFScholar
2022

Type-enriched Hierarchical Contrastive Strategy for Fine-Grained Entity Typing

COLING 2022main

Fine-grained entity typing (FET) aims to deduce specific semantic types of the entity mentions in the text. Modern methods for FET mainly focus on learning what a certain type looks like. And few works directly model the type differences, that is, let models know the extent that which one type is di…

Cited by 11SourcePDFScholar
2020

GP-SLAM+: real-time 3D lidar SLAM based on improved regionalized Gaussian process map reconstruction

IROS 2020poster

This paper presents a 3D lidar SLAM system based on improved regionalized Gaussian process (GP) map reconstruction to provide both low-drift state estimation and mapping in real-time for robotics applications. We utilize spatial GP regression to model the environment. This tool enables us to recover…

Cited by 16SourceScholar
2016

Learning models for constraint-based motion parameterization from interactive physics-based simulation

IROS 2016poster

For robotic agents to perform manipulation tasks in human environments at a human level or higher, they need to be able to relate the physical effects of their actions to how they are executing them; small variations in execution can have very different consequences. This paper proposes a framework…

Cited by 28SourceScholar
2016

Open robotics research using web-based knowledge services

ICRA 2016

In this paper we discuss how the combination of modern technologies in “big data” storage and management, knowledge representation and processing, cloud-based computation, and web technology can help the robotics community to establish and strengthen an open research discipline. We describe how we m

Cited by 22SourceScholar