← Search

Weijie Zhou

7 accepted papers

2026

ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response

ICML 2026poster

While passive agents merely follow instructions, proactive agents align with higher-level objectives, such as assistance and safety by continuously monitoring the environment to determine when and how to act. However, developing proactive agents is hindered by the lack of specialized resources. To a…

Cited by 0SourcecodeScholar
2025

LightPlanner: Unleashing the Reasoning Capabilities of Lightweight Large Language Models in Task Planning

IROS 2025

In recent years, lightweight large language models (LLMs) have garnered significant attention in the robotics field due to their low computational resource requirements and suitability for edge deployment. However, in task planning—particularly for complex tasks that involve dynamic semantic logic r

Cited by 4SourcecodeScholar
2025

PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments

NeurIPS 2025poster

Visual reasoning in multimodal large language models (MLLMs) has primarily been studied in passive, static settings, limiting their effectiveness in real-world physical environments where an embodied agent must contend with incomplete information due to occlusion or a limited field of view. Humans,…

Cited by 0SourceScholar
2025

PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability

CVPR 2025poster

Understanding the environment and a robot's physical reachability is crucial for task execution. While state-of-the-art vision-language models (VLMs) excel in environmental perception, they often generate inaccurate or impractical responses in embodied visual reasoning tasks due to a lack of underst…

2025

Projection, Interaction and Fusion: A Progressive Difference Fusion Network for Salient Object Detection

IJCAI 2025

In recent years, deep learning-based Salient Object Detection (SOD) methods have made tremendous progress; however, their performance in complex scenarios has reached a bottleneck. In this paper, we propose a novel Progressive Difference Fusion Network (PDFNet) based on fine-grained feature fusion.

2025

UMSSS: A Visual Scene Semantic Segmentation Dataset for Underground Mines

ICASSP 2025accepted

Specialized datasets designed for mining scenarios are the essential foundation for the development, operation, and research of intelligent mines. Currently, the available datasets focus primarily on open-pit mines, with a lack of specialized datasets for underground mines. This gap severely hinders…

Cited by 0SourceScholar
2025

You Should Learn to Stop Denoising on Point Clouds in Advance

AAAI 2025technical

Point clouds have become the preferred data format for a variety of tasks in 3D vision and graphics. However, raw point clouds often contain significant noise. This paper introduces the Adaptive Stop Denoising Network (ASDN), a novel approach aimed at restoring high-quality point clouds from noisy d…