← Search

Hao Wen

15 accepted papers

2025

An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint

EMNLP 2025

Recent work has demonstrated the remarkable potential of Large Language Models (LLMs) in test-time scaling. By making models think before answering, they are able to achieve much higher accuracy with extra inference computation.However, in many real-world scenarios, models are used under time constr

Cited by 0SourcePDFScholar
2025

GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration

CVPR 2025poster

GUI agents hold significant potential to enhance the experience and efficiency of human-device interaction. However, current methods face challenges in generalizing across applications (apps) and tasks, primarily due to two fundamental limitations in existing datasets. First, these datasets overlook…

2025

IDOL: Instant Photorealistic 3D Human Creation from a Single Image

CVPR 2025poster

Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the…

2025

MambaTrack: Exploiting Dual-Enhancement for Night UAV Tracking

ICASSP 2025accepted

Night unmanned aerial vehicle (UAV) tracking is impeded by the challenges of poor illumination, with previous daylight-optimized methods demonstrating suboptimal performance in low-light conditions, limiting the utility of UAV applications. To this end, we propose an efficient mamba-based tracker, l…

Cited by 0SourceScholar
2025

Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion

CVPR 2025poster

Existing image-to-3D creation methods typically split the task into two individual stage: multi-view image generation and 3D reconstruction, leading to two main limitations: (1) In multi-view generation stage, the multi-view generated images present a challenge to preserving 3D consistency;; (2) In…

2024

Development of Variable Transmission Series Elastic Actuator for Hip Exoskeletons

ICRA 2024poster

Series Elastic Actuator-based exoskeleton can offer precise torque control and transparency when interacting with human wearers. Accurate control of SEA-produced torques ensures the wearer’s voluntary motion and supports the implementation of multiple assistive paradigms. In this paper, a novel vari…

Cited by 1SourceScholar
2024

EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion

CVPR 2024poster

Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods that introduce 3D global representation into diffusion models have shown the potential to generate consistent multiviews but they have reduced generation speed a…

2024

WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark

NeurIPS 2024poster

Underwater Object Tracking (UOT) is essential for identifying and tracking submerged objects in underwater videos, but existing datasets are limited in scale, diversity of target categories and scenarios covered, impeding the development of advanced tracking algorithms. To bridge this gap, we take t…

2023

Crowd3D: Towards Hundreds of People Reconstruction From a Single Image

CVPR 2023poster

Image-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and…

Cited by 13SourcePDFScholar
2023

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

ICCV 2023poster

Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually notoriously expensive. Moreover, training via one or only a few traditional task…

Cited by 18PDFcodeScholar
2022

Point Primitive Transformer for Long-Term 4D Point Cloud Video Understanding

ECCV 2022poster

"This paper proposes a 4D backbone for long-term point cloud video understanding. A typical way to capture spatial-temporal context is using 4Dconv or transformer without hierarchy. However, those methods are neither effective nor efficient enough due to camera motion, scene changes, sampling patter…

2021

Design and Experimental Validation of a Robotic System for Reactor Core Detector Removal

ICRA 2021poster

The reactor power and the coolant level in the nuclear plant are monitored via the reactor core detectors. Every 4 to 5 years, the detectors with high-level radiation need to be removed, which is time-consuming and hazardous for workers. To address this issue, this paper introduces a novel robotic s…

Cited by 3SourceScholar
2021

End-to-End Semi-supervised Learning for Differentiable Particle Filters

ICRA 2021poster

Recent advances in incorporating neural networks into particle filters provide the desired flexibility to apply particle filters in large-scale real-world applications. The dynamic and measurement models in this framework are learnable through the differentiable implementation of particle filters. P…

Cited by 23SourcecodeScholar
2021

Virtual-Fixture Based Drilling Control for Robot-Assisted Craniotomy: Learning From Demonstration

RA-L 2021

One of the promising solutions for drilling craniotomy is robot-assisted surgery with human guidance. The present study deals with a piecewise collaborative drilling task assisted by a robot while containing aligning and drilling. It can enable surgeons to complete the operation more efficiently and

Cited by 37SourceScholar