← Search

Xiaoyan Guo

4 accepted papers

2026

Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO) have significantly advanced the reasoning capabilities of large language models. Extending these methods to multimodal settings, however, faces a critical challenge: the instability of std-based norma…

Cited by 0SourceScholar
2022

MobRecon: Mobile-Friendly Hand Mesh Reconstruction From Monocular Image

CVPR 2022poster

In this work, we propose a framework for single-view hand mesh reconstruction, which can simultaneously achieve high reconstruction accuracy, fast inference speed, and temporal coherence. Specifically, for 2D encoding, we propose lightweight yet effective stacked structures. Regarding 3D decoding, w…

Cited by 107PDFcodeScholar
2021

Camera-Space Hand Mesh Recovery via Semantic Aggregation and Adaptive 2D-1D Registration

CVPR 2021poster

Recent years have witnessed significant progress in 3D hand mesh recovery. Nevertheless, because of the intrinsic 2D-to-3D ambiguity, recovering camera-space 3D information from a single RGB image remains challenging. To tackle this problem, we divide camera-space mesh recovery into two sub-tasks, i…

Cited by 112PDFcodeScholar
2020

Improving Monocular Depth Estimation by Leveraging Structural Awareness and Complementary Datasets

ECCV 2020poster

Monocular depth estimation plays a crucial role in 3D recognition and understanding. One key limitation of existing approaches lies in their lack of structural information exploitation, which leads to inaccurate spatial layout, discontinuous surface, and ambiguous boundaries. In this paper, we tackl…

Cited by 35SourcePDFScholar