← Search

Hanhui Li

13 accepted papers

2026

Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration

ICML 2026poster

Reinforcement Learning with Verifiable Reward (RLVR) is a powerful method for enhancing the reasoning abilities of Large Language Models, but its full potential is limited by a lack of exploration in two key areas: \textbf{Depth} (the difficulty of problems) and \textbf{Breadth} (the number of train…

Cited by 0SourceScholar
2026

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

CVPR 2026

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex dynamics. This limitation arises primarily because existing approa

Cited by 0SourceScholar
2025

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

CVPR 2025poster

Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of visual language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-gra…

Cited by 2SourcePDFScholar
2025

GDrag:Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion

ICLR 2025poster

Recent interactive point-based image manipulation methods have gained considerable attention for being user-friendly. However, these methods still face two types of ambiguity issues that can lead to unsatisfactory outcomes, namely, intention ambiguity which misinterprets the purposes of users, and c…

Cited by 1SourcePDFScholar
2025

SeePhys: Does Seeing Help Thinking? – Benchmarking Vision-Based Physics Reasoning

NeurIPS 2025poster

We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In cont…

Cited by 0SourcecodeScholar
2024

3D Visibility-Aware Generalizable Neural Radiance Fields for Interacting Hands

AAAI 2024technical

Neural radiance fields (NeRFs) are promising 3D representations for scenes, objects, and humans. However, most existing methods require multi-view inputs and per-scene training, which limits their real-life applications. Moreover, current methods focus on single-subject cases, leaving scenes of inte…

2024

GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections

ECCV 2024poster

"General text-to-image models bring revolutionary innovation to the fields of arts, design, and media. However, when applied to garment generation, even the state-of-the-art text-to-image models suffer from fine-grained semantic misalignment, particularly concerning the quantity, position, and inter…

Cited by 5SourcePDFScholar
2024

Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand Avatars

NeurIPS 2024poster

In this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single subjects often yield unsatisfactory results due to limited input views, various hand poses, and occlusions. To address t…

2024

Monocular 3D Hand Mesh Recovery via Dual Noise Estimation

AAAI 2024technical

Current parametric models have made notable progress in 3D hand pose and shape estimation. However, due to the fixed hand topology and complex hand poses, current models are hard to generate meshes that are aligned with the image well. To tackle this issue, we introduce a dual noise estimation metho…

2022

BodyGAN: General-Purpose Controllable Neural Human Body Generation

CVPR 2022poster

Recent advances in generative adversarial networks (GANs) have provided potential solutions for photorealistic human image synthesis. However, the explicit and individual control of synthesis over multiple factors, such as poses, body shapes, and skin colors, remains difficult for existing methods.…

Cited by 11PDFScholar
2022

Multiple Temporal Context Embedding Networks for Unsupervised time Series Anomaly Detection

ICASSP 2022accepted

Unsupervised anomaly detection for time series signals is challenging, due to the imbalanced distribution of data and the lack of ground-truth labels. Current methods on this topic are mainly based on deep neural networks, which are optimized by heuristic constraints or empirical priors. However, va…

Cited by 0SourceScholar
2022

Towards Hard-pose Virtual Try-on via 3D-aware Global Correspondence Learning

NeurIPS 2022accept

In this paper, we target image-based person-to-person virtual try-on in the presence of diverse poses and large viewpoint variations. Existing methods are restricted in this setting as they estimate garment warping flows mainly based on 2D poses and appearance, which omits the geometric prior of the…

2021

Towards Complex and Continuous Manipulation: A Gesture Based Anthropomorphic Robotic Hand Design

RA-L 2021

Most current anthropomorphicrobotic hands can realize part of the human hand functions, particularly for object grasping. However, due to the complexity of the human hand, few current designs target at daily object manipulations, even for simple actions like rotating a pen. To tackle this problem, w

Cited by 14SourceScholar