← Search

Zuming Huang

3 accepted papers

2025

VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

NeurIPS 2025spotlight

Recently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through explicit reflection. They significantly outperform the best fast-thinking models, such as GPT-4o, on various math and science benchmarks. However, their multimodal re…

Cited by 0SourceScholar
2019

Look More Than Once: An Accurate Detector for Text of Arbitrary Shapes

CVPR 2019poster

Previous scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challengi…

Cited by 328PDFScholar
2015

Extraction of Virtual Baselines From Distorted Document Images Using Curvilinear Projection

ICCV 2015poster

The baselines of a document page are a set of virtual horizontal and parallel lines, to which the printed contents of document, e.g., text lines, tables or inserted photos, are aligned. Accurate baseline extraction is of great importance in the geometric correction of curved document images. In this…

Cited by 17PDFScholar