← Search

Jun Tang

10 accepted papers

2025

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

ICCV 2025poster

Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what extent capabilities in literacy with rich structure and fine-grained visual challenges. The current landscape lacks a compreh…

Cited by 0SourcePDFScholar
2024

Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading

NeurIPS 2024poster

Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy enviro…

2024

Platypus: A Generalized Specialist Model for Reading Text in Various Forms

ECCV 2024poster

"Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. Previously, individual specialist models are developed to tackle the sub-tasks of text reading (e.g., scene text recogni…

2022

PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer towards Video Object Detection

ECCV 2022poster

"Recent years have witnessed a trend of applying context frames to boost the performance of object detection as video object detection. Existing methods usually aggregate features at one stroke to enhance the feature. These methods, however, usually lack spatial information from neighboring frames a…

2022

Vision-Language Pre-Training for Boosting Scene Text Detectors

CVPR 2022poster

Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning for scene text detection, a task that intrinsically involves cross-modal interaction between the two modalities: vision…

Cited by 37PDFcodeScholar
2021

MOST: A Multi-Oriented Scene Text Detector With Localization Refinement

CVPR 2021poster

Over the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they might still fall short when handling text instances of extreme aspect ratios and varying scales. To tackle such difficult…

Cited by 117PDFScholar
2016

A simple 2D straight-leg passive dynamic walking model without foot-scuffing problem

IROS 2016poster

This paper presents a simple 2D passive dynamic walking model with straight legs based on a novel hip joint, called T-joint. The model directly solves the common foot-scuffing problem in straight-legged walkers without introducing any new degree of freedom or additional motion phase, which is unavoi…

Cited by 7SourceScholar