← Search

Kang Zheng

7 accepted papers

2025

Conditional Convolutions for End-to-End Single-Stage Video Text Detection

ICASSP 2025accepted

We propose a simple yet effective single-stage video text detection framework, termed CVTD (Conditional convolutions for Video Text Detection), which, to the best of our knowledge, is the first end-to-end single-stage video text detection framework.Most existing video text detection methods adopt te…

Cited by 0SourceScholar
2021

Automatic Vertebra Localization and Identification in CT by Spine Rectification and Anatomically-Constrained Optimization

CVPR 2021poster

Accurate vertebra localization and identification are required in many clinical applications of spine disorder diagnosis and surgery planning. However, significant challenges are posed in this task by highly varying pathologies (such as vertebral compression fracture, scoliosis, and vertebral fixati…

Cited by 35PDFcodeScholar
2020

Anatomy-Aware Siamese Network: Exploiting Semantic Asymmetry for Accurate Pelvic Fracture Detection in X-ray Images

ECCV 2020poster

Trauma PXR are essential for instantaneous pelvic bone fracture detection. However, small, pathologically critical fractures can be missed, even by experienced clinicians, under the very limited diagnosis times allowed in urgent care. As a result, fracture CAD has very high demands to save time and…

Cited by 43SourcePDFScholar
2020

Structured Landmark Detection via Topology-Adapting Deep Graph Learning

ECCV 2020poster

Image landmark detection aims to automatically identify the locations of predefined fiducial points. Despite recent success in this field, higher-ordered structural modeling to capture implicit or explicit relationships among anatomical landmarks has not been adequately exploited. In this work, we p…

Cited by 121SourcePDFScholar
2019

Visual Attention Consistency Under Image Transforms for Multi-Label Image Classification

CVPR 2019poster

Human visual perception shows good consistency for many multi-label image classification tasks under certain spatial transforms, such as scaling, rotation, flipping and translation. This has motivated the data augmentation strategy widely used in CNN classifier training -- transformed images are inc…

Cited by 307PDFScholar
2017

Learning View-Invariant Features for Person Identification in Temporally Synchronized Videos Taken by Wearable Cameras

ICCV 2017poster

In this paper, we study the problem of Cross-View Person Identification (CVPI), which aims at identifying the same person from temporally synchronized videos taken by different wearable cameras. Our basic idea is to utilize the human motion consistency for CVPI, where human motion can be computed by…

Cited by 28PDFScholar
2015

Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation

CVPR 2015poster

We propose a new learning-based method for estimating 2D human pose from a single image, using Dual-Source Deep Convolutional Neural Networks (DS-CNN). Recently, many methods have been developed to estimate human pose by using pose priors that are estimated from physiologically inspired graphical mo…

Cited by 299SourcePDFScholar