← Search

John See

12 accepted papers

2026

Unleashing Semantic and Geometric Priors for 3D Scene Completion

AAAI 2026technical

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving and robotic navigation. However, existing methods rely on a coupled encoder to deliver both semantic and geometric priors, which forces the model to make a trade-off between confli

Cited by 0SourcePDFScholar
2022

Speed Up Object Detection on Gigapixel-Level Images With Patch Arrangement

CVPR 2022poster

With the appearance of super high-resolution (e.g., gigapixel-level) images, performing efficient object detection on such images becomes an important issue. Most existing works for efficient object detection on high-resolution images focus on generating local patches where objects may exist, and th…

Cited by 13PDFScholar
2022

TA2N: Two-Stage Action Alignment Network for Few-Shot Action Recognition

AAAI 2022technical

Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarity between videos. Recently, it has been observed that directly measuring this si…

2021

Enhancing Self-Supervised Video Representation Learning via Multi-Level Feature Optimization

ICCV 2021poster

The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics and neglected lower-level representations and their temporal relationship which are crucial for general video understan…

Cited by 34PDFcodeScholar
2020

CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization

ECCV 2020poster

Most current pipelines for spatiotemporal action localization connect frame-wise or clip-wise detection results to generate action proposals. In this paper, we propose Coarse-to-Fine Action Detector (CFAD), an original end-to-end trainable framework for efficient spatiotemporal action localization.…

Cited by 30SourcePDFScholar
2020

Delving into the Cyclic Mechanism in Semi-supervised Video Object Segmentation

NeurIPS 2020poster

In this paper, we take attempt to incorporate the cyclic mechanism with the vision task of semi-supervised video object segmentation. By resorting to the accurate reference mask of the first frame, we try to mitigate the error propagation problem in most of current video object segmentation pipeline…

2020

PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments

ECCV 2020poster

Object detection using an oriented bounding box (OBB) can better target rotated objects by reducing the overlap with background areas. Existing OBB approaches are mostly built on horizontal bounding box detectors by introducing an additional angle dimension optimized by a distance loss. However, as…

2019

Towards Accurate One-Stage Object Detection With AP-Loss

CVPR 2019poster

One-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background class imbalance issue due to the large number of anchors. This paper alleviates this issue by proposing a novel framework t…

Cited by 173PDFcodeScholar
2016

Eulerian emotion magnification for subtle expression recognition

ICASSP 2016accepted

Subtle emotions are expressed through tiny and brief movements of facial muscles, called micro-expressions; thus, recognition of these hidden expressions is as challenging as inspection of microscopic worlds without microscopes. In this paper, we show that through motion magnification, subtle expres…

Cited by 0SourceScholar
2016

Intrinsic two-dimensional local structures for micro-expression recognition

ICASSP 2016accepted

An elapsed facial emotion involves changes of facial contour due to the motions (such as contraction or stretch) of facial muscles located at the eyes, nose, lips and etc. Thus, the important information such as corners of facial contours that are located in various regions of the face are crucial t…

Cited by 0SourceScholar