← Search

Junsong Yuan

92 accepted papers

2026

Learning 3D Shape Fidelity Metric from Real-world Distortions

CVPR 2026

3D generation and reconstruction have become essential in many computer vision applications, where the reconstructed or generated 3D shapes need to appear realistic to human perception. However, traditional metrics like Chamfer Distance to compare two 3D shapes focus primarily on matching accuracy o

Cited by 0SourceScholar
2026

SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation

AAAI 2026technical

3D generation and reconstruction techniques have been widely used in computer games, film, and other content creation areas. As the application grows, there is a growing demand for 3D shapes that look truly realistic. Traditional evaluation methods rely on a ground truth to measure mesh fidelity. Ho

Cited by 0SourcePDFScholar
2026

Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network

AAAI 2026technical

Textured high-fidelity 3D models are crucial for games, AR/VR, and film, but human-aligned evaluation methods still fall behind despite recent advances in 3D reconstruction and generation. Existing metrics, such as Chamfer Distance, often fail to align with how humans evaluate the fidelity of 3D sha

Cited by 0SourcePDFScholar
2025

CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation

ICCV 2025poster

In text-to-image (T2I) generation, achieving fine-grained control over attributes - such as age or smile - remains challenging, even with detailed text prompts. Slider-based methods offer a solution for precise control of image attributes.Existing approaches typically train individual adapter for ea…

Cited by 0SourcePDFScholar
2025

GeoRemover: Removing Objects and Their Causal Visual Artifacts

NeurIPS 2025spotlight

Towards intelligent image editing, object removal should eliminate both the target object and its causal visual artifacts, such as shadows and reflections. However, existing image appearance-based methods either follow strictly mask-aligned training and fail to remove these casual effects which are…

Cited by 0SourcecodeScholar
2025

PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions

ICCV 2025poster

Diffusion-based generative models have shown promise in synthesizing histopathology images to address data scarcity caused by privacy constraints. Diagnostic text reports provide high-level semantic descriptions, and masks offer fine-grained spatial structures essential for representing distinct mor…

2025

Recognizing Actions from Robotic View for Natural Human-Robot Interaction

ICCV 2025poster

Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationary. This setup is more flexible and practical than conventional human action recognition tasks. However, existing benchma…

2025

Text2Outfit: Controllable Outfit Generation with Multimodal Language Models

ICCV 2025poster

Existing outfit recommendation frameworks focus on outfit compatibility prediction and complementary item retrieval. We present a text-driven outfit generation framework, Text2Outfit, which generates outfits controlled by text prompts. Our framework supports two forms of outfit recommendation: 1) Te…

Cited by 0SourcePDFScholar
2025

dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data Analysis

CVPR 2025poster

Federated learning has wide applications in the medical field. It enables knowledge sharing among different healthcare institutes while protecting patients' privacy. However, existing federated learning systems are typically centralized, requiring clients to upload client-specific knowledge to a cen…

Cited by 0SourcePDFScholar
2024

Divide and Fuse: Body Part Mesh Recovery from Partially Visible Human Images

ECCV 2024poster

"We introduce a novel bottom-up approach for human body mesh reconstruction, specifically designed to address the challenges posed by partial visibility and occlusion in input images. Traditional top-down methods, relying on whole-body parametric models like SMPL, falter when only a small part of th…

Cited by 2SourcePDFScholar
2024

Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation

ECCV 2024poster

"In this paper, we explore the visual representations produced from a pre-trained text-to-video (T2V) diffusion model for video understanding tasks. We hypothesize that the latent representation learned from a pretrained generative T2V model encapsulates rich semantics and coherent temporal correspo…

2024

FSC: Few-point Shape Completion

CVPR 2024poster

While previous studies have demonstrated successful 3D object shape completion with a sufficient number of points they often fail in scenarios when a few points e.g. tens of points are observed. Surprisingly via entropy analysis we find that even a few points e.g. 64 points could retain substantial…

2024

IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation

ECCV 2024poster

"Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation methods may not generalize well to synthesized images or videos, and multi-view-based methods have difficulty controlli…

2024

Interaction-centric Spatio-Temporal Context Reasoning for Multi-Person Video HOI Recognition

ECCV 2024poster

"Understanding human-object interaction (HOI) in videos represents a fundamental yet intricate challenge in computer vision, requiring perception and reasoning across both spatial and temporal domains. Despite previous success of object detection and tracking, multi-person video HOI recognition stil…

2024

Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

NeurIPS 2024poster

Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, directly applying these techniques to video models results in unsatisfied frame quality. This issue arises from the limited frame appearance quality in public video datasets, affecting the performan…

2024

Spectrum AUC Difference (SAUCD): Human-aligned 3D Shape Evaluation

CVPR 2024poster

Existing 3D mesh shape evaluation metrics mainly focus on the overall shape but are usually less sensitive to local details. This makes them inconsistent with human evaluation as human perception cares about both overall and detailed shape. In this paper we propose an analytic metric named Spectrum…

Cited by 7SourcePDFScholar
2023

3D-Aware Facial Landmark Detection via Multi-View Consistent Training on Synthetic Data

CVPR 2023poster

Accurate facial landmark detection on wild images plays an essential role in human-computer interaction, entertainment, and medical applications. Existing approaches have limitations in enforcing 3D consistency while detecting 3D/2D facial landmarks due to the lack of multi-view in-the-wild training…

2023

High Fidelity 3D Hand Shape Reconstruction via Scalable Graph Frequency Decomposition

CVPR 2023poster

Despite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high fidelity hand modeling is required, e.g., personalized hand modeling. T…

2023

NeuRBF: A Neural Fields Representation with Adaptive Radial Basis Functions

ICCV 2023oral

We present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for storing local neural features and N-dimensional linear kernels for interpolating features at continuous query points. The…

Cited by 82PDFcodeScholar
2023

Neural Voting Field for Camera-Space 3D Hand Pose Estimation

CVPR 2023poster

We present a unified framework for camera-space 3D hand pose estimation from a single RGB image based on 3D implicit representation. As opposed to recent works, most of which first adopt holistic or pixel-level dense regression to obtain relative 3D hand pose and then follow with complex second-stag…

Cited by 5SourcePDFScholar
2023

POINTACL: Adversarial Contrastive Learning for Robust Point Clouds Representation Under Adversarial Attack

ICASSP 2023accepted

Adversarial contrastive learning (ACL) is considered an effective way to improve the robustness of pre-trained models. In contrastive learning, a projector which consists of multilayer perceptron (MLP) will project high dimension 3D point cloud feature into low dimension for calculating contrastive…

Cited by 0SourceScholar
2023

Progressive Multi-View Human Mesh Recovery with Self-Supervision

AAAI 2023technical

To date, little attention has been given to multi-view 3D human mesh estimation, despite real-life applicability (e.g., motion capture, sport analysis) and robustness to single-view ambiguities. Existing solutions typically suffer from poor generalization performance to new settings, largely due to…

Cited by 16SourcePDFScholar
2023

SOAR: Scene-debiasing Open-set Action Recognition

ICCV 2023poster

Deep models have the risk of utilizing spurious clues to make predictions, e.g., recognizing actions via classifying the background scene. This problem severely degrades the open-set action recognition performance when the testing samples exhibit scene distributions different from the training sampl…

Cited by 18PDFcodeScholar
2023

Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency Learning

ICCV 2023poster

As advanced image manipulation techniques emerge, detecting the manipulation becomes increasingly important. Despite the success of recent learning-based approaches for image manipulation detection, they typically require expensive pixel-level annotations to train, while exhibiting degraded performa…

Cited by 29PDFcodeScholar
2023

Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting

ICCV 2023poster

Hand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate for 3D real-world applications. In this paper, we set up an…

Cited by 19PDFcodeScholar
2022

AiATrack: Attention in Attention for Transformer Visual Tracking

ECCV 2022poster

"Transformer trackers have achieved impressive advancements recently, where the attention mechanism plays an important role. However, the independent correlation computation in the attention mechanism could result in noisy and ambiguous attention weights, which inhibits further performance improveme…

2022

Deformable VisTR: Spatio Temporal Deformable Attention for Video Instance Segmentation

ICASSP 2022accepted

Video instance segmentation (VIS) task requires classifying, segmenting, and tracking object instances over all frames in a video clip. Recently, VisTR [1] has been proposed as end-to-end transformer-based VIS framework, while demonstrating state-of-the-art performance. However, VisTR is slow to con…

Cited by 0SourceScholar
2022

Efficient Video Instance Segmentation via Tracklet Query and Proposal

CVPR 2022poster

Video Instance Segmentation (VIS) aims to simultaneously classify, segment, and track multiple object instances in videos. Recent clip-level VIS takes a short video clip as input each time showing stronger performance than frame-level VIS (tracking-by-segmentation), as more temporal context from mul…

Cited by 52PDFScholar
2022

Generation for Unsupervised Domain Adaptation: A Gan-Based Approach for Object Classification with 3D Point Cloud Data

ICASSP 2022accepted

Recent deep networks have achieved good performance on a variety of 3d points classification tasks. However, these models often face challenges in "wild tasks" where there are considerable differences between the labeled training/source data collected by one Lidar and unseen test/target data collect…

Cited by 0SourceScholar
2022

Joint Global-Local Alignment for Domain Adaptive Semantic Segmentation

ICASSP 2022accepted

Unsupervised domain adaptation has shown promising results in leveraging synthetic (source) images for semantic segmentation of real (target) images. One key issue is how to align data distributions between the source and target domains. Adversarial learning has been applied to align these distribut…

Cited by 0SourceScholar
2022

Learning Transferable Human-Object Interaction Detector With Natural Language Supervision

CVPR 2022poster

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors…

Cited by 66PDFcodeScholar
2022

MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video

CVPR 2022poster

Recent transformer-based solutions have been introduced to estimate 3D human pose from 2D keypoint sequence by considering body joints among all frames globally to learn spatio-temporal correlation. We observe that the motions of different joints differ significantly. However, the previous methods c…

Cited by 339PDFcodeScholar
2022

Neural Correspondence Field for Object Pose Estimation

ECCV 2022poster

"We propose a method for estimating the 6DoF pose of a rigid object with an available 3D model from a single RGB image. Unlike classical correspondence-based methods which predict 3D object coordinates at pixels of the input image, the proposed method predicts 3D object coordinates at 3D query point…

2022

OVIS: Open-Vocabulary Visual Instance Search via Visual-Semantic Aligned Representation Learning

AAAI 2022technical

We introduce the task of open-vocabulary visual instance search (OVIS). Given an arbitrary textual search query, Open-vocabulary Visual Instance Search (OVIS) aims to return a ranked list of visual instances, i.e., image patches, that satisfies the search intent from an image database. The term ``op…

Cited by 2SourcePDFScholar
2022

PREF: Predictability Regularized Neural Motion Fields

ECCV 2022poster

"Knowing the 3D motions in a dynamic scene is essential to many vision applications. Recent progress is mainly focused on estimating the activity of some specific elements like humans. In this paper, we leverage a neural motion field for estimating the motion of all points in a multiview setting. Mo…

Cited by 39SourcePDFScholar
2021

A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder

ICCV 2021poster

We present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requiremen…

Cited by 80PDFScholar
2021

ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization

AAAI 2021technical

The object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foregroun…

Cited by 86SourcePDFScholar
2021

Discovering Human Interactions With Large-Vocabulary Objects via Query and Multi-Scale Detection

ICCV 2021poster

In this work, we study the problem of human-object interaction (HOI) detection with large vocabulary object categories. Previous HOI studies are mainly conducted in the regime of limit object categories (e.g., 80 categories). Their solutions may face new difficulties in both object detection and int…

Cited by 32PDFScholar
2021

High Quality Disparity Remapping With Two-Stage Warping

ICCV 2021poster

A high quality disparity remapping method that preserves 2D shapes and 3D structures, and adjusts disparities of important objects in stereo image pairs is proposed. It is formulated as a constrained optimization problem, whose solution is challenging, since we need to meet multiple requirements of…

Cited by 2PDFScholar
2021

Model-Based 3D Hand Reconstruction via Self-Supervised Learning

CVPR 2021poster

Reconstructing a 3D hand from a single-view RGB image is challenging due to various hand configurations and depth ambiguity. To reliably reconstruct a 3D hand from a monocular image, most state-of-the-art methods heavily rely on 3D annotations at the training stage, but obtaining 3D annotations is e…

Cited by 124PDFcodeScholar
2021

Rethinking Soft Labels for Knowledge Distillation: A Bias–Variance Tradeoff Perspective

ICLR 2021poster

Knowledge distillation is an effective approach to leverage a well-trained network or an ensemble of them, named as the teacher, to guide the training of a student network. The outputs from the teacher network are used as soft labels for supervising the training of a new network. Recent studies (M…

2021

Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student Model

AAAI 2021technical

When adopting deep neural networks for a new vision task, a common practice is to start with fine-tuning some off-the-shelf well-trained network models from the community. Since a new task may require training a different network architecture with new domain data, taking advantage of off-the-shelf m…

Cited by 13SourcePDFScholar
2021

Stacked Homography Transformations for Multi-View Pedestrian Detection

ICCV 2021poster

Multi-view pedestrian detection aims to predict a bird's eye view (BEV) occupancy map from multiple camera views. This task is confronted with two challenges: how to establish the 3D correspondences from views to the BEV map and how to assemble occupancy information across views. In this paper, we p…

Cited by 56PDFScholar
2021

Track To Detect and Segment: An Online Multi-Object Tracker

CVPR 2021poster

Most online multi-object trackers perform object detection stand-alone in a neural net without any input from tracking. In this paper, we present a new online joint detection and tracking model, TraDeS (TRAck to DEtect and Segment), exploiting tracking clues to assist detection end-to-end. TraDeS in…

Cited by 452PDFcodeScholar
2021

Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and Context

AAAI 2021technical

Weakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classificat…

Cited by 32SourcePDFScholar
2020

3DV: 3D Dynamic Voxel for Action Recognition in Depth Video

CVPR 2020poster

For depth-based 3D action recognition, one essential issue is to represent 3D motion pattern effectively and efficiently. To this end, 3D dynamic voxel (3DV) is proposed as a novel 3D motion representation manner. With 3D space voxelization, the key idea of 3DV is to encode the 3D motion information…

Cited by 126PDFcodeScholar
2020

Clustering Driven Deep Autoencoder for Video Anomaly Detection

ECCV 2020poster

Because of the ambiguous definition of anomaly and the complexity of real data, anomaly detection in videos is one of the most challenging problems in intelligent video surveillance. Since the abnormal events are usually different from normal events in appearance and/or in motion behavior, we addres…

Cited by 290SourcePDFScholar
2020

Discovering Human Interactions With Novel Objects via Zero-Shot Learning

CVPR 2020poster

We aim to detect human interactions with novel objects through zero-shot learning. Different from previous works, we allow unseen object categories by using its semantic word embedding. To do so, we design a human-object region proposal network specifically for the human-object interaction detection…

Cited by 50PDFcodeScholar
2020

Hand-Transformer: Non-Autoregressive Structured Modeling for 3D Hand Pose Estimation

ECCV 2020poster

3D hand pose estimation is still far from a well-solved problem mainly due to the highly nonlinear dynamics of hand pose and the difficulties of modeling its inherent structural dependencies. To address this issue, we connect this structured output learning problem with the structured modeling frame…

Cited by 144SourcePDFScholar
2020

Learning Progressive Joint Propagation for Human Motion Prediction

ECCV 2020poster

Despite the great progress in human motion prediction, it remains a challenging task due to the complicated structural dynamics of human behaviors. In this paper, we address this problem in three aspects. First, to capture the long-range spatial correlations and temporal dependencies, we apply a tra…

Cited by 197SourcePDFScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

Temporal Distinct Representation Learning for Action Recognition

ECCV 2020poster

Motivated by the previous success of Two-Dimensional Convolutional Neural Network (2D CNN) on image recognition, researchers endeavor to leverage it to characterize videos. However, one limitation of applying 2D CNN to analyze videos is that different frames of a video share the same 2D CNN kernels,…

Cited by 38SourcePDFScholar
2020

Temporal-Context Enhanced Detection of Heavily Occluded Pedestrians

CVPR 2020poster

State-of-the-art pedestrian detectors have performed promisingly on non-occluded pedestrians, yet they are still confronted by heavy occlusions. Although many previous works have attempted to alleviate the pedestrian occlusion issue, most of them rest on still images. In this paper, we exploit the l…

Cited by 69PDFScholar
2020

Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization

ECCV 2020poster

Weakly-supervised Temporal Action Localization (W-TAL) aims to classify and localize all action instances in an untrimmed video under only video-level supervision. However, without frame-level annotations, it is challenging for W-TAL methods to identify false positive action proposals and generate a…

2019

3D Hand Shape and Pose Estimation From a Single RGB Image

CVPR 2019oral

This work addresses a novel and challenging problem of estimating the full 3D hand shape and pose from a single RGB image. Most current methods in 3D hand analysis from monocular RGB images only focus on estimating the 3D locations of hand keypoints, which cannot fully express the 3D shape of hand.…

Cited by 565PDFScholar
2019

A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth Image

ICCV 2019poster

For 3D hand and body pose estimation task in depth image, a novel anchor-based approach termed Anchor-to-Joint regression network (A2J) with the end-to-end learning ability is proposed. Within A2J, anchor points able to capture global-local spatial context information are densely set on depth image…

Cited by 221PDFcodeScholar
2019

Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional Networks

ICCV 2019poster

Despite great progress in 3D pose estimation from single-view images or videos, it remains a challenging task due to the substantial depth ambiguity and severe self-occlusions. Motivated by the effectiveness of incorporating spatial dependencies and temporal consistencies to alleviate these issues,…

Cited by 588PDFScholar
2019

Joint Representative Selection and Feature Learning: A Semi-Supervised Approach

CVPR 2019poster

In this paper, we propose a semi-supervised approach for representative selection, which finds a small set of representatives that can well summarize a large data collection. Given labeled source data and big unlabeled target data, we aim to find representatives in the target data, which can not onl…

Cited by 4PDFScholar
2019

SO-HandNet: Self-Organizing Network for 3D Hand Pose Estimation With Semi-Supervised Learning

ICCV 2019poster

3D hand pose estimation has made significant progress recently, where Convolutional Neural Networks (CNNs) play a critical role. However, most of the existing CNN-based hand pose estimation methods depend much on the training set, while labeling 3D hand pose on training data is laborious and time-co…

Cited by 105PDFScholar
2019

Temporal Structure Mining for Weakly Supervised Action Detection

ICCV 2019poster

Different from the fully-supervised action detection problem that is dependent on expensive frame-level annotations, weakly supervised action detection (WSAD) only needs video-level annotations, making it more practical for real-world applications. Existing WSAD methods detect action instances by sc…

Cited by 95PDFScholar
2018

Conditional Generative Adversarial Network for Structured Domain Adaptation

CVPR 2018poster

In recent years, deep neural nets have triumphed over many computer vision problems, including semantic segmentation, which is a critical task in emerging autonomous driving and medical image diagnostics applications. In general, training deep neural nets requires a humongous amount of labeled data,…

Cited by 374SourcePDFScholar
2018

Deformable Pose Traversal Convolution for 3D Action and Gesture Recognition

ECCV 2018poster

The representation of 3D pose plays a critical role for 3D body action and hand gesture recognition. Rather than directly representing the 3D pose using its joint locations, in this paper, we propose Deformable Pose Traversal Convolution which applies one-dimensional convolution to traverse the 3D p…

Cited by 87SourcePDFScholar
2018

Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals

CVPR 2018poster

In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods o…

Cited by 277SourcePDFScholar
2018

Salience Guided Depth Calibration for Perceptually Optimized Compressive Light Field 3D Display

CVPR 2018poster

Multi-layer light field displays are a type of computational three-dimensional (3D) display which has recently gained increasing interest for its holographic-like effect and natural compatibility with 2D displays. However, the major shortcoming, depth limitation, still cannot be overcome in the trad…

Cited by 27SourcePDFScholar
2018

Weakly-supervised 3D Hand Pose Estimation from Monocular RGB Images

ECCV 2018poster

Compared with depth-based 3D hand pose estimation, it is more challenging to infer 3D hand pose from monocular RGB images, due to substantial depth ambiguity and the difficulty of obtaining fully-annotated training data. Different from existing learning-based monocular RGB-input approaches that requ…

Cited by 365SourcePDFScholar
2017

3D Convolutional Neural Networks for Efficient and Robust Hand Pose Estimation From Single Depth Images

CVPR 2017poster

We propose a simple, yet effective approach for real-time hand pose estimation from single depth images using three-dimensional Convolutional Neural Networks (3D CNNs). Image based features extracted by 2D CNNs are not directly suitable for 3D hand pose estimation due to the lack of 3D spatial infor…

Cited by 356PDFScholar
2017

HOPE: Hierarchical Object Prototype Encoding for Efficient Object Instance Search in Videos

CVPR 2017poster

This paper tackles the problem of efficient and effective object instance search in videos. To effectively capture the relevance between a query and video frames and precisely localize the particular object, we leverage the object proposals to improve the quality of object instance search in videos.…

Cited by 16PDFScholar
2017

Spatio-Temporal Naive-Bayes Nearest-Neighbor (ST-NBNN) for Skeleton-Based Action Recognition

CVPR 2017poster

Motivated by previous success of using non-parametric methods to recognize objects, e.g., NBNN, we extend it to recognize actions using skeletons. Each 3D action is presented by a sequence of 3D poses. Similar to NBNN, our proposed Spatio-Temporal-NBNN applies stage-to-class distance to classify act…

Cited by 164PDFScholar
2016

From Keyframes to Key Objects: Video Summarization by Representative Object Proposal Selection

CVPR 2016poster

We propose to summarize a video into a few key objects by selecting representative object proposals generated from video frames. This representative selection problem is formulated as a sparse dictionary selection problem, i.e., choosing a few representatives object proposals to reconstruct the whol…

Cited by 136PDFScholar
2016

Robust 3D Hand Pose Estimation in Single Depth Images: From Single-View CNN to Multi-View CNNs

CVPR 2016poster

Articulated hand pose estimation plays an important role in human-computer interaction. Despite the recent progress, the accuracy of existing methods is still not satisfactory, partially due to the difficulty of embedded high-dimensional and non-linear regression problem. Different from the existing…

Cited by 375PDFScholar