← Search

Wenjun Zeng

81 accepted papers

2026

ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning

CVPR 2026

The introduction of negative labels (NLs) has proven effective in enhancing Out-of-Distribution (OOD) detection. However, existing methods often lack an understanding of OOD images, making it difficult to construct an accurate negative space. Furthermore, the absence of negative labels semantically

Cited by 0SourcecodeScholar
2026

Beyond Global Similarity: Multi-Conditional Retrieval for Fine-Grained Cross-Modal Understanding

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have substantially expanded the capabilities of multimodal retrieval, enabling systems to align and retrieve information across visual and textual modalities. Yet, existing benchmarks largely focus on coarse-grained or single-condition alig

Cited by 0SourcecodeScholar
2026

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

ICLR 2026poster

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma–misalignment of 2D image forecasting and 3D action prediction. Besides, such a vision-action entangled training manner limits model learning from large-scale, action-free web video…

Cited by 0SourcecodeScholar
2026

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning

CVPR 2026

Reinforcement Learning (RL) has achieved remarkable success in various domains, yet it often relies on carefully designed programmatic reward functions to guide agent behavior. Designing such reward functions can be challenging and may not generalize well across different tasks. To address this limi

Cited by 0SourceScholar
2026

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation

ICML 2026poster

Evaluating generative AI models is increasingly resource-intensive due to slow generation, expensive raters, and a rapidly growing landscape of models and benchmarks. We propose ProEval, a proactive evaluation framework that leverages transfer learning to efficiently estimate performance and identif…

Cited by 0SourceScholar
2026

PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations

CVPR 2026

Achieving efficient and robust whole-body control (WBC) is essential for enabling humanoid robots to perform complex tasks in dynamic environments. Despite the success of reinforcement learning (RL) in this domain, its sample inefficiency remains a significant challenge due to the intricate dynamics

Cited by 0SourcecodeScholar
2026

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

CVPR 2026

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in

Cited by 0SourceScholar
2026

The Overthinking Predicament: When Reasoning Hurts Ranking

ICLR 2026poster

Document reranking is a key component in information retrieval (IR), aimed at refining initial retrieval results to improve ranking quality for downstream tasks. Recent studies—motivated by large reasoning models (LRMs)—have begun incorporating explicit chain-of-thought (CoT) reasoning into LLM-base…

Cited by 0SourceScholar
2026

Tools are under-documented: Simple Document Expansion Boosts Tool Retrieval

ICLR 2026poster

Large Language Models (LLMs) have recently demonstrated strong capabilities in tool use, yet progress in tool retrieval remains hindered by incomplete and heterogeneous tool documentation. To address this challenge, we introduce Tool-DE, a new benchmark and framework that systematically enriches to…

Cited by 0SourcecodeScholar
2026

When MLLMs Meets Compression Distortion: A Coding Paradigm Tailored to MLLMs

ICLR 2026poster

The increasing deployment of powerful Multimodal Large Language Models (MLLMs), typically hosted on cloud platforms, urgently requires effective compression techniques to efficiently transmit signal inputs (e.g., images, videos) from edge devices with minimal bandwidth usage. However, conventional i…

Cited by 0SourcecodeScholar
2025

Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning

ICCV 2025poster

Training visual reinforcement learning (RL) in practical scenarios presents a significant challenge, i.e., RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these metho…

Cited by 0SourcePDFScholar
2025

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

NeurIPS 2025poster

Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to challenging image-based forecasting, which suffers from redundant i…

Cited by 0SourcecodeScholar
2025

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

ICCV 2025poster

Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that systematically integrates foundation models (e.g., CLIP and DINO…

Cited by 0SourcePDFScholar
2025

MultiConIR: Towards Multi-Condition Information Retrieval

EMNLP 2025

Multi-condition information retrieval (IR) presents a significant, yet underexplored challenge for existing systems. This paper introduces MultiConIR, the first benchmark specifically designed to evaluate retrieval and reranking models under nuanced multi-condition query scenarios across five divers

2025

Open-World Reinforcement Learning over Long Short-Term Imagination

ICLR 2025oral

Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be “short-sighted”, as they are typically trained on short snip…

2025

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

ICCV 2025poster

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-per…

2025

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

ICML 2025poster

User prompts for generative AI models are often underspecified, leading to a misalignment between the user intent and models' understanding. As a result, users commonly have to painstakingly refine their prompts. We study this alignment problem in text-to-image (T2I) generation and propose a prototy…

2025

RLLTE: Long-Term Evolution Project of Reinforcement Learning

AAAI 2025technical

We present RLLTE: a long-term evolution, extremely modular, and open-source framework for reinforcement learning (RL) research and application. Beyond delivering top-notch algorithm implementations, RLLTE also serves as a toolkit for developing algorithms. More specifically, RLLTE decouples the RL a…

2025

ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning

ICCV 2025poster

Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is consistently challenging due to its high non-stationarity a…

2025

UniScene: Unified Occupancy-centric Driving Scene Generation

CVPR 2025poster

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles…

2024

Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion

IJCAI 2024poster

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to predict accurate semantic scenes due to inherent geometric ambiguity and incomplete observations. In this paper, we resort…

2024

Closed-Loop Unsupervised Representation Disentanglement with $\\beta$-VAE Distillation and Diffusion Probabilistic Feedback

ECCV 2024poster

"Representation disentanglement may help AI fundamentally understand the real world and thus benefit both discrimination and generation tasks. It currently has at least three unresolved core issues: (i) heavy reliance on label annotation and synthetic data — causing poor generalization on natural sc…

Cited by 7SourcePDFScholar
2024

Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models

NeurIPS 2024poster

Disentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption that semantic factors are statistically independent. In reality…

Cited by 2SourcePDFScholar
2024

HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects

ECCV 2024poster

"Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of multiple objects. Thus, we propose HIMO, a large-scale MoCap da…

Cited by 3SourcePDFScholar
2024

Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion

ECCV 2024poster

"Camera-based 3D semantic scene completion (SSC) is pivotal for predicting complicated 3D layouts with limited 2D image observations. The existing mainstream solutions generally leverage temporal information by roughly stacking history frames to supplement the current frame, such straightforward tem…

2024

Inter-X: Towards Versatile Human-Human Interaction Analysis

CVPR 2024poster

The analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions lack of hand gestures and fine-grained textual descriptions. To better perceive and generate human-hum…

2024

Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning

NeurIPS 2024poster

Training offline RL models using visual inputs poses two significant challenges, *i.e.*, the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. T…

2024

One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene Perception

AAAI 2024technical

Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes directly with geometric correspondence, attempting to fully address…

Cited by 3SourcePDFScholar
2024

Rate-Distortion-Cognition Controllable Versatile Neural Image Compression

ECCV 2024poster

"Recently, the field of Image Coding for Machines (ICM) has garnered heightened interest and significant advances thanks to the rapid progress of learning-based techniques for image compression and analysis. Previous studies often require training separate codecs to support various bitrate levels, m…

Cited by 4SourcePDFScholar
2024

ReGenNet: Towards Human Action-Reaction Synthesis

CVPR 2024poster

Humans constantly interact with their surrounding environments. Current human-centric generative models mainly focus on synthesizing humans plausibly interacting with static scenes and objects while the dynamic human action-reaction synthesis for ubiquitous causal human-human interactions is less ex…

2024

Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation

NeurIPS 2024spotlight

There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a…

Cited by 2SourcePDFScholar
2023

ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation

ICCV 2023poster

We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equippe…

Cited by 74PDFScholar
2023

Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement Learning

ICML 2023poster

We present AIRS: **A**utomatic **I**ntrinsic **R**eward **S**haping that intelligently and adaptively provides high-quality intrinsic rewards to enhance exploration in reinforcement learning (RL). More specifically, AIRS selects shaping function from a predefined set based on the estimated task retu…

2023

NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation

ICCV 2023poster

3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general c…

Cited by 11PDFcodeScholar
2023

WEDGE: Web-Image Assisted Domain Generalization for Semantic Segmentation

ICRA 2023poster

Domain generalization for semantic segmentation is highly demanded in real applications, where a trained model is expected to work well in previously unseen domains. One challenge lies in the lack of data which could cover the diverse distributions of the possible unseen domains for training. In thi…

Cited by 27SourceScholar
2022

Learning Disentangled Representation by Exploiting Pretrained Generative Models: A Contrastive Learning View

ICLR 2022poster

From the intuitive notion of disentanglement, the image variations corresponding to different generative factors should be distinct from each other, and the disentangled representation should reflect those variations with separate dimensions. To discover the generative factors and learn disentangled…

2022

Lifelong Unsupervised Domain Adaptive Person Re-Identification With Coordinated Anti-Forgetting and Adaptation

CVPR 2022poster

Unsupervised domain adaptive person re-identification (ReID) has been extensively investigated to mitigate the adverse effects of domain gaps. Those works assume the target domain data can be accessible all at once. However, for the real-world streaming data, this hinders the timely adaptation to ch…

Cited by 42PDFScholar
2022

ReSTR: Convolution-Free Referring Image Segmentation Using Transformers

CVPR 2022poster

Referring image segmentation is an advanced semantic segmentation task where target is not a predefined class but is described in natural language. Most of existing methods for this task rely heavily on convolutional neural networks, which however have trouble capturing long-range dependencies betwe…

Cited by 176PDFScholar
2022

Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph

ICLR 2022poster

This paper addresses the unsupervised learning of content-style decomposed representation. We first give a definition of style and then model the content-style representation as a token-level bipartite graph. An unsupervised framework, named Retriever, is proposed to learn such representations. Firs…

2022

Robust Multi-Object Tracking by Marginal Inference

ECCV 2022poster

"Multi-object tracking in videos requires to solve a fundamental problem of one-to-one assignment between objects in adjacent frames. Most methods address the problem by first discarding impossible pairs whose feature distances are larger than a threshold, followed by linking objects using Hungarian…

Cited by 24SourcePDFScholar
2022

Sparse MLP for Image Recognition: Is Self-Attention Really Necessary?

AAAI 2022technical

Transformers have sprung up in the field of computer vision. In this work, we explore whether the core self-attention module in Transformer is the key to achieving excellent performance in image recognition. To this end, we build an attention-free network called sMLPNet based on the existing MLP-bas…

2022

Towards Building A Group-based Unsupervised Representation Disentanglement Framework

ICLR 2022poster

Disentangled representation learning is one of the major goals of deep learning, and is a key step for achieving explainable and generalizable models. The key idea of the state-of-the-art VAE-based unsupervised representation disentanglement methods is to minimize the total correlation of the joint…

2022

VirtualPose: Learning Generalizable 3D Human Pose Models from Virtual Data

ECCV 2022poster

"While monocular 3D pose estimation seems to have achieved very accurate results on the public datasets, their generalization ability is largely overlooked. In this work, we perform a systematic evaluation of the existing methods and find that they get notably larger errors when tested on different…

2022

When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism

AAAI 2022technical

Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. However, is the attention mechanism truly an indispensable part of ViT? Can it be replaced by some other alternatives? To de…

2021

An Empirical Study of the Collapsing Problem in Semi-Supervised 2D Human Pose Estimation

ICCV 2021poster

The state-of-the-art semi-supervised learning models are consistency-based which learn about unlabeled images by maximizing the similarity between different augmentations of an image. But when we apply the methods to human pose estimation which has extremely imbalanced class distribution, the models…

Cited by 37PDFcodeScholar
2021

Exploiting Sample Uncertainty for Domain Adaptive Person Re-Identification

AAAI 2021technical

Many unsupervised domain adaptive (UDA) person ReID approaches combine clustering-based pseudo-label prediction with feature fine-tuning. However, because of domain gap, the pseudo-labels are not always reliable and there are noisy/incorrect labels. This would mislead the feature representation lea…

Cited by 190SourcePDFScholar
2021

MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain Adaptation

CVPR 2021poster

For unsupervised domain adaptation (UDA), to alleviate the effect of domain shift, many approaches align the source and target domains in the feature space by adversarial learning or by explicitly aligning their statistics. However, the optimization objective of such domain alignment is generally no…

Cited by 137PDFScholar
2021

PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement Learning

NeurIPS 2021poster

Learning good feature representations is important for deep reinforcement learning (RL). However, with limited experience, RL often suffers from data inefficiency for training. For un-experienced or less-experienced trajectories (i.e., state-action sequences), the lack of data limits the use of them…

2021

Re-Energizing Domain Discriminator With Sample Relabeling for Adversarial Domain Adaptation

ICCV 2021poster

Many unsupervised domain adaptation (UDA) methods exploit domain adversarial training to align the features to reduce domain gap, where a feature extractor is trained to fool a domain discriminator in order to have aligned feature distributions. The discrimination capability of the domain classifier…

Cited by 18PDFScholar
2021

S2R-DepthNet: Learning a Generalizable Depth-Specific Structural Representation

CVPR 2021poster

Human can infer the 3D geometry of a scene from a sketch instead of a realistic image, which indicates that the spatial structure plays a fundamental role in understanding the depth of scenes. We are the first to explore the learning of a depth-specific structural representation, which captures the…

Cited by 61PDFcodeScholar
2021

Self-Supervised Visual Representations Learning by Contrastive Mask Prediction

ICCV 2021poster

Advanced self-supervised visual representation learning methods rely on the instance discrimination (ID) pretext task. We point out that the ID task has an implicit semantic consistency (SC) assumption, which may not hold in unconstrained datasets. In this paper, we propose a novel contrastive mask…

Cited by 49PDFcodeScholar
2021

ToAlign: Task-Oriented Alignment for Unsupervised Domain Adaptation

NeurIPS 2021poster

Unsupervised domain adaptive classifcation intends to improve the classifcation performance on unlabeled target domain. To alleviate the adverse effect of domain shift, many approaches align the source and target domains in the feature space. However, a feature is usually taken as a whole for alignm…

2021

Uncertainty-Aware Few-Shot Image Classification

IJCAI 2021poster

Few-shot image classification learns to recognize new categories from limited labelled data. Metric learning based approaches have been widely investigated, where a query sample is classified by finding the nearest prototype from the support set based on their feature similarities. A neural network…

Cited by 30SourcePDFScholar
2021

Unsupervised Visual Representation Learning by Tracking Patches in Video

CVPR 2021poster

Inspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to learn the visual representations. Modelled on the Catch game played by the children, we design a Catch-the-Patch (CtP) g…

Cited by 31PDFcodeScholar
2021

Very Important Person Localization in Unconstrained Conditions: A New Benchmark

AAAI 2021technical

This paper presents a new high-quality dataset for Very Important Person Localization (VIPLoc), named Unconstrained-7k. Generally, current datasets: 1) are limited in scale; 2) built under simple and constrained conditions, where the number of disturbing non-VIPs is not large, the scene is relativel…

2020

Beyond Intra-modality: A Survey of Heterogeneous Person Re-identification

IJCAI 2020poster

An efficient and effective person re-identification (ReID) system relieves the users from painful and boring video watching and accelerates the process of video analysis. Recently, with the explosive demands of practical applications, a lot of research efforts have been dedicated to heterogeneous pe…

2020

Fusing Wearable IMUs With Multi-View Images for Human Pose Estimation: A Geometric Approach

CVPR 2020poster

We propose to estimate 3D human pose from multi-view images and a few IMUs attached at person's limbs. It operates by firstly detecting 2D poses from the two signals, and then lifting them to the 3D space. We present a geometric approach to reinforce the visual features of each pair of joints based…

Cited by 80PDFcodeScholar
2020

Global Distance-distributions Separation for Unsupervised Person Re-identification

ECCV 2020poster

Supervised person re-identification (ReID) often has poor scalability and usability in real-world deployments due to domain gaps and the lack of annotations for the target domain data. Unsupervised person ReID through domain adaptation is attractive yet challenging. Existing unsupervised ReID approa…

Cited by 90SourcePDFScholar
2020

Joint Time-Frequency and Time Domain Learning for Speech Enhancement

IJCAI 2020poster

For single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framewo…

Cited by 0SourcePDFScholar
2020

Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-Identification

CVPR 2020poster

Video-based person re-identification (reID) aims at matching the same person across video clips. It is a challenging task due to the existence of redundancy among frames, newly revealed appearance, occlusion, and motion blurs. In this paper, we propose an attentive feature aggregation module, namely…

Cited by 138PDFScholar
2020

Multi-Scale Group Transformer for Long Sequence Modeling in Speech Separation

IJCAI 2020poster

In this paper, we introduce Transformer to the time-domain methods for single-channel speech separation. Transformer has the potential to boost speech separation performance because of its strong sequence modeling capability. However, its computational complexity, which grows quadratically with the…

Cited by 0SourcePDFScholar
2020

Relation-Aware Global Attention for Person Re-Identification

CVPR 2020poster

For person re-identification (re-id), attention mechanisms have become attractive as they aim at strengthening discriminative features and suppressing irrelevant ones, which matches well the key of re-id, i.e., discriminative feature learning. Previous approaches typically learn attention using loca…

Cited by 713PDFcodeScholar
2020

Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

CVPR 2020poster

Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this…

Cited by 635PDFcodeScholar
2020

Style Normalization and Restitution for Generalizable Person Re-Identification

CVPR 2020poster

Existing fully-supervised person re-identification (ReID) methods usually suffer from poor generalization capability caused by domain gaps. The key to solving this problem lies in filtering out identity-irrelevant interference and learning domain-invariant person representations. In this paper, we a…

Cited by 448PDFcodeScholar
2020

VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment

ECCV 2020poster

We present mph{VoxelPose} to estimate $3$D poses of multiple people from multiple camera views. In contrast to the previous efforts which require to establish cross-view correspondence based on noisy and incomplete $2$D pose estimates, mph{VoxelPose} directly operates in the $3$D space therefore avo…

2019

Moving Indoor: Unsupervised Video Depth Learning in Challenging Environments

ICCV 2019poster

Recently unsupervised learning of depth from videos has made remarkable progress and the results are comparable to fully supervised methods in outdoor scenes like KITTI. However, there still exist great challenges when directly applying this technology in indoor environments, e.g., large areas of no…

Cited by 87PDFScholar
2019

SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object Tracking

CVPR 2019poster

The greatest challenge facing visual object tracking is the simultaneous requirements on robustness and discrimination power. In this paper, we propose a SiamFC-based tracker, named SPM-Tracker, to tackle this challenge. The basic idea is to address the two requirements in two separate matching stag…

Cited by 290PDFScholar
2019

Unsupervised High-Resolution Depth Learning From Videos With Dual Networks

ICCV 2019poster

Unsupervised depth learning takes the appearance difference between a target view and a view synthesized from its adjacent frame as supervisory signal. Since the supervisory signal only comes from images themselves, the resolution of training data significantly impacts the performance. High-resoluti…

Cited by 77PDFScholar
2018

Adding Attentiveness to the Neurons in Recurrent Neural Networks

ECCV 2018poster

Recurrent neural networks (RNNs) are capable of modeling the temporal dynamics of complex sequential information. However, the structures of existing RNN neurons mainly focus on controlling the contributions of current and historical information but do not explore the different importance levels of…

Cited by 105SourcePDFScholar
2018

MiCT: Mixed 3D/2D Convolutional Tube for Human Action Recognition

CVPR 2018poster

Human actions in videos are three-dimensional (3D) signals. Recent attempts use 3D convolutional neural networks (CNNs) to explore spatio-temporal information for human action recognition. Though promising, 3D CNNs have not achieved high performanceon on this task with respect to their well-establis…

Cited by 295SourcePDFScholar
2017

Human Pose Estimation Using Global and Local Normalization

ICCV 2017poster

In this paper, we address the problem of estimating the positions of human joints, i.e., articulated pose estimation. Recent state-of-the-art solutions model two key issues, joint detection and spatial configuration refinement, together using convolutional neural networks. Our work mainly focuses on…

Cited by 82PDFScholar
2017

View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition From Skeleton Data

ICCV 2017poster

Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints du…

Cited by 683PDFcodeScholar
2015

High-Speed Hyperspectral Video Acquisition With a Dual-Camera Architecture

CVPR 2015poster

We propose a novel dual-camera design to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. Our work has two key technical contributions. First, we build a dual-camera system that simultaneously captures a panchromatic video at a high frame rate and a hypers…

Cited by 116SourcePDFScholar