← Search

Weiming Hu

74 accepted papers

2026

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval

CVPR 2026

Composed Image Retrieval (CIR) has demonstrated significant potential by enabling flexible, multimodal queries that combine a reference image and modification text. However, CIR inherently prioritizes semantic matching, struggling to reliably retrieve a user-specified instance across contexts. In pr

Cited by 0SourceScholar
2026

HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes

AAAI 2026technical

This paper tackles the challenging task of achieving storage-efficient yet high-fidelity motion representation in large-scale dynamic 3D Gaussian Splatting. Our motivation stems from the truth that existing urban-scale methods, which rely on massive and unstructured individual Gaussians for scene mo

Cited by 0SourcePDFScholar
2026

Integrating Diverse Assignment Strategies into DETRs

AAAI 2026technical

Label assignment is a critical component in object detectors, particularly within DETR-style frameworks where the one-to-one matching strategy, despite its end-to-end elegance, suffers from slow convergence due to sparse supervision. While recent works have explored one-to-many assignments to enrich

Cited by 0SourcePDFScholar
2026

MMhops-R1: Multimodal Multi-hop Reasoning

AAAI 2026technical

The ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-ste

Cited by 0SourcePDFScholar
2026

Resolving the Timestep Scaling Paradox in Spiking Neural Networks with a Timestep-Scalable Neuron Model

ICML 2026poster

Spiking Neural Networks (SNNs) have garnered increasing attention for their biological plausibility, energy efficiency, and temporal modeling capability. Due to the non-differentiability of spike generation, a widely used supervised training method for SNNs is backpropagation through time with surro…

Cited by 0SourceScholar
2026

SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker

CVPR 2026

Parameter-efficient fine-tuning (PEFT) in multimodal tracking reveals a concerning trend where recent performance gains are often achieved at the cost of inflated parameter budgets, which fundamentally erodes PEFT's efficiency promise. In this work, we introduce SEATrack, a Simple, Efficient, and Ad

Cited by 0SourcecodeScholar
2026

SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names

CVPR 2026

Open-vocabulary object detection (OVD) aims to detect objects described by arbitrary text, but most existing methods operate at a coarse category level and struggle with fine-grained, attribute-sensitive queries. We address this from both model and data perspectives. We propose a Semantic-Retrieval-

Cited by 0SourceScholar
2026

Temporal Representation Enhancement (TRE): Learning to Forget Dominant Patterns for Enhanced Temporal Spiking Features

CVPR 2026

Spiking Neural Networks (SNNs) naturally process visual inputs across multiple timesteps, offering rich temporal dynamics and energy-efficient computation. However, the temporally invariant supervision commonly used in training tends to reinforce the same dominant response patterns across timesteps,

Cited by 0SourceScholar
2025

D-RAG: Differentiable Retrieval-Augmented Generation for Knowledge Graph Question Answering

EMNLP 2025

Knowledge Graph Question Answering (KGQA) aims to answer natural language questions based on knowledge graphs.Recent approaches apply the Retrieval-Augmented Generation (RAG) paradigm to incorporate Large Language Models (LLMs) to this task, where a retriever selects a question-related subgraph and

Cited by 0SourcePDFScholar
2025

DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural Networks

ICLR 2025poster

Spiking Neural Networks (SNNs), with their biologically inspired spatio-temporal dynamics and spike-driven processing, are emerging as a promising low-power alternative to traditional Artificial Neural Networks (ANNs). However, the complex neuronal dynamics and non-differentiable spike communication…

Cited by 0SourcePDFScholar
2025

Each Complexity Deserves a Pruning Policy

NeurIPS 2025poster

The established redundancy in visual tokens within large vision–language models (LVLMs) allows for pruning to effectively reduce their substantial computational demands. Empirical evidence from previous works indicates that visual tokens in later decoder stages receive less attention than shallow la…

Cited by 0SourcecodeScholar
2025

Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection

ICCV 2025poster

We present the first work demonstrating that a pure Mamba block can achieve efficient Dense Global Fusion, meanwhile guaranteeing top performance for camera-LiDAR multi-modal 3D object detection. Our motivation stems from the observation that existing fusion strategies are constrained by their inabi…

2025

LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning

ICCV 2025poster

Implicit degradation estimation-based blind super-resolution (IDE-BSR) hinges on extracting the implicit degradation representation (IDR) of the LR image and adapting it to LR image features to guide HR detail restoration. Although IDE-BSR has shown potential in dealing with noise interference and c…

2025

MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural Networks

NeurIPS 2025poster

Brain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when proce…

Cited by 0SourcecodeScholar
2025

Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations

EMNLP 2025

Large Vision-Language Models (LVLMs) suffer from serious hallucination problems, where the model-generated responses are inconsistent with the visual inputs. Existing hallucination mitigation methods are mainly based on preference alignment and require external human annotations or auxiliary models

2025

Online Segment Any 3D Thing as Instance Tracking

NeurIPS 2025poster

Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to aggregate semantic information from Vision Foundation Models (VF…

Cited by 0SourcecodeScholar
2025

Reversing Flow for Image Restoration

CVPR 2025poster

Image restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, wh…

Cited by 0SourcePDFScholar
2025

SSTrack: Sample-interval Scheduling for Lightweight Visual Object Tracking

IJCAI 2025

In recent years, CPU real-time object tracking has gained significant attention due to its broad applications such as UAV-tracking. To maintain computational efficiency, most existing CPU real-time object trackers rely on lightweight backbones and employ a single initial template image without inter

2025

SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D Tracking

NeurIPS 2025poster

While existing query-based 3D end-to-end visual trackers integrate detection and tracking via the *tracking-by-attention* paradigm, these two chicken-and-egg tasks encounter optimization difficulties when sharing the same parameters. Our findings reveal that these difficulties arise due to two inher…

Cited by 0SourcecodeScholar
2025

Towards More Discriminative Feature Learning in SNNs with Temporal-Self-Erasing Supervision

AAAI 2025technical

Spiking Neural Networks (SNNs) are biologically inspired models that process visual inputs over multiple time steps. However, they often struggle with limited feature discrimination along the temporal dimension due to inherent spatiotemporal invariance. This limitation arises from the redundant acti…

Cited by 0SourcePDFScholar
2025

VisionMath: Vision-Form Mathematical Problem-Solving

ICCV 2025poster

Mathematical problems in real-world scenarios are often presented in a purely vision-form, where textual problem statement and accompanying math figures, e.g., geometry figures and functional graphs, are integrated into a single image. This vision-form problem-solving task requires precise comprehen…

2025

Visual-Instructed Degradation Diffusion for All-in-One Image Restoration

CVPR 2025poster

Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that…

2025

Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding

NeurIPS 2025poster

Speculative decoding improves LLM inference by generating and verifying multiple tokens in parallel, but existing systems suffer from suboptimal performance due to a mismatch between dynamic speculation and static runtime assumptions. We present Yggdrasil, a co-designed system that enables latency-o…

Cited by 0SourceScholar
2024

A-Teacher: Asymmetric Network for 3D Semi-Supervised Object Detection

CVPR 2024poster

This work proposes the first online asymmetric semi-supervised framework namely A-Teacher for LiDAR-based 3D object detection. Our motivation stems from the observation that 1) existing symmetric teacher-student methods for semi-supervised 3D object detection have characterized simplicity but impede…

Cited by 2SourcePDFScholar
2024

Animate3D: Animating Any 3D Model with Multi-view Video Diffusion

NeurIPS 2024poster

Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view attributes, and their results suffer from spatiotemporal inc…

Cited by 13SourcePDFScholar
2024

BEV2PR: BEV-Enhanced Visual Place Recognition with Structural Cues

IROS 2024

In this paper, we propose a new image-based visual place recognition (VPR) framework by exploiting the structural cues in bird’s-eye view (BEV) from a single monocular camera. The motivation arises from two key observations about place recognition methods based on both appearance and structure: 1) F

Cited by 4SourcecodeScholar
2024

Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video

ICLR 2024poster

In this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need for tedious multi-view data collection and camera calibration…

2024

EA-VTR: Event-Aware Video-Text Retrieval

ECCV 2024poster

"Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted video-level cross-modal contrastive learning also struggles to…

Cited by 3SourcePDFScholar
2024

How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?

CVPR 2024poster

Dominant dual-encoder models enable efficient image-text retrieval but suffer from limited accuracy while the cross-encoder models offer higher accuracy at the expense of efficiency. Distilling cross-modality matching knowledge from cross-encoder to dual-encoder provides a natural approach to harnes…

Cited by 2SourcePDFScholar
2024

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

EMNLP 2024main

Built on the power of LLMs, numerous multimodal large language models (MLLMs) have recently achieved remarkable performance on various vision-language tasks. However, most existing MLLMs and benchmarks primarily focus on single-image input scenarios, leaving the performance of MLLMs when handling re…

Cited by 10SourcePDFScholar
2024

PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts

ECCV 2024poster

"Due to the diversity of assessment requirements in various application scenarios for the IQA task, existing IQA methods struggle to directly adapt to these varied requirements after training. Thus, when facing new requirements, a typical approach is fine-tuning these models on datasets specifically…

2024

STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians

ECCV 2024poster

"Recent progress in pre-trained diffusion models and 3D generation have spurred interest in 4D content creation. However, achieving high-fidelity 4D generation with spatial-temporal consistency remains a challenge. In this work, we propose STAG4D, a novel framework that combines pre-trained diffusio…

Cited by 46SourcePDFScholar
2024

Self-Training with Pseudo-Label Scorer for Aspect Sentiment Quad Prediction

ACL 2024long

Aspect Sentiment Quad Prediction (ASQP) aims to predict all quads (aspect term, aspect category, opinion term, sentiment polarity) for a given review, which is the most representative and challenging task in aspect-based sentiment analysis. A key challenge in the ASQP task is the scarcity of labeled…

2024

Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training

COLING 2024main

In vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the reconstruction targets for MIM lack high-level semantics, and text is not sufficiently involved in masked modeling. These two…

Cited by 0SourcePDFScholar
2024

Set Prediction Guided by Semantic Concepts for Diverse Video Captioning

AAAI 2024technical

Diverse video captioning aims to generate a set of sentences to describe the given video in various aspects. Mainstream methods are trained with independent pairs of a video and a caption from its ground-truth set without exploiting the intra-set relationship, resulting in low diversity of generated…

Cited by 3SourcePDFScholar
2024

Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval

COLING 2024main

In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we prop…

2024

VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization

NeurIPS 2024poster

Bird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{g…

2023

A Closer Look at Self-Supervised Lightweight Vision Transformers

ICML 2023poster

Self-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms promote lightweight ViTs' performance is considerably less studied. In this work, we develop and benchmark several self…

2023

AUNet: Learning Relations Between Action Units for Face Forgery Detection

CVPR 2023poster

Face forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same domain. However, the problem remains chall…

Cited by 56SourcePDFScholar
2023

Exploiting Contextual Objects and Relations for 3D Visual Grounding

NeurIPS 2023poster

3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information…

2023

Learning To Exploit the Sequence-Specific Prior Knowledge for Image Processing Pipelines Optimization

CVPR 2023poster

The hardware image signal processing (ISP) pipeline is the intermediate layer between the imaging sensor and the downstream application, processing the sensor signal into an RGB image. The ISP is less programmable and consists of a series of processing modules. Each processing module handles a subta…

Cited by 7SourcePDFScholar
2023

Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action Recognition

ICASSP 2023accepted

Video action recognition is faced with the challenges of both huge computation burden and performance requirements. Using compressed domain data, which saves much decoding computation, is a possible solution. Unfortunately, existing compressed-domain-based (CD) methods fail to obtain high performanc…

Cited by 0SourceScholar
2023

Order-Prompted Tag Sequence Generation for Video Tagging

ICCV 2023poster

Video Tagging intends to infer multiple tags spanning relevant content for a given video. Typically, video tags are freely defined and uploaded by a variety of users, so they have two characteristics: abundant in quantity and disordered intra-video. It is difficult for the existing multi-label class…

Cited by 4PDFScholar
2023

PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer

AAAI 2023technical

3D object detection in autonomous driving aims to reason “what” and “where” the objects of interest present in a 3D world. Following the conventional wisdom of previous 2D object detection, existing methods often adopt the canonical Cartesian coordinate system with perpendicular axis. However, we co…

2023

ViLEM: Visual-Language Error Modeling for Image-Text Retrieval

CVPR 2023poster

Dominant pre-training works for image-text retrieval adopt "dual-encoder" architecture to enable high efficiency, where two encoders are used to extract image and text representations and contrastive learning is employed for global alignment. However, coarse-grained global alignment ignores detailed…

Cited by 14SourcePDFScholar
2023

ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual Tracking

NeurIPS 2023spotlight

Recently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions.…

2022

Attention-Aware Learning for Hyperparameter Prediction in Image Processing Pipelines

ECCV 2022poster

"Between the imaging sensor and the image applications, the hardware image signal processing (ISP) pipelines reconstruct an RGB image from the sensor signal and feed it into downstream tasks. The processing blocks in ISPs depend on a set of tunable hyperparameters that have a complex interaction wit…

Cited by 13SourcePDFScholar
2022

EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding Matching

CVPR 2022poster

Current metrics for video captioning are mostly based on the text-level comparison between reference and candidate captions. However, they have some insuperable drawbacks, e.g., they cannot handle videos without references, and they may result in biased evaluation due to the one-to-many nature of vi…

Cited by 43PDFcodeScholar
2022

Improving Visual Grounding With Visual-Linguistic Verification and Iterative Reasoning

CVPR 2022poster

Visual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated proposals or anchors, and fuse these features with the text em…

Cited by 147PDFcodeScholar
2022

Learning Target-aware Representation for Visual Tracking via Informative Interactions

IJCAI 2022poster

We introduce a novel backbone architecture to improve target-perception ability of feature representation for tracking. Having observed de facto frameworks perform feature matching simply using the backbone outputs for target localization, there is no direct feedback from the matching module to the…

Cited by 61SourcePDFScholar
2022

Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification

IJCAI 2022poster

Compared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On t…

Cited by 16SourcePDFScholar
2022

One More Check: Making “Fake Background” Be Tracked Again

AAAI 2022technical

The one-shot multi-object tracking, which integrates object detection and ID embedding extraction into a unified network, has achieved groundbreaking results in recent years. However, current one-shot trackers solely rely on single-frame detections to predict candidate bounding boxes, which may be u…

2022

Open-Vocabulary One-Stage Detection With Hierarchical Visual-Language Knowledge Distillation

CVPR 2022poster

Open-vocabulary object detection aims to detect novel object categories beyond the training set. The advanced open-vocabulary two-stage detectors employ instance-level visual-to-visual knowledge distillation to align the visual space of the detector with the semantic space of the Pre-trained Visual-…

Cited by 106PDFcodeScholar
2021

Channel-Wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition

ICCV 2021poster

Graph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. In GCNs, graph topology dominates feature aggregation and therefore is the key to extracting representative features. In this work, we propose a novel Channel-wise Topology…

Cited by 872PDFcodeScholar
2021

DPFPS: Dynamic and Progressive Filter Pruning for Compressing Convolutional Neural Networks from Scratch

AAAI 2021technical

Filter pruning is a commonly used method for compressing Convolutional Neural Networks (ConvNets), due to its friendly hardware supporting and flexibility. However, existing methods mostly need a cumbersome procedure, which brings many extra hyper-parameters and training epochs. This is because only…

2021

Differentiable Convolution Search for Point Cloud Processing

ICCV 2021poster

Exploiting convolutional neural networks for point cloud processing is quite challenging, due to the inherent irregular distribution and discrete shape representation of point clouds. To address these problems, many handcrafted convolution variants have sprung up in recent years. Though with elabora…

Cited by 10PDFScholar
2021

Learn To Match: Automatic Matching Network Design for Visual Tracking

ICCV 2021poster

Siamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remarkable success, it is important to note that the heuristic matching network design relies heavily on expert experience. M…

Cited by 239PDFcodeScholar
2021

Open-Book Video Captioning With Retrieve-Copy-Generate Network

CVPR 2021poster

In this paper, we convert traditional video captioning task into a new paradigm, i.e., Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a…

Cited by 125PDFScholar
2020

Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model

ECCV 2020poster

Recently, video streams have occupied a large proportion of Internet traffic, most of which contain human faces. Hence, it is necessary to predict saliency on multiple-face videos, which can provide attention cues for many content based applications. However, most of multiple-face prediction works o…

2020

Object Relational Graph With Teacher-Recommended Learning for Video Captioning

CVPR 2020poster

Taking full advantage of the information from both vision and language is critical for the video captioning task. Existing models lack adequate visual representation due to the neglect of interaction between object, and sufficient training for content-related words due to long-tailed problems. In th…

Cited by 387PDFScholar
2019

Anchor Diffusion for Unsupervised Video Object Segmentation

ICCV 2019poster

Unsupervised video object segmentation has often been tackled by methods based on recurrent neural networks and optical flow. Despite their complexity, these kinds of approach tend to favour short-term temporal dependencies and are thus prone to accumulating inaccuracies, which cause drift over time…

Cited by 142PDFcodeScholar
2019

Fast Online Object Tracking and Segmentation: A Unifying Approach

CVPR 2019poster

In this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach. Our method, dubbed SiamMask, improves the offline training procedure of popular fully-convolutional Siamese approaches for object trackin…

Cited by 1768PDFScholar
2019

Knowledge Distillation via Instance Relationship Graph

CVPR 2019poster

The key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including insta…

Cited by 371PDFScholar
2018

Deep Cost-Sensitive and Order-Preserving Feature Learning for Cross-Population Age Estimation

CVPR 2018poster

Facial age estimation from a face image is an important yet very challenging task in computer vision, since humans with different races and/or genders, exhibit quite different patterns in their facial aging processes. To deal with the influence of race and gender, previous methods perform age estima…

Cited by 36SourcePDFScholar
2018

Distractor-aware Siamese Networks for Visual Object Tracking

ECCV 2018poster

Recently, Siamese networks have drawn great attention in visual tracking community because of their balanced accuracy and speed. However, features used in most Siamese tracking approaches can only discriminate foreground from the non-semantic backgrounds. The semantic backgrounds are always consider…

2018

Interaction-aware Spatio-temporal Pyramid Attention Networks for Action Classification

ECCV 2018poster

Local features at neighboring spatial positions in feature maps have high correlation since their receptive fields are often overlapped. Self-attention usually uses the weighted sum (or other functions) with internal elements of each local feature to obtain its weight score, which ignores interactio…

Cited by 118SourcePDFScholar
2018

Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual Tracking

CVPR 2018poster

Offline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance obje…

2018

Visual Tracking via Spatially Aligned Correlation Filters Network

ECCV 2018poster

Correlation filters based trackers rely on a periodic assumption of the search sample to efficiently distinguish the target from the background. This assumption however yields undesired boundary effects and restricts aspect ratios of search samples. To handle these issues, an end-to-end deep archite…

2017

Spatio-Temporal Self-Organizing Map Deep Network for Dynamic Object Detection From Videos

CVPR 2017poster

In dynamic object detection, it is challenging to construct an effective model to sufficiently characterize the spatial-temporal properties of the background. This paper proposes a new Spatio-Temporal Self-Organizing Map (STSOM) deep network to detect dynamic objects in complex scenarios. The propos…

Cited by 18PDFScholar
2015

Multi-Feature Max-Margin Hierarchical Bayesian Model for Action Recognition

CVPR 2015poster

In this paper, a multi-feature max-margin hierarchical Bayesian model (M3HBM) is proposed for action recognition. Different from existing methods which separate representation and classification into two steps, M3HBM jointly learns a high-level representation by combining a hierarchical generative m…

Cited by 36SourcePDFScholar