← Search

Guangming Zhu

14 accepted papers

2026

Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions

CVPR 2026

The rapid movements and agile maneuvers of unmanned aerial vehicles (UAVs) induce significant observational challenges for multi-object tracking (MOT). However, existing UAV-perspective MOT benchmarks often lack these complexities, featuring predominantly predictable camera dynamics and linear motio

Cited by 0SourcecodeScholar
2026

FSD-CAP: Fractional Subgraph Diffusion with Class-Aware Propagation for Graph Feature Imputation

ICLR 2026poster

Imputing missing node features in graphs is challenging, particularly under high missing rates. Existing methods based on latent representations or global diffusion often fail to produce reliable estimates, and may propagate errors across the graph. We propose FSD-CAP, a two-stage framework designed…

Cited by 0SourceScholar
2026

SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition

CVPR 2026

Zero-shot skeleton-based action recognition aims to recognize unseen actions by transferring knowledge from seen categories through semantic descriptions. Most existing methods typically align skeleton features with textual embeddings within a shared latent space. However, the absence of contextual

Cited by 0SourcecodeScholar
2025

Intervening in Black Box: Concept Bottleneck Model for Enhancing Human Neural Network Mutual Understanding

ICCV 2025poster

Recent advances in deep learning have led to increasingly complex models with deeper layers and more parameters, reducing interpretability and making their decisions harder to understand. While many methods explain black-box reasoning, most lack effective interventions or only operate at sample-leve…

2025

Prompt-guided Disentangled Representation for Action Recognition

NeurIPS 2025poster

Action recognition is a fundamental task in video understanding. Existing methods typically extract unified features to process all actions in one video, which makes it challenging to model the interactions between different objects in multi-action scenarios. To alleviate this issue, we explore dise…

Cited by 0SourcecodeScholar
2025

VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding

ICCV 2025poster

3D Gaussian Splatting (3DGS) has become horsepower in high-quality, real-time rendering for novel view synthesis of 3D scenes. However, existing methods focus primarily on geometric and appearance modeling, lacking deeper scene understanding while also incurring high training costs that complicate t…

Cited by 0SourcePDFScholar
2024

DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition

ECCV 2024poster

"Neuromorphic sensors, specifically event cameras, revolutionize visual data acquisition by capturing pixel intensity changes with exceptional dynamic range, minimal latency, and energy efficiency, setting them apart from conventional frame-based cameras. The distinctive capabilities of event camera…

2024

Enhance Sketch Recognition’s Explainability via Semantic Component-Level Parsing

AAAI 2024technical

Free-hand sketches are appealing for humans as a universal tool to depict the visual world. Humans can recognize varied sketches of a category easily by identifying the concurrence and layout of the intrinsic semantic components of the category, since humans draw free-hand sketches based a common co…

2024

Language Model Guided Interpretable Video Action Reasoning

CVPR 2024poster

Although neural networks excel in video action recognition tasks their "black-box" nature makes it challenging to understand the rationale behind their decisions. Recent approaches used inherently interpretable models to analyze video actions in a manner akin to human reasoning. However it has been…

2023

UE4-NeRF:Neural Radiance Field for Real-Time Rendering of Large-Scale Scene

NeurIPS 2023poster

Neural Radiance Fields (NeRF) is a novel implicit 3D reconstruction method that shows immense potential and has been gaining increasing attention. It enables the reconstruction of 3D scenes solely from a set of photographs. However, its real-time rendering capability, especially for interactive real…

2021

Free-Form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud

ICCV 2021poster

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of…

Cited by 100PDFcodeScholar
2020

Efficient Scene Text Detection with Textual Attention Tower

ICASSP 2020accepted

Scene text detection has received attention for years and achieved an impressive performance across various benchmarks. In this work, we propose an efficient and accurate approach to detect multi-oriented text in scene images. The proposed feature fusion mechanism allows us to use a shallower networ…

Cited by 0SourceScholar
2018

Attention in Convolutional LSTM for Gesture Recognition

NeurIPS 2018poster

Convolutional long short-term memory (LSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into the LSTM or the convolutional LSTM (ConvLSTM) networks. Based on the previous gesture recognition architectures which combine the…

2016

Human activity recognition based on weighted limb features

IROS 2016poster

Human activity recognition plays an important role in personal assistive robot, being able to recognize human activity and perform corresponding assistive action is a great challenges for personal assistive robot. Human body is an articulated system of rigid segments that can be divided into five pa…

Cited by 3SourceScholar