← Search

Weihong Ren

14 accepted papers

2026

HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition

AAAI 2026technical

Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different hea

Cited by 0SourcePDFScholar
2026

LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation

CVPR 2026

Skeleton-based Temporal Action Segmentation (STAS) aims to densely parse untrimmed skeletal sequences into frame-level action categories. However, existing methods, while proficient at capturing spatio-temporal kinematics, neglect the underlying physical dynamics that govern human motion. This overs

Cited by 0SourcecodeScholar
2026

Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation

CVPR 2026

Skeleton-based Temporal Action Segmentation (STAS) seeks to densely segment and classify diverse actions within long, untrimmed skeletal motion sequences. However, existing STAS methodologies face challenges of limited inter-class discriminability and blurred segmentation boundaries, primarily due t

Cited by 0SourcecodeScholar
2025

A New Federated Learning Framework Against Gradient Inversion Attacks

AAAI 2025technical

Federated Learning (FL) aims to protect data privacy by enabling clients to collectively train machine learning models without sharing their raw data. However, recent studies demonstrate that information exchanged during FL is subject to Gradient Inversion Attacks (GIA) and, consequently, a variety…

2025

InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction Detection

NeurIPS 2025spotlight

Recently, Large Foundation Models (LFMs), e.g., CLIP and GPT, have significantly advanced the Human-Object Interaction (HOI) detection, due to their superior generalization and transferability. Prior HOI detectors typically employ single- or multi-modal prompts to generate discriminative representat…

Cited by 0SourceScholar
2024

Discovering Syntactic Interaction Clues for Human-Object Interaction Detection

CVPR 2024poster

Recently Vision-Language Model (VLM) has greatly advanced the Human-Object Interaction (HOI) detection. The existing VLM-based HOI detectors typically adopt a hand-crafted template (e.g. a photo of a person [action] a/an [object]) to acquire text knowledge through the VLM text encoder. However such…

Cited by 5SourcePDFScholar
2024

Exploiting Multi-Modal Synergies for Enhancing 3D Multi-Object Tracking

RA-L 2024

3D Multi-Object Tracking (MOT) aims to establish and maintain consistent object trajectories in continuously dynamic environments. At present, the tracking-by-detection has emerged as a dominant paradigm for 3D MOT, due to its simplicity and efficiency. However, this paradigm depends heavily on the

Cited by 5SourceScholar
2024

Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection

AAAI 2024technical

Human-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplet…

Cited by 5SourcePDFScholar
2024

GroupTrack: Multi-Object Tracking by Using Group Motion Patterns

IROS 2024poster

The main challenge of Multi-Object Tracking (MOT) lies in maintaining a distinctive identity for each target in dense crowds or occluded scenarios. Although the existing methods have achieved significantly progress by using robust object detectors or complex association strategies, they cannot effec…

Cited by 0SourceScholar
2024

Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation

ECCV 2024poster

"Skeleton-based Temporal Action Segmentation (STAS) aims to densely segment and classify human actions in long, untrimmed skeletal motion sequences. Existing STAS methods primarily model spatial dependencies among joints and temporal relationships among frames to generate frame-level one-hot classif…

2024

MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition

IROS 2024poster

In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Lang…

Cited by 1SourceScholar
2023

WSCFER: Improving Facial Expression Representations by Weak Supervised Contrastive Learning

IROS 2023poster

The major challenge of Facial Expression Recog-nition (FER) is to learn class discriminative representations, and the existing works mainly address it by designing various classification networks from class level. However, learning representations at class level is limited due to the inconspicuous c…

Cited by 3SourceScholar
2018

Fusing Crowd Density Maps and Visual Object Trackers for People Tracking in Crowd Scenes

CVPR 2018poster

While people tracking has been greatly improved over the recent years, crowd scenes remain particularly challenging for people tracking due to heavy occlusions, high crowd density, and significant appearance variation. To address these challenges, we first design a Sparse Kernelized Correlation Filt…

Cited by 30SourcePDFScholar
2017

Video Desnowing and Deraining Based on Matrix Decomposition

CVPR 2017poster

The existing snow/rain removal methods often fail for heavy snow/rain and dynamic scene. One reason for the failure is due to the assumption that all the snowflakes/rain streaks are sparse in snow/rain scenes. The other is that the existing methods often can not differentiate moving objects and snow…

Cited by 199PDFScholar