← Search

Jian-Fang Hu

25 accepted papers

2026

Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible

Cited by 0SourcecodeScholar
2026

MotionHiFlow: Text-to-Motion via Hierarchical Flow Matching

CVPR 2026

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which

Cited by 2SourcecodeScholar
2026

Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation

CVPR 2026

Referring Video Object Segmentation (RVOS) aims to segment objects in videos based on textual queries. Current methods mainly rely on large-scale supervised fine-tuning (SFT) of Multi-modal Large Language Models (MLLMs). However, this paradigm suffers from heavy data dependence and limited scalabili

Cited by 0SourcecodeScholar
2026

Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search

ICML 2026poster

Segmentation based on language has been a popular topic in computer vision. While recent advances in multimodal large language models (MLLMs) have endowed segmentation systems with reasoning capabilities, these efforts remain confined by the frozen internal knowledge of MLLMs, which limits their pot…

Cited by 0SourceScholar
2026

TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding

AAAI 2026technical

Spatio-Temporal Video Grounding (STVG) aims to localize a spatio-temporal tube that corresponds to a given language query in an untrimmed video. This is a challenging task since it involves complex vision-language understanding and spatiotemporal reasoning. Recent works have explored weakly-superv

Cited by 0SourcePDFScholar
2025

CLIP-RestoreX: Restore Image Structure and Perception in Exposure Correction

AAAI 2025technical

Exposure correction aims to adjust the exposure of an under- and over-exposed image to enhance its overall visual quality. The core challenge of this task lies in that it requires to faithfully restore both the structure and perception information. In this work, we present a novel exposure correctio…

2025

Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks

CVPR 2025poster

Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering errors or rely on static prototypes to represent normal actions. However, these approaches typically overlook the common sce…

2025

Panorama Generation From NFoV Image Done Right

CVPR 2025highlight

Generating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is not suitab…

2025

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

ICCV 2025poster

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existi…

Cited by 0SourcePDFScholar
2025

SAUGE: Taming SAM for Uncertainty-Aligned Multi-Granularity Edge Detection

AAAI 2025technical

Edge labels are typically at various granularity levels owing to the varying preferences of annotators, thus handling the subjectivity of per-pixel labels has been a focal point for edge detection. Previous methods often employ a simple voting strategy to diminish such label uncertainty or impose a…

2025

Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic

CVPR 2025poster

Action-driven stochastic human motion prediction aims to generate future motion sequences of a pre-defined target action based on given past observed sequences performing non-target actions. This task primarily presents two challenges. Firstly, generating smooth transition motions is hard due to the…

2025

ViSpeak: Visual Instruction Feedback in Streaming Videos

ICCV 2025poster

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming…

2024

AdvAD: Exploring Non-Parametric Diffusion for Imperceptible Adversarial Attacks

NeurIPS 2024poster

Imperceptible adversarial attacks aim to fool DNNs by adding imperceptible perturbation to the input data. Previous methods typically improve the imperceptibility of attacks by integrating common attack paradigms with specifically designed perception-based losses or the capabilities of generative mo…

2024

Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels

CVPR 2024poster

This paper focuses on open-ended video question answering which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task since a question may have multiple answers. However due to annotation costs the labe…

Cited by 2SourcePDFScholar
2024

Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion Model

CVPR 2024poster

Universal image restoration is a practical and potential computer vision task for real-world applications. The main challenge of this task is handling the different degradation distributions at once. Existing methods mainly utilize task-specific conditions (e.g. prompt) to guide the model to learn d…

2024

Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding

CVPR 2024poster

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However existing VPG approaches are heavily reliant on a considerable number of temporal labels that are…

Cited by 4SourcePDFScholar
2023

Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding

CVPR 2023poster

Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the la…

2023

Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding

CVPR 2023poster

Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex sem…

Cited by 24SourcePDFScholar
2023

Temporal Continual Learning with Prior Compensation for Human Motion Prediction

NeurIPS 2023poster

Human Motion Prediction (HMP) aims to predict future poses at different moments according to past motion sequences. Previous approaches have treated the prediction of various moments equally, resulting in two main limitations: the learning of short-term predictions is hindered by the focus on long-t…

2022

You Never Stop Dancing: Non-freezing Dance Generation via Bank-constrained Manifold Projection

NeurIPS 2022accept

One of the most overlooked challenges in dance generation is that the auto-regressive frameworks are prone to freezing motions due to noise accumulation. In this paper, we present two modules that can be plugged into the existing models to enable them to generate non-freezing and high fidelity dance…

Cited by 27SourcePDFScholar
2021

Action-guided 3D Human Motion Prediction

NeurIPS 2021poster

The ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human mo…

Cited by 10SourcePDFScholar
2021

Predictive Feature Learning for Future Segmentation Prediction

ICCV 2021poster

Future segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation models. However, these segmentation features are learned to be local discriminative…

Cited by 20PDFScholar
2019

Progressive Teacher-Student Learning for Early Action Prediction

CVPR 2019poster

The goal of early action prediction is to recognize actions from partially observed videos with incomplete action executions, which is quite different from action recognition. Predicting early actions is very challenging since the partially observed videos do not contain enough action information fo…

Cited by 166PDFcodeScholar
2018

Deep Bilinear Learning for RGB-D Action Recognition

ECCV 2018poster

In this paper, we focus on exploring modality-temporal mutual information for RGB-D action recognition. In order to learn time-varying information and multi-modal features jointly, we propose a novel deep bilinear learning framework. In the framework, we propose bilinear blocks that consist of two l…

Cited by 116SourcePDFScholar
2015

Jointly Learning Heterogeneous Features for RGB-D Activity Recognition

CVPR 2015poster

In this paper, we focus on heterogeneous feature learning for RGB-D activity recognition. Considering that features from different channels could share some similar hidden structures, we propose a joint learning model to simultaneously explore the shared and feature-specific components as an instanc…

Cited by 687SourcePDFScholar