← Search

Shengjin Wang

64 accepted papers

2026

HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene Interaction

CVPR 2026

Unified interpreting and synthesizing human behaviors within 3D environments is vital for advancing spatial intelligence and humanoid robotics. Despite recent advancements (e.g., HSI-GPT), two fundamental capabilities expected of a unified model--understanding and generation--still lag behind specia

Cited by 0SourceScholar
2026

LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction Detection

ICLR 2026poster

Human-Object Interaction (HOI) detection with vision-language models (VLMs) has progressed rapidly, yet a trade-off persists between specialization and generalization. Two major challenges remain: (1) the sparsity of supervision, which hampers effective transfer of foundation models to HOI tasks, a…

Cited by 0SourceScholar
2026

Preserving Topological and Geometric Embeddings for Point Cloud Recovery

AAAI 2026technical

Recovering point clouds involves the sequential process of sampling and restoration, yet existing methods struggle to effectively leverage both topological and geometric attributes. To address this, we propose an end-to-end architecture named TopGeoFormer, which maintains these critical properties t

Cited by 0SourcePDFScholar
2026

TCoT: Trajectory Chain-of-Thoughts for Robotic Manipulation with Failure Recovery in Vision-Language-Action Model

AAAI 2026technical

Recent advances in vision-language-action (VLA) models have demonstrated impressive generalization for robotic manipulation. However, these models often operate by directly mapping visual and linguistic inputs to subsequent actions, lacking intermediate task planning, along with failure detection an

Cited by 0SourcePDFScholar
2026

TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for general robotic manipulation. However, existing approaches typically omit intermediate reasoning steps and directly regress actions, limiting reasoning interpretability and performance in long-horizon or compositional tasks.

Cited by 0SourceScholar
2026

VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) often over-rely on linguistic priors even when images are provided, leading to object hallucinations. We revisit object-wise hallucination from the perspective of how visual evidence shapes the model's uncertainty. For each input, we measure decision uncertainty with an

Cited by 0SourceScholar
2025

Attention Augmented Structure-centric Bias Mitigation with Feature Disentanglement

ICASSP 2025accepted

Image classification models often rely on superficial visual features, such as textures or colors, leading to undesired bias. This can compromise the robustness and reliability of deep models, particularly their performance on out-of-distribution (o.o.d.) datasets. Existing approaches, focusing on d…

Cited by 0SourceScholar
2025

BookBot: A Robotic Manipulation Benchmark for Voice-Driven Book Recognition and Grasping in Cluttered Environments

IROS 2025

Books, as enduring repositories of cultural heritage as well as knowledge, play a fundamental role in human development. Although advances in embodied AI and robotics revolutionize automation in domains, e.g., manufacturing and logistics, robotic book manipulation remains an underexplored frontier.

Cited by 0SourcecodeScholar
2025

Dynamic Object Queries for Transformer-based Incremental Object Detection

ICASSP 2025accepted

Incremental object detection (IOD) aims to sequentially learn new classes, while maintaining the capability to locate and identify old ones. Prior methodologies mainly tackle catastrophic forgetting through knowledge distillation and exemplar replay, ignoring the conflict between limited model capac…

Cited by 0SourceScholar
2025

Find Details in Long Videos: Tower-of-Thoughts and Self-Retrieval Augmented Generation for Video Understanding

ICASSP 2025accepted

The Large Vision-Language Model (LVLM) has achieved impressive performance in the field of visual-language understanding. However, its ability to understand longer videos is still limited due to the length and information diversity of multi-modal videos. Moreover, accurately matching detailed conten…

Cited by 0SourceScholar
2025

HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction

CVPR 2025highlight

While flourishing developments have been witnessed in text-to-motion generation, synthesizing physically realistic, controllable, language-conditioned Human Scene Interactions (HSI) remains a relatively underexplored landscape. Current HSI methods naively rely on conditional Variational AutoEncoder…

Cited by 0SourcePDFScholar
2025

LIBA: Language Instructed Multi-granularity Bridge Assistant for 3D Visual Grounding

AAAI 2025technical

3D Vision Grounding (3D-VG) seeks to unravel referential language and identify targets in 3D physical world. Prevailing methods align with the 2D-VG's pipeline to pinpoint the referred object in a categorical multi-modal reasoning manner. However, the geometric complexities of 3D scenes and the nuan…

Cited by 0SourcePDFScholar
2024

Alice Benchmarks: Connecting Real World Re-Identification with the Synthetic

ICLR 2024poster

For object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap betwee…

Cited by 0SourcePDFScholar
2024

Exploring Pose-Aware Human-Object Interaction via Hybrid Learning

CVPR 2024poster

Human-Object Interaction (HOI) detection plays a crucial role in visual scene comprehension. In recent advancements two-stage detectors have taken a prominent position. However they are encumbered by two primary challenges. First the misalignment between feature representation and relation reasoning…

Cited by 4SourcePDFScholar
2024

FED-SDS: Adaptive Structured Dynamic Sparsity for Federated Learning Under Heterogeneous Clients

ICASSP 2024accepted

Federated Learning (FL) is a widely utilized distributed learning methodology that facilitates real-time continuous learning while preserving client privacy. In most FL implementations, it is assumed that all edge clients possess sufficient computational capabilities to participate in the training o…

Cited by 0SourceScholar
2024

G^3-LQ: Marrying Hyperbolic Alignment with Explicit Semantic-Geometric Modeling for 3D Visual Grounding

CVPR 2024poster

Grounding referred objects in 3D scenes is a burgeoning vision-language task pivotal for propelling Embodied AI as it endeavors to connect the 3D physical world with free-form descriptions. Compared to the 2D counterparts challenges posed by the variability of 3D visual grounding remain relatively u…

Cited by 8SourcePDFScholar
2024

Learning Generalizable Visual Representations via Self-Supervised Information Bottleneck

ICASSP 2024accepted

Numerous approaches have recently emerged in the realm of self-supervised visual representation learning. While these methods have demonstrated empirical success, a theoretical foundation that understands and unifies these diverse techniques remains to be established. In this work, we draw inspirati…

Cited by 0SourceScholar
2024

OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation

ECCV 2024poster

"In the current state of 3D object detection research, the severe scarcity of annotated 3D data, substantial disparities across different data modalities, and the absence of a unified architecture, have impeded the progress towards the goal of universality. In this paper, we propose OV-Uni3DETR, a u…

2024

One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection

NeurIPS 2024poster

The current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however,…

Cited by 2SourcePDFScholar
2024

RCIF: Towards Robust Distributed DNN Collaborative Inference Under Highly Lossy Networks

ICASSP 2024accepted

Collaborative Inference is a prospective paradigm for accelerating Deep Neural Network (DNN) inference by harnessing the computational resources of multiple devices. However, in highly lossy network environments, such as those encountered in wireless communication systems, the transmission loss of i…

Cited by 0SourceScholar
2023

Detecting Everything in the Open World: Towards Universal Object Detection

CVPR 2023poster

In this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We…

2023

Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-Identification

ICCV 2023oral

This paper aims to learn a domain-generalizable (DG) person re-identification (ReID) representation from large-scale videos without any annotation. Prior DG ReID methods employ limited labeled data for training due to the high cost of annotation, which restricts further advances. To overcome the bar…

Cited by 28PDFcodeScholar
2023

Uni3DETR: Unified 3D Detection Transformer

NeurIPS 2023poster

Existing point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, ther…

2023

VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes

IROS 2023poster

Robotic grasping faces new challenges in human-robot-interaction scenarios. We consider the task that the robot grasps a target object designated by human's language directives. The robot not only needs to locate a target based on vision-and-language information, but also needs to predict the reason…

Cited by 23SourcecodeScholar
2022

CRPN: Distinguish Novel Categories Via Class-Relevant Region Proposal Network for Few-Shot Object Detection

ICASSP 2022accepted

Few-shot object detection (FSOD) has attracted more attention in computer vision, where only very few training examples are presented during model learning process. A commonly-overlooked issue in FSOD is that novel classes are usually classified as background clutters in the pre-training process. An…

Cited by 0SourceScholar
2022

Delving into Probabilistic Uncertainty for Unsupervised Domain Adaptive Person Re-identification

AAAI 2022technical

Clustering-based unsupervised domain adaptive (UDA) person re-identification (ReID) reduces exhaustive annotations. However, owing to unsatisfactory feature embedding and imperfect clustering, pseudo labels for target domain data inherently contain an unknown proportion of wrong ones, which would mi…

2022

GraphCSPN: Geometry-Aware Depth Completion via Dynamic GCNs

ECCV 2022poster

"Image guided depth completion aims to recover per-pixel dense depth maps from sparse depth measurements with the help of aligned color images, which has a wide range of applications from robotics to autonomous driving. However, the 3D nature of sparse-to-dense depth completion has not been fully ex…

2022

Hybrid Physical Metric For 6-DoF Grasp Pose Detection

ICRA 2022poster

6-DoF grasp pose detection of multi-grasp and multi-object is a challenge task in the field of intelligent robot. To imitate human reasoning ability for grasping objects, data driven methods are widely studied. With the introduction of large-scale datasets, we discover that a single physical metric…

Cited by 22SourcecodeScholar
2022

Noisy Boundaries: Lemon or Lemonade for Semi-Supervised Instance Segmentation?

CVPR 2022poster

Current instance segmentation methods rely heavily on pixel-level annotated images. The huge cost to obtain such fully-annotated images restricts the dataset scale and limits the performance. In this paper, we formally address semi-supervised instance segmentation, where unlabeled images are employe…

Cited by 40PDFcodeScholar
2022

OSKDet: Orientation-Sensitive Keypoint Localization for Rotated Object Detection

CVPR 2022poster

Rotated object detection is a challenging issue in computer vision field. Inadequate rotated representation and the confusion of parametric regression have been the bottleneck for high performance rotated detection. In this paper, we propose an orientation-sensitive keypoint based rotated detector O…

Cited by 24PDFScholar
2022

Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-Identification

ICASSP 2022accepted

Person re-identification (re-ID) aims to match pedestrian images from non-overlapping cameras. It is a challenging task because of the feature misalignment problem caused by occlusion. In this paper, inspired by the coarse-to-fine nature of human perception, we propose a novel Progressive-Granularit…

Cited by 0SourceScholar
2022

Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval

ECCV 2022poster

"Current person image retrieval methods have achieved great improvements in accuracy metrics. However, they rarely describe the reliability of the prediction. In this paper, we propose an Uncertainty-Aware Learning (UAL) method to remedy this issue. UAL aims at providing reliability-aware prediction…

2021

A2-FPN: Attention Aggregation Based Feature Pyramid Network for Instance Segmentation

CVPR 2021poster

Learning pyramidal feature representations is crucial for recognizing object instances at different scales. Feature Pyramid Network (FPN) is the classic architecture to build a feature pyramid with high-level semantics throughout. However, intrinsic defects in feature extraction and fusion inhibit F…

Cited by 122PDFScholar
2021

Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification

NeurIPS 2021poster

Semi-supervised learning aims to leverage a large amount of unlabeled data for performance boosting. Existing works primarily focus on image classification. In this paper, we delve into semi-supervised learning for object detection, where labeled data are more labor-intensive to collect. Current met…

Cited by 33SourcePDFScholar
2021

Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection

CVPR 2021poster

In this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection models. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy lab…

Cited by 101PDFScholar
2021

Disentangled Representation for Age-Invariant Face Recognition: A Mutual Information Minimization Perspective

ICCV 2021poster

General face recognition has seen remarkable progress in recent years. However, large age gap still remains a big challenge due to significant alterations in facial appearance and bone structure. Disentanglement plays a key role in partitioning face representations into identity-dependent and age-de…

Cited by 35PDFScholar
2021

Do Different Tracking Tasks Require Different Appearance Models?

NeurIPS 2021poster

Tracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature…

2021

Multi-Target Domain Adaptation With Collaborative Consistency Learning

CVPR 2021poster

Recently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to the high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly ex…

Cited by 108PDFcodeScholar
2021

Partial Off-Policy Learning: Balance Accuracy and Diversity for Human-Oriented Image Captioning

ICCV 2021poster

Human-oriented image captioning with both high diversity and accuracy is a challenging task in vision+language modeling. The reinforcement learning (RL) based frameworks promote the accuracy of image captioning, yet seriously hurt the diversity. In contrast, other methods based on variational auto-e…

Cited by 9PDFScholar
2021

Towards Discriminative Representation Learning for Unsupervised Person Re-Identification

ICCV 2021poster

In this work, we address the problem of unsupervised domain adaptation for person re-ID where annotations are available for the source domain but not for target. Previous methods typically follow a two-stage optimization pipeline, where the network is first pre-trained on source and then fine-tuned…

Cited by 86PDFScholar
2020

CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions

ECCV 2020poster

This paper proposes a self-supervised learning method for the person re-identification (re-ID) problem, where existing unsupervised methods usually rely on pseudo labels, such as those from video tracklets or clustering. A potential drawback of using pseudo labels is that errors may accumulate and i…

Cited by 116SourcePDFScholar
2020

Video Super-Resolution With Temporal Group Attention

CVPR 2020poster

Video super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is d…

Cited by 220PDFcodeScholar
2020

Video Super-Resolution with Recurrent Structure-Detail Network

ECCV 2020poster

Most video super-resolution methods super-resolve a single reference frame with the help of neighboring frames in a temporal sliding window. They are less efficient compared to the recurrent-based methods. In this work, we propose a novel recurrent video super-resolution method which is both effecti…

2019

Perceive Where to Focus: Learning Visibility-Aware Part-Level Features for Partial Person Re-Identification

CVPR 2019poster

This paper considers a realistic problem in person re-identification (re-ID) task, i.e., partial re-ID. Under partial re-ID scenario, the images may contain a partial observation of a pedestrian. If we directly compare a partial pedestrian image with a holistic one, the extreme spatial misalignment…

Cited by 460PDFcodeScholar
2018

Beyond Part Models: Person Retrieval with Refined Part Pooling (and A Strong Convolutional Baseline)

ECCV 2018poster

Employing part-level features offers fine-grained information for pedestrian image description. A prerequisite of part discovery is that each part should be well located. Instead of using external resources like pose estimator, we consider content consistency within each part for precise part locati…

2018

Fast and Accurate Online Video Object Segmentation via Tracking Parts

CVPR 2018poster

Online video object segmentation is a challenging task as it entails to process the image sequence timely and accurately. To segment a target object through the video, numerous CNN-based methods have been developed by heavily finetuning on the object mask in the first frame, which is time-consuming…

2017

Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-Identification

ICCV 2017poster

In this paper, we tackle the vehicle Re-identification (ReID) problem which is of great importance in urban surveillance and can be used for multiple applications. In our vehicle ReID framework, an orientation invariant feature embedding module and a spatial-temporal regularization module are propos…

Cited by 458PDFScholar
2017

SegFlow: Joint Learning for Video Object Segmentation and Optical Flow

ICCV 2017poster

This paper proposes an end-to-end trainable network, SegFlow, for simultaneously predicting pixel-wise object segmentation and optical flow in videos. The proposed SegFlow has two branches where useful information of object segmentation and optical flow is propagated bidirectionally in a unified fra…

Cited by 431PDFcodeScholar
2016

Bagging regularized common spatial pattern with hybrid motor imagery and myoelectric signal

ICASSP 2016accepted

Common Spatial Pattern(CSP) is a widely used algorithm in BCI application. However, it is sensitive to noise and artifact. In this paper, we propose a bagging regularized common spatial pattern (Bagging RCSP) approach for BCI with hybrid motor imagery and myoelectric signal. We divide the training s…

Cited by 0SourceScholar
2016

Weakly Supervised Object Localization With Progressive Domain Adaptation

CVPR 2016poster

We address the problem of weakly supervised object localization where only image-level annotations are available for training. Many existing approaches tackle this problem through object proposal mining. However, a substantial amount of noise in object proposals causes ambiguities for learning discr…

Cited by 257PDFScholar
2015

Fast Orthogonal Projection Based on Kronecker Product

ICCV 2015poster

We propose a family of structured matrices to speed up orthogonal projections for high-dimensional data commonly seen in computer vision applications. In this, a structured matrix is formed by the Kronecker product of a series of smaller orthogonal matrices. This achieves O(dlogd) computational comp…

Cited by 55PDFScholar
2015

Query-Adaptive Late Fusion for Image Search and Person Re-Identification

CVPR 2015poster

Feature fusion has been proven effective [31, 32] in image search. Typically, it is assumed that the to-be-fused heterogeneous features work well by themselves for the query. However, in a more realistic situation, one does not know in advance whether a feature is effective or not for a given query.…

Cited by 383SourcePDFScholar