← Search

Ruiping Wang

34 accepted papers

2026

RoboPCA: Pose-Centered Affordance Learning from Human Demonstrations for Robot Manipulation

ICRA 2026poster

Understanding spatial affordances---comprising the contact regions of object interaction and the corresponding contact poses---is essential for robots to effectively manipulate objects and accomplish diverse tasks. However, existing spatial affordance prediction methods mainly focus on locating the …

2026

SkillNet: Hierarchical Skill Modeling for Compositional Generalization in Vision-Language Action Models

ICML 2026poster

Transfer across diverse task compositions and unseen behaviors remains a significant challenge for vision-language action (VLA) models. Skills are repeatable and atomic components for various tasks, and similarities shared with different skills provide evidence for transferability across behaviors. …

Cited by 0SourceScholar
2026

Tell as You Want: Customizing Image Narrative with Knowledge and Thoughts

AAAI 2026technical

With the advancement of vision-language models, image captioning has made significant progress, leading to the generation of more accurate and detailed descriptions. Current image captioning primarily focuses on describing the apparent visual characteristics, which are easily observed by most humans

Cited by 0SourcePDFScholar
2026

Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation

ICRA 2026poster

While skill-centric approaches leverage foundation models to enhance generalization in compositional tasks, they often rely on fixed skill libraries, limiting adaptability to new tasks without manual intervention. To address this, we propose Uni-Skill, a Unified Skill-centric framework that supports…

2025

Dynamic Behavior Cloning With Temporal Feature Prediction: Enhancing Robotic Arm Manipulation in Moving Object Tasks

RA-L 2025

In numerous real-world applications, the ability to accurately perceive and respond to dynamic changes in the environment, while also maintaining the flexibility to transfer learned skills across different tasks, is crucial for the effective operation of robotic arms. Behavior cloning is particularl

Cited by 5SourceScholar
2025

EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment

ICLR 2025poster

Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanc…

2025

FreeMask3D: Zero-Shot Point Cloud Instance Segmentation Without 3D Training

RA-L 2025

Point cloud instance segmentation is crucial for 3D scene understanding in robotics. However, existing methods heavily rely on learning-based approaches that require large amounts of annotated 3D data, resulting in high annotation costs. Therefore, developing cost-effective and data-efficient soluti

Cited by 0SourceScholar
2025

OV3D-CG: Open-vocabulary 3D Instance Segmentation with Contextual Guidance

ICCV 2025poster

Open-vocabulary 3D instance segmentation (OV-3DIS), which aims to segment and classify objects beyond predefined categories, is a critical capability for embodied AI applications. Existing methods rely on pre-trained 2D foundation models, focusing on instance-level features while overlooking context…

2025

R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action Planner

CVPR 2025poster

This paper explores using large language models (LLMs) as low-level action planners for embodied tasks. While LLMs excel as the robot's "brain" for high-level planning, they face challenges in directly controlling the "body" by generating precise low-level actions. This limitation arises from LLMs'…

2025

SocialMOIF: Multi-Order Intention Fusion for Pedestrian Trajectory Prediction

CVPR 2025poster

The analysis and prediction of agent trajectories are crucial for decision-making processes in intelligent systems, with precise short-term trajectory forecasting being highly significant across a range of applications. Agents and their social interactions have been quantified and modeled by researc…

2024

Autonomous Interactive Correction MLLM for Robust Robotic Manipulation

CoRL 2024poster

The ability to reflect on and correct failures is crucial for robotic systems to interact stably with real-life objects. Observing the generalization and reasoning capabilities of Multimodal Large Language Models (MLLMs), previous approaches have aimed to utilize these models to enhance robotic syst…

Cited by 4SourceScholar
2024

GPSFormer: A Global Perception and Local Structure Fitting-based Transformer for Point Cloud Understanding

ECCV 2024poster

"Despite the significant advancements in pre-training methods for point cloud understanding, directly capturing intricate shape information from irregular point clouds without reliance on external data remains a formidable challenge. To address this problem, we propose GPSFormer, an innovative Globa…

2024

Point2Real: Bridging the Gap between Point Cloud and Realistic Image for Open-World 3D Recognition

AAAI 2024technical

Recognition in open-world scenarios is an important and challenging field, where Vision-Language Pre-training paradigms have greatly impacted the 2D domain. This inspires a growing interest in introducing 2D pre-trained models, such as CLIP, into the 3D domain to enhance the ability of point cloud u…

2023

Glance and Focus: Memory Prompting for Multi-Event Video Question Answering

NeurIPS 2023poster

Video Question Answering (VideoQA) has emerged as a vital tool to evaluate agents’ ability to understand human daily behaviors. Despite the recent success of large vision language models in many multi-modal tasks, complex situation reasoning over videos involving multiple human-object interaction ev…

2022

Implicit-Part Based Context Aggregation for Point Cloud Instance Segmentation

IROS 2022poster

Context information is important for instance segmentation on point clouds. Existing methods either only use local surroundings by stacking multiple convolution layers or use non-local methods to model long-range interactions. However, they usually directly operate on points which is an unstructured…

Cited by 0SourcecodeScholar
2021

Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic Environments

ICCV 2021poster

Visual understanding goes well beyond the study of images or videos on the web. To achieve complex tasks in volatile situations, the human can deeply understand the environment, quickly perceive events happening around, and continuously track objects' state changes, which are still challenging for c…

Cited by 37PDFScholar
2021

FAIEr: Fidelity and Adequacy Ensured Image Caption Evaluation

CVPR 2021poster

Image caption evaluation is a crucial task, which involves the semantic perception and matching of image and text. Good evaluation metrics aim to be fair, comprehensive, and consistent with human judge intentions. When humans evaluate a caption, they usually consider multiple aspects, such as whethe…

Cited by 41PDFcodeScholar
2021

Holistic Pose Graph: Modeling Geometric Structure Among Objects in a Scene Using Graph Inference for 3D Object Prediction

ICCV 2021poster

Due to the missing depth cues, it is essentially ambiguous to detect 3D objects from a single RGB image. Existing methods predict the 3D pose for each object independently or merely by combining local relationships within limited surroundings, but rarely explore the inherent geometric relationships…

Cited by 5PDFScholar
2020

Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene Text

CVPR 2020poster

Answering questions that require reading texts in an image is challenging for current models. One key difficulty of this task is that rare, polysemous, and ambiguous words frequently appear in images, e.g., names of places, products, and sports teams. To overcome this difficulty, only resorting to p…

Cited by 150PDFScholar
2020

Sketching Image Gist: Human-Mimetic Hierarchical Scene Graph Generation

ECCV 2020poster

Scene graph aims to faithfully reveal humans' perception of image content. When humans analyze a scene, they usually prefer to describe image gist first, namely major objects and key relations in a scene graph. This humans' inherent perceptive habit implies that there exists a hierarchical structure…

2019

Exploring Context and Visual Pattern of Relationship for Scene Graph Generation

CVPR 2019poster

Relationship is the core of scene graph, but its prediction is far from satisfying because of its complex visual diversity. To alleviate this problem, we treat relationship as an abstract object, exploring not only significative visual pattern but contextual information for it, which are two key asp…

Cited by 115PDFScholar
2019

Transferable Contrastive Network for Generalized Zero-Shot Learning

ICCV 2019poster

Zero-shot learning (ZSL) is a challenging problem that aims to recognize the target categories without seen data, where semantic information is leveraged to transfer knowledge from some source classes. Although ZSL has made great progress in recent years, most existing approaches are easy to overfit…

Cited by 241PDFScholar
2018

Learning Class Prototypes via Structure Alignment for Zero-Shot Recognition

ECCV 2018poster

Zero-shot learning (ZSL) aims to recognize objects of novel classes without any training samples of specific classes, which is achieved by exploiting the semantic information and auxiliary datasets. Recently most ZSL approaches focus on learning visual-semantic embeddings to transfer knowledge from…

Cited by 162SourcePDFScholar
2018

Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level Relationships

CVPR 2018poster

Context is important for accurate visual recognition. In this work we propose an object detection algorithm that not only considers object visual appearance, but also makes use of two kinds of context including scene contextual information and object relationships within a single image. Therefore, o…

Cited by 298SourcePDFScholar
2017

Discriminative Covariance Oriented Representation Learning for Face Recognition With Image Sets

CVPR 2017poster

For face recognition with image sets, while most existing works mainly focus on building robust set models with hand-crafted feature, it remains a research gap to learn better image representations which can closely match the subsequent image set modeling and classification. Taking sample covariance…

Cited by 49PDFScholar
2017

Learning Discriminative Latent Attributes for Zero-Shot Classification

ICCV 2017poster

Zero-shot learning (ZSL) aims to transfer knowledge from observed classes to the unseen classes, based on the assumption that both the seen and unseen classes share a common semantic space, among which attributes enjoy a great popularity. However, few works study whether the human-designed semantic…

Cited by 128PDFScholar
2017

Learning Multifunctional Binary Codes for Both Category and Attribute Oriented Retrieval Tasks

CVPR 2017poster

In this paper we propose a unified framework to address multiple realistic image retrieval tasks concerning both category and attributes. Considering the scale of modern datasets, hashing is favorable for its low complexity. However, most existing hashing methods are designed to preserve one single…

Cited by 41PDFScholar
2015

Discriminant Analysis on Riemannian Manifold of Gaussian Distributions for Face Recognition With Image Sets

CVPR 2015poster

This paper presents a method named Discriminant Analysis on Riemannian manifold of Gaussian distributions (DARG) to solve the problem of face recognition with image sets. Our goal is to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end,…

Cited by 186SourcePDFScholar
2015

Face Video Retrieval With Image Query via Hashing Across Euclidean Space and Riemannian Manifold

CVPR 2015poster

Retrieving videos of a specific person given his/her face image as query becomes more and more appealing for applications like smart movie fast-forwards and suspect searching. It also forms an interesting but challenging computer vision task, as the visual data to match, i.e., still image and video…

Cited by 82SourcePDFScholar
2015

Log-Euclidean Metric Learning on Symmetric Positive Definite Manifold with Application to Image Set Classification

ICML 2015poster

The manifold of Symmetric Positive Definite (SPD) matrices has been successfully used for data representation in image set classification. By endowing the SPD manifold with Log-Euclidean Metric, existing methods typically work on vector-forms of SPD matrix logarithms. This however not only inevitabl…

Cited by 313SourcePDFScholar
2015

Projection Metric Learning on Grassmann Manifold With Application to Video Based Face Recognition

CVPR 2015poster

In video based face recognition, great success has been made by representing videos as linear subspaces, which typically lie in a special type of non-Euclidean space known as Grassmann manifold. To leverage the kernel-based methods developed for Euclidean space, several recent methods have been prop…

Cited by 297SourcePDFScholar
2015

Two Birds, One Stone: Jointly Learning Binary Code for Large-Scale Face Image Retrieval and Attributes Prediction

ICCV 2015poster

We address the challenging large-scale content-based face image retrieval problem, intended as searching images based on the presence of specific subject, given one face image of him/her. To this end, one natural demand is a supervised binary code learning method. While the learned codes might be di…

Cited by 65PDFScholar