← Search

Peng Zhou

56 accepted papers

2026

Contractive Anchor Resolvent Diffusion for Incomplete Multi-View Clustering

ICML 2026poster

Incomplete Multi-View Clustering (IMVC) is fundamentally challenged by structural degradation induced by missing views, rather than the absence of feature values. Existing graph-based approaches either rely on costly data imputation or adopt first-order linear fusion, which acts as a weak low-pass f…

Cited by 0SourceScholar
2026

Detection-Explanation-Improvement: A Closed-Loop Framework of Enhancing Anomaly Detection with Counterfactual Explanations

IJCAI 2026

Many state‑of‑the‑art anomaly detection models operate as black boxes, limiting interpretability and hindering reliable deployment. While recent advances in explainable artificial intelligence have focused on explaining why individual instances are detected as anomalous, comparatively little attenti

Cited by 0Scholar
2026

FairGB: A Fair Granular-Ball Generation Method for Data Classification

ICML 2026poster

With the widespread application of data-driven classifiers in high-risk domains, group fairness has increasingly become a key research focus. However, most existing methods rely on model constraints or data reweighting, which often suffer from limited interpretability and may distort the original da…

Cited by 0SourceScholar
2026

Iterative Shaping of Multi-Particle Aggregates Based on Action Trees and VLM

ICRA 2026poster

In this paper, we address the problem of manipulating multi-particle aggregates using a bimanual robotic system. Our approach enables the autonomous transport of dispersed particles through a series of shaping and pushing actions using robotically-controlled tools. Achieving this advanced manipulati…

2026

PoinnCARE: Hyperbolic Multi-Modal Learning for Enzyme Classification

ICLR 2026poster

Enzyme Commission (EC) number prediction is vital for elucidating enzyme functions and advancing biotechnology applications. However, current methods struggle to capture the hierarchical relationships among enzymes and often overlook critical structural and active site features. To bridge this gap,…

Cited by 0SourceScholar
2026

Preference-Enhanced Reinforcement Learning for Pluralistic Image Inpainting

ICML 2026poster

Existing image inpainting frameworks rely on strictly supervised training paradigms, often suffering from an over-reliance on ground-truth reconstruction, which leads to conservative outputs with misaligned creativity and limited diversity. To this end, we propose the first framework to explore Grou…

Cited by 0SourceScholar
2026

ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence On Mobile Devices

CVPR 2026

Multimodal large language models (MLLMs) have made significant progress in mobile agent development, yet their capabilities are predominantly confined to a reactive paradigm, where they merely execute explicit user commands. The emerging paradigm of proactive intelligence, where agents autonomously

Cited by 0SourcecodeScholar
2025

A Joint Learning of Force Feedback of Robotic Manipulation and Textual Cues for Granular Materials Classification

RA-L 2025

Granular materials (GMs) are formed by a collection of particles. Even if their visual representation is straightforward, it can be seriously affected in the visually constrained environment. Based on frequency features observed in force signals, this paper proposes a non-visual classifier, <bold xm

Cited by 22SourceScholar
2025

A2ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization

ACL 2025finding

Long context large language models (LLMs) pose significant challenges for efficient serving due to the large memory footprint and high access overhead of KV cache.Retrieval-based KV cache reduction methods can mitigate these challenges, typically by offloading the complete KV cache to CPU and retrie…

2025

AlignBot: Aligning VLM-Powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots

ICRA 2025

This paper presents AlignBot, a novel framework designed to optimize VLM-powered customized task planning for household robots by effectively aligning with user reminders. In domestic settings, aligning task planning with user reminders poses significant challenges due to the limited quantity, diver

Cited by 9SourceScholar
2025

BagIt! An Adaptive Dual-Arm Manipulation of Fabric Bags for Object Bagging

RA-L 2025

Bagging tasks, commonly found in industrial scenarios, are challenging considering deformable bags' complicated and unpredictable nature. This paper presents an automated bagging system from the proposed adaptive Structure-of-Interest (SOI) manipulation strategy for dual robot arms. The system dynam

Cited by 0SourceScholar
2025

Collaborative Similarity Fusion and Consistency Recovery for Incomplete Multi-view Clustering

AAAI 2025technical

As partial samples are often absent in certain views, incomplete multi-view clustering has become a challenging task. To tackle data with missing views, current methods either utilize the data similarity relations to recover missing samples or primarily consider the available information of existing…

Cited by 0SourcePDFScholar
2025

Consensus Graph Filter Learning for Multiple Graph Clustering

ICASSP 2025accepted

Multi-view Clustering (MVC) has gained significant attention for its ability to utilize consistent and complementary information from multiple views. Graph filter-based MVC methods have recently demonstrated promising performance, attracting growing interest. However, existing graph filter-based met…

Cited by 0SourceScholar
2025

Gradient-based Causal Feature Selection

IJCAI 2025

Causal feature selection leverages causal discovery techniques to identify critical features associated with a target variable using observational data. Traditional methodologies primarily rely on constraint-based or score-based techniques, which are fraught with limitations. For example, conditiona

2025

Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile Manipulation

AAAI 2025technical

Enabling humanoid robots to perform long-horizon mobile manipulation planning in real-world environments based on embodied perception and comprehension abilities has been a longstanding challenge. With the recent rise of large language models (LLMs), there has been a notable increase in the developm…

Cited by 1SourcePDFScholar
2025

Iterative Shaping of Multi-Particle Aggregates Based on Action Trees and VLM

RA-L 2025

In this paper, we address the problem of manipulating multi- particle aggregates using a bimanual robotic system. Our approach enables the autonomous transport of dispersed particles through a series of shaping and pushing actions using robotically controlled tools. Achieving this advanced manipulat

Cited by 1SourceScholar
2025

Large Language and Protein Assistant for Protein-Protein Interactions Prediction

ACL 2025long

Predicting the types and affinities of protein-protein interactions (PPIs) is crucial for understanding biological processes and developing novel therapeutic approaches. While encoding proteins themselves is essential, PPI networks can also provide rich prior knowledge for these predictive tasks. Ho…

2025

Learning to Hang Crumpled Garments with Confidence-Guided Grasping and Active Perception

IROS 2025

Accurately recognizing the structural regions of targeted objects is crucial for successful manipulation. In this study, we concentrate on the task of hanging crumpled garments on a rack, a common scenario in household environments. This context presents two primary challenges: (1) perceiving and gr

Cited by 0SourceScholar
2025

Local Causal Discovery Without Causal Sufficiency

AAAI 2025technical

Local causal discovery is crucial for revealing the causal relationships between specific variables from data. Existing local causal discovery algorithms are designed under the assumption of causal sufficiency, which states that there are no latent common causes for two or more of the observed varia…

Cited by 0SourcePDFScholar
2025

M-LLM Based Video Frame Selection for Efficient Video Understanding

CVPR 2025poster

Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context vid…

Cited by 3SourcePDFScholar
2025

Multi-view Clustering via Multi-granularity Ensemble

IJCAI 2025

Multi-view clustering aims to integrate complementary information from multiple views to improve clustering performance. However, existing ensemble-based methods suffer from information loss due to their reliance on single-granularity labels, limiting the discriminative capability of learned represe

Cited by 0SourcePDFScholar
2025

Sharper Error Bounds in Late Fusion Multi-view Clustering with Eigenvalue Proportion Optimization

AAAI 2025technical

Multi-view clustering (MVC) aims to integrate complementary information from multiple views to enhance clustering performance. Late Fusion Multi-View Clustering (LFMVC) has shown promise by synthesizing diverse clustering results into a unified consensus. However, current LFMVC methods struggle with…

2025

Unsupervised Multi-View Outlier Detection via Optimal Graph Filtering

ICASSP 2025accepted

Unsupervised multi-view outlier detection has garnered increasing attention in recent years, yet existing methods face persistent challenges. Many approaches rely predominantly on first-order neighborhood information, overlooking the richer insights offered by higher-order structures, which can degr…

Cited by 0SourceScholar
2024

A + B: A General Generator-Reader Framework for Optimizing LLMs to Unleash Synergy Potential

ACL 2024findings

Retrieval-Augmented Generation (RAG) is an effective solution to supplement necessary knowledge to large language models (LLMs). Targeting its bottleneck of retriever performance, “generate-then-read” pipeline is proposed to replace the retrieval stage with generation from the LLM itself. Although p…

2024

Adaptive Shape Servoing of Elastic Rods Using Parameterized Regression Features and Auto-Tuning Motion Controls

RA-L 2024

The robotic manipulation of deformable linear objects has shown great potential in a wide range of real-world applications. However, it presents many challenges due to the objects' non-linear properties and high-dimensional geometric configuration. In this letter, we propose an efficient shape servo

Cited by 33SourceScholar
2024

Efficient Multi-view Unsupervised Feature Selection with Adaptive Structure Learning and Inference

IJCAI 2024poster

As data with diverse representations become high-dimensional, multi-view unsupervised feature selection has been an important learning paradigm. Generally, existing methods encounter the following challenges: (i) traditional solutions either concatenate different views or introduce extra parameters…

Cited by 11SourcePDFScholar
2024

Efficient Planar Fabric Repositioning: Deformation-Aware RRT* for Non-Prehensile Fabric Manipulation

RA-L 2024

Fabrics present significant challenges to robotic manipulation due to their complex dynamics and infinite degrees of freedom. This letter proposes a non-prehensile approach to aligning a fabric cut piece to a specified target pose, which is a common step for many garment manufacturing tasks. Compare

Cited by 4SourceScholar
2024

FocalDreamer: Text-Driven 3D Editing via Focal-Fusion Assembly

AAAI 2024technical

While text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape w…

Cited by 56SourcePDFScholar
2024

Gated Slot Attention for Efficient Linear-Time Sequence Modeling

NeurIPS 2024poster

Linear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch. This paper introduces Gated…

2024

Higher Order Multiple Graph Filtering for Structured Graph Learning

ICASSP 2024accepted

In the field of machine learning, multi-view clustering aims to reveal hidden clustering patterns across different data perspectives. However, traditional methods often struggle due to their reliance on low-order similarity data. To overcome this, we propose a new approach that integrates the learni…

Cited by 0SourceScholar
2024

Interactive Perception for Deformable Object Manipulation

RA-L 2024

Interactive perception enables robots to manipulate the environment and objects to bring them into states that benefit the perception process. Deformable objects pose challenges to this due to manipulation difficulty and occlusion in vision-based perception. In this work, we address such a problem w

Cited by 12SourceScholar
2024

K-Means Clustering Based on Chebyshev Polynomial Graph Filtering

ICASSP 2024accepted

Clustering, a key unsupervised learning method, is widely used in various fields. While the classic K-means algorithm is popular, it often neglects valuable high-order information in data. To address this, we have proposed an improved K-means algorithm incorporating a Chebyshev polynomial approximat…

Cited by 0SourceScholar
2024

Learning to Rank Patches for Unbiased Image Redundancy Reduction

CVPR 2024poster

Images suffer from heavy spatial redundancy because pixels in neighboring regions are spatially correlated. Existing approaches strive to overcome this limitation by reducing less meaningful image regions. However current leading methods rely on supervisory signals. They may compel models to preserv…

2024

Puff-Net: Efficient Style Transfer with Pure Content and Style Feature Fusion Network

CVPR 2024poster

Style transfer aims to render an image with the artistic features of a style image while maintaining the original structure. Various methods have been put forward for this task but some challenges still exist. For instance it is difficult for CNN-based methods to handle global information and long-r…

2024

Self-Training Based Few-Shot Node Classification by Knowledge Distillation

AAAI 2024technical

Self-training based few-shot node classification (FSNC) methods have shown excellent performance in real applications, but they cannot make the full use of the information in the base set and are easily affected by the quality of pseudo-labels. To address these issues, this paper proposes a new self…

2024

Semantics-Aware Receding Horizon Planner for Object-Centric Active Mapping

RA-L 2024

The escalating demands for real-time scene comprehension in modern industries underscore the growing significance of semantic information in the daily tasks of robots, particularly in areas like autonomous inspection and target searching. This letter introduces a semantics-aware receding horizon pla

Cited by 14SourceScholar
2024

VMCC-NET: Uncovering Challenging Regions in Semi-Supervised Medical Image Segmentation with Voxel Mask Based Cyclic-Consistency Network

ICASSP 2024accepted

Semi-supervised learning has been widely used to train models by utilizing a small amount of labeled data and a large amount of unlabeled data, especially in medical image analysis. However, existing methods still do not fully reveal the complex regions (e.g., small branches or fuzzy edges). These c…

Cited by 0SourceScholar
2023

Resolving Task Confusion in Dynamic Expansion Architectures for Class Incremental Learning

AAAI 2023technical

The dynamic expansion architecture is becoming popular in class incremental learning, mainly due to its advantages in alleviating catastrophic forgetting. However, task confu- sion is not well assessed within this framework, e.g., the discrepancy between classes of different tasks is not well learne…

2023

Totally Dynamic Hypergraph Neural Networks

IJCAI 2023poster

Recent dynamic hypergraph neural networks (DHGNNs) are designed to adaptively optimize the hypergraph structure to avoid the dependence on the initial hypergraph structure, thus capturing more hidden information for representation learning. However, most existing DHGNNs cannot adjust the hyperedge n…

2022

Information Augmentation for Few-shot Node Classification

IJCAI 2022poster

Although meta-learning and metric learning have been widely applied for few-shot node classification (FSNC), some limitations still need to be addressed, such as expensive time costs for the meta-train and difficult of exploring the complex structure inherent the graph data. To address in issues, th…

Cited by 11SourcePDFScholar
2022

Keypoint-Based Planar Bimanual Shaping of Deformable Linear Objects Under Environmental Constraints With Hierarchical Action Framework

RA-L 2022

This letter addresses the problem of contact-based manipulation of deformable linear objects (DLOs) towards desired shapes with a dual-arm robotic system. To alleviate the burden of high-dimensional continuous state-action spaces, we model DLOs as kinematic multibody systems via our proposed keypoin

Cited by 42SourceScholar
2021

LaSeSOM: A Latent and Semantic Representation Framework for Soft Object Manipulation

RA-L 2021

Soft object manipulation has recently gained popularity within the robotics community due to its potential applications in many economically important areas. Although great progress has been recently achieved in these types of tasks, most state-of-the-art methods are case-specific; They can only be

Cited by 43SourceScholar
2021

Path Planning With Automatic Seam Extraction Over Point Cloud Models for Robotic Arc Welding

RA-L 2021

This letter presents a point cloud based robotic system for arc welding. Using hand gesture controls, the system scans partial point cloud views of workpiece and reconstructs them into a complete 3D model by a linear iterative closest point algorithm. Then, a bilateral filter is extended to denoise

Cited by 125SourceScholar
2021

Tri-level Robust Clustering Ensemble with Multiple Graph Learning

AAAI 2021technical

Clustering ensemble generates a consensus clustering result by integrating multiple weak base clustering results. Although it often provides more robust results compared with single clustering methods, it still suffers from the robustness problem if it does not treat the unreliability of base result…

Cited by 42SourcePDFScholar
2018

Learning Rich Features for Image Manipulation Detection

CVPR 2018poster

Image manipulation detection is different from traditional semantic object detection because it pays more attention to tampering artifacts than to image content, which suggests that richer features need to be learned. We propose a two-stream Faster R-CNN network and train it end-to- end to detect th…

Cited by 797SourcePDFScholar
2018

Pose Transferrable Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is an important task in the field of intelligent security. A key challenge is how to capture human pose variations, while existing benchmarks (i.e., Market1501, DukeMTMC-reID, CUHK03, etc.) do NOT provide sufficient pose coverage to train a robust ReID system. To add…

Cited by 456SourcePDFScholar