← Search

Pengfei Zhang

29 accepted papers

2026

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

ICML 2026poster

REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuri…

Cited by 0SourceScholar
2026

Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal Retrieval

AAAI 2026technical

Cross-modal hashing has emerged as a pivotal solution for efficient retrieval across diverse modalities, such as images and texts, by mapping them into compact binary hash spaces. However, in real-world scenarios, the modalities data is often missing or misaligned. Existing methods are most rely on

Cited by 0SourcePDFScholar
2026

Efficient, Secure, Differentially Private Deep Learning in the Two-Server Model

AAAI 2026technical

Existing solutions on differentially private deep learning (DPDL) either require the assumption of a trusted data server (centralized DPDL) or suffer from poor utility (local DPDL); and hence their adoptions are hampered in real-world scenarios.We present CRYPTDP, a crypto-assisted differentially pr

Cited by 0SourcePDFScholar
2026

High Dimensional Distributed Gradient Descent with Arbitrary Number of Byzantine Attackers

AAAI 2026technical

Adversarial attacks pose a major challenge to distributed learning systems, prompting the development of numerous robust learning methods. However, most existing approaches suffer from the curse of dimensionality, i.e. the error increases with the number of model parameters. In this paper, we make a

Cited by 0SourcePDFScholar
2026

MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative Recommendation

AAAI 2026technical

Generative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language M

Cited by 0SourcePDFScholar
2026

PISA: Privacy-Preserving Split Adaptation with Model IP Protection

ICML 2026poster

Fine-tuning Large Language Models (LLMs) enables data holders to construct proprietary, task-specific models by leveraging external high-performance computing infrastructure. However, existing paradigms typically address data privacy and model intellectual property (IP) in isolation, failing to simu…

Cited by 0SourceScholar
2026

PrivSV: Differentially Private Steering Vector for Large Language Models

AAAI 2026technical

Steering Vector (SV) is a powerful technique for controlling Large Language Models (LLMs) by manipulating their activations without altering model weights. However, when constructed from sensitive data, SV poses significant privacy risks, as it may leak private information. Existing differential pri

Cited by 0SourcePDFScholar
2026

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

ICLR 2026poster

Deep learning-based respiratory auscultation is currently hindered by two fundamental challenges: (i) inherent information loss, as converting signals into spectrograms discards transient acoustic events and clinical context; (ii) limited data availability, exacerbated by severe class imbalance. To…

Cited by 0SourcecodeScholar
2026

Stabilizing Cross-Modal Bidirectional Attribution: Few-Shot Adversarial Prompt Tuning for Robust Vision-Language Models

AAAI 2026technical

Large-scale pre-trained vision-language models (VLMs) like CLIP show exceptional performance and zero-shot generalization. However, their reliability may be severely undermined by a critical vulnerability to subtle adversarial perturbations. Our work reveals a critical cross-modal vulnerability: vis

Cited by 0SourcePDFScholar
2026

Towards Vision-Spatiotemporal Fusion in Traffic Forecasting: A Survey on Cross-Modal Alignment

IJCAI 2026

Traffic forecasting is evolving, with world models emerging as a powerful framework applicable to tasks such as core state, trajectory, event, and demand forecasting. These tasks involve both visual and spatiotemporal data, yet most existing methods treat them separately, hindering a unified underst

Cited by 0Scholar
2025

Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval

NeurIPS 2025poster

The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existi…

Cited by 0SourceScholar
2025

EVICheck: Evidence-Driven Independent Reasoning and Combined Verification Method for Fact-Checking

IJCAI 2025

Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have demonstrated significant potential in automated fact-checking. However, existing methods face limitations in insufficient evidence utilization and lack of explicit verification criteria. Specifically, these approaches aggrega

2025

Jumping Mechanism Assists Takeoff for Large-Sized Flapping-Wing Robots

IROS 2025

Flapping-wing robots exhibit numerous advantages in flight performance, which mimic the natural flight of birds or insects. However, autonomous takeoff remains a significant challenge for large-sized bird-like flapping-wing robots. To address this challenge, we design a jumping mechanism based on a

Cited by 0SourceScholar
2025

KinMo: Kinematic-aware Human Motion Understanding and Generation

ICCV 2025poster

Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as "run", fails to capture essential details like variations in speed, limb positioning, and kine…

2025

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

EMNLP 2025

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has be

2024

EVSMap: An Efficient Volumetric-Semantic Mapping Approach for Embedded Systems

IROS 2024poster

Despite significant progress in perception tasks such as 3D scene mapping and semantic information extraction using SLAM and deep learning, applying these techniques within computationally constrained embedded systems remains a challenge. In this work, we introduce a novel end-to-end framework for e…

Cited by 0SourceScholar
2024

Enabling Few-Shot Learning with PID Control: A Layer Adaptive Optimizer

ICML 2024poster

Model-Agnostic Meta-Learning (MAML) and its variants have shown remarkable performance in scenarios characterized by a scarcity of labeled data during the training phase of machine learning models. Despite these successes, MAMLbased approaches encounter significant challenges when there is a substan…

2024

Fine-Grained Bipartite Concept Factorization for Clustering

CVPR 2024poster

In this paper we propose a novel concept factorization method that seeks factor matrices using a cross-order positive semi-definite neighbor graph which provides comprehensive and complementary neighbor information of the data. The factor matrices are learned with bipartite graph partitioning which…

Cited by 2SourcePDFScholar
2022

Performance Improvement of a High-Speed Swimming Robot for Fish-Like Leaping

RA-L 2022

Many aquatic animals are able to leap out of water effortlessly, however, it is still exceedingly challenging for a swimming robot. Inspired by the fast-swimming mechanism of fish, in this letter, we develop an untethered high-speed swimming robot with the integration of high-frequency oscillation a

Cited by 17SourceScholar
2021

An Open-Source, Fiducial-Based, Underwater Stereo Visual-Inertial Localization Method with Refraction Correction

IROS 2021poster

Underwater visual localization is an essential technique for the autonomous operation of underwater robots. However, the unique underwater image characteristics, including refraction, sparse features, and severe noise, pose an enormous challenge to it. For addressing these issues, this paper propose…

Cited by 10SourceScholar
2020

Navigation Command Matching for Vision-based Autonomous Driving

ICRA 2020poster

Learning an optimal policy for autonomous driving task to confront with complex environment is a long- studied challenge. Imitative reinforcement learning is accepted as a promising approach to learn a robust driving policy through expert demonstrations and interactions with environments. However, t…

Cited by 9SourceScholar
2020

Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

CVPR 2020poster

Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this…

Cited by 635PDFcodeScholar
2019

SR-LSTM: State Refinement for LSTM Towards Pedestrian Trajectory Prediction

CVPR 2019poster

In crowd scenarios, reliable trajectory prediction of pedestrians requires insightful understanding of their social behaviors. These behaviors have been well investigated by plenty of studies, while it is hard to be fully expressed by hand-craft rules. Recent studies based on LSTM networks have show…

Cited by 636PDFScholar
2018

Adding Attentiveness to the Neurons in Recurrent Neural Networks

ECCV 2018poster

Recurrent neural networks (RNNs) are capable of modeling the temporal dynamics of complex sequential information. However, the structures of existing RNN neurons mainly focus on controlling the contributions of current and historical information but do not explore the different importance levels of…

Cited by 105SourcePDFScholar
2017

View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition From Skeleton Data

ICCV 2017poster

Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints du…

Cited by 683PDFcodeScholar