← Search

peng sun

53 accepted papers

2026

DynaMem: Consistent Long Video Generation via Hierarchical Memory and Motion Priors

ICML 2026poster

Recent text-to-video diffusion models can synthesize visually compelling clips from natural language prompts. However, practical applications increasingly demand long-form videos with evolving narratives and persistent identity. A common solution is autoregressive generation, where the video is prod…

Cited by 0SourceScholar
2026

GSON: A Group-Based Social Navigation Framework with Large Multimodal Model

ICRA 2026poster

With the increasing presence of service robots and autonomous vehicles in human environments, navigation systems need to evolve beyond simple destination reach to incorporate social awareness. This paper introduces GSON, a novel group-based social navigation framework that leverages Large Multimodal…

2026

PACE: Proactive Agent-Level Admission Control for Efficient Agentic Batch Inference

ICML 2026poster

Batch inference for agentic workloads stresses the GPU key–value (KV) cache in a sustained and cumulative manner, often causing severe throughput degradation well before memory capacity is exhausted. We identify this phenomenon as middle-phase thrashing, a previously under-characterized pathology in…

Cited by 0SourceScholar
2026

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

ICLR 2026poster

Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where the model develops an implicit preference for certain modalities, leading to the…

Cited by 0SourceScholar
2026

Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping

ICLR 2026poster

Reinforcement learning (RL) has recently become the core paradigm for aligning and strengthening large language models (LLMs). Yet, applying RL in off-policy settings—where stale data from past policies are used for training—improves sample efficiency, but remains challenging: policy entropy decline…

Cited by 0SourcecodeScholar
2026

TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows

ICLR 2026poster

Recent advances in large multi-modal generative models have demonstrated impressive capabilities in multi-modal generation, including image and video generation. These models are typically built upon multi-step frameworks like diffusion and flow matching, which inherently limits their inference effi…

Cited by 0SourcecodeScholar
2026

UNIVERSAL AND EFFICIENT LOADING BALANCING FOR RL TRAINING OF LARGE MULTIMODAL MODELS

ICLR 2026poster

Reinforcement learning (RL) is crucial for aligning Vision-Language Models (VLMs), but its practical application is hampered by significant system-level bottlenecks. The typical RL pipeline, encompassing data loading, inference-based rollouts, and model updates, suffers from severe inefficiencies wh…

Cited by 0SourceScholar
2026

VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection

AAAI 2026technical

To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervisio

Cited by 0SourcePDFScholar
2025

Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios

ICCV 2025poster

With the rapid advancement of generative models, highly realistic image synthesis has posed new challenges to digital security and media credibility. Although AI-generated image detection methods have partially addressed these concerns, a substantial research gap remains in evaluating their performa…

Cited by 0SourcePDFScholar
2025

DSRC: Learning Density-Insensitive and Semantic-Aware Collaborative Representation Against Corruptions

AAAI 2025technical

As a potential application of Vehicle-to-Everything (V2X) communication, multi-agent collaborative perception has achieved significant success in 3D object detection. While these methods have demonstrated impressive results on standard benchmarks, the robustness of such approaches in the face of com…

2025

GSON: A Group-Based Social Navigation Framework With Large Multimodal Model

RA-L 2025

With the increasing presence of service robots and autonomous vehicles in human environments, navigation systems need to evolve beyond simple destination reach to incorporate social awareness. This paper introduces <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/199

Cited by 10SourceScholar
2025

Multi-agent KTO: Enhancing Strategic Interactions of Large Language Model in Language Game

NeurIPS 2025poster

Achieving Artificial General Intelligence (AGI) requires AI agents that can not only make strategic decisions but also engage in flexible and meaningful communication. Inspired by Wittgenstein's language game theory, we propose that language agents can learn through in-context interaction rather tha…

Cited by 0SourcecodeScholar
2025

Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning

EMNLP 2025

Natural language chain-of-thought (N-CoT) and Program chain-of-thought (P-CoT) have emerged as two primary paradigms for large language models (LLMs) to solve mathematical reasoning problems. Current research typically endeavors to achieve unidirectional enhancement: P-CoT enhanced N-CoT or N-CoT en

2025

Towards Annotation-Free Evaluation: KPAScore for Human Keypoint Detection

ICCV 2025poster

Human keypoint detection is fundamental in computer vision, with applications in pose estimation and action recognition. However, existing evaluation metrics (e.g., OKS, PCP, PDJ) rely on human-annotated ground truth, a labor-intensive process that increases costs, limits scalability. To address thi…

Cited by 0SourcePDFScholar
2025

Uncertainty-Aware Crime Prediction With Spatial Temporal Multivariate Graph Neural Networks

ICASSP 2025accepted

Crime prediction (CP) plays a pivotal role in urban analytics, contributing significantly to personal safety and societal stability. Unlike conventional time series forecasting, CP faces unique difficulties due to the inherent sparsity of crime incidents, particularly within small spatial regions an…

Cited by 0SourceScholar
2025

Unlocking the Power of LSTM for Long Term Time Series Forecasting

AAAI 2025technical

Traditional recurrent neural network architectures, such as long short-term memory neural networks (LSTM), have historically held a prominent role in time series forecasting (TSF) tasks. While the recently introduced sLSTM for Natural Language Processing (NLP) introduces exponential gating and memor…

2025

Unveiling the Pruning Risks on Privacy Vulnerabilities of Deep Neural Networks

ICASSP 2025accepted

Large-scale deep neural networks (DNNs), such as large language models, have gained immense popularity due to their outstanding performance across various tasks. However, their application in resource-constrained scenarios faces significant challenges due to the high computational costs and memory u…

Cited by 0SourceScholar
2024

A Dual Stealthy Backdoor: From Both Spatial and Frequency Perspectives

AAAI 2024technical

Backdoor attacks pose serious security threats to deep neural networks (DNNs). Backdoored models make arbitrarily (targeted) incorrect predictions on inputs containing well-designed triggers, while behaving normally on clean inputs. Prior researches have explored the invisibility of backdoor trigger…

2024

Align before Collaborate: Mitigating Feature Misalignment for Robust Multi-Agent Perception

ECCV 2024oral

"Collaborative perception has received widespread attention recently since it enhances the perception ability of autonomous vehicles via inter-agent information sharing. However, the performance of existing systems is hindered by the unavoidable collaboration noises, which induce feature-level spati…

Cited by 0SourcePDFScholar
2024

Byzantine-robust Decentralized Federated Learning via Dual-domain Clustering and Trust Bootstrapping

CVPR 2024poster

Decentralized federated learning (DFL) facilitates collaborative model training across multiple connected clients without a central coordination server thereby avoiding the single point of failure in traditional centralized federated learning (CFL). However DFL exhibits heightened susceptibility to…

Cited by 7SourcePDFScholar
2024

ERMVP: Communication-Efficient and Collaboration-Robust Multi-Vehicle Perception in Challenging Environments

CVPR 2024poster

Collaborative perception enhances perception performance by enabling autonomous vehicles to exchange complementary information. Despite its potential to revolutionize the mobile industry challenges in various environments such as communication bandwidth limitations localization errors and informatio…

2024

Energy-based Backdoor Defense without Task-Specific Samples and Model Retraining

ICML 2024poster

Backdoor defense is crucial to ensure the safety and robustness of machine learning models when under attack. However, most existing methods specialize in either the detection or removal of backdoors, but seldom both. While few works have addressed both, these methods rely on strong assumptions or e…

Cited by 4SourcePDFScholar
2024

ReFT: Reasoning with Reinforced Fine-Tuning

ACL 2024long

One way to enhance the reasoning capability of Large Language Models (LLMs) is to conduct Supervised Fine-Tuning (SFT) using Chain-of-Thought (CoT) annotations. This approach does not show sufficiently strong generalization ability, however, because the training only relies on the given CoT data. In…

2024

Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning

ICML 2024poster

In this paper, we propose **R**$^3$: Learning **R**easoning through **R**everse Curriculum **R**einforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reas…

2024

Unleashing the Potential of Large Language Models through Spectral Modulation

EMNLP 2024finding

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, garnering significant attention from both academia and industry. However, enhancing the performance of LLMs typically requires scaling up model sizes or fine-tuning with additional datasets, which results…

2023

A Novel Efficient Multi-View Traffic-Related Object Detection Framework

ICASSP 2023accepted

With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Acc…

Cited by 0SourceScholar
2023

Boosting Signal Modulation Few-Shot Learning with Pre-Transformation

ICASSP 2023accepted

The recent flourish of deep learning on various tasks is largely accredited to the rich and high-quality labeled data. Nonetheless, collecting sufficient labeled samples is not very practical for many real applications. Few-shot Learning (FSL) provides a promising solution that allows a model to lea…

Cited by 0SourceScholar
2023

Event-Based Frame Interpolation With Ad-Hoc Deblurring

CVPR 2023poster

The performance of video frame interpolation is inherently correlated with the ability to handle motion in the input scene. Even though previous works recognize the utility of asynchronous event information for this task, they ignore the fact that motion may or may not result in blur in the input vi…

2023

Privacy-Preserving Adversarial Facial Features

CVPR 2023poster

Face recognition service providers protect face privacy by extracting compact and discriminative facial features (representations) from images, and storing the facial features for real-time recognition. However, such features can still be exploited to recover the appearance of the original face by b…

Cited by 22SourcePDFScholar
2023

Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception

ICCV 2023poster

Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this em…

Cited by 68PDFcodeScholar
2023

Towards Generation and Transition of Diverse Gaits for Quadrupedal Robots Based on Trajectory Optimization and Whole-Body Impedance Control

RA-L 2023

Trajectory optimization (TO) combined with whole-body control (WBC) have been a widely accepted approach for dynamic gait control of quadruped robots. However, there are still open issues in this framework, one is the lack of a unified description of intrinsic inter-limb coordination for wide range

Cited by 19SourceScholar
2023

Towards Transferable Targeted Adversarial Examples

CVPR 2023poster

Transferability of adversarial examples is critical for black-box deep learning model attacks. While most existing studies focus on enhancing the transferability of untargeted adversarial attacks, few of them studied how to generate transferable targeted adversarial examples that can mislead models…

2022

A Unified Diversity Measure for Multiagent Reinforcement Learning

NeurIPS 2022accept

Promoting behavioural diversity is of critical importance in multi-agent reinforcement learning, since it helps the agent population maintain robust performance when encountering unfamiliar opponents at test time, or, when the game is highly non-transitive in the strategy space (e.g., Rock-Paper-Sc…

Cited by 16SourcePDFScholar
2022

Event-Based Fusion for Motion Deblurring with Cross-Modal Attention

ECCV 2022poster

"Traditional frame-based cameras inevitably suffer from motion blur due to long exposure times. As a kind of bio-inspired camera, the event camera records the intensity changes in an asynchronous way with high temporal resolution, providing valid image degradation information within the exposure tim…

2022

FedInv: Byzantine-Robust Federated Learning by Inversing Local Model Updates

AAAI 2022technical

Federated learning (FL) is a privacy-preserving distributed machine learning paradigm that enables multiple clients to collaboratively train statistical models without disclosing raw training data. However, the inaccessible local training data and uninspectable local training process make FL suscept…

Cited by 61SourcePDFScholar
2021

Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion

CVPR 2021poster

RGB-D salient object detection (SOD) is usually formulated as a problem of classification or regression over two modalities, i.e., RGB and depth. Hence, effective RGB-D feature modeling and multi-modal feature fusion both play a vital role in RGB-D SOD. In this paper, we propose a depth-sensitive RG…

Cited by 214PDFcodeScholar
2021

Towards Distraction-Robust Active Visual Tracking

ICML 2021spotlight

In active visual tracking, it is notoriously difficult when distracting objects appear, as distractors often mislead the tracker by occluding the target or bringing a confusing appearance. To address this issue, we propose a mixed cooperative-competitive multi-agent game, where a target and multiple…

Cited by 45SourcePDFScholar
2020

Graph-Guided Architecture Search for Real-Time Semantic Segmentation

CVPR 2020poster

Designing a lightweight semantic segmentation network often requires researchers to find a trade-off between performance and speed, which is always empirical due to the limited interpretability of neural networks. In order to release researchers from these tedious mechanical trials, we propose a Gra…

Cited by 120PDFScholar
2019

AD-VAT: An Asymmetric Dueling mechanism for learning Visual Active Tracking

ICLR 2019poster

Visual Active Tracking (VAT) aims at following a target object by autonomously controlling the motion system of a tracker given visual observations. Previous work has shown that the tracker can be trained in a simulator via reinforcement learning and deployed in real-world scenarios. However, during…

2019

Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI

ICML 2019oral

We consider the problem of multi-agent reinforcement learning (MARL) in video game AI, where the agents are located in a spatial grid-world environment and the number of agents varies both within and across episodes. The challenge is to flexibly control an arbitrary number of agents while achieving…

Cited by 72SourcePDFScholar
2018

An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method

ICML 2018oral

We propose a novel algorithmic framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-gradient (VMOR-HPE) method with a global convergence guarantee for the maximal monotone operator inclusion problem. Its iteration complexities and local linear convergence rate are provided, which theoreti…

Cited by 4SourcePDFScholar
2018

End-to-end Active Object Tracking via Reinforcement Learning

ICML 2018oral

We study active object tracking, where a tracker takes as input the visual observation (i.e. frame sequence) and produces the camera control signal (e.g., move forward, turn left, etc). Conventional methods tackle the tracking and the camera control separately, which is challenging to tune jointly.…

Cited by 113SourcePDFScholar
2018

Exponentially Weighted Imitation Learning for Batched Historical Data

NeurIPS 2018poster

We consider deep policy learning with only batched historical trajectories. The main challenge of this problem is that the learner no longer has a simulator or ``environment oracle'' as in most reinforcement learning settings. To solve this problem, we propose a monotonic advantage reweighted imitat…

2018

Tagging Like Humans: Diverse and Distinct Image Annotation

CVPR 2018poster

In this work we propose a new automatic image annotation model, dubbed diverse and distinct image annotation (D2IA). The generative model D2IA is inspired by the ensemble of human annotations, which create semantically relevant, yet distinct and diverse tags. In D2IA, we generate a relevant and dist…