← Search

Peng Chen

75 accepted papers

2026

Aurora: Towards Universal Generative Multimodal Time Series Forecasting

ICLR 2026poster

Cross-domain generalization is very important in Time Series Forecasting because similar historical information may lead to distinct future trends due to the domain-specific characteristics. Recent works focus on building unimodal time series foundation models and end-to-end multimodal supervised mo…

Cited by 0SourcecodeScholar
2026

Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning

CVPR 2026

Diffusion Models have emerged as a leading class of generative models, yet their iterative sampling process remains computationally expensive. Timestep distillation is a promising technique to accelerate generation, but it often requires extensive training and leads to image quality degradation. Fur

Cited by 0SourceScholar
2026

LD-EnSF: Synergizing Latent Dynamics with Ensemble Score Filters for Fast Data Assimilation with Sparse Observations

ICLR 2026poster

Data assimilation techniques are crucial for accurately tracking complex dynamical systems by integrating observational data with numerical forecasts. Recently, score-based data assimilation methods emerged as powerful tools for high-dimensional and nonlinear data assimilation. However, these method…

Cited by 0SourceScholar
2026

LIF Recurrent Memory Enables Long-Horizon Spiking Computation

ICML 2026poster

Processing long sequence data such as speech requires models to maintain long-term dependencies, which is challenging for recurrent spiking neural networks due to high temporal dynamics in neuron models that leak stored information in their membrane potentials, and due to vanishing gradients during …

Cited by 0SourceScholar
2026

Learning Rollout from Sampling: An R1-Style Tokenized Traffic Simulation Model

RA-L 2026

Learning diverse and high-fidelity traffic simulations from human driving demonstrations is crucial for autonomous driving evaluation. The recent next-token prediction (NTP) paradigm, widely adopted in large language models (LLMs), has been applied to traffic simulation and achieves iterative improv

Cited by 0SourceScholar
2026

Origami-Inspired 3-DoF Robotic Joint Design With Pneumatic Actuators

RA-L 2026

Soft pneumatic joints, as the core components of soft robotic manipulators, often face a trade-off between multiple degrees of freedom (DoF) mobility and load capacity. Inspired by the Yoshimura origami pattern, this paper presents a rigid–soft hybrid pneumatic joint. Through spatially ordered place

Cited by 0SourceScholar
2026

PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering

ICML 2026poster

Time series reasoning demands both the perception of complex dynamics and logical depth. However, existing LLM-based approaches exhibit two limitations: they often treat time series merely as text or images, failing to capture the patterns like trends and seasonalities needed to answer specific ques…

Cited by 0SourceScholar
2026

Performance-Driven Demonstration Selection for In-Context Learning

IJCAI 2026

In-context learning (ICL) enables large language models (LLMs) to adapt to new tasks with considerable performance gains, yet its effectiveness is highly sensitive to the choice of demonstrations. Most existing selection methods rely on heuristic or proxy signals (e.g., similarity, diversity, or unc

Cited by 0Scholar
2026

SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models

CVPR 2026

Diffusion models have demonstrated exceptional generative capabilities but are computationally intensive, posing significant challenges for deployment in resource-constrained or latency-sensitive environments.Quantization offers an effective means to reduce model size and computational cost, with po

Cited by 0SourcecodeScholar
2026

SpecExit: Accelerating Large Reasoning Model via Speculative Exit

ICML 2026poster

Despite their strong performance on reasoning tasks, large reasoning models (LRMs) often suffer from overthinking, producing unnecessarily long outputs and incurring high end-to-end latency, a significant limitation to their real-world deployment. To address overthinking, early-exit mechanisms have …

Cited by 0SourceScholar
2026

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

ICML 2026poster

Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video generation. However, we find that directly applying these techniques to image-to-video (I2V) models often fails to yield …

Cited by 0SourceScholar
2026

Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language Models

ICML 2026poster

Although reinforcement learning (RL) enhances the reasoning capabilities of large language models (LLMs), it is primarily learned from the model's self-generated distribution, limiting its ability to acquire reasoning skills beyond its initial knowledge. To overcome this, we propose a Difficulty-Awa…

Cited by 0SourceScholar
2026

Tequila: Deadzone-free Ternary Quantization for Large Language Models

ICLR 2026poster

Quantization techniques are essential for the deployment of Large Language Models (LLMs) on edge devices. However, prevailing methods often rely on mixed-precision multiplication that lacks efficient hardware support, making it not feasible. Ternary weight quantization addresses this by constraining…

Cited by 0SourcecodeScholar
2026

Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed Interaction

ICLR 2026poster

Time series anomaly detection plays a critical role in many dynamic systems. However, previous approaches have primarily relied on unimodal numerical data, overlooking the importance of complementary information from other modalities. In this paper, we propose a novel multimodal time series anomaly…

Cited by 0SourcecodeScholar
2026

Towards Non-Stationary Time Series Forecasting with Temporal Stabilization and Frequency Differencing

AAAI 2026technical

Time series forecasting is critical for decision making across dynamic domains such as energy, finance, transportation, and cloud computing. However, real-world time series often exhibit non-stationarity, including temporal distribution shifts and spectral variability, which poses significant challe

Cited by 0SourcePDFScholar
2026

Towards Optimal Robustness in Learning-Augmented Paging

ICML 2026spotlight

Learning-augmented paging has been extensively studied in recent years. A key advantage over naive ML-based approaches is \emph{bounded robustness}, which guarantees worst-case performance even when predictions are inaccurate, making these algorithms valuable for real-world systems. Prior work achie…

Cited by 0SourceScholar
2026

Unlocking the Value of Text: Event-Driven Reasoning and Multi-Level Alignment for Time Series Forecasting

ICLR 2026poster

Existing time series forecasting methods primarily rely on the numerical data itself. However, real-world time series exhibit complex patterns associated with multimodal information, making them difficult to predict with numerical data alone. While several multimodal time series forecasting methods…

Cited by 0SourcecodeScholar
2026

ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative Pruning

AAAI 2026technical

Large Vision-Language Models (LVLMs) incur high computational costs due to significant redundancy in their visual tokens. To effectively reduce this cost, researchers have proposed various visual token pruning methods. However, existing methods are generally limited, either losing critical visual in

Cited by 0SourcePDFScholar
2025

A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solution

EMNLP 2025

Low-resource language understanding is challenging, even for large language models (LLMs). An epitome of this problem is the CompRehensive lIterary chineSe readIng comprehenSion (CRISIS), whose difficulties include limited linguistic data, long input, and insight-required questions. Besides the comp

2025

A Gait Phase Detection and Gait Spatio-temporal Features Extraction Method Based on the Inertial Measurement Unit*

IROS 2025

The quantitative evaluation of the improvement of physical function is crucial for patients with impaired motor function, such as stroke, in conducting related rehabilitation training activities. Specially, a practical and easy-to-operate gait feature detection and extraction system for a home is ur

Cited by 0SourceScholar
2025

A Synchronous-Optimized and Safety-Improved Framework for Human-Robot Interaction in Robot-Assisted Knee Arthroplasty

RA-L 2025

Shared control is the most commonly used physical human-robot interaction (pHRI) in robot-assisted orthopedic surgery. However, challenges persist in tasks such as robotassisted knee arthroplasty, particularly in terms of motion delays and inadequate three-dimensional (3D) constraints. In this lette

Cited by 0SourceScholar
2025

Air Quality Prediction with Physics-Guided Dual Neural ODEs in Open Systems

ICLR 2025poster

Air pollution significantly threatens human health and ecosystems, necessitating effective air quality prediction to inform public policy. Traditional approaches are generally categorized into physics-based and data-driven models. Physics-based models usually struggle with high computational demands…

Cited by 3SourcePDFScholar
2025

Backdoor Attack on Vertical Federated Graph Neural Network Learning

IJCAI 2025

Federated Graph Neural Network (FedGNN) integrate federated learning (FL) with graph neural networks (GNNs) to enable privacy-preserving training on distributed graph data. Vertical Federated Graph Neural Network (VFGNN), a key branch of FedGNN, handles scenarios where data features and labels are d

Cited by 0SourcePDFScholar
2025

CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games

ICCV 2025poster

Recent advances in Vision-Language-Action models (VLAs) have expanded the capabilities of embodied intelligence. However, significant challenges remain in real-time decision-making in complex 3D environments, which demand second-level responses, high-resolution perception, and tactical reasoning und…

2025

Context-Aware Multi-Scale Polyp Segmentation Network

ICASSP 2025accepted

Colonoscopy is the gold standard for detecting colorectal lesions and is critical for early screening and prevention of colorectal cancer. However, accurate polyp segmentation remains a challenging task due to the diverse morphology, varying sizes and indistinct boundaries of polyps. To address thes…

Cited by 0SourceScholar
2025

Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-Training

IJCAI 2025

Multimodal protein features play a crucial role in protein function prediction. However, these features encompass a wide range of information, ranging from structural data and sequence features to protein attributes and interaction networks, making it challenging to decipher their complex interconne

2025

GazeGaussian: High-Fidelity Gaze Redirection with 3D Gaussian Splatting

ICCV 2025poster

Gaze estimation encounters generalization challenges when dealing with out-of-distribution data. To address this problem, recent methods use neural radiance fields (NeRF) to generate augmented data. However, existing methods based on NeRF are computationally expensive and lack facial details. 3D Gau…

2025

GeSubNet: Gene Interaction Inference for Disease Subtype Network Generation

ICLR 2025oral

Retrieving gene functional networks from knowledge databases presents a challenge due to the mismatch between disease networks and subtype-specific variations. Current solutions, including statistical and deep learning methods, often fail to effectively integrate gene interaction knowledge from data…

Cited by 0SourcePDFScholar
2025

GraphAvatar: Compact Head Avatars with GNN-Generated 3D Gaussians

AAAI 2025technical

Rendering photorealistic head avatars from arbitrary viewpoints is crucial for various applications like virtual reality. Although previous methods based on Neural Radiance Fields (NeRF) can achieve impressive results, they lack fidelity and efficiency. Recent methods using 3D Gaussian Splatting (3D…

2025

Hierarchical Decision-Making for Autonomous Navigation: Integrating Deep Reinforcement Learning and Fuzzy Logic in Four-Wheel Independent Steering and Driving Systems

IROS 2025

This paper presents a hierarchical decision-making framework for autonomous navigation in four-wheel independent steering and driving (4WISD) systems. The proposed approach integrates deep reinforcement learning (DRL) for high-level navigation with fuzzy logic for low-level control to ensure both ta

Cited by 2SourceScholar
2025

Improving Sidescan Sonar Performance Using Array Upsampling Beamforming Synthetic Aperture

ICASSP 2025accepted

Sidescan sonar technology plays an important role in seabed topography detection and mapping, but the performance of mainstream matched filter technology in long-distance image and towing speed limits its application scenarios. In addition, due to the slow propagation speed of sound waves in water,…

Cited by 0SourceScholar
2025

Industrial-Grade Sensor Simulation via Gaussian Splatting: A Modular Framework for Scalable Editing and Full-Stack Validation

IROS 2025

Sensor simulation is pivotal for scalable validation of autonomous driving systems, yet existing Neural Radiance Fields (NeRF) based methods face applicability and efficiency challenges in industrial workflows. This paper introduces a Gaussian Splatting (GS) based system to address these challenges:

Cited by 3SourceScholar
2025

Latent-EnSF: A Latent Ensemble Score Filter for High-Dimensional Data Assimilation with Sparse Observation Data

ICLR 2025poster

Accurate modeling and prediction of complex physical systems often rely on data assimilation techniques to correct errors inherent in model simulations. Traditional methods like the Ensemble Kalman Filter (EnKF) and its variants as well as the recently developed Ensemble Score Filters (EnSF) face si…

Cited by 4SourcePDFScholar
2025

Learning Cocoercive Conservative Denoisers via Helmholtz Decomposition for Poisson Imaging Inverse Problems

NeurIPS 2025poster

Plug-and-play (PnP) methods with deep denoisers have shown impressive results in imaging problems. They typically require strong convexity or smoothness of the fidelity term and a (residual) non-expansive denoiser for convergence. These assumptions, however, are violated in Poisson inverse problems,…

Cited by 0SourceScholar
2025

LightGTS: A Lightweight General Time Series Forecasting Model

ICML 2025poster

Existing works on general time series forecasting build foundation models with heavy model parameters through large-scale multi-source pretraining. These models achieve superior generalization ability across various datasets at the cost of significant computational burdens and limitations in resourc…

Cited by 0SourcePDFScholar
2025

Robustifying Learning-Augmented Caching Efficiently without Compromising 1-Consistency

NeurIPS 2025poster

The online caching problem aims to minimize cache misses when serving a sequence of requests under a limited cache size. While naive learning-augmented caching algorithms achieve ideal $1$-consistency, they lack robustness guarantees. Existing robustification methods either sacrifice $1$-consistency…

Cited by 0SourceScholar
2025

SHF: Symmetrical Hierarchical Forest with Pretrained Vision Transformer Encoder for High-Resolution Medical Segmentation

NeurIPS 2025spotlight

This paper presents a novel approach to addressing the long-sequence problem in high-resolution medical images for Vision Transformers (ViTs). Using smaller patches as tokens can enhance ViT performance, but quadratically increases computation and memory requirements. Therefore, the common practice…

Cited by 0SourceScholar
2025

Symmetry and Fusion Data Augmentation for Semi-Supervised Medical Segmentation

ICASSP 2025accepted

In semi-supervised medical image segmentation, appropriately merging labeled and unlabeled data before network training instead of using them separately can effectively reduce knowledge loss, mitigate distribution discrepancies and promote efficient knowledge transfer to unlabeled data. However, exi…

Cited by 0SourceScholar
2025

Towards a General Time Series Forecasting Model with Unified Representation and Adaptive Transfer

ICML 2025poster

With the growing availability of multi-domain time series data, there is an increasing demand for general forecasting models pre-trained on multi-source datasets to support diverse downstream prediction scenarios. Existing time series foundation models primarily focus on scaling up pre-training data…

Cited by 0SourcePDFScholar
2025

Universal Backdoor Defense via Label Consistency in Vertical Federated Learning

IJCAI 2025

Backdoor attacks in vertical federated learning (VFL) are particularly concerning as they can covertly compromise VFL decision-making, posing a severe threat to critical applications of VFL. Existing defense mechanisms typically involve either label obfuscation during training or model pruning durin

Cited by 0SourcePDFScholar
2024

AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis

CVPR 2024highlight

Neural implicit fields have been a de facto standard in novel view synthesis. Recently there exist some methods exploring fusing multiple modalities within a single field aiming to share implicit features from different modalities to enhance reconstruction performance. However these modalities often…

2024

Enabling Secure Wireless Communications via Movable Antennas

ICASSP 2024accepted

A pioneering secure transmission scheme is proposed, which harnesses movable antennas (MAs) to optimize antenna positions for augmenting the physical layer security. Particularly, an MA-enabled secure wireless system is considered, where a multi-antenna transmitter communicates with a single-antenna…

Cited by 0SourceScholar
2024

LiT: Unifying LiDAR "Languages" with LiDAR Translator

NeurIPS 2024poster

LiDAR data exhibits significant domain gaps due to variations in sensors, vehicles, and driving environments, creating “language barriers” that limit the effective use of data across domains and the scalability of LiDAR perception models. To address these challenges, we introduce the LiDAR Translato…

2024

Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series Forecasting

ICLR 2024poster

Transformers for time series forecasting mainly model time series from limited or fixed scales, making it challenging to capture different characteristics spanning various scales. We propose Pathformer, a multi-scale Transformer with adaptive pathways. It integrates both temporal resolution and temp…

2024

Programmable Pressure Control in Pneumatic Soft Robots With 2-Way 2-State Solenoid Valves

RA-L 2024

Pneumatic soft robotics is known for its high adaptability and strong load capacity, and the pneumatic control system is regarded as the neural system of pneumatic soft robots to transmit signals and control pressure. The control accuracy and responsiveness of the system are critical to the performa

Cited by 10SourceScholar
2024

Real Appearance Modeling for More General Deepfake Detection

ECCV 2024poster

"Recent studies in deepfake detection have shown promising results when detecting deepfakes of the same type as those present in training. However, their ability to generalize to unseen deepfakes remains limited. This work improves the generalizable deepfake detection from a simple principle: an ide…

Cited by 3SourcePDFScholar
2024

Similarity Knowledge Distillation with Calibrated Mask

ICASSP 2024accepted

In this paper, we propose a novel and efficient method for knowledge distillation, which is structurally simple and requires negligible computation overhead. Our method includes three modules. The first module is the calibrated mask, which avoids the teacher model’s incorrect representation to distu…

Cited by 0SourceScholar
2024

Task-Agnostic Self-Distillation for Few-Shot Action Recognition

IJCAI 2024poster

Task-oriented matching is one of the core aspects of few-shot Action Recognition. Most previous works leverage the metric features within the support and query sets of individual tasks, without considering the metric information across different matching tasks. This oversight represents a significan…

Cited by 1SourcePDFScholar
2023

BEVHeight: A Robust Framework for Vision-Based Roadside 3D Object Detection

CVPR 2023poster

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision…

2023

DEHRFormer: Real-Time Transformer for Depth Estimation and Haze Removal from Varicolored Haze Scenes

ICASSP 2023accepted

Varicolored haze caused by chromatic casts poses haze removal and depth estimation challenges. Recent learning-based depth estimation methods are mainly targeted at dehazing first and estimating depth subsequently from haze-free scenes. This way, the inner connections between colored haze and scene…

Cited by 0SourceScholar
2023

MSP-Former: Multi-Scale Projection Transformer for Single Image Desnowing

ICASSP 2023accepted

Snow removal causes challenges due to its characteristic of complex degradations. To this end, targeted treatment of multi-scale snow degradations is critical for the network to learn effective snow removal. In order to handle the diverse scenes, we propose a multi-scale projection transformer (MSP-…

Cited by 0SourceScholar
2023

PRODIGY: Enabling In-context Learning Over Graphs

NeurIPS 2023spotlight

In-context learning is the ability of a pretrained model to adapt to novel and diverse downstream tasks by conditioning on prompt examples, without optimizing any parameters. While large language models have demonstrated this ability, how in-context learning could be performed over graphs is unexpl…

Cited by 79SourcePDFScholar
2021

Explainable Subgraph Reasoning for Forecasting on Temporal Knowledge Graphs

ICLR 2021poster

Modeling time-evolving knowledge graphs (KGs) has recently gained increasing interest. Here, graph representation learning has become the dominant paradigm for link prediction on temporal KGs. However, the embedding-based approaches largely operate in a black-box fashion, lacking the ability to inte…

Cited by 229SourcePDFScholar
2021

Focusing-Based Wideband Adaptive Beamforming Using Covariance Matrix Reconstruction

ICASSP 2021accepted

Most focusing-based beamforming methods are devoted to solely focusing matrix designing, which aims to minimize the overall focusing error. However, these methods may suffer from performance degradation when steering vector (SV) error exists. To maximize the overall performance of focusing-based bea…

Cited by 0SourceScholar
2021

Locate Globally, Segment Locally: A Progressive Architecture With Knowledge Review Network for Salient Object Detection

AAAI 2021technical

Salient object location and segmentation are two different tasks in salient object detection (SOD). The former aims to globally find the most attractive objects in an image, whereas the latter can be achieved only using local regions that contain salient objects. However, previous methods mainly acc…

Cited by 172SourcePDFScholar
2021

SA-BNN: State-Aware Binary Neural Network

AAAI 2021technical

Binary Neural Networks (BNNs) have received significant attention due to the memory and computation efficiency recently. However, the considerable accuracy gap between BNNs and their full-precision counterparts hinders BNNs to be deployed to resource-constrained platforms. One of the main reasons fo…

Cited by 24SourcePDFScholar
2019

Learning to Capture a Film-Look Video with a Camera Drone

ICRA 2019poster

The development of intelligent drones has simplified aerial filming and provided smarter assistant tools for users to capture a film-look footage. Existing methods of autonomous aerial filming either specify predefined camera movements for a drone to capture a footage, or employ heuristic approaches…

Cited by 47SourceScholar
2019

Learning to Film From Professional Human Motion Videos

CVPR 2019poster

We investigate the problem of 6 degrees of freedom (DOF) camera planning for filming professional human motion videos using a camera drone. Existing methods either plan motions for only a pan-tilt-zoom (PTZ) camera, or adopt ad-hoc solutions without carefully considering the impact of video content…

Cited by 41PDFScholar
2019

Projected Stein Variational Newton: A Fast and Scalable Bayesian Inference Method in High Dimensions

NeurIPS 2019poster

We propose a projected Stein variational Newton (pSVN) method for high-dimensional Bayesian inference. To address the curse of dimensionality, we exploit the intrinsic low-dimensional geometric structure of the posterior distribution in the high-dimensional parameter space via its Hessian (of the lo…

2018

3D Exterior Soundfield Reproduction Using a Planar Loudspeaker Array

ICASSP 2018accepted

In this paper, we propose a planar array of dipole (or first-order) loudspeakers to reproduce a full three dimensional (3D) exterior soundfield over a desired spatial region. The proposed method is inspired by the spherical harmonics based solution for 3D soundfield reproduction. When decomposed in…

Cited by 10SourceScholar
2018

ACT: An Autonomous Drone Cinematography System for Action Scenes

ICRA 2018poster

Drones are enabling new forms of cinematography. Aerial filming via drones in action scenes is difficult because it requires users to understand the dynamic scenarios and operate the drone and camera simultaneously. Existing systems allow the user to manually specify the shots and guide the drone to…

Cited by 97SourceScholar
2018

Design of a 2 Motor 2 Degrees-of-Freedom Coupled Tendon-driven Joint Module

IROS 2018poster

A 2 motor 2 degrees-of-freedom (2M2D) coupled tendon driven joint module is proposed as a basic component for robot arms. Torque reallocation via tendon coupling can enhance the output torque of one single joint. According to the motor position, the joint module is classified into four types: the ex…

Cited by 9SourceScholar
2017

REDBEE: A visual-inertial drone system for real-time moving object detection

IROS 2017poster

Aerial surveillance and monitoring demand both real-time and robust motion detection from a moving camera. Most existing techniques for drones involve sending a video data streams back to a ground station with a high-end desktop computer or server. These methods share one major drawback: data transm…

Cited by 29SourceScholar