← Search

Jun Peng

10 accepted papers

2026

ADAFLOW: EFFICIENT LONG VIDEO EDITING VIA ADAPTIVE ATTENTION SLIMMING AND KEYFRAME SELECTION

ICASSP 2026poster

Despite great progress, text-driven long video editing is still notoriously challenging mainly due to excessive memory overhead. Although recent efforts have simplified this task into a two-step process of keyframe translation and interpolation generation, the token-wise keyframe translation still p…

Cited by 0SourcePDFScholar
2026

KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs

AAAI 2026technical

Training-free video understanding methods leverage the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating videos as a sequences of static frames, thus obviating the need for costly video-specific training. However, this paradigm often suffers from severe

Cited by 0SourcePDFScholar
2025

Bioinspired Microrobot Climbing on Fabrics Using a Single Actuator

RA-L 2025

Microrobots climbing on fabrics are well-suited for reconnaissance and rescue tasks in indoor environments. However, achieving stable adhesion while maintaining a simplified locomotion mechanism remains a formidable challenge. This study presents a 4cm, 18.2g climbing robot designed with a single-ac

Cited by 1SourceScholar
2025

From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors via LLM-guided Symbolic Reasoning

ICCV 2025poster

Current object detectors excel at entity localization and classification, yet exhibit inherent limitations in event recognition capabilities. This deficiency arises from their architecture's emphasis on discrete object identification rather than modeling the compositional reasoning, inter-object cor…

Cited by 0SourcePDFScholar
2025

TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning

AAAI 2025technical

Despite the efficiency of prompt learning in transferring vision-language models (VLMs) to downstream tasks, existing methods mainly learn the prompts in a coarse-grained manner where the learned prompt vectors are shared across all categories. Consequently, the tailored prompts often fail to discer…

2024

Efficient Event Stream Super-Resolution with Recursive Multi-Branch Fusion

IJCAI 2024poster

Current Event Stream Super-Resolution (ESR) methods overlook the redundant and complementary information present in positive and negative events within the event stream, employing a direct mixing approach for super-resolution, which may lead to detail loss and inefficiency. To address these issues,…

2024

Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization

ICML 2024poster

Text-to-3D-aware face (T3D Face) generation and manipulation is an emerging research hot spot in machine learning, which still suffers from low efficiency and poor quality. In this paper, we propose an ***E**nd-to-End **E**fficient and **E**ffective* network for fast and accurate T3D face generation…

2022

PixelFolder: An Efficient Progressive Pixel Synthesis Network for Image Generation

ECCV 2022poster

"Pixel synthesis is a promising research paradigm for image generation, which can well exploit pixel-wise prior knowledge for generation. However, existing methods still suffer from excessive memory footprint and computation overhead. In this paper, we propose a progressive pixel synthesis network t…

2021

Hierarchical Context Guided Aggregation Network for Stereo Matching

ICASSP 2021accepted

Nowadays, CNN-based stereo matching methods achieved remarkable performance, and how to efficiently exploit contextual information in cost aggregation stage is the key to improve performance. In this paper, we propose a simple yet efficient network named Hierarchical Context Guided Aggregation Netwo…

Cited by 0SourceScholar
2019

Towards Cross-modality Topic Modelling via Deep Topical Correlation Analysis

ICASSP 2019accepted

The cross-modality topic detection in social media retains as an open problem mainly due to the difficulty of dealing with modality independence and modality missing. In this paper, we present a novel Deep Topical Correlation Analysis (DTCA) approach, which achieves robust and accurate topic detecti…

Cited by 0SourceScholar