← Search

Xinxin Zhu

10 accepted papers

2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

CVPR 2026

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or inhibit long-range temporal modeling via rigid, predefined sparse patterns. This

Cited by 0SourceScholar
2026

PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization

AAAI 2026technical

Reinforcement Fine-tuning (RFT) methods such as Group Relative Policy Optimization (GRPO) have demonstrated strong capabilities in aligning Large Language Models with human preferences. However, these approaches often suffer from limited data efficiency, necessitating extensive on-policy rollouts to

Cited by 0SourcePDFScholar
2026

W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image Editing

ICLR 2026poster

While recent advances in Diffusion Transformers (DiTs) have significantly advanced text-to-image generation, text-driven image editing remains challenging. Existing approaches either struggle to balance structural preservation with flexible modifications or require costly fine-tuning of large models…

Cited by 0SourceScholar
2025

Resource Allocation for Semantic Segmentation Tasks in Autonomous Driving: A Likelihood Active Inference Approach

ICASSP 2025accepted

The latest Segment Anything Model enables realtime scene annotation and understanding for autonomous driving systems, enhancing driving safety. However, effectively allocating resources for real-time performance and accuracy remains challenging in edge-cloud architectures. Traditional reinforcement…

Cited by 0SourceScholar
2023

MOSO: Decomposing MOtion, Scene and Object for Video Prediction

CVPR 2023poster

Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their dynamics. Based on this insight, we propose a two-stage MOtion, Scene and Object decomposition framework (MOSO) for video…

2023

SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on Hierarchies

AAAI 2023technical

Multivariate time series forecasting with hierarchical structure is widely used in real-world applications, e.g., sales predictions for the geographical hierarchy formed by cities, states, and countries. The hierarchical time series (HTS) forecasting includes two sub-tasks, i.e., forecasting and rec…

Cited by 4SourcePDFScholar
2023

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

NeurIPS 2023poster

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audi…

2021

Consistent-Separable Feature Representation for Semantic Segmentation

AAAI 2021technical

Cross-entropy loss combined with softmax is one of the most commonly used supervision components in most existing segmentation methods. The softmax loss is typically good at optimizing the inter-class difference, but not good at reducing the intra-class variation, which can be suboptimal for semanti…

Cited by 3SourcePDFScholar
2020

Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning

IJCAI 2020poster

Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently, non-autoregressive decoding has been proposed in machine translation to speed up the inference time by generating all…

Cited by 0SourcePDFScholar
2020

Normalized and Geometry-Aware Self-Attention Network for Image Captioning

CVPR 2020poster

Self-attention (SA) network has shown profound value in image captioning. In this paper, we improve SA from two aspects to promote the performance of image captioning. First, we propose Normalized Self-Attention (NSA), a reparameterization of SA that brings the benefits of normalization inside SA. W…

Cited by 284PDFScholar