← Search

Zhenyu Yang

24 accepted papers

2026

Beyond Quadratic: Linear-Time Change Detection with RWKV

AAAI 2026technical

Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this confli

Cited by 0SourcePDFScholar
2026

Capacity-Agnostic Parameter Isolation for Continual Graph Learning

ICML 2026poster

Existing parameter isolation-based methods in continual learning employ diverse designs to learn more tasks within a limited model capacity. However, most of their designs inevitably incur substantial computational overhead if their model capacity is enlarged to accommodate further tasks as the task…

Cited by 0SourceScholar
2026

FedHPro: Federated Hyper-Prototype Learning via Gradient Matching

ICML 2026poster

Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prototype-based FL is in the spotlight, since shared global prototypes offer semantic anchors for aligning client-specific local prototypes. However, ex…

Cited by 0SourceScholar
2026

LacTokGen: Latent Consistency Tokenizer for 1024-pixel Image Generation by 256 Tokens

CVPR 2026

Image tokenization has significantly advanced visual generation and multimodal modeling, particularly when paired with autoregressive models. However, current methods face challenges in balancing efficiency and quality: high-resolution image generation either requires an excessive number of tokens o

Cited by 0SourcecodeScholar
2026

MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments

RA-L 2026

Autonomous navigation in unknown environments requires multi-scale spatial understanding that captures geometric details, topological connectivity, and global structure to support high-level decision making under partial observability. Existing approaches struggle to efficiently capture such multi-s

Cited by 1SourceScholar
2026

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

ICML 2026poster

Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically-aligned surface-based models. We argue this is an artifact of inefficient surface tokenization and the failure to use anatomy as a predictive signal. We present **NeurIPS**, a framework t…

Cited by 0SourceScholar
2026

QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response

ICLR 2026poster

The increasing demand for real-time interaction in online video scenarios necessitates a new class of efficient streaming video understanding models. However, existing approaches often rely on a flawed, query-agnostic ``change-is-important'' principle, which conflates visual dynamics with semantic r…

Cited by 0SourceScholar
2025

FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality

ICLR 2025poster

In this paper, we present \textbf{\textit{FasterCache}}, a novel training-free strategy designed to accelerate the inference of video diffusion models with high-quality generation. By analyzing existing cache-based methods, we observe that \textit{directly reusing adjacent-step features degrades vid…

Cited by 6SourcePDFScholar
2025

GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models

AAAI 2025technical

Posters serve an essential function in marketing and advertising by improving visual communication and brand visibility, thus significantly contributing to industrial design. With the latest developments in controllable T2I diffusion models, research interest has surged in text rendering within synt…

2025

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

NeurIPS 2025poster

Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and determine optimal response timing, often compromising real-time responsiveness and na…

Cited by 0SourcecodeScholar
2025

MsRAG: Knowledge Augumented Image Captioning with Object-level Multi-source RAG

IJCAI 2025

Language-Visual Large Models (LVLMs) have made significant strides in enhancing visual understanding capabilities. However, these models often struggle with knowledge-based visual tasks due to constrains in their pre-training data scope and timeliness. Existing Retrieval-Augmented Generation (RAG) m

Cited by 0SourcePDFScholar
2025

Multimodal Dialogue Emotion Recognition Based on Label Optimization and Coarse-Grained Assisted Fine-Grained

ICASSP 2025accepted

Multimodal dialogue emotion recognition integrates data from multiple modalities to accurately identify emotional states in conversations. However, differences in expression and information density across modalities complicate the fusion of features. Traditional methods may introduce redundant infor…

Cited by 0SourceScholar
2025

Non-Autoregressive Multimodal Machine Translation

ICASSP 2025accepted

Performing better text translation by integrating auxiliary inputs from visual information has gained widespread attention in recent years. While existing methods outperform the text-only translation models, the step-by-step generative style reduces the inference speed, which limits their applicabil…

Cited by 0SourceScholar
2025

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

ICLR 2025spotlight

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding ty…

2025

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

ICCV 2025poster

Text-to-image (T2I) models are well known for their ability to produce highly realistic images, while multimodal large language models (MLLMs) are renowned for their proficiency in understanding and integrating multiple modalities. However, currently there is no straightforward and efficient framewo…

2023

Generating Coherent Narratives by Learning Dynamic and Discrete Entity States with a Contrastive Framework

AAAI 2023technical

Despite advances in generating fluent texts, existing pretraining models tend to attach incoherent event sequences to involved entities when generating narratives such as stories and news. We conjecture that such issues result from representing entities as static embeddings of superficial words, whi…

2022

Curriculum-Based Self-Training Makes Better Few-Shot Learners for Data-to-Text Generation

IJCAI 2022poster

Despite the success of text-to-text pre-trained models in various natural language generation (NLG) tasks, the generation performance is largely restricted by the number of labeled data in downstream tasks, particularly in data-to-text generation tasks. Existing works mostly utilize abundant unlabel…

2022

Dual-discriminative Graph Neural Network for Imbalanced Graph-level Anomaly Detection

NeurIPS 2022accept

Graph-level anomaly detection aims to distinguish anomalous graphs in a graph dataset from normal graphs. Anomalous graphs represent a very few but essential patterns in the real world. The anomalous property of a graph may be referable to its anomalous attributes of particular nodes and anomalous s…

Cited by 44SourcePDFScholar
2022

LaMemo: Language Modeling with Look-Ahead Memory

NAACL 2022long

Although Transformers with fully connected self-attentions are powerful to model long-term dependencies, they are struggling to scale to long texts with thousands of words in language modeling. One of the solutions is to equip the model with a recurrence memory. However, existing approaches directly…

2019

Learning to Capture a Film-Look Video with a Camera Drone

ICRA 2019poster

The development of intelligent drones has simplified aerial filming and provided smarter assistant tools for users to capture a film-look footage. Existing methods of autonomous aerial filming either specify predefined camera movements for a drone to capture a footage, or employ heuristic approaches…

Cited by 47SourceScholar
2019

Learning to Film From Professional Human Motion Videos

CVPR 2019poster

We investigate the problem of 6 degrees of freedom (DOF) camera planning for filming professional human motion videos using a camera drone. Existing methods either plan motions for only a pan-tilt-zoom (PTZ) camera, or adopt ad-hoc solutions without carefully considering the impact of video content…

Cited by 41PDFScholar
2018

ACT: An Autonomous Drone Cinematography System for Action Scenes

ICRA 2018poster

Drones are enabling new forms of cinematography. Aerial filming via drones in action scenes is difficult because it requires users to understand the dynamic scenarios and operate the drone and camera simultaneously. Existing systems allow the user to manually specify the shots and guide the drone to…

Cited by 97SourceScholar