← Search

Zirui Ge

4 accepted papers

2026

Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics

ICML 2026poster

Video action models are a promising foundation for Vision–Language–Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynami…

Cited by 0SourceScholar
2026

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

AAAI 2026technical

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how

Cited by 0SourcePDFScholar
2025

GNCL: A Graph Neural Network with Consistency Loss for Segment-Level Spoofed Speech Detection

ICASSP 2025accepted

Segment-level spoofed speech detection focuses on recognizing fake or synthetic segments within identifying partially spoofed speech. Nevertheless, existing models for this segment-level task usually overlook latent local relationships between fake and bona fide segments, and further, a lack of inte…

Cited by 0SourceScholar
2025

Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement

ICASSP 2025accepted

Time-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates an…

Cited by 0SourceScholar