← Search

Yunke Wang

16 accepted papers

2026

Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation

ICLR 2026poster

Robotic manipulation with Vision-Language-Action models requires efficient inference over long-horizon multi-modal context, where attention to dense visual tokens dominates computational cost. Existing methods optimize inference speed by reducing visual redundancy within VLA models, but they overloo…

Cited by 0SourcecodeScholar
2026

Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation

CVPR 2026

Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they remain brittle under distribution shifts: when test scenarios change, VLAs often reproduce memorized trajectories inste

Cited by 0SourcecodeScholar
2026

D2Cache: Second-Order Delta Caching for Higher Video Diffusion Acceleration

CVPR 2026

Video diffusion models achieve impressive visual fidelity but remain computationally prohibitive for real-time or interactive generation due to their sequential denoising process. Recent caching methods accelerate inference by reusing outputs across timesteps, typically estimating each new output fr

Cited by 0SourcecodeScholar
2026

GeoCoT: Towards Reliable Remote Sensing Reasoning with Manifold Perspective

CVPR 2026

Multimodal Large Language Models (MLLMs) have shown strong potential in remote sensing (RS) through multi-task reasoning and cross-modal generalization.However, existing RS-MLLMs mainly rely on a single shared expert for all tasks, making it hard to produce reliable results. Meanwhile, the intrinsic

Cited by 0SourceScholar
2026

Motion Dynamics Learning for Few-Shot Embodied Adaptation

ICML 2026poster

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, yet adapting pretrained models to novel tasks typically relies on substantial task-specific demonstrations, limiting scalability. Current VLA methods mostly focus on action imitation, which ignores the richer s…

Cited by 0SourceScholar
2026

See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model

ICML 2026poster

Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deployment. Existing token pruning methods suffer from a fundamental trade-off: aggressive compression using pruning inevitably discards critical geometric …

Cited by 0SourceScholar
2026

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

ICML 2026poster

Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain gap and limited environmental diversity, resulting in weak real-world generaliza…

Cited by 0SourceScholar
2025

VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching

NeurIPS 2025poster

Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions in an end-to-end manner. However, their substantial computational cost poses a challenge for real-time robotic control,…

Cited by 0SourcecodeScholar
2025

WaterDiffusion: Learning a Prior-involved Unrolling Diffusion for Joint Underwater Saliency Detection and Visual Restoration

AAAI 2025technical

Underwater salient object detection (USOD) plays a pivotal role in various vision-based marine exploration tasks. However, existing USOD techniques face the dilemma of object mislocalization and imprecise boundaries due to the complex underwater environment. The quality degradation of raw underwater…

Cited by 0SourcePDFScholar