← Search

Yihan Lin

7 accepted papers

2026

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

ICML 2026oral

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of…

Cited by 0SourceScholar
2026

Mind Dreamer: Untethering Imagination via Active Counterfactual Reasoning on Latent Manifolds

ICML 2026poster

Model-Based Reinforcement Learning (MBRL) leverages latent imagination for sample efficiency, yet remains constrained by **Historical Tethering**: imagination is typically initialized from observed states. This creates a learning asymmetry, where the world model’s manifold discovery outpaces the pol…

Cited by 0SourceScholar
2026

Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor

CVPR 2026

Motion blur arises when rapid scene changes occur during the exposure period, collapsing rich intra-exposure motion into a single RGB frame. Without explicit structural or temporal cues, RGB-only deblurring is highly ill-posed and often fails under extreme motion. Inspired by the human visual system

Cited by 0SourcecodeScholar
2025

CSVO: Complementary-Pathway Spatial-Enhanced Visual Odometry for Extreme Environments with Brain-Inspired Vision Sensors

IROS 2025

Visual Odometry (VO) estimates the pose and motion trajectory of the camera based on visual input, serving as a fundamental technique for robotic positioning and navigation. However, existing VO methods face challenges in visual degradation in extreme environments, e.g., high dynamic range or fast-m

Cited by 0SourcecodeScholar
2025

Diffusion-Based Extreme High-speed Scenes Reconstruction with the Complementary Vision Sensor

ICCV 2025poster

Recording and reconstructing high-speed scenes poses a significant challenge. While high-speed cameras can capture fine temporal details, their extremely high bandwidth demands make continuous recording unsustainable. Conversely, traditional RGB cameras, typically operating at 30 FPS, rely on frame…

2023

Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature Learning

ICASSP 2023accepted

This paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a deep neural network model infers attribute information that…

Cited by 0SourceScholar
2021

Temporal-Wise Attention Spiking Neural Networks for Event Streams Classification

ICCV 2021poster

How to effectively and efficiently deal with spatio-temporal event streams, where the events are generally sparse and non-uniform and have the us temporal resolution, is of great value and has various real-life applications. Spiking neural network (SNN), as one of the brain-inspired event-triggered…

Cited by 224PDFScholar