← Search

Yicheng Feng

13 accepted papers

2026

Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild

CVPR 2026

Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small, precisely-labeled datasets and vast in-the-wild footage with unreliable

Cited by 0SourcecodeScholar
2026

LERD: Latent Event-Relational Dynamics for Neurodegenerative Classification

ICML 2026poster

Alzheimer’s disease (AD) alters brain electrophysiology and disrupts multichannel EEG dynamics, making accurate and clinically useful EEG-based diagnosis increasingly important for screening and disease monitoring. However, many existing approaches rely on black-box classifiers and do not explicitly…

Cited by 0SourceScholar
2026

Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos

CVPR 2026

Vision-Language-Action (VLA) models provide a promising paradigm for robot learning by integrating visual perception with language-guided policy learning. However, most existing approaches rely on 2D visual inputs to perform actions in 3D physical environments, creating a significant gap between per

Cited by 0SourcecodeScholar
2026

Vision-Language-Action Pretraining from Large-Scale Human Videos

ICML 2026poster

Existing Vision-Language-Action (VLA) models struggle with complex manipulation tasks requiring high dexterity and generalization, primarily due to their reliance on synthetic data with significant sim-to-real gaps or limited teleoperated demonstrations. To address this bottleneck, we propose levera…

Cited by 0SourceScholar
2025

From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities

ICLR 2025poster

Multimodal Large Language Models have made significant strides in integrating visual and textual information, yet they often struggle with effectively aligning these modalities. We introduce a novel image tokenizer that bridges this gap by applying the principle of Byte-Pair Encoding (BPE) to visual…

2025

OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and Data

NeurIPS 2025poster

Recent advances in large multimodal models have significantly advanced video comprehension, yet their performance remains limited in first-person scenarios. The interactive nature of egocentric videos is critical for applications like embodied intelligence, but introduces complex visual contexts tha…

Cited by 0SourcecodeScholar
2025

Unified Multimodal Understanding via Byte-Pair Visual Encoding

ICCV 2025poster

Multimodal large language models (MLLMs) have made significant progress in vision-language understanding, yet effectively aligning different modalities remains a fundamental challenge. We present a framework that unifies multimodal understanding by applying byte-pair encoding to visual tokens. Unlik…

Cited by 0SourcePDFScholar
2025

VideoOrion: Tokenizing Object Dynamics in Videos

ICCV 2025poster

We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos--the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to extract object dynamics through a detect-segment-track pipeline,…

Cited by 0SourcePDFScholar
2024

LLaMA-Rider: Spurring Large Language Models to Explore the Open World

NAACL 2024findings

Recently, various studies have leveraged Large Language Models (LLMs) to help decision-making and planning in environments and try to align the LLMs’ knowledge with the world conditions. Nonetheless, the capacity of LLMs to continuously acquire environmental knowledge and adapt in an open world rema…

2024

Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

ICLR 2024poster

Recent studies have presented compelling evidence that large language models (LLMs) can equip embodied agents with the self-driven capability to interact with the world, which marks an initial step toward versatile robotics. However, these efforts tend to overlook the visual richness of open worlds,…

Cited by 26SourcePDFScholar
2024

UniCode : Learning a Unified Codebook for Multimodal Large Language Models

ECCV 2024poster

"In this paper, we propose UniCode, a novel approach within the domain of multimodal large language models (MLLMs) that learns a unified codebook to efficiently tokenize visual, text, and potentially other types of signals. This innovation addresses a critical limitation in existing MLLMs: their rel…

Cited by 12SourcePDFScholar