← Search

Siteng Huang

26 accepted papers

2026

Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models

AAAI 2026technical

Large vision-language models (LVLMs) excel at visual understanding but face efficiency challenges due to quadratic complexity when processing long multimodal contexts. While token compression can reduce computational costs, existing approaches are designed for single-view LVLMs and fail to account f

Cited by 0SourcePDFScholar
2026

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In

Cited by 0SourcecodeScholar
2026

High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning With Gaussian Splatting

RA-L 2026

The scalability of robotic learning is fundamentally bottlenecked by the significant cost and labor of real-world data collection. While simulated data offers a scalable alternative, it often fails to generalize to the real world due to significant gaps in visual appearance, physical properties, and

Cited by 7SourceScholar
2026

RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation

ICRA 2026poster

This paper presents RynnVLA-001, a vision-language-action (VLA) model built upon large-scale video generative pretraining from human demonstrations. We propose a novel two-stage pretraining methodology. The first stage, Ego-Centric Video Generative Pretraining, trains an Image-to-Video model to pred…

2026

Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors

AAAI 2026technical

A dexterous hand capable of generalizable grasping objects is fundamental for the development of general-purpose embodied AI. However, previous methods focus narrowly on low-level grasp stability metrics, neglecting affordance-aware positioning and human-like poses which are crucial for downstream m

Cited by 0SourcePDFScholar
2026

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

AAAI 2026technical

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how

Cited by 0SourcePDFScholar
2026

Variation-aware Vision Token Dropping for Faster Large Vision-Language Models

CVPR 2026

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding tasks. However, the increasing demand for high-resolution image and long-video understanding results in substantial token counts, consequently leading to reduced inference efficiency. Token com

Cited by 0SourcecodeScholar
2025

Accelerating Diffusion Transformers with Token-wise Feature Caching

ICLR 2025poster

Diffusion transformers have shown significant effectiveness in both image and video synthesis at the expense of huge computation costs. To address this problem, feature caching methods have been introduced to accelerate diffusion transformers by caching the features in previous timesteps and reusing…

2025

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

ICCV 2025accepted

In robotic visuomotor policy learning, diffusion-based models have achieved significant success in improving the accuracy of action trajectory generation compared to traditional autoregressive models. However, they suffer from inefficiency due to multiple denoising steps and limited flexibility from…

Cited by 0SourcePDFScholar
2025

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

AAAI 2025technical

In recent years, applying multi-modal large language models (MLLMs) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, MLLMs comprise the well-known Transformer network, which has a less efficient quadratic computation complexity. In this s…

2025

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

CoRL 2025poster

Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing VLA frameworks primarily address short-horizon tasks, and their effectiveness on long-horizon, multi-step robotic manipu…

Cited by 0SourceScholar
2025

Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning

ICRA 2025

This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the l

Cited by 1SourcecodeScholar
2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

NeurIPS 2025poster

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either require specialized sensors or fail to effectively exploit d…

Cited by 0SourcecodeScholar
2024

Check Locate Rectify: A Training-Free Layout Calibration System for Text-to-Image Generation

CVPR 2024poster

Diffusion models have recently achieved remarkable progress in generating realistic images. However challenges remain in accurately understanding and synthesizing the layout requirements in the textual prompts. To align the generated image with layout instructions we present a training-free layout c…

2024

Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation

CVPR 2024poster

This study focuses on a novel task in text-to-image (T2I) generation namely action customization. The objective of this task is to learn the co-existing action from limited data and generalize it to unseen humans or even animals. Experimental results show that existing subject-driven customization m…

2024

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

ECCV 2024oral

"Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video making it challenging for LVLMs to perform adequately due to the complexity of the relationship between language and s…

2024

Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation

AAAI 2024technical

Recently, despite the unprecedented success of large pre-trained visual-language models (VLMs) on a wide range of downstream tasks, the real-world unsupervised domain adaptation (UDA) problem is still not well explored. Therefore, in this paper, we first experimentally demonstrate that the unsupervi…

2024

QUAR-VLA: Vision-Language-Action Model for Quadruped Robots

ECCV 2024poster

"The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional quadruped robot learning typically handles language interaction and visual autonomous perception separately, which, while simplifying system design, also limits the sy…

Cited by 19SourcePDFScholar
2024

Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning

CVPR 2024poster

Recent compositional zero-shot learning (CZSL) methods adapt pre-trained vision-language models (VLMs) by constructing trainable prompts only for composed state-object pairs. Relying on learning the joint representation of seen compositions these methods ignore the explicit modeling of the state and…

2024

VGDIFFZERO: Text-To-Image Diffusion Models Can Be Zero-Shot Visual Grounders

ICASSP 2024accepted

Large-scale text-to-image diffusion models have shown impressive capabilities for generative tasks by leveraging strong vision-language alignment from pre-training. However, most vision-language discriminative tasks require extensive fine-tuning on carefully-labeled datasets to acquire such alignmen…

Cited by 0SourceScholar
2023

VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval

CVPR 2023poster

Many recent studies leverage the pre-trained CLIP for text-video cross-modal retrieval by tuning the backbone with additional heavy modules, which not only brings huge computational burdens with much more parameters, but also leads to the knowledge forgetting from upstream models. In this work, we p…

2022

Domain Generalized Few-Shot Image Classification via Meta Regularization Network

ICASSP 2022accepted

In few-shot image classification scenarios, meta-learning methods aim to learn transferable feature representations extracted from seen domains (base classes) in the meta-training phase and quickly adapt to unseen domains (novel classes) in the meta-testing phase. However, when seen and unseen domai…

Cited by 0SourceScholar
2022

Tree Structure-Aware Few-Shot Image Classification via Hierarchical Aggregation

ECCV 2022poster

"In this paper, we mainly focus on the problem of how to learn additional feature representations for few-shot image classification through pretext tasks (e.g., rotation or color permutation and so on). This additional knowledge generated by pretext tasks can further improve the performance of few-s…

2021

Attributes-Guided and Pure-Visual Attention Alignment for Few-Shot Recognition

AAAI 2021technical

The purpose of few-shot recognition is to recognize novel categories with a limited number of labeled examples in each class. To encourage learning from a supplementary view, recent approaches have introduced auxiliary semantic modalities into effective metric-learning frameworks that aim to learn a…