← Search

Zicong Hong

6 accepted papers

2026

HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning

ICML 2026poster

Vision–Language–Action (VLA) models have shown strong performance in robotic manipulation, but often struggle in long-horizon or out-of-distribution scenarios due to the lack of explicit mechanisms for multimodal reasoning and anticipating how the world will evolve under action. Recent works introdu…

Cited by 0SourceScholar
2026

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization

ICML 2026poster

4-bit quantization reduces the memory footprint and latency of large language model inference, but its aggressive precision reduction can severely degrade accuracy. Prior methods address this by decomposing each weight matrix into two components (e.g., via singular value decomposition) and quantizin…

Cited by 0SourceScholar
2026

WMPO: World Model-based Policy Optimization for Vision-Language-Action Models

ICLR 2026poster

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation, but their reliance on expert demonstrations limits their ability to learn from failures and perform self-corrections. Reinforcement learning (RL) addresses these through self-improving interact…

Cited by 0SourcecodeScholar
2025

DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning

NeurIPS 2025poster

Despite the significant breakthrough of Mixture-of-Experts (MoE), the increasing scale of these MoE models presents huge memory and storage challenges. Existing MoE pruning methods, which involve reducing parameter size with a uniform sparsity across all layers, often lead to suboptimal outcomes and…

Cited by 0SourceScholar
2024

Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language Models

ICML 2024poster

Fine-tuning the learnable prompt for a pre-trained vision-language model (VLM), such as CLIP, has demonstrated exceptional efficiency in adapting to a broad range of downstream tasks. Existing prompt tuning methods for VLMs do not distinguish spurious features introduced by biased training data from…

Cited by 4SourcePDFScholar
2024

Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation Editing

NeurIPS 2024poster

Recent advancements in vision-language-to-image (VL2I) diffusion generation have made significant progress. While generating images from broad vision-language inputs holds promise, it also raises concerns about potential misuse, such as copying artistic styles without permission, which could have le…

Cited by 0SourcePDFScholar