← Search

Yun Cheng

9 accepted papers

2026

Cut Less, Fold More: Model Compression through the Lens of Projection Geometry

ICLR 2026poster

Compressing neural networks without retraining is vital for deployment at scale. We study calibration-free compression through the lens of projection geometry: structured pruning is an axis-aligned projection, whereas model folding performs a low-rank projection via weight clustering. We formalize b…

Cited by 0SourceScholar
2026

DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question Answering

AAAI 2026technical

Video question answering (VideoQA), whose goal is to produce answers through the integration of linguistic and visual understanding, has emerged as a significant research focus. Although Large Multimodal Models (LMMs) and autonomous agent methods have achieved notable advances in VideoQA, excessive

Cited by 0SourcePDFScholar
2025

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

ICML 2025poster

Vision Language Models (VLMs) are impressive at visual question answering and image captioning. But they underperform on multi-step visual reasoning---even compared to LLMs on the same tasks presented in text form---giving rise to perceptions of *modality imbalance* or *brittleness*. Towards a syste…

2024

ESP-PCT: Enhanced VR Semantic Performance through Efficient Compression of Temporal and Spatial Redundancies in Point Cloud Transformers

IJCAI 2024poster

Semantic recognition is pivotal in virtual reality (VR) applications, enabling immersive and interactive experiences. A promising approach is utilizing millimeter-wave (mmWave) signals to generate point clouds. However, the high computational and memory demands of current mmWave point cloud models h…

2024

LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition

AAAI 2024technical

The Vision Transformer (ViT) excels in accuracy when handling high-resolution images, yet it confronts the challenge of significant spatial redundancy, leading to increased computational and memory requirements. To address this, we present the Localization and Focus Vision Transformer (LF-ViT). This…

2024

Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications

ICLR 2024poster

In many machine learning systems that jointly learn from multiple modalities, a core research question is to understand the nature of multimodal interactions: how modalities combine to provide new task-relevant information that was not present in either alone. We study this challenge of interaction…

2023

Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework

NeurIPS 2023poster

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain fundamental research questions: How can we quantify the interact…

2021

MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

NeurIPS 2021poster

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. It is a challenging yet crucial area with numerous real-world applications in multimedia, affective computing, robotics, finance, human-computer interaction, and healthcare. Unfortunatel…

Cited by 186SourceScholar