← Search

Yunze Liu

13 accepted papers

2026

MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

ICLR 2026poster

The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. Nevertheless, visual language models (VLMs) still face significant computational overhead when scaled from images to the video domain. When video data is too large (due to high frame rates and long dura…

Cited by 0SourceScholar
2025

CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene Rendering

ICCV 2025poster

Current state-of-the-art 3D reconstruction models face limitations in building extra-large scale outdoor scenes, primarily due to the lack of sufficiently large-scale and detailed datasets. In this paper, we present a extra-large fine-grained dataset with 10 billion points composed of 41,006 drone-c…

Cited by 0SourcePDFScholar
2025

MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining

CVPR 2025poster

Hybrid Mamba-Transformer networks have recently garnered broad attention. These networks can leverage the scalability of Transformers while capitalizing on Mamba's strengths in long-context modeling and computational efficiency. However, the challenge of effectively pretraining such hybrid networks…

2025

MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic Data

CVPR 2025poster

This paper introduces MobileH2R, a framework for learning generalizable vision-based human-to-mobile-robot (H2MR) handover skills. Unlike traditional fixed-base handovers, this task requires a mobile robot to reliably receive objects in a large workspace enabled by its mobility. Our key insight is t…

Cited by 0SourcePDFScholar
2025

MutualNeRF: Improve the Performance of NeRF under Limited Samples with Mutual Information Theory

UAI 2025

This paper introduces MutualNeRF, a framework enhancing Neural Radiance Field (NeRF) performance under limited samples using Mutual Information Theory. While NeRF excels in 3D scene synthesis, challenges arise with limited data and existing methods that aim to introduce prior knowledge lack theoreti

Cited by 0SourcePDFScholar
2025

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

EMNLP 2025

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and personalized assistive technologies. However, existing benchmark

2024

CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding

ICRA 2024poster

This paper introduces a novel approach named CrossVideo, which aims to enhance self-supervised cross-modal contrastive learning in the field of point cloud video understanding. Traditional supervised learning methods encounter limitations due to data scarcity and challenges in label acquisition. To…

Cited by 5SourceScholar
2023

Complete-to-Partial 4D Distillation for Self-Supervised Point Cloud Sequence Representation Learning

CVPR 2023poster

Recent work on 4D point cloud sequences has attracted a lot of attention. However, obtaining exhaustively labeled 4D datasets is often very expensive and laborious, so it is especially important to investigate how to utilize raw unlabeled data. However, most existing self-supervised point cloud repr…

2022

HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction

CVPR 2022poster

We present HOI4D, a large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction. HOI4D consists of 2.4M RGB-D egocentric video frames over 4000 sequences collected by 9 participants interacting with 800 different object instances from…

Cited by 183PDFcodeScholar
2022

Point Primitive Transformer for Long-Term 4D Point Cloud Video Understanding

ECCV 2022poster

"This paper proposes a 4D backbone for long-term point cloud video understanding. A typical way to capture spatial-temporal context is using 4Dconv or transformer without hierarchy. However, those methods are neither effective nor efficient enough due to camera motion, scene changes, sampling patter…

2021

Contrastive Multimodal Fusion With TupleInfoNCE

ICCV 2021poster

This paper proposes a method for representation learning of multimodal data using contrastive losses. A traditional approach is to contrast different modalities to learn the information shared between them. However, that approach could fail to learn the complementary synergies between modalities tha…

Cited by 83PDFcodeScholar