← Search

Zhimin Chen

6 accepted papers

2026

Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders

ICLR 2026poster

Modern large-scale recommendation systems rely heavily on user interaction history sequences to enhance the model performance. The advent of large language models and sequential modeling techniques, particularly transformer architectures, has led to significant advancements (e.g., HSTU, SIM, and TW…

Cited by 0SourcecodeScholar
2026

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

CVPR 2026

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical instruments, and human speech). Prior video-conditioned audio genera

Cited by 0SourcecodeScholar
2026

REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion

CVPR 2026

As an important and challenging problem in computer vision, Panoramic Semantic Segmentation (PASS) aims to provide complete scene perception based on an ultra-wide angle of view. Most PASS methods often focus on spherical geometry with RGB input or use the depth information in original or HHA format

Cited by 0SourceScholar
2025

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Learner

ICCV 2025poster

Recently, multi-modal masked autoencoders (MAE) has been introduced in 3D self-supervised learning, offering enhanced feature learning by leveraging both 2D and 3D data to capture richer cross-modal representations. However, these approaches have two limitations: (1) they inefficiently require both…

Cited by 0SourcePDFScholar
2024

SAM-Guided Masked Token Prediction for 3D Scene Understanding

NeurIPS 2024poster

Foundation models have significantly enhanced 2D task performance, and recent works like Bridge3D have successfully applied these models to improve 3D scene understanding through knowledge distillation, marking considerable advancements. Nonetheless, challenges such as the misalignment between 2D an…

Cited by 1SourcePDFScholar
2023

Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation Models

NeurIPS 2023poster

Foundation models have achieved remarkable results in 2D and language tasks like image segmentation, object detection, and visual-language understanding. However, their potential to enrich 3D scene representation learning is largely untapped due to the existence of the domain gap. In this work, we p…