← Search

Yubei Chen

15 accepted papers

2026

Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

CVPR 2026

A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical state of both the embodied agent and its environment. Accurate world models are essential for enabling agents to think, plan, and reason effectively in comple

Cited by 0SourcecodeScholar
2026

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

AAAI 2026technical

Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown im

Cited by 0SourcePDFScholar
2025

Neural Motion Simulator Pushing the Limit of World Models in Reinforcement Learning

CVPR 2025poster

An embodied system must not only model the patterns of the external world but also understand its own motion dynamics. A motion dynamic model is essential for efficient skill acquisition and effective planning. In this work, we introduce the neural motion simulator (MoSim), a world model that predic…

2025

Recover Biological Structure from Sparse-View Diffraction Images with Neural Volumetric Prior

ICCV 2025poster

Volumetric reconstruction of label-free living cells from non-destructive optical microscopic images reveals cellular metabolism in native environments. However, current optical tomography techniques require hundreds of 2D images to reconstruct a 3D volume, hindering them from intravital imaging of…

Cited by 0SourcePDFScholar
2025

URLOST: Unsupervised Representation Learning without Stationarity or Topology

ICLR 2025poster

Unsupervised representation learning has seen tremendous progress. However, it is constrained by its reliance on domain specific stationarity and topology, a limitation not found in biological intelligence systems. For instance, unlike computer vision, human vision can process visual signals sampled…

Cited by 1SourcePDFScholar
2024

Pose-Aware Self-Supervised Learning with Viewpoint Trajectory Regularization

ECCV 2024oral

"Learning visual features from unlabeled images has proven successful for semantic categorization, often by mapping different views of the same object to the same feature to achieve recognition invariance. However, visual recognition involves not only identifying what an object is but also understan…

2024

Unsupervised Feature Learning with Emergent Data-Driven Prototypicality

CVPR 2024poster

Given a set of images our goal is to map each image to a point in a feature space such that not only point proximity indicates visual similarity but where it is located directly encodes how prototypical the image is according to the dataset. Our key insight is to perform unsupervised feature learnin…

Cited by 4SourcePDFScholar
2023

Minimalistic Unsupervised Representation Learning with the Sparse Manifold Transform

ICLR 2023top-25%

We describe a minimalistic and interpretable method for unsupervised representation learning that does not require data augmentation, hyperparameter tuning, or other engineering designs, but nonetheless achieves performance close to the state-of-the-art (SOTA) SSL methods. Our approach leverages the…

Cited by 8SourcePDFScholar
2023

On the duality between contrastive and non-contrastive self-supervised learning

ICLR 2023top-5%

Recent approaches in self-supervised learning of image representations can be categorized into different families of methods and, in particular, can be divided into contrastive and non-contrastive approaches. While differences between the two families have been thoroughly discussed to motivate new a…

Cited by 116SourcePDFScholar
2022

Decoupled Contrastive Learning

ECCV 2022poster

"Contrastive learning (CL) is one of the most successful paradigms for self-supervised learning (SSL). In a principled way, it considers two augmented views of the same image as positive to be pulled closer, and all other images negative to be pushed further apart. However, behind the impressive suc…

2019

Superposition of many models into one

NeurIPS 2019poster

We present a method for storing multiple models within a single set of parameters. Models can coexist in superposition and still be retrieved individually. In experiments with neural networks, we show that a surprisingly large number of models can be effectively stored within a single parameter inst…

Cited by 154SourcePDFScholar