← Search

Joy Hsu

16 accepted papers

2026

Discovering Hybrid World Representations with Co-Evolving Foundation Models

AAAI 2026technical

This perspective article discusses an emerging research direction: to what extent can foundation models yield usable structure for modeling the physical world? We offer a Markovian formulation of structured world models and outline the notion of multi-level hybrid world representations that support

Cited by 0SourcePDFScholar
2026

Learning Situated Awareness in the Real World

ICML 2026spotlight

A core aspect of human perception is *situated awareness*, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models (MFMs) emphasize **environment-centric** spatial relation…

Cited by 0SourceScholar
2025

From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries

NeurIPS 2025poster

Real-world scenes, such as those in ScanNet, are difficult to capture, with highly limited data available. Generating realistic scenes with varied object poses remains an open and challenging task. In this work, we propose FactoredScenes, a framework that synthesizes realistic 3D scenes by leveragin…

Cited by 0SourceScholar
2024

Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners

CVPR 2024poster

3D visual grounding is a challenging task that often requires direct and dense supervision notably the semantic label for each object in the scene. In this paper we instead study the naturally supervised setting that learns from only 3D scene and QA pairs where prior works underperform. We propose t…

Cited by 7SourcePDFScholar
2023

Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

ICLR 2023top-25%

Robots operating in the real world require both rich manipulation skills as well as the ability to semantically reason about when to apply those skills. Towards this goal, recent works have integrated semantic representations from large-scale pretrained vision-language (VL) models into manipulation…

2023

What’s Left? Concept Grounding with Logic-Enhanced Foundation Models

NeurIPS 2023poster

Recent works such as VisProg and ViperGPT have smartly composed foundation models for visual reasoning—using large language models (LLMs) to produce programs that can be executed by pre-trained vision-language models. However, they operate in limited domains, such as 2D images, not fully exploiting…

2021

Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations

NeurIPS 2021poster

We consider the task of representation learning for unsupervised segmentation of 3D voxel-grid biomedical images. We show that models that capture implicit hierarchical relationships between subvolumes are better suited for this task. To that end, we consider encoder-decoder architectures with a hyp…

Cited by 35SourcePDFScholar
2021

DARCNN: Domain Adaptive Region-Based Convolutional Neural Network for Unsupervised Instance Segmentation in Biomedical Images

CVPR 2021poster

In the biomedical domain, there is an abundance of dense, complex data where objects of interest may be challenging to detect or constrained by limits of human knowledge. Labelled domain specific datasets for supervised tasks are often expensive to obtain, and furthermore discovery of novel distinct…

Cited by 43PDFcodeScholar