← Search

Ta-Ying Cheng

9 accepted papers

2026

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

AAAI 2026technical

Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown im

Cited by 0SourcePDFScholar
2024

Learning Continuous 3D Words for Text-to-Image Generation

CVPR 2024poster

Current controls over diffusion models (e.g. through text or ControlNet) for image generation fall short in recognizing abstract continuous attributes like illumination direction or non-rigid shape change. In this paper we present an approach for allowing users of text-to-image models to have fine-g…

2024

SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

NeurIPS 2024poster

Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such as articulating dynamic scene changes and motion planning, require a fundamental and explicit 3D understanding beyond c…

Cited by 5SourcePDFScholar
2024

Towards Learning Group-Equivariant Features for Domain Adaptive 3D Detection

NeurIPS 2024poster

The performance of 3D object detection in large outdoor point clouds deteriorates significantly in an unseen environment due to the inter-domain gap. To address these challenges, most existing methods for domain adaptation harness self-training schemes and attempt to bridge the gap by focusing on a…

Cited by 0SourcePDFScholar
2024

ZeST: Zero-Shot Material Transfer from a Single Image

ECCV 2024poster

"We propose , a method for zero-shot material transfer to an object in the input image given a material exemplar image. leverages existing diffusion adapters to extract implicit material representation from the exemplar image. This representation is used to transfer the material using pre-trained in…

2023

3DMiner: Discovering Shapes from Large-Scale Unannotated Image Datasets

ICCV 2023poster

We present 3DMiner -- a pipeline for mining 3D shapes from challenging large-scale unannotated image datasets. Unlike other unsupervised 3D reconstruction methods, we assume that, within a large-enough dataset, there must exist images of objects with similar shapes but varying backgrounds, textures,…

Cited by 0PDFcodeScholar
2023

Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion Estimation

NeurIPS 2023poster

A truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion estimates, we present an SE(3) equivariant architecture and a traini…

2022

Meta-Sampler: Almost-Universal yet Task-Oriented Sampling for Point Clouds

ECCV 2022poster

"Sampling is a key operation in point-cloud task and acts to increase computational efficiency and tractability by discarding redundant points. Universal sampling algorithms (e.g., Farthest Point Sampling) work without modification across different tasks, models, and datasets, but by their very natu…

2022

Pose Adaptive Dual Mixup for Few-Shot Single-View 3D Reconstruction

AAAI 2022technical

We present a pose adaptive few-shot learning procedure and a two-stage data interpolation regularization, termed Pose Adaptive Dual Mixup (PADMix), for single-image 3D reconstruction. While augmentations via interpolating feature-label pairs are effective in classification tasks, they fall short in…

Cited by 9SourcePDFScholar