← Search

Michael Ying Yang

19 accepted papers

2026

4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation

AAAI 2026technical

Remarkable advances in recent 2D image and 3D shape generation have induced a significant focus on dynamic 4D content generation. However, previous 4D generation methods commonly struggle to maintain spatial-temporal consistency and adapt poorly to rapid temporal variations, due to the lack of effec

Cited by 5SourcePDFScholar
2026

ARFlow: Auto-regressive Optical Flow Estimation for Arbitrary-Length Videos via Progressive Next-Frame Forecasting

ICLR 2026poster

Optical flow estimation is a fundamental computer vision task that predicts per-pixel displacements from consecutive images. Recent works attempt to exploit temporal cues to improve the estimation performance. However, their temporal modeling is restricted to short video sequences due to the unaffor…

Cited by 0SourceScholar
2026

Query-Guided Spatial–Temporal–Frequency Interaction for Music Audio–Visual Question Answering

ICLR 2026poster

Audio–Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natural language questions. Inspired by recent advances in Video QA, many existing AVQA approaches primarily focus on visual…

Cited by 0SourcecodeScholar
2026

StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift Compensation

CVPR 2026

We propose StreamVLO, a streaming visual-LiDAR odometry framework that performs unified spatio-temporal correlation with Mamba models and tackles the long-standing cumulative drift problem via an online Cumulative Drift Compensation scheme for localization in 4D dynamic environments. Specifically, S

Cited by 0SourceScholar
2025

DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-Temporal Fusion

ICRA 2025

Visual-LiDAR odometry is a critical component for autonomous system localization, yet achieving high accuracy and strong robustness remains a challenge. Traditional approaches commonly struggle with sensor misalignment, fail to fully leverage temporal information, and require extensive manual tuning

Cited by 2SourceScholar
2021

CABiNet: Efficient Context Aggregation Network for Low-Latency Semantic Segmentation

ICRA 2021poster

With the increasing demand of autonomous machines, pixel-wise semantic segmentation for visual scene understanding needs to be not only accurate but also efficient for any potential real-time applications. In this paper, we propose CABiNet (Context Aggregated Bi-lateral Network), a dual branch convo…

Cited by 74SourceScholar
2021

Context-Aware Layout to Image Generation With Enhanced Object Appearance

CVPR 2021poster

A layout to image (L2I) generation model aims to generate a complicated image containing multiple objects (things) against natural background (stuff), conditioned on a given layout. Built upon the recent advances in generative adversarial networks (GANs), recent L2I models have made great progress.…

Cited by 65PDFcodeScholar
2021

Cuboids Revisited: Learning Robust 3D Shape Fitting to Single RGB Images

CVPR 2021poster

Humans perceive and construct the surrounding world as an arrangement of simple parametric models. In particular, man-made environments commonly consist of volumetric primitives such as cuboids or cylinders. Inferring these primitives is an important step to attain high-level, abstract scene descrip…

Cited by 32PDFcodeScholar
2021

Exploring Dynamic Context for Multi-path Trajectory Prediction

ICRA 2021poster

To accurately predict future positions of different agents in traffic scenarios is crucial for safely deploying intelligent autonomous systems in the real-world environment. However, it remains a challenge due to the behavior of a target agent being affected by other agents dynamically and there bei…

Cited by 49SourcecodeScholar
2021

Spatial-Temporal Transformer for Dynamic Scene Graph Generation

ICCV 2021poster

Dynamic scene graph generation aims at generating a scene graph of the given video. Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic inte…

Cited by 175PDFcodeScholar
2020

CONSAC: Robust Multi-Model Fitting by Conditional Sample Consensus

CVPR 2020poster

We present a robust estimator for fitting multiple parametric models of the same form to noisy measurements. Applications include finding multiple vanishing points in man-made scenes, fitting planes to architectural imagery, or estimating multiple rigid motions within the same sequence. In contrast…

Cited by 73PDFcodeScholar
2020

NODIS: Neural Ordinary Differential Scene Understanding

ECCV 2020poster

Semantic image understanding is a challenging topic in computer vision. It requires to detect all objects in an image, but also to identify all the relations between them. Detected objects, their labels and the discovered relations can be used to construct a scene graph which provides an abstract se…

2017

Analyzing modular CNN architectures for joint depth prediction and semantic segmentation

ICRA 2017poster

This paper addresses the task of designing a modular neural network architecture that jointly solves different tasks. As an example we use the tasks of depth estimation and semantic segmentation given a single RGB image. The main focus of this work is to analyze the cross-modality influence between…

Cited by 82SourceScholar
2016

Uncertainty-Driven 6D Pose Estimation of Objects and Scenes From a Single RGB Image

CVPR 2016poster

In recent years, the task of estimating the 6D pose of object instances and complete scenes, i.e. camera localization, from a single input image has received considerable attention. Consumer RGB-D cameras have made this feasible, even for difficult, texture-less objects and scenes. In this work, we…

Cited by 627PDFScholar
2015

Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images

ICCV 2015poster

Analysis-by-synthesis has been a successful approach for many tasks in computer vision, such as 6D pose estimation of an object in an RGB-D image which is the topic of this work. The idea is to compare the observation with the output of a forward process, such as a rendered image of the object of in…

Cited by 263PDFScholar