← Search

Tim K. Marks

19 accepted papers

2026

AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects

CVPR 2026

Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios, overlooking shape complexities and assembly trajectories in industri

Cited by 0SourceScholar
2024

Disentangled Acoustic Fields For Multimodal Physical Scene Understanding

IROS 2024poster

We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward neural networks to directly regress the variables from sound, lea…

Cited by 0SourceScholar
2024

RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation

CVPR 2024poster

We leverage Large Language Models (LLM) for zeroshot Semantic Audio Visual Navigation (SAVN). Existing methods utilize extensive training demonstrations for reinforcement learning yet achieve relatively low success rates and lack generalizability. The intermittent nature of auditory signals further…

Cited by 7SourcePDFScholar
2024

TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models

CVPR 2024poster

Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g. a woman's photo) and a text description (e.g. "a woman is drinking water."). Existing TI2V frameworks often require costly training on video-text datasets and specific model desig…

2023

Discriminative 3D Shape Modeling for Few-Shot Instance Segmentation

ICRA 2023poster

In this paper, we present a simple and efficient scheme for segmenting approximately convex 3D object instances in depth images in a few-shot setting via discriminatively modeling the 3D shape of the object using a neural network. Our key idea is to select pairs of 3D points on the depth image betwe…

Cited by 2SourceScholar
2023

H-SAUR: Hypothesize, Simulate, Act, Update, and Repeat for Understanding Object Articulations from Interactions

ICRA 2023poster

The world is filled with articulated objects that are difficult to determine how to use from vision alone, e.g., a door might open inwards or outwards. Humans handle these objects with strategic trial-and-error: first pushing a door then pulling if that doesn't work. We enable these capabilities in…

Cited by 3SourceScholar
2023

Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image Synthesis

ICCV 2023poster

Conditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models that perform plug-and-play generation, i.e., to use a predefined or pretrained model, which is not explicitly trained o…

Cited by 13PDFcodeScholar
2022

(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering

AAAI 2022technical

Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approaches often ignore the fact that videos are essentially sequences of 2D ``views'' of events happening in a 3D space, and tha…

Cited by 25SourcePDFScholar
2022

Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher Learning

ICASSP 2022accepted

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges (DSTC7, DSTC8). In these challenges, the best-performing systems…

Cited by 0SourceScholar
2022

MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image Manipulation

AAAI 2022technical

Recent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to control the characteristics of the generated faces in a meaningf…

Cited by 35SourcePDFScholar
2021

InSeGAN: A Generative Approach to Segmenting Identical Instances in Depth Images

ICCV 2021poster

In this paper, we present InSeGAN an unsupervised 3D generative adversarial network (GAN) for segmenting (nearly) identical instances of rigid objects in depth images. Using an analysis-by-synthesis approach, we design a novel GAN architecture to synthesize a multiple-instance depth image with indep…

Cited by 2PDFScholar
2020

LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood

CVPR 2020poster

Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting land…

Cited by 196PDFcodeScholar
2019

Audio Visual Scene-Aware Dialog

CVPR 2019poster

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video whi…

Cited by 226PDFcodeScholar
2019

End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features

ICASSP 2019accepted

In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differenti…

Cited by 0SourceScholar
2017

Attention-Based Multimodal Fusion for Video Description

ICCV 2017poster

Current methods for video description are based on encoder-decoder sentence generation using recurrent neural networks (RNNs). Recent work has demonstrated the advantages of integrating temporal attention mechanisms into these models, in which the decoder network predicts each word in the descriptio…

Cited by 469PDFScholar
2016

A Multi-Stream Bi-Directional Recurrent Neural Network for Fine-Grained Action Detection

CVPR 2016poster

We present a multi-stream bi-directional recurrent neural network for fine-grained action detection. Recently, two-stream convolutional neural networks (CNNs) trained on stacked optical flow and image frames have been successful for action recognition in videos. Our system uses a tracking algorithm…

Cited by 606PDFScholar
2015

Real-Time 3D Head Pose and Facial Landmark Estimation From Depth Images Using Triangular Surface Patch Features

CVPR 2015poster

We present a real-time system for 3D head pose estimation and facial landmark localization using a commodity depth sensor. We introduce a novel triangular surface patch (TSP) descriptor, which encodes the shape of the 3D surface of the face within a triangular area. The proposed descriptor is viewpo…

Cited by 120SourcePDFScholar