← Search

Katerina Fragkiadaki

55 accepted papers

2026

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning

CVPR 2026

Humanoid robotics has strong potential to transform daily service and caregiving applications. Although recent advances in general motion tracking within physics engines (GMT) have enabled virtual characters and humanoid robots to reproduce a broad range of human motions, these behaviors are primari

Cited by 0SourceScholar
2026

MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second

CVPR 2026

We present MoVieS, a Motion-aware View Synthesis model that reconstructs 4D dynamic scenes from monocular videos in one second. It represents dynamic 3D scenes with pixel-aligned Gaussian primitives and explicitly supervises their time-varying motions. This allows, for the first time, the unified mo

Cited by 0SourcecodeScholar
2026

RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph

CVPR 2026

Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on top of 2D visual backbones and depend heavily on labeled data for training, which is often scarce in real-world scenarios, causing a sim-to-real gap. Moreover

Cited by 0SourceScholar
2026

RobotArena $\infty$: Unlimited Robot Benchmarking via Real-to-Sim Translation

ICLR 2026poster

The pursuit of robot generalists—instructable agents capable of performing diverse tasks across diverse environments—demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to re…

Cited by 0SourcecodeScholar
2026

Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

ICML 2026poster

We have witnessed remarkable advances in LLM reasoning capabilities with the advent of DeepSeek-R1. However, much of this progress has been fueled by the abundance of internet question–answer (QA) pairs—a major bottleneck going forward, since such data is limited in scale and concentrated mainly in …

Cited by 0SourceScholar
2025

AllTracker: Efficient Dense Point Tracking at High Resolution

ICCV 2025poster

We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing point tracking methods, our approach delivers high-resolution and dense (all-pixel) correspondence fields, which can be…

2025

Diffusion Beats Autoregressive in Data-Constrained Settings

NeurIPS 2025poster

Autoregressive (AR) models have long dominated the landscape of large language models, driving progress across a wide range of tasks. Recently, diffusion-based language models have emerged as a promising alternative, though their advantages over AR models remain underexplored. In this paper, we syst…

Cited by 0SourcecodeScholar
2025

Grounded Reinforcement Learning for Visual Reasoning

NeurIPS 2025poster

While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual reasoning introduces added complexity by requiring models to direct visual attention, interpret perceptual inputs, and ground abstract reasoning in spat…

Cited by 0SourcecodeScholar
2025

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers

NeurIPS 2025poster

We introduce PartCrafter, the first structured 3D generative model that jointly synthesizes multiple semantically meaningful and geometrically distinct 3D meshes from a single RGB image. Unlike existing methods that either produce monolithic 3D shapes or follow two-stage pipelines, i.e. first segmen…

Cited by 0SourceScholar
2025

Robust Multi-Object 4D Generation for In-the-wild Videos

CVPR 2025poster

We address the challenge of generating dynamic 4D scenes from monocular multi-object videos with heavy occlusions and introduce Robust4DGen, a novel approach that integrates rendering-based deformable 3D Gaussian optimization with generative priors for view synthesis. While existing view-synthesis m…

Cited by 0SourcePDFScholar
2025

TAPIP3D: Tracking Any Point in Persistent 3D Geometry

NeurIPS 2025poster

We introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos. TAPIP3D represents videos as camera-stabilized spatio-temporal feature clouds, leveraging depth and camera motion information to lift 2D video features into a 3D world space where camera moveme…

Cited by 0SourcecodeScholar
2025

Unifying 2D and 3D Vision-Language Understanding

ICML 2025poster

Progress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture for 2D and 3D vision-language understanding that bridges the gap between existing 2D-centric models and the rich 3D sensory data available in embodied sys…

2024

3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

CoRL 2024poster

Diffusion policies are conditional diffusion models that learn robot action distributions conditioned on the robot and environment state. They have recently shown to outperform both deterministic and alternative action distribution learning formulations. 3D robot policies use 3D scene feature repres…

Cited by 117SourceScholar
2024

Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous and Instruction-guided Driving

CVPR 2024poster

Diffusion models excel at modeling complex and multimodal trajectory distributions for decision-making and control. Reward-gradient guided denoising has been recently proposed to generate trajectories that maximize both a differentiable reward function and the likelihood under the data distribution…

Cited by 6SourcePDFScholar
2024

DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos

NeurIPS 2024poster

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object dynamic scenes? There are two challenges in this direction: First…

2024

ODIN: A Single Model for 2D and 3D Segmentation

CVPR 2024highlight

State-of-the-art models on contemporary 3D segmentation benchmarks like ScanNet consume and label dataset-provided 3D point clouds obtained through post processing of sensed multiview RGB-D images. They are typically trained in-domain forego large-scale 2D pre-training and outperform alternatives th…

2024

RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

ICML 2024poster

We present RoboGen, a generative robotic agent that automatically learns diverse robotic skills at scale via generative simulation. RoboGen leverages the latest advancements in foundation and generative models. Instead of directly adapting these models to produce policies or low-level actions, we ad…

Cited by 88SourcePDFScholar
2024

Tractable Joint Prediction and Planning over Discrete Behavior Modes for Urban Driving

ICRA 2024poster

Significant progress has been made in training multimodal trajectory forecasting models for autonomous driving. However, effectively integrating these models with downstream planners and model-based control approaches is still an open problem. Although these models have conventionally been evaluated…

Cited by 1SourceScholar
2024

VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought

NeurIPS 2024spotlight

Large-scale generative language and vision-language models (LLMs and VLMs) excel in few-shot in-context learning for decision making and instruction following. However, they require high-quality exemplar demonstrations to be included in their context window. In this work, we ask: Can LLMs and VLMs g…

Cited by 5SourcePDFScholar
2024

Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models

ICRA 2024poster

Object tracking is central to robot perception and scene understanding, allowing robots to parse a video stream in terms of moving objects with names. Tracking-by-detection has long been a dominant paradigm for object tracking of specific object categories [1], [2]. Recently, large-scale pre-trained…

Cited by 12SourceScholar
2023

Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

CoRL 2023poster

3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning. Many manipulation tasks require high spatial precision in end-effector pose prediction, which typically demands high-resolution 3D feature grids that are computationa…

Cited by 71SourcecodeScholar
2023

Analogy-Forming Transformers for Few-Shot 3D Parsing

ICLR 2023poster

We present Analogical Networks, a model that segments 3D object scenes with analogical reasoning: instead of mapping a scene to part segments directly, our model first retrieves related scenes from memory and their corresponding part structures, and then predicts analogous part structures in the inp…

Cited by 5SourcePDFScholar
2023

Brain Dissection: fMRI-trained Networks Reveal Spatial Selectivity in the Processing of Natural Images

NeurIPS 2023poster

The alignment between deep neural network (DNN) features and cortical responses currently provides the most accurate quantitative explanation for higher visual areas. At the same time, these model features have been critiqued as uninterpretable explanations, trading one black box (the human brain) f…

Cited by 9SourcePDFScholar
2023

ChainedDiffuser: Unifying Trajectory Diffusion and Keypose Prediction for Robotic Manipulation

CoRL 2023poster

We present ChainedDiffuser, a policy architecture that unifies action keypose prediction and trajectory diffusion generation for learning robot manipulation from demonstrations. Our main innovation is to use a global transformer-based action predictor to predict actions at keyframes, a task that req…

Cited by 88SourceScholar
2023

Diffusion-TTA: Test-time Adaptation of Discriminative Models via Generative Feedback

NeurIPS 2023poster

The advancements in generative modeling, particularly the advent of diffusion models, have sparked a fundamental question: how can these models be effectively used for discriminative tasks? In this work, we find that generative models can be great test-time adapters for discriminative models. Our me…

2023

Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement

RSS 2023poster

Language is compositional; an instruction can express multiple relation constraints to hold among objects in a scene that a robot is tasked to rearrange. Our focus in this work is an instructable scene-rearranging framework that generalizes to longer instructions and to spatial concept compositions…

2023

FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation

ICLR 2023top-25%

Humans manipulate various kinds of fluids in their everyday life: creating latte art, scooping floating objects from water, rolling an ice cream cone, etc. Using robots to augment or replace human labors in these daily settings remain as a challenging task due to the multifaceted complexities of flu…

2023

Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language Models

EMNLP 2023long findings

Pre-trained and frozen LLMs can effectively map simple scene re-arrangement instructions to programs over a robot's visuomotor functions through appropriate few-shot example prompting. To parse open-domain natural language and adapt to a user's idiosyncratic procedures, not known during prompt engin…

Cited by 0SourcecodeScholar
2023

Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?

ICRA 2023poster

Building 3D perception systems for autonomous vehicles that do not rely on high-density LiDAR is a critical research problem because of the expense of LiDAR systems compared to cameras and other sensors. Recent research has developed a variety of camera-only methods, where features are differentiabl…

Cited by 143SourceScholar
2023

Test-time Adaptation with Slot-Centric Models

ICML 2023poster

Current visual detectors, though impressive within their training distribution, often fail to parse out-of-distribution scenes into their constituent entities. Recent test-time adaptation methods use auxiliary self-supervised losses to adapt the network parameters to each test example independently…

2022

Bottom Up Top down Detection Transformers for Language Grounding in Images and Point Clouds

ECCV 2022poster

"Most models tasked to ground referential utterances in 2D and 3D scenes learn to select the referred object from a pool of object proposals provided by a pre-trained detector. This is limiting because an utterance may refer to visual entities at various levels of granularity, such as the chair, the…

2022

Particle Video Revisited: Tracking through Occlusions Using Point Trajectories

ECCV 2022poster

"Tracking pixels in videos is typically studied as an optical flow estimation problem, where every pixel is described with a displacement vector that locates it in the next frame. Even though wider temporal context is freely available, prior efforts to take this into account have yielded only small…

2022

Planning with Spatial-Temporal Abstraction from Point Clouds for Deformable Object Manipulation

CoRL 2022poster

Effective planning of long-horizon deformable object manipulation requires suitable abstractions at both the spatial and temporal levels. Previous methods typically either focus on short-horizon tasks or make strong assumptions that full-state information is available, which prevents their use on de…

Cited by 39SourceScholar
2022

TIDEE: Tidying Up Novel Rooms Using Visuo-Semantic Commonsense Priors

ECCV 2022poster

"We introduce TIDEE, an embodied agent that tidies up a disordered scene based on learned commonsense object placement and room arrangement priors. TIDEE explores a home environment, detects objects that are out of their natural place, infers plausible object contexts for them, localizes such contex…

2021

CoCoNets: Continuous Contrastive 3D Scene Representations

CVPR 2021poster

This paper explores self-supervised learning of amodal 3D feature representations from RGB and RGB-D posed images and videos, agnostic to object and scene semantic content, and evaluates the resulting scene representations in the downstream tasks of visual correspondence, object tracking, and object…

Cited by 27PDFScholar
2021

Disentangling 3D Prototypical Networks for Few-Shot Concept Learning

ICLR 2021poster

We present neural architectures that disentangle RGB-D images into objects’ shapes and styles and a map of the background scene, and explore their applications for few-shot 3D object detection and few-shot concept classification. Our networks incorporate architectural biases that reflect the image f…

2021

HyperDynamics: Meta-Learning Object and Agent Dynamics with Hypernetworks

ICLR 2021poster

We propose HyperDynamics, a dynamics meta-learning framework that conditions on an agent’s interactions with the environment and optionally its visual observations, and generates the parameters of neural dynamics models based on inferred properties of the dynamical system. Physical and visual proper…

Cited by 27SourcePDFScholar
2021

Track, Check, Repeat: An EM Approach to Unsupervised Tracking

CVPR 2021poster

We propose an unsupervised method for detecting and tracking moving objects in 3D, in unlabelled RGB-D videos. The method begins with classic handcrafted techniques for segmenting objects using motion cues: we estimate optical flow and camera motion, and conservatively segment regions that appear to…

Cited by 9PDFScholar
2021

Visually-Grounded Library of Behaviors for Manipulating Diverse Objects across Diverse Configurations and Views

CoRL 2021poster

We propose a visually-grounded library of behaviors approach for learning to manipulate diverse objects across varying initial and goal configurations and camera placements. Our key innovation is to disentangle the standard image-to-action mapping into two separate modules that use different types o…

Cited by 1SourceScholar
2020

3D-OES: Viewpoint-Invariant Object-Factorized Environment Simulators

CoRL 2020

We propose an action-conditioned dynamics model that predicts scene changes caused by object and agent interactions in a viewpoint-invariant 3D neural scene representation space, inferred from RGB-D videos. In this 3D feature space, objects do not interfere with one another and their appearance pers

Cited by 0SourcePDFScholar
2020

Embodied Language Grounding With 3D Visual Feature Representations

CVPR 2020poster

We propose associating language utterances to 3D visual abstractions of the scene they describe. The 3D visual abstractions are encoded as 3-dimensional visual feature maps. We infer these 3D visual scene feature maps from RGB images of the scene via view prediction: when the generated 3D scene feat…

Cited by 23PDFScholar
2020

Learning from Unlabelled Videos Using Contrastive Predictive Neural 3D Mapping

ICLR 2020poster

Predictive coding theories suggest that the brain learns by predicting observations at various levels of abstraction. One of the most basic prediction tasks is view prediction: how would a given scene look from an alternative viewpoint? Humans excel at this task. Our ability to imagine and fill in m…

Cited by 30SourcecodeScholar
2020

Tracking Emerges by Looking Around Static Scenes, with Neural 3D Mapping

ECCV 2020poster

with Neural 3D Mapping","We hypothesize that an agent that can look around in static scenes can learn rich visual representations applicable to 3D object tracking in complex dynamic scenes. We are motivated in this pursuit by the fact that the physical world itself is mostly static, and multiview co…

2018

Geometry-Aware Recurrent Neural Networks for Active Visual Recognition

NeurIPS 2018poster

We present recurrent geometry-aware neural networks that integrate visual in- formation across multiple views of a scene into 3D latent feature tensors, while maintaining an one-to-one mapping between 3D physical locations in the world scene and latent feature locations. Object detection, object seg…

Cited by 44SourcePDFScholar
2018

Reinforcement Learning of Active Vision for Manipulating Objects under Occlusions

CoRL 2018

We consider artificial agents that learn to jointly control their gripper and camera in order to reinforcement learn manipulation policies in the presence of occlusions from distractor objects. Distractors often occlude the object of interest and cause it to disappear from the field of view. We prop

2018

Reward Learning From Narrated Demonstrations

CVPR 2018poster

Humans effortlessly “program” one another by communicating goals and desires in natural language. In contrast, humans program robotic behaviours by indicating desired object locations and poses to be achieved [5], by providing RGB images of goal configurations [19], or supplying a demonstration to b…

Cited by 46SourcePDFScholar
2017

Adversarial Inverse Graphics Networks: Learning 2D-To-3D Lifting and Image-To-Image Translation From Unpaired Supervision

ICCV 2017poster

Researchers have developed excellent feed-forward models that learn to map images to desired outputs, such as to the images' latent factors, or to other images, using supervised learning. Learning such mappings from unlabelled data, or improving upon supervised models by exploiting unlabelled data,…

Cited by 170PDFScholar
2017

Self-supervised Learning of Motion Capture

NeurIPS 2017spotlight

Current state-of-the-art solutions for motion capture from a single camera are optimization driven: they optimize the parameters of a 3D human model so that its re-projection matches measurements in the video (e.g. person segmentation, optical flow, keypoint detections etc.). Optimization models are…

Cited by 369SourcePDFScholar
2016

Human Pose Estimation With Iterative Error Feedback

CVPR 2016spotlight

Hierarchical feature extractors such as Convolutional Networks (ConvNets) have achieved impressive performance on a variety of classification tasks using purely feedforward processing. Feedforward architectures can learn rich representations of the input space but do not explicitly model dependencie…

Cited by 1074PDFcodeScholar