← Search

Daniel Yang

12 accepted papers

2025

SeaSplat: Representing Underwater Scenes with 3D Gaussian Splatting and a Physically Grounded Image Formation Model

ICRA 2025

We introduce SeaSplat, a method to enable real-time rendering of underwater scenes leveraging recent advances in 3D radiance fields. Underwater scenes are challenging visual environments, as rendering through a medium such as water introduces both range and color dependent effects on image capture.

Cited by 38SourcecodeScholar
2025

TrajFlow: Multi-modal Motion Prediction via Flow Matching

IROS 2025

Efficient and accurate motion prediction is crucial for ensuring safety and informed decision-making in autonomous driving, particularly under dynamic real-world conditions that necessitate multi-modal forecasts. We introduce TrajFlow, a novel flow matching-based motion prediction framework that add

Cited by 5SourcecodeScholar
2024

Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization

ICASSP 2024accepted

Video summarization remains a huge challenge in computer vision due to the size of the input videos to be summarized. We propose an efficient, language-only video summarizer that achieves competitive accuracy with high data efficiency. Using only textual captions obtained via a zero-shot approach, w…

Cited by 0SourceScholar
2024

Probabilistic Feature Matching for Fast Scalable Visual Prompting

IJCAI 2024poster

In this work, we propose a novel framework for image segmentation guided by visual prompting which leverages the power of vision foundation models. Inspired by recent advancements in computer vision, our approach integrates multiple large-scale pretrained models to address the challenges of segment…

Cited by 1SourcePDFScholar
2024

Rank2Reward: Learning Shaped Reward Functions from Passive Video

ICRA 2024poster

Teaching robots novel skills with demonstrations via human-in-the-loop data collection techniques like kinesthetic teaching or teleoperation puts a heavy burden on human supervisors. In contrast to this paradigm, it is often significantly easier to provide raw, action-free visual data of tasks being…

Cited by 5SourcecodeScholar
2021

Robotic Grasping through Combined Image-Based Grasp Proposal and 3D Reconstruction

ICRA 2021poster

We present a novel approach to robotic grasp planning using both a learned grasp proposal network and a learned 3D shape reconstruction network. Our system generates 6-DOF grasps from a single RGB-D image of the target object, which is provided as input to both networks. By using the geometric recon…

Cited by 52SourceScholar
2020

Reward Prediction Error as an Exploration Objective in Deep RL

IJCAI 2020poster

A major challenge in reinforcement learning is exploration, when local dithering methods such as epsilon-greedy sampling are insufficient to solve a given task. Many recent methods have proposed to intrinsically motivate an agent to seek novel states, driving the agent to discover improved reward. H…

Cited by 0SourcePDFScholar
2019

Modeling and state estimation of a Micro Ball-balancing Robot using a high yaw-rate dynamic model and an Extended Kalman Filter

ICRA 2019poster

The state estimation and control of a ball-balancing robot under high yaw rate is a challenging problem due to its highly nonlinear 3D dynamic. The small size and low-cost components in our Micro Ball-Balancing Robot makes the system inherently very noisy which further increases the complexity of th…

Cited by 0SourceScholar
2018

A minimalist Stair Climbing Robot (SCR) formed as a leg balancing & climbing Mobile Inverted Pendulum (MIP)

IROS 2018poster

This paper presents a (patent-pending) small, quasi-static, minimal-complexity Stair Climbing Robot (SCR). The vehicle design is given simply by adding a third motor to a (Segway-like) Mobile Inverted Pendulum (MIP), enabling it to maneuver up stairs, leveraging feedback control, by planting it's “f…

Cited by 12SourceScholar
2018

Demo2Vec: Reasoning Object Affordances From Online Videos

CVPR 2018poster

Watching expert demonstrations is an important way for humans and robots to reason about affordances of unseen objects. In this paper, we consider the problem of reasoning object affordances through the feature embedding of demonstration videos. We design the Demo2Vec model which learns to extract e…

Cited by 132SourcePDFScholar
2015

Design and control of a micro ball-balancing robot (MBBR) with orthogonal midlatitude omniwheel placement

IROS 2015poster

Ball-balancing robots (BBRs) are endowed with rich dynamics. When properly designed and stabilized via feedback to eliminate jitter, and intuitively coordinated with a well-designed smartphone interface, BBRs exhibit a uniquely fluid and organic motion. Unlike mobile inverted pendulums (MIPs, akin t…

Cited by 11SourceScholar