← Search

Yoshimitsu Aoki

16 accepted papers

2026

Geometric-Photometric Event-based 3D Gaussian Ray Tracing

CVPR 2026

Event cameras offer a high temporal resolution over traditional frame-based cameras, which makes them suitable for motion and structure estimation. However, it has been unclear how event-based 3D Gaussian Splatting (3DGS) approaches could leverage fine-grained temporal information of sparse events.

Cited by 0SourcecodeScholar
2026

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning

CVPR 2026

Humanoid robotics has strong potential to transform daily service and caregiving applications. Although recent advances in general motion tracking within physics engines (GMT) have enabled virtual characters and humanoid robots to reproduce a broad range of human motions, these behaviors are primari

Cited by 0SourceScholar
2025

Formula-Supervised Sound Event Detection: Pre-Training Without Real Data

ICASSP 2025accepted

In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effec…

Cited by 0SourceScholar
2025

Simultaneous Motion And Noise Estimation with Event Cameras

ICCV 2025poster

Event cameras are emerging vision sensors whose noise is challenging to characterize. Existing denoising methods for event cameras are often designed in isolation and thus consider other tasks, such as motion estimation, separately (i.e., sequentially after denoising). However, motion is an intrinsi…

2024

Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor Navigation

ICRA 2024poster

In outdoor environments, Vision-and-Language Navigation (VLN) requires an agent to rely on multi-modal cues from real-world urban environments and natural language instructions. While existing outdoor VLN models predict actions using a combination of panorama and instruction features, this approach…

Cited by 1SourceScholar
2024

PCT: Perspective Cue Training Framework for Multi-Camera BEV Segmentation

IROS 2024

Generating annotations for bird’s-eye-view (BEV) segmentation presents significant challenges due to the scenes’ complexity and the high manual annotation cost. In this work, we address these challenges by leveraging the abundance of unlabeled data available. We propose the Perspective Cue Training

Cited by 5SourceScholar
2024

Rethinking Image Super Resolution from Training Data Perspectives

ECCV 2024poster

"In this work, we investigate the understudied effect of the training data used for image super-resolution (SR). Most commonly, novel SR methods are developed and benchmarked on common training datasets such as DIV2K and DF2K. However, we investigate and rethink the training data from the perspectiv…

2023

Listening Human Behavior: 3D Human Pose Estimation With Acoustic Signals

CVPR 2023poster

Given only acoustic signals without any high-level information, such as voices or sounds of scenes/actions, how much can we infer about the behavior of humans? Unlike existing methods, which suffer from privacy issues because they use signals that include human speech or the sounds of specific actio…

Cited by 19SourcePDFScholar
2022

Diverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background Creation

CVPR 2022poster

We address the problem of generating a 360-degree image from a single image with a narrow field of view by estimating its surroundings. Previous methods suffered from overfitting to the training resolution and deterministic generation. This paper proposes a completion method using a transformer for…

Cited by 36PDFcodeScholar
2020

Joint Pedestrian Detection and Risk-level Prediction with Motion-Representation-by-Detection

ICRA 2020poster

The paper presents a pedestrian near-miss detector with temporal analysis that provides both pedestrian detection and risk-level predictions which are demonstrated on a self-collected database. Our work makes three primary contributions: (i) The framework of pedestrian near-miss detection is propose…

Cited by 6SourceScholar
2018

Anticipating Traffic Accidents With Adaptive Loss and Large-Scale Incident DB

CVPR 2018poster

In this paper, we propose a novel approach for traffic accident anticipation through (i) Adaptive Loss for Early Anticipation (AdaLEA) and (ii) a large-scale self-annotated incident database. The proposed AdaLEA allows us to gradually learn an earlier anticipation as training progresses. The loss fu…

Cited by 157SourcePDFScholar