← Search

YUSUKE Mukuta

14 accepted papers

2026

Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets

ICLR 2026poster

Scalable robot policy pre-training has been hindered by the high cost of collecting high-quality demonstrations for each platform. In this study, we address this issue by uniting offline reinforcement learning (offline RL) with cross-embodiment learning. Offline RL leverages both expert and abundant…

Cited by 0SourceScholar
2025

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning

ICML 2025poster

For continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality operator, RL algorithms for continuous actions typically model Q…

2025

Intend to Move: A Multimodal Dataset for Intention-Aware Human Motion Understanding

NeurIPS 2025poster

Human motion is inherently intentional, yet most motion modeling paradigms focus on low-level kinematics, overlooking the semantic and causal factors that drive behavior. Existing datasets further limit progress: they capture short, decontextualized actions in static scenes, providing little groundi…

Cited by 0SourceScholar
2024

Content-Specific Humorous Image Captioning Using Incongruity Resolution Chain-of-Thought

NAACL 2024findings

Although automated image captioning methods have benefited considerably from the development of large language models (LLMs), generating humorous captions is still a challenging task. Humorous captions generated by humans are unique to the image and reflect the content of the image. However, caption…

2024

Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning

AAAI 2024technical

In deep reinforcement learning, estimating the value function to evaluate the quality of states and actions is essential. The value function is often trained using the least squares method, which implicitly assumes a Gaussian error distribution. However, a recent study suggested that the error distr…

Cited by 0SourcePDFScholar
2022

Unsupervised Pose-Aware Part Decomposition for Man-Made Articulated Objects

ECCV 2022poster

"Man-made articulated objects exist widely in the real world. However, previous methods for unsupervised part decomposition are unsuitable for such objects because they assume a spatially fixed part location, resulting in inconsistent part parsing. In this paper, we propose PPD (unsupervised Pose-aw…

Cited by 13SourcePDFScholar
2021

Real-Time Mesh Extraction from Implicit Functions via Direct Reconstruction of Decision Boundary

ICRA 2021poster

The ability to estimate 3D object shape from a single image is vital to robotics and manufacturing. For instance, it enables iterative trial-and-error in simulated environments. In single-view reconstruction, implicit functions have demonstrated superior results over traditional methods. However, im…

Cited by 0SourceScholar
2021

Spherical Image Generation from a Single Image by Considering Scene Symmetry

AAAI 2021technical

Spherical images taken in all directions (360 degrees by 180 degrees) allow the full surroundings of a subject to be represented, providing an immersive experience to viewers. Generating a spherical image from a single normal-field-of-view (NFOV) image is convenient and expands the usage scenarios c…

Cited by 19SourcePDFScholar
2019

Multi-Stage Pathological Image Classification Using Semantic Segmentation

ICCV 2019accepted

Histopathological image analysis is an essential process for the discovery of diseases such as cancer. However, it is challenging to train CNN on whole slide images (WSIs) of gigapixel resolution considering the available memory capacity. Most of the previous works divide high resolution WSIs into s…

Cited by 54SourcePDFScholar
2015

Common Subspace for Model and Similarity: Phrase Learning for Caption Generation From Images

ICCV 2015poster

Generating captions to describe images is a fundamental problem that combines computer vision and natural language processing. Recent works focus on descriptive phrases, such as "a white dog" to explain the visual composites of an input image. The phrases can not only express objects, attributes, ev…

Cited by 81PDFcodeScholar