← Search

Osamu Yoshie

15 accepted papers

2026

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

ICRA 2026poster

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT ) tend to treat multi-view features equally and directly concatenate them for policy learning. How ever, it will introduce redundant visual information and bring hig…

2025

ARM : nnU-Net with Arena Mechanism for Medical Image Segmentation

ICASSP 2025accepted

The success of nnU-Net proves the significance of the rationality of workflow architecture and configuration settings in improving segmentation accuracy. However, since that, most efforts to improve U-Net have continued to address CNN inner limitations caused by architecture. These methods encounter…

Cited by 0SourceScholar
2025

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

RA-L 2025

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT [1]) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it will introduce redundant visual information and bring h

Cited by 8SourceScholar
2025

FinLLM-B: When Large Language Models Meet Financial Breakout Trading

NAACL 2025industry

Trading range breakout is a key method in the technical analysis of financial trading, widely employed by traders in financial markets such as stocks, futures, and foreign exchange. However, distinguishing between true and false breakout and providing the correct rationale cause significant challeng…

Cited by 0SourcePDFScholar
2025

GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning

NeurIPS 2025poster

With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary,…

Cited by 0SourceScholar
2025

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

EMNLP 2025

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting

Cited by 0SourcePDFScholar
2025

MetaCert: Metabolic Attention Network Utilizing Uncertainty Estimation for Multimodal Aspect-Category-Sentiment Triple Extraction

ICASSP 2025accepted

Multimodal Aspect-Category-Sentiment Triple Extraction (MACSTE) is a highly complex subtask within Multimodal Aspect-Based Sentiment Analysis (MABSA), requiring simultaneous attribute extraction and sentiment polarity prediction from image-text pairs. While existing research often emphasizes modalit…

Cited by 0SourceScholar
2025

Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn Guidance

EMNLP 2025

The emergence of text-to-image synthesis (TIS) models has significantly influenced digital image creation by producing high-quality visuals from written descriptions. Yet these models are sensitive on textual prompts, posing a challenge for novice users who may not be familiar with TIS prompt writin

2023

A Simple Framework for Text-Supervised Semantic Segmentation

CVPR 2023poster

Text-supervised semantic segmentation is a novel research topic that allows semantic segments to emerge with image-text contrasting. However, pioneering methods could be subject to specifically designed network architectures. This paper shows that a vanilla contrastive language-image pre-training (C…

2022

Contrastive Vision-Language Pre-training with Limited Resources

ECCV 2022poster

"Pioneering dual-encoder pre-training works (e.g., CLIP and ALIGN) have revealed the potential of aligning multi-modal representations with contrastive learning. However, these works require a tremendous amount of data and computational resources (e.g., billion-level web data and hundreds of GPUs),…

2022

Discriminability-Transferability Trade-Off: An Information-Theoretic Perspective

ECCV 2022poster

"This work simultaneously considers the discriminability and transferability properties of deep representations in the typical supervised learning task, i.e., image classification. By a comprehensive temporal analysis, we observe a trade-off between these two properties. The discriminability keeps i…

2020

ExchNet: A Unified Hashing Network for Large-Scale Fine-Grained Image Retrieval

ECCV 2020poster

Retrieving content relevant images from a large-scale fine-grained dataset could suffer from intolerably slow query speed and highly redundant storage cost, due to high-dimensional real-valued embeddings which aim to distinguish subtle visual differences of fine-grained objects. In this paper, we st…

Cited by 52SourcePDFScholar
2020

NMS by Representative Region: Towards Crowded Pedestrian Detection by Proposal Pairing

CVPR 2020poster

Although significant progress has been made in pedestrian detection recently, pedestrian detection in crowded scenes is still challenging. The heavy occlusion between pedestrians imposes great challenges to the standard Non-Maximum Suppression (NMS). A relative low threshold of intersection over uni…

Cited by 202PDFScholar