← Search

Yuji Matsumoto

15 accepted papers

2026

Docora: A System for Interactive Knowledge Extraction and Visualization from Scientific PDFs

AAAI 2026technical

Scientific research articles, typically distributed in PDF format, contain valuable knowledge but remain challenging to convert into structured datasets due to fragmented workflows that separate parsing, annotation, and visualization. Existing annotation platforms operate on plain text, which requir

Cited by 0SourcePDFScholar
2025

Entity Profile Generation and Reasoning with LLMs for Entity Alignment

EMNLP 2025

Entity alignment (EA) involves identifying and linking equivalent entities across different knowledge graphs (KGs). While knowledge graphs provide structured information about real-world entities, only a small fraction of these entities are aligned. The entity alignment process is challenging due to

Cited by 0SourcePDFScholar
2025

PolyMinder: A Support System for Entity Annotation and Relation Extraction in Polymer Science Documents

COLING 2025system demonstrations

The growing volume of scientific literature in polymer science presents a significant challenge for researchers attempting to extract and annotate domain-specific entities, such as polymer names, material properties, and related information. Manual annotation of these documents is both time-consumin…

2025

Post Persona Alignment for Multi-Session Dialogue Generation

EMNLP 2025

Multi-session persona-based dialogue generation presents challenges in maintaining long-term consistency and generating diverse, personalized responses. While large language models (LLMs) excel in single-session dialogues, they struggle to preserve persona fidelity and conversational coherence acros

Cited by 0SourcePDFScholar
2025

Zero-Shot Entailment Learning for Ontology-Based Biomedical Annotation Without Explicit Mentions

COLING 2025main

Automatic biomedical annotation is essential for advancing medical research, diagnosis, and treatment. However, it presents significant challenges, especially when entities are not explicitly mentioned in the text, leading to difficulties in extraction of relevant information. These challenges are i…

2024

A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages

COLING 2024main

User-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilin…

2024

PolyNERE: A Novel Ontology and Corpus for Named Entity Recognition and Relation Extraction in Polymer Science Domain

COLING 2024main

Polymers are widely used in diverse fields, and the demand for efficient methods to extract and organize information about them is increasing. An automated approach that utilizes machine learning can accurately extract relevant information from scientific papers, providing a promising solution for a…

Cited by 1SourcePDFScholar
2024

Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations

COLING 2024main

Enhancing user engagement through personalization in conversational agents has gained significance, especially with the advent of large language models that generate fluent responses. Personalized dialogue generation, however, is multifaceted and varies in its definition – ranging from instilling a…

Cited by 16SourcePDFScholar
2023

Unsupervised Paraphrasing of Multiword Expressions

ACL 2023findings

We propose an unsupervised approach to paraphrasing multiword expressions (MWEs) in context. Our model employs only monolingual corpus data and pre-trained language models (without fine-tuning), and does not make use of any external resources such as dictionaries. We evaluate our method on the SemEv…

2022

Unsupervised Lexical Substitution with Decontextualised Embeddings

COLING 2022main

We propose a new unsupervised method for lexical substitution using pre-trained language models. Compared to previous approaches that use the generative capability of language models to predict substitutes, our method retrieves substitutes based on the similarity of contextualised and decontextualis…

2021

Nested Named Entity Recognition via Explicitly Excluding the Influence of the Best Path

ACL 2021long

This paper presents a novel method for nested named entity recognition. As a layered method, our method extends the prior second-best path recognition method by explicitly excluding the influence of the best path. Our method maintains a set of hidden states at each time step and selectively leverage…

2020

Coordination Boundary Identification without Labeled Data for Compound Terms Disambiguation

COLING 2020main

We propose a simple method for nominal coordination boundary identification. As the main strength of our method, it can identify the coordination boundaries without training on labeled data, and can be applied even if coordination structure annotations are not available. Our system employs pre-train…

Cited by 4SourcePDFScholar