SearcharxivSearch

arXiv subjects

Yuji Matsumoto

Publications and source records attributed to Yuji Matsumoto.

At least 19 recordsLinked to original sources

Applicability Condition Extraction for Therapeutic Drug-Disease Relations

Identifying conditions that a certain drug takes therapeutic effect on a target disease is crucial for clinical decision-making support. However, most existing biomedical information extraction methods have focused on identifying only relations between drugs and diseases, while largely overlooking the context-specific conditions where such relations can apply. To address this problem, we introduce the task of applicability condition extraction for therapeutic drug-disease relations from biomedical research literature. We create the first dataset that has manually annotated triples of drugs, diseases, and applicability conditions on biomedical paper abstracts with 1,119 drug-disease pairs. Using this dataset, we systematically evaluate the performance of a range of existing methods. In addition, we propose a new method that enhances LoRA to consider relations between drugs and diseases. Our method consistently outperforms strong baselines across different evaluation settings.

cs.AI

PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature

Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stages because of nonspecific clinical presentations. However, to our knowledge, there is no publicly available prion-disease-focused dataset designed to capture a broad range of clinically relevant entities from the biomedical literature. We introduce PrionNER, a manually annotated named entity recognition dataset for prion disease clinical information in PubMed abstracts. The current release comprises 317 abstracts, 2,943 sentences, and 6,955 text-bound entity annotations spanning 15 coarse-grained and 31 fine-grained clinically oriented entity types covering diseases, symptoms, diagnostics, findings, anatomy, treatments, and temporal and statistical evidence. Inter-annotator agreement reaches 81.78 exact-match F1, indicating strong annotation consistency. We benchmark supervised BERT baselines, W2NER, and zero-shot extractors on PrionNER. W2NER is the strongest supervised model, and Gemma-4-31B is the strongest zero-shot model, but the benchmark remains challenging, especially for structurally complex mentions and fine-grained clinically adjacent label distinctions. PrionNER provides a clinically grounded benchmark for prion-disease information extraction and supports research on rare-disease biomedical NLP under low-resource, fine-grained, and non-flat extraction conditions. The dataset, annotation guidelines, and evaluation scripts are available at https://github.com/daotuanan/PrionNER/.

cs.CL

Top-down string-to-dependency Neural Machine Translation

Most of modern neural machine translation (NMT) models are based on an encoder-decoder framework with an attention mechanism. While they perform well on standard datasets, they can have trouble in translation of long inputs that are rare or unseen during training. Incorporating target syntax is one approach to dealing with such length-related problems. We propose a novel syntactic decoder that generates a target-language dependency tree in a top-down, left-to-right order. Experiments show that the proposed top-down string-to-tree decoding generalizes better than conventional sequence-to-sequence decoding in translating long inputs that are not observed in the training data.

cs.CL

The Wisdom of Many Queries: Complexity-Diversity Principle for Dense Retriever Training

Synthetic query generation has become essential for training dense retrievers, yet prior methods generate one query per document, focusing solely on query quality. We are the first to systematically study multi-query synthesis and discover a quality-diversity trade-off: high-quality queries benefit in-domain tasks, while diverse queries benefit out-of-domain (OOD) generalization. Through controlled experiments on 4 benchmark types across Contriever, RetroMAE, and Qwen3-Embedding, we find that diversity benefit strongly correlates with query complexity (r$\geq$0.95, p<0.05), approximated by content words (CW). We formalize this as the Complexity-Diversity Principle (CDP): query complexity determines optimal diversity. Based on CDP, we propose complexity-aware training: multi-query synthesis for high-complexity tasks and CW-weighted training for existing data. Both strategies improve OOD performance on reasoning-intensive benchmarks, with compounded gains when combined.

cs.IR

Better Generalizing to Unseen Concepts: An Evaluation Framework and An LLM-Based Auto-Labeled Pipeline for Biomedical Concept Recognition

Generalization to unseen concepts is a central challenge due to the scarcity of human annotations in Mention-agnostic Biomedical Concept Recognition (MA-BCR). This work makes two key contributions to systematically address this issue. First, we propose an evaluation framework built on hierarchical concept indices and novel metrics to measure generalization. Second, we explore LLM-based Auto-Labeled Data (ALD) as a scalable resource, creating a task-specific pipeline for its generation. Our research unequivocally shows that while LLM-generated ALD cannot fully substitute for manual annotations, it is a valuable resource for improving generalization, successfully providing models with the broader coverage and structural knowledge needed to approach recognizing unseen concepts. Code and datasets are available at https://github.com/bio-ie-tool/hi-ald.

cs.CL

Overview of SCIDOCA 2025 Shared Task on Citation Prediction, Discovery, and Placement

We present an overview of the SCIDOCA 2025 Shared Task, which focuses on citation discovery and prediction in scientific documents. The task is divided into three subtasks: (1) Citation Discovery, where systems must identify relevant references for a given paragraph; (2) Masked Citation Prediction, which requires selecting the correct citation for masked citation slots; and (3) Citation Sentence Prediction, where systems must determine the correct reference for each cited sentence. We release a large-scale dataset constructed from the Semantic Scholar Open Research Corpus (S2ORC), containing over 60,000 annotated paragraphs and a curated reference set. The test set consists of 1,000 paragraphs from distinct papers, each annotated with ground-truth citations and distractor candidates. A total of seven teams registered, with three submitting results. We report performance metrics across all subtasks and analyze the effectiveness of submitted systems. This shared task provides a new benchmark for evaluating citation modeling and encourages future research in scientific document understanding. The dataset and task materials are publicly available at https://github.com/daotuanan/scidoca2025-shared-task.

cs.DL

A Scaling Law for the Orbital Architecture of Planetary Systems Formed by Gravitational Scattering and Collisions

In the standard formation models of terrestrial planets in the solar system and close-in super-Earths in non-resonant orbits recently discovered by exoplanet observations, planets are formed by giant impacts of protoplanets or planetary embryos after the dispersal of protoplanetary disk gas in the final stage. This study aims to theoretically clarify a fundamental scaling law for the orbital architecture of planetary systems formed by giant impacts. In the giant impact stage, protoplanets gravitationally scatter and collide with one another to form planets. Using {\em N}-body simulations, we investigate the orbital architecture of planetary systems formed from protoplanet systems by giant impacts. As the orbital architecture parameters, we focus on the mean orbital separation between two adjacent planets and the mean orbital eccentricity of planets in a planetary system. We find that the orbital architecture is determined by the ratio of the two-body surface escape velocity of planets $v_\mathrm{esc}$ to the Keplerian circular velocity $v_\mathrm{K}$, $k$ = The mean orbital separation and eccentricity are about $2 ka$ and $0.3 k$, respectively, where $a$ is the system semimajor axis. With this scaling, the orbital architecture parameters of planetary systems are nearly independent of their total mass and semimajor axis.

astro-ph.EP

Semi-analytical model for the dynamical evolution of planetary system II: Application to systems formed by a planet formation model

The standard formation model of close-in low-mass planets involves efficient inward migration followed by growth through giant impacts after the protoplanetary gas disk disperses. While detailed N-body simulations have enhanced our understanding, their high computational cost limits statistical comparisons with observations. In our previous work, we introduced a semi-analytical model to track the dynamical evolution of multiple planets through gravitational scattering and giant impacts after the gas disk dispersal. Although this model successfully reproduced N -body simulation results under various initial conditions, our validation was still limited to cases with compact, equally-spaced planetary systems. In this paper, we improve our model to handle more diverse planetary systems characterized by broader variations in planetary masses, semi-major axes, and orbital separations and validate it against recent planet population synthesis results. Our enhanced model accurately reproduces the mass distribution and orbital architectures of the final planetary systems. Thus, we confirm that the model can predict the outcomes of post-gas disk dynamical evolution across a wide range of planetary system architectures, which is crucial for reducing the computational cost of planet formation simulations.

astro-ph.EP

Post Persona Alignment for Multi-Session Dialogue Generation

Multi-session persona-based dialogue generation presents challenges in maintaining long-term consistency and generating diverse, personalized responses. While large language models (LLMs) excel in single-session dialogues, they struggle to preserve persona fidelity and conversational coherence across extended interactions. Existing methods typically retrieve persona information before response generation, which can constrain diversity and result in generic outputs. We propose Post Persona Alignment (PPA), a novel two-stage framework that reverses this process. PPA first generates a general response based solely on dialogue context, then retrieves relevant persona memories using the response as a query, and finally refines the response to align with the speaker's persona. This post-hoc alignment strategy promotes naturalness and diversity while preserving consistency and personalization. Experiments on multi-session LLM-generated dialogue data demonstrate that PPA significantly outperforms prior approaches in consistency, diversity, and persona relevance, offering a more flexible and effective paradigm for long-term personalized dialogue generation.

cs.CL

Semi-analytical model for the dynamical evolution of planetary systems via giant impacts

In the standard model of terrestrial planet formation, planets are formed through giant impacts of planetary embryos after the dispersal of the protoplanetary gas disc. Traditionally, $N$-body simulations have been used to investigate this process. However, they are computationally too expensive to generate sufficient planetary populations for statistical comparisons with observational data. A previous study introduced a semi-analytical model that incorporates the orbital and accretionary evolution of planets due to giant impacts and gravitational scattering. This model succeeded in reproducing the statistical features of planets in $N$-body simulations near 1 au around solar-mass stars. However, this model is not applicable to close-in regions (around 0.1 au) or low-mass stars because the dynamical evolution of planetary systems depends on the orbital radius and stellar mass. This study presents a new semi-analytical model applicable to close-in orbits around stars of various masses, validated through comparison with $N$-body simulations. The model accurately predicts the final distributions of planetary mass, semi-major axis, and eccentricity for the wide ranges of orbital radius, initial planetary mass, and stellar mass, with significantly reduced computation time compared to $N$-body simulations. By integrating this model with other planet-forming processes, a computationally low-cost planetary population synthesis model can be developed.

astro-ph.EP

MA-COIR: Leveraging Semantic Search Index and Generative Models for Ontology-Driven Biomedical Concept Recognition

Recognizing biomedical concepts in the text is vital for ontology refinement, knowledge graph construction, and concept relationship discovery. However, traditional concept recognition methods, relying on explicit mention identification, often fail to capture complex concepts not explicitly stated in the text. To overcome this limitation, we introduce MA-COIR, a framework that reformulates concept recognition as an indexing-recognition task. By assigning semantic search indexes (ssIDs) to concepts, MA-COIR resolves ambiguities in ontology entries and enhances recognition efficiency. Using a pretrained BART-based model fine-tuned on small datasets, our approach reduces computational requirements to facilitate adoption by domain experts. Furthermore, we incorporate large language models (LLMs)-generated queries and synthetic data to improve recognition in low-resource settings. Experimental results on three scenarios (CDR, HPO, and HOIP) highlight the effectiveness of MA-COIR in recognizing both explicit and implicit concepts without the need for mention-level annotations during inference, advancing ontology-driven concept recognition in biomedical domain applications. Our code and constructed data are available at https://github.com/sl-633/macoir-master.

cs.CL

Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations

Enhancing user engagement through personalization in conversational agents has gained significance, especially with the advent of large language models that generate fluent responses. Personalized dialogue generation, however, is multifaceted and varies in its definition -- ranging from instilling a persona in the agent to capturing users' explicit and implicit cues. This paper seeks to systemically survey the recent landscape of personalized dialogue generation, including the datasets employed, methodologies developed, and evaluation metrics applied. Covering 22 datasets, we highlight benchmark datasets and newer ones enriched with additional features. We further analyze 17 seminal works from top conferences between 2021-2023 and identify five distinct types of problems. We also shed light on recent progress by LLMs in personalized dialogue generation. Our evaluation section offers a comprehensive summary of assessment facets and metrics utilized in these works. In conclusion, we discuss prevailing challenges and envision prospect directions for future research in personalized dialogue generation.

cs.CL

Chondrule Destruction via Dust Collisions in Shock Waves

A leading candidate for the heating source of chondrules and igneous rims is shock waves. This mechanism generates high relative velocities between chondrules and dust particles. We have investigated the possibility of the chondrule destruction in collisions with dust particles behind a shock wave using a semianalytical treatment. We find that the chondrules are destroyed during melting in collisions. We derive the conditions for the destruction of chondrules and show that the typical size of the observed chondrules satisfies the condition. We suggest that the chondrule formation and rim accretion are different events if they are heated by shock waves.

astro-ph.EP

A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages

User-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilingual corpus of texts concerning ADRs gathered from diverse sources, including patient fora, social media, and clinical reports in German, French, and Japanese. Our corpus contains annotations covering 12 entity types, four attribute types, and 13 relation types. It contributes to the development of real-world multilingual language models for healthcare. We provide statistics to highlight certain challenges associated with the corpus and conduct preliminary experiments resulting in strong baselines for extracting entities and relations between these entities, both within and across languages.

cs.CL

Magnetic Order in Honeycomb Layered U$_2$Pt$_6$Ga$_{15}$ Studied by Resonant X-ray and Neutron Scatterings

Antiferromagnetic (AF) order of U$_{2}$Pt$_{6}$Ga$_{15}$ with the ordering temperature $T_{\rm N}$ = 26 K was investigated by resonant X-ray scattering and neutron diffraction on single crystals. This compound possesses a unique crystal structure in which uranium ions form honeycomb layers and then stacks along the $c$-axis with slight offset, which gives rise to a stacking disorder. The AF order can be described with the propagation vector of $q = (1/6, 1/6, 0)$ in the hexagonal notation. The ordered magnetic moments orient perpendicular to the honeycomb layers, indicating a collinear spin structure consistent with Ising-like anisotropy. The magnetic reflections are found to be broadened along $c^*$ indicating that the stacking disorder results in anisotropic correlation lengths. The semi-quantitative analysis of neutron diffraction intensity, combined with group theory considerations based on the crystallographic symmetry, suggests a zig-zag type magnetic structure for the AF ground state, in which the AF coupling runs perpendicular to the stacking offset, characterized as $q = (1, 0, 0)_{\rm orth}$. The realization of the zig-zag magnetic structure implies the presence of frustrating intralayer exchange interactions involving both ferromagnetic (FM) first-neighbor and AF second and third-neighbor interactions in this compound.

cond-mat.str-el

Unsupervised Paraphrasing of Multiword Expressions

We propose an unsupervised approach to paraphrasing multiword expressions (MWEs) in context. Our model employs only monolingual corpus data and pre-trained language models (without fine-tuning), and does not make use of any external resources such as dictionaries. We evaluate our method on the SemEval 2022 idiomatic semantic text similarity task, and show that it outperforms all unsupervised systems and rivals supervised systems.

cs.CL

Igneous Rim Accretion on Chondrules in Low-Velocity Shock Waves

Shock wave heating is a leading candidate for the mechanisms of chondrule formation. This mechanism forms chondrules when the shock velocity is in a certain range. If the shock velocity is lower than this range, dust particles smaller than chondrule precursors melt, while chondrule precursors do not. We focus on the low-velocity shock waves as the igneous rim accretion events. Using a semi-analytical treatment of the shock-wave heating model, we found that the accretion of molten dust particles occurs when they are supercooling. The accreted igneous rims have two layers, which are the layers of the accreted supercooled droplets and crystallized dust particles. We suggest that chondrules experience multiple rim-forming shock events.

astro-ph.EP

Is In-hospital Meta-information Useful for Abstractive Discharge Summary Generation?

During the patient's hospitalization, the physician must record daily observations of the patient and summarize them into a brief document called "discharge summary" when the patient is discharged. Automated generation of discharge summary can greatly relieve the physicians' burden, and has been addressed recently in the research community. Most previous studies of discharge summary generation using the sequence-to-sequence architecture focus on only inpatient notes for input. However, electric health records (EHR) also have rich structured metadata (e.g., hospital, physician, disease, length of stay, etc.) that might be useful. This paper investigates the effectiveness of medical meta-information for summarization tasks. We obtain four types of meta-information from the EHR systems and encode each meta-information into a sequence-to-sequence model. Using Japanese EHRs, meta-information encoded models increased ROUGE-1 by up to 4.45 points and BERTScore by 3.77 points over the vanilla Longformer. Also, we found that the encoded meta-information improves the precisions of its related terms in the outputs. Our results showed the benefit of the use of medical meta-information.

cs.CL