Searcharxiv⌕ Search

arXiv subjects

Hiroaki Sugiyama

Publications and source records attributed to Hiroaki Sugiyama.

At least 19 recordsLinked to original sources

Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models

Recent studies have shown that Theory of Mind (ToM) in large language models (LLMs) has not reached human-level performance yet. Since fine-tuning LLMs on ToM datasets often degrades their generalization, several inference-time methods have been proposed to enhance ToM in LLMs. However, existing inference-time methods for ToM are specialized for inferring beliefs from contexts involving changes in the world state. In this study, we present a new inference-time method for ToM, Shoes-of-Others (SoO) prefilling, which makes fewer assumptions about contexts and is applicable to broader scenarios. SoO prefilling simply specifies the beginning of LLM outputs with ``Let's put ourselves in A's shoes.'', where A denotes the target character's name. We evaluate SoO prefilling on two benchmarks that assess ToM in conversational and narrative contexts without changes in the world state and find that it consistently improves ToM across five categories of mental states. Our analysis suggests that SoO prefilling elicits faithful thoughts, thereby improving the ToM performance.

cs.CL↗

Enhancing Impression Change Prediction in Speed Dating Simulations Based on Speakers' Personalities

This paper focuses on simulating text dialogues in which impressions between speakers improve during speed dating. This simulation involves selecting an utterance from multiple candidates generated by a text generation model that replicates a specific speaker's utterances, aiming to improve the impression of the speaker. Accurately selecting an utterance that improves the impression is crucial for the simulation. We believe that whether an utterance improves a dialogue partner's impression of the speaker may depend on the personalities of both parties. However, recent methods for utterance selection do not consider the impression per utterance or the personalities. To address this, we propose a method that predicts whether an utterance improves a partner's impression of the speaker, considering the personalities. The evaluation results showed that personalities are useful in predicting impression changes per utterance. Furthermore, we conducted a human evaluation of simulated dialogues using our method. The results showed that it could simulate dialogues more favorably received than those selected without considering personalities.

cs.CL↗

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression

To improve user engagement during conversations with dialogue systems, we must improve individual dialogue responses and dialogue impressions such as consistency, personality, and empathy throughout the entire dialogue. While such dialogue systems have been developing rapidly with the help of large language models (LLMs), reinforcement learning from AI feedback (RLAIF) has attracted attention to align LLM-based dialogue models for such dialogue impressions. In RLAIF, a reward model based on another LLM is used to create a training signal for an LLM-based dialogue model using zero-shot/few-shot prompting techniques. However, evaluating an entire dialogue only by prompting LLMs is challenging. In this study, the supervised fine-tuning (SFT) of LLMs prepared reward models corresponding to 12 metrics related to the impression of the entire dialogue for evaluating dialogue responses. We tuned our dialogue models using the reward model signals as feedback to improve the impression of the system. The results of automatic and human evaluations showed that tuning the dialogue model using our reward model corresponding to dialogue impression improved the evaluation of individual metrics and the naturalness of the dialogue response.

cs.CL↗

ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind

Existing Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challenges, we introduce ToMATO, a new ToM benchmark formulated as multiple-choice QA over conversations. ToMATO is generated via LLM-LLM conversations featuring information asymmetry. By employing a prompting method that requires role-playing LLMs to verbalize their thoughts before each utterance, we capture both first- and second-order mental states across five categories: belief, intention, desire, emotion, and knowledge. These verbalized thoughts serve as answers to questions designed to assess the mental states of characters within conversations. Furthermore, the information asymmetry introduced by hiding thoughts from others induces the generation of false beliefs about various mental states. Assigning distinct personality traits to LLMs further diversifies both utterances and thoughts. ToMATO consists of 5.4k questions, 753 conversations, and 15 personality trait patterns. Our analysis shows that this dataset construction approach frequently generates false beliefs due to the information asymmetry between role-playing LLMs, and effectively reflects diverse personalities. We evaluate nine LLMs on ToMATO and find that even GPT-4o mini lags behind human performance, especially in understanding false beliefs, and lacks robustness to various personality traits.

cs.CL↗

LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs

This paper introduces LLM-jp, a cross-organizational project for the research and development of Japanese large language models (LLMs). LLM-jp aims to develop open-source and strong Japanese LLMs, and as of this writing, more than 1,500 participants from academia and industry are working together for this purpose. This paper presents the background of the establishment of LLM-jp, summaries of its activities, and technical reports on the LLMs developed by LLM-jp. For the latest activities, visit https://llm-jp.nii.ac.jp/en/.

cs.CL↗

User-Specific Dialogue Generation with User Profile-Aware Pre-Training Model and Parameter-Efficient Fine-Tuning

This paper addresses user-specific dialogs. In contrast to previous research on personalized dialogue focused on achieving virtual user dialogue as defined by persona descriptions, user-specific dialogue aims to reproduce real-user dialogue beyond persona-based dialogue. Fine-tuning using the target user's dialogue history is an efficient learning method for a user-specific model. However, it is prone to overfitting and model destruction due to the small amount of data. Therefore, we propose a learning method for user-specific models by combining parameter-efficient fine-tuning with a pre-trained dialogue model that includes user profiles. Parameter-efficient fine-tuning adds a small number of parameters to the entire model, so even small amounts of training data can be trained efficiently and are robust to model destruction. In addition, the pre-trained model, which is learned by adding simple prompts for automatically inferred user profiles, can generate speech with enhanced knowledge of the user's profile, even when there is little training data during fine-tuning. In experiments, we compared the proposed model with large-language-model utterance generation using prompts containing users' personal information. Experiments reproducing real users' utterances revealed that the proposed model can generate utterances with higher reproducibility than the compared methods, even with a small model.

cs.CL↗

Bipartite-play Dialogue Collection for Practical Automatic Evaluation of Dialogue Systems

Automation of dialogue system evaluation is a driving force for the efficient development of dialogue systems. This paper introduces the bipartite-play method, a dialogue collection method for automating dialogue system evaluation. It addresses the limitations of existing dialogue collection methods: (i) inability to compare with systems that are not publicly available, and (ii) vulnerability to cheating by intentionally selecting systems to be compared. Experimental results show that the automatic evaluation using the bipartite-play method mitigates these two drawbacks and correlates as strongly with human subjectivity as existing methods.

cs.CL↗

Spoken Dialogue Strategy Focusing on Asymmetric Communication with Android Robots

Humans are easily conscious of small differences in an android robot's (AR's) behaviors and utterances, resulting in treating the AR as not-human, while ARs treat us as humans. Thus, there exists asymmetric communication between ARs and humans. In our system at Dialogue Robot Competition 2022, this asymmetry was a considerable research target in our dialogue strategy. For example, tricky phrases such as questions related to personal matters and forceful requests for agreement were experimentally used in AR's utterances. We assumed that these AR phrases would have a reasonable chance of success, although humans would likely hesitate to use the phrases. Additionally, during a five-minute dialogue, our AR's character, such as its voice tones and sentence expressions, changed from mechanical to human-like type in order to pretend to tailor to customers. The characteristics of the AR developed by our team, DSML-TDU, are introduced in this paper.

cs.RO↗

Empirical Analysis of Training Strategies of Transformer-based Japanese Chit-chat Systems

In recent years, several high-performance conversational systems have been proposed based on the Transformer encoder-decoder model. Although previous studies analyzed the effects of the model parameters and the decoding method on subjective dialogue evaluations with overall metrics, they did not analyze how the differences of fine-tuning datasets affect on user's detailed impression. In addition, the Transformer-based approach has only been verified for English, not for such languages with large inter-language distances as Japanese. In this study, we develop large-scale Transformer-based Japanese dialogue models and Japanese chit-chat datasets to examine the effectiveness of the Transformer-based approach for building chit-chat dialogue systems. We evaluated and analyzed the impressions of human dialogues in different fine-tuning datasets, model parameters, and the use of additional information.

cs.CL↗

New model for radiatively generated Dirac neutrino masses and lepton flavor violating decays of the Higgs boson

We propose a new mechanism to explain neutrino masses with lepton number conservation, in which the Dirac neutrino masses are generated at the two-loop level involving a dark matter candidate. In this model, branching ratios of lepton flavor violating decays of the Higgs boson can be much larger than those of lepton flavor violating decays of charged leptons. If lepton flavor violating decays of the Higgs boson are observed at future collider experiments without detecting lepton flavor violating decays of charged leptons, most of the models previously proposed for tiny neutrino masses are excluded while our model can still survive. We show that the model can be viable under constraints from current data for neutrino experiments, searches for lepton flavor violating decays of charged leptons and dark matter experiments.

hep-ph↗

Neutrino Mass, Dark Matter and Baryon Asymmetry without Lepton Number Violation

We propose a model to explain tiny masses of neutrinos with the lepton number conservation, where neither too heavy particles beyond the TeV-scale nor tiny coupling constants are required. Assignments of conserving lepton numbers to new fields result in an unbroken $Z_2$ symmetry that stabilizes the dark matter candidate (the lightest $Z_2$-odd particle). In this model, $Z_2$-odd particles play an important role to generate the mass of neutrinos. The scalar dark matter in our model can satisfy constraints on the dark matter abundance and those from direct searches. It is also shown that the strong first-order phase transition, which is required for the electroweak baryogenesis, can be realized in our model. In addition, the scalar potential can in principle contain CP-violating phases, which can also be utilized for the baryogenesis. Therefore, three problems in the standard model, namely absence of neutrino masses, the dark matter candidate, and the mechanism to generate baryon asymmetry of the Universe, may be simultaneously resolved at the TeV-scale. Phenomenology of this model is also discussed briefly.

hep-ph↗

Testing neutrino mass generation mechanisms from the lepton flavor violating decay of the Higgs boson

We investigate how observations of the lepton flavor violating decay of the Higgs boson ($h \to \ell\ell^\prime$) can narrow down models of neutrino mass generation mechanisms, which were systematically studied in Refs. [1,2] by focusing on the combination of new Yukawa coupling matrices with leptons. We find that a wide class of models for neutrino masses can be excluded if evidence for $h \to \ell\ell^\prime$ is really obtained in the current or future collider experiments. In particular, simple models of Majorana neutrino masses cannot be compatible with the observation of $h \to \ell\ell^\prime$. It is also found that some of the simple models to generate masses of Dirac neutrinos radiatively can be compatible with a significant rate of the $h \to \ell\ell^\prime$ process.

hep-ph↗

Probing Models of Dirac Neutrino Masses via the Flavor Structure of the Mass Matrix

We classify models of the Dirac neutrino mass by concentrating on flavor structures of the mass matrix. The advantage of our classification is that we do not need to specify detail of models except for Yukawa interactions because flavor structures can be given only by products of Yukawa matrices. All possible Yukawa interactions between leptons (including the right-handed neutrino) are taken into account by introducing appropriate scalar fields. We also take into account the case with Yukawa interactions of leptons with the dark matter candidate. Then, we see that flavor structures can be classified into seven groups. The result is useful for the efficient test of models of the neutrino mass. One of seven groups can be tested by measuring the absolute neutrino mass. Other two can be tested by probing the violation of the lepton universality in $\ell \to \ell^\prime ν\overlineν$. In order to test the other four groups, we can rely on searches for new scalar particles at collider experiments.

hep-ph↗

R-Parity Conserving Supersymmetric Extension of the Zee Model

We extend the Zee model, where tiny neutrino masses are generated at the one loop level, to a supersymmetric model with R-parity conservation. It is found that the neutrino mass matrix can be consistent with the neutrino oscillation data thanks to the nonholomorphic Yukawa interaction generated via one-loop diagrams of sleptons. We find a parameter set of the model, where in addition to the neutrino oscillation data, experimental constraints from the lepton flavor violating decays of charged leptons and current LHC data are also satisfied. In the parameter set, an additional CP-even neutral Higgs boson other than the standard-model-like one, a CP-odd neutral Higgs boson, and two charged scalar bosons are light enough to be produced at the LHC and future lepton colliders. If the lightest charged scalar bosons are mainly composed of the SU(2)_L-singlet scalar boson in the model, they would decay into e nu and mu nu with 50% of a branching ratio for each. In such a case, the relation among the masses of the charged scalar bosons and the CP-odd Higgs in the minimal supersymmetric standard model approximately holds with a radiative correction. Our model can be tested by measuring the specific decay patterns of charged scalar bosons and the discriminative mass spectrum of additional scalar bosons.

hep-ph↗

Probing Models of Neutrino Masses via the Flavor Structure of the Mass Matrix

We discuss what kinds of combinations of Yukawa interactions can generate the Majorana neutrino mass matrix. We concentrate on the flavor structure of the neutrino mass matrix because it does not depend on details of the models except for Yukawa interactions while determination of the overall scale of the mass matrix requires to specify also the scalar potential and masses of new particles. Thus, models to generate Majorana neutrino mass matrix can be efficiently classified according to the combination of Yukawa interactions. We first investigate the case where Yukawa interactions with only leptons are utilized. Next, we consider the case with Yukawa interactions between leptons and gauge singlet fermions, which have the odd parity under the unbroken Z_2 symmetry. We show that combinations of Yukawa interactions for these cases can be classified into only three groups. Our classification would be useful for the efficient discrimination of models via experimental tests for not each model but just three groups of models.

hep-ph↗

Radiative Neutrino Mass Models

In this short review, we see some typical models in which light neutrino masses are generated at the loop level. These models involve new Higgs bosons whose Yukawa interactions with leptons are constrained by the neutrino oscillation data. Predictions about flavor structures of $\ell \to \overline{\ell}_1 \ell_2 \ell_3$ and leptonic decays of new Higgs bosons via the constrained Yukawa interactions are briefly summarized in order to utilize such Higgs as a probe of $ν$ physics.

hep-ph↗

Neutrino Mass and Dark Matter from Gauged $U(1)_{B-L}$ Breaking

We propose a new model where the Dirac mass term for neutrinos, the Majorana mass term for right-handed neutrinos, and the other new fermion masses arise via the spontaneous breakdown of the $U(1)_{B-L}$ gauge symmetry. The anomaly-free condition gives four sets of assignment of the B-L charge to new particles, and three of these sets have an associated global $U(1)_{DM}$ symmetry which stabilizes dark matter candidates. The dark matter candidates contribute to generating the Dirac mass term for neutrinos at the one-loop level. Consequently, tiny neutrino masses are generated at the two-loop level via a Type-I-Seesaw-like mechanism. We show that this model can satisfy current bounds from neutrino oscillation data, the lepton flavor violation, the relic abundance of the dark matter, and the direct search for the dark matter. This model would be tested at future collider experiments and dark matter experiments.

hep-ph↗

Discrimination of Models Including Doubly Charged Scalar Bosons by Using Tau Lepton Decay Distributions

The doubly-charged scalar boson H^{--} is involved in several new physics models for generating neutrino masses. Depending on new physics models, H^{--} has the Yukawa interaction with a pair of left-handed charged leptons or a pair of right-handed ones. In this talk, we see that these two Yukawa interactions can be distinguished by measuring energy distributions of charged pions produced by decays of tau leptons if pair-produced H^{++} H^{--} can decay sufficiently into four charged leptons which involve one or two tau leptons. The information on the Yukawa interaction of H^{--} will help to discriminate models in which neutrino masses are generated.

hep-ph↗