SearcharxivSearch

arXiv subjects

Yuta Yamamoto

Publications and source records attributed to Yuta Yamamoto.

5 recordsLinked to original sources

LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) is a framework in which a Generator, such as a Large Language Model (LLM), produces answers by retrieving documents from an external collection using a Retriever. In practice, Generators must integrate evidence from long contexts, perform multi-step reasoning, interpret tables, and abstain when evidence is missing. However, existing benchmarks for Generators provide limited coverage, with none enabling simultaneous evaluation of multiple capabilities under unified conditions. To bridge the gap between existing evaluations and practical use, we introduce LIT-RAGBench (the Logic, Integration, Table, Reasoning, and Abstention RAG Generator Benchmark), which defines five categories: Integration, Reasoning, Logic, Table, and Abstention, each further divided into practical evaluation aspects. LIT-RAGBench systematically covers patterns combining multiple aspects across categories. By using fictional entities and scenarios, LIT-RAGBench evaluates answers grounded in the provided external documents. The dataset consists of 114 human-constructed Japanese questions and an English version generated by machine translation with human curation. We use LLM-as-a-Judge for scoring and report category-wise and overall accuracy. Across API-based and open-weight models, no model exceeds 90% overall accuracy. By making strengths and weaknesses measurable within each category, LIT-RAGBench serves as a valuable metric for model selection in practical RAG deployments and for building RAG-specialized models. We release LIT-RAGBench, including the dataset and evaluation code, at https://github.com/Koki-Itai/LIT-RAGBench.

cs.CL

A Wide and Deep Exploration of Radio Galaxies with Subaru HSC (WERGS). X. The Massive and Passive Nature of Radio Galaxies at $z \sim 4$

High-$z$ radio galaxies (HzRGs) are considered important objects for understanding the formation and evolution of massive galaxies in the early universe. However, till date, detailed studies of the stellar population of HzRGs such as the star-formation history have been scarce. Therefore, this study conducted a new survey to establish a less-biased sample of HzRGs and consequently investigate their properties. We utilized a sample of $g$-dropout Lyman break galaxies (LBGs) obtained from an optical wide and deep imaging survey made by Subaru Hyper Suprime-Cam (HSC). Based on the cross-matching of this LBG sample with the VLA FIRST radio survey data, we constructed a photometric sample of high-redshift radio galaxies (HzRGs) at $z \sim 4$ for $\sim$560 deg$^2$ survey field. Consequently, we identified 146 HzRG candidates. To analyze the characteristics of these candidates, we focus on objects exhibiting the near-infrared photometry of VIKING or UKIDSS and the mid-infrared photometry of unWISE (28 objects). The results indicate that 7 objects exhibit SEDs consistent with galaxies at $z \sim 4$. The HzRG candidates have very large stellar masses with $\sim 4.2 \times 10^{11} M_{\odot}$ on average. This stellar mass is similar to that of previously discovered USS HzRGs at $z \sim 4$, though our sample is affected by a sample selection bias that selects only HzRGs with $M_{\star} > 10^{11} M_{\odot}$. Further, the SEDs of those HzRG candidates suggest a past fast quenching with a rough timescale of $\sim$0.1 Gyr, as evidenced from the rest-frame UVJ diagram.

astro-ph.GA

New technique to select recent fast-quenching galaxies at $z\sim2$ using the optical colors

Many massive quiescent galaxies have been discovered at $z>2$ thanks to multi-wavelength deep and wide surveys, however, substantial deep near-infrared spectroscopic observations are needed to constrain their star-formation histories statistically. Here, we present a new technique to select quiescent galaxies with a short quenching timescale ($\leq0.1$ Gyr) at $z\sim2$ photometrically. We focus on a spectral break at $\sim1600$ Å~that appears for such fast-quenching galaxies $\sim1$ Gyr after quenching when early A-type stars go out, but late A-type stars still live. This spectral break at $z\sim2$ is similar to a Lyman break at $z\sim4$. We construct a set of color criteria for $z\sim2$ fast-quenching galaxies on $g-r$ vs. $r-i$ and $i-J$ vs. $J-H$ or $\rm i-[3.6]$ vs. $\rm [3.6]-[4.5]$ color diagrams, which are available with the existing and/or future wide imaging surveys, by simulating various model galaxy spectra and test their robustnesses using the COSMOS2020 catalog. Galaxies with photometric and/or spectroscopic redshifts $z\sim2$ and low specific star formation rates are successfully selected using these colors. The number density of these fast-quenching galaxy candidates at $z\sim2$ suggests that massive galaxies not so far above the star-formation main sequence at $z=3-4$ should be their progenitors.

astro-ph.GA

Investigating Quantitative-Qualitative Topical Preference: A Comparative Study of Early and Late Engagers in Japanese ChatGPT Conversations

This study investigates engagement patterns related to OpenAI's ChatGPT on Japanese Twitter, focusing on two distinct user groups - early and late engagers, inspired by the Innovation Theory. Early engagers are defined as individuals who initiated conversations about ChatGPT during its early stages, whereas late engagers are those who began participating at a later date. To examine the nature of the conversations, we employ a dual methodology, encompassing both quantitative and qualitative analyses. The quantitative analysis reveals that early engagers often engage with more forward-looking and speculative topics, emphasizing the technological advancements and potential transformative impact of ChatGPT. Conversely, the late engagers intereact more with contemporary topics, focusing on the optimization of existing AI capabilities and considering their inherent limitations. Through our qualitative analysis, we propose a method to measure the proportion of shared or unique viewpoints within topics across both groups. We found that early engagers generally concentrate on a more limited range of perspectives, whereas late engagers exhibit a wider range of viewpoints. Interestingly, a weak correlation was found between the volume of tweets and the diversity of discussed topics in both groups. These findings underscore the importance of identifying semantic bias, rather than relying solely on the volume of tweets, for understanding differences in communication styles between groups within a given topic. Moreover, our versatile dual methodology holds potential for broader applications, such as studying engagement patterns within different user groups, or in contexts beyond ChatGPT.

cs.SI

Bicategorical Models of Classical Propositional Logic

Führmann and Pym constructed models of classical propositional logic in an order-enriched categorical setting, whose typical example is the category $\mathbf{Rel}$ of sets and relations. It is remarkable in that they are both non-degenerate and symmetric, i.e., free from the choices of the reduction strategy. As a furter categorification of this direction, we give bicategorical models of classical propositional logic that is also symmetric and non-degenerate. Primal examples of our models include $\mathbf{Rel}$, $\mathbf{Span}$, and $\mathbf{Prof}$, which shows that we can construct models that are non-degenerate not only for $1$-cells but also for $2$-cells and the logical negations.

math.CT