SearcharxivSearch

arXiv subjects

Yo Nakawake

Publications and source records attributed to Yo Nakawake.

3 recordsLinked to original sources

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems. We introduce the NazoNazo Benchmark, a renewable and extensible evaluation dataset derived from Japanese children's riddles that isolates a specific class of reasoning processes: insight-like representational restructuring and metacognitive evaluation. Rather than modeling reasoning in general, these tasks provide a focused test of failure modes that are difficult to detect in standard benchmarks. We curate 201 riddles and establish a human reference on a 120-item subset (n = 126; mean accuracy 52.9%). The benchmark is fully open, low-cost to refresh, and designed for continual evaluation under reduced contamination risk. We evaluate 38 frontier LLMs (2023-2025) under a strict retrieval-free, zero-shot protocol. On the human-comparison subset, non-reasoning models achieve 7.6% accuracy and reasoning-oriented models reach 17.6%, compared with a human mean of 52.9%, although performance varies substantially across models. Beyond accuracy, qualitative analysis of model-generated thought-logs identifies a distinctive failure mode, which we call verification failure: models generate a correct intermediate candidate but fail to endorse it as their final answer. This dissociation between candidate generation and endorsement reveals a metacognitive bottleneck: across the models with usable thought-logs, verification failures account for between 5% and 39% of a model's incorrect answers. By isolating the gap between generation and verification, this work provides a practical framework for diagnosing reasoning reliability and suggests concrete directions for improvement, including better calibration, structured verification, and stopping mechanisms.

cs.AI

Prestige bias drives the viral spread of content reposted by influencers in online communities

Cultural evolution theory suggests that prestige bias - whereby individuals preferentially learn from prestigious figures - has played a key role in human ecological success. However, its impact within online environments remains unclear, particularly with respect to whether reposts by prestigious individuals amplify diffusion more effectively than reposts by noninfluential users. We analyzed over 55 million posts and 520 million reposts on Twitter (currently X) to examine whether users with high influence scores (hg indices) more effectively amplified the reach of others' content. Our findings indicate that posts shared by influencers are more likely to be further shared than those shared by non-influencers. This effect persisted over time, especially in viral posts. Moreover, a small group of highly influential users accounted for approximately half of the information flow within repost cascades. These findings demonstrate a prestige bias in information diffusion within the digital society, suggesting that cognitive biases shape content spread through reposting.

cs.SI

Systematic quantitative analyses reveal the folk-zoological knowledge embedded in folktales

Cultural learning is a unique human capacity essential for a wide range of adaptations. Researchers have argued that folktales have the pedagogical function of transmitting the essential information for the environment. The most important knowledge for foraging and pastoral society is folk-zoological knowledge, such as the predator-prey relationship among wild animals, or between wild and domesticated animals. Here, we analysed the descriptions of the 382 animal folktales using the natural language processing method and descriptive statistics listed in a worldwide tale-type index (Aarne-Thompson-Uther type index). Our analyses suggested that first, the predator-prey relationship frequently appeared in a co-occurrent animal pair within a folktale (e.g., cat and mouse or wolf and pig), and second, the motif of 'deception', describing the antagonistic behaviour among animals, appeared relatively higher in 'wild and domestic animals' and 'wild animals' than other types. Furthermore, the motif of 'deception' appeared more frequently in pairs, corresponding to the predator-prey relationship. These results corresponded with the hypothesis that the combination of animal characters and what happens in stories represented relationships in the real world. The present study demonstrated that the combination of quantitative methods and qualitative data broaden our understanding of the evolutionary aspects of human cultures.

cs.CL