Searcharxiv⌕ Search

arXiv subjects

Makoto Fukushima

Publications and source records attributed to Makoto Fukushima.

6 recordsLinked to original sources

Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

Whether human or large language model (LLM), an agent in a discussion reads only a few of the others' contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the bound with one number, the message capacity, which sets how many of the others' messages an agent reads, and generate the communication network from it. Over 31,824 randomized queries, we found that an 8-billion-parameter model's judgment of a claim effectively reduces to a logistic function of a weighted sum of its inbox, the update rule of a stochastic binary neuron with divisively normalized weights. From these weights and the network's degree statistics alone, the wrong consensus should become unreachable from any start once agents read, on average, fewer than 6.4 of their 31 sources. In 1,414 episodes with assigned starts the prediction failed: the correct side won in fewer than 50% of episodes from every start, and in only 28-45% when 75% of agents started correct. The failure traces to the field, the threshold that a claim's wording sets for the agent's answer before any message is read: the experimental claims' fields lay below the calibration mean, and with each claim's own field the same weights reproduce the outcomes. Reversing the wording showed that the threshold follows what a claim asserts, not whether it is true. On a second 8B model the pipeline predicts claim-dependent bistability; transition points appeared where computed, and an eight-claim calibration matched in 15 of 16 conditions. At 70B the assertion bias is not detected. Thus a collective's fate is largely set by two single-agent measurements: the threshold a claim's wording sets, and the message capacity that sets the transition point.

cs.MA↗

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions -- shared protocols for reading meaning beyond the literal message -- which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (an online Hanabi platform), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, -0.7 pp in AI pairs, and +16.4 pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46 pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38-41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6 pp at the convention-free level, rising monotonically to +21.7 pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.

cs.AI↗

Analytically tractable model of synaptic crowding explains emergent small-world structure and network dynamics

Neural circuits must balance local connectivity constraints against the need for global integration. Here we introduce a minimal wiring rule motivated by synaptic crowding: as a neuron accumulates incoming connections, each additional synapse becomes progressively harder to form. This single-parameter model admits an exact finite-size solution for the induced in-degree distribution and yields simple scaling laws: mean connectivity grows only logarithmically with network size while variance remains bounded -- consistent with homeostatic regulation of synaptic density. When candidates are encountered in order of spatial proximity, the crowding rule produces a broad, approximately power-law distribution of connection lengths without prescribing any explicit distance-dependent wiring law; combined with shortcut rewiring, this yields networks with small-world characteristics. We further show that the induced degree statistics largely determine attractor basin boundaries in threshold network dynamics, while local clustering primarily modulates the prevalence of long-lived non-absorbing outcomes near these boundaries. The model provides testable predictions linking local developmental constraints to macroscopic network organization and dynamics.

q-bio.NC↗

Advancements and limitations of LLMs in replicating human color-word associations

Color-word associations play a fundamental role in human cognition and design applications. Large Language Models (LLMs) have become widely available and have demonstrated intelligent behaviors in various benchmarks with natural conversation skills. However, their ability to replicate human color-word associations remains understudied. We compared multiple generations of LLMs (from GPT-3 to GPT-4o) against human color-word associations using data collected from over 10,000 Japanese participants, involving 17 colors and 80 words (10 word from eight categories) in Japanese. Our findings reveal a clear progression in LLM performance across generations, with GPT-4o achieving the highest accuracy in predicting the best voted word for each color and category. However, the highest median performance was approximately 50% even for GPT-4o with visual inputs (chance level of 10%). Moreover, we found performance variations across word categories and colors: while LLMs tended to excel in categories such as Rhythm and Landscape, they struggled with categories such as Emotions. Interestingly, color discrimination ability estimated from our color-word association data showed high correlation with human color discrimination patterns, consistent with previous studies. Thus, despite reasonable alignment in basic color discrimination, humans and LLMs still diverge systematically in the words they assign to those colors. Our study highlights both the advancements in LLM capabilities and their persistent limitations, raising the possibility of systematic differences in semantic memory structures between humans and LLMs in representing color-word associations.

cs.CL↗

Fluctuations between high- and low-modularity topology in time-resolved functional connectivity

Modularity is an important topological attribute for functional brain networks. Recent studies have reported that modularity of functional networks varies not only across individuals being related to demographics and cognitive performance, but also within individuals co-occurring with fluctuations in network properties of functional connectivity, estimated over short time intervals. However, characteristics of these time-resolved functional networks during periods of high and low modularity have remained largely unexplored. In this study we investigate spatiotemporal properties of time-resolved networks in the high and low modularity periods during rest, with a particular focus on their spatial connectivity patterns, temporal homogeneity and test-retest reliability. We show that spatial connectivity patterns of time-resolved networks in the high and low modularity periods are represented by increased and decreased dissociation of the default mode network module from task-positive network modules, respectively. We also find that the instances of time-resolved functional connectivity sampled from within the high (low) modularity period are relatively homogeneous (heterogeneous) over time, indicating that during the low modularity period the default mode network interacts with other networks in a variable manner. We confirmed that the occurrence of the high and low modularity periods varies across individuals with moderate inter-session test-retest reliability and that it is correlated with previously-reported individual differences in the modularity of functional connectivity estimated over longer timescales. Our findings illustrate how time-resolved functional networks are spatiotemporally organized during periods of high and low modularity, allowing one to trace individual differences in long-timescale modularity to the variable occurrence of network configurations at shorter timescales.

q-bio.NC↗

Dynamic fluctuations coincide with periods of high and low modularity in resting-state functional brain networks

We investigate the relationship of resting-state fMRI functional connectivity estimated over long periods of time with time-varying functional connectivity estimated over shorter time intervals. We show that using Pearson's correlation to estimate functional connectivity implies that the range of fluctuations of functional connections over short time scales is subject to statistical constraints imposed by their connectivity strength over longer scales. We present a method for estimating time-varying functional connectivity that is designed to mitigate this issue and allows us to identify episodes where functional connections are unexpectedly strong or weak. We apply this method to data recorded from $N=80$ participants, and show that the number of unexpectedly strong/weak connections fluctuates over time, and that these variations coincide with intermittent periods of high and low modularity in time-varying functional connectivity. We also find that during periods of relative quiescence regions associated with default mode network tend to join communities with attentional, control, and primary sensory systems. In contrast, during periods where many connections are unexpectedly strong/weak, default mode regions dissociate and form distinct modules. Finally, we go on to show that, while all functional connections can at times manifest stronger (more positively correlated) or weaker (more negatively correlated) than expected, a small number of connections, mostly within the visual and somatomotor networks, do so a disproportional number of times. Our statistical approach allows the detection of functional connections that fluctuate more or less than expected based on their long-time averages and may be of use in future studies characterizing the spatio-temporal patterns of time-varying functional connectivity

q-bio.NC↗