SearcharxivSearch

arXiv subjects

Kohei Hayashi

Publications and source records attributed to Kohei Hayashi.

At least 19 recordsLinked to original sources

Steering Recurrent Reasoners at Inference Time with Readout Feedback

Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more steps or sampling more trajectories, but ignore information revealed within each trajectory. Here we show that recurrent models can be improved at inference time by using their own readout probabilities to steer latent dynamics without retraining. We introduce Readout Feedback (RoFB), a test-time intervention that converts intermediate predictions into token-wise pairwise coupling forces injected into the latent dynamics. Across three recurrent models (AKOrN, ItrSA++, TRM) on Sudoku and Maze, RoFB yields clear gains in four of six model-task pairs, achieving performance unattainable by merely running more steps or selecting from multiple trajectories, at comparable or lower computational cost. These results suggest that closed-loop steering of latent dynamics can serve as a complementary inference-time control mechanism for recurrent reasoning models.

cs.LG

Looped Transformers with Source-Centered State Evolution

Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective depth at a fixed parameter count. However, that shared block must then govern an entire trajectory of varying hidden states over trained and extrapolated depths. Furthermore, in additive-injection looped Transformers, an input-conditioned signal is reintroduced at every recurrent step, so applying the shared transition at an input-conditioned reference can still move the hidden state. In this paper, we propose Source-Centered State Evolution (SCSE), which is designed to reconcile input conditioning with reference-preserving shared recurrence. Specifically, SCSE retains input dependence through its learned anchor and initial deviation, allows nonzero deviations to drive recurrent computation while mapping zero deviation to zero, and guarantees exact anchor invariance through its zero-deviation mask. The designated anchor is thereby a one-step fixed point by construction. The zero-deviation forcing bias is the next deviation produced from the anchor itself and vanishes in SCSE, while nonzero deviations remain active and support state-dependent recurrent computation. Our theory shows that the zero-deviation forcing bias is a design degree of freedom whose task effect can be harmful, neutral, or beneficial; SCSE resolves this choice in favor of exact anchor invariance by setting the bias to zero. Across WikiText-2, WikiText-103, direct web-corpus pretraining, held-out web-text transfer, and LAMBADA completion, SCSE improves the controlled recurrent quality frontier. Ablation studies identify the learned anchor and the anchor-coordinate deviation recurrence as the primary contributors to the gain, and a trained-model case study grounds the anchor-response diagnostic in observed recurrent motion.

cs.LG

Dark Matter in Draco and Bo\"otes I: Hints of a Core in an Ultra-Faint Dwarf from Simulation-Based Inference

The density profiles of dwarf spheroidal galaxies are among the most sensitive probes of dark matter physics, yet extracting them from noisy stellar kinematics remains a fundamental obstacle. We present GraphNPE, a simulation-based inference method for dynamical mass modeling that incorporates measurement uncertainties and spectroscopic selection functions in the forward model. Using mock data, we show that methods relying solely on line-of-sight velocity dispersion are biased toward cuspy density profiles, even in the absence of the mass-anisotropy degeneracy. By accessing higher-order velocity moments, particularly line-of-sight kurtosis, GraphNPE breaks key degeneracies and recovers density profiles with substantially less bias. We apply GraphNPE to Draco and Bo\"otes I using MMT/Hectochelle and DESI for Draco, and the S5 survey for Bo\"otes I. For each, we report density profiles and dark matter $J$- and $D$-factors. For Draco, GraphNPE yields consistent results across datasets, marginally preferring a cuspy inner profile ($\rho_{150} \sim 1.6-1.9 \times 10^8\,\mathrm{M}_\odot\,\mathrm{kpc}^{-3}$) in agreement with literature. On DESI, however, second-order Jeans modeling fits the dispersion but fails to reproduce the kurtosis, demonstrating higher-order moments are essential. For Bo\"otes I, limited statistical power prevents definitive determination of the inner slope. GraphNPE recovers $\rho_{150} = 0.36^{+0.15}_{-0.11} \times 10^8\,\mathrm{M}_\odot\,\mathrm{kpc}^{-3}$, significantly lower than literature and consistent with a cored inner profile. This places Bo\"otes I among the lowest density dwarfs at comparable stellar masses.

astro-ph.GA

Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables

We use a mean-field-based transformer model to theoretically investigate how auxiliary variables, such as positional encoding, prevent mode collapse of self-attention mechanisms. The use of mean-field transformers to analyze the properties of self-attention mechanisms has garnered significant attention in recent years due to their ability to comprehensively analyze token interactions. However, analysis of this simple model suggests that mode collapse, where token distributions degenerate to a single point, occurs during long inferences (i.e., many layers), indicating a discrepancy with reality. This study investigates this mean-field transformer model and demonstrates that the introduction of auxiliary variables, such as positional encoding, acts as a counterforce against theoretical mode collapse. Specifically, we show that in the theoretical scheme, the energy-maximizing distribution does not degenerate to a single point; instead, it is characterized by a pushforward of the auxiliary variable distribution, thereby avoiding concentration in the Dirac measure. Our main examples are the positional encoding and the fixed prompt insertion treated as a parallel auxiliary-variable mechanism. Furthermore, we demonstrate that positional encoding and prompt insertion possess universality of representation in the limit, meaning that the limit distribution of inference can exactly represent a wide class of distributions. We also analyze several key properties of positional encoding and metastability, and validate our theoretical results through mathematical experiments.

cs.LG

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions

Neural network models with latent recurrent processing, where identical layers are recursively applied to the latent state, have gained attention as promising models for performing reasoning tasks. A strength of such models is that they enable test-time scaling, where the models can enhance their performance in the test phase without additional training. Models such as the Hierarchical Reasoning Model (HRM) and Artificial Kuramoto Oscillatory Neurons (AKOrN) can facilitate deeper reasoning by increasing the number of recurrent steps, thereby enabling the completion of challenging tasks, including Sudoku, Maze solving, and AGI benchmarks. In this work, we introduce confidence-based voting (C-voting), a test-time scaling strategy designed for recurrent models with multiple latent candidate trajectories. Initializing the latent state with multiple candidates using random variables, C-voting selects the one maximizing the average of top-1 probabilities of the predictions, reflecting the model's confidence. Additionally, it yields 4.9% higher accuracy on Sudoku-hard than the energy-based voting strategy, which is specific to models with explicit energy functions. An essential advantage of C-voting is its applicability: it can be applied to recurrent models without requiring an explicit energy function. Finally, we introduce a simple attention-based recurrent model with randomized initial values named ItrSA++, and demonstrate that when combined with C-voting, it outperforms HRM on Sudoku-extreme (95.2% vs. 55.0%) and Maze (78.6% vs. 74.5%) tasks.

cs.LG

Galactic Archaeology with the Subaru `\=Onohi`ula Prime Focus Spectrograph Strategic Program

The recently commissioned Subaru `\=Onohi`ula Prime Focus Spectrograph (PFS) will obtain spectra from nearly 2,400 fibers that cover 1.24 square degrees. The 360 night Subaru Strategic Program for PFS is dedicating approximately one-third of its allocation (130 nights) to study the structure and evolution of galaxies in the Local Group. This Galactic Archaeological survey has three pillars. (1) We will determine whether the mass density profiles of dwarf galaxies are consistent with cusps, as expected for cold dark matter, or cores, as expected from alternative dark matter theories or baryonic feedback. We will deduce the density profiles as a function of radius from modeling of the full line-of-sight velocity and abundance distributions for six dwarf galaxies. Our total sample will consist of 18,000 member stars to beyond the nominal tidal radius of each system. (2) From measurements of the [alpha/Fe] abundance ratio, we will learn the difference in assembly history of the two most massive galaxies in the Local Group: M31 and the Milky Way. We will observe 30,000 member stars over 45 square degrees of M31's halo and outer disk. (3) We will uncover how the most fragile (outer) part of the Milky Way responded to accretion events both in the distant past (such as Gaia-Sausage Enceladus) and in more recent history (such as the Sagittarius dwarf spheroidal galaxy). To support this study, PFS will provide velocities and metallicities--from which, in combination with photometry, we will deduce ages--for tens of thousands of main-sequence stars out to a Galactocentric distance of ~30 kpc.

astro-ph.GA

Degree-preserving conservative processes and a unified approach for their hydrodynamics

We investigate a broad class of large-scale one-dimensional interacting systems characterized by a single conservation law and satisfying the "degree-preserving property". Under mild and natural assumptions, we establish a unified framework for the analysis of both invariant measures and hydrodynamic limits. In particular, we prove that when the generator preserves the degree of polynomials of the state variables up to order two, the marginals of any product invariant measure must belong to a family of six specific distributions. This classification is shown to be consistent with a classical result on univariate natural exponential families due to C.N. Morris, which we apply here for the first time in the context of microscopic stochastic systems. As a consequence, we construct a new interacting particle system whose invariant measure is given by the generalized hyperbolic secant distribution. Furthermore, we prove that, despite the generality of the dynamics, the macroscopic behavior of all models in this class is governed by the classical heat equation, with a diffusion coefficient depending explicitly on the underlying microscopic interactions.

math.PR

Exploration of Fast-Slow Latent Recurrence for Train-Short, Test-Long Generalization

We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory. Our focus is on a persistent fast slow recurrent formulation in which a latent state is maintained across observations rather than reset at each stream step. For each incoming observation, the model performs multiple weight-shared latent updates with a recurrent core and then carries the resulting state forward to the next observation. This allows the model to maintain and refine a compact stream-level state without reprocessing a growing context. We evaluate this formulation across symbolic sequence prediction, supervised navigation, and partially observable reinforcement learning tasks. Across these settings, persistent latent recurrence improves OOD generalization over recurrent, state-space, and Transformer baselines. Through recurrent-core ablations, we identify architectural ingredients that are consistently associated with strong OOD performance, including state-dependent transitions and feature-wise nonlinear mixing. Together, these results highlight the value of revisiting persistent recurrence as an architectural bias for more generalizable sequence prediction.

cs.LG

A Long Stellar Stream in M83: Possible Connection Between XUV Disks and Minor Mergers?

We present the confirmation and characterization of a long stream (S-stream) in the southern part of M83. This feature is revealed using deep wide-field photometric data obtained by the Hyper Suprime-Cam (HSC) mounted on the Subaru Telescope. Using individual red giant branch (RGB) stars, we successfully trace the stream over a large length of $\sim 81$~kpc and a considerable width of $\sim 9$ kpc. With a mean surface brightness of ${\langle \mu_{\it V} \rangle} \sim 31.8_{-1.9}^{+1.3}$ mag arcsec$^{-2}$, it is one of the most diffuse extragalactic streams currently known. The mean photometric metallicity of the stream is $\langle[{\rm M/H}]\rangle = -1.23\pm0.02$ dex with a standard deviation of $0.28\pm0.01$ dex, and we estimate the stellar mass to be $(8.5_{-2.8}^{+4.2}) \times 10^6~{\rm M_\odot}$ from the luminosity of RGB stars. Compared to its well-known northern counterpart, the S-stream is slightly more metal-poor, but our large-area RGB map shows compelling evidence that these two features are related, originating from a single low-mass merger event. We identify density variations along the S-stream, which more likely reflect intrinsic density structure within the progenitor rather than the interaction with dark matter subhalos. Similarities between the morphology of the S-stream and some features in the \HI distribution suggest that a minor merger event may have disturbed and redistributed M83's outer \HI gas, leading to triggered star formation and the formation of the XUV disk.

astro-ph.GA

Robust Sim-to-Real Cloth Untangling through Reduced-Resolution Observations via Adaptive Force-Difference Quantization

Robotic cloth untangling requires progressively disentangling fabric by adapting pulling actions to changing contact and tension conditions. Because large-scale real-world training is impractical due to cloth damage and hardware wear, sim-to-real policy transfer is a promising solution. However, cloth manipulation is highly sensitive to interaction dynamics, and policies that depend on precise force magnitudes often fail after transfer because similar force responses cannot be reproduced due to the reality gap. We observe that untangling is largely characterized by qualitative tension transitions rather than exact force values. This indicates that directly minimizing the sim-to-real gap in raw force measurements does not necessarily align with the task structure. We therefore hypothesize that emphasizing coarse force-change patterns while suppressing fine environment-dependent variations can improve robustness of sim-to-real transfer. Based on this insight, we propose Adaptive Force-Difference Quantization (ADQ), which reduces observation resolution by representing force inputs as discretized temporal differences and learning state-dependent quantization thresholds adaptively. This representation mitigates overfitting to environment-specific force characteristics and facilitates direct sim-to-real transfer. Experiments in both simulation and real-world cloth untangling demonstrate that ADQ achieves higher success rates and exhibits greater robustness in sim-to-real transfer than policies using raw force inputs. Supplementary video is available at https://youtu.be/ZeoBs-t0AWc

cs.RO

Fuzzy Dark Matter and the Impact of Core--Halo Diversity on Its Particle Mass Constraints

We investigate how diversity in the core--halo mass relation and the inclusion of higher-order velocity moments affect constraints on the fuzzy dark matter particle mass ($m_\psi$) inferred from the internal kinematics of dwarf galaxies. Using stellar line-of-sight velocities and projected positions for eight Milky Way dwarf spheroidal galaxies, we model their dark matter halos as solitonic cores embedded within outer Navarro--Frenk--White envelopes. We apply both second- and fourth-order Jeans analyses to derive the posterior distribution of $m_\psi$. Our results show that there are two ranges of $m_\psi$ consistent with the observed kinematics: $\log_{10}(m_\psi/\mathrm{eV}) = -19.72^{+0.64}_{-0.56}$, and a narrower low-mass window $\log_{10}(m_\psi/\mathrm{eV}) = -21.81^{+0.39}_{-0.26}$, both within the 68\% credible intervals. The latter becomes prominent only when core--halo diversity is taken into account, which highlights the sensitivity of the inferred fuzzy dark matter particle mass constraints to our understanding of the core--halo relation. Future observations, providing larger stellar samples and more precise kinematic measurements, will be essential for clarifying the allowed parameter space of fuzzy dark matter.

astro-ph.GA

The Characteristic Mass and Energy Conversion Efficiency in the Cusp-Core Transition of Dark Matter Haloes: Implications for Scaling Relations and Supernova feedbacks

Galaxies in the nearby Universe, particularly dwarf systems, exhibit inner mass profiles of dark matter haloes that systematically depart from canonical cold dark matter expectations, signalling an interplay between baryonic feedback and the collisionless halo. We update an analytical cusp-core transition model by incorporating the effect of supernova-driven mass loss. Adapting this model to SPARC galaxies, we measure the energy conversion efficiency epsilon, defined as the fraction of supernova feedback energy that is used to change the central dark-matter potential. We find epsilon ~ 0.01 for nearby SPARC galaxies. Building on these measurements, we compare the dynamical energy required for a cusp-core transformation with the feedback energy available over burst cycles and identify a cusp-core transition forbidden region on the halo-stellar mass plane where transformation cannot occur. Galaxies with halo masses from 10^8 to 10^11 M_sun lie outside the forbidden region, whereas ultra-faint dwarf galaxies < 10^8 M_sun, galaxy groups and clusters > 10^11 M_sun fall within it, consistent with their high central densities and the inefficiency of core formation at very low and very high masses. This approach also explains the observed diversity of inner density profiles in low-mass systems, showing that both the star formation rate and the energy conversion efficiency govern them, with the latter emerging as a key parameter setting the strength of the cusp-core transition. Beyond the cusp-core problem, our observationally inferred energy conversion efficiency provides a model independent benchmark that strongly constrains galaxy formation models.

astro-ph.GA

Nonlinear fluctuations for a chain of weakly anharmonic oscillators with stochastic perturbation

We study the fluctuations of the phonon modes in a one-dimensional chain of anharmonic oscillators where the deterministic Hamiltonian dynamics is perturbed by random exchanges of momentum between nearest neighbor particles. There are three locally conserved quantities: volume, momentum and energy. We study the evolution in equilibrium of the fluctuation fields of the two phonon modes (linear combination of the volume stretch and momentum), on a diffusive space-time scale after recentering on their sound velocities. We show that, weakening the anharmonicity with the scale parameter, the recentered phonon fluctuations fields converge to the stationary solutions of two uncoupled stochastic Burgers equations. The nonlinearity in the Burgers equation depends on the presence of a cubic term in the anharmonic potential (corresponding to the $\alpha$-FPUT dynamics). Main ingredients of the proof, based on a compactness argument for the Dynkin's martingale decomposition, are the second-order Boltzmann-Gibbs principle, as well as equipartition of energy, to characterize the nonlinear term and Riemann-Lebesgue estimates showing that fields with diverging velocity to different directions have no interaction in the limit.

math.PR

Chemodynamics of Bo\"otesI with $S^{5}$: Revised Velocity Gradient, Dark Matter Density, and Galactic Chemical Evolution Constraints

We combine new spectroscopic observations of the ultra faint dwarf galaxy (UFD) Bo\"otes I (Boo I) from the Southern Stellar Stream Spectroscopic Survey ($S^{5}$) with $\sim$15 years of archival spectroscopic data to create the largest sample of stellar kinematics and metallicities to date in any Milky Way UFD. Our combined sample includes 148 members extending out to $\sim$7 half-light radii ($r_h$), including 24 newly confirmed members, 18 binary candidates, 15 RR Lyrae stars, and 92 [Fe/H] measurements. Using this larger and more spatially extended sample, we provide updated constraints on Boo I's systemic properties, including its radial population gradients. Properly accounting for perspective rotation effects in a UFD for the first time, we detect a $4\sigma$ line-of-sight velocity gradient of $1.2\pm0.3$ km s$^{-1}$ $r_h^{-1}$ aligned along Boo I's orbit and discuss its potential tidal origins. We also infer a metallicity gradient of $-0.10\pm0.02$ dex $r_h^{-1}$ in agreement with previous studies. Using an axisymmetric Jeans model, we provide updated constraints on Boo I's dark matter density profile, which weakly favor a cusped ($\gamma=1.0^{+0.5}_{-0.6}$) dark matter profile. Lastly, we re-analyze Boo I's metallicity distribution function with a one-zone galactic chemical evolution model and place new constraints on its rapid, inefficient star formation and strong galactic outflows.

astro-ph.GA

Cusp-to-Core Transition of Dark Matter Halos across Galaxy Mass Scales

We investigate the diversity of dark matter (DM) density profiles in a large sample of late-type galaxies from the SPARC database, with the goal of testing whether a cusp-to-core transition occurs across galaxy mass scales. We perform Bayesian fits to high-quality rotation curves using flexible halo models that allow for variations in the inner slopes of DM density profiles. We quantify the central dark matter structure using the surface density within the inner region of the halo, defined as $\Sigma_{\rm DM}(<0.01r_{V_{\rm max}})$, and compare the SPARC galaxies with Milky Way dwarf satellites as well as galaxy groups and clusters. Our results reveal significant diversity in the inner density slopes of SPARC galaxies, ranging from steep cusps to shallow cores, and show that many of them lie below the cuspy profiles predicted by the cold dark matter model, consistent with core-like structures. In contrast, both lower-mass dwarf galaxies and higher-mass galaxy clusters tend to follow the cuspy DM halos. These findings suggest that baryonic feedback may induce a cusp-to-core transition in Milky Way-mass galaxies, as predicted by hydrodynamical simulations. However, observational limitations and modeling uncertainties still prevent a definitive conclusion. This study provides new empirical insights into the halo mass-dependent nature of DM inner structures and the role of baryonic processes in shaping them.

astro-ph.GA

Hydrodynamic limit for some gradient and attractive spin models

We study the hydrodynamic limit for three gradient spin models: generalized Kipnis-Marchioro-Presutti (KMP), its discrete version and a family of harmonic models, under symmetric and nearest-neighbor interactions. These three models share some universal properties: occupation variables are unbounded, all these processes are of gradient type, their invariant measures are product with spatially homogeneous weights, and, notably, they are all attractive, meaning that the process preserves the partial order of measures along the dynamics. In view of hydrodynamics of large-scale interacting systems, dealing with processes taking values in unbounded configuration spaces is known to be a challenging problem. In the present paper, we show the hydrodynamic limit for all three models listed above in a comprehensive way, and show as a main result, that, under the diffusive time scaling, the hydrodynamic equation is given by the heat equation with model-dependent diffusion coefficient. Our novelty is showing the attractiveness for each model, which is crucial for the proof of hydrodynamics.

math.PR

Stochastic oscillators out of equilibrium: scaling limits and correlation estimates

We consider a purely harmonic chain of oscillators which is perturbed by a stochastic noise. Under this perturbation, the system exhibits two conserved quantities: the volume and the energy. At the level of the hydrodynamic limit, under diffusive scaling, we show that depending on the strength of the Hamiltonian dynamics, energy and volume evolve according to either a system of autonomous heat equations or a non-linear system of coupled parabolic equations. Moreover, for general initial measures, under diffusive scaling, we can characterize the non-equilibrium volume fluctuations. The proofs are based on precise bounds on the two-point volume correlation function and a uniform fourth-moment estimate.

math.PR

JFlow: Model-Independent Spherical Jeans Analysis using Equivariant Continuous Normalizing Flows

The kinematics of stars in dwarf spheroidal galaxies have been studied to understand the structure of dark matter halos. However, the kinematic information of these stars is often limited to celestial positions and line-of-sight velocities, making full phase space analysis challenging. Conventional methods rely on projected analytic phase space density models with several parameters and infer dark matter halo structures by solving the spherical Jeans equation. In this paper, we introduce an unsupervised machine learning method for solving the spherical Jeans equation in a model-independent way as a first step toward model-independent analysis of dwarf spheroidal galaxies. Using equivariant continuous normalizing flows, we demonstrate that spherically symmetric stellar phase space densities and velocity dispersions can be estimated without model assumptions. As a proof of concept, we apply our method to Gaia challenge datasets for spherical models and measure dark matter mass densities for given velocity anisotropy profiles. Our method can identify halo structures accurately, even with a small number of tracer stars.

astro-ph.GA