SearcharxivSearch

arXiv subjects

Kaito Takanami

Publications and source records attributed to Kaito Takanami.

4 recordsLinked to original sources

An Asymptotic Theory of Chain-of-Thought in In-Context Learning

Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermediate reasoning steps at inference time. Yet the scaling behavior of generalization with CoT depth remains poorly understood. To address this question, we study a theoretically solvable model of CoT for in-context weight prediction in linear regression, where test-time reasoning is represented as an iterative refinement of the weight-parameter estimate. Using tools from random matrix theory under high-dimensional asymptotics, we derive an exact formula for the generalization error as a function of reasoning depth, pretraining data amount, and context length. Our analysis reveals a sharp phase transition separating exponential and polynomial improvement, saturation, and overthinking, and characterizes how the optimal reasoning depth scales. We further show that deeper reasoning is most effective with sufficiently rich pretraining and in-context information, whereas limited pretraining or context makes longer reasoning prone to error amplification or saturation. We also validate these predictions through experiments on fully learned linear attention and softmax attention models. Our results provide a unified theoretical account of how test-time CoT depth affects generalization.

stat.ML

Learning Linear Regression with Low-Rank Tasks in-Context

In-context learning (ICL) is a key building block of modern large language models, yet its theoretical mechanisms remain poorly understood. It is particularly mysterious how ICL operates in real-world applications where tasks have a common structure. In this work, we address this problem by analyzing a linear attention model trained on low-rank regression tasks. Within this setting, we precisely characterize the distribution of predictions and the generalization error in the high-dimensional limit. Moreover, we find that statistical fluctuations in finite pre-training data induce an implicit regularization. Finally, we identify a sharp phase transition of the generalization error governed by task structure. These results provide a framework for understanding how transformers learn to learn the task structure.

cond-mat.dis-nn

The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture Model

Self-distillation (SD), a technique where a model improves itself using its own predictions, has attracted attention as a simple yet powerful approach in machine learning. Despite its widespread use, the mechanisms underlying its effectiveness remain unclear. In this study, we investigate the efficacy of hyperparameter-tuned multi-stage SD with a linear classifier for binary classification on noisy Gaussian mixture data. For the analysis, we employ the replica method from statistical physics. Our findings reveal that the primary driver of SD's performance improvement is denoising through hard pseudo-labels, with the most notable gains observed in moderately sized datasets. We also identify two practical heuristics to enhance SD: early stopping that limits the number of stages, which is broadly effective, and bias parameter fixing, which helps under label imbalance. To empirically validate our theoretical findings derived from our toy model, we conduct additional experiments on CIFAR-10 classification using pretrained ResNet backbone. These results provide both theoretical and practical insights, advancing our understanding and application of SD in noisy settings.

stat.ML

Detection of diffusion anisotropy from an individual short particle trajectory

In parallel with advances in microscale imaging techniques, the fields of biology and materials science have focused on precisely extracting particle properties based on their diffusion behavior. Although the majority of real-world particles exhibit anisotropy, their behavior has been studied less than that of isotropic particles. In this study, we introduce a new method for estimating the diffusion coefficients of individual anisotropic particles using short-trajectory data on the basis of a maximum likelihood framework. Traditional estimation techniques often use mean-squared displacement (MSD) values or other statistical measures that inherently remove angular information. Instead, we treated the angle as a latent variable and used belief propagation to estimate it while maximizing the likelihood using the expectation-maximization algorithm. Compared to conventional methods, this approach facilitates better estimation of shorter trajectories and faster rotations, as confirmed by numerical simulations and experimental data involving bacteria and quantum rods. Additionally, we performed an analytical investigation of the limits of detectability of anisotropy and provided guidelines for the experimental design. In addition to serving as a powerful tool for analyzing complex systems, the proposed method will pave the way for applying maximum likelihood methods to more complex diffusion phenomena.

cond-mat.mes-hall