SearcharxivSearch

arXiv subjects

Xisong Dong

Publications and source records attributed to Xisong Dong.

3 recordsLinked to original sources

Unveiling the Reasoning Process of Large Language Models

Large language models often reason beyond surface tokens, but the internal stage at which token-level information becomes abstract relational structure remains unclear. We investigate this question by analyzing how attention heads and layers transform information during autoregressive reasoning. Across mathematical and symbolic reasoning tasks, we observe a consistent layer-wise division of labor: outer layers mainly preserve and route input-related features, whereas middle layers reorganize them into more transferable rule-level representations. This interpretation is supported by representation geometry: middle-layer states occupy lower-dimensional manifolds and show stronger alignment across disjoint vocabularies that instantiate the same symbolic rules. It is further supported by causal interventions: removing middle-layer components identified by our interaction-based criterion produces substantially larger downstream changes and accuracy drops than removing components from other regions or at random. Together, these results suggest that abstract reasoning is not uniformly distributed across transformer layers, but is preferentially formed in a middle-layer computation stage that converts token-level information into reusable relational structure.

cs.AI

Grokking From Abstraction to Intelligence

Grokking in modular arithmetic has established itself as the quintessential fruit fly experiment, serving as a critical domain for investigating the mechanistic origins of model generalization. Despite its significance, existing research remains narrowly focused on specific local circuits or optimization tuning, largely overlooking the global structural evolution that fundamentally drives this phenomenon. We propose that grokking originates from a spontaneous simplification of internal model structures governed by the principle of parsimony. We integrate causal, spectral, and algorithmic complexity measures alongside Singular Learning Theory to reveal that the transition from memorization to generalization corresponds to the physical collapse of redundant manifolds and deep information compression, offering a novel perspective for understanding the mechanisms of model overfitting and generalization.

cs.AI

The process of 3D-printed skull models for the anatomy education

Objective The 3D printed medical models can come from virtual digital resources, like CT scanning. Nevertheless, the accuracy of CT scanning technology is limited, which is 1mm. In this situation, the collected data is not exactly the same as the real structure and there might be some errors causing the print to fail. This study presents a common and practical way to process the skull data to make the structures correctly. And then we make a skull model through 3D printing technology, which is useful for medical students to understand the complex structure of skull. Materials and Methods The skull data is collected by the CT scan. To get a corrected medical model, the computer-assisted image processing goes with the combination of five 3D manipulation tools: Mimics, 3ds Max, Geomagic, Mudbox and Meshmixer, to reconstruct the digital model and repair it. Subsequently, we utilize a low-cost desktop 3D printer, Ultimaker2, with polylactide filament (PLA) material to print the model and paint it based on the atlas. Result After the restoration and repairing, we eliminate the errors and repair the model by adding the missing parts of the uploaded data within 6 hours. Then we print it and compare the model with the cadaveric skull from frontal, left, right and anterior views respectively. The printed model can show the same structures and also the details of the skull clearly and is a good alternative of the cadaveric skull.

cs.OH