Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 919 records · Page 51Linked to original sources

Localization Transition in Kinetically Deformed one-dimensional Aubry-André Model

We propose a $q$-deformation in the single-particle kinetic energy and investigate how it modifies the localization in the one-dimensional Aubry-André (AA) model. We construct a Hermitian $q$-deformed kinetic operator as a nonlinear function of the lattice translation operator, preserving the uniform lattice and recovering the conventional AA Hamiltonian continuously in the undeformed limit $q\to1$. The deformation generates an infinite set of correlated odd-range kinetic processes, controlled by a single parameter q, rather than phenomenologically involving long-range hopping. Under the dual transformation, this long-range hopping appears as higher harmonics of the dual quasiperiodic potential, providing a controlled route for breaking the exact self-duality of the AA model, with $q$ as the control parameter, consequently modifying the localization structure for $q\neq1$. In contrast to the conventional AA model, where all eigenstates localize simultaneously at $λ_c=2$, the deformed model exhibits a fraction of delocalized states even beyond $λ_c =2$. An intermediate regime also emerges in the $q-λ$ plane where localized and extended eigenstates coexist across the spectrum. We also propose a possible experimental realization of the hierarchy produced in the $q$-deformed kinetic setting in a periodically driven AA model.

cond-mat.stat-mech↗

Derivation of the General Solution of the Black-Scholes Boundary-Value Problem

There are infinitely many functions that satisfy the Black-Scholes partial differential equation and the terminal condition corresponding to the European call option. This means that the Black-Scholes formula, which led to the award of the 1997 Nobel Prize in Economic Sciences, is not the unique solution, as was once assumed. Consequently, it violates the law of one price, one of the fundamental laws of economics and finance. In this article, we present a rigorous derivation of these solutions to the Black-Scholes boundary value problem.

cond-mat.stat-mech↗

Theory on Attention Dynamics for Out-of-Distribution In-Context Learning

Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. However, their performance often deteriorates when encountering out-of-distribution (OOD) inputs that deviate from the training distribution, and the underlying theory remains poorly understood. To fill this gap, we characterize the OOD error under the input distribution shift through the interplay between the dynamics of the so-called $α$-type and $β$-type attention weights, which represent the transformer's confidence in identifying the correct and incorrect features, respectively. Our results indicate that the OOD error for each feature depends on all pairwise interactions between the training features and OOD features, and under certain cases the transformer performs no better than random guessing. To improve the OOD generalization performance, we next investigate the impact of model finetuning with the OOD data, and particularly, characterize the model forgetting performance on the source domain. Interestingly, the performance on the source domain may not always degrade after finetuning, which highly depends on the nature of the feature shift: finetuning on OOD domain keeps enhancing the confidence of identifying correct features from the original distribution, while the interference from other incorrect features may either increase or decrease. Extensive experiments on both synthetic and real data are conducted to corroborate the theoretical insights.

cs.LG↗

Quantifying the Impact of Upright Patient Positioning on Cardiac Substructures Using Deep Learning

Upright patient positioners with diagnostic-quality vertical CT at treatment isocenter may improve image-guided radiation therapy (RT). However, cardiac substructure (CS) geometry in upright patients remains insufficiently characterized. This work evaluated if a supine-trained deep-learning (DL) CS segmentation model generalizes to upright CT images and quantified posture-dependent CS changes in thoracic patients, to assess potential CS-sparing benefits. 8 thoracic proton therapy patients underwent paired supine/upright 4DCT imaging. 20 CS and lungs were manually labeled on both datasets for positional comparisons. A previously developed supine-trained, DL pipeline generated the same CS, and performance was evaluated using Dice similarity coefficient (DSC) and 95% Hausdorff distance (HD95). Upright CTs were rigidly registered to corresponding supine CTs by aligning the thoracic vertebrae, and CS centroid shifts were measured in the registered coordinate frame and relative to the carina. Paired differences were assessed using Wilcoxon signed-rank tests (p<0.05). The DL model successfully predicted all 20 CS on both upright (DSC, 0.65(0.24); HD95, 7.6(5.6)mm) and supine (DSC, 0.72(0.19); HD95, 5.8(2.6)mm) images yet with lower (p<0.05) performance upright. Upright positioning significantly increased median lung volume by 20.6% (range, -8.9%-42.8%). After vertebral alignment, most CS centroids shifted significantly inferior (median heart shift, 23mm; range, 18-37mm) and anterior (median heart shift, 5.0mm; range, 1.0-13.0mm) when upright. Relative to the carina, most CS shifted significantly inferior and closer anterior-posterior. A supine-trained DL model generalized to upright CT images for CS segmentation. Upright positioning produced increased lung volumes and significant inferior CS displacement, suggesting favorable geometry changes that may support cardiac-sparing workflows.

physics.med-ph↗

Orbital-engineered px,y-kagome lattice in a halogen monolayer

Multi-orbital kagome lattices with explicit orbital degrees of freedom remain largely unexplored, as most experimentally realized systems rely on complex d-electron manifolds that are approximated by isotropic single-orbital models. Here, we overcome this limitation by realizing a px,y-orbital kagome lattice through deposition of a Br monolayer on Ag(111), where orbital filtering selectively suppresses the pz channel. Scanning tunneling microscopy, angle-resolved photoemission spectroscopy, and density-functional-theory calculations reveal a large-area, highly ordered kagome structure whose band dispersions quantitatively match the anisotropic px,y tight-binding model. To extract the intrinsic manifold from the substrate background, we construct an effective H-passivated model, which uncover the intrinsic electronic structure and reveals nontrivial topological characteristics of the px,y kagome manifold driven by first-order spin-orbit coupling effect. Our work establishes Br/Ag(111) as an experimentally accessible platform for multi-orbital kagome physics, extending the kagome paradigm from the conventional d-orbital regime to an orbitally engineered topological setting.

cond-mat.mtrl-sci↗

Invariant Atoms: Sparse Coordinates of Local Semantic Geometry in Language Model Representations

Large language models often preserve meaning despite substantial changes in wording, style, and syntax, while small semantic edits can systematically alter their hidden representations. This suggests that semantic variation may be organized along recurring local directions. We propose the Invariant Atom Hypothesis: local semantic motion admits preferred sparse coordinates along directions that remain stable under meaning-preserving transformations. We learn a shared semantic frame and sparse coordinates that reconstruct semantic displacements while suppressing nuisance variation, with anchor-dependent diagonal modulation adjusting atom strengths without sample-specific rotations. Empirically, the atoms exhibit strong semantic--nuisance separation, sparse reconstruction, reproducible directions, and causal effects on model predictions. The learned geometry generalizes to unseen semantic neighborhoods and nuisance families, while local reweighting improves semantic selectivity and preserves a consistent global-to-local structure. Atom signatures also remain stable under model modification. These findings support reusable invariant directions as a sparse coordinate system for local semantic geometry in language models.

cs.LG↗

Reliable Parallel Decoding in Masked Diffusion Language Models

Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.

cs.CL↗

Channel-Dependent State Space Model for Multivariate Time Series Forecasting

Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and model only temporal dynamics, while channel-dependent (CD) methods consider both but typically rely on architectural compromises to mitigate overfitting and computational overhead. We therefore propose Chameleon, a specialized CD state space model (SSM) that enables data-dependent, fine-grained interactions across variables while scaling linearly with their number. By connecting selective SSMs with the Kalman filter, we leverage the missing measurement update in the former for cross-variable modeling while preserving the SSM backbone for robust temporal modeling. We further identify favorable inductive biases of GatedDeltaNet for time series, adapt it as our backbone, and improve generalization through additional techniques, including a previously unexplored stochastic perturbation of reversible instance normalization. On strongly dependent ODE and PEMS datasets, Chameleon achieves the best MSE and MAE across all settings, while its CI ablation and prior CD methods incur 61-178% higher MSE on average. Across 28 standard benchmark settings, Chameleon also achieves better MSE and MAE than each baseline in at least 27 and 22 cases, respectively. Training-time and peak-memory analyses on Traffic and ETT further demonstrate competitive efficiency and favorable memory scalability across different variable counts.

cs.LG↗

DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction

We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the underlying interaction dynamics insufficiently constrained. We propose DynamicHOI, a physics-aware HOI reconstruction framework combining geometry-grounded diffusion refinement with coupled hand-object dynamics. Geometry spatially grounds visual evidence for trajectory refinement, while articulated inverse dynamics and Newton-Euler dynamics derive hand generalized forces and object wrenches for dynamics-level supervision. We further couple hand and object dynamics through contact-force transfer and recover active hand actuation as an interaction-level physical quantity. We formulate its empirical magnitude distribution into a probabilistic prior that penalizes unlikely actuation and suppresses mechanically implausible reconstructed motion. Experiments on three HOI datasets show consistent improvements in both hand and object reconstruction. The reconstructed trajectories further benefit downstream applications including hand world-model generation and robotic manipulation learning, demonstrating the value of physics-aware HOI modeling beyond reconstruction.

cs.CV↗

Emergent phases of superposition: from partial to full representation

Large language models are thought to represent features by vectors in a hidden space of dimension given by the model's width. Superposition, in which more features are represented than the width by letting representation vectors overlap, is a leading account of how representation vectors are organized. However, how model width and data statistics determine the configuration of representation vectors and the resulting loss when the number of features and the width are large remains less understood. Here we show, in Anthropic's toy model of superposition, that increasing the width drives a continuous phase transition from a partial-representation phase, where only a subset of features receives appreciable representation vectors while the rest vanish, to a full-representation phase, where every feature is represented. Our theory via a partial random projection approximation predicts, and experiments confirm, that the critical width grows linearly with the number of active features up to a logarithmic factor. The loss scaling changes across the transition: below the critical width, the loss grows linearly with the number of active features and depends weakly on the width in a form set by data statistics; above it, the loss grows approximately quadratically with the number of active features and decays inversely with the width. Non-uniform firing probabilities delay the transition and lower the loss, as more frequent features occupy more space. Our results provide an account of how model width and data statistics jointly shape representations and loss, a step toward understanding representation scaling in large models.

cs.LG↗

Nonexistence of maximal curves of genus five over $\F_{64}$

We show that there is no maximal curve of genus five over $\F_{64}$. As a consequence, $N_{64}(5)=140$, and the genus spectrum of maximal curves over $\F_{64}$ is determined. The proof uses the vanishing of the third iterate of the Cartier operator. We prove that a nonhyperelliptic curve of genus five in characteristic two satisfying this condition is nontrigonal. Its canonical theta characteristic defines a separable cover of degree four with one geometric branch value. The two possible ramification types give either a rational subcanonical point or a point bound obtained from the cubic resolvent. Both cases exclude maximality over $\F_{64}$.

math.AG↗

Memory Consolidation Flattens the Temporal Shape of User Facts

Long-term memory systems turn conversations into short stored notes. A note can keep a user fact while losing evidence about whether the fact still holds. For example, "I am driving a Peugeot" can become "The user drives a Peugeot," which drops the cue that the activity is ongoing. We call this aspectual flattening and measure it with LAPSE, a benchmark of matched user statements that differ only in temporal form. We find that memory writers flatten aspect selectively. Three writer models flattened the progressive statement but kept its simple-present match in 244 of 381 pairs, never the reverse. The asymmetry holds in all 11 model configurations tested and in the installed pipelines mem0, Graphiti, and Letta. The lost cue matters to later readers. In exploratory tests, changing only the stored verb shifted all three readers' estimates that a fact still holds. When readers could ask the user before acting, two of three acted without asking more often on flattened notes. Our planned memory-use task could not detect this, because readers there acted on almost every stored fact, even expired ones. Memory writing can thus remove evidence that later models use to decide whether to act.

cs.CL↗

Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models

Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions, raising a basic question: what local metric best captures semantically consequential variation? We propose Fisher-induced invariant representation geometry (Fisher-IRG), which measures local representation directions through their predictive sensitivity. Around each representation, we construct semantic-preserving and semantic-changing neighborhoods, aggregate their local Fisher information, and recover invariant directions through a contrastive generalized eigenvalue problem. Controlled displacement analyses first show that comparable Euclidean motion can have substantially different predictive consequences, supporting the need for a predictive geometry. Across language and vision models, Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions. Representation interventions further localize semantic effects to the Fisher-derived subspace, and held-out separation and retrieval show that the recovered geometry generalizes beyond the discovery neighborhoods. These results support Fisher-IRG as a principled framework for characterizing local invariant representation geometry.

cs.LG↗

Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension

Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as skip-gram have been used to learn chord embeddings from symbolic corpora, but their ability to recover tonal structure and its relation to tonal tension remains underexplored. In this work, we investigate how skip-gram chord embeddings reflect tonal structure and whether they provide a useful basis for analyzing structural aspects of tonal tension. Using chord sequences with and without transposition-based augmentation, we evaluate the learned spaces from geometric, functional, and tension-related perspectives. We show that augmented embeddings exhibit strong transposition equivariance, recover a clear circle-of-fifths structure, and support interpretable shifts between key-related regions of the learned space. We then derive embedding-based measures from chord-to-key distance and contextual chord-distance relations, and show that they capture meaningful aspects of tonal tension structure through correspondence with matched tonal measures and moderate alignment with human tension profiles. Across analyses, transposition-based augmentation generally improves the stability, tonal coherence, and interpretability of the learned space.

cs.SD↗

Rethinking Reasoning Paths as Phase-Structured Trajectories

Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.

cs.AI↗

Everything Everywhere All At Once: A Hierarchical Framework for Mapping Anisotropies in the Local Universe

We present a Bayesian hierarchical forward model for mapping the local luminosity-density field from heterogeneous redshift surveys. By dividing the survey volume into angular cells and redshift shells, we model the observed galaxy distribution directly in apparent-magnitude and redshift space. This allows us to infer cell-level luminosity functions and densities while preserving each survey's unique selection function. Simultaneously, the ensemble of cells constrains the cosmic-mean luminosity function, its redshift evolution, and the hyperparameters describing cell-to-cell variation. We apply the framework to 2MRS, 6dFGS, and GAMA over $0.005<z<0.065$, combining wide sky coverage with deeper and fainter galaxy samples. The model successfully recovers global $J$-band luminosity-function parameters and a mean luminosity density consistent with previous low-redshift measurements. The inferred luminosity-density scatter decreases with volume, consistent with the cosmic-variance amplitude expected in a $Λ$CDM universe. Within the volume we probe, we find no evidence for a large coherent underdensity. Instead, the local Universe is broadly consistent with smooth evolution and cosmic variance. The nearest shell at $z \simeq 0.01$ is the clearest exception, lying $\simeq 0.12$~dex ($\simeq 25\%$) below the smooth global model, a $-2.6σ$ underdensity with a one-sided significance of $\simeq 99.5\%$. This framework offers a scalable architecture for mapping cosmic density fields with next-generation wide and deep surveys.

astro-ph.CO↗

Moments of the Cross-Sectional Siegel-Veech Transfrom

We study the Siegel--Veech transform on the Poincaé section for the horocycle flow consisting of lattice surfaces with a visible horizontal holonomy vector of length at most one. We compute its first moment and derive formulas for higher and factorial moments. We emphasize these formulas in the case of the square torus of unit area and, as an application, use its first moment formula to recover the Boca--Zaharescu pair-correlation density for Farey fractions.

math.DS↗

Degree Balance as a Fine-Grained Complexity Boundary for Quantum SAT

The local Hamiltonian problem is the canonical $\mathsf{QMA}$-complete problem, and $O(2^n)$ time classical algorithms and $O(2^{n/2})$ time quantum algorithms are known to solve the problem in the worst case. It is not clear how to improve these brute force strategies for a broad class of the problem because ground states are highly entangled in general, and we cannot directly apply known strategies for classical CSPs. In this work, we present exponentially faster classical and quantum algorithms under two mild assumptions: (1) the Hamiltonian is frustration-free on YES instances, and (2) it is approximately regular, meaning that every qubit is acted upon by approximately the same number of constraints. We complement these upper bounds by showing that, assuming (Q)SETH, quantum 5-SAT admits no non-trivial worst-case speedup. Our lower bound further demonstrates that the dependence of our algorithms on regularity is in some sense nearly optimal. Specifically, quantum 5-SAT remains (Q)SETH-hard even for Hamiltonians in which all but $O(\sqrt{n})$ qubits participate in only constantly many constraints, while the remaining $O(\sqrt{n})$ qubits each participate in $O(\sqrt{n})$ constraints. By contrast, if either the size of this high-degree subset or the degrees of its qubits is reduced by a factor of $n^δ$, for any $δ>0$, our algorithm solves the problem in time $O(2^{(1- \varepsilon)n})$ for some $\varepsilon>0$. Together, our upper and lower bounds establish a fine-grained complexity dichotomy for quantum satisfiability.

quant-ph↗