Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 847 records · Page 47Linked to original sources

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

Mixed-reality headsets such as Apple Vision Pro replace the touch screen with gaze-as-pointer interaction: the wearer looks at a target and confirms with an air pinch. Because the display is inside the headset and the eye tracker is walled off from third-party software, such input is widely assumed to be unobservable to bystanders---a built-in defense against the shoulder-surfing that plagues phones and laptops. We present GAZEleak, a side-channel attack that recovers gaze-driven input from external video of head motion alone, under a strictly local, physical-observer threat model: the adversary only films the wearer from across the room and installs no software on the device. The attack exploits the centrally coupled eye-head motor program: gaze shifts recruit small, target-dependent head reorientations that project into sub-degree pose changes recoverable from commodity video. GAZEleak implements a measurement-based, sparse-optical-flow inference pipeline for users' 6-digit device passcodes. On a preliminary front-view dataset from three author-subjects who were aware of the attack hypothesis, GAZEleak places the true code within the top ten guesses for 56% (10 of 18) of test codes under a cross-person protocol with no labeled victim data, and for every code of the most exposed subject. The performance is subject-dependent: no passcode from the least exposed subject reaches the top ten, although its median guessed passcode rank is 12,786 rather than 500,000 expected from an uninformative ordering. These results provide preliminary evidence that gaze-coupled head motion can expose passcode information under controlled conditions, while motivating broader evaluation across users, behaviors, and capture settings.

cs.CR↗

Quorum sensing with density-enhanced motility and size regulation

Motivated by biological systems and my previous studies of quorum sensing with density-enhanced motility, I study a model of density-enhanced size and motility, in which both particle size and motility increase when the local density (over a lengthscale larger than particle diameter) exceeds the global density. The emergent structures along different size ratios are spots, holes, bands and labyrinths controlled by several distinct observations. First, the characteristic pattern length scale is controlled by the sensing radius and remains approximately independent of the size ratio. Second, the total area occupied by the particles decreases with increasing size ratio. Third, the active fraction barely depends on the size ratio. These results demonstrate how both particle size and motility can couple microscopic regulation to large-scale pattern formation. I conclude by discussing possible extensions of the model.

cond-mat.soft↗

Landau-de Gennes corrections to the Oseen-Frank limit: Anchoring-induced tilt modes

An asymptotic analysis of the Landau--de Gennes framework is performed to compute higher-order corrections to the Oseen--Frank limit in a bounded three-dimensional domain under appropriately scaled surface anchoring energy. A systematic decomposition of the $\mathbf{Q}$-tensor into three mutually orthogonal subspaces-the uniaxial scalar, geometric tilt vector, and transverse anisotropy tensor-reveals that the leading $\mathcal{O}(\varepsilon)$ correction to the Oseen--Frank director field $\mathbf{n}_0(\mathbf{x})$ is dominated by a non-vanishing tilt field $\mathbf{p}_1(\mathbf{x})$, where $\varepsilon$ represents the ratio of the nematic coherence length to the characteristic domain size. This macroscopic variation constitutes a soft mode released from the boundary once the surface anchoring energy is retained at its physical scaling rather than driven to an infinite strength. We show that this tilt field is governed by the linear Jacobi equation, $\mathcal{J}_{\mathbf{n}_0}(\mathbf{p}_1)=\mathbf{0}$, subject to a non-trivial, anchoring-driven Dirichlet boundary condition, where $\mathcal{J}_{\mathbf{n}_0}$ is the on-shell Jacobi operator of the harmonic map $\mathbf{n}_0$ on $\mathbb{S}^2$. The two fields are accompanied at $\mathcal{O}(\varepsilon^2)$ by an off-shell correction to both the uniaxial scalar and the transverse anisotropy tensor, passively induced by the elastic non-uniformity $(\nabla\mathbf{n}_0\neq\mathbf{0})$ and the boundary-driven tilt $(\mathbf{p}_1\neq\mathbf{0})$. Through $\mathcal{O}(\varepsilon^2)$, the tilt enters the energy only through surface terms, not the bulk, providing a pathway for the system to lower its energy . Under the conventional benchmark of rigid Dirichlet conditions, this response is annihilated outright, demonstrating that corrections built upon infinite energy barriers obscure the underlying physics of anchoring-driven tilt modes.

cond-mat.soft↗

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.

cs.LG↗

Potential-energy surfaces of water and its molecular ions up to H$_2$O$^{3+}$

We compute the potential-energy surfaces for molecular ions up to H$_2$O$^{3+}$ with any combinations of outer-valence, inner-valence and core holes. To obtain the potential-energy surfaces, we consider several ab initio techniques based on the CASSCF method and benchmark them with respect to the literature. We thus identify the techniques which produce accurate potential-energy surfaces as a function of the two bond lengths and the angle between the two bond lengths. These potential-energy surfaces will be used in future studies of the interaction of water with free-electron laser pulses. Furthermore, the techniques employed in this work in the context of water can be used to obtain the potential-energy surfaces of other triatomic molecules.

physics.atom-ph↗

Robust identification of drive-by sensing ride-hailing market with Points of Interest monitoring under uncertain ride demand

Taxi-based mobile sensing has emerged as a cost-efficient paradigm for large-scale urban environmental monitoring. In practice, both passive and active sensing strategies are adopted by ride-hailing platforms. Passive sensing, conducted during passenger-serving trips, is constrained by stochastic and spatially imbalanced ride demand, leading to limited and uneven coverage. Active sensing, executed by vacant taxis, provides greater control over sensing operations but incurs additional operational costs, and is thus typically treated as a supplementary strategy. To address the inefficiencies of the conventional "passive-first, active-second" paradigm, which may delay the monitoring of critical Points of Interest (POIs) and increase system-wide costs, we propose a Distributionally Robust Optimization (DRO)-based framework for active mobile sensing. First, we develop an enhanced A*-based drive-by sensing routing policy that integrates vehicle-task matching while capturing both global routing efficiency and local sensing opportunities. Second, we formulate the active sensing problem as a DRO model that explicitly accounts for uncertainty in ride demand through an ambiguity set of probability distributions, enabling robust decision-making for vacant taxi routing and vehicle-task matching. The proposed framework is evaluated on both static and dynamic mobile sensing settings using real-world data. Computational results demonstrate that our approach achieves better performance in terms of sensing coverage, operational cost, and robustness compared to benchmark strategies, highlighting the value of integrating distributional robustness into taxi-based sensing operations.

math.OC↗

Zero-Shot Reactive Obstacle Avoidance for Generative Robot Policies

We propose NUDGE (Nudge Update via Differentiable GEometry), a training-free obstacle-avoidance procedure that can be incorporated in any robot policy based on diffusion or flow matching, including diffusion policies and vision-language-action models. Our work injects gradients from a signed distance field, a function returning each point's distance to the nearest obstacle, into the policy at inference time to steer it away from obstacles. It supports any common action parameterization, from absolute or relative joint poses to end-effector poses, through a differentiable joint-trajectory decoder. Experiments show that NUDGE preserves the policy's task distribution and runs reactively in real time.

cs.RO↗

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under $23\times$--$64\times$ compression and remains competitive at $144\times$, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to approximately 36% end-to-end speedup and using $16.6\times$/$78.8\times$ lower compressor latency/FLOPs.

cs.CV↗

$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of $λ$-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) $λ$-balance prevents collapse, and (ii) our regularizer applied to the backbone induces $λ$-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations' ranks. We apply SACReg to JEPA and propose $λ$-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, $λ$-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.

cs.LG↗

Who Drives the System? Classifying Mean-Field Particle Systems from Trajectory Data

We study the supervised classification problem for interacting particle systems (IPS) from trajectory observations. The IPS belongs to one of K classes, each characterized by a distinct interaction drift. Given a learning dataset, the goal is then to predict the class label of a newly observed system based on the trajectory of a single particle. This setting raises several statistical challenges. First, the particles within each system interact and are therefore dependent. Second, although the full particle system is Markovian, the trajectory of a single particle is not. To address these difficulties, we exploit the McKean-Vlasov limit of the IPS, which describes the dynamics of a typical particle as the number of particles tends to infinity. We propose a plug-in classification procedure based on estimating the interaction drift associated with each class from discretely observed trajectory data. Hence, for each class, we provide a new nonparametric estimator of the drifts by minimizing a ridge-regularized least-squares contrast over a B-spline basis. In particular, our theoretical findings reveal that the convergence rate of the resulting classifier is of order N -1/6+$ε$ for any $ε$ > 0, where N denotes the number of particles in each system. Numerical experiments illustrate the performance of the proposed method and show that it outperforms an end-to-end neural network baseline that does not exploit the underlying particle-system structure.

math.ST↗

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.

cs.SE↗

Jordan induction for $\mathrm{GL}_n(\mathcal{O})$

Let $\mathcal{O}$ be a complete discrete valuation ring with a finite residue field. We introduce a simple framework of constructing smooth representations of $\mathrm{GL}_n(\mathcal{O})$ modulo the knowledge of nilpotent orbits, called Jordan induction, which enjoys several remarkable properties: It yields only irreducible representations, it yields all the even level irreducible representations, and up to conjugation it yields non-isomorphic representations. In the even level case, this gives an explicit realisation of Hill's analogue of Lusztig's Jordan decomposition, as well as a vast generalisation of Gérardin's construction for $\mathrm{GL}_n(\mathcal{O})$. As a simple application, we construct an explicit section to the orbit map. We also propose a conjecture linking Jordan induction and Lusztig induction, aiming at generalising the algebraisation theorem obtained in recent joint works with Stasinski.

math.RT↗

Finite-Time Blowup for Navier-Stokes with Smooth Forcing, Part I: Construction of Self-Similar Solutions with Admissible Stress and Flat Remainder

This paper gives a readable and accessible version of the profile-construction part of OpenAI's manuscript "Finite Time Blowup for Navier-Stokes". We regard OpenAI's work as a major advance on the Navier-Stokes Millennium Prize Problem. The profiles are smooth and axi-symmetric. Inserting them into the Navier-Stokes equations gives a residual consisting of a divergence-form term and a remainder vanishing to infinite order in $1-t$ on fixed similarity sectors. The associated stress and radial shear satisfy the admissible cone condition. The stress need not be small. We introduce a new linear model that provides a clearer explanation of the inner core construction. The cancellation by oscillatory pulses will be treated in the companion paper "Finite-Time Blowup for Navier-Stokes with Smooth Forcing, Part II: Residual Correction via Oscillatory Pulses".

math.AP↗

Building Transformation Layers for Riemannian Neural Networks

Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclidean geometries. However, previous approaches typically focus on a few selected manifolds and rely on specific properties of the target manifold. In contrast, this work proposes a framework for constructing FC and convolutional layers over computationally tractable Riemannian spaces. This framework incorporates several previous FC layers across different geometries as special cases and is instantiated on ten representative manifolds, including three hyperbolic models, five geometries of the symmetric positive definite (SPD) manifold, and two Grassmannian perspectives. Experiments on different manifolds demonstrate the effectiveness and applicability of our approach. Code can be found at https://github.com/GitZH-Chen/RieTrans.

cs.AI↗

Just Initialize: A Training-Free Initialization Component for Large-Scale Routing Optimization

Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational cost. We instead shift the focus to a useful initialization that can be refined into a high-quality solution with limited downstream refinement. We propose Just Initialize, a training-free and solver-agnostic initialization component for large-scale routing optimization. Just Initialize compresses a large routing instance into a compact surrogate space, optimizes its global routing structure, and recovers the resulting solution as an optimization-friendly starting point in the original space. Extensive experiments on Traveling Salesman Problems (TSPs), Capacitated Vehicle Routing Problems (CVRPs), Vehicle Routing Problems with Time Windows (VRPTWs), and Prize-Collecting Traveling Salesman Problems (PCTSPs) demonstrate that Just Initialize achieves high-quality solutions comparable to or better than state-of-the-art methods while substantially reducing computational cost across instances ranging from 1K to 100K nodes, including an average speedup of approximately 70$\times$, sub-second runtimes on 10K-node instances, and runtimes within tens of seconds on 100K-node instances.

cs.AI↗

NeuronSifter: Intervention Planning in CNS Microenvironments

Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regimen to an identity token or a scalar exposure, discarding where and when the target is engaged; handing a point estimate to a separate planner then discards the joint uncertainty that makes a measurement worth running. We therefore treat decision quality as a property of the intervention interface, not of controller placement. NeuronSifter compiles regimens into state-conditional target-occupancy fields with support masks, propagates them through microenvironment dynamics with an occupancy-conditioned diffusion operator, and selects measurements by their expected reduction in intervention loss, assimilating typed outcomes into the same posterior. In a declared synthetic Alzheimer's disease (AD) evaluation over 64 paired scenario blocks, occupancy conditioning lowers trajectory continuous ranked probability score from 0.165 to 0.110 and raises intervention ordering accuracy from 0.760 to 0.880, and every paired benchmark contrast remains separated after Holm correction. Decision-directed acquisition attains terminal risk 0.160 against 0.166 for a matched numerical Bayesian experimental design planner, and reaches the target risk at 0.796 $[0.732,0.873]$ of an earlier design control's cost, while the corresponding ratio against the matched planner, 0.963 $[0.907,1.025]$, is not separated from equality; point-state and dependence-ablated interfaces instead raise risk to 0.220 and 0.199, and a full-posterior external controller ties exactly. Published AD trials supply a separate retrospective endpoint bridge.

cs.LG↗

Complete classification of $7$-adic Galois images for non-CM elliptic curves over $\Q$

We solve Conjecture 1.6 of Furio and Lombardo [On $7$-adic Galois representations for elliptic curves over $\Q$, Proc.\ London Math.\ Soc.\ (3) \textbf{133} (2026), e70193.]. As a consequence, we complete the classification of $7$-adic Galois images for non-CM elliptic curves over $\Q$. This is the first solved case in Problem 6.2 from Balakrishnan [\emph{Chabauty and beyond: Explicit p-adic methods for rational points}, Proceedings of the International Congress of Mathematicians 2026 - Volume 3: Invited Lectures (Sections 1-4), 321-341, 2026]. The proof is the result of a long AI-training period by the author from June 2026 to September 2026. The author has been working on this project alone without the help of any other.

math.NT↗

From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.

cs.CV↗