Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Global existence of weak solutions to a nematic electrolyte model in 2D

In this paper, we study a system of partial differential equations modeling the dynamics of nematic electrolytes on a two-dimensional torus $\mathbb{T}^2$. The model couples the Poisson--Nernst--Planck equations for the evolution of ion concentrations and electrostatic potential with the Ericksen-Leslie equations for the flow and orientation of nematic liquid crystals. We prove the global existence of weak solutions to this system. The proof relies on a Ginzburg-Landau approximation scheme. Key steps include deriving uniform energy estimates and establishing the strong convergence of the director field by utilizing a Pohozaev-type identity to handle the lack of compactness in the Ericksen stress tensor.

math.AP↗

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however, attenuates this change anisotropically: much of the change encoded in the teacher's hidden states reaches the logits at a small fraction of its weight, and the sampled-token log-probability ratios on which output-space extrapolation relies inject noise that the extrapolation amplifies, making training unstable. We observe that reinforcement learning (RL) shifts a model's internal representations relative to its base checkpoint, and that the direction of this shift can be measured at every layer. Motivated by this observation, we propose RIDE (RL-Induced Direction Extrapolation), which extrapolates the RL-induced change directly in representation space: at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher, which makes explicit how the objective moves the student along the RL-induced direction while limiting its deviation from the teacher. Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base. Project page: https://github.com/xixixixixxxx/RIDE.

cs.LG↗

From Reconnaissance to Response: Quantitative Risk Parameterization and Game Theoretic Containment in Modern Enterprise Attack

Modern Security Operations Centers struggle with delayed manual incident response, enabling adversaries to advance through the Cyber Kill Chain during early stage reconnaissance. While classical game theoretic defense models optimize strategic resource allocation, they rely on static utility matrices that fail to adapt to dynamic telemetry. This paper presents an integrated, metrics driven decision engine that bridges quantitative risk parameterization and continuous automated response time. Common Vulnerability Scoring Systems exploitability parameters are mapped to attacker success probabilities and evaluate defender log distributions via Factor Analysis of Information Risk Monte Carlo simulations. Real time SIEM logs streams are modeled as Poisson process arrival rates, dynamically updating defender posterior threat belief through sequential Bayesian filtering. A closed form threshold is derived by framing the interaction as a dynamic Bayesian Stackelberg game, where the expected unmitigated risk exceeds proactive containment cost. Parameterized against empirical data from the 2023 MGM Resorts and Caesars Entertainment cyber incident, simulation results demonstrate that the engine suppresses transient background noise while triggering automated SOAR network isolation within seconds of adversarial probing. Multi parameter sensitivity analysis confirms that the decision boundary dynamically adjusts to live perimeter vulnerability, offering a control theoretic foundation for sub minute automated threat containment.

cs.GT↗

Optimal Multi-Reward Reinforcement Learning

We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $ε$-optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error $V_{0}^{*, m} - V_{0}^{\widehatπ^{m}, m}$ where $m\in [M]$ represents the reward function and $V_{0}^{*, m}=\mathbb{E}_{s_1\sim μ}[V_{1}^{*, m}(s_1)]$. Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of $$ O\left(\frac{SAH^3}{ε^2}\log M \mathrm{polylog}\left(\frac{SAH\log M}{\min\left\{ε, 1\right\}δ}\right)\right)$$ episodes, with no additional burn-in cost. This matches the information-theoretic lower bound up to a factor of $ \mathrm{polylog}(SAH\log M/(\min\left\{ε, 1\right\}δ))$. Our method combines three technical ingredients. First, we adapt MVP to reward-switching learning to construct optimistic value estimates. Second, we use fresh replay samples to conservatively evaluate the candidate policies. Third, gap-based multiplicative weights updates adjust the reward-sampling distribution using the differences between these estimates, converting weighted learning progress into simultaneous guarantees for all rewards.

cs.LG↗

TT-FDTD: Tensor Train Accelerated Three-Dimensional FDTD With Logarithmic Cost of Spatial Operators

Quantized tensor-train (QTT) compression is incorporated into a full-vector three-dimensional scattered-field finite-difference time-domain (FDTD) formulation on uniform Yee grids. All six electromagnetic-field components, material-dependent update coefficients, equivalent-current sources, and staggered finite-difference operators are represented in compatible QTT form. Gaussian regularization of voxelized material interfaces is used to reduce the coefficient ranks generated by abrupt dielectric and conductivity transitions. The formulation is evaluated for an anatomically heterogeneous human-head model and a homogeneous dielectric sphere on grids containing up to $512^3$ spatial cells. The reported results show that interface smoothing substantially reduces material-coefficient ranks and that the TT--FDTD solution reproduces the full-grid transient fields with pointwise absolute errors on the order of $10^{-4}$ in the examined slices. Compared with conventional FDTD, the tensor representation greatly reduces storage at fine discretizations, although tensor contractions and recompression introduce additional per-step computational cost. These results demonstrate the feasibility and memory--time tradeoff of QTT-accelerated three-dimensional FDTD for large structured-grid simulations.

physics.comp-ph↗

AdaptArena: Evaluating Test-Time Personalization of Web Agents

Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent preferences from implicit signals. Despite its importance for deployment, this problem setting is largely underexplored in existing benchmarks. To address this gap, we introduce AdaptArena, a benchmark for evaluating test-time personalization of web agents via implicit preference inference. AdaptArena consists of 480 tasks, featuring both single-preference and double-preference scenarios. Each evaluation task must be solved by retrieving and leveraging the most relevant historical user trajectory that implicitly encodes the target preference. In addition, we introduce AdaptiveAgent, a retrieval-based framework for standardized evaluation of implicit preference inference. Experiments reveal a substantial performance gap: while oracle agents with access to ground-truth preferences achieve an 82.92% success rate, the evaluated LLM agents using our framework reach at most 15.62%. Furthermore, we find that correctly inferring user preferences is necessary but not sufficient for task success, as execution failures in downstream web interactions remain a significant bottleneck even when agents align with the target preference. These findings highlight implicit preference inference and robust action grounding as key challenges for deploying reliable, user-facing web agents. We release our code: https://github.com/McGill-NLP/web-agents-test-time-adaptations

cs.LG↗

Effects of Longitudinal Spin Current Density Gradient on Spin-Orbit Torque Switching of Perpendicular Magnetization

It has remained a critical open question as to whether a longitudinal spin current gradient (e.g., due to a gradient in the thickness, composition, or width) can replace the longitudinal magnetic field required for deterministic spin-orbit torque switching of normal metal/ferromagnet heterostructures with perpendicular magnetic anisotropy. Here, we report robust micromagnetic and experimental evidence that any realistic longitudinal spin-current density gradient cannot be an effective replacement for a longitudinal magnetic field to enable deterministic switching of a perpendicular magnetic anisotropy device. Instead, the longitudinal spin current density can only modify the nucleation and pulse-timing-sensitive oscillations of magnetic domains and thus the picosecond-scale strip-like indeterministic switching windows. The same conclusions hold robustly when the transverse effective field associated with the spin-current pulse (the sum of the Oersted field and field-like torque) is taken into account. The indeterministic switching prevents applications in the presence of finite device-to-device variations and drifts in write current pulse and thermal fluctuations of device parameters. We also experimentally show that, instead of a longitudinal spin current density gradient, perpendicular spins are much more effective in deterministic switching of perpendicular spin torque devices.

cond-mat.mes-hall↗

LLMs Learn to Evade Latent Monitors from Prior Feedback Alone

Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated. We ask whether an agent can infer the monitor's decision rule from this feedback and then selectively edit its activations to evade detection. Unlike prior evasion attacks, the model is never explicitly told what the monitor detects. Surprisingly, off-the-shelf models already produce activation edits aligned with the monitored direction, but at insufficient magnitude for evasion. Simply scaling up these edits by a factor of 8 reduces the monitor's TPR from 100% to 27%. A rank-1 LoRA amplifies this behavior into effective evasion within the forward pass, reducing TPR further to 4% on held-out concept monitors while leaving other concepts at their normal detection rates. Capabilities on standard benchmarks are retained under this finetuning, and the evasion skill survives retraining the monitors on the new activations. Mechanistically, we find evidence that the model computes its activation edit from the prior in-context turns, and show that the edit becomes more aligned with the monitored direction as more examples are provided. These results demonstrate feedback-conditioned control over activations and suggest that latent monitoring should be treated as an interactive process in which agents can observe and respond to oversight measures.

cs.LG↗

On the Willmore energy of flat $n$-tori in $\mathbb{R}^N$ and Chen's conjecture for $n$-tori

This paper establishes the sharp lower bound $(4nπ^2)^{n/2}$ for the Willmore energy $\mathcal{W}$ of flat $n$-tori in the Euclidean space. Up to Möbius transformations, the Clifford $n$-torus $\mathbb{S}^1\bigl(\sqrt{1/n}\,\bigr) \times \cdots \times \mathbb{S}^1\bigl(\sqrt{1/n}\,\bigr) \subset \mathbb{S}^{2n-1} \subset \mathbb{R}^{2n}$ is shown to be the unique minimizer attaining this bound. This also confirms Chen's conjecture for flat $n$-tori. However, when $n \geq3 $, we show that Chen's conjecture fails on the total mean curvature of general immersed $n$-tori: certain Möbius transformations of the Clifford $n$-torus strictly decrease the total mean curvature.

math.DG↗

Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics

We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.

cs.CV↗

Enabling the Ambient Pressure Growth of ScB2 Crystals for AlGaN Power Electronics

Here we report the growth of single crystalline ScB2, an ultrahigh-temperature ceramic, at ambient pressure in a laser-heated Optical Floating Zone via the travelling solvent method. Crystals have been grown from both Sc-rich (55-65 at% Sc) and B-rich self-flux (80-83 at% B) at growth rates in the range of 0.2-2 mm/hr. The structure of grown crystals is in good agreement with an AlB2-type layered hexagonal phase, space group P6/mmm, with lattice constants a = 3.1423(2) Å (resp. 3.1502(3) Å) and c = 3.5084(3) Å (resp. 3.5041(3) Å) for crystals grown under Sc-rich (resp. B-rich) conditions. Crystals natively grow along the in-plane [100] direction. Electron backscattered diffraction shows that Sc-flux growth results in boules with multiple domains containing Sc inclusions, with the domains highly aligned. In contrast, B-flux boules are single domain after the initial nucleation region. Rocking-curve measurements of B-flux crystals for the (h000) and (000l) reflections show single peaks, establishing that the crystals are free from grain boundaries; the asymmetry in the scattered-intensity tails suggests the presence of point defects. Surface X-ray photoemission spectroscopy shows that the electronic environment in B-flux crystals is superior to that of Sc-flux crystals and produces highly resolved binding-energy peaks for B 1s and Sc 2p. Work-function measurements for the (11-20) plane give a value of approximately 5 eV, consistent with the highly electrically conductive nature of ScB2. These results demonstrate the viable ambient-pressure growth of ScB2, establish it as a lattice-matched substrate candidate for Al-rich AlGaN power microelectronics, and show that this growth route enables scalable manufacturing of ScB2 substrates.

cond-mat.mtrl-sci↗

Know the Normal, Track the Attack: Context-Grounded and Stateful LLM Investigation over System Provenance

Provenance-based intrusion detection systems (PIDSs) identify suspicious activity in audit streams, but their outputs remain difficult to turn into coherent attack narratives. Direct LLM analyses of local anomalous subgraphs lack deployment-specific normal-behavior knowledge and validated attack state across evidence fragments. This can cause unsupported attack interpretations of routine activities and incorrect attribution of temporally dispersed evidence to attack stages. We present ANCHOR, an investigation-oriented provenance system that combines evidence curation with context-grounded LLM reasoning. It calibrates anomaly judgments by relation type and links anomalous windows through rare relation-role patterns. The resulting evidence queues preserve causal structure, temporal boundaries, and cross-window continuity. The investigator interprets process-centered evidence using two complementary forms of context. Deployment Context combines environment-specific interaction and object baselines with high-risk security knowledge. Case Context uses a confidence-gated Attack-Tracking Cache to maintain investigation state across windows. Correlating current evidence with high-confidence prior findings, ANCHOR incrementally reconstructs attack narratives organized by kill-chain stages. We evaluate ANCHOR on six DARPA Transparent Computing E3/E5 datasets across three operating systems. Controlled evidence-level and end-to-end comparisons show improved overall IoC recovery and attack-stage attribution over state-of-the-art provenance-based baselines. These gains persist under a fixed LLM backbone in our evaluation. ANCHOR processes a full audit day at dollar-level API cost, supporting practical, context-grounded investigation across windows.

cs.CR↗

Closed-Form Cartesian Forward Kinetostatics for Spatial Multi-Segment Tendon-Driven Continuum Robots

Forward kinetostatics of spatial tendon-driven continuum robots typically requires a nonlinear equilibrium solve for each actuation input. This paper develops a force-to-Cartesian-configuration model with a closed-form solution in quadratures for spatial multi-segment robots under tendon actuation. The Cartesian backbone centerline and accumulated material twist serve as generalized coordinates, from which the strain measures and tendon geometry are derived. Variational equilibrium yields explicit axial and bending relations and establishes zero equilibrium material twist within the proposed model for admissible longitudinal non-helical routing. The solution is propagated segment by segment without an iterative equilibrium solve, while retaining axial deformation, spatially varying axial and bending stiffnesses and tendon-routing diameter, and segment-dependent tendon participation. Numerical comparisons with a full-strain geometric variable-strain model (GVS) yield maximum length-normalized tip-position discrepancies of 8.91 x 10^-6 and 1.01 x 10^-5 for the single- and three-segment robots, respectively. Mean evaluation times of 1.52 μs and 2.94 μs, with corresponding speedups of approximately 1864x and 3348x over the baseline, demonstrate the computational advantage of the explicit force-to-configuration mapping in the reported benchmark.

cs.RO↗

Reimagine Video Dynamics

Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context. We learn this token through self-supervised reconstruction: given the first frame as visual context, a renderer must recover the original video from the dynamics token, encouraging it to capture how the scene evolves rather than how it looks. This disentanglement allows video dynamics to be edited directly while preserving visual context. We develop a language-guided dynamics-token editor that transforms source dynamics into target dynamics, and train it with a scalable counterfactual video-pair pipeline and a two-stage training strategy. Extensive experiments show that RVD enables effective video dynamics editing, training-free retiming, and appearance-controlled re-rendering.

cs.CV↗

Evanescent-wave Johnson Noise from Superconductors

We compute the evanescent-wave Johnson noise (EWJN) in the vacuum half-space above a superconductor, and the resulting relaxation time ($T_1$) of spin and charge qubits placed at nanometer distances from the surface. The electromagnetic response is described by a single microscopic transverse current-response kernel $Q(q, ω)$ for a BCS superconductor. This is computed for varying densities of both non-magnetic impurities and magnetic impurities, for arbitrary frequency and temperature and for wave vectors $q \ll k_F$ (the Fermi wavevector). When combined with the fluctuation-dissipation theorem and the nonlocal surface impedances of the half-space, this yields the magnetic and electric field noise at any distance $z \gg k_F^{-1}$ from the surface, from which we obtain $T_1$. Just below $T_c$ the magnetic noise is enhanced relative to the normal state by the coherence (Hebel-Slichter-type) peak of the dissipative conductivity and drops exponentially at lower temperatures; the electric noise shows no coherence peak. The theory predicts that there is a zero-temperature noise floor induced by magnetic impurities. In the gapless regime produced by pair breaking, the finite subgap density of states $ν(0)$ yields a temperature-independent noise spectral density and a relaxation rate bounded by $T_1^{-1}(T)\le[ν(0)/ν_F]^{2}\,T_{1,N}^{-1}(T)$ for $T\ll T_c$, with equality in the extreme nonlocal regime. Here $ν(0)$ and $ν_F$ are the superconducting and normal-state densities of states at the Fermi energy, and $T_{1,N}(T)$ is the relaxation time the same electrode would produce in its normal state at the same temperature.

cond-mat.supr-con↗

A Polynomial Time Characterization For Strongly EFX Orientable Graphs

Discrete fair division is the problem of dividing a discrete set of goods among agents in a fair manner. In this setting, one of the most sought-after notions of fairness is envy-freeness up to any good (EFX). In 2023, Christodoulou, Fiat, Koutsoupias, and Sgouritsa introduced the idea of a graphical valuation, where the fair division problem is represented by a simple graph where vertices are the agents and the edges are the goods, and each vertex only values incident edges. They showed that an EFX allocation always exists, while determining the existence of an EFX orientation is NP-hard. They posed an open question of determining which graphs always admitted an EFX orientation regardless of valuation. These graphs, called strongly EFX orientable graphs, were first studied by Zeng and Mehta in 2025, who demonstrated that all such graphs have chromatic number at most 3, and bipartite graphs always admit an EFX orientation regardless of valuation. In this manuscript, we finish resolving this question by giving a polynomial time characterization of strongly EFX orientable graphs. In particular, we show that a connected graph $G$ is strongly EFX orientable if and only if either of the following is true: (1) $G$ is bipartite, or (2) the block decomposition of $G$ contains exactly one nonbipartite block $B$, and there exists a vertex $v \in B$ such that the degree of $v$ within $B$ is 2 and $G-v$ is bipartite. This proof was discovered by AI, with human intervention to break the problem into the appropriate subproblems.

cs.GT↗

Gorenstein homological properties of n-trivial extensions of rings

We explicitly investigate Gorenstein projective, injective and flat modules over the $n$-trivial extension $R\ltimes_{n}M$ of a ring $R$ by an $R$-bimodule $M$. Assume that $fd(M^{\otimes_{R}i}_{R})<\infty$ and $pd(_{R}M^{\otimes_{R}i})<\infty$ for any $1\leq i\leq n$, $fd(\textbf{Z}(R)_{R\ltimes_{n}M})<\infty$ and $pd(_{R\ltimes_{n}M}\textbf{Z}(R))<\infty$. It is proven that a left $R\ltimes_{n}M$-module $(X,f)$ is Gorenstein projective if and only if the sequence $M^{\otimes_{R}n+1}\otimes_{R} X\stackrel{(M\otimes f) \cdots(M^{\otimes_{R}n}\otimes f)}\longrightarrow M\otimes_{R} X\stackrel{f}\longrightarrow X$ is exact and coker$(f)$ is a Gorenstein projective left $R$-module. As a consequence, we characterize Gorenstein projective, injective and flat modules over tensor rings.

math.RA↗

InterBias-SV: Compound Conditions in Speaker Verification

Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trained encoders, their mean contrasts are +0.0026, +0.0088, and +0.0024 in equal error rate (EER), with larger variation across settings. These descriptive averages do not establish equivalence to additivity: trial matching, checkpoint identity, and parts of the condition metadata remain unverified. We also examine two interpretation problems. Near-chance EER can make additive predictions difficult to interpret, but chance performance is not a hard EER ceiling, and correlation with the prediction does not identify a saturation mechanism. Ratios of demographic gaps are unstable when their clean reference is near zero; absolute gaps provide a more direct summary. The benchmark provides condition definitions, analysis scripts, and explicit requirements for interpretable compound-condition comparisons, while separating recomputable summaries from claims that require further experimental validation.

cs.SD↗