Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,225 records · Page 68Linked to original sources

The Reach of Abelian Covers in Hypergraphs

Covers in hypergraphs are frequently studied to capture various forms of dependence between hyperedges. For example, even covers--which check if each vertex appears in an even number of hyperedges--have found much success recently in the study of locally decodable codes. Inspired by a recently-emerging line of work on the non-redundancy of constraint satisfaction problems (CSPs), we introduce and study two novel families of covers of hypergraphs which are stricter than even covers: \emph{Abelian} covers and Catalan covers. Abelian covers are similar to even covers, except that arithmetic is now done over the integers rather than modulo 2, allowing us to capture dependences over arbitrary Abelian groups. Catalan covers capture the behavior of non-Abelian groups by only allowing local cancellations in a sequence of hyperedges. We prove three main results about Abelian and Catalan covers. First, using tools from lattice theory, we show that any $r$-uniform hypergraph with $n$ vertices and $n \log(r)$ hyperedges has an Abelian cover. Second, using tools from algebraic topology, we show that in any $3$-uniform hypergraph, Abelian covers and Catalan covers are equivalent; thereby showing that Catalan covers emerge after $O(n)$ hyperedges in $3$-uniform hypergraphs. Finally, using the theory of nilpotent groups, we show that there exists a $4$-uniform hypergraph which has an Abelian cover but not a Catalan cover. Collectively, these results exactly characterize the reach that Abelian covers have in deducing dependences in hypergraphs. As our primary application, we show that any arity-$3$ CSP with an infinite-domain Mal'tsev extension has linear non-redundancy. This implies near optimal streaming, sparsification, and kernelization algorithms for this family of CSPs. Previously, such a result was only known for the much simpler case of arity-$2$ CSPs.

cs.DM↗

Federated Clustering with Unknown Local and Global Cluster Cardinalities

Federated clustering methods that do not require the global number of clusters $K$ still assume that each client knows its local number $K_g$. This assumption is hard to justify when clients know no more about their data than the server does, as in fault diagnosis across independently operated industrial sites. We propose a two-phase framework in which neither count is known: each client first estimates $K_g$ from its own data, and an aggregator that requires local counts, such as FedGEM, then uses these estimates in place of the true values. For the first phase we introduce Adaptive Split--Merge (ASM), which grows a spherical Gaussian mixture by BIC-driven splitting and then merges excess components. ASM uses no labels, selects its hyperparameters on held-out client data only, and makes no assumption about how clusters are shared across clients. We derive a closed-form split criterion whose critical cluster size falls with anisotropy and rises with dimension, and show empirically that over-fragmentation grows with the number of points per cluster, which federation divides among clients. Across eight datasets, ASM with FedGEM attains a mean ARI of 0.333, against 0.256 for the next best label-free estimator and 0.361 when the true local counts are supplied. It also gives the most reliable global estimates of $K$ and is robust when client size is decoupled from local cardinality.

cs.LG↗

SCORAS-MoE: Joint Compression and Resource-Adaptive Deployment of MoE-VLMs in LEO Satellite Networks

Deploying large vision-language models (VLMs) onboard satellites enables onboard data processing and reduces raw data downlink. However, onboard inference faces two resource challenges. Limited onboard memory and energy require model compression and distributed deployment. Dynamic resource availability requires fast deployment decisions as illumination, battery levels, and communication conditions change. We present SCORAS-MoE, a joint compression and deployment framework for mixture-of-experts (MoE) VLMs in low Earth orbit (LEO) satellite networks. To address limited resources, SCORAS-MoE measures the perturbation of the routed MoE output caused by low-rank approximation, assigns higher ranks to more sensitive experts, and distributes compressed model shards across satellites for cooperative inference. The compressed models yield profiles of measured accuracy and inference energy. To adapt to dynamic resources, the online scheduler selects profile compositions and shard placements in each slot. For each candidate composition, it reduces placement to a minimum-cost assignment problem solved by the Hungarian algorithm, while enumerating the compositions yields the optimal deployment for the current-slot objective. Experiments on Qwen3-VL-30B-A3B-Instruct show that allocating ranks based on output perturbation is particularly effective under aggressive compression, with an absolute gain of $3.7\%$ in mean accuracy over uniform rank allocation when expert projections retain $30\%$ of their original parameters. The fixed-profile scheduler achieves higher throughput with fewer service switches and lower battery impact than the evaluated proximal policy optimization (PPO) and evolutionary baselines, with respective speedups of $8.7\times$ and $183.5\times$. Adaptive profile selection further improves the balance between service quality and energy use.

cs.NI↗

Dirichlet gravitons in AdS$_4$ and $\mathcal{L}_Λw_{1+\infty}$ wedge algebra

We derive the AdS deformation of the flat-space soft-graviton algebra with Dirichlet boundary conditions. Starting from the AdS$_4$ spinor-helicity representation, we Mellin-transform linearized graviton solutions and construct the AdS wedge, including its Laurent completion. The Dirichlet boundary condition pairs opposite helicities, while AdS covariance fixes their relative normalization. Assuming linear closure and normalizing the global modes to reproduce the geometric AdS isometries, covariance and the Jacobi identities uniquely determine the soft-mode bracket. The result is the wedge of $\mathcal L_Λw_{1+\infty}$, where $Λ=-\ell^{-2}$. This provides a bulk derivation of the cosmological-constant deformation that incorporates the AdS boundary condition, without relying on collinear splitting functions or a self-dual truncation.

hep-th↗

Graph-Spectral Flow Matching for Multivariate Time Series Anomaly Detection

Multivariate time series anomaly detection typically relies on evaluating discrepancies between observations and outputs produced by models trained on normal data. An alternative perspective is to characterize the distribution of normal data through the generative dynamics, i.e., the velocity field, of flow matching models. However, standard flow matching typically adopts linear probability paths that overlook dependencies among variables, leading to a misalignment with the structured data distribution. To address this issue, we propose GRASP, a flow matching framework with a graph-spectral path for multivariate time series anomaly detection. GRASP incorporates graph structure into the probability path by minimizing a fixed-endpoint action that combines kinetic energy with graph Dirichlet energy. This formulation yields a closed-form path based on graph-frequency-dependent hyperbolic interpolation. A velocity predictor trained on normal data then detects anomalies using weighted velocity discrepancies aggregated across source samples, flow times, and graph frequencies. Theoretically, we establish that GRASP is invariant to the choice of Laplacian eigenbasis and decompose its expected oracle anomaly score into bounded endpoint uncertainty and graph-frequency-weighted Fisher discrepancy. Experiments on four benchmarks demonstrate the superior anomaly detection performance of GRASP and validate the effectiveness of its graph-spectral path and weighting mechanism.

cs.LG↗

When Can Prefixes Compile LoRA? Exact Resource-Capped Tests for Frozen Attention

Can a fixed continuous prefix replace a given low-rank adapter while the attention head stays frozen? In this research, we show that the answer depends on the adapter's target through three conditions. First, observability: at one causal readout, every independent key--value prefix sees the content only through the query, attention partition, and value numerator, so a target that differs on two inputs with equal summaries incurs an error floor at every prefix length; norm caps extend this floor to nearly equal summaries. Second, realizability: at a common query, any prefix reduces exactly to two aggregate variables, and the norm-capped optimum is an attained second-order-cone program, also after a fixed output projection; it places two equal-norm rank-one value updates on opposite sides of compilability. Third, implementation: under affine query exposure, $2r$ signed slots approximate a rank-$r$ value update, but their values grow as $O(ε^{-3/2})$, and the construction passes all 400 tolerance checks in float64 yet only 38 in bfloat16. A first-layer GPT-2 readout with fixed token and position meets the common-query condition without clamping activations; at three such heads, the capped optimum leaves 18.4\% to 74.2\% of the projected adapter effect uncompiled, with a head-dependent value--query ordering. All claims concern local approximation at one head, not whole-network equivalence.

cs.LG↗

Correspondence between multileaf topology of closed geodesics and spatiotemporal autocorrelations of hotspot images in Schwarzschild spacetime

The classification of relativistic closed orbits by their multileaf structures provides a framework for studying strong-field dynamics. Identifying these structures in astronomical images remains challenging. Using ray tracing, we construct the spatiotemporal autocorrelations of the primary image of a pointlike hotspot moving along bound closed geodesics around a Schwarzschild black hole. These orbits are classified by three integers $(z,w,v)$, where $z$ counts the leaves, $w$ counts the additional whirls during each radial period, and $v$ specifies the order in which the orbital leaves are traced. For the orbit families examined, our numerical results establish a correspondence between the topological integers $(z,w,v)$ and the numbers of correlation bands $N_{\rm band}$ and recurrence points $N_{\rm rec}$. The integers are recovered as $z=N_{\rm rec}-1$, $w=\left\lfloor (N_{\rm band}+1)/(N_{\rm rec}-1)\right\rfloor-1$, and $v=(N_{\rm band}+1)\bmod(N_{\rm rec}-1)$. Here, $\lfloor x\rfloor$ denotes the greatest integer not exceeding $x$, and $\bmod$ denotes the remainder operation. These relations provide a quantitative method for recovering closed-orbit topology from hotspot image autocorrelations.

gr-qc↗

Testing KMTNet--PRIME Optical--Near-Infrared Source-Color Constraints in KMT-2024-BLG-0211 and KMT-2024-BLG-1522

We present analyses of two 2024 microlensing events, KMT-2024-BLG-0211 and KMT-2024-BLG-1522, jointly observed by KMTNet in the optical and PRIME in the near-infrared. For KMT-2024-BLG-0211, the KMTNet--PRIME $I-H$ color provides a useful constraint on the angular source radius for a short-timescale finite-source event with a giant source. The event is consistent with a low-mass stellar lens, although the weak parallax constraint leaves the lens mass and distance uncertain. For KMT-2024-BLG-1522, we compare the results obtained from $V-I$, $I-H$, $V-H$, and $J-H$ color constraints. The microlensing parameters are nearly identical among the four analyses, yielding a robust binary-lens solution with nearly equal masses. However, the inferred source properties and lens physical parameters depend on the adopted source color because the different color estimates imply different angular source radii. Comparing the $(V-I)_{\rm KMT}$ versus $(I-H)_{\rm KMT,PRIME}$ plane shows that the inferred source colors lie off empirical color--color relations but within the scatter of observed stars. The prevalence of such deviations should be tested with a larger sample of KMTNet--PRIME events and this specific case can be further tested with future adaptive-optics follow-up, which can directly measure the lens--source relative proper motion and lens flux.

astro-ph.GA↗

Development and Performance Study of a Capillary Liquid Scintillator Neutron Detector

Capillary liquid scintillator detectors are promising for high-resolution neutron imaging, yet experimental data on their light spread mechanism and spatial performance remain limited. Here, we report a neutron detector based on a hexagonal capillary array filled with EJ-309 liquid scintillator, with an inner diameter of about 50 um and a camera readout of 9 um pixels. Laser experiments show that the FWHM of the full light spot decreases from 260 um to 90 um with a metal light absorber, confirming effective suppression of lateral light spread. Using an AmBe neutron source, an effective field of view with a 5-sigma threshold was established from background frames. For single-capillary events, the pulse height spectrum follows a Landau distribution with a most probable value of 0.133 +/- 0.001 (stat.), and the intrinsic detection efficiency is 10.07% +/- 1.26% (stat.) +/- 1.43% (syst.), corresponding to about 13.55% when normalized to the active liquid scintillator area. The point spread function core yields a radial FWHM of 12.8 um and a centroid positioning precision of approximately 5.5 um (1 sigma), while the intrinsic position resolution is limited by the capillary pitch to 54 um. Linearity is good for 1 to 2 capillaries, with deviation appearing for 2 to 3 capillaries due to additional capture of spread light. These results provide experimental basis and physical understanding for imaging applications of capillary liquid scintillator neutron detectors.

physics.ins-det↗

Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.

cs.AI↗

Beyond Conditional Independence: Root Cause Analysis with Deep Causal Models

Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, existing approaches rely either on heuristic methods or on conditional independence tests with a strong unconfoundedness assumption, and thus fail to exploit other complicated distributional constraints in the presence of latent variables. To relax these assumptions, we model the underlying system as a causal model and the anomalous system as a change in the structural functions of the same causal model. Specifically, to handle unobserved confounders, we establish an implicit connection between distributional constraint testing and root cause analysis. To adapt our approach to data generated from arbitrary causal models, we employ the deep causal model (DCM) framework, in which we design the causal model using neural networks. Finally, we illustrate how our method, RCA-DCM, can utilize different levels of partial graphical knowledge to perform RCA. We evaluate RCA-DCM against state-of-the-art baselines on simulated datasets, a physics-based causal chamber and two micro-service applications. RCA-DCM improves top-1 accuracy over the strongest baseline on both Sock Shop (0.880 vs. 0.752) and Online Boutique (0.776 vs. 0.712), and when the true root cause in the causal chamber is unobserved and acts as a latent confounder, it recovers the exact root-cause set more often than any competing method (perfect recovery rate (PRR) 0.846 vs. 0.731).

cs.LG↗

Quantitative finiteness of monic characters of knots

Dunfield, Friedl, and Jackson showed that if the $SL(2, \mathbb{C})$-character variety of a knot has an irreducible curve component that contains the character of an irreducible representation and a character with nonmonic twisted Alexander polynomial, then this component has only finitely many characters with monic twisted Alexander polynomials. In this paper, we give explicit upper bounds on the number of such characters in terms of a presentation of the knot group. In particular, we give upper bounds in terms of the crossing number of the knot.

math.GT↗

NeuroDyn-EEG: An Interpretable Pre-trained Model for EEG Based on Neural Dynamics

Clinical scalp electroencephalography (EEG) offers a noninvasive window into neural dynamics of neuropsychiatric disorders. However, discriminative deep models often lack anatomically indexed physiological interpretability. We propose NeuroDyn-EEG, a pretraining framework integrating generative priors from neural dynamics. It couples an extended Jansen-Rit neural mass model, leadfield-based source projection, and simulation-based parameter inversion. Trained on synthetic parameter-EEG pairs within physiological ranges, NeuroDyn-EEG estimates 11 regional parameter families across 90 AAL regions plus one global parameter from standard 19-channel EEG, using only ~2.43M trainable parameters. We evaluate the framework across three levels. First, controlled simulations demonstrate robust parameter recovery under diverse noise conditions, while real resting-state EEG evaluations confirm spectral and phase consistency in an inverse-forward closed loop. Second, on four clinical benchmarks (AD65, PD31, Figshare MDD, and TUAB), NeuroDyn-EEG achieves competitive classification performance, securing the highest BACC, AUROC, and AUCPR on PD31 and MDD, and highest BACC on AD65. Third, post hoc regional analyses reveal disease-specific alterations: local synaptic connectivity C_1 involves the most altered regions in AD65, whereas the firing threshold theta ranks first in MDD, offering testable mechanistic hypotheses. Overall, NeuroDyn-EEG maps scalp EEG to anatomically indexed dynamical parameters, bridging representation learning and mechanistic neurophysiology. Code: https://github.com/Gnosis-Neurodynamics/NeuroDyn-EEG.

cs.NE↗

LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks

Vision-language-action (VLA) models struggle to reuse recurring interactions in unseen tasks. Our diagnostic study reveals that reliable task completion does not imply consistent execution of constituent atomic actions across task contexts. We present LexiconVLA, a retrievable atomic-action lexicon for cross-task reuse. Global and detail codebooks capture shared interaction structure and fine-grained execution variation, respectively, preserving both reusable patterns and execution details. Visual-Atomic Action Alignment couples trajectory reconstruction from visual state changes with visual outcome prediction from action codes, grounding the lexicon in motion and its effects. We learn these codebooks with trajectory reconstruction and visual alignment on our AtomAction Dataset of 57,803 segments from 69 tasks. A planner and scene-aware adapter translate new goals into code-conditioned subtasks for a shared policy, without skill-specific experts or deployment-time parameter updates. Across five policy backbones on 26 RLBench tasks, LexiconVLA largely maintains performance on 18 seen tasks while improving success on 8 tasks held out from policy training. With BridgeVLA, unseen-task success rises from 16.67% to 34.17% (+17.50 percentage points), and overall success reaches 71.08%, the highest among methods with reported results. Real-robot experiments demonstrate stepwise execution and failure recovery.

cs.RO↗

DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning

Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8\% on average cross-domain accuracy than EMO-R3.

cs.CV↗

Causal-EVC: Breaking Emotional Spurious Causality via Spatiotemporal Grounding and Counterfactual Intervention

Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple attention matching, which inevitably suffers from {causal redundancy and spurious correlations} in co-occurrence bias (e.g., misclassifying ``sadness'' as ``joy'' on a sunny beach), leading to severe shortcut learning from confusing backgrounds. Furthermore, existing evaluations fail to verify whether models have genuinely mastered causal reasoning or merely exploited background confounders. To address these limitations, we first construct {EVC-CauseGround}, a comprehensive benchmark with dense spatio-temporal causal annotations. Crucially, it introduces a carefully selected {Causal-Faithfulness Subset} to explicitly quantify genuine emotion-cause attribution. Second, we propose {Causal-EVC}, an emotion-grounding captioning framework, which introduces a Motion-guided Causal Spatiotemporal Localization module to precisely decouple causal triggers from background confounders. Besides, we introduce an Interpretable Sparse Emotion Routing module. By synthesizing counterfactual representations and formulating a novel counterfactual contrastive objective, we enforce the model to anchor its emotion predictions strictly on authentic causal triggers instead of confusing background. Extensive experiments show that Causal-EVC not only achieves the best performance on semantic metrics but also exhibits significant advantages in the causal-faithfulness subset, which demonstrates that our model could mine emotional cues from genuine visual causes and mitigate co-occurrence bias for interpretable multimodal emotion understanding.

cs.CV↗

Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. Construction tests scene-level spatial reasoning from open-ended language specifications, where many realizations are valid; editing tests precise control of scene state, where the agent must recover the target scene from reference images while preserving everything else. Rather than scoring code or rendered views, Code4Scene evaluates the generated engine-native scene for task fulfillment, artifact integrity, and static physical validity, with edits additionally compared against withheld ground truth. Across 14 coding-agent configurations on the 95-case public set, construction and editing performance are strongly correlated but not interchangeable (Spearman $ρ= 0.78$): Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.

cs.AI↗

XBDD: A Highly Optimized ROBDD with Per-Edge Variable-Flip Maps

The Reduced Ordered Binary Decision Diagram (ROBDD) is a canonical representation of Boolean functions and is widely used in tasks such as equivalence checking and satisfiability checking of combinational circuits. Classical ROBDD packages greatly improve the efficiency of building ROBDDs through a series of optimization techniques, and compress the node scale of the ROBDD through complement edges. However, existing implementations do not take into account the local polarity differences of isomorphic Boolean functions, and still produce a distinct node for each polarity combination, thereby causing an explosion in the number of nodes. This paper proposes XBDD, a highly optimized ROBDD that, on the basis of fully implementing complement edges and their accompanying engineering techniques, introduces a per-edge variable-flip map. XBDD attaches a flip map to each edge to indicate which input variables must be negated when that edge is followed. This allows nodes that differ only in local input polarities to be merged, further reducing the node count. For certain function families, this sharing even yields exponential compression. We also propose methods that use a bitmap and a map pool to substantially reduce the extra overhead brought by the map, and propose normalization and cofactor operators for the map. In addition, XBDD implements several other engineering optimizations to further improve both time and space efficiency. Experiments show that XBDD trades a controllable time cost for a significant space gain, validating the effectiveness of the per-edge variable-flip map.

cs.DS↗