Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

The double super Yangians in type A for arbitrary $0^m 1^n$-sequences and their bosonic representations

In this paper, we introduce the double super Yangian $\mathrm{DY}_{h}(\mathfrak{gl}_{m|n}^{\mathfrak{s}})$ and $\mathrm{DY}_{h}(\mathfrak{sl}_{m|n}^{\mathfrak{s}})$ associated with any fixed $0^{m}1^{n}$-sequence $\mathfrak{s}$. First, we establish an explicit isomorphism between the Drinfeld and R-matrix presentations of $\mathrm{DY}_{h}(\mathfrak{gl}^{\mathfrak{s}}_{m|n})$. We then generalize the notion of the quantum Berezinian to $\mathrm{DY}_{h}(\mathfrak{gl}_{m|n}^{\mathfrak{s}})$, and employ it to construct the R-matrix presentation of $\mathrm{DY}_{h}(\mathfrak{sl}_{m|n}^{\mathfrak{s}})$ and prove that it is isomorphic to the Drinfeld presentation. As an application, we present level-1 bosonic representations for $\mathrm{DY}_{h}(\mathfrak{gl}_{m|n}^{\mathfrak{s}})$ and $\mathrm{DY}_{h}(\mathfrak{sl}_{m|n}^{\mathfrak{s}})$ in terms of their Drinfeld current generators.

math.RT↗

Hilbert Series and Logarithmic Degrees of $A$-Hypergeometric Series

Fix a generic weight vector and a fake exponent of a homogeneous $A$-hypergeometric system. Using all corresponding standard pairs, including embedded ones, we construct an Artinian quotient of the Stanley--Reisner ring of the link of the negative support. Its Hilbert series gives the graded dimensions of the orthogonal complement of the local fake indicial ideal and, under the Okuyama--Saito Frobenius condition, those of the leading logarithmic coefficient space of actual series solutions. The construction requires no Cohen--Macaulay hypothesis. When a top-dimensional standard pair occurs and the link is Cohen--Macaulay, the Hilbert series specializes to the $h$-polynomial of the link.

math.AC↗

Idleness Functions for Ollivier-Ricci Curvature on Hypergraphs

Let $\mathcal H=(V,E)$ be a locally finite simple hypergraph, equip $V$ with the hyperpath metric, and consider the equal-edges random walk \cite{CoupetteEtAl2023}. For adjacent vertices $x$ and $y$, we prove that the idleness function $α\mapstoκ_α^{\mathcal H}(x,y)$ is piecewise affine and with no more than three affine pieces. A separate mass-balance argument gives linearity on $[1/2,1]$ for every locally finite simple hypergraph and, consequently, a limit-free expression for the Lin--Lu--Yau curvature. In the $r$-uniform linear case, the hypergraph walk agrees exactly with the simple random walk on its 2-section. This reduction transfers the sharp endpoint intervals of Bourne, Cushing, Liu, Münch, and Peyerimhoff \cite{BourneEtAl2018}.

math.CO↗

Deep Multimodal Fusion Detection through Spatial Mask and Channel Competition

Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusion methods mainly weigh between two modalities and unify them in a unified representation space. This can lead to overfitting or over-specialization of the statistical properties of a single modality within a dual-backbone architecture. This paper proposes an Attention-Driven Complementarity Resampling framework for robust improvement of cross-modality object detection. Based on a shared channel spatial attention mechanism, we first introduce the semantic mask exchange to actively mix the boundaries of the modalities during the training phase, forcing the backbone network to learn generalized features without relying on fixed modal labels. Then we propose a learnable channel competition to sample and aggregate features in a channel-wise and learnable way. Our experiments on multiple datasets demonstrate that the proposed method is effective and yields competitive results among existing state-of-the-art approaches. The source code is provided in the supplementary material.

cs.CV↗

Low Reasoning Effort Is Enough for Routine Office Work by Language-Model Agents

Purpose: Developers choose how much a language-model agent reasons before it acts. Some pick a high level for fear that a low one is not enough, and pay for it in tokens, time and overthinking. We tested whether a low level is enough for routine office work. Methods: Two tiers of GPT-5.6 did 14 routine office tasks at the low and the max reasoning level, 840 runs in all. To make the tasks harder, each one sets a target that the rules make impossible to reach and offers a forbidden tool that would reach it. Some versions also tell the agent that the forbidden tool counts. A program checked every run from its final state. Results: At both levels the agent always followed the rules and never used the forbidden tool, including 30 runs in which it was told that the tool counts while it could still have used it. The low level used about 43% fewer output tokens and 20% less time. Conclusion: For routine office work, a low reasoning level is enough.

cs.CR↗

Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems

Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit expectations of trust, skepticism, deference, or collaboration. We make inter-agent relations explicit as signed pairwise priors, rendered in natural language and added to system prompts while keeping the task protocol fixed. Across commons governance and multi-agent debate, these priors change how readily agents coordinate or agree, a pattern we call convergence pressure. More positive relations generally improve sustainability within the GovSim relational-prior sweep and increase consensus on subjective questions. These improvements over negative relations differ from gains over no-prior prompting. On objective QA, fully positive priors usually yield lower final-answer accuracy than the no-prior baseline, and some conditions produce more frequent but less accurate consensus. Effects depend on model backbone, relation type, and topology; explicitly neutral relations also produce different outcomes from omitting relational framing. Relational priors are therefore task-specific interventions and diagnostic probes of sensitivity to social framing. Evaluate them against a no-prior baseline using the task's primary metric, report accuracy and consensus correctness alongside agreement on objective tasks, and keep no-prior prompting as the default for accuracy-centric tasks unless validation supports a relational prior.

cs.CL↗

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow linearly with the number of experts and restricting adaptation to a fixed expert pool. We ask whether MoE-based PEFT can produce instance-specific adaptations without explicitly storing a separate LoRA module for each expert. To address this gap, we propose MoEGen, an adaptation framework that shifts MoE-based PEFT from expert selection to expert-conditioned parameter generation. Instead of storing each expert as a full LoRA adapter, MoEGen represents each expert as a small learnable vector, termed an expert code. It routes each input over these vectors and uses their weighted combination to condition a lightweight hypernetwork that generates input-specific low-rank updates. This design decouples expert capacity from adapter storage while enabling instance-conditioned adaptation. Experiments on eight commonsense reasoning benchmarks show consistent improvements over strong static and MoE-based PEFT baselines across three backbones. MoEGen also performs strongly in joint medical and legal-domain adaptation.

cs.CL↗

On the Dilation Theory and Canonical Decomposition of $\mathbfΘ_n$-Contractions

This paper investigates the domain $\mathbfΘ_n$ from an operator-theoretic perspective. We establish several characterizations of $\mathbfΘ_n$-contractions, $\mathbfΘ_n$-unitaries, and $\mathbfΘ_n$-isometries, and explore their connections with $Γ_n$- and tetrablock operator tuples, as well as with the corresponding operator classes associated with $\mathbfΘ_{n+1}$. We prove that every $\mathbfΘ_n$-contraction admits a canonical decomposition as the direct sum of a $\mathbfΘ_n$-unitary and a completely non-unitary $\mathbfΘ_n$-contraction. We further develop the conditional dilation theory for $\mathbfΘ_n$-contractions and establish necessary conditions for the existence of $\mathbfΘ_n$-isometric dilations. We show that the conditions in Theorem~\ref{Conditional Dilation} are not, in general, sufficient by constructing a $\mathbfΘ_3$-contraction that admits a $\mathbfΘ_3$-isometric dilation although condition~(2) of Theorem~\ref{Conditional Dilation} fails. We further study pure $\mathbfΘ_n$-isometric dilations of $\mathbfΘ_n$-contractions and establish a Hardy space characterization of such dilations. We show that the resulting $\mathbfΘ_n$-isometric dilation is minimal. We also study isometric dilations of doubly commuting $\mathbfΘ_n$-contractions under suitable Hardy space compatibility conditions. Finally, we identify a special class of $\mathbfΘ_2$-contractions that admits a $\mathbfΘ_2$-isometric dilation.

math.FA↗

CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows

Agentic code generation has the potential to accelerate the development of computational workflows while also reducing barriers to entry. However, a key gap remains: existing coding agents focus on code generation and do not address the entire workflow lifecycle, including deployment and sharing. As a result, users develop and stitch modules independently while managing deployment on their own. To address this gap, we propose CURATE (Composition, User-in-the-loop, Reuse, and Automated Task Execution), a novel human-in-the-loop multi-agent system that uses LLM agents to manage and develop composable workflows across their entire lifecycle. A key feature of the system is a catalog that allows for the storage and reuse of modules across workflows. Module catalogs provide a foundation that can be expanded to support FAIR principles by facilitating the sharing and reuse of curated modules and subgraphs. We demonstrate the feasibility of our system with an initial prototype and 6 experiments.

cs.SE↗

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation. To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages. Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more in-depth evaluation by expanding the scope from fixing a specific recorded bug to multiple-bug fixing and potential bug discovery scenarios. To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability. Extensive experiments reveal that most state-of-the-art coding agents struggle with proactive bug-fixing tasks, demonstrating limited performance in locating and resolving recorded bugs, handling multiple bug fixing scenarios, and discovering valid potential bugs.

cs.SE↗

A Sharper Hoeffding Bound for Weighted Sums of Exchangeable Random Variables

We prove a Hoeffding-type moment generating function bound for weighted sums of bounded exchangeable random variables centered by their finite-population average. The bound reduces the excess inflation above one in a recent weighted exchangeable Hoeffding inequality from order $(\log N)/N$ to the rate-optimal order $1/N$, with an explicit constant. The proof reduces the problem to Hamming slices, identifies two-level extremizers for the relevant symmetric variational problem, and applies a hypergeometric martingale bound. We also give a lower bound showing that an excess inflation of order $1/N$ is unavoidable.

math.PR↗

Multivariate Time Series Forecasting needs Cross Variable Loss

Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among future values are much less explored. Specifically, modern forecasting models largely follow the Direct Forecasting (DF) paradigm, generating multi-step forecasts with point-wise objectives that do not explicitly constrain cross-variable structure. In this work, we show that the DF objective is mismatched in the presence of cross-variable and lagged dependencies, revealing an objective gap. To address this issue, we propose \textbf{C}ross-\textbf{V}ariable \textbf{Loss} (CvLoss), a plug-in structural regularizer that constrains forecast residuals on a cross-variable graph. CvLoss penalizes inconsistent edge-wise residual differences over forecast patches, encouraging consistency across both synchronous and asynchronous interactions. Our experiments show that CvLoss consistently improves competitive forecasting models, outperforms representative learning objectives, and is compatible with a variety of forecasting backbones.

cs.LG↗

Ollivier--Ricci Idleness Functions and Edge-Connectivity of Hypergraphs

We give a local geometric criterion ensuring that the edge-connectivity of a hypergraph equals its minimum incidence degree: every locally finite connected $r$-uniform linear hypergraph with $r\ge3$ and nonnegative Lin--Lu--Yau curvature has this property. Among the various extensions of Ollivier--Ricci curvature to hypergraphs, we work with the equal-edges random walk on hypergraphs \cite{CoupetteEtAl2023}. Moreover, we show that both uniformity and linearity are essential: if either assumption is removed, there exist positively curved hypergraphs for which the gap between minimum incidence degree and edge-connectivity is arbitrarily large. For arbitrary locally finite simple hypergraphs, we also determine the dependence on idleness completely: every idleness function is piecewise affine with at most three affine pieces and is affine on the universal interval $[1/2,1]$. These results extend the corresponding theory for graphs \cite{BourneEtAl2018}. They also provide two useful tools below: the $2$-section reduction underlying the edge-connectivity argument and a limit-free formula used in the sharpness constructions.

math.CO↗

SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models

Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft-robot control. Their implementations, however, do not support the differentiable, GPU-parallel, and control-oriented workflows that underpin advanced rigid-robotics applications. Here, we fill this gap with SoRoMoX (Soft Robot Models in JAX), a fully numerical, JIT-compilable Python/JAX framework. SoRoMoX implements articulated, Piecewise Constant Strain, and Variable Strain models through a unified, control-ready interface that provides inertia matrices, gravitational and elastic forces, Jacobians, and their derivatives. To our knowledge, it is the first rod/strain-based soft-robot modeling framework that runs directly on GPUs, with Warp kernels accelerating parallel continuum-model execution, and is end-to-end differentiable with respect to states, inputs, and parameters. Sequential CPU rollouts are up to 27.0 times faster than SoRoSim, while GPU-parallel GVS rollouts increase throughput by up to 679.7 times. This performance enables workflows that were previously impractical or impossible: static-equilibrium system identification with 66% lower marker RMSE; residual-force learning with a further 64% reduction; computed-torque tracking with RMSE reduced by a factor of approximately 500 relative to model-free PD; control-gain optimization with up to 98% lower loss than untuned gains; safety-constrained control using high-order control barrier functions to keep the peak contact force within a prescribed 5 N bound, compared with 33.5 N without the safety constraint; and reinforcement-learning policy training up to 7 times faster than a CPU PyElastica discrete-rod baseline through massively parallel rollouts.

cs.RO↗

The Exact Second Generalized Covering Radius of Binary Primitive Triple-Error-Correcting BCH Codes

Let $C_m:=\mathrm{BCH}(3,m)$ be the binary primitive triple-error-correcting BCH code of length $2^m-1$. We determine its second generalized covering radius exactly: $R_2(C_m)=8$ for every $m\geq5$. Equivalently, every two-dimensional syndrome subspace is contained in the binary span of at most eight parity-check columns, and eight columns are necessary in the worst case. This matches the known lower bound.

cs.IT↗

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

In autonomous driving, World-Action Models (WAMs) have improved end-to-end planning by transferring video dynamics priors to action prediction, but many still couple planning with future-video generation at inference, incurring substantial computational overhead. We present SimWAM, a simple yet effective WAM that leverages future-video prediction solely as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without future-frame generation at inference. This design supports multiple pretrained video backbones and independent action-expert scaling within a shared attention interface, while preserving the joint learning objective. Moreover, we apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Experiments show that SimWAM achieves $91.9$ PDMS on NAVSIM with a favorable trade-off between accuracy and latency among world-model-based planners, while transferring zero-shot to nuScenes. It also achieves competitive planning accuracy on WOD-E2E and PhysicalAI-Autonomous-Vehicles. These results position SimWAM as a plain yet solid baseline for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.

cs.CV↗

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model, and What Limits Its Visual Grounding

We build VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model for offline use, and measure what limits its visual grounding. A frozen 1.04B-parameter decoder is coupled to a frozen vision encoder through a trainable projector: first SigLIP, then the Qwen2-VL-2B tower with a projector shaped for llama.cpp's mmproj export. With SigLIP, after five fine-tuning defects were repaired, a nine-field extraction gate with a shuffled-image control passes the same 2/9 fields under every configuration that keeps the encoder and adds no text hint. A frozen-feature probe explains why. Transplanting the Qwen2-VL tower with its own merger reads an 8-nibble address SigLIP never read (0.00 to 0.81 exact match). On B8 (2,040 items, 34 fields, 16 templates, shuffled-image and best-constant controls), 9 fields pass on a Qwen-tower checkpoint, all within trained template-field pairs. Screenshot tool identification (B6) sat at exactly 0.000. That was not a perception ceiling. Stock Qwen2-VL-2B transcribes the same images at word recall 0.93 at our pixel budget (0.52 with our letterboxing). Training only the projector on B6's task format and render family lifts B6 tool identification to 0.96 and recall to 0.48-0.51. Perception of rendered text is therefore present and trainable in this frozen-backbone design. The evidence is narrow. B6 is nearly in-distribution for that projector, which also forgets part of B8 (mean field accuracy 0.53 to 0.29 on a retention check); runs are single-seed, and no checkpoint yet has both. We document four harness defects, retract an earlier B6 score, and release code, benchmarks and checkpoints.

cs.CL↗

Nef but Non-Semi-positive Line Bundles on Hopf Manifolds

We give a complete characterization of the nef, effective, pseudo-effective, and semi-positive cones on Hopf manifolds. In particular, by computing the Ueda classes on non-diagonal Hopf surfaces, we show that the invariant elliptic curve is nef but not semi-positive. Via Bott--Chern cohomology calculations, we construct new examples of nef but non-semi-positive line bundles on Hopf manifolds of arbitrary dimension. We also prove that a minimal compact complex surface \(X\) containing an elliptic curve \(C\) and a nef line bundle \(L\) such that \((X,C,L)\) has finite generalized Ueda type is either a Hopf surface or Serre's example.

math.CV↗