Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Bayesian-Guided Cooperative RL Beamforming for Wireless Adversarial User Detection

In next-generation wireless networks, communication systems are expected to go beyond simple data transmission and simultaneously provide high data rates, efficiency, and security. This requirement has motivated the extensive adoption of machine learning methods to develop intelligent and real-time network management frameworks, enabling the system to continuously monitor and react to channel variations and user behavior while maintaining efficient information delivery. In this context, the integration of machine learning with beamforming enables adaptive and data-driven beam direction selection, improving both the efficiency and security of wireless links. In this work, a 3GPP-based system model is first implemented under a no-attacker scenario, and an exhaustive search is employed as a reference to identify the best beamforming configurations. The proposed framework is then evaluated in the presence of an attacker and under different network scalability conditions. We demonstrate that the reinforcement learning-based approaches, namely Q-learning and SARSA (State-Action-Reward-State-Action), consistently outperform random selection in terms of total channel capacity, attacker detection accuracy, and performance stability. Among the evaluated reinforcement learning methods, Q-learning achieves the best overall trade-off between detection accuracy and computational efficiency. Our results indicate that the proposed framework provides a stable, scalable, and effective solution for joint beamforming and security-aware decision-making in dynamic and adversarial wireless environments.

cs.IT↗

On the Realizability of Edge-Girth Sequences

The edge-girth of an edge $e$ in a simple connected graph is the length of a shortest cycle containing $e$, with $g_e = \infty$ if no such cycle exists, and the edge-girth sequence of a graph is the nondecreasing sequence of the edge-girths of its edges. We characterize the sequences that are realizable as the edge-girth sequence of a simple connected graph. A constant sequence $(g^{(m)})$ with $g$ finite and $m \geq 1$ is realizable if and only if $m = g$ or $m \geq \lceil 3g/2 \rceil$. Otherwise, writing $S = S_0 \uplus (g^{(m)})$, where $g$ is the maximum value of $S$ and $m$ its multiplicity, we prove that $S$ is realizable if and only if $S_0$ is realizable and $m$ lies in an explicit set determined by $g$ and by the maximum diameters of the graphs realizing $S_0$, alone or extended by fewer than $g$ edges of edge-girth $g$. The proof rests on a chaining lemma, which turns a shortest cycle of a realization into a lower bound on the maximum diameter of the realizations of a shorter sequence.

math.CO↗

S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and $π_0$, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.

cs.RO↗

OneLatent: Latent Reasoning for Efficient Foundation Recommendation Models

Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their use as the backbone of foundation recommendation models (FRMs). Existing methods enhance recommendations through explicit Chain-of-Thought (CoT) reasoning under a Think-then-Answer paradigm. However, explicit CoT incurs substantial inference overhead by generating lengthy reasoning traces and relies on manually designed templates that struggle to capture diverse, dynamic user interests. We propose OneLatent, an efficient latent reasoning framework that compresses explicit reasoning traces into several learnable latent tokens, enabling Latent-Reason-then-Answer inference without generating verbose traces. OneLatent first introduces Multi-View Adaptive CoT (MV-ACoT), which creates diverse, high-quality teacher-generated supervision by exploring user interests from multiple perspectives and automatically adapting reasoning complexity to each instance. Building on pretrained FRMs, it then uses a three-stage latent-token alignment paradigm to progressively internalize CoT traces into learnable latent tokens. Finally, a multistage curriculum-based post-training strategy activates latent-token reasoning for downstream recommendation tasks. Experiments on an industrial-scale Kuaishou dataset and the public Kuaishou LLM-Rec benchmark show that OneLatent consistently outperforms explicit CoT-based methods and traditional baselines. Compared with the Think and No-Think variants of FRMs, OneLatent improves SID@64 by 17.44% and 9.33%, respectively, while achieving over 17x higher online inference throughput. We further develop a production serving system for scalable, real-time FRM inference. An online A/B test in Kuaishou's local-services advertising scenario shows that deploying OneLatent with this system yields an estimated 9.6% revenue lift over strong online baselines, including OneRec and OneReason.

cs.IR↗

The HRT Conjecture for Symmetric Configurations and Real-Valued Functions

The Heil-Ramanathan-Topiwala conjecture asserted that every finite collection of distinct time-frequency shifts of a nonzero square-integrable function is linearly independent. Recent counterexamples with Schwartz windows leave open the classification of configurations and windows for which independence holds. We prove linear independence for every nonzero window in $L^2(\mathbb{R})$ and every configuration consisting of finitely many commensurably spaced points on one line together with two distinct points on a parallel line. As a consequence, we obtain independence for every nonzero real-valued window and every four point configuration $\{(0,0),(0,1),(s,0),(a,b)\}$ with $abs\ne0$. The configurations underlying the known four-point counterexamples can be put in this form by translation and dilation, so they cannot yield dependence with nonzero real-valued windows. The proof combines factorization of Laurent polynomials with product estimates along orbits of irrational rotations. Finally, for a fixed nonzero window and a fixed three-point system, we show that the set of time-frequency parameters producing a shift in its span is compact, has Lebesgue measure zero, has Hausdorff dimension at most one, and is totally disconnected.

math.FA↗

SIGIL: Skill Compilation for Reliable and Efficient Agent Execution

Agent skills describe reusable procedures, but runtime models must still interpret their instructions and coordinate execution. We introduce Skill Compilation, which translates the procedure prescribed by an authored skill into an executable harness while preserving decisions left to the model. Our compiler, SIGIL, translates skills and their resources into a typed intermediate representation, validates it, and deterministically generates the harness. We evaluate SIGIL on 11 compliance-critical skills, where following the prescribed procedure is part of correctness. SIGIL improves adherence across all four runtime models, achieving up to 100% measured mean Skill Adherence compared with 26.2-54.0% for direct SKILL.md execution. On this suite, SIGIL reduces total runtime tokens by 21-45% for three of the four models. On SkillsBench, SIGIL also improves task performance across all four models, with absolute gains of 3.5-35.7 percentage points. These findings suggest that compiling reusable skill procedures into executable harnesses can improve adherence, task completion, and runtime efficiency.

cs.SE↗

Quantum vs Classical Erasure: Equal Bounds but Unequal Costs

Irreversibility has a fundamental thermodynamic cost, erasing information inevitably generates heat. This connection is quantified by the Landauer bound, which gives the minimum dissipation needed to erase a single bit of information. While this bound applies in both classical and quantum settings, it is saturated only in idealised limits of infinite resources. Here, we provide a unified first principles description of finite-resource erasure in both classical and quantum systems. We begin by proving the communal folklore that in the idealised regime the erasure cost of a bit encoded in a quantum or classical system is the same. Despite this, we show that their practical implementation differs substantially: achieving comparable erasure quality in quantum systems requires more control, larger accessible energy gaps and longer operation times. Classical protocols can achieve the erasure of a comparable quantum protocol under far weaker constraints which we expose in trade-off relations. Our results explain why practical erasure schemes fall short of Landauer's bound and show that classical systems enjoy several fundamental thermodynamic advantages.

quant-ph↗

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Self-Verification

Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multi-modal reasoning. However, recent studies have revealed that such models often use tools unfaithfully. Many process images are irrelevant to the question (e.g., the crops miss the queried target), yet the tool call still receives full credit and the model still answers correctly. Such decorative or misaligned tool calls waste computation and reveal that the model does not faithfully use the evidence it retrieves. This may stem from two limitations of prevailing methods: the tool reward fails to distinguish useful from useless calls, and tool feedback carries no signal of usefulness. To this end, we introduce FaithEyes, a multi-agent self-judging framework. Concretely, we use a VLM to judge whether each process image helps answer the question. The judgement is injected into the reasoning context as part of the tool observation to help subsequent reasoning, and meanwhile is used to scale the tool reward by the helpful-tool ratio to suppress reward hacking. To keep judgement available at evaluation, we further design a multi-agent framework where the model itself serves as a subagent to judge the tool calls from the main agent, eliminating any dependence on external models at inference. Training via a two-stage SFT + RL pipeline on adapted open-source data, FaithEyes attains competitive or superior accuracy across visual perception and reasoning benchmarks, while substantially improving tool faithfulness and reducing inference cost. The homepage is at https://github.com/Mosi-AI/FaithEyes.

cs.CV↗

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs 1,079 RCA tasks with six days of metrics, logs, and traces collected from an OpenTelemetry-instrumented microservice system under continuous simulated user load. Agents investigate this recorded history through real observability interfaces---Prometheus, Jaeger, and OpenSearch via Grafana---with full access to application source code. Tasks systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $κ_w = 0.91$). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard---a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access reduces RCA accuracy and increases the hallucination rate for every evaluated model. These results come from a curated 50 GB / six-day testbed of standalone tasks on a system whose code and instrumentation are public. Since real production systems are orders of magnitude larger, more dynamic, and more idiosyncratic, the gap we report underscores the engineering work still needed before agents can be entrusted with production reliability. We release the public set at https://hub.harborframework.com/datasets/orca-bench/orca-bench.

cs.CL↗

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks. We rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness and Tool Effect. Mode Adaptiveness characterizes whether an MLLM recognizes when tools are necessary and invokes them accordingly, avoiding unnecessary computational overhead while improving performance on problems requiring tool assistance. Tool Effect characterizes whether tools extend the model's capabilities on problems unsolvable through tool-free reasoning without introducing errors on problems it can already solve. Our analysis quantifies these properties and reveals that existing models exhibit limited Mode Adaptiveness, while tool-use gains on hard examples are largely offset by harm on easy ones. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model trained with supervised fine-tuning (SFT) and reinforcement learning (RL). Its RL stage combines Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion. Necessity-Aware Adaptive Reward encourages tool-free solutions when they succeed while preserving full reward for successful tool use when tool-free rollouts fail. Hint-Guided Capability Expansion uses verified, answer-free expert hints to recover learning signals from all-wrong rollout groups, aiming to extend tool-use capability on the hardest problems. Across 13 benchmarks, Beacon achieves the highest average score among the evaluated open-source models and ranks first on 11 benchmarks. On five diagnostic benchmarks, it improves the average tool-available accuracy over its tool-free accuracy by 1.96 points and achieves the largest tool-gain minus tool-harm score (+3.14 points). These results show Beacon's advanced performance, Mode Adaptiveness, and the net benefit of tool use.

cs.CV↗

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted camera, while training-time CBF guidance represents clearance to every body link, and an adversarial motion prior regularizes the resulting evasive reflexes. We evaluate on a controlled any-link contact benchmark with seeded throws in two regimes: single throws and a deployment loop in which the robot walks back to its station and recovers between throws. On this benchmark, the policy comes within a few points of a privileged state oracle: a fixed onboard camera alone is adequate for evasion. We find that usable barrier structure depends on perceptual observability: Joint-CBF gives the best performance with accurate ball states, degrades under fixed-camera observations when used only as training guidance, and recovers with a ball-tracking gimbal or privileged runtime filter. We therefore deploy a lightweight Link-CBF policy zero-shot on the Unitree G1 in the real world, where it tolerates imperfect perception, succeeds on 95% of throws, and uses semantic segmentation to dodge different balls.

cs.RO↗

Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study

Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. While performance gains from parameter-efficient fine-tuning are well characterised, the corresponding impact on trustworthiness (factual calibration and adversarial robustness) remains poorly understood. This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation across three SLM architectures (TinyLlama 1B, Gemma-2 2B, Llama 3.2 1B), three domains (healthcare, legal, finance), two training-data conditions (benign and adversarially perturbed), and four fine-tuning strategies (baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA, TA-LoRA). Trustworthiness is evaluated through TruthfulQA MC2 (factual calibration) and HarmBench ASR (adversarial robustness) across all 216 experimental configurations with three random seeds. Three principal findings emerge. First, baseline QLoRA domain adaptation produces minimal TruthfulQA MC2 change across all model-domain combinations (mean |Delta TQA| < 0.02). Second, adversarially perturbed training data consistently improves domain adaptation quality (Delta loss approximately -0.040) without worsening trustworthiness benchmarks. Third, none of the three safety-preserving strategies reduced adversarial harm susceptibility: Safety-DPO was effectively neutral (mean Delta ASR < 0.001), while Dark ER and TA-LoRA increased mean HarmBench ASR by +0.171 and +0.155 respectively in safety-aligned models (Gemma-2 2B, Llama 3.2 1B), with individual configurations exceeding +0.45. These results challenge the assumption that replay-based and arithmetic-merge strategies transfer alignment to domain-adapted SLMs.

cs.CL↗

C-Instrument: Automating RL Data Generation and Hillclimbing with a Constitution-Grid Instrument

Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts. Our finding is that over-refusal improves 22.4% to 12.8%, while under-refusal on adversarial attacks silently worsens 0.27 to 0.33. We present C-Instrument, a constitution-grid data instrument that generates the RL training data, and C-LIM, a per-cell learnability score that decides each cell's move: prune, densify, amend, expand. C-LIM flags the dead-weight data region before any training budget is spent: 187 untargeted rows had bought zero gain, and our method lifts the same region's learning impact 0.733 to 0.80. Code and the constitution are open-sourced.

cs.CL↗

Verifier-Induced Support Reshaping in On-Policy Optimization

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce. We call this verifier-induced support reshaping and define effective rewardable support as successful trajectories reachable within a fixed rollout budget. Across two model families, we study this effect through repeated verifier-scored sampling and bidirectional training on mathematical reasoning and constrained instruction following, including sequential training with the opposite verifier. Math-RLVR raises average instruction-following success but reduces the number of prompts with any successful response under repeated sampling. On IFEval with Qwen3-8B-Base, pass@1 rises by 6.5 percentage points while best@32 falls by 9.8 percentage points, and the same divergence appears across both models and IF benchmarks. Conversely, IF-RLVR shifts math responses from step-by-step openings toward direct answers, lowers best@k across sampling budgets, and reduces reward variation for later Math-RLVR. Token-distribution analyses and controlled opening interventions show that these changes concentrate in the first few response tokens. RLVR mainly reranks openings already available in the base policy, and the selected opening causally affects math searchability. The tested reference-policy constraints, routing priors, and on-policy distillation preserve cross-task support only partially; MathIF and ReasonIF show that marginal gains translate only partly into responses that are both correct and constraint-following. Therefore, endpoint improvements do not guarantee future trainability or joint capability under on-policy optimization. Code is available at https://github.com/sylvain-wei/VISR

cs.LG↗

Learning Latent Algebraic Structure from Ambiguous Set Observations

We study when statistically learnable latent structure can also be recovered efficiently, and how membership queries change the answer. An unknown support $A\subseteq\mathbb F_2^n$ has small additive doubling and is observed through a fixed set $B$ satisfying $|A\triangle B|\leη|A|$. We seek one linear subspace $V$ such that every compatible support $A$ is covered by few $V$-cosets and satisfies $|V|\le|A|$. For every $η<1$, polynomially many uniform samples suffice statistically, with cost polynomial in the doubling constant and proportional to $(1-η)^{-1}$; this radius dependence is sharp. Under a specified hardness assumption for learning parities with noise (search-LPN), however, no polynomial-time sample-only learner achieves even constant covering cost, including when the latent support is unique. At fixed structural parameters and the same constant covering budget, adding exact membership queries to $B$ permits polynomial-time recovery. The general query learner constructs a short structural list and uses fresh samples to select one common output through a majority-coverage rule. Persistent structured cores make this candidate construction possible. At doubling one, a complementary distinction appears at $η=1/3$: coarse recovery remains polynomial time, while exact recovery requires exponentially many accesses in the worst case when latent cardinality is unknown.

cs.DS↗

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments. Code is avaible: https://github.com/AI45Lab/OpenART#

cs.CL↗

Power-Logistic Shocks in a Lossy Nonlinear Transmission Line

We consider a lossy transmission line with a nonlinear voltage--charge relation. We derive an equation for a traveling front with the charge approaching constant asymptotic values on both sides of the front and solve the inverse problem for this equation exactly. Starting from a prescribed monotonic front profile and a prescribed front speed, we determine the local sound speed within the front. This quantity is the central object of our analysis and allows us to determine the voltage--charge relation of the transmission line in which the front propagates. The squared sound speed, averaged uniformly over the charge interval spanned by the front, is equal to the squared front speed. We specifically consider fronts with a power-logistic profile and its special case, the hyperbolic-tangent profile. All physically admissible fronts of this form are shocks rather than kinks. The voltage--charge relation of the transmission line in which the power-logistic shocks propagate is expressed in terms of the lower incomplete beta function. We propose an approximate method of solving the traveling front equation in a transmission line with the prescribed voltage--charge relation by the inverse-design approach. We examine the local spectral properties of the front in terms of the original partial differential equations and show that the short-wavelength perturbation modes are damped. The kink--shock distinction is considered in the framework of the long-wavelength characteristic analysis.

nlin.PS↗