Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,099 records · Page 61Linked to original sources

Farey Symbols for the Picard Group

We introduce Picard Farey symbols for finite-index subgroups of \(G=\PSL_2(\ZZ[i])\), using the Gaussian Farey tessellation of \(\HH^3\) by ideal octahedra. A symbol records finitely many Gaussian-rational octahedral occurrences, their local stabilizers, and ordered cooriented face pairings. We prove that a valid marked symbol reconstructs the finite Picard-cell action and hence the subgroup, and we work out a range of exact examples. For a torsion-free subgroup of index \(n\), every fundamental domain formed from complete Farey octahedra contains \(n/12\) octahedra; a spanning-tree construction gives one with at most \(n/4+1\) side-pairing transformations. We also identify the Gaussian octahedral edge--face complex with the integral rank-two Steinberg presentation for \(\QQ(i)\). Thus the same decorated geometry directly determines a finite presentation of coefficient-valued Bianchi modular symbols; for \(Γ_1(2+i)\) we give the complete \(3\times4\) group-ring boundary matrix. Canonical reduction and an effective Hecke-compatible reduction remain open.

math.NT↗

Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses

Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imitates a teacher model. We define Agent Distillation as the persistent transfer of task-solving knowledge from a teacher agent to a student agent. Our study organizes the field by where transferred knowledge is retained: within the model, as artifacts, through the execution harness, or across substrates. This perspective separates transfer evidence from its outcome and clarifies how knowledge moves between agent components. We develop an evaluation framework that relates retention to causal contribution and deployed utility. Together, these contributions establish a foundation for the reliable, maintainable, and safe development of increasingly complex agentic systems.

cs.AI↗

A Spread-Gated Hawkes-Flocking Model for Best Bid and Ask Dynamics, with an Application to Limit Order Placement

We study the joint dynamics of the best bid and ask prices with a spread-gated Hawkes-flocking model. The model tracks four types of best-quote movements: spread-narrowing movements are switched off when the spread is at its one-tick minimum, and a cross-side excitation term, whose activation depends on the prevailing spread, links the two sides of the book. We show that the process is non-explosive on every finite horizon, give an $O(N)$ recursive likelihood, and validate the maximum likelihood estimator by simulation. On real intraday limit order book data for two large-tick stocks, INTC and MSFT, the restriction that removes the cross-side term is rejected, and the full model improves fit substantially by AIC and BIC; the likelihood is multimodal on a single day, so estimation uses a multi-start search. As an application, we derive the closed-form optimal size of a single-period limit order placed at the best or second-best quote, given the model's next-event probabilities and externally supplied execution probabilities.

q-fin.CP↗

Combined constraints on the diffuse flux of cosmic neutrinos between 10^16 eV and 10^26 eV

Experimental constraints on the diffuse neutrino flux above the PeV range have been obtained with Cherenkov neutrino telescopes, air-shower arrays, radio detectors in and above polar ice, and radio observations of the Moon. Published results use different flavor, energy-bin, and statistical conventions, so their flux limits cannot be combined directly. We reconstruct energy-dependent exposures of the published searches and place them in a common convention for the total all-flavor $ν+\barν$ flux over $10^{16}$-$10^{26}$ eV. Published event counts and expected backgrounds are combined with a Poisson likelihood, and 90% C.L. quasi-differential limits are obtained for one-decade $E_ν^{-1}$ test spectra using a one-sided profile-likelihood construction. Folding theoretical spectra of cosmogenic neutrinos and selected new-physics scenarios with the combined energy-dependent exposure yields constraints on their flux normalizations. This homogeneous analysis provides a reproducible observational benchmark across ten decades in neutrino energy and a common reference for current and projected searches.

astro-ph.HE↗

Physics-Aware Machine Unlearning for Cyber-Physical Systems

This paper proposes a physics-guided gradient-ascent-based machine unlearning method that couples the forgetting signal with the physical residual of the target cyber-physical systems, ensuring that weight updates during unlearning are steered toward physically feasible regions of the weight space. The physics residual acts as a safety fence during gradient ascent: the model is steered away from the poisoned behavioral basin and simultaneously toward physics-compliant territory, rather than toward an arbitrary alternative that may still violate domain constraints. We evaluate the proposed method against four baselines: naive gradient ascent, exact unlearning, SISA, and full retraining on an IEEE 34-bus distribution system, driven by two physics-informed neural network-based distribution energy resource controllers and validated through high-fidelity OpenDSS power-flow co-simulation. From the evaluation, we found that our proposed physics-guided model simultaneously removes poison and restores physical compliance, which are essential for the safe deployment of safety-critical cyber-physical systems

cs.LG↗

Lossless Compression of Lookup Tables for Hardware Applications

Large lookup tables are widely used in hardware to store constant-valued arrays for applications ranging from elementary mathematical operations, such as constant-coefficient multiplication and nonlinear function evaluation, to emerging machine learning models, including table-based neural networks (NNs) and Kolmogorov-Arnold networks (KANs). However, storing extensive tables of constant values can lead to excessive hardware costs in resource-constrained edge devices such as FPGAs. In this paper, we propose CompressedLUT, a lossless compression scheme and its decoder hardware architecture for the efficient storage and retrieval of arbitrary data in hardware. Our method combines decomposition, self-similarities, higher-bit compression, and multilevel compression techniques to maximize table size savings without accuracy loss. Its hardware decoder primarily uses addition, arithmetic right shift, and several small lookup tables, ensuring low area and high throughput. We evaluated CompressedLUT on FPGAs by implementing multiple nonlinear functions, constant-coefficient multipliers (CCMs), and KANs at 12-bit resolution. CompressedLUT is available as an open-source tool.

cs.AR↗

WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.

cs.SE↗

What Makes Recurrence Effective in Looped Language Models?

Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. (2) Performance also depends on how distinct layers and recurrent iterations are allocated, showing that effective depth alone is insufficient to predict behavior. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget. (3) Finally, we find that conventional initial-state injection offers limited robustness to varying recurrence depth. We therefore propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets. Overall, our results clarify when recurrent computation helps, where it fails, and offer practical guidelines for designing LoopLMs across variable inference budgets.

cs.LG↗

Decompose Dynamics Before Learning Dependencies in Spatiotemporal Systems

Relations in networked spatiotemporal systems are often learned from observations that entangle dynamics governed by different mechanisms, obscuring what evolves locally and how it propagates across nodes. We introduce Component-Aware Network Dynamics with Ordered Relations (CANDOR), which decomposes local dynamics before learning their dependencies. CANDOR represents each trajectory through a persistent background, gradual accumulation and release, and sparse shocks. Conditioned on these components, a delay-aware physical branch models edge and sample-dependent propagation over directed topology, while a topology-unconstrained functional branch discovers latent dependencies from background dynamics. Context-adaptive fusion combines functional, forward-propagation, and reverse-support forecasts, with training objectives encouraging specialized and semantically consistent representations. Experiments on two traffic benchmarks and three long-horizon water-quality datasets span two distinct spatiotemporal systems: human-driven urban traffic and naturally evolving river water quality. CANDOR consistently outperforms the strongest baselines, reducing MAE by up to 4.31% for traffic and MSE by up to 5.47% for water-quality forecasting. These results establish decomposition before dependency learning as an effective principle for spatiotemporal representation learning.

cs.CE↗

PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation

Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.

cs.LG↗

A Digital Simulation Toolkit for Physics-Based Generation of Realistic Experimental Scanning Tunneling Microscopy Images

Scanning Tunneling Microscopy (STM) is a widely used tool for characterizing surfaces of materials at the atomic scale, playing a crucial role in discoveries across condensed matter physics and materials science. Despite its extreme spatial resolution, STM is one of the most sensitive microscopy techniques and is highly prone to noise. While existing unsupervised denoising methods are very cheap to train, these are primarily focused on removing the noise with minimal recovery of key physical information. While supervised methods can offer superior performance, the major bottleneck is that a large amount of paired clean-noisy experimental images is required which are impractical to obtain. Thus, we developed a low-cost physics-driven digital toolkit to rapidly generate large volume of realistic STM images. Firstly, we simulate clean images from a chosen material system. Then, with prior knowledge of the physical characteristics of the artifacts and noise present in STM experiments, we formulate several artifact-noise functions such as Gaussian electronic noise, 1/f flicker noise, scan-line noise, background tilt and mechanical drift. These physically informed noise components are then added to the simulated clean images to generate realistic STM images. We demonstrated the capability of the proposed digital toolkit to generate AI-ready data for denoising images of the (111) surfaces of copper and lead, while preserving atoms, defects, and electron waves. We also validated the quality of the downstream image analysis of learning electron wave patterns induced by quantum interference from Cu(111) images. Results show that the supervised models trained on digitally generated AI-ready data can more effectively denoise and learn electron wave patterns on Cu(111) images than benchmarked unsupervised approaches, indicating that the proposed toolkit facilitate scientific discovery.

cs.LG↗

Foundation-Model-Guided Topology-Aware Semantic Risk Fields for Manipulation

Robot motion planning in everyday environments must satisfy hard geometric constraints while accounting for context-dependent semantic risk. We present a foundation-model-guided, topology-aware semantic risk field that extends manipulation safety beyond collision avoidance. For each manipulated-object/scene-object pair, a foundation model provides six directional risk weights and a pair-specific spatial decay scale. The method combines these priors with voxelized 3D scene geometry using topology-aware shielding and geodesic spatial decay. A GPU-parallel backend batches object-level distance and risk computations to construct a dense 3D field that serves as a modular cost for downstream motion planning. We evaluate the field's shielding behavior under full and partial barriers and compare its 3D workspace representation with a pixel-wise semantic-prior baseline. Across three household simulation scenarios, trajectories optimized with the proposed field have lower semantic exposure than collision-only trajectories under the same geometric constraints. We also evaluate the computational practicality and reliability of the supporting pipeline. Together, these results support the proposed field as a practical topology-aware semantic cost representation for manipulation planning beyond collision avoidance.

cs.RO↗

Inducing Process Supervision from Outcome-Only Reinforcement Learning

Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relative advantage is used to optimize the entire generated response. Intuitively, when checking intermediate steps helps determine the outcome, more accurate checks can lead to better outcome judgments and higher rewards. Outcome-only RL can therefore reinforce step-level verification without explicit process supervision. We validate the effectiveness of TIPS across math and agent benchmarks and four backbone families. Notably, TIPS-Qwen3-4B-Thinking-2507 reaches 85.2 F1 on ProcessBench with only 3.2K outcome-labeled trajectories, surpassing all evaluated trained PRMs and strong prompt-only judges such as GPT-5.4-Instruct and Claude-4.7-Opus, while still trailing o1-mini. Code and data are available at https://github.com/RUCBM/TIPS.

cs.LG↗

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advantage: a skill in context lifts WebShop success from 42.2% to 56.2%, yet changes the probabilities of fewer than a quarter of the sampled tokens. Much to Align: a skill changes the hidden states of over 80% of response tokens, in a way that linear probes can trace back to the specific skill. To exploit this, we propose Privileged Representation On-policy Self-Distillation (PR-OPD). After a GRPO warm start, the policy writes a hindsight skill for each trajectory, re-reads its own responses with that skill as a stop-gradient teacher, and aligns its projected hidden states to the teacher's at every layer alongside the reward objective, with no external skill library, separate teacher, or inference overhead. On ALFWorld and WebShop with two backbones, PR-OPD achieves the best overall results in every setting, improving over GRPO by up to 4.7 points in ALFWorld success and 14.0 points in WebShop accuracy. Code is available at https://github.com/balibata/PR-OPD.

cs.LG↗

Realistic Characterization and Modeling of Vehicle-Mounted Antenna Radiation Patterns Based on Large-Scale Full-Vehicle Measurements

Vehicle-mounted antenna radiation patterns can be significantly altered after installation due to interactions with the vehicle body. However, existing studies mainly focus on simulations or isolated antenna characterization, while large-scale experimental investigations remain limited. This paper presents a comprehensive experimental analysis of vehicle-mounted antenna radiation patterns based on full-vehicle spherical near-field measurements. A large-scale dataset covering ten commercial vehicle platforms, multiple antenna mounting locations, and various wireless functions, including 4G, 5G, V2X, and GNSS, is established under a unified measurement framework. Based on the measured results, representative case studies and statistical characterization are performed to investigate the impacts of vehicle installation on antenna radiation behavior. The results reveal substantial variations across mounting locations and wireless functions, highlighting the strong influence of vehicle integration on realistic antenna performance. Furthermore, a novel statistical model is proposed to generate realistic vehicle-mounted antenna patterns using a small set of physically interpretable parameters. The obtained dataset, characterization results, and modeling approach provide a foundation for realistic antenna evaluation and wireless system simulations in vehicular communication applications.

eess.SP↗

OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration

Cross-attention is a crucial component in learning-based image-to-point-cloud (I2P) registration. Although existing cross-attention mechanisms have achieved promising progress, attention ambiguity remains a fundamental challenge that hinders the learning of discriminative 2D-3D correspondences. To address this problem, we revisit cross-attention and establish ordinary differential equations (ODEs) to model the ideal I2P feature interaction. Based on this formulation, we develop an ODE-driven cross-attention (OCA) module that refines feature representations and attention matrices through ODEs. In practice, OCA can be seamlessly integrated into existing I2P registration frameworks. To validate its effectiveness, we incorporate OCA into five state-of-the-art baselines and evaluate on four public benchmark datasets. Experimental results demonstrate that OCA improves registration recall by up to 5\%, 9\%, and 15\% under the standard, fine-tuning, and zero-shot settings, respectively.

cs.CV↗

Where Predictive Supervision Goes Shapes What VLA Policies Learn

Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.

cs.RO↗

TRACE: A Trilinear Channel Estimation Framework for Tri-Hybrid Architectures with Reconfigurable Antennas

Tri-hybrid beamforming with reconfigurable antennas (RAs) adds an electromagnetic (EM)-domain degree of freedom to radio-frequency (RF) analog combining and baseband processing, but it also couples the EM angular response, the spatial steering matrix, and the path gains inside every received pilot. This letter develops TRACE, a trilinear channel estimation framework for the single-user uplink of a tri-hybrid receiver. Cycling the radiation pattern over the pilot dimension yields a third-order PARAFAC tensor of the received pilots, with factors corresponding to the EM-domain angular factor, the RF-projected steering factor, and the pilot-weighted fading factor. TRACE fits this tensor by alternating least squares and then recovers the physical channel parameters in closed form. Kruskal-based identifiability translates into design rules for the number of pilot slots, RF chains, and EM probing states. In a compressive configuration, TRACE attains the SNR slope of an oracle-support bound, staying within a factor of $1.4$ of it above $10$~dB, while outperforming a dictionary-based greedy baseline.

eess.SP↗