SearcharxivSearch

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 181 records · Page 10Linked to original sources

Explanation-Guided Medical Named Entity Recognition with Stability and Boundary Awareness for Atopic Dermatitis

Objective: This study aims to improve the reliability and robustness of medical named entity recognition (NER) in Chinese atopic dermatitis (AD) clinical texts through explanation-guided learning. Methods: We propose a stability and boundary-aware explanation-guided NER framework. Perturbation-based analysis is used to evaluate explanation stability and entity boundary sensitivity. An adaptive fusion strategy dynamically combines local and global explanation to generate more reliable token-level explanations. The fused explanation signals are further incorporated into model training through stability, boundary-aware, and consistency constraints. Results: Experiments on Chinese AD NER datasets show that the proposed framework improves explanation robustness and achieves consistent performance gains across multiple NER models. The adaptive fusion strategy also provides more stable explanations and stronger boundary perception than individual explanation methods. Conclusion: The proposed method effectively integrates reliable explanation signals into medical NER training, improving both recognition performance and explanation reliability. The framework provides a practical and generalizable solution for explainable medical NER and offers reliable support for downstream clinical decision-making and medical knowledge applications.

cs.CL

IViT: A Novel Interpretable Visual Transformer for Skin Disease Detection

The clinical diagnosis of skin diseases is susceptible to interference from inter-class similarity of skin lesions, and over-reliance on clinicians'experience easily leads to subjective bias. Although existing deep learning aided diagnosis methods achieve competitive accuracy, they suffer from the black-box opacity of Vision Transformer (ViT) and poor adaptability to medical few-shot scenarios. Moreover, mainstream explainable algorithms generally face the bottleneck of significant accuracy degradation when improving interpretability. This paper proposes an interpretable ViT (IViT) constrained by Quadratic Programming (QP). The introduced pre-trained transfer learning adapts to few-shot feature extraction. A discrete QP feature selection framework is constructed to screen generic and discriminative features consistent with clinical diagnostic logic. A multi-objective loss function is designed to reduce feature redundancy and optimize activation distribution while preserving classification performance. Experimental results on six standard skin disease datasets show that IViT achieves an accuracy of 93.80%, only 0.21% lower than the baseline, with feature redundancy reduced by 29.5%. Its core activation regions are consistent with clinically concerned lesion areas. The proposed model balances accuracy and interpretability, providing a reliable solution for the clinical deployment of few-shot intelligent skin disease diagnosis.

eess.IV

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models (LLMs). In parallel, diffusion-oriented RL algorithms, e.g., DanceGRPO and FlowGRPO, have rapidly expanded the scope of RL from language reasoning to diffusion-based visual and flow-based generation. However, efficient RL systems for diffusion generative LLMs remain underexplored. Existing implementations, e.g., veRL-Omni, still rely on colocated execution, which simplifies synchronization but couples rollout and training resources, limits heterogeneous deployment, and constrains independent scaling. To this end, we introduce DigenRL, a disaggregated RL framework for diffusion-based generative LLMs that supports flexible resource allocation, accommodates heterogeneous GPUs, and facilitates efficient task scheduling. To maximally reduce the execution bubbles in the disaggregated architecture, we propose: 1) a generation-axis pipeline (GAP) and time-step parallelism (TSP) in the diffusion architecture to enable finer-grained pipelining between rollout and training; 2) an elastic trainer-assisted generation (TAG) approach to enable the trainer GPU resources to dynamically assist in executing rollout generations; and 3) a tightly one-step constrained asynchronous strategy to further utilize the tail bubble in the pipeline. Extensive experiments are conducted on three hardware testbeds with 16-32 GPUs using HunyuanVideo-13B, Wan2.1-14B, FLUX.1-12B, and QwenImage-20B generative models. Experimental results show that DigenRL achieves 1.56-2.10x throughput improvements over state-of-the-art diffusion RL systems, veRL-Omni and GenRL.

cs.AI

SOLAR: AI-Powered Speed-of-Light Performance Analysis

How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are central to software, hardware, and algorithm optimizations. Speed-of-Light (SOL) analysis answers them by computing a workload's theoretical minimum execution time on a given architecture. Yet deriving SOL bounds remains manual, error-prone, and disconnected from rapid model development. To close this gap, we introduce SOLAR, a framework that automatically derives validated SOL bounds from PyTorch and JAX source code. SOLAR leverages both generative and deterministic components in its flow: an LLM frontend translates any source programs into an executable Affine Loop IR, validated by output comparison; a deterministic flow lifts the IR into an einsum graph; and an analytical backend computes unfused, fused, and cache-aware SOL bounds. SOLAR provides comprehensive operator and language coverage, produces validated bounds with zero observed SOL violations, and offers multi-fidelity analysis that tightens bounds and surfaces optimization insights. We evaluate SOLAR across KernelBench, JAX/Flax models, and robotics workloads. These experiments demonstrate four use cases: headroom analysis at multiple fidelity levels, identifying optimization opportunities, cross-platform exploration, and inverse-roofline hardware provisioning.

cs.LG

From Approximate Floquet Engineering to Full Floquet Theory: Coherent Control of Chiral Spin Systems in Spintronics

Coherent control of interacting spin systems under time-periodic driving is a central challenge in spin-based quantum technologies. Here we demonstrate the applicability of a full Floquet-space formalism, adapted from Nuclear Magnetic Resonance (NMR) methodologies, to model the dynamics of driven coupled electron spins in the presence of a static magnetic field B0 and a transverse oscillating field B1. The framework explicitly includes isotropic exchange coupling J and the chiral Dzyaloshinskii-Moriya antisymmetric exchange interaction (DMI), and its numerical convergence is systematically validated with respect to Fourier-space truncation. In the non-interacting limit, the expected driven-spin dynamics is recovered, with the oscillation periodicity governed by B1. Exchange coupling alone does not modify the collective spin expectation values under the chosen initial condition, consistent with symmetry considerations. In contrast, increasing DMI generates a finite Sy component, suppresses Sz, and produces tilted, elliptical Bloch-sphere trajectories, reflecting the emergence of chiral spin-spin correlations. These effects are pronounced for open boundary conditions, while remaining nearly negligible in the periodic boundary case. When exchange coupling and DMI coexist, the dynamics becomes strongly perturbed and multi-frequency in nature. Together, these results demonstrate that full Floquet-space modeling provides a robust and predictive framework for analyzing and engineering coherent dynamics in driven interacting spin systems beyond simple coherent-rotation regimes.

quant-ph

Quantum Variational Approaches to the Maximum Independent Set Problem at Utility Scale

Near-optimal solutions to Maximum Independent Set on dense graphs sit in local optima that greedy correction and maximality extension cannot escape. We encode near-optimal seeds as a uniform quantum superposition on ancilla qubits and evolve them under an excitation-preserving variational ansatz that holds the search inside the feasible Hamming-weight subspace. A preprocessing stage of spectral reordering and distance-based sparsification, together with history-guided post-processing of the sampled bitstrings, takes the method to 200 nodes. The ansatz entangles the seed branches and the bond dimension of the simulated state grows with circuit depth, which places deeper circuits outside the reach of exact matrix product state simulation at the bond dimensions available to us. Larger instances therefore need quantum hardware. Measured on the data register alone, the superposition behaves as a classical mixture over the seeds, so coherence between the branches has to be created and then looked for. We do this with a CRZ phase layer and post-selection on the ancilla qubits, which brings the branches into interference and exposes the cross terms. The idea is to see if interference between near-optimal seeds widens the range of independent sets the circuit returns. Standard VQE with this pipeline recovers the certified optimum for instances up to 125 nodes, and these run on ibm_marrakesh with parameters transferred from noiseless simulation. The ancilla construction is introduced for the sizes past that point. On a 180-node hard instance the superposition recovers the certified MIS. Five 200-node instances are solved to the certified optimum, and on the 400-node brock400-1 benchmark the method reaches size 25 against a certified optimum of 27. These are the largest hard general-graph instances we know of where a gate-based variational algorithm optimises the full circuit directly.

quant-ph

Prediction of biological radiation effects based on ionization clusters (nanodosimetry)

This article reviews approaches that link the formation of ionization clusters in nanometric volumes to radiobiological effectiveness. The corresponding models were developed as the field of nanodosimetry developed. Some address early biological radiation effects, such as DNA damage, while most aim to predict cell survival or inactivation. The models also differ in the nanodosimetric quantities considered, with many based on the probability distribution of ionization cluster formation in a single target. Some models account for the synergistic effects of pairs of ionization clusters formed in different targets. Several models feature macroscopic aggregation frameworks based on particle fluence, which are proposed for use in radiotherapy treatment planning, particularly in ion-beam radiotherapy. The models are presented here using harmonized terminology and notation for nanodosimetric quantities. An extension of the conceptual framework of nanodosimetry is also discussed. This extension transitions from a target-centered description to a track-centered description. It also introduces nanodosimetry-based analogs of dosimetric concepts, such as dose and linear energy transfer. This paper traces and summarizes the historical development of nanodosimetry-based biological effect models and discusses conceptual aspects of the models to reveal their underlying assumptions and the extent to which they are mechanistic or merely elucidate correlations. Eventually, an attempt is made to identify the key open questions in this field that still need to be addressed.

physics.med-ph

Risk-Sensitive Learning in Population Games under Extreme Events: Bifurcations and Chaotic Dynamics

Inspired by nonequilibrium phenomena in game dynamics and behavioral evidence on the impact of extreme events on decision making, we investigate the nonlinear dynamics of a discrete-time multiagent learning rule in population congestion games under extreme events affecting one of the actions. The population state, following a risk-sensitive variant of the Multiplicative Weights Update (MWU), is coupled with a belief variable capturing the agents perceived risk and updated through an adaptive expectation rule. We perform a two-parameter bifurcation analysis with respect to the agents controlled parameters, identifying regions of qualitatively distinct behavior. Equilibria are studied first from both game-theoretic and dynamical perspectives. The resulting two-dimensional system exhibits complex behavior, including multi-stability among fixed points, invariant curves, periodic and chaotic attractors. Despite this complexity, the attractors can be grouped into distinct families, while the Cesàro averages of the trajectories are shown to converge to the stationary equilibrium. The incorporation of risk associated with the extreme event leads to new dynamical phenomena: attracting invariant curves arise and give rise to phase-locking Arnold tongues, within which the dynamics is qualitatively similar. In this setting, codimension-two resonances are identified as organizing centers, both within individual tongues and along the bifurcation curves associated with the fixed-point family. Chaotic attractors emerge and are destroyed through Feigenbaum cascades and forward or reverse boundary crises, with interior and merging crises also observed, along with transient chaos and narrow periodic windows. For each qualitatively distinct region, representative phase portraits and the associated basins of attraction are examined.

nlin.CD

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

Strong speech-to-text (S2T) LLMs already provide robust speech perception and text reasoning, but adding speech-to-speech (S2S) output is challenging: fine-tuning the backbone can degrade the original S2T performance, while attaching a downstream talker reintroduces a serial text-to-speech bottleneck. We present PRIME-Speech, a frozen-backbone S2S conversion framework that trains only speech-generation modules. PRIME-Speech synchronizes a causal audio post-decoder with intermediate hidden states of the frozen backbone, so codec tokens are generated from the model's evolving reasoning trajectory rather than from completed text chunks. The post-decoder uses mixed hidden-state, text, and audio-history conditioning, and a training-time packing strategy with turn-level audio KV-cache and position reset stabilizes multi-turn spoken interaction without additional multi-turn S2S training data. Multi-token prediction further reduces the effective codec prediction rate and improves first-audio latency without modifying the reasoning path. Across speech translation, spoken QA, speech understanding, and multi-turn dialogue, PRIME-Speech preserves the S2T behavior of the frozen backbone while producing accurate, low-WER spoken responses.

eess.AS

Position--momentum uncertainty relation for the average confidence width

We introduce the average confidence width $Δ_{a} x=\int_0^1Δ_{c} x (θ)\,\mathrm{d}θ$: the confidence width $Δ_{c} x(θ)$---the measure of the smallest region of position space carrying probability $θ$---averaged over all confidence levels. It is the first moment of the decreasing rearrangement of the position density, an $L^1$ (mean-absolute-deviation) measure of localization, and $Δ_{a} x\,Δ_{a} p$ is dilation invariant. We study the uncertainty relation $Δ_{a} x\,Δ_{a} p\ge c^\ast\hbar$ for pure and mixed states. A mean--entropy argument with the Bialynicki-Birula--Mycielski relation gives $c^\ast\geπ/e$, while the ground state $ψ_0$ of the Fourier-invariant operator $|x|+|p|$, with $E_0=1.1040744$, gives $c^\ast\le E_0^2\approx1.2190<4/π$: Gaussian states, which saturate the Heisenberg--Kennard and entropic relations, are not optimal here. We prove that $ψ_0$ is nondegenerate, nodeless, even, self-dual, and symmetric decreasing, with algebraic tails, and that the sharp hybrid relations $Δ_{a} x \cdot 2 \langle |p-p_0| \rangle \ge E_0^2\hbar$ and $2\langle|x-x_0| \rangle \cdotΔ_{a} p\ge E_0^2\hbar$ hold for all states, so $Δ_{a} x\,Δ_{a} p\ge E_0^2\hbar$ whenever either density is symmetric unimodal; the second variation of $Δ_{a} x\,Δ_{a} p$ at $ψ_0$ is positive definite apart from the symmetry directions. Extensive numerical minimization, including coherent states, never finds a product below $E_0^2\hbar$ and consistently returns $ψ_0$, supporting the conjecture $c^\ast=E_0^2$. In two dimensions the relation is sharp and saturated by Gaussians, identifying the dimension as the origin of the non-Gaussian optimizer in one dimension. Finally, we discuss other complementary pairs and the relation of the width to interferometric predictability and visibility.

quant-ph

Rethinking Post-Hoc Calibration in Semantic Segmentation

Reliable confidence estimates are essential in semantic segmentation, yet modern models often remain miscalibrated. We investigate two overlooked issues in post-hoc calibration. First, adding a constant to all logits leaves softmax probabilities unchanged, but several standard calibrators depend on this arbitrary offset. In segmentation, this offset can vary across pixels or voxels, introducing spatially varying representation dependence. We characterize translation-invariant (TI) calibrators and construct TI counterparts of shift-sensitive methods. Second, calibrating with cross-entropy can degrade segmentation quality due to mismatched training and calibration objectives and limited calibration data. We investigate decision-preserving calibration under argmax- and order-preservation constraints. Since these constraints restrict affine softmax calibrators to temperature scaling, we introduce more expressive class-conditional affine calibrators that preserve decisions. Across natural-image and medical segmentation benchmarks, including corruption-based covariate shift, TI variants generally improve calibration, while decision-preserving variants prevent segmentation degradation by construction and retain strong calibration performance. Our findings provide practical design principles for post-hoc calibration in semantic segmentation.

cs.CV

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits

Mechanistic interpretability seeks to explain transformer behavior through circuits: sets of internal components that causally support a behavior. However, self-repair creates a blind spot: ablating a primary component can activate a dormant backup, so a circuit that explains behavior in the intact model can become incomplete under the intervention used to test it. We formulate this gap as conditional circuit completion: given a primary set, identify components that become causally important after its removal. We introduce conditional co-ablation (CoAx), which ranks candidates by growth in ablation effect after primary-set removal. We show that a perfectly dormant backup can be indistinguishable from an irrelevant component to per-unit intact-state scores, whereas its conditional effect change exactly aggregates all interaction orders linking it to the removed set. On GPT-2-small's Indirect Object Identification (IOI) circuit, CoAx recovers the documented backup heads at 0.941 ROC-AUC, versus 0.815 for the strongest intact-state attribution baseline and 0.758 for the matched conditional-energy control. Recovery drops to 0.40 +/- 0.13 AUC for alternative component sets matched in behavioral effect, output displacement, and depth, showing that recovery is specific to the removed circuit. Beyond recovery, the CoAx-selected heads are causally load-bearing: freezing them after primary removal sharply reduces the IOI margin, while adding them to the incomplete circuit reduces incompleteness from 0.75 to 0.21. More broadly, conditional growth aligns with intervention-derived repair in 11/12 held-out instances across 4 mechanism clusters, and CoAx completions outperform matched random completions on all 8 non-GPT-2 models spanning 6 architecture families. Together, causal explanations of self-repairing transformers must account for backup circuitry when primary components fail.

cs.LG

Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training

Scientific Fitness Coaching (SFC) is typically delivered by human professionals, making it costly and inaccessible to many. While recent advances in Large Language Models (LLMs) show considerable promise for more inclusive fitness coaching, directly deploying prevailing general-purpose LLMs in SFC reveals critical limitations. These models often lack sufficient domain-specific knowledge integration, leading to weak performance on complex SFC scenarios. In this paper, we introduce FitOne, a series of fitness LLMs (with 8B and 32B parameters) designed to improve reliability and domain specialization for SFC applications. Built upon the Qwen3 foundation models, FitOne is developed through a three-stage post-training pipeline consisting of continual pre-training, supervised fine-tuning, and reinforcement learning, using large-scale, high-quality datasets derived from rigorous knowledge engineering. We conduct comprehensive evaluations of FitOne on professional fitness certification exams, including ACSM-EP and NSCA-CSCS, as well as general capabilities such as knowledge reasoning and instruction following. Experimental results show that, while retaining strong general capabilities, FitOne-8B/32B achieves average improvements of up to 10.09%/9.29% and 12.73%/7.01% on the ACSM-EP and NSCA-CSCS exams, respectively, compared with the Qwen3 base models. Furthermore, in-depth ablation studies confirm the necessity of each training stage, highlighting the pipeline's effectiveness in balancing domain expertise enhancement with general ability retention. We believe this research advances LLM systems toward more reliable fitness intelligence and will inspire future research on developing domain-specific LLMs.

cs.AI

Contextual Cellular Growth (ConCeG) of neural cells for realistic grey matter tissue generation for diffusion MRI simulations

Accurate interpretation of diffusion magnetic resonance imaging (dMRI) signals in grey matter (GM) remains challenging due to the complex, heterogeneous, and densely packed cellular environment. Numerical phantoms provide a controlled framework for investigating the relationship between microstructure and diffusion signals, yet existing approaches often lack the morphological realism and multi-cellular organisation required to faithfully represent GM tissue. In this work, we introduce Contextual Cellular Growth (ConCeG), a generative framework for creating individual cells or constructing dense, three-dimensional, multi-cellular GM substrates informed by real neuronal and glial morphologies. The method combines topological neuron synthesis with a spatially constrained growth network, allowing for the controlled generation of heterogeneous cellular environments with realistic intra- and extracellular compartments. Synthetic cells are generated using morphological and topological characteristics derived from biological reconstructions. We validate the framework through comparisons of structural features with real cellular data, demonstrating strong agreement in branch order, length, angle, and tortuosity distributions. Power spectrum analysis further shows that both intracellular compartments reproduce the spatial correlations observed in biological tissue. Together, these results show ConCeG provides a biologically grounded framework for generating grey matter substrates suitable for large scale diffusion MRI simulation.

physics.med-ph

Controlling Atom Array in an Ultra-high-cooperativity Optical Cavity

Neutral-atom array and cavity quantum electrodynamics offer complementary strengths for quantum science: scalable, reconfigurable qubit architectures and strong coherent light-matter coupling. Combining them in a single platform requires an optical cavity with simultaneously high cooperativity, sufficient mode volume to accommodate atom array, and ample side optical access for atom trapping, imaging, cooling, and rearrangement, a combination that is challenging to achieve. Here we realize an atomic array integrated with a millimeter-scale Fabry--Pérot cavity whose optically-characterized single-atom cooperativity reaches $η_{\mathrm{cav}}=125\pm13$. Atom-cavity transmission spectra of trapped atoms yield an effective spectroscopic cooperativity $η_{\mathrm{spec}}=106\pm 3$, providing an in-situ verification of strong coupling in the integrated platform. In this integrated system, we achieve a state-measurement fidelity of $99.925^{+15}_{-18}\%$ for a single atom with a 4-$μ$s readout time. We also demonstrate simultaneous coupling of up to 21 individually trapped atoms to the antinode of the cavity mode. The key technical advance is a two-step mirror-fabrication method combining precision mechanical shaping and carbon-dioxide laser polishing, which produces concave fused-silica mirrors with sub-millimeter radii of curvature and residual roughness below $2~Å$. Our results establish a regime of cavity-integrated atomic array that simultaneously provides high cooperativity, large mode volume, and flexible manipulation of individual atoms, opening opportunities for cavity-assisted quantum state readout and long-range entanglement-engineering in atom-array platforms.

quant-ph

When arrow patterns meet classical patterns

Seeking to bridge the structural divide between a permutation's cycle notation and its one-line notation, Berman and Tenner introduced a novel notion of permutation pattern known as the arrow pattern. Recently, Archer and Laudone initiated a systematic study of arrow pattern avoidance, leaving behind three intriguing conjectures. In this paper, we resolve all three conjectures. First, we enumerate all six subclasses of permutations that simultaneously avoid a classical pattern of length 3 and a fixed arrow pattern of length 3, thereby confirming the first two conjectures. Second, we settle the third conjecture (which involves a different arrow pattern) by providing two independent proofs. These proofs rely on a restriction of Biane's bijection to non-nesting involutions and Krattenthaler's bijection from 321-avoiding permutations to Dyck paths, respectively.

math.CO

Optimal rates of uniform convergence for weighted Birkhoff averages via almost all rotations

In this paper, we investigate weighted Birkhoff averages for toral translations associated with compactly supported weighting functions. By introducing several new analytical techniques, we establish optimal uniform convergence rates for almost all rotations and specific (or even all) initial points. Unlike the $\mathcal{O}(N^{-1})$ rate best achieved in classical ergodic theory, we show that these weighted averages exhibit polynomial or even exponential convergence. We establish the optimality of these convergence rates in multiple aspects, particularly concerning regularity indices across four distinct cases: finite differentiability, the $C^\infty$ class, logarithmic $C^\infty$ classes, and Gevrey classes. Our results demonstrate that the regularity of the observable essentially dictates the convergence rate; furthermore, we prove that no admissible choice of weighting function can, in general, overcome the lower bounds imposed by this regularity. In contrast to the generically slow convergence of standard time averages, this work provides an optimal and nearly complete characterization of rapid convergence for weighted Birkhoff averages.

math.DS

Exoplanet Detection Using Adaptive Quantum-Optimal Measurement

Detecting terrestrial exoplanets in the habitable zones of nearby stars remains a critical challenge. Such planets can be \(10^8\) to \(10^{10}\) times fainter than their host stars and lie at diffraction-limited angular separations, where starlight strongly obscures the companion signal. Here we present an adaptive quantum measurement method for estimating the number, positions, and brightnesses of mutually incoherent point sources in the sub-Rayleigh, ultra-high-contrast regime, operating at contrasts down to \(10^{-8}\) -- five orders of magnitude beyond previous quantum imaging approaches to exoplanet detection. The method adopts a spatial-mode basis that is updated to maximize the quantum Fisher information per detected photon. Estimation is performed by maximum likelihood in log-brightness coordinates, and the source count is determined by Bayesian-information-criterion (BIC) model selection directly from photon-count statistics, without a tunable detection threshold. For point sources within sub-Rayleigh separations and with brightness ratios spanning eight orders of magnitude, the method reconstructs complete scenes with a mean success rate of \(72.5\%\). Furthermore, it is robust to misalignment, maintaining a \(71.3\%\) success rate under offsets of up to six pixels. These results demonstrate that terrestrial exoplanets can be detected below the Rayleigh limit, a regime previously inaccessible to direct imaging.

physics.optics