SearcharxivSearch

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 91 records · Page 5Linked to original sources

Efficient third-order iterative algorithms for computing zeros of special functions

This manuscript presents a novel and reliable third-order iterative procedure for computing the zeros of solutions to second-order ordinary differential equations. By approximating the solution of the related Riccati differential equation using the trapezoidal rule, this study has derived the proposed third-order method. This work establishes sufficient conditions to ensure the theoretical non-local convergence of the proposed method. This study provides suitable initial guesses for the proposed third-order iterative procedure to compute all zeros in a given interval of the solutions to second-order ordinary differential equations. The orthogonal polynomials like Legendre and Hermite, as well as the special functions like Bessel, Coulomb wave, confluent hypergeometric, and cylinder functions, satisfy the proposed conditions for convergence. Numerical simulations demonstrate the effectiveness of the proposed theory. This work also presents a comparative analysis with recent studies.

math.NA

Robotic Tele-Operation for Upper Aerodigestive Tract Microsurgery: System Design and Validation

Upper aerodigestive tract (UADT) treatments frequently employ transoral laser microsurgery (TLM) for procedures such as the removal of tumors or polyps. In TLM, a laser beam is used to cut target tissue, while forceps are employed to grasp, manipulate, and stabilize tissue within the UADT. Although TLM systems may rely on different technologies and interfaces, forceps manipulation is still predominantly performed manually, introducing limitations in ergonomics, precision, and controllability. This paper proposes a novel robotic system for tissue manipulation in UADT procedures, based on a novel end-effector designed for forceps control. The system is integrated within a teleoperation framework that employs a robotic manipulator with a programmed remote center of motion (RCM), enabling precise and constrained instrument motion while improving surgeon ergonomics. The proposed approach is validated through two experimental studies and a dedicated usability evaluation, demonstrating its effectiveness and suitability for UADT surgical applications.

cs.RO

Learning Password Best Practices Through In-Task Instruction

Users often make security- and privacy-relevant decisions without a clear understanding of the rules that govern safe behavior. We introduce pedagogical friction, a design approach that inserts brief, instructional interactions at the moment of action. We evaluate this approach in the context of password creation, a familiar task with clear quality criteria. We conducted a randomized study with 128 participants across four interface conditions that varied the depth and interactivity of guidance. We assessed three outcomes: (1) rule compliance in a subsequent password task without guidance, (2) accuracy on survey questions tied to password rules, and (3) behavior-knowledge alignment, which captures whether participants who correctly followed a rule also recognized it on the survey. Across the guided conditions, participants corrected most rule violations in the follow-up task and showed high behavior-knowledge alignment. Survey results suggested clearer advantages for some rule types, especially symbol related questions. These results position pedagogical friction as a lightweight intervention for security- and privacy-critical interfaces.

cs.HC

Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model

Multimodal aspect-based sentiment analysis (MABSA) aims to identify aspect-level sentiments by jointly modeling textual and visual information, which is essential for fine-grained opinion understanding in social media. Existing approaches mainly rely on discriminative classification with complex multimodal fusion, yet they lack explicit sentiment explainability. In this paper, we reformulate MABSA as a generative and explainable task, proposing a unified framework that simultaneously predicts aspect-level sentiment and generates natural language explanations. Based on multimodal large language models (MLLMs), our approach employs a prompt-based generative paradigm, jointly producing sentiment and explanation. To further enhance aspect-oriented reasoning capabilities, we propose a dependency-syntax-guided sentiment cue strategy. This strategy prunes and textualizes the aspect-centered dependency syntax tree, guiding the model to distinguish different sentiment aspects and enhancing its explainability. To enable explainability, we use MLLMs to construct explanation-augmented datasets for fine-tuning. Experiments show that our approach not only achieves overall gains in sentiment classification accuracy, but also produces coherent and aspect-grounded explanations.

cs.CL

BEAT-Net: Injecting Biomimetic Spatio-Temporal Priors for Interpretable ECG Diagnosis

Automated electrocardiogram diagnosis using deep learning remains limited by signal-agnostic representations that treat multi-lead recordings as undifferentiated time-series or images, forcing models to rediscover physiological structure implicitly. This leads to data inefficiency, poor generalization, and opaque decision boundaries misaligned with clinical reasoning. We present BEAT-Net, a supervised biomimetic framework that integrates QRS-centered biological tokenization with a hierarchical architecture mirroring the cardiologist's workflow. A QRS tokenizer converts continuous signals into semantically complete heartbeat sequences, which are processed through four specialized stages: morphological feature extraction via a Word Encoder, lead-invariant normalization through a Spatial Operator, temporal context injection by a Temporal Operator, and global reasoning using a Transformer-based Sentence Encoder. Evaluated across three large-scale benchmarks including PTB-XL, CPSC2018, and CSN, BEAT-Net achieves diagnostic accuracy of 0.924 AUC, comparable to dominant CNN baselines at 0.925 AUC, while reducing parameters by 95 percent from 2.06 million to 0.7 million. Critically, BEAT-Net surpasses the 39.5-million-parameter foundation model HeartLang on morphological Form classification, reaching 0.901 AUC compared to HeartLang's 0.832 AUC, while attaining full CNN-level performance using only 35 percent of training data and exhibiting superior cross-dataset generalization. Learned attention patterns spontaneously align with established clinical heuristics, demonstrating that explicit physiological structure provides a more efficient and interpretable alternative to massive pre-training for clinical deployment.

cs.LG

Channel Estimation in MIMO Systems Aided by Microwave Linear Analog Computers (MiLACs)

Microwave linear analog computers (MiLACs) have recently emerged as a promising solution for future gigantic multiple-input multiple-output (MIMO) systems, enabling beamforming with greatly reduced hardware and computational cost. However, channel estimation for MiLAC-aided systems remains an open problem. Conventional least squares (LS) and minimum mean square error (MMSE) estimation rely on intensive digital computation, which undermines the computational advantage offered by MiLACs. In this letter, we propose efficient LS and MMSE channel estimation schemes for MiLAC-aided MIMO systems. By designing the training precoder and combiner implemented by lossless and reciprocal MiLACs, the proposed schemes perform LS and MMSE estimation in the analog domain, leaving only simple digital scaling. They achieve identical estimation performance to their digital counterparts while significantly reducing computational complexity. Numerical results verify the effectiveness of the proposed schemes.

eess.SP

MemeLens: Multilingual Multitask VLMs for Memes

Memes are a dominant medium for online communication and manipulation because meaning emerges from interactions between embedded text, imagery, and cultural context. Existing meme research is distributed across tasks (e.g., \textit{hate, misogyny, propaganda, sentiment, humour}) and languages, which limits cross-domain generalization. To address this gap, we propose \textsc{MemeLens}, a unified multilingual, multitask explanation-enhanced Vision-Language Model (VLM) for meme understanding. We consolidate $38$ public meme datasets, filter and map dataset-specific labels into a shared taxonomy of $20$ tasks spanning harm, targets, figurative/pragmatic intent, and affect. We present a comprehensive empirical analysis across modeling paradigms, task categories, and datasets. Our findings suggest that robust meme understanding requires multimodal training, varies substantially across semantic categories, and remains sensitive to over-specialization when models are fine-tuned on individual datasets rather than trained in a unified setting. We make the experimental resources (https://github.com/MohamedBayan/MemeLens), model (https://huggingface.co/QCRI/MemeLens-VLM) and datasets (https://huggingface.co/datasets/QCRI/MemeLens) publicly available to the community.

cs.AI

Sensitivity Analysis of Singlet Vector-Like B Quarks via Photon-Induced and Z-Initiated Processes at FCC-$μp$

This study presents a systematic sensitivity analysis of singlet-type vector-like $B$ quark production at an FCC--$μp$ collider with a centre-of-mass energy of $\sqrt{s}=24.5~\mathrm{TeV}$ through photon- and $Z$-initiated production mechanisms. The analysis focuses on the $B\to Zb$ decay channel, considering the leptonic decay of the $Z$ boson and the hadronic decay of the accompanying $W$ boson, leading to the final state $\ell^+\ell^-bjj$. Detector-resolution effects are incorporated through a simplified Gaussian smearing procedure, and the discovery and exclusion sensitivities are evaluated using an Asimov-based statistical framework. For the photon-induced channel, the most favourable sensitivity is obtained for $R_L=0.05$ and $\mathcal{L}=1000~\mathrm{fb^{-1}}$. Over the mass range $M_B=2$--$3~\mathrm{TeV}$, the expected $5σ$ discovery reach is approximately $g^\ast\simeq0.263$--$0.328$, while the $95\%$ C.L. exclusion sensitivity extends to $g^\ast\simeq0.162$--$0.197$. On the other hand, the $Z$-initiated channel provides a substantially stronger sensitivity and extends the investigated mass range up to $M_B=4.5~\mathrm{TeV}$. For $R_L=0.05$ and $\mathcal{L}=500~\mathrm{fb^{-1}}$, the $5σ$ discovery reach is approximately $g^\ast\simeq0.038$--$0.056$, while the corresponding $95\%$ C.L. exclusion sensitivity reaches $g^\ast\simeq0.021$--$0.032$. These results demonstrate that photon- and $Z$-initiated single production at an FCC--$μp$ collider provide complementary probes of heavy vector-like $B$ quarks. Moreover, the $Z$-initiated channel offers particularly strong sensitivity to small effective couplings in the multi-TeV mass region beyond the present direct LHC reach.

hep-ph

Accelerator and Brake: Dynamic Persuasion with Dead Ends

This paper studies dynamic persuasion in a strategic-experimentation relationship in which the principal has a single-peaked preference over the agent's stopping time. Excessive experimentation may end in a dead end. The principal privately observes project quality, which determines the agent's payoff conditional on success, while both parties learn about feasibility only through the agent's experimentation. We show that an optimal policy uses at most two one-shot disclosures: an accelerator before the principal's ideal stopping time and a brake afterward. A local Arrow--Pratt comparison of induced payoffs over stopping time determines whether the accelerator is concentrated or gradual. Under common discounting, the comparison yields a one-shot accelerator. Under heterogeneous discounting, the one-shot result remains robust unless the agent is sufficiently more impatient than the principal, in which case the ranking reverses over an interval and the accelerator can take a one-shot--gradual--one-shot form.

econ.TH

DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response

Vision--language models (VLMs) show promise for disaster-response remote sensing, but existing benchmarks mainly emphasize scene-level or damage-centric assessment. To study this building-centric gap, we introduce \method{}, a diagnostic benchmark built on xBD, a pre/post-disaster satellite dataset with building-level damage labels. \method{} enriches building instances with OpenStreetMap-derived functional labels and contains 134{,}108 task-specific instruction records across 15 task types, spanning instance-level assessment, scene-level counting, multi-instance reasoning, and structured report generation. The benchmark supports RGB pre/post-disaster imagery, single- and multi-view instance formulations, and scene-level RGB/SAR diagnostic inputs. Experiments with general-domain and remote-sensing VLMs show that models perform better on visible damage cues than on building-function understanding, multi-instance reasoning, counting, and grounded reporting. Instruction tuning improves performance on several tasks but does not close this building-centric gap.

cs.CV

Machine-Learning-Enhanced Discretize-then-Project Reduced-Order Modeling of Turbulent Flows on Collocated Grids

This study presents a hybrid reduced-order modeling (ROM) framework for incompressible flows on collocated finite-volume grids, combining a discretize-then-project consistent-flux formulation for velocity and pressure with a non-intrusive neural-network closure for turbulent viscosity. The intrusive formulation preserves discrete mass conservation and pressure-velocity coupling, while a reduced pressure reference-cell constraint fixes pressure gauge ambiguity. We evaluate Multilayer Perceptron (MLP), Transformer, and Long Short-Term Memory (LSTM) closures. For a three-dimensional lid-driven cavity at $Re=100$, the LSTM-based ROM achieves relative errors of 0.7% in velocity and 4% in turbulent viscosity. At $Re=3200$, a mode-sensitivity study identifies $N=15$ POD modes as the best overall configuration, balancing accuracy, dimension, robustness, and cost. It yields a final relative velocity error of approximately 12.3% and an online wall-clock speedup of approximately $50\times$ over the full-order model; energy and enstrophy errors remain below 11% for all three architectures. This regime requires case-specific neural-network retraining and pressure reference-cell parameter retuning. In a time-extrapolation test trained on $t\in[0,3]$,s and rolled out to $t=6$,s, the ROM remains bounded, although velocity and pressure errors increase beyond the training window. The LSTM turbulent-viscosity closure remains robust, identifying long-horizon pressure accuracy as the main limitation. These results demonstrate the potential of consistent projection-based modeling combined with data-driven turbulence closure for efficient reduced-order simulation.

math.NA

Collective excitations in chiral spin liquid: chiral roton and long-wavelength nematic mode

Chiral spin liquid (CSL) is a magnetic analogue of the fractional quantum Hall (FQH) liquid. Collective excitations play a vital role in shaping our understanding of these exotic quantum phases of matter and their quantum phase transitions. While the magneto-roton and long-wavelength chiral graviton modes in the FQH and fractional Chern insulator (FCI) liquids have been extensively explored, whether CSLs host analogous or qualitatively different modes remains elusive. Here we explore the collective excitations in the SU(2) symmetric CSL phase. Combining exact diagonalization and time-dependent variational principle calculations, we identify two spin-singlet collective modes: a chiral p-wave roton mode at finite momentum, and a elliptically polarized d-wave nematic mode at zero momentum, both of which are prominent across the CSL phase. The chiral p-wave singlet roton has no counterpart in FQH of FCI systems, and the q = 0 d-wave mode also exhibits fingerprint distinct from those of FQH/FCI liquids. We also elucidate that both singlet modes are general for CSLs on various lattice models. By tuning J2, we find the nematic mode to be pronouncedly soft, together with the spin-triplet two-spinon bound states, potentially promoting strong nematic and spin stripe instabilities. Our work paves the way for further understanding CSL from the dynamical perspective and provides new spectroscopic signatures for future experiments of CSL candidates.

cond-mat.str-el

Fermionic magic resources in disordered quantum spin chains

Fermionic non-Gaussianity quantifies a quantum state's deviation from a classically tractable free-fermionic description, constituting a necessary resource for computational quantum advantage. Here we use fermionic antiflatness (FAF) to measure this deviation across ergodic and many-body localized (MBL) regimes. We focus on the paradigmatic disordered spin-$1\!/2$ XXZ chain and its impurity variant with local interactions. Across highly excited eigenstates, FAF evolves from typical-state behavior at weak disorder to strongly suppressed values deep in the MBL regime, with volume-law scaling in the XXZ chain and an area-law bound in the impurity setting. Rare long-range cat-like eigenstates exhibit a pronounced enhancement of FAF, making it a sensitive diagnostic of mechanisms proposed to destabilize MBL. Starting from product states, we find that in the MBL regime FAF grows slowly in time, approaching saturation via a power-law relaxation. Overall, our results show that MBL suppresses fermionic non-Gaussianity, and the associated complexity beyond free fermions, while ergodicity restores it, motivating explorations of fermionic non-Gaussianity in other ergodicity-breaking phenomena.

quant-ph

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

End-to-end autonomous driving models increasingly benefit from large vision-language models for semantic understanding, yet safe and reliable planning under long-tail conditions remains challenging, particularly in mixed-traffic environments involving heterogeneous road users and rare safety-critical interactions. This paper proposes HERMES, a holistic risk-aware end-to-end multimodal driving framework that explicitly incorporates long-tail semantic knowledge into trajectory planning. HERMES employs a foundation-model-assisted annotation pipeline to construct structured Long-Tail Scene Context and Long-Tail Planning Context, capturing hazard-centric scene information, maneuver intent, and risk-aware planning guidance. A Tri-Modal Driving Module then integrates multi-view visual observations, historical ego-motion, and long-tail semantic instructions through intent- and risk-aware conditioning for trajectory generation. Extensive experiments on a large-scale real-world long-tail driving benchmark demonstrate consistent improvements over representative recent baselines in overall planning performance and across diverse safety-critical scenarios. Ablation studies further validate the effectiveness and complementary roles of the major components within HERMES.

cs.RO

RankSteer: Can Pointwise LLM Rankers Be Calibrated at the Representation Level?

Large language models (LLMs) are strong zero-shot pointwise rankers, but lag behind pairwise and listwise methods. Beyond missing comparative signals, we identify a \textit{calibration gap}: ranking-relevant information encoded in hidden states is not fully captured by the scalar output head. We propose RankSteer, a post-hoc activation-steering framework that calibrates ranking via projection-based interventions along multiple directions at inference time: decision, evidence, and, optionally, role. This is achieved without updating model weights or introducing cross-document comparisons. We instantiate RankSteer on two structurally distinct pointwise variants and observe improvements over their respective baselines on most TREC DL and BEIR datasets across three backbones. This suggests that the calibration gap is a general property of pointwise rankers. Our additional geometric analysis shows that steering improves ranking by concentrating each query's document representations along an existing ranking geometry, offering new insight into how LLMs internally represent and calibrate relevance judgments.

cs.IR

MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

Audio large language models (AudioLLMs) enable instruction following over speech and general audio, but progress is limited by the scarcity of diverse, conversational, and instruction-aligned speech--text data. This gap is particularly pronounced for persona-grounded and dialectal interactions, where collecting real multi-speaker recordings remains costly and slow. We introduce MENASpeechBank, a reference speech bank comprising ~18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. We develop a controllable data pipeline that (i) constructs persona profiles enriched with World Values Survey (WVS) inspired attributes, (ii) defines a taxonomy driven ~5Kconversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates ~417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) produces speaker-conditioned user-turn audio (synthetic) from reference recordings to preserve speaker diversity. We evaluate synthetic and human recorded conversations and provide an analysis. We will make the MENASpeechBank available for the community.(\href{https://huggingface.co/datasets/QCRI/MenaSpeechBank)

cs.SD

Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling

Time series (TS) modeling has come a long way from early statistical, mainly linear, approaches to the current trend in TS foundation models. With a lot of hype and industrial demand in this field, it is not always clear how much progress there really is. To advance TS forecasting and analysis to the next level, here we argue that the field needs a dynamical systems (DS) perspective. TS of observations from natural or engineered systems almost always originate from some underlying DS, and arguably access to its governing equations would yield theoretically optimal forecasts. This is the promise of DS reconstruction (DSR), a class of ML/AI approaches that aim to infer surrogate models of the underlying DS from data. But models based on DS principles offer other profound advantages: Beyond short-term forecasts, they enable to predict the long-term statistics of an observed system, which in many practical scenarios may be the more relevant quantities. DS theory furthermore provides domain-independent theoretical insight into mechanisms underlying TS generation, and thereby will inform us, e.g., about upper bounds on performance of any TS model, generalization into unseen regimes as in tipping points, or potential control strategies. After reviewing some of the central concepts, methods, measures, and models in DS theory and DSR, we will discuss how insights from this field can advance TS modeling in crucial ways, enabling better forecasting with much lower computational and memory footprints. We conclude with a number of specific suggestions for translating insights from DSR into TS modeling.

cs.LG

Progressive Binarization - Pauli Correlation Encoding: a Continuation Method for Constrained Optimization

Pauli Correlation Encoding (PCE) reduces the qubit requirements of quantum optimization by embedding the problem variables into the expectation values of Pauli observables, so that the number of qubits can be much smaller than the number of variables. PCE has not yet been studied for constrained optimization. We extend it to constrained combinatorial problems, using the budget-constrained MinCut as a case study, and show that the standard formulation fails to reliably enforce the constraint: feasibility hinges on the binarization of the encoded variables, which depends sensitively on hyperparameters that are hard to tune and do not transfer across instances. To address this, we introduce Progressive-Binarization PCE (PB-PCE), an adaptive continuation scheme that progressively increases the binarization parameter while re-optimizing the circuit from the previous solution, driving the variables towards the binary domain. PB-PCE attains near-complete constraint satisfaction (88--100\%) and smaller cut sizes than standard PCE, with a number of stages (10--20) essentially independent of problem size, solving instances of up to 300 variables with only 9-qubit circuits.

quant-ph