SearcharxivSearch

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 109 records · Page 6Linked to original sources

A variational framework for modal estimation

Multivariate mode estimation arises in many statistical problems such as inverse problems, multimodal sampling, and density-based clustering, but becomes challenging in moderate to high dimensions, especially when the underlying density is not directly evaluable. We introduce GERVE (Gibbs-measure Entropy-Regularized Variational Estimation), a sample-based method for estimating multivariate modes by approximating Gibbs distributions directly from samples, without estimating or evaluating the density. GERVE uses Gaussian-mixture variational annealing and natural-gradient optimization, producing a mixture concentrated in high-density regions whose component responsibilities also provide a clustering of the observations. We prove theoretical guarantees in two regimes: as the Gibbs temperature goes to zero, the optimal variational mixture concentrates around the global modes of the population density; at fixed positive temperature, we prove existence, consistency, and asymptotic normality of empirical maximizers and propose a bootstrap procedure for uncertainty quantification. Simulations and a real-data experiment show that GERVE accurately recovers modes and produces meaningful clusters.

stat.ME

EgoPush: Egocentric Multi-Object Rearrangement for Mobile Robots via Constrained Teacher Observability

Humans rearrange objects in cluttered environments using egocentric perception, actively moving to keep task-relevant spatial cues in view. Mobile robots have not matched this: rearrangement is usually built on a global pose estimate or a map, which is exactly what a robot carrying one camera lacks, while pushing keeps changing the scene it would have to be built from. We present EgoPush, which pushes objects into anchor-relative formations from onboard RGB-D alone, with no global localization, external tracking, or map at deployment, and transfers zero-shot to a TurtleBot in controlled and visually cluttered scenes. What makes this learnable turns out to be a property of the teacher rather than of the student: three privileged teachers trained with identical rewards, architecture, and hyperparameters all exceed $98\%$ success, yet their distilled egocentric students reach $0\%$, $54.8\%$, and $87.3\%$, the only variable being the teacher's observation function. EgoPush therefore trains the teacher under egocentric observability constraints, restricting it to visibility-limited cues and revealing target references only when the anchor is centrally visible, so that its supervision is recoverable by a depth-based student distilled online. Making the teacher trainable in the first place needs two further pieces: a role-grouped object-centric interface shared by teacher and student, and stage-wise temporally decayed rewards for long-horizon credit assignment. Videos, the playable task, and code are available at https://ai4ce.github.io/EgoPush/.

cs.RO

Learning End-to-End Control for Omnidirectional Aerial Motion on Overactuated Tilt-rotor Quadrotors

While reinforcement learning (RL) has been successfully applied to conventional quadrotors for agile and robust flight, whether actuator-level RL can be reliably deployed on tilt-rotor aerial robots remains an open question, as the hybrid actuation coupling brushless rotors with rotational joints introduces a substantially harder sim-to-real gap. In this work, we propose an end-to-end RL framework for omnidirectional motion control on overactuated tilt-rotor quadrotors, directly mapping target poses to joint and rotor commands. The learning framework combines actuator-level simulation with an asymmetric actor-critic architecture for 6D pose-reaching. For reliable sim-to-real transfer on the hybrid actuation, we integrate system identification with minimal yet physically grounded domain randomization. The trained policy is deployed zero-shot on real hardware and evaluated across waypoint hovering, external disturbances, payload variation and trajectory tracking, together with simulated traversal of allocation-singular configurations. The policy is compared with a state-of-the-art NMPC baseline: NMPC attains lower steady-state position error, whereas the RL policy offers a more uniform orientation error across evaluations, transitions between poses faster, and requires less onboard computation.

cs.RO

VideoPulse: Neonatal heart rate and peripheral capillary oxygen saturation (SpO2) estimation from contact free video

Remote photoplethysmography (rPPG) enables contact free monitoring of vital signs and is especially valuable for neonates, since conventional methods often require sustained skin contact with adhesive probes that can irritate fragile skin and increase infection control burden. We present VideoPulse, a neonatal dataset and an end to end pipeline that estimates neonatal heart rate and peripheral capillary oxygen saturation (SpO2) from facial video. VideoPulse contains 157 recordings totaling 2.6 hours from 52 neonates with diverse face orientations. Our pipeline performs face alignment and artifact aware supervision using denoised pulse oximeter signals, then applies 3D CNN backbones for heart rate and SpO2 regression with label distribution smoothing and weighted regression for SpO2. Predictions are produced in 2 second windows. On the NBHR neonatal dataset, we obtain heart rate MAE 2.97 bpm using 2 second windows (2.80 bpm at 6 second windows) and SpO2 MAE 1.69 percent. Under cross dataset evaluation, the NBHR trained heart rate model attains 5.34 bpm MAE on VideoPulse, and fine tuning an NBHR pretrained SpO2 model on VideoPulse yields MAE 1.68 percent. These results indicate that short unaligned neonatal video segments can support accurate heart rate and SpO2 estimation, enabling low cost non invasive monitoring in neonatal intensive care.

eess.IV

Joint Subcarrier Phase Recovery for Nonlinearity Mitigation

We propose a low-complexity phase recovery scheme that simultaneously mitigates laser phase noise and fiber nonlinearity across several subcarriers. In a long single-span link with Raman amplification, the scheme achieves 0.9 dB gain with 99 real multiplications per complex symbol.

eess.SP

ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation

Achieving autonomous and versatile whole-body loco-manipulation remains a central barrier to making humanoids practically useful. Yet existing approaches are fundamentally constrained: retargeted data are often scarce or low-quality; methods struggle to scale to large skill repertoires; and, most importantly, they rely on tracking predefined motion references rather than generating behavior from perception and high-level task specifications. To address these limitations, we propose ULTRA, a unified framework with two key components. First, we introduce a physics-driven neural retargeting algorithm that translates large-scale motion capture to humanoid embodiments while preserving physical plausibility for contact-rich interactions. Second, we learn a unified multimodal controller that supports both dense references and sparse task specifications, under sensing ranging from accurate motion-capture state to noisy egocentric visual inputs. We distill a universal tracking policy into this controller, compress motor skills into a compact latent space, and apply reinforcement learning finetuning to expand coverage and improve robustness under out-of-distribution scenarios. This enables coordinated whole-body behavior from sparse intent without test-time reference motions. We evaluate ULTRA in simulation and on a real Unitree G1 humanoid. Results show that ULTRA generalizes to autonomous, goal-conditioned whole-body loco-manipulation from egocentric perception, consistently outperforming tracking-only baselines with limited skills.

cs.RO

Continuous-time multi-armed bandits under random intervention times

This paper examines multi-armed bandits with $J$ independent arms in which actions are taken at random discrete times. When an arm is operated, it must remain active for a renewal inter-arrival time. For arms evolving as a Lévy process, we provide an explicit characterization of the Gittins index, known to yield an optimal strategy. Furthermore, when the inter-arrival times are exponential and the arms evolve as a spectrally negative Lévy process, a reflected spectrally negative Lévy process, or a diffusion process, the Gittins index is explicitly characterized in terms of the scale function or diffusion characteristics, respectively. Convergence analysis and numerical experiments are performed to support the theoretical results.

math.OC

EventGeM: Global-to-Local Feature Matching for Event-Based Visual Place Recognition

Event cameras are rapidly rising in popularity for robotic and computer vision tasks because their sparse activation delivers energy-efficient, high-dynamic-range, and fast sensing. Event cameras have been used in robotic navigation and localization tasks where positioning must occur in real time with sufficient accuracy. However, current event-based localization methods suffer from poor spatial understanding and are not viewpoint tolerant. In this paper, we address the problem of viewpoint-robust place recognition directly from event streams. We present EventGeM, a global-to-local feature fusion pipeline for event-based visual place recognition that combines whole-image feature detection to shortlist top candidates for 2D homography-based re-ranking with random sample consensus (RANSAC). We also contribute a regional generalized mean pooling (GeM) layer that learns to return the most relevant spatial features using per-row exponents to pool event streams into a compact global descriptor, trained on the NYC-Event-VPR dataset. These contributions overcome shortfalls in currently available event-based localization methods that fail to recognize similar places with large changes in viewpoint. To evaluate viewpoint-robust localization, we contribute a new event-based dataset that includes repeated traverses with a severe lateral shift. EventGeM improves absolute Recall@1 by 7 to 43 percentage points over the strongest baseline in each experiment. We also deploy EventGeM on a robotic platform, demonstrating real-time performance of our hierarchical pipeline. The code for EventGeM is available at https://github.com/AdamDHines/Event-GeM.

cs.CV

Introduction to Generalized Symmetries

These notes were prepared for a series of intensive lectures delivered at Hokkaido University, Nagoya University, Kyoto University, and Kyushu University. We begin with a brief review of higher-form symmetries, anomalies, and discrete gauge theories, before introducing non-invertible symmetries in $(1+1)$-dimensional systems. The basic structure of fusion categories is then discussed, including a discussion of categorical analogs of discrete gauging and representation theory. We subsequently turn to $(3+1)$-dimensional theories, where several physical applications of non-invertible symmetries are discussed. These notes are intended to be largely self-contained, and require no prior familiarity with subjects such as conformal field theory or lattice models.

hep-th

High-optical-depth, sub-Doppler-width absorption lines at telecom wavelengths in hot, optically driven rubidium vapor

Doppler broadening presents a major limitation for high-resolution spectroscopy and nonlinear optics in room-temperature atomic vapors. Here, we demonstrate the suppression of Doppler broadening accompanied by pronounced absorption on the upper transition of a three-level ladder system, achieved by dressing the intermediate state with a strong control field. As a concrete realization, we study a hot vapor of $^{87}$Rb where the lower transition is driven by a strong control field resonant with the D2 line at a wavelength of 780 nm, while a weak counter-propagating probe field at the telecom C-band wavelength of 1529 nm ($5P_{(3/2)}\leftrightarrow 4D_{(5/2)}$) interrogates the dressed states. We observe absorption features with a resonant optical depth of approximately 4 and a full width at half maximum of about 17 MHz. Remarkably, this corresponds to an order-of-magnitude reduction relative to the Doppler width, while the optical depth on the upper transition of the ladder scheme exceeds that of the Doppler-broadened lower transition. The measured spectra are in good agreement with theoretical modeling. Combining high optical density with sub-Doppler-width absorption lines typically requires laser-cooled atoms, while our approach profits from the experimental simplicity of a hot-vapor platform.

physics.atom-ph

Naturally Light Distortion

In the most general formulation of gravity, the metric and connection are independent degrees of freedom, and the connection may include torsion and non-metricity (or distortion, collectively) degrees of freedom, resulting in a huge number of possible dynamical fields. However, most fields are either non-dynamical or extremely heavy and the general relativity is recovered at low energy. We find a unique naturally light vector- or scalar-like distortion field, which can be dynamical and have phenomenological implications. In particular, a light scalar particle that mixes with the Higgs boson naturally appears.

gr-qc

Taming the Adversary: A Cost-to-Disturbance Ratio Approach to Adversarial Reinforcement Learning

Reinforcement learning (RL) policies trained in simulation often degrade once deployed on real systems, where the controller must reject external disturbances that were never encountered in simulation. Robust RL addresses this by exposing the controller to perturbations while it learns, through domain randomization, adversarial minimax formulations, or probabilistic mixtures of protagonist and adversarial behavior. However, an unregulated disturbance mechanism destabilizes training and often collapses nominal performance relative to standard, non-robust methods. We propose cost-to-disturbance ratio adversarial training (CoDRA), a framework that expresses the controller--adversary trade-off as a ratio of accumulated cost to accumulated squared disturbance norm, and optimizes it through a self-normalized actor--critic update. In this algorithm, each value term is scaled by a stop-gradient normalization constant computed from the current batch. This moderates the adversary's incentive without altering the controller's own update, and requires neither an explicit disturbance penalty nor an auxiliary trade-off parameter. We evaluate CoDRA on two MuJoCo pendulum environments under force and mass sweeps. On InvertedDoublePendulum, CoDRA attains the lowest cost at every force level, including a force outside the range seen during training, and in all but one cell of the mass grid, whereas its advantage is less pronounced on the milder InvertedPendulum.

cs.LG

LLMs and Speech: Integration vs. Combination

In this work, we study different approaches to utilize large language models (LLMs) for automatic speech recognition (ASR). Specifically, we compare the tight integration of an acoustic model (AM) with the LLM ("speech LLM") to the traditional way of combining AM and LLM via shallow fusion and provide ablations on the effect of different label units and LLM sizes. For tight integration, we further examine the effect of attention interfaces, encoder downsampling, and length normalization. Furthermore, we investigate joint recognition with a CTC model to mitigate hallucinations of speech LLMs and present effective optimizations. We train and evaluate on LibriSpeech and Loquacious and additionally evaluate on the HuggingFace ASR leaderboard. Across model sizes, we find that shallow fusion consistently outperforms tight integration of AM and LLM on in-domain data, highlighting the importance of strong shallow-fusion baselines when evaluating speech LLMs for ASR. On the more heterogeneous HuggingFace ASR leaderboard, however, the integrated prefix LLM achieves lower average WER than shallow fusion, with gains concentrated on out-of-domain corpora.

eess.AS

Constrained Feedback Control of Nonlinear Systems via Approximate HJB and Control Barrier Functions

This paper presents a two-stage framework for constrained feedback control of input-affine nonlinear systems. Offline, an approximate value function for the unconstrained problem is computed, for example using Hamilton--Jacobi--Bellman (HJB)-based policy iteration. Online, the proposed quadratic program (QP) minimizes the pre-Hamiltonian evaluated using the approximate value-function gradient subject to safety constraints enforced by control barrier functions (CBFs). This architecture decouples performance optimization from constraint enforcement, allowing constraints to be modified without recomputing the value function. As in CBF-QP architectures based on control Lyapunov functions (CLFs), safety is enforced as a hard constraint; however, the performance objective targets approximate optimality rather than a prescribed Lyapunov decay. Numerical results on a linear 2-state hovercraft and a nonlinear 9-state spacecraft attitude-control problem show agreement with the constrained open-loop optimal control problem (OCP) benchmark in the linear case, and performance close to the OCP benchmark, improving on CLF-based controllers, in the nonlinear case.

eess.SY

How do LLMs Compute Verbal Confidence

Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs internally generate such scores remains unknown. We address two questions: first, when confidence is computed -- just-in-time when requested, or automatically during answer generation and cached for later retrieval; and second, what verbal confidence represents -- token log-probabilities, or a richer evaluation of answer quality? Focusing on Gemma 3 27B (across TriviaQA, BigMath, and MMLU), Qwen 2.5 7B, and the reasoning model Magistral Small 24B, we provide convergent evidence for cached retrieval. Activation steering, patching, noising, and swap experiments reveal that confidence representations emerge at answer-adjacent positions before appearing at the verbalization site. Attention blocking pinpoints the information flow: confidence is gathered from answer tokens, cached at the first post-answer position, then retrieved for output. Critically, linear probing and variance partitioning reveal that these cached representations explain substantial variance in verbal confidence beyond token log-probabilities, suggesting a richer answer-quality evaluation rather than a simple fluency readout. These findings demonstrate that verbal confidence reflects automatic, sophisticated self-evaluation -- not post-hoc reconstruction -- with implications for understanding metacognition in LLMs and improving calibration.

cs.CL

Mid-infrared reconfiguration of population flow in lanthanide nanocrystals

Converting mid-infrared (MIR) radiation to visible or near-infrared wavelengths is essential for imaging and sensing, yet achieving sensitive, low-power, and scalable detection remains challenging. Lanthanide nanocrystals provide an alternative through ratiometric luminescence but are typically constrained by Boltzmann statistics, which tie population distributions to lattice temperature and limit signal contrast. Here we show that MIR irradiation rebalances dissipative relaxation pathways, driving lanthanide emitters into a non-Boltzmann steady state that enables non-thermal control of population distributions. This allows emission behaviors inaccessible under thermal equilibrium. We exploit this regime to achieve linear MIR detection with respect to MIR power across 6.8 to 8.6 micrometers. The ratiometric response is intrinsically independent of the pump power, enabling operation at an ultralow excitation power of 10 uW, several orders of magnitude lower than conventional approaches. Using standard silicon photodetectors, we then demonstrate room-temperature MIR imaging with detection limits approaching 4 nW um-2. Our results establish lanthanide nanoparticles as an efficient platform for MIR conversion and sensing in nanophotonic systems.

physics.optics

Objective Model Prior Probabilities in Variable Selection

For many years it was routine to use equal model prior probabilities in Bayesian model uncertainty analysis. At least twenty years ago it became clear that this was problematic, leading to support of much too large models in the increasingly huge model spaces being considered in genomics and other fields. A popular replacement was to adopt a suggestion of Harold Jeffreys for the variable selection problem in which a total of $k$ possible variables are being considered for inclusion in the model: give the collection of all models containing $d$ variables ($d = 0, . . . , k$) prior probability $1/(k + 1)$ and then divide this prior probability equally among the models in the collection. Many other choices of model prior probabilities that impose severe parsimony have also been introduced. We begin by reviewing the problems with using equal model prior probabilities and then discuss some serious problems with the Jeffreys choice. Finally, we introduce and study a number of objective alternative choices of model prior probabilities, from both numerical and theoretical perspectives.

stat.ME

Causal Evidence that Language Models use Confidence to Drive Behavior

Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can be extracted from language model outputs, yet a fundamental question remains: do models actually use these signals to control behavior, such as deciding whether to answer or abstain? To investigate, we developed a four-phase paradigm. Phase~1 elicited baseline confidence estimates without an abstention option. Phase~2 revealed that LLMs apply an implicit threshold to internal confidence when deciding to abstain, with confidence effect sizes approximately an order of magnitude larger than alternative mechanisms. Phase~3 provided direct causal evidence through activation steering: boosting or suppressing confidence signals correspondingly decreased or increased abstention rates. Phase~4 extended this by systematically varying instructed thresholds, demonstrating that LLMs actively deploy confidence signals to implement abstention policies. Critically, beyond calibrated log-probability based confidence derived from the output distribution, verbal confidence independently predicted abstention across all models, despite being objectively less discriminatory of answer correctness. Activation decoding at the last pre-answer token further showed that both observable measures are lossy readouts of a richer internal representation. Together, these results suggest that abstention is not fully captured by the strength of evidence in the output distribution alone, but is better explained by the joint operation of a multidimensional internal confidence representation and threshold-based policies -- consistent with structured metacognitive control in LLMs, a capacity of growing importance as models transition to autonomous agents that must recognize their own uncertainty.

cs.LG