SearcharxivSearch

arXiv subjects

Yu-Feng Li

Publications and source records attributed to Yu-Feng Li.

At least 19 recordsLinked to original sources

Quantifying Information Hierarchy for Neutrino Oscillation Parameters at JUNO

Since neutrinos are quantum systems inherently, the precision with which oscillation parameters can be estimated ultimately depends on how much information about these parameters is encoded in the neutrino state and how efficiently that information can be extracted through measurement. In this work, we quantify how information encoded in reactor antineutrino states flows through the measurement process to the events observed at the detector, using quantum and classical Fisher information. We establish the information ladder for JUNO, revealing that the loss of precision across different information levels is strongly parameter dependent. We demonstrate that the JUNO configuration approaches the optimal statistical limit for the oscillation parameters of the solar sector, while information on $\theta_{13}$ and $\Delta m_{31}^{2}$ is significantly degraded by the measurement strategy and detector effects. Despite this information loss, the remaining information is sufficient for JUNO to achieve sub-percent precision on $\Delta m_{31}^{2}$ within six years.

hep-ph

LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory

Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen environments. Recent progress has been largely driven by Multimodal Large Language Models (MLLMs). Existing methods follow a next-step action prediction paradigm, supervising only the expert action, which requires a high quantity of data for training. They also rely on cognitive maps, accumulated historical frames, or external 3D tools to maintain states, leading to high computational and memory overhead. To realize resource efficiency VLN, we propose LookStep, a unified end-to-end framework that combines Language Centric Future State Modeling and Event Driven Rolling Memory that uses language labels to generate coarse-grained navigation progress and future states for each candidate action, while autonomously deciding whether to write each observation into a bounded rolling memory with a semantic role. We validate LookStep empirically. On VLN-CE tasks, LookStep outperforms existing methods under the same training settings, achieving a 49.7\% success rate on R2R-CE Val-Unseen with better memory efficiency and less data usage. Code and model is available at https://github.com/kunyang-YU/LookStep.

cs.CV

Characterizing Single-Signal Events from Atmospheric-Neutrino Neutral-Current Interactions in Large Liquid Scintillator Detectors

Neutral-current interactions of atmospheric neutrinos in large liquid scintillator detectors offer a new opportunity to study single-signal events (hereafter singles), characterized by a prompt energy deposition on the MeV-to-GeV scale and no identified delayed signal. In this work, we systematically investigate the model dependence of atmospheric-neutrino singles due to the primary neutrino-nucleus interaction, residual-nucleus de-excitation, and secondary interactions in the scintillator. Our results show that the dominant model dependence originates from the primary neutrino-nucleus interaction, especially for neutral-current processes on carbon, whereas de-excitation is essential for the singles selection yet leads to relatively small spectral variations among realistic models. Secondary-interaction effects are also subdominant overall. We further present the predicted event rates and prompt-energy spectra for neutral-current singles, along with the charged-current contribution. Separately, we estimate the low-energy contribution from elastic scattering of sub-\SI{100}{\MeV} atmospheric neutrinos on free protons. These results highlight the physics potential of current and future large liquid scintillator detectors, such as the Jiangmen Underground Neutrino Observatory, to study atmospheric-neutrino singles, probe neutrino-nucleus interaction models, and improve background estimates for rare-event searches.

hep-ex

SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data

Tabular data, as a core data format in machine learning, often lacks the discriminative power needed for high-performance modeling due to insufficient feature informativeness. Automated Feature Engineering (AutoFE) overcomes this by automating feature generation and selection, ensuring both model performance and operational efficiency. However, traditional AutoFE often yield features with poor interpretability because they rely on blind mathematical transformations, while large language models (LLM)-based AutoFE faces challenges in requiring costly multi-round iterations to generate high-utility features to effectively enhance model performance, compounded by inherent risks of bias and hallucination. In this paper, we combine symbolic regression with LLMs for feature engineering (SymboLLM-FE) to solve these challenges. We extract mathematically expressive formulas strongly correlated with the target via symbolic regression, which can enhance model performance, then refine them by LLMs with rich prior knowledge to ensure interpretability. Empirical results on six real-world datasets and four Kaggle competitions demonstrate that SymboLLM-FE outperforms existing AutoFE. SymboLLM-FE also addresses the dual challenges of poor interpretability and numerous iterations by employing a statistical prior-grounded LLM refinement mechanism and single-digit LLM calls.

cs.LG

Binary Hypothesis Testing: A Robust Framework Against the Look Elsewhere Effect

In particle physics, discovery claims conventionally require an observed significance exceeding $5\sigma$. However, the interpretation of a $5\sigma$ result depends critically on the testing procedure, namely whether the hypothesis is tested at a single pre-specified point in parameter space or by scanning over a range of possible signal locations. This distinction gives rise to the look-elsewhere effect, a concept that is widely used but often interpreted as a simple penalty for scanning. In this work, we reinterpret the look-elsewhere effect as a correction for procedural inconsistency arising when the null distribution is generated under one procedure while the test statistic is evaluated under another. Within the framework of hypothesis testing, we examine its implications and clarify the distinct roles of the look-elsewhere effect in binary and peak-search scenarios. In particular, we show that binary test is robust against the look-elsewhere effect, whereas peak searches require an explicit correction for the search over signal locations. Using a moderate trial factor of approximately 26, calibrated from the ATLAS Higgs search, we show that a $3\sigma$ global significance of peak search can correspond to approximately $4\sigma$ significance of the binary test for the same value of the observed test statistic. This reformulation provides a clearer statistical interpretation of the look-elsewhere effect and offers a more coherent framework for understanding significance claims in particle physics.

hep-ex

Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning

Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that demand both precise spatial perception and fine-grained geometric computation beyond end-to-end generation. Tool augmentation offers a natural solution, while existing methods either plan tool calls from scratch without explicit dependency constraints or rely on fixed pipelines that are redundant and generalize poorly across spatial tasks. An effective spatial reasoning agent should instead accumulate reusable experience and adaptively compose it for new problems. To this end, we propose NeSy-Spatial, a neuro-symbolic framework for self-evolving spatial skills. NeSy-Spatial abstracts tool interactions and geometric operations into typed executable atomic instructions and composes them into two complementary skill types: Tool-Use Skills for organizing tool execution and Geometry Skills for structured geometric reasoning. During inference, NeSy-Spatial retrieves and executes relevant skills in a closed-loop process. During evolution, it analyzes buffered successful and failed trajectories to refine skill structures and prune unreliable or inactive entries. Experiments on three spatial reasoning benchmarks show that NeSy-Spatial consistently improves reasoning accuracy with more precise tool utilization.

cs.AI

Precision three-Dimensional Atmospheric Neutrino Flux Calculation Based on Honda Flux Model

We present a comprehensive three-dimensional atmospheric neutrino flux calculation based on the well-recognized simulation framework develeped by Honda and his collaborators, incorporating for the first time the muon propagation inside the Earth and its subsequent decay or nuclear capture. Other updates of essential input models include: the AMS02-based primary cosmic ray model, IGRF2020 geomagnetic field, and muon-recalibrated hadronic interaction model. The calculation covers seven detector sites across diverse geomagnetic environments, spanning 10~MeV to $10^4$~GeV. Significant site-dependent differences appear at $E_\nu < 10$~GeV, with $\nu_\mu$ flux at IceCube approximately twice that at JUNO below 1~GeV. Compared to HKKMS15, deviations of 2\%--10\% are attributed to the updated input models. Below 100~MeV, we present precise flux results, revealing that muon propagation contributes a globally significant component to the low-energy neutrino flux at all sites, with an approximately site-independent absolute increment. The hadronic uncertainty is re-estimated across the energy range using the updated hadronic interaction model, with significant reduction of the systematic error compared to previous calculations. These results provide essential inputs for neutrino oscillation and rare-event search experiments including JUNO, Super-Kamiokande/Hyper-Kamiokande, DUNE, KM3NeT/ORCA, and IceCube, as well as direct dark matter detection experiments facing the neutrino fog.

hep-ph

Leptonic CP Phase Determination from Fisher Information in NO$\nu$A and T2K

The precise determination of the leptonic CP phase $\delta_{\rm CP}$ remains one of the central objectives of current and future long-baseline (LBL) neutrino oscillation experiments. Quantum estimation theory provides a natural framework to quantify the ultimate precision limits for estimating physical parameters encoded in quantum states. In this work, we employ the quantum Fisher information to investigate how much information about $\delta_{\rm CP}$ is intrinsically encoded in neutrino states and how efficiently it is extracted in present LBL experiments such as T2K and NO$\nu$A. We first analyze the intrinsic quantum sensitivity of neutrino and antineutrino states and demonstrate how matter effects generate a neutrino mass-ordering dependent information structure. To compare the intrinsic information content of the quantum state with the information experimentally accessible through flavor measurements, we compute the event-level Fisher information from reconstructed event spectra using Poisson statistics. We find that both experiments extract only a small fraction of the total information available in the underlying quantum state. This extraction efficiency becomes particularly suppressed near maximally CP-violating regions, where the reconstructed event spectra exhibit reduced sensitivity to small variations in $\delta_{\rm CP}$. Our analysis provides a complementary information-theoretic perspective on precise estimation of oscillation parameters in LBL neutrino experiments.

hep-ph

On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective

Test-time adaptation (TTA) aims to adapt models to maintain reliable performance on non-stationary test streams without requiring labeled data. Despite its empirical success, the learnability of TTA under non-stationary streams remains unexplored. A key challenge is the lack of a principled theoretical framework that simultaneously aligns with the TTA objective and captures both continuously evolving distribution shifts and intrinsic information constraints. To address this gap, we propose the first theoretical framework for studying the learnability of TTA and introduce $(\epsilon,\delta)$-Recovery Complexity and $(\epsilon,\rho)$-TTA Learnability. Recovery complexity measures the post-shift time needed to maintain excess risk below a target level with high probability, and is further extended to TTA learnability, which measures the long-term reliability of TTA. Within this framework, we introduce a novel discrete surrogate for non-stationary test streams, enabling a unified and tractable analysis of both gradual and abrupt shifts. We derive order-wise matching lower and upper bounds on recovery complexity, revealing fundamental limits of TTA and an intrinsic adaptivity-information trade-off. These results provide unified learnability guarantees for TTA that complement regret-based analyses.

cs.LG

Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models

Looped Language Models (LoopLMs) enable efficient latent reasoning through depth recurrence, yet exhibit unreliable test-time scaling behavior: performance often peaks at a certain iteration depth and then collapses with further recurrence. Through latent dynamics analysis, we find an inherent trade-off between stability and effectiveness in existing architectures and strategies. By conceptualizing reasoning as uncertainty reduction, we propose that convergence toward stable fixed points while preserving effectiveness represents a promising way. To this end, we propose STARS (STAbility-driven Recurrent Scaling), a training framework that constrains latent states to approach asymptotically stable fixed points. This is realized via efficient Jacobian Spectral Radius Regularization with random loop sampling, enabling STARS to maximize effectiveness while ensuring rigorous stability. Experiments on arithmetic tasks show that STARS achieves reliable test-time scaling, and on complex mathematical reasoning it substantially mitigates performance degradation as recurrence depth increases while also improving peak performance.

cs.LG

Revival of the Reactor Antineutrino Anomaly

The Reactor Antineutrino Anomaly (RAA) refers to the deficit observed between the average event rate measured in reactor antineutrino experiments with respect to the theoretical prediction. This anomaly was first identified in 2011 ($2.5\,\sigma$) as a consequence of the Huber-Muller reactor antineutrino flux calculation. It was thought to be resolved in 2021 as a result of previous reactor antineutrino flux calculations, with a reduction to about $1\,\sigma$. In this work, we present the RAA obtained with the latest reactor antineutrino flux calculation published in 2023 by a French research group, which was never used before for the calculation of the RAA. It is the first summation flux model which includes a comprehensive uncertainty budget. The result indicates a revival of the RAA at the level of $2.2\,\sigma$. We also consider the usual simplest explanation of the RAA by active-sterile neutrino oscillations. We present the constraints on the oscillation parameters and we derive a tension of $3.8\,\sigma$ with the results of gallium source experiments (Gallium Anomaly) taking into account also the solar neutrino and KATRIN bounds, that of the combined short-baseline reactor spectral ratio measurements, and that of the Daya Bay search for a sub-eV sterile neutrino. Since the tension may be due to underestimated systematic uncertainties and the main tension is between the gallium data and the other data, we finally present the results of a global analysis with enlarged gallium uncertainties, which reduce the global tension to $1.3\,\sigma$.

hep-ph

Revisiting the Travel Planning Capabilities of Large Language Models

Travel planning serves as a critical task for long-horizon reasoning, exposing significant deficits in LLMs. However, existing benchmarks and evaluations primarily assess final plans in an end-to-end manner, which lacks interpretability and makes it difficult to analyze the root causes of failures. To bridge this gap, we decompose travel planning into five constituent atomic sub-capabilities, including \emph{Constraint Extraction}, \emph{Tool Use}, \emph{Plan Generation}, \emph{Error Identification}, and \emph{Error Correction}. We implement a decoupled evaluation protocol leveraging oracle intermediate contexts to rigorously isolate these components, thereby measuring the atomic performance boundary without the noise of cascading errors. Our results highlight a clear contrast in performance: while LLMs are proficient in extracting explicit constraints, they struggle to infer implicit, open-world requirements. Furthermore, they exhibit structural biases in plan generation and suffer from ineffective self-correction, characterized by excessive sensitivity and erroneous persistence. These findings offer precise directions for improving LLM reasoning and planning abilities.

cs.AI

Programmatic Context Augmentation for LLM-based Symbolic Regression

Symbolic regression (SR), the task of discovering mathematical expressions that best describe a given dataset, remains a fundamental challenge in scientific discovery. Traditional approaches, primarily based on genetic algorithms and related evolutionary methods, have proven useful but suffer from scalability and expressivity limitations. Recently, large language model (LLM)-based evolutionary search methods have been introduced into SR and show promise. However, existing LLM-based approaches typically rely on scalar evaluation metrics, such as mean squared error, as the sole source of feedback during the search process, thereby overlooking the rich information embedded in the dataset. To address this limitation, we propose a novel LLM-based evolutionary search framework that incorporates programmatic context augmentation. By enabling code-based interactions with the dataset, our method can actively perform data analysis and extract informative signals, beyond aggregated evaluation scores. We evaluate our framework on advanced benchmarks, such as LLM-SRBench, and demonstrate superior efficiency and accuracy compared to strong baselines.

cs.AI

VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industry, remains underexplored. In this paper, we introduce \textit{VT-Bench}, the first unified benchmark for standardizing vision-tabular discriminative prediction and generative reasoning tasks. VT-Bench aggregates 14 datasets across 9 domains (medical-centric, while covering pets, media, and transportation) with over 756K samples. We evaluate 23 representative models, including unimodal experts, specialized visual-tabular models, general-purpose vision-language models (VLMs), and tool-augmented methods, highlighting substantial challenges of visual-tabular learning. We believe VT-Bench will stimulate the community to build more powerful multi-modal vision-tabular foundation models. Benchmark: https://github.com/Ziyi-Jia990/VT-Bench

cs.CV

Activation Compression in LLMs: Theoretical Analysis and Efficient Algorithm

Training large language models (LLMs) is highly memory-intensive, as training must store not only weights and optimizer states but also intermediate activations for backpropagation. While existing memory-efficient methods largely focus on gradients and optimizer states, activation compression is less well established due to the lack of LLM-tailored theory and guarantees. In this work, we develop a theoretical framework showing that activation compression is safe for linear operators when activation compression is unbiased, but problematic for nonlinear ones. We further derive gradient variance bound and establish convergence guarantees for applying activation compression to all linear operators under the standard $L$-smoothness assumption, showing that it does not change the convergence rate. Guided by the theory, we propose an activation-gradient co-compression method that reuses low-rank activation factors to compress linear-layer gradients without extra computation or additional gradient error. We conduct extensive experiments on Qwen and LLaMA models using a pretraining benchmark and multiple fine-tuning benchmarks to validate our theory and demonstrate competitive performance of our method in both accuracy and compression efficiency. We provide our code in the supplementary material for reproducibility.

cs.LG

Lifting Traces to Logic: Programmatic Skill Induction with Neuro-Symbolic Learning for Long-Horizon Agentic Tasks

Foundation model-driven agents often struggle with long-horizon planning due to the transient nature of purely prompting-based reasoning. While existing skill induction methods mitigate this by distilling experience into state-blind parameterized scripts, they fail to capture the conditional logic required for robust execution in dynamic environments. In this paper, we propose Neuro-Symbolic Skill Induction (NSI), a framework that lifts interaction traces into modular, \textit{logic-grounded} programs. By synthesizing explicit control flows and dynamic variable binding, NSI empowers agents to discover \textit{when} and \textit{why} to act. This paradigm enables the efficient generalization, allowing agents to induce skills from few-shot examples and flexibly adapt to unseen goals. Experiments on a series of agentic tasks demonstrate that NSI consistently outperforms state-of-the-art baselines, empowering agents to self-evolve into architects of logic-grounded skills.

cs.AI

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models

Spatial reasoning is a cornerstone capability for intelligent systems to perceive and interact with the physical world. However, multimodal large language models (MLLMs) frequently suffer from hallucinations and imprecision when parsing complex geometric layouts. As data-driven scaling struggles to internalize structured geometric priors and spatial constraints, integrating mature, specialized vision models presents a compelling alternative. Despite its promise, applying this paradigm to spatial reasoning is hindered by two key challenges: The difficulty of invoking heterogeneous, parameter-rich tools, as well as the challenge of understanding and effectively leveraging their diverse low-level outputs (e.g., segmentation masks, depth maps) in high-level reasoning. To address these challenges, we propose LAST, a unified framework for tool-augmented spatial reasoning. LAST features an extensible interactive sandbox, termed LAST-Box, which abstracts heterogeneous tool invocations into atomic instructions and reusable spatial skills, returning multimodal hints (e.g., annotated images and textual descriptions) that can be directly consumed by LLMs. We further design a three-stage progressive training strategy that guides models from understanding tool outputs to proficient and adaptive tool invocation. Experiments on four datasets show that LAST-7B achieves around 20\% performance gains over its backbone and outperforms strong proprietary closed-source LLMs, substantially enhancing reasoning on complex spatial tasks.

cs.CV

Thinking with Tables: Enhancing Multi-Modal Tabular Understanding via Neuro-Symbolic Reasoning

Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across modalities such as images and text. However, tabular data, despite being a critical real-world modality, remains relatively underexplored in multimodal learning. In this paper, we focus on the task of Tabular-Vision Multi-Modal Understanding (TVMU) and identify three core challenges: (1) high structural variability and data incompleteness in tables, (2) implicit and complex feature dependencies, and (3) significant heterogeneity in problem-solving pipelines across downstream tasks. To address these issues, we propose Thinking with Tables (TWT). TWT employs a program-aided code-based neuro-symbolic reasoning mechanism that facilitates key operations, such as information extraction and element modeling, by interacting with external environments. We evaluate TWT on eight representative datasets. Experimental results demonstrate that TWT consistently outperforms existing baselines by an average of 10\% in accuracy, achieving performance comparable to, or even surpassing, proprietary commercial SOTA LLMs on TVMU tasks. Models and codes are available at https://github.com/kunyang-YU/Thinking-with-Tables

cs.CL