SearcharxivSearch

arXiv subjects

Qi You

Publications and source records attributed to Qi You.

11 recordsLinked to original sources

ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives

Humanities and social science research requires close reading of long narrative materials such as novels, scripts, archives, and case reports, yet many users have limited access to costly proprietary long-context models. Compact, locally deployable language models are a practical alternative, but directly feeding them an entire long context remains costly, hard to inspect, and prone to missing sparse evidence. We present ClueWeaver, an evidence-aware dual-agent framework for long-narrative question answering with compact local models. A Finder identifies passages containing answer-critical clues through retrieval-guided segmentation, while an Interpreter derives the answer from the selected evidence, produces rationales with paragraph-ID citations, and applies an internal self-calibration pass for high-risk questions. Both agents are optimized with reward-guided reinforcement learning: Finder rewards emphasize evidence retention and faithful paragraph-ID references, and Interpreter rewards emphasize correctness, grounding, and concise explanations. This decomposition makes evidence selection and reasoning more inspectable than end-to-end prompting. Experiments across multiple long-context narrative question answering and claim verification settings show that ClueWeaver substantially improves local end-to-end language models while providing evidence coverage and paragraph-referenced reasoning traces. Code is available at https://github.com/Ameame1/ClueWeaver.

cs.CL

Ai2-Kit: Streamlining AI-Accelerated Ab Initio Workflows for Complex Chemical Systems

Molecular simulations of complex chemical systems, such as catalysis, electrochemistry, and energy storage, often need to capture the interplay of effects such as electronic structure, finite-temperature fluctuations, and electric-field response. Such complexity is difficult to address with traditional ab initio calculations, which are limited by the time and length scales they can reach. AI-accelerated ab initio (AI2) methods use machine learning potentials trained on first-principles data to replace expensive electronic-structure calculations, extending ab initio accuracy to these regimes, but their routine application requires reliable workflows that connect first-principles calculations, model training, molecular dynamics, enhanced sampling, trajectory analysis, and HPC orchestration. Here we present ai2-kit, a software toolkit for developing accessible, reproducible, and extensible AI2 workflows. ai2-kit provides high-semantic-density command-line interfaces and Python APIs for structure and dataset conversion, batch task generation, active-learning screening, job orchestration, and workflow recovery. We demonstrate ai2-kit in four representative applications: active-learning-based machine learning potential construction, free-energy perturbation for redox and acid-base processes, electrochemical machine learning potentials for electrified interfaces, and spectroscopies from machine learning molecular dynamics. ai2-kit also provides AI-agent skills that help users adapt these use cases into customized workflows for their own chemical systems and computational software stacks. Together, ai2-kit helps turn AI2 methods from bespoke computational protocols into reusable and extensible workflows for complex chemical systems, from model construction to property prediction.

physics.chem-ph

Orientifolds of Gepner models without K\"ahler moduli

One of the main challenges in string theory Calabi-Yau compactifications to four dimensions is the stabilization of the massless complex structure and K\"ahler moduli. In type IIB string theory, complex structure moduli can be stabilized perturbatively by turning on fluxes on the internal space, while there is no perturbative mechanism for K\"ahler moduli stabilization. Since every Calabi-Yau manifold has at least one K\"ahler modulus (the overall volume), there is no hope to stabilize all moduli perturbatively. A way out is given by Landau-Ginzburg/Gepner models string vacua which can have no K\"ahler moduli. To identify the most promising candidates for fully stabilized perturbative string vacua, we provide an exhaustive list of Landau-Ginzburg orientifold models with no K\"ahler moduli, and compute for each model the number of complex structure moduli together with the tadpole charge. From this, we can identify which of these models are genuine candidates for phenomenologically relevant string vacua.

hep-th

Anomaly-Free Spectra, Unimodular Lattices and 6D R-Symmetry Gauged Supergravity

We study the classification problem for anomaly-free 6D $\mathcal N=(1,0)$ supergravities with a gauged abelian R-symmetry and one tensor multiplet. We present eleven new models with gauge group $G_{\mathrm{non-Abelian}}\times U(1)_R$ that satisfy the local Green--Schwarz factorization condition, together with several recently proposed global consistency conditions. In particular, the low-rank models we found are precisely where some of the recent enumeration literature is least directly applicable. These examples suggest that the landscape of anomaly-free gauged $U(1)_R$ supergravities may be richer than previously recognized while still remaining highly constrained. We analyze the arithmetic structure of the anomaly coefficients, including their integral pairings, embeddability into rank-two unimodular charge lattices, the characteristic-vector condition and ghost-free gauge-field conditions. We show that $n_V \equiv 8 \pmod{12}$ is necessary and sufficient for the unimodular embeddability in the rank-two case, when the gauge group does not contain $SU(2)$, $SU(3)$ and $G_2$. For the characteristic-vector condition we verify sufficiency for the branches realized by our examples and identify a remaining branch requiring additional exclusion. We also present a detailed discussion of the contribution to the anomaly polynomial when the $D_4$ Lie algebra is present. These results sharpen the boundary between anomaly-free 6D spectra, global-consistency constraints, and possible UV realization in string theory or F-theory.

hep-th

A Contrastive Learning Framework Empowered by Attention-based Feature Adaptation for Street-View Image Classification

Street-view image attribute classification is a vital downstream task of image classification, enabling applications such as autonomous driving, urban analytics, and high-definition map construction. It remains computationally demanding whether training from scratch, initialising from pre-trained weights, or fine-tuning large models. Although pre-trained vision-language models such as CLIP offer rich image representations, existing adaptation or fine-tuning methods often rely on their global image embeddings, limiting their ability to capture fine-grained, localised attributes essential in complex, cluttered street scenes. To address this, we propose CLIP-MHAdapter, a variant of the current lightweight CLIP adaptation paradigm that appends a bottleneck MLP equipped with multi-head self-attention operating on patch tokens to model inter-patch dependencies. With approximately 1.4 million trainable parameters, CLIP-MHAdapter achieves superior or competitive accuracy across eight attribute classification tasks on the Global StreetScapes dataset, attaining new state-of-the-art results while maintaining low computational cost. The code is available at https://github.com/SpaceTimeLab/CLIP-MHAdapter.

cs.CV

LitVISTA: A Benchmark for Narrative Orchestration in Literary Text

Computational narrative analysis aims to capture rhythm, tension, and emotional dynamics in literary texts. Existing large language models can generate long stories but overly focus on causal coherence, neglecting the complex story arcs and orchestration inherent in human narratives. This suggests a structural misalignment between model- and human-generated narratives. We therefore position narrative analysis as a diagnostic proxy for generation and propose VISTA Space, a high-dimensional framework for narrative orchestration that unifies human and model perspectives while jointly characterizing narrative function and structure in a common space. We further introduce LitVISTA, a structurally annotated benchmark grounded in literary texts, which operationalizes VISTA Space for systematic evaluation of models' narrative orchestration capabilities. Under an oracle setting with gold event anchors, we evaluate frontier LLMs including GPT, Claude, Grok, and Gemini. Results reveal systematic deficiencies, as current models struggle to jointly capture narrative function and structure and fail to form an integrated global view of literary narrative orchestration. End-to-end analysis further shows that failures are dominated by anchor identification and localization errors. Even advanced thinking modes yield mixed and often limited gains for literary narrative understanding.

cs.CL

State Space and Self-Attention Collaborative Network with Feature Aggregation for DOA Estimation

Accurate direction-of-arrival (DOA) estimation for sound sources is challenging due to the continuous changes in acoustic characteristics across time and frequency. In such scenarios, accurate localization relies on the ability to aggregate relevant features and model temporal dependencies effectively. In time series modeling, achieving a balance between model performance and computational efficiency remains a significant challenge. To address this, we propose FA-Stateformer, a state space and self-attention collaborative network with feature aggregation. The proposed network first employs a feature aggregation module to enhance informative features across both temporal and spectral dimensions. This is followed by a lightweight Conformer architecture inspired by the squeeze-and-excitation mechanism, where the feedforward layers are compressed to reduce redundancy and parameter overhead. Additionally, a temporal shift mechanism is incorporated to expand the receptive field of convolutional layers while maintaining a compact kernel size. To further enhance sequence modeling capabilities, a bidirectional Mamba module is introduced, enabling efficient state-space-based representation of temporal dependencies in both forward and backward directions. The remaining self-attention layers are combined with the Mamba blocks, forming a collaborative modeling framework that achieves a balance between representation capacity and computational efficiency. Extensive experiments demonstrate that FA-Stateformer achieves superior performance and efficiency compared to conventional architectures.

eess.SP

Decoding the Competing Effects of Dynamic Solvation Structures on Nuclear Magnetic Resonance Chemical Shifts of Battery Electrolytes via Machine Learning

Understanding the solvation structure of electrolytes is critical for optimizing the electrochemical performance of rechargeable batteries, as it directly influences properties such as ionic conductivity, viscosity, and electrochemical stability. The highly complex structures and strong interactions in high-concentration electrolytes make accurate modeling and interpretation of their ``structure-property" relationships even more challenging with spectroscopic methods. In this study, we present a machine learning-based approach to predict dynamic $^7$Li NMR chemical shifts in LiFSI/DME electrolyte solutions. Additionally, we provide a comprehensive structural analysis to interpret the observed chemical shift behavior in our experiments, particularly the abrupt changes in $^7$Li chemical shifts at high concentrations. Using advanced modeling techniques, we quantitatively establish the relationship between molecular structure and NMR spectra, offering critical insights into solvation structure assignments. Our findings reveal the coexistence of two competing local solvation structures that shift in dominance as electrolyte concentration approaches the concentrated limit, leading to anomalous reverse of $^7$Li NMR chemical shift in our experiment. This work provides a detailed molecular-level understanding of the intricate solvation structures probed by NMR spectroscopy, leading the way for enhanced electrolyte design.

physics.chem-ph

Tadpole conjecture in non-geometric backgrounds

Calabi-Yau compactifications have typically a large number of complex structure and/or K\"ahler moduli that have to be stabilised in phenomenologically-relevant vacua. The former can in principle be done by fluxes in type IIB solutions. However, the tadpole conjecture proposes that the number of stabilised moduli can at most grow linearly with the tadpole charge of the fluxes required for stabilisation. We scrutinise this conjecture in the $2^6$ Gepner model: a non-geometric background mirror dual to a rigid Calabi-Yau manifold, in the deep interior of moduli space. By constructing an extensive set of supersymmetric Minkowski flux solutions, we spectacularly confirm the linear growth, while achieving a slightly higher ratio of stabilised moduli to flux charge than the conjectured upper bound. As a byproduct, we obtain for the first time a set of solutions within the tadpole bound where all complex structure moduli are massive. Since the $2^6$ model has no K\"ahler moduli, these show that the massless Minkowski conjecture does not hold beyond supergravity.

hep-th

High-resolution Power Doppler Using Null Subtraction Imaging

To improve the spatial resolution of power Doppler (PD) imaging, we explored null subtraction imaging (NSI) as an alternative beamforming technique to delay-and-sum (DAS). NSI is a nonlinear beamforming approach that uses three different apodizations on receive and incoherently sums the beamformed envelopes. NSI uses a null in the beam pattern to improve the lateral resolution, which we apply here for improving PD spatial resolution both with and without contrast microbubbles. In this study, we used NSI with three types of singular value decomposition (SVD)-based clutter filters and noise equalization to generate high-resolution PD images. An element sensitivity correction scheme was also proposed as a crucial component of NSI-based PD imaging. First, a microbubble trace experiment was performed to evaluate the resolution improvement of NSI-based PD over traditional DAS-based PD. Then, both contrast-enhanced and contrast free ultrasound PD images were generated from the scan of a rat brain. The cross-sectional profile of the microbubble traces and microvessels were plotted. FWHM was also estimated to provide a quantitative metric. Furthermore, iso-frequency curves were calculated to provide a resolution evaluation metric over the global field of view. Up to six-fold resolution improvement was demonstrated by the FWHM estimate and four-fold resolution improvement was demonstrated by the iso-frequency curve from the NSI-based PD microvessel images compared to microvessel images generated by traditional DAS-based beamforming. A resolvability of 39 um was measured from the NSI-based PD microvessel image. The computational cost of NSI-based PD was only increased by 40 percent over the DAS-based PD.

eess.SP

High-level synthesis design of scalable ultrafast ultrasound beamformer with single FPGA

Ultrafast ultrasound imaging is essential for advanced ultrasound imaging techniques such as ultrasound localization microscopy (ULM) and functional ultrasound (fUS). Current ultrafast ultrasound imaging is challenged by the ultrahigh data bandwidth associated with the radio frequency (RF) signal, and by the latency of the computationally expensive beamforming process. As such, continuous ultrafast data acquisition and beamforming remain elusive with existing software beamformers based on CPUs or GPUs. To address these challenges, the proposed work introduces a novel method of implementing an ultrafast ultrasound beamformer specifically for ultrafast plane wave imaging (PWI) on a field programmable gate array (FPGA) by using high-level synthesis. A parallelized implementation of the beamformer on a single FPGA was proposed by 1) utilizing a delay compression technique to reduce the delay profile size, which enables both run-time pre-calculated delay profile loading from external memory and delay reuse 2) vectorizing channel data fetching which is enabled by delay reuse, and 3) using fixed summing networks to reduce consumption of logic resources. Our proposed method presents two unique advantages over current FPGA beamformers: 1) high scalability that allows fast adaptation to different FPGA resources and beamforming speed demands by using Xilinx High-Level Synthesis as the development tool, and 2) allow a compact form factor design by using a single FPGA to complete the beamforming instead of multiple FPGAs. With the proposed method, a sustainable average beamforming rate of 4.83 G samples/second in terms of input raw RF sample was achieved. The resulting image quality of the proposed beamformer was compared with the software beamformer on the Verasonics Vantage system for both phantom imaging and in vivo imaging of a mouse brain.

eess.SP