SearcharxivSearch

arXiv subjects

Dawei Fu

Publications and source records attributed to Dawei Fu.

9 recordsLinked to original sources

SE-GoS: Self-Evolving Graph-of-Skills for Skill Library at Scale

Modern LLM agents increasingly rely on reusable skills, yet as skill libraries scale to thousands of entries, effective retrieval becomes a bottleneck. Graph-of-Skills (GoS) addresses this challenge by exploiting dependency-aware graph structure for scalable skill retrieval, while SkillDAG further demonstrates that skill graphs can accumulate execution-backed structure online. However, these approaches leave open whether historical execution traces can be systematically distilled into a better retrieval graph that generalizes to unseen tasks. We present Self-Evolving Graph-of-Skills (SE-GoS), a training-free framework that evolves an existing GoS graph from execution traces while preserving the original retrieval pipeline. SE-GoS performs three complementary updates: topology evolution that discovers and prunes skill relationships from execution evidence, edge-weight evolution that reinforces retrieval-relevant relationships based on historical effectiveness, and description evolution that optimizes retrieval-facing skill descriptions using execution feedback. Across three LLMs on SkillsBench, SE-GoS consistently improves task reward while reducing input tokens relative to full skill loading, with gains varying across model families. In a representative setting, one evolution round improves reward from 52.4\% to 59.4\% while reducing input tokens by approximately one-third relative to full skill loading, and the resulting graph transfers to a disjoint held-out split with a 5.4-point improvement over the static GoS baseline. These results show that skill graphs can be improved from execution experience without model training, changes to the retrieval algorithm, or modifications to skill content, turning a static retrieval graph into an evolving retrieval infrastructure.

cs.AI

PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark

Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contracts: legible labels, prescribed aspect ratios, and -- above all -- abstaining from fabricated scientific figures. We present POSTERHARNESS, an auditable harness reframing poster generation as measurable instruction-following tasks, with a pilot benchmark and failure taxonomy. POSTERHARNESS uses a placeholder-first contract to separate two jobs models otherwise conflate. The model performs visual-summary design: typography, reading path, color, and background -- but never draws data-bearing figures. Every figure region must be an empty labeled placeholder; a deterministic compositor inserts real source-paper figures at detected coordinates. This makes properties measurable: placeholder count and ID accuracy, blankness, aspect-ratio compliance, abstention from synthesized graphics, public-text hygiene, and source-figure provenance -- with failures logged as explicit rejections, not hidden in plausible-looking output. We instantiate the harness on 12 papers (6 HEP, 6 AI/ML-adjacent) and report three findings. (i) A counterfactual probe shows the placeholder contract drives VLM-counted synthesized figures from 34 to 0 across three papers. (ii) A failure taxonomy identifies blocking contracts: placeholder geometry, placeholder QA, template critic, and public text. (iii) Comparison with Paper2Poster shows a trade-off: PosterHarness yields higher-resolution artifacts, lower white-canvas fraction, and stronger VLM visual preference; the deterministic baseline retains slightly more PosterQuiz-style information and runs faster. We report this as regime characterization, not a superiority claim. All artifacts, prompts, manifests, and audit scripts are released as a reusable evaluation component.

cs.CV

A Reproducible Benchmark and Evidence-Retrieval Software Framework for Silicon Detector R&D Literature

Silicon pixel detector R&D depends on a large and rapidly growing technical literature, including beam-test and irradiation studies, performance measurements, simulation, and design reports. Locating the supporting evidence passage for a measurement, operating condition, or design decision is therefore a computing and data-science challenge for detector-development workflows. General-purpose language models are insufficient unless grounded in traceable primary sources, particularly in a domain with specialised terminology, configuration-dependent measurements, and rapidly evolving experimental results. We address this with a reproducible, general-purpose framework for evidence-grounded retrieval over technical literature, using silicon pixel detector R&D as a demanding validation domain. The framework combines sparse lexical retrieval, dense semantic retrieval, and hybrid reciprocal-rank fusion, with an optional graph-guided exploration layer and grounded, abstention-aware response generation. The accompanying benchmark provides manually curated chunk-level evidence annotations, source-level diagnostics, semantic relevance checks, and negative-query abstention tests over two detector query sets. We evaluate six retrieval configurations across 378 source documents and 8,442 indexed chunks. Hybrid sparse-dense retrieval gives the strongest strict evidence recovery, achieving Hit@5 of 0.917 on the core benchmark and 0.951 on the curated extension benchmark, while graph-based methods are more effective for literature exploration and source discovery. Graph expansion is therefore best employed as a discovery layer over the hybrid retrieval backbone. The framework provides reusable software for traceable, 1 evidence-grounded knowledge access in silicon detector R&D and high-energy physics instrumentation.

physics.ins-det

Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis

Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a rapidly expanding and heterogeneous body of scientific literature. As high-energy physics (HEP) increasingly explores agent-assisted analysis workflows, efficiently locating, integrating, and verifying scientific evidence becomes an essential capability. While retrieval-augmented generation (RAG) offers a promising framework for scientific question answering, integrating agentic reasoning without compromising retrieval precision remains a key challenge. In this work, we present agentic hybrid RAG, an evidence-grounded RAG framework for muon collider research. The framework combines a hybrid retriever, integrating sparse lexical and dense semantic retrieval, with an agentic reasoning module for query decomposition, evidence expansion, and grounded answer generation. To enable systematic evaluation, we construct the first benchmark for retrieval-augmented scientific question answering in the muon collider domain, comprising a curated literature corpus together with dedicated retrieval and answer-generation benchmarks covering major detector and physics research topics. Extensive evaluation shows that hybrid retrieval provides the strongest retrieval backbone, while agentic reasoning is most effective for controlled evidence expansion and answer synthesis. Built on this principle, agentic hybrid RAG consistently outperforms representative retrieval and RAG baselines in retrieval effectiveness, answer quality, evidence coverage, and factual grounding. Together, the benchmark and framework provide a foundation for evidence-grounded scientific question answering and future HEP analysis agents operating over large-scale scientific literature.

hep-ex

Novel $|V_{cb}|$ extraction method via boosted $bc$-tagging with in-situ calibration

We present a novel method for measuring $|V_{cb}|$ at the LHC using an advanced boosted-jet tagger to identify "$bc$ signatures". By associating boosted $W \rightarrow bc$ signals with $bc$-matched jets from top-quark decays, we enable an in-situ calibration of the tagger. This approach significantly suppresses backgrounds while reducing uncertainties in flavor tagging efficiencies -- key to improving measurement precision. Our study is enabled by the development of realistic, AI-based large- and small-radius taggers, Sophon and the newly introduced SophonAK4, validated to match ATLAS and CMS's state-of-the-art taggers. The method complements the conventional small radius jet approach and enables a ~30% improvement in $|V_{cb}|$ precision under HL-LHC projections. As a byproduct, it enhances $H^{\pm} \rightarrow bc$ search sensitivity by a factor of 2--5 over the recent ATLAS result based on Run 2 data. Our work offers a new perspective for the precision $|V_{cb}|$ measurement and highlights the potential of using advanced tagging models to probe unexplored boosted regimes at the LHC.

hep-ph

Accelerating Resonance Searches via Signature-Oriented Pre-training

The search for heavy resonances beyond the Standard Model (BSM) is a key objective at the LHC. While the recent use of advanced deep neural networks for boosted-jet tagging significantly enhances the sensitivity of dedicated searches, it is limited to specific final states, leaving vast potential BSM phase space underexplored. We introduce a novel experimental method, Signature-Oriented Pre-training for Heavy-resonance ObservatioN (Sophon), which leverages deep learning to cover an extensive number of boosted final states. Pre-trained on the comprehensive JetClass-II dataset, the Sophon model learns intricate jet signatures, ensuring the optimal constructions of various jet tagging discriminates and enabling high-performance transfer learning capabilities. We show that the method can not only push widespread model-specific searches to their sensitivity frontier, but also greatly improve model-agnostic approaches, accelerating LHC resonance searches in a broad sense.

hep-ph

Knowledge-enhanced Memory Model for Emotional Support Conversation

The prevalence of mental disorders has become a significant issue, leading to the increased focus on Emotional Support Conversation as an effective supplement for mental health support. Existing methods have achieved compelling results, however, they still face three challenges: 1) variability of emotions, 2) practicality of the response, and 3) intricate strategy modeling. To address these challenges, we propose a novel knowledge-enhanced Memory mODEl for emotional suppoRt coNversation (MODERN). Specifically, we first devise a knowledge-enriched dialogue context encoding to perceive the dynamic emotion change of different periods of the conversation for coherent user state modeling and select context-related concepts from ConceptNet for practical response generation. Thereafter, we implement a novel memory-enhanced strategy modeling module to model the semantic patterns behind the strategy categories. Extensive experiments on a widely used large-scale dataset verify the superiority of our model over cutting-edge baselines.

cs.CL

Muon Beam for Neutrino CP Violation: connecting energy and neutrino frontiers

We propose here a proposal to connect neutrino and energy frontiers, by exploiting collimated muon beams for neutrino oscillations, which generate symmetric neutrino and antineutrino sources: $\mu^+\rightarrow e^+\,\bar{\nu}_{\mu}\, \nu_{e}$ and $\mu^-\rightarrow e^-\, \nu_{\mu} \,\bar{\nu}_{e}$. Interfacing with long baseline neutrino detectors such as DUNE and T2K, this experiment can be applicable to measure tau neutrino properties, and also to probe neutrino CP phase, by measuring muon electron (anti-)neutrino mixing or tau (anti-)neutrino appearance, and differences between neutrino and antineutrino rates. There are several significant benefits leading to large neutrino flux and high sensitivity on CP phase, including 1) collimated and manipulable muon beams, which lead to a larger acceptance of neutrino sources in the far detector side; 2) symmetric $\mu^+$ and $\mu^-$ beams, and thus symmetric neutrino and antineutrino sources, which make this proposal ideally useful for measuring neutrino CP violation. More importantly, $\bar{\nu}_{e,\mu}\rightarrow\bar{\nu}_\tau$ and $\nu_{e,\mu}\rightarrow \nu_\tau$, and, $\bar{\nu}_{e}\rightarrow\bar{\nu}_\mu$ and $\nu_{e}\rightarrow \nu_\mu$ oscillation signals can be collected simultaneously, with no needs for separate specific runs for neutrinos or antineutrinos. Based on a simulation of neutrino oscillation experiment, we estimate $10^4$ tau (anti-) neutrinos can be collected within 5 years which makes this proposal suitable for a brighter tau neutrino factory. Moreover, more than 7 standard deviations of sensitivity can be reached for $\dcp = |\pi/2|$, within only five ears of data taking, by combining tau and muon (anti-) neutrino appearances. With the development of a more intensive muon beam targeting future muon collider, the neutrino potential of the current proposal will surely be further improved.

hep-ph

New methods to achieve meson, muon and gamma light sources through asymmetric electron positron collisions

We propose methods to produce energetic meson beams such as charged and neutral Kaons, which are boosted to be collimated and with relatively long life time. The first type of methods is based on asymmetric electron positron collisions with a center of mass energy of, e.g., 1020 MeV, and Kaons can be produced at a rate of $10^{4-5}/s$. The electron and positron beams are either asymmetric in energy, e.g., 10 GeV electron beam with 26 MeV positron beam, or asymmetric in space, e.g., 10 GeV electron and positron beams collisions separated with a angle around 0.05 radius. Such proposals should be able to be achieved with a reasonable budget. The other type of method is relying on TeV positron on target experiment, where Kaon beams can be achieved at around $10^{7}$ per bunch crossing. Such Kaon beams are clean with small contamination, and can have great physics potential on, e.g., hyperon searches through Kaon nuclei collision, Kaon rare decay measurement, and Kaon proton or Kaon lepton collisions. The same technique with very asymmetric electron positron collisions can also be extended to other final states such as pions and tau leptons.

hep-ph