SearcharxivSearch

arXiv subjects

Ruobing Jiang

Publications and source records attributed to Ruobing Jiang.

16 recordsLinked to original sources

PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark

Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contracts: legible labels, prescribed aspect ratios, and -- above all -- abstaining from fabricated scientific figures. We present POSTERHARNESS, an auditable harness reframing poster generation as measurable instruction-following tasks, with a pilot benchmark and failure taxonomy. POSTERHARNESS uses a placeholder-first contract to separate two jobs models otherwise conflate. The model performs visual-summary design: typography, reading path, color, and background -- but never draws data-bearing figures. Every figure region must be an empty labeled placeholder; a deterministic compositor inserts real source-paper figures at detected coordinates. This makes properties measurable: placeholder count and ID accuracy, blankness, aspect-ratio compliance, abstention from synthesized graphics, public-text hygiene, and source-figure provenance -- with failures logged as explicit rejections, not hidden in plausible-looking output. We instantiate the harness on 12 papers (6 HEP, 6 AI/ML-adjacent) and report three findings. (i) A counterfactual probe shows the placeholder contract drives VLM-counted synthesized figures from 34 to 0 across three papers. (ii) A failure taxonomy identifies blocking contracts: placeholder geometry, placeholder QA, template critic, and public text. (iii) Comparison with Paper2Poster shows a trade-off: PosterHarness yields higher-resolution artifacts, lower white-canvas fraction, and stronger VLM visual preference; the deterministic baseline retains slightly more PosterQuiz-style information and runs faster. We report this as regime characterization, not a superiority claim. All artifacts, prompts, manifests, and audit scripts are released as a reusable evaluation component.

cs.CV

A Reproducible Benchmark and Evidence-Retrieval Software Framework for Silicon Detector R&D Literature

Silicon pixel detector R&D depends on a large and rapidly growing technical literature, including beam-test and irradiation studies, performance measurements, simulation, and design reports. Locating the supporting evidence passage for a measurement, operating condition, or design decision is therefore a computing and data-science challenge for detector-development workflows. General-purpose language models are insufficient unless grounded in traceable primary sources, particularly in a domain with specialised terminology, configuration-dependent measurements, and rapidly evolving experimental results. We address this with a reproducible, general-purpose framework for evidence-grounded retrieval over technical literature, using silicon pixel detector R&D as a demanding validation domain. The framework combines sparse lexical retrieval, dense semantic retrieval, and hybrid reciprocal-rank fusion, with an optional graph-guided exploration layer and grounded, abstention-aware response generation. The accompanying benchmark provides manually curated chunk-level evidence annotations, source-level diagnostics, semantic relevance checks, and negative-query abstention tests over two detector query sets. We evaluate six retrieval configurations across 378 source documents and 8,442 indexed chunks. Hybrid sparse-dense retrieval gives the strongest strict evidence recovery, achieving Hit@5 of 0.917 on the core benchmark and 0.951 on the curated extension benchmark, while graph-based methods are more effective for literature exploration and source discovery. Graph expansion is therefore best employed as a discovery layer over the hybrid retrieval backbone. The framework provides reusable software for traceable, 1 evidence-grounded knowledge access in silicon detector R&D and high-energy physics instrumentation.

physics.ins-det

Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis

Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a rapidly expanding and heterogeneous body of scientific literature. As high-energy physics (HEP) increasingly explores agent-assisted analysis workflows, efficiently locating, integrating, and verifying scientific evidence becomes an essential capability. While retrieval-augmented generation (RAG) offers a promising framework for scientific question answering, integrating agentic reasoning without compromising retrieval precision remains a key challenge. In this work, we present agentic hybrid RAG, an evidence-grounded RAG framework for muon collider research. The framework combines a hybrid retriever, integrating sparse lexical and dense semantic retrieval, with an agentic reasoning module for query decomposition, evidence expansion, and grounded answer generation. To enable systematic evaluation, we construct the first benchmark for retrieval-augmented scientific question answering in the muon collider domain, comprising a curated literature corpus together with dedicated retrieval and answer-generation benchmarks covering major detector and physics research topics. Extensive evaluation shows that hybrid retrieval provides the strongest retrieval backbone, while agentic reasoning is most effective for controlled evidence expansion and answer synthesis. Built on this principle, agentic hybrid RAG consistently outperforms representative retrieval and RAG baselines in retrieval effectiveness, answer quality, evidence coverage, and factual grounding. Together, the benchmark and framework provide a foundation for evidence-grounded scientific question answering and future HEP analysis agents operating over large-scale scientific literature.

hep-ex

Weighted Graph Clustering via Scale Contraction and Graph Structure Learning

Graph clustering aims to partition nodes into distinct clusters based on their similarity, thereby revealing relationships among nodes. Nevertheless, most existing methods do not fully utilize these edge weights. Leveraging edge weights in graph clustering tasks faces two critical challenges. (1) The introduction of edge weights may significantly increase storage space and training time, making it essential to reduce the graph scale while preserving nodes that are beneficial for the clustering task. (2) Edge weight information may inherently contain noise that negatively impacts clustering results. However, few studies can jointly optimize clustering and edge weights, which is crucial for mitigating the negative impact of noisy edges on clustering task. To address these challenges, we propose a contractile edge-weight-aware graph clustering network. Specifically, a cluster-oriented graph contraction module is designed to reduce the graph scale while preserving important nodes. An edge-weight-aware attention network is designed to identify and weaken noisy connections. In this way, we can more easily identify and mitigate the impact of noisy edges during the clustering process, thus enhancing clustering effectiveness. We conducted extensive experiments on three real-world weighted graph datasets. In particular, our model outperforms the best baseline, demonstrating its superior performance. Furthermore, experiments also show that the proposed graph contraction module can significantly reduce training time and storage space.

cs.LG

Prospects for measuring electroweak production of $Z\gamma\gamma$ and 2 jets at the LHC

Vector boson scattering (VBS) serves as a powerful channel for probing the Standard Model, particularly the electroweak symmetry breaking mechanism. Currently, studies of VBS mainly focus on $2 \to 2$ scattering. In this study, we investigate the $2\to 3$ VBS process of $\text{p p} \to Z\gamma\gamma + 2~\text{jets}$ through Monte Carlo simulations, including signal generation and background analysis. The signal significance is evaluated across different phase-space regions. With an integrated luminosity of 500 fb$^{-1}$, the signal significance can reach about 4.5 $\sigma$. These results suggest that, given the ongoing release of the LHC Run 3 dataset, there will be promising opportunities to explore and potentially discover a series of $2\to 3$ VBS processes, starting with the $Z\gamma\gamma$ channel.

hep-ph

Hierarchy-Consistent Learning and Adaptive Loss Balancing for Hierarchical Multi-Label Classification

Hierarchical Multi-Label Classification (HMC) faces critical challenges in maintaining structural consistency and balancing loss weighting in Multi-Task Learning (MTL). In order to address these issues, we propose a classifier called HCAL based on MTL integrated with prototype contrastive learning and adaptive task-weighting mechanisms. The most significant advantage of our classifier is semantic consistency including both prototype with explicitly modeling label and feature aggregation from child classes to parent classes. The other important advantage is an adaptive loss-weighting mechanism that dynamically allocates optimization resources by monitoring task-specific convergence rates. It effectively resolves the "one-strong-many-weak" optimization bias inherent in traditional MTL approaches. To further enhance robustness, a prototype perturbation mechanism is formulated by injecting controlled noise into prototype to expand decision boundaries. Additionally, we formalize a quantitative metric called Hierarchical Violation Rate (HVR) as to evaluate hierarchical consistency and generalization. Extensive experiments across three datasets demonstrate both the higher classification accuracy and reduced hierarchical violation rate of the proposed classifier over baseline models.

cs.LG

Exploring the Tradeoff Between Diversity and Discrimination for Continuous Category Discovery

Continuous category discovery (CCD) aims to automatically discover novel categories in continuously arriving unlabeled data. This is a challenging problem considering that there is no number of categories and labels in the newly arrived data, while also needing to mitigate catastrophic forgetting. Most CCD methods cannot handle the contradiction between novel class discovery and classification well. They are also prone to accumulate errors in the process of gradually discovering novel classes. Moreover, most of them use knowledge distillation and data replay to prevent forgetting, occupying more storage space. To address these limitations, we propose Independence-based Diversity and Orthogonality-based Discrimination (IDOD). IDOD mainly includes independent enrichment of diversity module, joint discovery of novelty module, and continuous increment by orthogonality module. In independent enrichment, the backbone is trained separately using contrastive loss to avoid it focusing only on features for classification. Joint discovery transforms multi-stage novel class discovery into single-stage, reducing error accumulation impact. Continuous increment by orthogonality module generates mutually orthogonal prototypes for classification and prevents forgetting with lower space overhead via representative representation replay. Experimental results show that on challenging fine-grained datasets, our method outperforms the state-of-the-art methods.

cs.CV

Incorporating Attributes and Multi-Scale Structures for Heterogeneous Graph Contrastive Learning

Heterogeneous graphs (HGs) are composed of multiple types of nodes and edges, making it more effective in capturing the complex relational structures inherent in the real world. However, in real-world scenarios, labeled data is often difficult to obtain, which limits the applicability of semi-supervised approaches. Self-supervised learning aims to enable models to automatically learn useful features from data, effectively addressing the challenge of limited labeling data. In this paper, we propose a novel contrastive learning framework for heterogeneous graphs (ASHGCL), which incorporates three distinct views, each focusing on node attributes, high-order and low-order structural information, respectively, to effectively capture attribute information, high-order structures, and low-order structures for node representation learning. Furthermore, we introduce an attribute-enhanced positive sample selection strategy that combines both structural information and attribute information, effectively addressing the issue of sampling bias. Extensive experiments on four real-world datasets show that ASHGCL outperforms state-of-the-art unsupervised baselines and even surpasses some supervised benchmarks.

cs.LG

Weighted Graph Structure Learning with Attention Denoising for Node Classification

Node classification in graphs aims to predict the categories of unlabeled nodes by utilizing a small set of labeled nodes. However, weighted graphs often contain noisy edges and anomalous edge weights, which can distort fine-grained relationships between nodes and hinder accurate classification. We propose the Edge Weight-aware Graph Structure Learning (EWGSL) method, which combines weight learning and graph structure learning to address these issues. EWGSL improves node classification by redefining attention coefficients in graph attention networks to incorporate node features and edge weights. It also applies graph structure learning to sparsify attention coefficients and uses a modified InfoNCE loss function to enhance performance by adapting to denoised graph weights. Extensive experimental results show that EWGSL has an average Micro-F1 improvement of 17.8% compared with the best baseline.

cs.LG

Testing Bell inequalities and probing quantum entanglement at CEPC

We study quantum entanglement and test violation of Bell-type inequality at the Circular Electron Positron Collider (CEPC), which is one of the most attractive future colliders. It's a promising particle collider designed to search new physics, make Standard Model (SM) precision measurements, and serving as a Higgs factory. Our study is based on a fast simulation of the $Z$ boson pair production from Higgs boson decay at $\sqrt{s} = 250$ GeV. The detector effects are also included in the simulation. The spin density matrix of the joint $ZZ$ system is parametrized using irreducible tensor operators and reconstructed from the spherical coordinates of the decay leptons. To test Bell inequalities, we construct observable quantities for the $H \to ZZ*$ process in CEPC by using the (Collins-Gisin-Linden-Massar-Popescu) CGLMP inequality, whose value is determined from the density matrix of the Z boson pairs. The sensitivity of the Bell inequality violation is observed with more than 1$\sigma$ and the presence of the quantum entanglement is probed with more than 2$\sigma$ confidence level.

hep-ph

Searches for multi-Z boson productions and anomalous gauge boson couplings at a muon collider

Multi-boson productions can be exploited as novel probes either for standard model precision tests or new physics searches, and have become one of those popular topics in the ongoing LHC experiments, and in future collider studies, including those for electron-positron and muon-muon colliders. Here we focus on two examples, i.e., ZZZ direct productions through $\mu^{+}\mu^{-}$ annihilation at a 1 TeV muon collider, and ZZ productions through vector boson scattering at a 10 TeV muon collider, with an integrated luminosity of $10 \, \text{ab}^{-1}$. Various channels are considered, including, such as $ZZZ \rightarrow 4l2\nu$ and $ZZZ \rightarrow 4l + 2 \text{ jets}$, etc. Expected significance on these multi-Z boson production processes are provided based on a detailed Monte Carlo study and signal background analysis. Sensitives on anomalous gauge boson couplings are also presented.

hep-ex

Incorporating Higher-order Structural Information for Graph Clustering

Clustering holds profound significance in data mining. In recent years, graph convolutional network (GCN) has emerged as a powerful tool for deep clustering, integrating both graph structural information and node attributes. However, most existing methods ignore the higher-order structural information of the graph. Evidently, nodes within the same cluster can establish distant connections. Besides, recent deep clustering methods usually apply a self-supervised module to monitor the training process of their model, focusing solely on node attributes without paying attention to graph structure. In this paper, we propose a novel graph clustering network to make full use of graph structural information. To capture the higher-order structural information, we design a graph mutual infomax module, effectively maximizing mutual information between graph-level and node-level representations, and employ a trinary self-supervised module that includes modularity as a structural constraint. Our proposed model outperforms many state-of-the-art methods on various datasets, demonstrating its superiority.

cs.LG

A proposed PKU-Muon experiment for muon tomography and dark matter search

We propose here a set of new methods to directly detect light mass dark matter through its scattering with abundant atmospheric muons or accelerator beams. Firstly, we plan to use the free cosmic-ray muons interacting with dark matter in a volume surrounded by tracking detectors, to trace possible interaction between dark matter and muons. Secondly, we will interface our device with domestic or international muon beams. Due to much larger muon intensity and focused beam, we anticipate the detector can be made further compact and the resulting sensitivity on dark matter searches will be improved. Furthermore, we will measure precisely directional distributions of cosmic-ray muons, either at mountain or sea level, and the differences may reveal possible information of dark matter distributed near the earth. Specifically, our methods can have advantages over `exotic' dark matters which are either muon-philic or slowed down due to some mechanism, and sensitivity on dark matter and muon scattering cross section can reach as low as microbarn level.

hep-ex

Modeling Multi-aspect Preferences and Intents for Multi-behavioral Sequential Recommendation

Multi-behavioral sequential recommendation has recently attracted increasing attention. However, existing methods suffer from two major limitations. Firstly, user preferences and intents can be described in fine-grained detail from multiple perspectives; yet, these methods fail to capture their multi-aspect nature. Secondly, user behaviors may contain noises, and most existing methods could not effectively deal with noises. In this paper, we present an attentive recurrent model with multiple projections to capture Multi-Aspect preferences and INTents (MAINT in short). To extract multi-aspect preferences from target behaviors, we propose a multi-aspect projection mechanism for generating multiple preference representations from multiple aspects. To extract multi-aspect intents from multi-typed behaviors, we propose a behavior-enhanced LSTM and a multi-aspect refinement attention mechanism. The attention mechanism can filter out noises and generate multiple intent representations from different aspects. To adaptively fuse user preferences and intents, we propose a multi-aspect gated fusion mechanism. Extensive experiments conducted on real-world datasets have demonstrated the effectiveness of our model.

cs.IR

Searching for Majorana Neutrinos at a Same-Sign Muon Collider

Majorana properties of neutrinos have long been a focus in the pursuit of possible new physics beyond the standard model, which has motivated lots of dedicated theoretical and experimental studies. A future same-sign muon collider is an ideal platform to search for Majorana neutrinos through the Lepton Number Violation process. Specifically, this t-channel kind of process is less kinematically suppressed and has a good advantage in probing Majorana neutrinos at high mass regions up to 10 TeV. In this paper, we perform a detailed fast Monte Carlo simulation study through examining three different final states: 1) pure-leptonic state with electrons or muons, 2) semi-leptonic state, and 3) pure-hadronic state in the resolved or merged categories. Furthermore, we perform a full simulation study on the pure-leptonic final state to validate our fast simulation results.

hep-ph

Enhancing Marine Data Transmission with Socially-Aware Resilient Vessel Networks

With the multi-dimensional exploration towards oceans, enormous sensing data has been generated with significant volume, velocity, variety and heterogeneity. The resulted Big Marine Data (BMD) thus issue unprecedented architectural challenges on existing marine communication systems. Current dominant marine communication technologies, e.g., shore-based cellular stations, high frequency radio, and expensive satellites, extremely suffer from short coverage, low bandwidth, insecurity, and unavailable cross-domain transmission. In this paper, Resilient Vessel Network (RVN) is proposed to fundamentally enhance BMD transmission. RVNs with widespread self-organized vessels and opportunistic connections reveal advantages of ubiquity, resilience, low cost and cross-domain transmission. To efficiently manage opportunistic vessel-to-vessel (V2V) connections for optimal routing, Social Network Analysis (SNA) on historical vessel interactions is applied for vessel familiarity measurement and community detection. The performance of the proposed community-based routing (CBR) is comprehensively evaluated with real datasets of fishing vessel trajectories. It is demonstrated that CBR achieves much lower transmission cost with comparable delivery ratio compared to typical routing algorithms.

cs.NI