SearcharxivSearch

arXiv subjects

Zepeng Li

Publications and source records attributed to Zepeng Li.

17 recordsLinked to original sources

GateDiffInt: Gate-Mediated Controllable Diffusion and Multi-Intent LLM Distillation for User Behavior Modeling

Existing ranking models encode intent only implicitly, making it hard to disentangle structured intents of varying strength and temporal scale. Noise and intent in behavior sequences are mutually reinforcing---we call this Noise--Intent Coupling (NIC). Noise dilutes true intents, while the lack of structured intent priors leaves denoising without a clear target. To address NIC, we propose GateDiffInt, an intent interaction framework for industrial ranking. It uses the final conversion signal to jointly align sequence denoising and intent extraction. GateDiffInt applies a controllable forward diffusion process with dual gating to enhance and denoise behavior sequences. A large language model then acts as teacher to distill four structured intents---long-term, short-term, latent, and conversion into a lightweight student model. The enhanced sequence and structured intent representations are deeply fused via attention to produce intent-aware representations for conversion-rate prediction. Extensive experiments on public and large-scale industrial datasets show consistent gains over strong baselines. In online A/B tests serving hundreds of millions of daily active users, GateDiffInt delivers substantial GMV improvements and has been deployed to primary traffic, confirming both effectiveness and production readiness.

cs.IR

Differentiable Routability-Driven Package Floorplanning with Pin Assignment

As advanced packaging technology evolves, increasing interconnect density in redistribution layers (RDLs) makes routability critical to package floorplanning. Meanwhile, power integrity requirements often reserve fan-in regions for the power delivery network (PDN), forcing signal nets through fan-out regions and complicating routability estimation. Existing uniform grid-based congestion models cannot accurately characterize fan-out congestion, while previous pin assignment methods struggle to evaluate net crossings. We propose a differentiable routability-driven floorplanning and pin assignment algorithm for advanced packaging with fan-out routing. First, a differentiable wirelength minimization method directly models discrete chip orientations and back-propagates wirelength gradients to chip locations and orientations. It reduces wirelength under fixed pin selection while avoiding the bias of continuous-angle modeling. Second, a crossing-aware pin assignment method incorporates net-crossing cost into a multi-strategy DPSO algorithm and uses GPU-parallel cost evaluation to reduce wirelength efficiently. Finally, a differentiable routability maximization method constructs a congestion model tailored to fan-out routing and establishes a back-propagation path from congestion information to chip locations, thereby guiding routability optimization. Experimental results show that our method achieves 100% routability on all benchmarks. For cases successfully routed by the baselines, it reduces wirelength by up to approximately 23% compared with a leading floorplanning method equipped with our pin assignment flow.

cs.AR

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Large Language Models (LLMs) show promise for code compilation tasks, but applying them to runtime performance tuning is difficult due to complex microarchitectural effects and noisy runtime measurements. We present AutoPass, a multi-agent framework for compiler performance tuning that uses compiler and runtime evidence to guide LLM-generated optimization decisions. Rather than treating the compiler as a black box like prior auto-tuning schemes, AutoPass opens up the compiler to the LLM, enabling it to query compiler-internal optimization states and analyze the intermediate representation to orchestrate compiler options. The search process iteratively refines optimization configurations using measured runtime feedback to diagnose regressions and guide latency-improving edits. AutoPass operates in an inference-only, training-free setting and requires no offline training or task-specific fine-tuning, making it readily applicable to new benchmarks and platforms. We implement AutoPass on the LLVM compiler and evaluate it on server-grade x86-64 and embedded ARM64 systems. AutoPass outperforms expert-tuned heuristics and classical autotuning methods, achieving geometric-mean speedups of 1.043x and 1.117x over LLVM -O3 on x86-64 and ARM64, respectively.

cs.SE

Towards Clinical Practice in CT-Based Pulmonary Disease Screening: An Efficient and Reliable Framework

Deep learning models for pulmonary disease screening from Computed Tomography (CT) scans promise to alleviate the immense workload on radiologists. Still, their high computational cost, stemming from processing entire 3D volumes, remains a major barrier to widespread clinical adoption. Current sub-sampling techniques often compromise diagnostic integrity by introducing artifacts or discarding critical information. To overcome these limitations, we propose an Efficient and Reliable Framework (ERF) that fundamentally improves the practicality of automated CT analysis. Our framework introduces two core innovations: (1) A Cluster-based Sub-Sampling (CSS) method that efficiently selects a compact yet comprehensive subset of CT slices by optimizing for both representativeness and diversity. By integrating an efficient k-nearest neighbor search with an iterative refinement process, CSS bypasses the computational bottlenecks of previous methods while preserving vital diagnostic features. (2) An Ambiguity-aware Uncertainty Quantification (AUQ) mechanism, which enhances reliability by specifically targeting data ambiguity arising from subtle lesions and artifacts. Unlike standard uncertainty measures, AUQ leverages the predictive discrepancy between auxiliary classifiers to construct a specialized ambiguity score. By maximizing this discrepancy during training, the system effectively flags ambiguous samples where the model lacks confidence due to visual noise or intricate pathologies. Validated on two public datasets with 2,654 CT volumes across diagnostic tasks for 3 pulmonary diseases, ERF achieves diagnostic performance comparable to the full-volume analysis (over 90% accuracy and recall) while reducing processing time by more than 60%. This work represents a significant step towards deploying fast, accurate, and trustworthy AI-powered screening tools in time-sensitive clinical settings.

eess.IV

Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML), silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.

physics.ins-det

ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection

Zero-shot anomaly detection (ZSAD) aims to detect anomalies without any target domain training samples, relying solely on external auxiliary data. Existing CLIP-based methods attempt to activate the model's ZSAD potential via handcrafted or static learnable prompts. The former incur high engineering costs and limited semantic coverage, whereas the latter apply identical descriptions across diverse anomaly types, thus fail to adapt to complex variations. Furthermore, since CLIP is originally pretrained on large-scale classification tasks, its anomaly segmentation quality is highly sensitive to the exact wording of class names, severely constraining prompting strategies that depend on class labels. To address these challenges, we introduce ViP$^{2}$-CLIP. The key insight of ViP$^{2}$-CLIP is a Visual-Perception Prompting (ViP-Prompt) mechanism, which fuses global and multi-scale local visual context to adaptively generate fine-grained textual prompts, eliminating manual templates and class-name priors. This design enables our model to focus on precise abnormal regions, making it particularly valuable when category labels are ambiguous or privacy-constrained. Extensive experiments on 15 industrial and medical benchmarks demonstrate that ViP$^{2}$-CLIP achieves state-of-the-art performance and robust cross-domain generalization.

cs.CV

Fluorescence time profile measurement of LAB based liquid scintillator in response to medium relativistic ion particles

Liquid scintillator is widely used in particle physics experiments due to its high light yield, good timing resolution, scalability and low cost. Certain liquid scintillators exhibit pulse shape discrimination capabilities because of difference in fluorescence timing properties induced by different particles. Its fluoresence timing properties have been measured mostly for radioactive decay sources at MeV energies. We present a novel measurement of fluorescence time properties of LAB based liquid scintillator in response to high-energy ions of hydrogen (Z = 1), helium (Z = 2) and Krypton at around 200-300 MeV/u for the first time. We compared the results to those from radioactive sources and observed a distinct $dE/dX$ dependence, regardless of the particle type. These findings are essential for physics searches such as the diffuse supernova neutrino background in large liquid scintillator detectors like JUNO, and are also critical towards understanding the underlying scintillation timing mechanism.

physics.ins-det

MM-DADM: Multimodal Drug-Aware Diffusion Model for Virtual Clinical Trials

High failure rates in cardiac drug development necessitate virtual clinical trials via electrocardiogram (ECG) generation to reduce risks and costs. However, existing ECG generation models struggle to balance morphological realism with pathological flexibility, fail to disentangle demographics from genuine drug effects, and are severely bottlenecked by early-phase data scarcity. To overcome these hurdles, we propose the Multimodal Drug-Aware Diffusion Model (MM-DADM), the first generative framework for generating individualized drug-induced ECGs. Specifically, our proposed MM-DADM integrates a Dynamic Cross-Attention (DCA) module that adaptively fuses External Physical Knowledge (EPK) to preserve morphological realism while avoiding the suppression of complex pathological nuances. To resolve feature entanglement, a Causal Feature Encoder (CFE) actively filters out demographic noise to extract pure pharmacological representations. These representations subsequently guide a Causal-Disentangled ControlNet (CDC-Net), which leverages counterfactual data augmentation to explicitly learn intrinsic pharmacological mechanisms despite limited clinical data. Extensive experiments on $9,443$ ECGs across $8$ drug regimens demonstrate that MM-DADM outperforms $10$ state-of-the-art ECG generation models, improving simulation accuracy by at least $6.13\%$ and recall by $5.89\%$, while providing highly effective data augmentation for downstream classification tasks.

cs.LG

Ion manipulation from liquid Xe to vacuum: Ba-tagging for a nEXO upgrade and future $0 νββ$ experiments

Neutrinoless double beta decay {($0νββ$)} provides a way to probe physics beyond the Standard Model of particle physics. The upcoming nEXO experiment will search for $0νββ$ decay in $^{136}$Xe with a projected half-life sensitivity exceeding $10^{28}$ years at the 90\% confidence level using a liquid xenon (LXe) Time Projection Chamber (TPC) filled with 5 tonnes of Xe enriched to $\sim$90\% in the {$ββ$}-decaying isotope $^{136}$Xe. In parallel, a potential future upgrade to nEXO is being investigated with the aim to further suppress radioactive backgrounds and to confirm $ββ$-decay events. This technique, known as Ba-tagging, comprises extracting and identifying the $ββ$-decay daughter $^{136}$Ba ion. One tagging approach being pursued involves extracting a small volume of LXe in the vicinity of a potential $ββ$-decay using a capillary tube and facilitating a liquid-to-gas phase transition by heating the capillary exit. The Ba ion is then separated from the accompanying Xe gas using a radio-frequency (RF) carpet and RF funnel, conclusively identifying the ion as $^{136}$Ba via laser-fluorescence spectroscopy and mass spectrometry. Simultaneously, an accelerator-driven Ba ion source is being developed to validate and optimize this technique. The motivation for the project, the development of the different aspects, along with the current status and results, are discussed here.

physics.ins-det

Enhancing Semi-Supervised Learning via Representative and Diverse Sample Selection

Semi-Supervised Learning (SSL) has become a preferred paradigm in many deep learning tasks, which reduces the need for human labor. Previous studies primarily focus on effectively utilising the labelled and unlabeled data to improve performance. However, we observe that how to select samples for labelling also significantly impacts performance, particularly under extremely low-budget settings. The sample selection task in SSL has been under-explored for a long time. To fill in this gap, we propose a Representative and Diverse Sample Selection approach (RDSS). By adopting a modified Frank-Wolfe algorithm to minimise a novel criterion $α$-Maximum Mean Discrepancy ($α$-MMD), RDSS samples a representative and diverse subset for annotation from the unlabeled data. We demonstrate that minimizing $α$-MMD enhances the generalization ability of low-budget learning. Experimental results show that RDSS consistently improves the performance of several popular SSL frameworks and outperforms the state-of-the-art sample selection approaches used in Active Learning (AL) and Semi-Supervised Active Learning (SSAL), even with constrained annotation budgets.

cs.LG

Real-time Position Reconstruction for the KamLAND-Zen Experiment using Hardware-AI Co-design

Monolithic liquid scintillator detector technology is the workhorse for detecting neutrinos and exploring new physics. The KamLAND-Zen experiment exemplifies this detector technology and has yielded top results in the quest for neutrinoless double-beta ($0νββ$) decay. To understand the physical events that occur in the detector, experimenters must reconstruct each event's position and energy from the raw data produced. Traditionally, this information has been obtained through a time-consuming offline process, meaning that event position and energy would only be available days after data collection. This work introduces a new pipeline to acquire this information quickly by implementing a machine learning model, PointNet, onto a Field Programmable Gate Array (FPGA). This work outlines a successful demonstration of the entire pipeline, showing that event position and energy information can be reliably and quickly obtained as physics events occur in the detector. This marks one of the first instances of applying hardware-AI co-design in the context of $0νββ$ decay experiments.

hep-ex

Evaluation of the KLauS ASIC at low temperature

The Taishan Antineutrino Observatory (TAO) is proposed to first use a cold liquid scintillator detector (-50~$^\circ$C) equipped with large-area silicon photomultipliers (SiPMs) ($\sim$10~m$^2$) to precisely measure the reactor antineutrino spectrum with a record energy resolution of < 2\% at 1 MeV. The KLauS ASIC shows excellent performance at room temperature and is a potential readout solution for TAO. In this work, we report evaluations of the fifth version of the KLauS ASIC (KLauS5) from room temperature to -50~$^\circ$C with inputs of injected charge or SiPMs. Our results show that KLauS5 has good performance at the tested temperatures with no significant degradation of the charge noise, charge linearity, gain uniformity or recovery time. Meanwhile, we also observe that several key parameters degrade when the chip operates in cold conditions, including the dynamic range and power consumption. However, even with this degradation, a good signal-to-noise ratio and good resolution of a single photoelectron can still be achieved for the tested SiPM with a gain of greater than 1.5$\times$10$^6$ and even an area of SiPM up to 1~cm$^2$ in one channel, corresponding to an input capacitance of approximately 5~nF. Thus, we conclude that KLauS5 can fulfill the TAO requirements for precise charge measurement.

physics.ins-det

$r$-strongly vertex-distinguishing total coloring of graphs

Inspired by the phenomenon of co-channel interference in communication network, a novel graph parameter, called $r$-vertex-strongly-distinguishing total coloring (abbreviate as $D(r)$-VSDTC), is proposed in this paper. Given a graph $G$, an $r$-VSDTC is an assignment of $k$ colors to $V(G)\cup E(G)$ such that any two adjacent or incident elements receive different colors and any two vertices with distance at most $r$ have distinct color-set, where the color-set of a vertex $u$ is the set of colors assigned on $u$ and its neighborhoods and incident edges. The \emph{$r$-vertex-strongly-distinguishing total chromatic number} of $G$, denoted by $χ_{r-vsdt}(G)$, is the minimum integer $k$ for which $G$ admits a $k$-$D(r)$-VSDTC. We show that $χ_{1-vsdt}(G)\leq 4Δ(G)$ for every graph $G$ without isolated edges and $χ_{1-vsdt}(G)\le kΔ(G)+3$ for a $k$-degenerated graph $G$ without isolated edges, where $1\le k\le 3$.

math.CO

The characterization of perfect Roman domination stable trees

A \emph{perfect Roman dominating function} (PRDF) on a graph $G = (V, E)$ is a function $f : V \rightarrow \{0, 1, 2\}$ satisfying the condition that every vertex $u$ for which $f(u) = 0$ is adjacent to exactly one vertex $v$ for which $f(v) = 2$. The weight of a PRDF is the value $w(f) = \sum_{u \in V}f(u)$. The minimum weight of a PRDF on a graph $G$ is called the \emph{perfect Roman domination number $γ_R^p(G)$} of $G$. A graph $G$ is perfect Roman domination domination stable if the perfect Roman domination number of $G$ remains unchanged under the removal of any vertex. In this paper, we characterize all trees that are perfect Roman domination stable.

math.CO

On uniquely 3-colorable plane graphs without prescribed adjacent faces

A graph $G$ is \emph{uniquely k-colorable} if the chromatic number of $G$ is $k$ and $G$ has only one $k$-coloring up to permutation of the colors. For a plane graph $G$, two faces $f_1$ and $f_2$ of $G$ are \emph{adjacent $(i,j)$-faces} if $d(f_1)=i$, $d(f_2)=j$ and $f_1$ and $f_2$ have a common edge, where $d(f)$ is the degree of a face $f$. In this paper, we prove that every uniquely 3-colorable plane graph has adjacent $(3,k)$-faces, where $k\leq 5$. The bound 5 for $k$ is best possible. Furthermore, we prove that there exist a class of uniquely 3-colorable plane graphs having neither adjacent $(3,i)$-faces nor adjacent $(3,j)$-faces, where $i,j\in \{3,4,5\}$ and $i \neq j$. One of our constructions implies that there exist an infinite family of edge-critical uniquely 3-colorable plane graphs with $n$ vertices and $\frac{7}{3}n-\frac{14}{3}$ edges, where $n(\geq 11)$ is odd and $n\equiv 2\pmod{3}$.

math.CO

Tree-colorable maximal planar graphs

A tree-coloring of a maximal planar graph is a proper vertex $4$-coloring such that every bichromatic subgraph, induced by this coloring, is a tree. A maximal planar graph $G$ is tree-colorable if $G$ has a tree-coloring. In this article, we prove that a tree-colorable maximal planar graph $G$ with $δ(G)\geq 4$ contains at least four odd-vertices. Moreover, for a tree-colorable maximal planar graph of minimum degree 4 that contains exactly four odd-vertices, we show that the subgraph induced by its four odd-vertices is not a claw and contains no triangles.

math.CO

Size of edge-critical uniquely 3-colorable planar graphs

A graph $G$ is \emph{uniquely k-colorable} if the chromatic number of $G$ is $k$ and $G$ has only one $k$-coloring up to permutation of the colors. A uniquely $k$-colorable graph $G$ is edge-critical if $G-e$ is not a uniquely $k$-colorable graph for any edge $e\in E(G)$. Mel'nikov and Steinberg [L. S. Mel'nikov, R. Steinberg, One counterexample for two conjectures on three coloring, Discrete Math. 20 (1977) 203-206] asked to find an exact upper bound for the number of edges in a edge-critical 3-colorable planar graph with $n$ vertices. In this paper, we give some properties of edge-critical uniquely 3-colorable planar graphs and prove that if $G$ is such a graph with $n(\geq6)$ vertices, then $|E(G)|\leq \frac{5}{2}n-6 $, which improves the upper bound $\frac{8}{3}n-\frac{17}{3}$ given by Matsumoto [N. Matsumoto, The size of edge-critical uniquely 3-colorable planar graphs, Electron. J. Combin. 20 (3) (2013) $\#$P49]. Furthermore, we find some edge-critical 3-colorable planar graphs which have $n(=10,12, 14)$ vertices and $\frac{5}{2}n-7$ edges.

math.CO