SearcharxivSearch

arXiv subjects

Jia Yang

Publications and source records attributed to Jia Yang.

At least 19 recordsLinked to original sources

GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis

Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranked above it or what evidence a tool-using model examines before making its decision. We introduce GraphRareBench, a provenance-preserving benchmark containing 2,365 ontology-derived cases and 18,093 target-confounder pairs. Each case includes a coarsened HPO query, a fixed candidate pool, graph-defined hard confounders, and source-linked evidence records. On the 237-case gene-component-disjoint test split, supervised rankers using a shared 21-feature interface achieved MRRs ranging from 0.640 to 0.740 and case-averaged target-over-confounder accuracies ranging from 0.898 to 0.916. Agents instantiated with Agents-A1 and DeepSeek-V4-Flash achieved MRRs of 0.746 and 0.718, respectively. Their paired MRR difference was not statistically significant, whereas their target-evidence coverage differed by 0.561. Together with the observation that 22.1% to 43.7% of selected Hit@10 successes still ranked at least one graph-defined hard confounder above the target, these results indicate that full-pool retrieval, hard-confounder discrimination, and observable evidence access capture complementary aspects of model behavior. GraphRareBench therefore provides a foundation for more transparent and evidence-aware evaluation of phenotype-driven diagnostic systems. Code and data are available at https://github.com/GUI0609/GraphRareBench.

q-bio.QM

The Equality Cases for the Grone-Merris-Bai Theorem

The Grone--Merris inequality, conjectured by Grone and Merris~(1994) and first proved by Bai~(2011), states that for every graph $G$ of order $n$ and every $1\le k\le n$, $\sum_{i=1}^k\lambda_i(G)\le\sum_{i=1}^k d_i^*(G)$, where $\lambda_1\ge\cdots\ge\lambda_n$ are the Laplacian eigenvalues and $d_1^*\ge\cdots\ge d_n^*$ is the conjugate degree sequence. In this paper we determine exactly when equality holds. Using the split-graph trace inequality developed by Kothari and Tudose~(2026) in their proof of Brouwer's Laplacian conjecture---which relies on Bai's theorem and also establishes the equivalence between the two conjectures---together with the recent characterization of the Brouwer equality cases by Cai, Chen, Yang and Zhang~(2027), we prove that equality holds in the Grone--Merris inequality if and only if the graph $G$ belongs to one of two explicitly described families. Both families are obtained from a threshold graph by a surgical operation at one terminal block: in the first family, edges are removed from the initial dominating block; in the second, edges are added inside the initial isolated block. Our analysis yields a complete combinatorial description of all pairs $(G,k)$ for which the Grone--Merris bound is tight.

math.CO

The Equality Cases For the Laplacian Conjecture of Brouwer

The Laplacian conjecture of Brouwer asserts that for any graph \(G\) of order n with \(m\) edges, the sum of the \(k\) largest Laplacian eigenvalues satisfies \(s_k(G) \le m + \binom{k+1}{2}\) for $k=1, \ldots, n$. Later, Li and Guo in 2022 further proposed the full Brouwer's Laplacian spectrum conjecture. Recently, Kothari and Tudose in 2026 proved the Brouwer's conjecture. Motivated by their perfect proof and methods, we proved that for a simple graph of order $n$ with $m$ edges and $1\le k\le n-1$, \(s_k(G) = m + \binom{k+1}{2}\) if and only if $G$ is a threshold graph with clique number \(k+1\), which confirms the full Brouwer conjecture proposed by Li and Guo.

math.CO

A Survey on GNN-based Link Prediction: Techniques, Applications, and Challenges

Graph Neural Networks (GNNs) have emerged as the leading paradigm for link prediction, enabling the inference of missing connections and the anticipation of potential future links. However, existing reviews lack systematic exploration specifically targeting underlying GNN architectures and diverse graph structures. To address this critical gap, this paper provides a comprehensive review of GNN-based link prediction from a novel and dedicated GNN perspective. We propose an innovative taxonomy that categorizes recent advancements based on techniques and applications. From a technique perspective, we focus on key GNN encoder architectures, including GCN-based, GAE-based, GAT-based, and GFormer-based methods, discussing their strengths and limitations. From an application perspective, we highlight prominent use cases of link prediction in knowledge graphs and recommendation systems, demonstrating their real-world impact. In addition, we examine the current challenges and discuss promising future directions.

cs.AI

Investigation of the Physical Mechanism behind Retention Loss in FeFETs with MIFIFIS Gate Structure

A Metal-Gate Blocking Layer (GBL)- Ferroelectric-Tunnel Dielectric Layer (TDL)-Ferroelectric -Channel Insulator (Ch.IL)-Si (MIFIFIS) structure is proposed to achieve a larger MW for applications in Fe-NAND. However, the large retention loss (RL) in the MIFIFIS structure restricts its application. In this work, we vary the physical thickness of the GBL and TDL, and conduct an in-depth analysis of the energy bands of the gate structure to investigate the physical mechanism behind the RL in FeFETs with the MIFIFIS structure. The physical origin of the RL is that the electric field direction across the TDL reduces the potential barrier provided by the ferroelectric near the silicon substrate. Based on the above physical mechanism, the RL can be reduced to 12% and 0.2% by redesigning the gate structure or reducing the pulse amplitude, respectively. Our work contributes to a deeper understanding of the physical mechanism behind the RL in FeFETs with the MIFIFIS gate structure. It provides guidance for enhancing the reliability of FeFETs.

physics.app-ph

Unveiling Retention Loss Mechanism in FeFETs with Gate-side Interlayer by Decoupling Trapped Charges and Ferroelectric Polarization

We propose a direct experimental extraction technique for trapped charges and quantitative energy band diagrams in the FeFETs with metal-insulator-ferroelectric-insulator-semiconductor (MIFIS) structure, derived from the physical relationship between Vth and gate-side interlayer (G.IL) thickness. By decoupling trapped charges and ferroelectric polarization, we reveal that: (i) The gateinjected charges and channel-injected charges are excessive and maintain consistent ratios to ferroelectric polarization (~170% and ~130%, respectively). (ii) Retention loss originates from the detrapping of gate-injected charges rather than ferroelectric depolarization. (iii) As the G.IL thickens, the gate-injected charge de-trapping path transforms from gate-side to channel-side. To address the retention loss, careful material design, optimization, and bandgap engineering in the MIFIS structure are crucial. This work advances the understanding of high retention strategies for MIFIS-FeFETs in 3D FE NAND.

physics.app-ph

A Computation of Tamarkin-Tsygan Calculus

We compute the full Tamarkin-Tsygan calculus of a Koszul algebra whose global dimension exceeds the number of generators. Our results show that even for algebras possessing an economic presentation and agreeable homological properties, the Hochschild (co)homology, as well as the structure of the Tamarkin--Tsygan calculus may exhibit a rather intricate behavior.

math.RA

ALOHA: Empowering Multilingual Agent for University Orientation with Hierarchical Retrieval

The rise of Large Language Models~(LLMs) revolutionizes information retrieval, allowing users to obtain required answers through complex instructions within conversations. However, publicly available services remain inadequate in addressing the needs of faculty and students to search campus-specific information. It is primarily due to the LLM's lack of domain-specific knowledge and the limitation of search engines in supporting multilingual and timely scenarios. To tackle these challenges, we introduce ALOHA, a multilingual agent enhanced by hierarchical retrieval for university orientation. We also integrate external APIs into the front-end interface to provide interactive service. The human evaluation and case study show our proposed system has strong capabilities to yield correct, timely, and user-friendly responses to the queries in multiple languages, surpassing commercial chatbots and search engines. The system has been deployed and has provided service for more than 12,000 people.

cs.CL

Pseudo-Knowledge Graph: Meta-Path Guided Retrieval and In-Graph Text for RAG-Equipped LLM

The advent of Large Language Models (LLMs) has revolutionized natural language processing. However, these models face challenges in retrieving precise information from vast datasets. Retrieval-Augmented Generation (RAG) was developed to combining LLMs with external information retrieval systems to enhance the accuracy and context of responses. Despite improvements, RAG still struggles with comprehensive retrieval in high-volume, low-information-density databases and lacks relational awareness, leading to fragmented answers. To address this, this paper introduces the Pseudo-Knowledge Graph (PKG) framework, designed to overcome these limitations by integrating Meta-path Retrieval, In-graph Text and Vector Retrieval into LLMs. By preserving natural language text and leveraging various retrieval techniques, the PKG offers a richer knowledge representation and improves accuracy in information retrieval. Extensive evaluations using Open Compass and MultiHop-RAG benchmarks demonstrate the framework's effectiveness in managing large volumes of data and complex relationships.

cs.IR

DisCo: Graph-Based Disentangled Contrastive Learning for Cold-Start Cross-Domain Recommendation

Recommender systems are widely used in various real-world applications, but they often encounter the persistent challenge of the user cold-start problem. Cross-domain recommendation (CDR), which leverages user interactions from one domain to improve prediction performance in another, has emerged as a promising solution. However, users with similar preferences in the source domain may exhibit different interests in the target domain. Therefore, directly transferring embeddings may introduce irrelevant source-domain collaborative information. In this paper, we propose a novel graph-based disentangled contrastive learning framework to capture fine-grained user intent and filter out irrelevant collaborative information, thereby avoiding negative transfer. Specifically, for each domain, we use a multi-channel graph encoder to capture diverse user intents. We then construct the affinity graph in the embedding space and perform multi-step random walks to capture high-order user similarity relationships. Treating one domain as the target, we propose a disentangled intent-wise contrastive learning approach, guided by user similarity, to refine the bridging of user intents across domains. Extensive experiments on four benchmark CDR datasets demonstrate that DisCo consistently outperforms existing state-of-the-art baselines, thereby validating the effectiveness of both DisCo and its components.

cs.IR

Effect of Top Al$_2$O$_3$ Interlayer Thickness on Memory Window and Reliability of FeFETs With TiN/Al$_2$O$_3$/Hf$_{0.5}$Zr$_{0.5}$O$_2$/SiO$_x$/Si (MIFIS) Gate Structure

We investigate the effect of top Al2O3 interlayer thickness on the memory window (MW) of Si channel ferroelectric field-effect transistors (Si-FeFETs) with TiN/Al$_2$O$_3$/Hf$_{0.5}$Zr$_{0.5}$O$_2$/SiO$_x$/Si (MIFIS) gate structure. We find that the MW first increases and then remains almost constant with the increasing thickness of the top Al2O3. The phenomenon is attributed to the lower electric field of the ferroelectric Hf$_{0.5}$Zr$_{0.5}$O$_2$ in the MIFIS structure with a thicker top Al2O3 after a program operation. The lower electric field makes the charges trapped at the top Al2O3/Hf0.5Zr0.5O$_2$ interface, which are injected from the metal gate, cannot be retained. Furthermore, we study the effect of the top Al$_2$O$_3$ interlayer thickness on the reliability (endurance characteristics and retention characteristics). We find that the MIFIS structure with a thicker top Al$_2$O$_3$ interlayer has poorer retention and endurance characteristics. Our work is helpful in deeply understanding the effect of top interlayer thickness on the MW and reliability of Si-FeFETs with MIFIS gate stacks.

cond-mat.mtrl-sci

Impact of the Top SiO2 Interlayer Thickness on Memory Window of Si Channel FeFET with TiN/SiO2/Hf0.5Zr0.5O2/SiOx/Si (MIFIS) Gate Structure

We study the impact of top SiO2 interlayer thickness on the memory window (MW) of Si channel ferroelectric field-effect transistor (FeFET) with TiN/SiO2/Hf0.5Zr0.5O2/SiOx/Si (MIFIS) gate structure. We find that the MW increases with the increasing thickness of the top SiO2 interlayer, and such an increase exhibits a two-stage linear dependence. The physical origin is the presence of the different interfacial charges trapped at the top SiO2/Hf0.5Zr0.5O2 interface. Moreover, we investigate the dependence of endurance characteristics on initial MW. We find that the endurance characteristic degrades with increasing the initial MW. By inserting a 3.4 nm SiO2 dielectric interlayer between the gate metal TiN and the ferroelectric Hf0.5Zr0.5O2, we achieve a MW of 6.3 V and retention over 10 years. Our work is helpful in the device design of FeFET.

cond-mat.mtrl-sci

Impact of Top SiO2 interlayer Thickness on Memory Window of Si Channel FeFET with TiN/SiO2/Hf0.5Zr0.5O2/SiOx/Si (MIFIS) Gate Structure

We study the impact of top SiO2 interlayer thickness on memory window of Si channel FeFET with TiN/SiO2/Hf0.5Zr0.5O2/SiOx/Si (MIFIS) gate structure. The memory window increases with thicker top SiO2. We realize the memory window of 6.3 V for 3.4 nm top SiO2. Moreover, we find that the endurance characteristic degrades with increasing the initial memory window.

physics.app-ph

Crimson: Empowering Strategic Reasoning in Cybersecurity through Large Language Models

We introduces Crimson, a system that enhances the strategic reasoning capabilities of Large Language Models (LLMs) within the realm of cybersecurity. By correlating CVEs with MITRE ATT&CK techniques, Crimson advances threat anticipation and strategic defense efforts. Our approach includes defining and evaluating cybersecurity strategic tasks, alongside implementing a comprehensive human-in-the-loop data-synthetic workflow to develop the CVE-to-ATT&CK Mapping (CVEM) dataset. We further enhance LLMs' reasoning abilities through a novel Retrieval-Aware Training (RAT) process and its refined iteration, RAT-R. Our findings demonstrate that an LLM fine-tuned with our techniques, possessing 7 billion parameters, approaches the performance level of GPT-4, showing markedly lower rates of hallucination and errors, and surpassing other models in strategic reasoning tasks. Moreover, domain-specific fine-tuning of embedding models significantly improves performance within cybersecurity contexts, underscoring the efficacy of our methodology. By leveraging Crimson to convert raw vulnerability data into structured and actionable insights, we bolster proactive cybersecurity defenses.

cs.CR

Compound Attention and Neighbor Matching Network for Multi-contrast MRI Super-resolution

Multi-contrast magnetic resonance imaging (MRI) reflects information about human tissue from different perspectives and has many clinical applications. By utilizing the complementary information among different modalities, multi-contrast super-resolution (SR) of MRI can achieve better results than single-image super-resolution. However, existing methods of multi-contrast MRI SR have the following shortcomings that may limit their performance: First, existing methods either simply concatenate the reference and degraded features or exploit global feature-matching between them, which are unsuitable for multi-contrast MRI SR. Second, although many recent methods employ transformers to capture long-range dependencies in the spatial dimension, they neglect that self-attention in the channel dimension is also important for low-level vision tasks. To address these shortcomings, we proposed a novel network architecture with compound-attention and neighbor matching (CANM-Net) for multi-contrast MRI SR: The compound self-attention mechanism effectively captures the dependencies in both spatial and channel dimension; the neighborhood-based feature-matching modules are exploited to match degraded features and adjacent reference features and then fuse them to obtain the high-quality images. We conduct experiments of SR tasks on the IXI, fastMRI, and real-world scanning datasets. The CANM-Net outperforms state-of-the-art approaches in both retrospective and prospective experiments. Moreover, the robustness study in our work shows that the CANM-Net still achieves good performance when the reference and degraded images are imperfectly registered, proving good potential in clinical applications.

eess.IV

Distributionally Robust Optimal Power Flow with Uncertain Renewable Energy Output

Optimal power flow (OPF) is an important tool for Independent System Operators (ISOs) to deal with the power generation management. With the increasing penetration of renewable energy into power grids, challenges arise in tackling the OPF problem due to the intermittent nature of renewable energy output. To address these challenges, we develop a multi-stage distributionally robust approach for the direct-current optimal power flow (DC-OPF) problem to minimize total generation cost under renewable energy uncertainty. In our model, we assume the renewable energy output follows an ambiguous distribution that can be characterized by a confidence set. By utilizing the revealed data sequentially, the proposed approach can provide a reliable and robust optimal OPF decision without restricting the renewable energy output distribution to any particular distribution class. The computational results also verify the effectiveness of our approach to reduce the conservativeness and meanwhile maintain the reliability.

math.OC

Turbulence-assisted formation of bacterial cellulose

Bacterial cellulose is an important class of biomaterials which can be grown in well-controlled laboratory and industrial conditions. The cellulose structure is affected by several biological, chemical and environmental factors, including hydrodynamic flows in bacterial suspensions. In this work, we explore the possibility of using well controlled turbulent flow to control the bacterial cellulose production. Turbulent flows leading to random motion of fluid elements may affect the structure of the extracellular polymeric matrix produced by bacteria. Here we show that two-dimensional turbulence at the air-liquid interface generates chaotic rotation at a well-defined scale and random persistent stretching of the fluid elements. The results offer new approaches to engineering of the bacterial cellulose structure by controlling turbulence parameters.

physics.bio-ph

Molecular Dynamics Simulation of the Interaction between Cracks in Single-Crystal Aluminum

The interaction between cracks, as well as their propagation, in single-crystal aluminum is investigated at the atomic scale using the molecular dynamics method and the modified embedded atom method. The results demonstrated that the crack propagation in aluminum is a quite complex process, which is accompanied by micro-crack growth, merging, stress shielding, dislocation emission, and phase transformation of the crystal structure. The main deformation mechanisms at the front of the fatigue crack are holes, slip bands, and cross-slip bands. During crack propagation, there are interactions between cracks. Such interactions inhibit the phase transition at the crack tip, and also affect the direction and speed of crack propagation. The intensity of the interaction between two cracks decreases with the increase of the distance between them and increases with increasing crack size. Moreover, while this interaction has no effect on the elastic modulus of the material, it affects its strength limit.

cond-mat.mtrl-sci