SearcharxivSearch

arXiv subjects

Shuyan Chen

Publications and source records attributed to Shuyan Chen.

10 recordsLinked to original sources

Cyclic Sources of Strong Domination in Graph Norms

Conlon and Lee asked for strongly dominating graphs beyond norming graphs and even paths. We construct a two-parameter family of pairwise non-isomorphic $2$-connected strongly dominating graphs that are not seminorming, and hence lie outside the two classes of examples previously identified for signed strong domination. The construction uses cyclic amalgamation of two-rooted blocks. For root-reversible blocks, we characterize the generation of all even cyclic amalgams by local even-Schatten inequalities for transfer operators. We determine this criterion for $K_{2,m}$, with the roots in the part of size $m$: it holds exactly when $m$ is even. We also classify the connected outerplanar strongly dominating graphs and the connected root-reversible outerplanar blocks satisfying the universal cyclic criterion.

math.CO

Odd-Cycle Span Defect: A Polynomial Lower Bound and a Square-Root Upper Bound

For a graph $G$, let $\psi(G)=\max\{\chi(G[V(C)]):C$ is an odd cycle of $G\}$, with $\psi(G)=0$ when $G$ is bipartite. For positive integers $N$, set $F(N)=\max\{\chi(G)-\psi(G):|V(G)|\le N\}$. The function $F$ measures the finite-order additive gap arising from an open problem of Erdos and Hajnal. We prove $N^{1/6-o(1)}\le F(N)<\sqrt{6N}$. The lower bound raises the finite-order scale supplied by the Cameron-Clow path-colour construction from $\log N/\log\log N$ to a fixed power of $N$. Its proof constructs a palette-code graph from a binary covering code $\mathcal{C}\subseteq\{0,1\}^p$ and establishes the exact identities $\chi(G)=2p+\ell-\rho(\mathcal{C})$ and $\psi(G)=2p$. Near-middle Hamming coverings yield the exponent $1/6$. The upper bound combines Polavarapu's connectivity theorem, the Chvatal-Erdos Hamiltonicity theorem, and maximum-independent-set stripping.

math.CO

Binary smoothing and relative Turan densities of ordered triangle-tails

For every $b\ge1$, let $Q_{2,b}$ be the ordered graph obtained from a transitive ordered triangle by attaching a monotone tail of length $b$ at its rightmost vertex. We prove $\rho_{<}(Q_{2,b})=\frac12$ for all $b\ge1$. Thus the previously isolated case $Q_{2,2}$ is one member of an exact infinite triangle-tail family. The lower bound is the sharp forward-template identity $\lambda(Q_{2,b})=1/2$, and it is realized already inside binary-level hosts: a parity-cut construction gives $Q_{2,b}$-free binary-level graphs with density exactly $1/2$ on every level. The upper bound uses the binary rich-level reduction. Its key input is the intrinsic decomposition of a $Q_{2,b}$-free graph into tail-starting vertices $T_b$ and the complement $R$: there are no forward edges from $R$ to $T_b$, $G[T_b]$ is ordered-triangle-free, and $G[R]$ is monotone-$\vec{P}_{b+1}$-free. We control these pieces by weighted binary $\vec{P}_{b+1}$ smoothing and weighted binary Mantel smoothing, the latter following from a binary ultrametric cut-domination theorem. We also record exact path-blow-up values, giving a reusable template/rich-host calculus for ordered relative densities.

math.CO

CoLaDAG: Compositional Latent Log-ratio DAG Analysis of the Gut Microbiome under Silver Nanoparticle Exposure

Directed network analysis of microbiome counts is complicated by compositional sampling, high dimensionality, and limited biological replication. We present CoLaDAG, a fixed-reference latent additive log-ratio (ALR) estimator for generating sparse directed conditional-dependence hypotheses from compositional counts. The method combines a multinomial observation model, a working linear Gaussian structural equation model, nonconvex DC-ADMM optimization, and post-estimation thresholding with greedy acyclic projection. Under simulations aligned with this observation model, CoLaDAG obtained the largest mean exact-direction Matthews correlation and the smallest mean false discovery rate among the evaluated implementations; performance deteriorated under continuous-data and dropout misspecification. In the 12-mouse silver-nanoparticle (AgNP) case study, the 58-node fitted graph was sensitive to block resampling and ALR reference choice: 60 of 284 primary edges attained a mouse-block selection frequency of at least 0.60. The reported orientations and dose-stratified slopes are exploratory, coordinate-specific hypotheses rather than identified causal or exposure effects. The leading stable relations prioritize anaerobic gut taxa for targeted abundance, metabolite, and perturbation studies, but do not establish cross-feeding or toxicological mechanisms.

stat.AP

From Literature to Lab: Closed-Loop Advancement of Perovskite Solar Cells via Domain Knowledge Guided LLM

Perovskite solar cells (PSCs) have been considered as a next-generation disruptive photovoltaic technology, yet their advancement is constrained by the complexity of perovskite recipe with high-dimensional material and process design space. Despite the impressive general reasoning of Large Language Models (LLMs), they struggle with two limitations for application in PSCs: an inability to align general semantics with the perovskite domain knowledge, and an inefficiency in navigating high-dimensional perovskite material and recipe design spaces. To address these limitations, we introduce a domain-knowledge-guided framework PVK-LLM, a specialized model to serve as an expert to bridge general semantics with perovskite domain knowledge. By integrating this domain knowledge into a hierarchical Bayesian Optimization workflow, our approach efficiently navigates the high-dimension design space on a solar cell simulator platform. The domain knowledge resolves cold-start problems while dynamically adapting to simulator feedback. Moreover, in an individual wet-lab experiment aimed at maximizing power conversion efficiency (PCE), our framework autonomously proposes a novel synergistic four-component recipe comprising specialized organic passivation recipe (3MTPAI, PDAI2, EDAI2, and PipDI) which has not been reported in existing literature. This AI-designed recipe effectively achieves a champion PCE value of over 26.0 %, approaching world records achieved through extensive expert trial-and-error. Our approach can effectively enable LLM comprehend the domain knowledge, which can efficiently navigate in a high-dimensional, capable to accelerate the advancement in real-world perovskite as well as other material science development.

cond-mat.mtrl-sci

Perovskite-LLM: Knowledge-Enhanced Large Language Models for Perovskite Solar Cell Research

The rapid advancement of perovskite solar cells (PSCs) has led to an exponential growth in research publications, creating an urgent need for efficient knowledge management and reasoning systems in this domain. We present a comprehensive knowledge-enhanced system for PSCs that integrates three key components. First, we develop Perovskite-KG, a domain-specific knowledge graph constructed from 1,517 research papers, containing 23,789 entities and 22,272 relationships. Second, we create two complementary datasets: Perovskite-Chat, comprising 55,101 high-quality question-answer pairs generated through a novel multi-agent framework, and Perovskite-Reasoning, containing 2,217 carefully curated materials science problems. Third, we introduce two specialized large language models: Perovskite-Chat-LLM for domain-specific knowledge assistance and Perovskite-Reasoning-LLM for scientific reasoning tasks. Experimental results demonstrate that our system significantly outperforms existing models in both domain-specific knowledge retrieval and scientific reasoning tasks, providing researchers with effective tools for literature review, experimental design, and complex problem-solving in PSC research.

cs.AI

A penalized online sequential test of heterogeneous treatment effects for generalized linear models

Identification of heterogeneous treatment effects (HTEs) has been increasingly popular and critical in various penalized strategy decisions using the A/B testing approach, especially in the scenario of a consecutive online collection of samples. However, in high-dimensional settings, such an identification remains challenging in the sense of lack of detection power of HTEs with insufficient sample instances for each batch sequentially collected online. In this article, a novel high-dimensional test is proposed, named as the penalized online sequential test (POST), to identify HTEs and select useful covariates simultaneously under continuous monitoring in generalized linear models (GLMs), which achieves high detection power and controls the Type I error. A penalized score test statistic is developed along with an extended p-value process for the online collection of samples, and the proposed POST method is further extended to multiple online testing scenarios, where both high true positive rates and under-controlled false discovery rates are achieved simultaneously. Asymptotic results are established and justified to guarantee properties of the POST, and its performance is evaluated through simulations and analysis of real data, compared with the state-of-the-art online test methods. Our findings indicate that the POST method exhibits selection consistency and superb detection power of HTEs as well as excellent control over the Type I error, which endows our method with the capability for timely and efficient inference for online A/B testing in high-dimensional GLMs framework.

stat.ME

A Dynamic Model for Frequency Response Optimization in Photovoltaic Visible Light Communication

Photovoltaic (PV) modules are recently employed in photovoltaic visible light communication (PVLC) for simultaneous energy harvesting and visible light communication. A PV-based receiver features large signal output, easy optical alignment, and self-powered operation. However, PV modules usually have a severe bandwidth limitation when used as passive photodetectors. In this paper, we systematically investigate the internal impedance dynamic of PV modules and how that affects their frequency response characteristics under different illuminances. We propose a simplified yet accurate dynamic PV mode AC detection model to capture the frequency response characteristics of a PVLC receiver. The model is validated with the impedance spectroscopy characterization methodologies. Experimental results show that a PV module's internal resistance and capacitance depend on incident illuminance, affecting PV's frequency response. The bandwidth is exacerbated under indoor environments with low illuminance levels due to the increment of internal resistance for PV modules. The RC constant can be reduced for PVLC receivers working near open-circuit voltage conditions by adding a moderate local light to decrease the internal resistance value. For practical implementation, PVLC receivers will employ a load for data recovery. We show that adjusting the forward bias conditions can simultaneously reduce the resistance and capacitance values. With the optimization of equivalent trans-impedance, the data rate of a Cadmium telluride (CdTe) PV module achieves a 3.8 times enhancement under 200 lux. We also demonstrate that the BER of a 5-Mbit/s eight-level pulse amplitude modulation (PAM8) signal can be reduced from 9.8*10-2 to 1.4*10-3 by maximizing the transimpedance gain-bandwidth product.

eess.SP

On the Nonlinear Distortion Characterization in Photovoltaic Modules for Visible Light Communication

Photovoltaic (PV) modules have been employed in visible light communication (VLC) for simultaneous energy harvesting and data reception. A PV-based receiver features easy optical alignment and self-powered operation. It is commonly assumed that PV modules in VLC have a linear optical-electrical response, which is generally true under high illumination levels. This paper will illustrate the exacerbated PV's nonlinear distortion when under typical indoor illumination. The nonlinearity of a PV module for different numbers of PV cells is also characterized. We investigated the transmission performance of a 1-Mbit/s PAM4 signal under different illuminance. Experimental results show that the bit-error rate (BER) first decreases and then increases with the increasing illuminance. Thus, an optimal illuminance to minimize BER exists. In addition, we demonstrated two distortion mitigation methods, namely localized distortion compensation lighting and post-distortion compensation. BER reduction from 3.2x10-1 to 2.6x10-3 and 2.8x10-2 to 1.5x10-2 are achieved with the two respective schemes.

eess.SP