SearcharxivSearch

arXiv subjects

Solomon Eshun

Publications and source records attributed to Solomon Eshun.

3 recordsLinked to original sources

Agentic Self-Healing for Data and AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software

Modern organizations rely on data, machine learning, and software delivery pipelines to move data, train models, deploy applications, refresh dashboards, and support business-critical decisions. However, these pipelines often fail because of data quality issues, schema changes, upstream source changes, infrastructure problems, orchestration failures, and model workflow issues. Existing ZeroOps, observability, and AI operations platforms can help teams detect incidents, investigate root causes, and in some cases recommend or execute fixes. However, many of these solutions are expensive, vendor-specific, or difficult for smaller teams to adapt across different tools and environments. This paper first compares existing off-the-shelf solutions for AI-assisted pipeline monitoring, root-cause analysis, and automated remediation, including their strengths, limitations, and practical trade-offs. Based on this comparison, we find that the main gap is architectural rather than technological: the required ingredients for self-healing pipelines already exist, but they are fragmented across vendor-specific platforms, observability tools, incident systems, and open-source components. We therefore propose an affordable, vendor-agnostic reference architecture for agentic self-healing pipelines using open-source and low-cost tools. The proposed architecture combines monitoring, pipeline metadata, incident history, deterministic policy checks, AI-assisted diagnosis, approval workflows, and controlled remediation actions to help teams detect, diagnose, repair, verify, and learn from pipeline issues with less manual effort. The goal is to provide a practical reference architecture that can be adapted across data engineering, machine learning operations, and software delivery environments.

cs.ET

Can Large Language Models Design Biological Weapons? Evaluating Moremi Bio

Advances in AI, particularly LLMs, have dramatically shortened drug discovery cycles by up to 40% and improved molecular target identification. However, these innovations also raise dual-use concerns by enabling the design of toxic compounds. Prompting Moremi Bio Agent without the safety guardrails to specifically design novel toxic substances, our study generated 1020 novel toxic proteins and 5,000 toxic small molecules. In-depth computational toxicity assessments revealed that all the proteins scored high in toxicity, with several closely matching known toxins such as ricin, diphtheria toxin, and disintegrin-based snake venom proteins. Some of these novel agents showed similarities with other several known toxic agents including disintegrin eristostatin, metalloproteinase, disintegrin triflavin, snake venom metalloproteinase, corynebacterium ulcerans toxin. Through quantitative risk assessments and scenario analyses, we identify dual-use capabilities in current LLM-enabled biodesign pipelines and propose multi-layered mitigation strategies. The findings from this toxicity assessment challenge claims that large language models (LLMs) are incapable of designing bioweapons. This reinforces concerns about the potential misuse of LLMs in biodesign, posing a significant threat to research and development (R&D). The accessibility of such technology to individuals with limited technical expertise raises serious biosecurity risks. Our findings underscore the critical need for robust governance and technical safeguards to balance rapid biotechnological innovation with biosecurity imperatives.

q-bio.QM

Equity in Focus : Investigating Gender Disparities in Glioblastoma via Propensity Score Matching

Gender disparities in health outcomes have garnered significant attention, prompting investigations into their underlying causes. Glioblastoma (GBM), a devastating and highly aggressive form of brain tumor, serves as a case for such inquiries. Despite the mounting evidence on gender disparities in GBM outcomes, investigations specific at the molecular level remain scarce and often limited by confounding biases in observational studies. In this study, I aimed to investigate the gender-related differences in GBM outcomes using propensity score matching (PSM) to control for potential confounding variables. The data used was accessed from the Cancer Genome Atlas (TCGA), encompassing factors such as gender, age, molecular characteristics and different glioma grades. Propensity scores were calculated for each patient using logistic regression, representing the likelihood of being male based on the baseline characteristics. Subsequently, patients were matched using the nearest-neighbor (with a restricted caliper) matching to create a balanced male-female group. After PSM, 303 male-female pairs were identified, with similar baseline characteristics in terms of age and molecular features. The analysis revealed a higher incidence of GBM in males compared to females, after adjusting for potential confounding factors. This study contributes to the discourse on gender equity in health, paving the way for targeted interventions and improved outcomes, and may guide efforts to improve gender-specific treatment strategies for GBM patients. However, further investigations and prospective studies are warranted to validate these findings and explore additional factors that might contribute to the observed gender-based differences in GBM outcomes aside from the molecular characteristics.

stat.AP