Searcharxiv⌕ Search

arXiv subjects

Hassan Eshkiki

Publications and source records attributed to Hassan Eshkiki.

4 recordsLinked to original sources

Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition

Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference rather than automated reasoning. LLMs could help fill these gaps, but single-pass outputs are unreliable and can introduce silent errors into downstream computation. We present a quality-controlled LLM pipeline for ingredient data acquisition that combines robust statistical estimation, domain-specific invariant checks, and a web-fetch fallback. An illustrative Heap's Law fit to 233 recipes suggests that unique-ingredient growth is sub-linear and front-loaded: the projected ratio of unique ingredients to recipes falls from 1.74 at 100 recipes to 0.19 at 5,000. For each ingredient attribute, repeated LLM queries are treated as samples from a model-induced answer distribution, and we apply robust point estimators and normalised confidence scores across numerical, Boolean, multiple-choice, open categorical, and optional integer types. An invariant guard layer enforces nutritional and logical self-consistency within each ingredient record. Minor numeric inconsistencies are reconciled via a linear program that minimises worst-case percentage deviation while preserving semantic zeros, and major violations are escalated to web-evidence-grounded repair, then human review only if that fails. On a curated 30-ingredient reference set, the pipeline achieves 98.4% exact match on nutrient flags and cuts median absolute percentage error on nutrient ratios from 31.9% for the median-aggregated baseline to 10.1%, a reduction of 21.8 percentage points, at an API cost of about $1 per ingredient. This frames LLM-assisted database construction as a controlled data-engineering workflow that makes uncertainty operational rather than discarding it.

cs.IR↗

Automated detection of circadian-dependent epileptic biomarkers for seizure localization using machine learning and signal processing

Accurate localization of the seizure onset zone (SOZ) is essential for successful epilepsy surgery, yet the reliability of commonly used interictal biomarkers is limited by temporal variability and behavioral state. This study aims to investigate the circadian and sleep-dependent dynamics of epileptic biomarkers and to identify conditions that maximize seizure localization precision. Longterm intracranial EEG recordings from nine patients with drug-resistant focal epilepsy were retrospectively analyzed using automated signal processing and machine learning techniques. Interictal spikes, spike sequences, high-frequency oscillations (HFOs), and pathological HFOs were automatically detected, while sleep and wake states were classified using the alpha-delta power ratio. Biomarker rates, spatial distributions, and localization accuracy were quantitatively evaluated using Euclidean distance relative to the clinically defined SOZ. The results show that all biomarkers exhibit significantly higher rates during sleep, with pronounced early-morning peaks. Importantly, spike sequences and pathological HFOs demonstrated superior spatial precision compared to conventional spikes or HFOs alone. Mean distances to the SOZ were substantially lower for pathological HFOs and spike sequences, with statistically significant differences among biomarkers (ANOVA, p < 0.001). These findings demonstrate that sleep-state analysis, particularly using propagated spike sequences and pathological HFOs, substantially improves SOZ localization accuracy. The proposed framework provides practical guidance for sleep-focused presurgical EEG analysis and supports the development of automated and clinically efficient seizure localization systems.

eess.SP↗

Multi-Temporal Frames Projection for Dynamic Processes Fusion in Fluorescence Microscopy

Fluorescence microscopy is widely employed for the analysis of living biological samples; however, the utility of the resulting recordings is frequently constrained by noise, temporal variability, and inconsistent visualisation of signals that oscillate over time. We present a unique computational framework that integrates information from multiple time-resolved frames into a single high-quality image, while preserving the underlying biological content of the original video. We evaluate the proposed method through an extensive number of configurations (n = 111) and on a challenging dataset comprising dynamic, heterogeneous, and morphologically complex 2D monolayers of cardiac cells. Results show that our framework, which consists of a combination of explainable techniques from different computer vision application fields, is capable of generating composite images that preserve and enhance the quality and information of individual microscopy frames, yielding 44% average increase in cell count compared to previous methods. The proposed pipeline is applicable to other imaging domains that require the fusion of multi-temporal image stacks into high-quality 2D images, thereby facilitating annotation and downstream segmentation.

cs.CV↗

Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative Analysis

This work contributes towards balancing the inclusivity and global applicability of natural language processing techniques by proposing the first 'name entity recognition' dataset for Kurdish Sorani, a low-resource and under-represented language, that consists of 64,563 annotated tokens. It also provides a tool for facilitating this task in this and many other languages and performs a thorough comparative analysis, including classic machine learning models and neural systems. The results obtained challenge established assumptions about the advantage of neural approaches within the context of NLP. Conventional methods, in particular CRF, obtain F1-scores of 0.825, outperforming the results of BiLSTM-based models (0.706) significantly. These findings indicate that simpler and more computationally efficient classical frameworks can outperform neural architectures in low-resource settings.

cs.CL↗