SearcharxivSearch

arXiv subjects

Ahmed Wali

Publications and source records attributed to Ahmed Wali.

3 recordsLinked to original sources

Before You Poll with LLMs: A Deliberative Diagnostic Framework

Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan opinions can still misrepresent how those opinions change. Applying the framework to five frontier models using data from America in One Room (526 personas, 72 questions), we find that every model fails, each in a unique manner. GPT-5.1 exhibits reversal: its personas become more hostile toward the opposing party after balanced information, while humans become less so. This reversal is selective (80% on outgroup vs. 26% on policy questions) and symmetric across partisan identities. Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B exhibit overshoot, shifting correctly but at 5-7x human magnitude. DeepSeek V3 exhibits rigidity with near-zero change. Targeted ablations reveal that policy content triggers these failures and that they are identity-specific: GPT-5.1 reverses on outgroup questions but overshoots on ingroup; Gemini shows the inverse. We term this signature self-sycophancy: conformity to the model's internal stereotype of the persona rather than reasoning from the information provided. Our framework offers a concrete protocol: run the deliberative diagnostic before trusting LLM personas to mimic revised beliefs.

cs.CL

A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification

Although supervised machine learning is popular for information extraction from clinical notes, creating large annotated datasets requires extensive domain expertise and is time-consuming. Meanwhile, large language models (LLMs) have demonstrated promising transfer learning capability. In this study, we explored whether recent LLMs can reduce the need for large-scale data annotations. We curated a manually-labeled dataset of 769 breast cancer pathology reports, labeled with 13 categories, to compare zero-shot classification capability of the GPT-4 model and the GPT-3.5 model with supervised classification performance of three model architectures: random forests classifier, long short-term memory networks with attention (LSTM-Att), and the UCSF-BERT model. Across all 13 tasks, the GPT-4 model performed either significantly better than or as well as the best supervised model, the LSTM-Att model (average macro F1 score of 0.83 vs. 0.75). On tasks with high imbalance between labels, the differences were more prominent. Frequent sources of GPT-4 errors included inferences from multiple samples and complex task design. On complex tasks where large annotated datasets cannot be easily collected, LLMs can reduce the burden of large-scale data labeling. However, if the use of LLMs is prohibitive, the use of simpler supervised models with large annotated datasets can provide comparable results. LLMs demonstrated the potential to speed up the execution of clinical NLP studies by reducing the need for curating large annotated datasets. This may result in an increase in the utilization of NLP-based variables and outcomes in observational clinical studies.

cs.CL

Efficient and Sustainable Treatment of Tannery Wastewater by a Sequential Electrocoagulation-UV Photolytic Process

Tannery wastewater contains large amounts of pollutants that, if directly discharged into ecosystems, can generate an environmental hazard. The present investigation has focused the attention to the remediation of wastewater originated from tanned leather in Tunisia. The analysis revealed wastewater with a high level of chemical oxygen demand (COD) of 7376 mgO2/L. The performance in reduction of COD, via electrocoagulation (EC) or UV photolysis or, finally, operating electrocoagulation and photolysis in sequence was examined. The effect of voltage and reaction time on COD reduction, as well as the phytotoxicity were determined. Treated effluents were analysed by UV spectroscopy, extracting the organic components with solvents differing in polarity. A sequential EC and UV treatment of the tannery wastewater has been proven effective in the reduction of COD. These treatments combined afforded 94.1 % of COD reduction, whereas the single EC and UV treatments afforded respectively 85.7 and 55.9 %. The final COD value of 428.7 mg/L was found largely below the limit of 1000 mg/L for admission of wastewater in public sewerage network. Germination tests of Hordeum Vulgare seeds indicated reduced toxicity for the remediated water. Energy consumptions of 33.33 kWh/m3 and 314.28 kWh/m3 were determined for the EC process and for the same followed by UV treatment. Both those technologies are yet available and ready for scale-up.

physics.bio-ph