SearcharxivSearch

arXiv subjects

Sahar Maleki

Publications and source records attributed to Sahar Maleki.

4 recordsLinked to original sources

Beyond the Interface: Redefining UX for Society-in-the-Loop AI Systems

Artificial intelligence systems increasingly operate in decision-critical environments where probabilistic outputs and Human-in-the-Loop (HITL) interactions reshape user engagement. Traditional user experience (UX) frameworks, designed for deterministic systems, fail to capture these evolving sociotechnical dynamics. This paper argues that in AI-enabled HITL systems, UX must transcend frontend usability to encompass backend performance, organizational workflows, and decision making structures. We employ a mixed-methods approach, combining an inductive social construction analysis of 269 stakeholder insights with the deployment of an operational HITL video anomaly detection system. Our findings reveal that stakeholders experience AI through multifaceted themes: risk, governance, and organizational capacity. Experimental results further demonstrate how detection behavior and alert routing directly calibrate human oversight and workload. Grounded in these results, we formalize a new evaluative framework centered on four sociotechnical metrics: Accuracy (FPR/FNR), Operational Latency (response time), Adaptation Time (deployment burden), and Trust (validated automation scales). This framework redefines UX as a multi-layered construct spanning infrastructure and governance, providing a rigorous foundation for evaluating AI systems embedded within complex real-world ecosystems.

cs.HC

A Case Study in Responsible AI-Assisted Video Solutions: Multi-Metric Behavioral Insights in a Public Market Setting

Despite recent advances in Computer Vision and Artificial Intelligence (AI), AI-assisted video solutions have struggled to penetrate real-world urban environments due to significant concerns regarding privacy, ethical risks, and technical challenges like bias and explainability. This work addresses these barriers through a case study in a city-center public market, demonstrating a pathway for the responsible deployment of AI in community spaces. By adopting a user-centric methodology that prioritizes public trust and privacy safeguards, we show that detailed, operationally relevant behavioral insights can be derived from abstract data representations without compromising ethical standards. The study focuses on generating Multi-Metric Behavioral Insights through the extraction of three complementary signals: customer directional flow, dwell duration, and movement patterns. Utilizing human pose detection and complex behavioral analysis - processed through geometric normalization and motion modeling - the system remains robust under tracking fragmentation and occlusion. Data collected over 18 days, spanning routine operations and a festival window from May 2-4, reveals a consistently right-skewed dwell-time behavior. While most visits last approximately 3-4 minutes, peak activity periods increase the mean to roughly 22 minutes. Furthermore, movement analysis indicates uneven circulation, with over 60% of traffic concentrated in approximately 30% of the venue space. By mapping popular thoroughfares and high-traffic storefronts, this case study provides venue managers and business owners with objective, measurable information to optimize foot traffic. Ultimately, these results demonstrate that AI-enabled video solutions can be successfully integrated into urban environments to provide high-fidelity spatial analytics while maintaining strict adherence to privacy and social responsibility.

cs.CY

Multilayer Perceptron Neural Network Model: A Novel Approach for LFP Contrast Sensitivity Tuning

Local field potentials (LFPs) have been demonstrated to be an important measurement to study the activity of a local population of neurons. The response tunings of LFPs have been mostly reported as weaker and broader than spike tunings. Therefore, selecting optimized tuning methods is essential for appropriately evaluating the LFP responses and comparing them with neighboring spiking activity. In this paper, new models for tuning of the contrast response functions (CRFs) are proposed. To this end, luminance contrast-evoked LFP responses recorded in primate primary visual cortex (V1) are first analyzed. Then, supersaturating CRFs are distinguished from linear and saturating CRFs by using monotonicity index (MI). The supersaturated recording data are then identified through static identification methods including multilayer perceptron (MLP) neural network, radial basis function (RBF) neural network, fuzzy model, neuro-fuzzy model, and the local linear model tree (LOLIMOT) algorithm. Our results demonstrate that the MLP neural network, compared to traditional and modified hyperbolic Naka-Rushton functions, exhibits superior performance in tuning the local field potential responses to luminance contrast stimuli, resulting in successful tuning of a significantly higher number of neural recordings of all three types. These results suggest that the MLP neural network model can be used as a novel approach to measure a better fitted contrast sensitivity tuning curve of a population of neurons than other currently used models.

eess.SP

FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models

Research on evaluating and analyzing large language models (LLMs) has been extensive for resource-rich languages such as English, yet their performance in languages such as Persian has received considerably less attention. This paper introduces FarsEval-PKBETS benchmark, a subset of FarsEval project for evaluating large language models in Persian. This benchmark consists of 4000 questions and answers in various formats, including multiple choice, short answer and descriptive responses. It covers a wide range of domains and tasks,including medicine, law, religion, Persian language, encyclopedic knowledge, human preferences, social knowledge, ethics and bias, text generation, and respecting others' rights. This bechmark incorporates linguistics, cultural, and local considerations relevant to the Persian language and Iran. To ensure the questions are challenging for current LLMs, three models -- Llama3-70B, PersianMind, and Dorna -- were evaluated using this benchmark. Their average accuracy was below 50%, meaning they provided fully correct answers to fewer than half of the questions. These results indicate that current language models are still far from being able to solve this benchmark

cs.CL