SearcharxivSearch

arXiv subjects

Sara Khosravi

Publications and source records attributed to Sara Khosravi.

10 recordsLinked to original sources

Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($ρ= 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\ GPT-4o and DeepSeek-R1 vs.\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\%), whereas DeepSeek-R1 trailed V3 (42.9\%), and inter-reader agreement was near chance (Krippendorff's $α= 0.070$), with within-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding. The judge, by contrast, favored reasoning outputs in 89.0\% of dimension-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity. These results suggest that development-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point-estimate judges can obscure the heterogeneity in subjective human evaluation.

cs.LG

Practical Policy Distillation for Reinforcement Learning in Radio Access Networks

Adopting artificial intelligence (AI) in radio access networks (RANs) presents several challenges, including limited availability of link-level measurements (e.g., CQI reports), stringent real-time processing constraints (e.g., sub-1 ms per TTI), and network heterogeneity (different spectrum bands, cell types, and vendor equipment). A critical yet often overlooked barrier lies in the computational and memory limitations of RAN baseband hardware, particularly in legacy 4th Generation (4G) systems, which typically lack on-chip neural accelerators. As a result, only lightweight AI models (under 1 Mb and sub-100~μs inference time) can be effectively deployed, limiting both their performance and applicability. However, achieving strong generalization across diverse network conditions often requires large-scale models with substantial resource demands. To address this trade-off, this paper investigates policy distillation in the context of a reinforcement learning-based link adaptation task. We explore two strategies: single-policy distillation, where a scenario-agnostic teacher model is compressed into one generalized student model; and multi-policy distillation, where multiple scenario-specific teachers are consolidated into a single generalist student. Experimental evaluations in a high-fidelity, 5th Generation (5G)-compliant simulator demonstrate that both strategies produce compact student models that preserve the teachers' generalization capabilities while complying with the computational and memory limitations of existing RAN hardware.

cs.LG

AI Agents in Drug Discovery

Artificial intelligence (AI) agents are emerging as transformative tools in drug discovery, with the ability to autonomously reason, act, and learn through complicated research workflows. Building on large language models (LLMs) coupled with perception, computation, action, and memory tools, these agentic AI systems could integrate diverse biomedical data, execute tasks, carry out experiments via robotic platforms, and iteratively refine hypotheses in closed loops. We provide a conceptual and technical overview of agentic AI architectures, ranging from ReAct and Reflection to Supervisor and Swarm systems, and illustrate their applications across key stages of drug discovery, including literature synthesis, toxicity prediction, automated protocol generation, small-molecule synthesis, drug repurposing, and end-to-end decision-making. To our knowledge, this represents the first comprehensive work to present real-world implementations and quantifiable impacts of agentic AI systems deployed in operational drug discovery settings. Early implementations demonstrate substantial gains in speed, reproducibility, and scalability, compressing workflows that once took months into hours while maintaining scientific traceability. We discuss the current challenges related to data heterogeneity, system reliability, privacy, and benchmarking, and outline future directions towards technology in support of science and translation.

cs.LG

Beam Alignment Using Trajectory Information in Mobile Millimeter-wave Networks

Millimeter-wave and terahertz systems rely on beamforming/combining codebooks to determine the best beam directions during the initial access and data transmission. Existing approaches suffer from large codebook sizes and high beam searching overhead in the presence of mobile devices. To address this issue, we utilize the similarity of the channel in adjacent locations to divide the user trajectory into a set of separate regions and maintain a set of candidate beams for each region in a database. Due to the tradeoff between the number of regions and the signalling overhead, i.e., the greater number of regions results in a higher signal-to-noise ratio (SNR) but also a larger signalling overhead for the database, we propose an optimization framework to find the minimum number of regions based on the trajectory of a mobile device. Using a ray tracing tool, we demonstrate that the proposed method provides high SNR while being more robust to the location information accuracy in comparison to the lookup table baseline and fixed size region baseline.

cs.IT

Reinforcement Learning-based Joint Handover and Beam Tracking in Millimeter-wave Networks

In this paper, we develop an algorithm for joint handover and beam tracking in millimeter-wave (mmWave) networks. The aim is to provide a reliable connection in terms of the achieved throughput along the trajectory of the mobile user while preventing frequent handovers. We model the association problem as an optimization problem and propose a reinforcement learning-based solution. Our approach learns whether and when beam tracking and handover should be performed and chooses the target base stations. In the case of beam tracking, we propose a tracking algorithm based on measuring a small spatial neighbourhood of the optimal beams in the previous time slot. Simulation results in an outdoor environment show the superior performance of our proposed solution in achievable throughput and the number of handovers needed in comparison to a multi-connectivity baseline and a learning-based handover baseline.

eess.SY

Location-Aided Beamforming in Mobile Millimeter-Wave Networks

Due to the large bandwidth available, millimeter-Wave (mmWave) bands are considered a viable opportunity to significantly increase the data rate in cellular and wireless networks. Nevertheless, the need for beamforming and directional communication between the transmitter and the receiver increases the complexity of the channel estimation and link establishment phase. Location-aided beamforming approaches have the potential to enable fast link establishment in mmWave networks. However, these are often very sensitive to location errors. In this work, we propose a beamforming algorithm based on tracking spatial correlation of the available strong paths between the transmitter and the receiver. We show that our method is robust to uncertainty in location information, i.e., location error and can provide a reliable connection to a moving user along a trajectory. The numerical results show that our approach outperforms benchmarks on various levels of error in the location information accuracy. The gain is more prominent in high location error scenarios.

eess.SY

Learning Enhancement in Higher Education with Wearable Technology

Wearable technologies have traditionally been used to measure and monitor vital human signs for well-being and healthcare applications. However, there is a growing interest in using and deploying these technologies to facilitate teaching and learning, particularly in a higher education environment. The aim of this paper is therefore to systematically review the range of wearable devices that have been used for enhancing the teaching and delivery of engineering curricula in higher education. Moreover, we compare the advantages and disadvantages of these devices according to the location in which they are worn on the human body. According to our survey, wearable devices for enhanced learning have mainly been worn on the head (e.g. eyeglasses), wrist (e.g. watches) and chest (e.g. electrocardiogram patch). In fact, among those locations, head-worn devices enable better student engagement with the learning materials, improved student attention as well as higher spatial and visual awareness. We identify the research questions and discuss the research inclusion and exclusion criteria to present the challenges faced by researchers in implementing learning technologies for enhanced engineering education. Furthermore, we provide recommendations on using wearable devices to improve the teaching and learning of engineering courses in higher education.

cs.HC

Learning-based Load Balancing Handover in Mobile Millimeter Wave Networks

Millimeter-wave (mmWave) communication is a promising solution to the high data rate demands in the upcoming 5G and beyond communication networks. When it comes to supporting seamless connectivity in mobile scenarios, resource and handover management are two of the main challenges in mmWave networks. In this paper, we address these two problems jointly and propose a learning-based load balancing handover in multi-user mobile mmWave networks. Our handover algorithm selects a backup base station and allocates the resource to maximize the sum rate of all the users while ensuring a target rate threshold and preventing excessive handovers. We model the user association as a non-convex optimization problem. Then, by applying a deep deterministic policy gradient (DDPG) method, we approximate the solution of the optimization problem. Through simulations, we show that our proposed algorithm minimizes the number of the events where a user's rate is less than its minimum rate requirement and minimizes the number of handovers while increasing the sum rate of all users.

eess.SP

Learning-based Handover in Mobile Millimeter-wave Networks

Millimeter-wave (mmWave) communication is considered as a key enabler of ultra-high data rates in the future cellular and wireless networks. The need for directional communication between base stations (BSs) and users in mmWave systems, that is achieved through beamforming, increases the complexity of the channel estimation. Moreover, in order to provide better coverage, dense deployment of BSs is required which causes frequent handovers and increased association overhead. In this paper, we present an approach that jointly addresses the beamforming and handover problems. Our solution entails an efficient beamforming method with a minimum number of pilots and a learning-based handover method supporting mobile scenarios. We use reinforcement learning algorithm to learn the optimal choices of the backup BSs in different locations of a mobile user. We show that our method provides high rate and reliability in all locations of the user's trajectory with a minimal number of handovers. Simulation results in an outdoor environment based on geometric mmWave channel modeling and real building map data show the superior performance of our proposed solution in achievable instantaneous rate and trajectory rate.

eess.SP

Efficient Beamforming for Mobile mmWave Networks

We design a lightweight beam-searching algorithm for mobile millimeter-wave systems. We construct and maintain a set of path skeletons, i.e., potential paths between a user and the serving base station to substantially expedite the beam-searching process. To exploit the spatial correlations of the channels, we propose an efficient algorithm that measures the similarity of the skeletons and re-executes the beam-searching procedure only when the old one becomes obsolete. We identify and optimize several tradeoffs between: i) the beam-searching overhead and the instantaneous rate of the users, and ii) the number of users and the update overhead of the path skeletons. Simulation results in an outdoor environment with real building map data show that the proposed method can significantly improve the performance of beam-searching in terms of latency, energy consumption and achievable throughout.

eess.SP