SearcharxivSearch

arXiv subjects

Emily Wang

Publications and source records attributed to Emily Wang.

4 recordsLinked to original sources

From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation

Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decision to the LLM, which tends to perform according to its built-in capabilities even when instructed to simulate a student with low mastery. As a result, these approaches may have difficulty distinguishing students with low and high levels of mastery. We demonstrate this limitation using 379 College Board-calibrated SAT Algebra items and five archetypal mastery profiles. Three LLMs from three vendors (Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini) achieve 96.8-100% accuracy across all profiles. To address this limitation, we introduce a method grounded in a Stochastic Student Knowledge Graph (SSKG). A curriculum knowledge graph (CKG) is extracted from an open algebra textbook, and each SAT solution is decomposed into a chain of required triples. The SSKG assigns a mastery probability to each triple, which is sampled to determine question correctness. An LLM then generates a first-person rationale consistent with the outcome. The simulation reduces accuracy to 44.1-85.2% across profiles and produces a clear monotone mastery gradient.

cs.AI

Uncertainty-Aware Graph Self-Training with Expectation-Maximization Regularization

In this paper, we propose a novel \emph{uncertainty-aware graph self-training} approach for semi-supervised node classification. Our method introduces an Expectation-Maximization (EM) regularization scheme to incorporate an uncertainty mechanism during pseudo-label generation and model retraining. Unlike conventional graph self-training pipelines that rely on fixed pseudo-labels, our approach iteratively refines label confidences with an EM-inspired uncertainty measure. This ensures that the predictive model focuses on reliable graph regions while gradually incorporating ambiguous nodes. Inspired by prior work on uncertainty-aware self-training techniques~\cite{wang2024uncertainty}, our framework is designed to handle noisy graph structures and feature spaces more effectively. Through extensive experiments on several benchmark graph datasets, we demonstrate that our method outperforms strong baselines by a margin of up to 2.5\% in accuracy while maintaining lower variance in performance across multiple runs.

cs.LG

Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People

Navigating unfamiliar environments presents significant challenges for blind and low-vision (BLV) individuals. In this work, we construct a dataset of images and goals across different scenarios such as searching through kitchens or navigating outdoors. We then investigate how grounded instruction generation methods can provide contextually-relevant navigational guidance to users in these instances. Through a sighted user study, we demonstrate that large pretrained language models can produce correct and useful instructions perceived as beneficial for BLV users. We also conduct a survey and interview with 4 BLV users and observe useful insights on preferences for different instructions based on the scenario.

cs.CL

Star formation and UV colors of the brightest Cluster Galaxies in the representative XMM-Newton Cluster Structure Survey

We present UV broadband photometry and optical emission-line measurements for a sample of 32 Brightest Cluster Galaxies (BCGs) in clusters of the Representative XMM-Newton Cluster Structure Survey (REXCESS) with z = 0.06-0.18. The REXCESS clusters, chosen to study scaling relations in clusters of galaxies, have X-ray measurements of high quality. The trends of star formation and BCG colors with BCG and host properties can be investigated with this sample. The UV photometry comes from the XMM Optical Monitor, supplemented by existing archival GALEX photometry. We detected Hαand forbidden line emission in 7 (22%) of these BCGs, in optical spectra. All of the emission-line BCGs occupy clusters classified as cool cores, for an emission-line incidence rate of 70% for BCGs in cool core clusters. Significant correlations between the Hαequivalent widths, excess UV production in the BCG, and the presence of dense, X-ray bright intracluster gas with a short cooling time are seen, including the fact that all of the Hαemitters inhabit systems with short central cooling times and high central ICM densities. Estimates of the star formation rates based on Hαand UV excesses are consistent with each other in these 7 systems, ranging from 0.1-8 solar masses per year. The incidence of emission-line BCGs in the REXCESS sample is intermediate, somewhat lower than in other X-ray selected samples (-35%), and somewhat higher than but statistically consistent with optically selected, slightly lower redshift BCG samples (-10-15%). The UV-optical colors (UVW1-R-4.7\pm0.3) of REXCESS BCGs without strong optical emission lines are consistent with those predicted from templates and observations of ellipticals dominated by old stellar populations. We see no trend in UV-optical colors with optical luminosity, R-K color, X-ray temperature, redshift, or offset between X-ray centroid and X-ray peak ( ).

astro-ph.CO