SearcharxivSearch

arXiv subjects

Seth Spielman

Publications and source records attributed to Seth Spielman.

5 recordsLinked to original sources

How people use Copilot for Health

We analyze over 500,000 de-identified health-related conversations with Microsoft Copilot from January 2026 to characterize what people ask conversational AI about health. We develop a hierarchical intent taxonomy of 12 primary categories using privacy-preserving LLM-based classification validated against expert human annotation, and apply LLM-driven topic-clustering for prevalent themes within each intent. Using this taxonomy, we characterize the intents and topics behind health queries, identify who these queries are about, and analyze how usage varies by device and time of day. Five findings stand out. First, nearly one in five conversations involve personal symptom assessment or condition discussion, and even the dominant general information category (40%) is concentrated on specific treatments and conditions, suggesting that this is a lower bound on personal health intent. Second, one in seven of these personal health queries concern someone other than the user, such as a child, a parent, a partner, suggesting that conversational AI can be a caregiving tool, not just a personal one. Third, personal queries about symptoms and emotional health queries increase markedly in the evening and nighttime hours, when traditional healthcare is most limited. Fourth, usage diverges sharply by device: mobile concentrates on personal health concerns, while desktop is dominated by professional and academic work. Fifth, a substantial share of queries focuses on navigating healthcare systems such as finding providers, and understanding insurance, highlighting friction in the delivery of existing healthcare. These patterns have direct implications for platform-specific design, safety considerations, and the responsible development of health AI.

cs.HC

It's About Time: The Temporal and Modal Dynamics of Copilot Usage

We analyze 37.5 million deidentified conversations with Microsoft's Copilot between January and September 2025. Unlike prior analyses of AI usage, we focus not just on what people do with AI, but on how and when they do it. We find that how people use AI depends fundamentally on context and device type. On mobile, health is the dominant topic, which is consistent across every hour and every month we observed - with users seeking not just information but also advice. On desktop, the pattern is strikingly different: work and technology dominate during business hours, with "Work and Career" overtaking "Technology" as the top topic precisely between 8 a.m. and 5 p.m. These differences extend to temporal rhythms: programming queries spike on weekdays while gaming rises on weekends, philosophical questions climb during late-night hours, and relationship conversations surge on Valentine's Day. These patterns suggest that users have rapidly integrated AI into the full texture of their lives, as a work aid at their desks and a companion on their phones.

cs.CY

Large language models can accurately predict searcher preferences

Relevance labels, which indicate whether a search result is valuable to a searcher, are key to evaluating and optimising search systems. The best way to capture the true preferences of users is to ask them for their careful feedback on which results would be useful, but this approach does not scale to produce a large number of labels. Getting relevance labels at scale is usually done with third-party labellers, who judge on behalf of the user, but there is a risk of low-quality data if the labeller doesn't understand user needs. To improve quality, one standard approach is to study real users through interviews, user studies and direct feedback, find areas where labels are systematically disagreeing with users, then educate labellers about user needs through judging guidelines, training and monitoring. This paper introduces an alternate approach for improving label quality. It takes careful feedback from real users, which by definition is the highest-quality first-party gold data that can be derived, and develops an large language model prompt that agrees with that data. We present ideas and observations from deploying language models for large-scale relevance labelling at Bing, and illustrate with data from TREC. We have found large language models can be effective, with accuracy as good as human labellers and similar capability to pick the hardest queries, best runs, and best groups. Systematic changes to the prompts make a difference in accuracy, but so too do simple paraphrases. To measure agreement with real searchers needs high-quality "gold" labels, but with these we find that models produce better labels than third-party workers, for a fraction of the cost, and these labels let us train notably better rankers.

cs.IR

Weather of the Dorm WIFI Ecosystem at the University of Colorado Boulder for Fall Semester 2019 to Spring Semester 2020 a Case Study of WIFI and a Campus Response to the COVID-19 Perturbation

Growing use of network technology in Higher Education means that there has been increasing demand to adapt technology platforms and tools that transform student learning strategies, faculty teaching, research modalities, as well as general operations. Many of the new modalities are necessary for IHE business. In August 2019, we began collecting and analyzing data from the campus WIFI network. A goal of the research was to answer question like what passive sensing of the IHE WIFI might tell us about the dynamics of the WIFI weather in the IHE ecosystem and what does anonymized data tell us about the IHE ecosystem. The analogy with weather prediction seemed appropriate and a viable approach. Starting Fall 2019, data were collected in the observational phase. In the analysis phase, we applied Singular Spectrum Analysis decomposition, to deconstruct WIFI data from dorms, the central campus dining cafeteria, the recreation center, and other buildings on campus. That analysis led to the identification of clusters of buildings that behaved similarly. Just as in the case of models of the weather, a final component of this research was forecasting. We found that weekly forecast of WIFI behavior in the Fall 2019, were straight forward using SSA and seemed to present behavior of a low dimensional dynamical system. However, in Spring 2020, and the COVID perturbation, the campus ecosystem received a shock and data show that the campus changed very quickly. We found that as the campus moved to conduct remote learning, teaching, the closure of research labs, and the edict to work remotely, SSA forecasting techniques not trained on the Spring 2020, data after the shock, performed poorly. While SSA forecasting trained on a portion of the data did better.

cs.CY

Exploring the Usage of Online Food Delivery Data for Intra-Urban Job and Housing Mobility Detection and Characterization

Human mobility plays a critical role in urban planning and policy-making. However, at certain spatial and temporal resolutions, it is very challenging to track, for example, job and housing mobility. In this study, we explore the usage of a new modality of dataset, online food delivery data, to detect job and housing mobility. By leveraging millions of meal orders from a popular online food ordering and delivery service in Beijing, China, we are able to detect job and housing moves at much higher spatial and temporal resolutions than using traditional data sources. Popular moving seasons and origins/destinations can be well identified. More importantly, we match the detected moves to both macro- and micro-level factors so as to characterize job and housing dynamics. Our findings suggest that commuting distance is a major factor for job and housing mobility. We also observe that: (1) For home movers, there is a trade-off between lower housing cost and shorter commuting distance given the urban spatial structure; (2) For job hoppers, those who frequently work overtime are more likely to reduce their working hours by switching jobs. While this new modality of dataset has its limitations, we believe that ensemble approaches would be promising, where a mash-up of multiple datasets with different characteristic limitations can provide a more comprehensive picture of job and housing dynamics. Our work demonstrates the effectiveness of utilizing food delivery data to detect and analyze job and housing mobility, and contributes to realizing the full potential of ensemble-based approaches.

cs.CY