SearcharxivSearch

arXiv subjects

Panos Ipeirotis

Publications and source records attributed to Panos Ipeirotis.

9 recordsLinked to original sources

Natural Language Interfaces for Databases: What Changes for SQL-Literate Users?

Natural Language Interfaces for Databases (NLIDBs) let users query data in everyday language instead of SQL, and recent systems translate those questions accurately. Accuracy says little about the work of querying: does an NLIDB remove that work or only reallocate it? We report a mixed-method, between-subjects study comparing SQL-LLM, a GPT-4o-backed NLIDB, with Snowflake, a traditional SQL platform. Twenty SQL-literate professionals and graduate students, ten per interface, each completed 12 tasks drawn from the BIRD benchmark. SQL-LLM cut completion time per query by about 31%, but the speedup did not buy accuracy: graded against the BIRD gold answers, SQL-LLM users were correct on 46% of queries versus 64% for Snowflake. Two analysts independently coded the think-aloud sessions to locate the effort. SQL-LLM users left schema navigation to the model and spent it verifying that the generated SQL matched their intent; Snowflake users explored the schema and built syntax by hand. Per-minute rates of the coded behaviors did not differ between interfaces. The interface reallocated the work of querying rather than removing it: users still had to verify the generated SQL, and an NLIDB that hid it would remove the step that let them trust the answer.

cs.DB

AI Strategy: How to Choose What AI Product to Implement

Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite decisions. At the residential real-estate brokerage Compass, one AI product (Likely-to-Sell recommendations) flagged sales outreach opportunities and went on to account for nine figures in annual gross commission revenue. Another championed AI product (a Time-on-Market pricing tool) was rightly shelved. A simple ROI estimate could not distinguish the two. We present expected ROI (eROI), a framework that decomposes each bet into three components and rates them separately: Value if Successful, Likelihood of Success, and Investment Required. Each maps to a question executives can answer before building: How valuable would it be if it worked? How likely is it to work? And what would it cost to implement? Separating the three breaks a common catch-22: teams cannot estimate ROI until they know whether a project will work, yet cannot know whether it will work without building it. Judging Value if Successful on its own dissolves the loop, letting a team argue that a product would be valuable if it worked while it weighs how likely that is. The framework also asks, before ranking anything, whether there are enough good ideas on the table. After ranking, it guides assembling a portfolio of bets rather than funding only the single top-ranked project. We illustrate eROI on Compass's candidate AI products. Precise ROI estimates are hard to make given the inherent uncertainty of AI projects. Coarse business-level ratings of the three components are enough to tell strong bets from weak ones.

cs.CY

Scalable and Personalized Oral Assessments Using Voice AI

Written work no longer certifies that a student understands it: a polished analysis now says little about who did the thinking. Oral examinations restore that evidentiary link, but they have never scaled, because conducting and grading them is expensive. We report on a system in which voice AI conducts a personalized oral exam and a council of three large language models (LLMs) grades the transcript, each model scoring independently and then revising after reading the others. Across two undergraduate cohorts at NYU Stern (36 students in Fall 2025, 37 in Spring 2026), a voice subscription covered all speaking time and grading stayed under one dollar per exam. The deployments yield five practical engineering lessons that should generalize wherever understanding must be tested under questioning, from job interviews to professional certification.

cs.CY

Theoretical Foundations of δ-margin Majority Voting

In high-stakes ML applications such as fraud detection, medical diagnostics, and content moderation, practitioners rely on consensus-based approaches to control prediction quality. A particularly valuable technique -- δδδ-margin majority voting -- collects votes sequentially until one label exceeds alternatives by a threshold δδδ, offering stronger confidence than simple majority voting. Despite widespread adoption, this approach has lacked rigorous theoretical foundations, leaving practitioners reliant on heuristics for key metrics like expected accuracy and cost. This paper establishes a comprehensive theoretical framework for δδδ-margin majority voting by formulating it as an absorbing Markov chain and leveraging Gambler's Ruin theory. Our contributions form a practical \emph{design calculus} for δδδ-margin voting: (1)~Closed-form expressions for consensus accuracy, expected voting duration, variance, and the stopping-time PMF, enabling model-based design rather than trial-and-error. (2)~A Bayesian extension handling uncertainty in worker accuracy, supporting real-time monitoring of expected quality and cost as votes arrive, with single-Beta and mixture-of-Betas priors. (3)~Cost-calibration methods for achieving equivalent quality across worker pools with different accuracies and for setting payment rates accordingly. We validate our predictions on two real-world datasets, demonstrating close agreement between theory and observed outcomes. The framework gives practitioners a rigorous toolkit for designing δδδ-margin voting processes, replacing ad-hoc experimentation with model-based design where quality control and cost transparency are essential.

stat.AP

Learning to Pay Attention: Unsupervised Modeling of Attentive and Inattentive Respondents in Survey Data

The integrity of behavioral and social-science surveys depends on detecting inattentive respondents who provide random or low-effort answers. Traditional safeguards, such as attention checks, are often costly, reactive, and inconsistent. We propose a unified, label-free framework for inattentiveness detection that scores response coherence using complementary unsupervised views: geometric reconstruction (Autoencoders) and probabilistic dependency modeling (Chow-Liu trees). While we introduce a "Percentile Loss" objective to improve Autoencoder robustness against anomalies, our primary contribution is identifying the structural conditions that enable unsupervised quality control. Across nine heterogeneous real-world datasets, we find that detection effectiveness is driven less by model complexity than by survey structure: instruments with coherent, overlapping item batteries exhibit strong covariance patterns that allow even linear models to reliably separate attentive from inattentive respondents. This reveals a critical ``Psychometric-ML Alignment'': the same design principles that maximize measurement reliability (e.g., internal consistency) also maximize algorithmic detectability. The framework provides survey platforms with a scalable, domain-agnostic diagnostic tool that links data quality directly to instrument design, enabling auditing without additional respondent burden.

cs.HC

Algorithmic Hiring and Diversity: Reducing Human-Algorithm Similarity for Better Outcomes

Algorithmic tools are increasingly used in hiring to improve fairness and diversity, often by enforcing constraints such as gender-balanced candidate shortlists. However, we show theoretically and empirically that enforcing equal representation at the shortlist stage does not necessarily translate into more diverse final hires, even when there is no gender bias in the hiring stage. We identify a crucial factor influencing this outcome: the correlation between the algorithm's screening criteria and the human hiring manager's evaluation criteria -- higher correlation leads to lower diversity in final hires. Using a large-scale empirical analysis of nearly 800,000 job applications across multiple technology firms, we find that enforcing equal shortlists yields limited improvements in hire diversity when the algorithmic screening closely mirrors the hiring manager's preferences. We propose a complementary algorithmic approach designed explicitly to diversify shortlists by selecting candidates likely to be overlooked by managers, yet still competitive according to their evaluation criteria. Empirical simulations show that this approach significantly enhances gender diversity in final hires without substantially compromising hire quality. These findings highlight the importance of algorithmic design choices in achieving organizational diversity goals and provide actionable guidance for practitioners implementing fairness-oriented hiring algorithms.

cs.LG

On the Predictability of Utilizing Rank Percentile to Evaluate Scientific Impact

Bibliographic metrics are commonly utilized for evaluation purposes within academia, often in conjunction with other metrics. These metrics vary widely across fields and change with the seniority of the scholar; consequently, the only way to interpret these values is by comparison with other academics within the same field who are of similar seniority. Among the field- and time- normalized indicators, rank percentile has grown in popularity, and it is preferred over other types of indicators. In this paper, we propose and justify a novel rank percentile indicator for scholars. Furthermore, we emphasize on the time factor that is built into the rank percentile, and we demonstrate that the rank percentile is highly predictable. The publication percentile is highly stable over time, while the scholar percentile exhibits short-term stability and can be predicted via a simple linear regression model. More advanced models that utilize extensive lists of features offer slightly superior performance; however, the simplicity and interpretability of the simple model impose significant advantages over the additional complexity of other models.

cs.DL

What do crowd workers think about creative work?

Crowdsourcing platforms are a powerful and convenient means for recruiting participants in online studies and collecting data from the crowd. As information work is being more and more automated by Machine Learning algorithms, creativity $-$ that is, a human's ability for divergent and convergent thinking $-$ will play an increasingly important role on online crowdsourcing platforms. However, we lack insights into what crowd workers think about creative work. In studies in Human-Computer Interaction (HCI), the ability and willingness of the crowd to participate in creative work seems to be largely unquestioned. Insights into the workers' perspective are rare, but important, as they may inform the design of studies with higher validity. Given that creativity will play an increasingly important role in crowdsourcing, it is imperative to develop an understanding of how workers perceive creative work. In this paper, we summarize our recent worker-centered study of creative work on two general-purpose crowdsourcing platforms (Amazon Mechanical Turk and Prolific). Our study illuminates what creative work is like for crowd workers on these two crowdsourcing platforms. The work identifies several archetypal types of workers with different attitudes towards creative work, and discusses common pitfalls with creative work on crowdsourcing platforms.

cs.HC

Creativity on Paid Crowdsourcing Platforms

General-purpose crowdsourcing platforms are increasingly being harnessed for creative work. The platforms' potential for creative work is clearly identified, but the workers' perspectives on such work have not been extensively documented. In this paper, we uncover what the workers have to say about creative work on paid crowdsourcing platforms. Through a quantitative and qualitative analysis of a questionnaire launched on two different crowdsourcing platforms, our results revealed clear differences between the workers on the platforms in both preferences and prior experience with creative work. We identify common pitfalls with creative work on crowdsourcing platforms, provide recommendations for requesters of creative work, and discuss the meaning of our findings within the broader scope of creativity-oriented research. To the best of our knowledge, we contribute the first extensive worker-oriented study of creative work on paid crowdsourcing platforms.

cs.HC