SearcharxivSearch

arXiv subjects

Karl T. Ulrich

Publications and source records attributed to Karl T. Ulrich.

4 recordsLinked to original sources

AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas

This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs.

cs.AI

Lucky or Good? Outcome Noise, Effective Sample Size, and the Attribution of Skill

When do outcome records carry enough signal to support reliable inferences about skill? When they do not, what should evaluators substitute? The framework answering the first question characterizes any decision domain with two parameters: the noise reflected in each outcome and the effective number of independent outcomes that are available over an observation window. When domains are positioned in a two-dimensional space of noise versus number of outcomes, those in which capital, prestige, and political power are routinely allocated on the basis of realized outcomes (e.g., mutual fund management, venture capital, executive performance) fall in the region where outcome records contain too little signal to support reliable individual-level inferences. Evaluating actors when outcome records are insufficient can be done by adopting the populationlevel empirical validation methods long used in medicine: has the actor adopted the practices that, at the population level, are associated with better outcomes?

econ.GN

Dead Reckoning: Counting Your Customers Who Never Say Goodbye

Firms in non-contractual commerce face the challenge of knowing how many customers they actually have because customers can stop buying without ever saying they have left. Buy-Till-You-Die models address this by estimating each customer's probability of being alive, a quantity called P(alive) and used in every major software tool for dashboards, churn, customer equity, and enterprise valuation. We show this practice confounds two distinct quantities. Within the beta-geometric family, P(alive) is the infinite-horizon limit of an observable family of finite-horizon repeat-purchase probabilities. Every finite-horizon estimate, such as the probability of repeat purchase within 12 months, is a forecast of a verifiable event. The infinite-time limit can only be reached by extrapolation, a customer count obtained by dead reckoning. The implied count is therefore only partially identified: realized returners are the lower bound, and estimation conventions determine the reported point estimate above that. On a seven-year panel of 31,683 customers, specifications with nearly identical observable forecasts estimate the number of alive customers anywhere from 3,654 to 27,734, a factor of 7.6; a default software weighting parameter alone swings the count 42 percent; and five years of later purchases falsify the maximum-likelihood count from below. The patterns replicate on the CDNOW benchmark, with a 2.4x spread. Most of what practice calls miscalibration is instead a category error: summed P(alive) overshoots realized eighteen-month returners by 2.25x, while the same model's own eighteen-month forecast errs by just 1.18x. The remedy is to report an auditable horizon count, estimate return probabilities at a stated horizon, audit them across scoring dates and horizons, recalibrate as cohorts drift, and report the total count, if at all, as an interval rather than a point.

q-fin.GN

The Satoshi Overhang: Why the Bear Case is Bounded

Renewed attention to the identity of Bitcoin's pseudonymous creator has revived an old worry: that the roughly 1.148 million BTC mined by Satoshi and never moved represent a major tail risk for bitcoin. This paper argues that the worry is overstated. The mechanical downside of selling the position is bounded well below the feared collapse, and the outcomes most consistent with sixteen years of observed behavior are not bearish for bitcoin's effective supply. We analyze the position in two ways. First, we model the case of a purely financial holder. Multiple sale scenarios, checked against both a square-root-law estimate and the historical record of large sales, suggest that bitcoin's current market liquidity could absorb a patient multi-year sale with a cumulative price impact centered around 10 to 13 percent relative to a no-sale case. The same arithmetic also links the downside from a surprise sale to the upside from a confirmed burn: both are bounded by the same effective-supply adjustment, so the doom case and the burn-rally case cannot both be large. Second, we consider the preferences implied by the sixteen-year record. Ideological restraint, privacy, already having enough, and preserving the myth all point toward further dormancy, permanent loss of access, or a deliberate burn. A sale or an act of sabotage remains possible, but the record supports it less strongly. Under both approaches, the mechanical bear case is bounded, and the likeliest outcomes are neutral to mildly positive for bitcoin's effective supply. The argument does not rule out transient overshoot or leverage-driven amplification; it bounds the durable repricing the coins themselves can cause.

q-fin.GN