SearcharxivSearch

arXiv subjects

Andrew Kim

Publications and source records attributed to Andrew Kim.

10 recordsLinked to original sources

Google's AI & Economy ATLAS v1.0: Mapping Gemini Usage in the Economy

This paper introduces the AI & Economy ATLAS (Activity, Task, Landscape, and Adoption Study), an ongoing economic research initiative using Google AI usage data. The first iteration of ATLAS is built on 15 million de-identified interactions across the Gemini App, Google AI Mode, and Gemini API. Using privacy-preserving algorithms as well as established and bespoke classification methods, we map AI usage to over 800 occupations, 4000 tasks, 300 household activities, 150 countries, and 140 languages. We then make a number of observations on what the data reveals about AI's diffusion, and its usage at work and in day-to-day life. In the workplace, we show that while AI adoption spans occupations covering just above 88% of US employment, penetration remains shallow and overwhelmingly collaborative in nature, with end-to-end task automation limited in scope. Outside of work, AI spans activities making up about 98% of Americans' non-sleep time, with disproportionately high use in high-friction tasks such as engaging with government and professional service providers, likely delivering economic value that standard national accounts may miss. Globally, adoption scales with national wealth and has broad linguistic distribution, with English queries representing only around a third of volume. As we build upon ATLAS and expand its scope and capabilities, we will continue to provide large-scale empirical evidence to inform the public, policy and academic questions about the ongoing AI transformation.

econ.GN

Causal Risk Minimization for High-Dimensional Treatments

Predicting the effect of interventions with many possible variations, e.g., therapeutic content that affects mental health outcomes or an earnings call transcript that drives movement in share price, is useful across several domains. However, classical causal estimators tend to assume that all possible interventions are observed, which is infeasible when interventions vary widely, for instance, in the space of all text strings. We adapt a well-known approach of recasting causal inference as a learning problem, to address high-dimensional treatment spaces. Specifically, under standard assumptions like no unobserved confounding, we show that causal error decomposes into a series of moment-balancing errors of increasing order, and design objectives that directly improve causal estimation. We also show how to project the effect of a high-dimensional treatment onto lower-dimensional treatment attributes, which allows a single model to answer several causal questions without additional attribute-specific training. We empirically evaluate our estimators in settings with high-dimensional continuous, discrete, and text treatments, the last of which used a semi-synthetic dataset of Amazon Reviews. Our experiments demonstrate the benefit of higher-order balance error optimization and competitive performance of projected causal estimates with attribute-specific estimators.

cs.LG

Biphasic Meniscus Coating for Scalable and Material Efficient Quantum Dot Films

Colloidal quantum dots (cQDs) have emerged as a cornerstone of next-generation optoelectronics, offering unparalleled spectral tunability and solution-processability. However, the transition from laboratory-scale devices to sustainable industrial manufacturing is fundamentally hindered by spin-coating workflows, which are intrinsically wasteful and restricted to planar geometries. These limitations are particularly acute for high-performance cQDs containing regulated elements such as lead, cadmium, or mercury, where poor material utilization exacerbates both environmental burden and cost. Here we report a biphasic dip-coating strategy that redefines the material efficiency of nanocrystal film fabrication. By utilizing an immiscible underlayer to displace ~88% of the active reservoir volume, we demonstrate a deposition geometry that decouples material consumption from total precursor volume. Infrared PbS photodetectors fabricated via this approach maintain their performance against spin-coated benchmarks while reducing ink consumption by up to 20-fold. Our technoeconomic analysis reveals that this biphasic architecture achieves cost parity at film thicknesses an order of magnitude lower than conventional monophasic dip-coating. Our results establish a low-waste framework for solution-processed materials, providing a viable pathway for the resource-efficient manufacturing of optoelectronic devices.

cond-mat.mtrl-sci

Machine Learning-guided accelerated discovery of structure-property correlations in lean magnesium alloys for biomedical applications

Magnesium alloys are emerging as promising alternatives to traditional orthopedic implant materials thanks to their biodegradability, biocompatibility, and impressive mechanical characteristics. However, their rapid in-vivo degradation presents challenges, notably in upholding mechanical integrity over time. This study investigates the impact of high-temperature thermal processing on the mechanical and degradation attributes of a lean Mg-Zn-Ca-Mn alloy, ZX10. Utilizing rapid, cost-efficient characterization methods like X-ray diffraction and optical, we swiftly examine microstructural changes post-thermal treatment. Employing Pearson correlation coefficient analysis, we unveil the relationship between microstructural properties and critical targets (properties): hardness and corrosion resistance. Additionally, leveraging the least absolute shrinkage and selection operator (LASSO), we pinpoint the dominant microstructural factors among closely correlated variables. Our findings underscore the significant role of grain size refinement in strengthening and the predominance of the ternary Ca2Mg6Zn3 phase in corrosion behavior. This suggests that achieving an optimal blend of strength and corrosion resistance is attainable through fine grains and reduced concentration of ternary phases. This thorough investigation furnishes valuable insights into the intricate interplay of processing, structure, and properties in magnesium alloys, thereby advancing the development of superior biodegradable implant materials.

cond-mat.mtrl-sci

The EDGE-CALIFA Survey: An Extragalactic Database for Galaxy Evolution Studies

The EDGE-CALIFA survey provides spatially resolved optical integral field unit (IFU) and CO spectroscopy for 125 galaxies selected from the CALIFA Data Release 3 sample. The Extragalactic Database for Galaxy Evolution (EDGE) presents the spatially resolved products of the survey as pixel tables that reduce the oversampling in the original images and facilitate comparison of pixels from different images. By joining these pixel tables to lower dimensional tables that provide radial profiles, integrated spectra, or global properties, it is possible to investigate the dependence of local conditions on large-scale properties. The database is freely accessible and has been utilized in several publications. We illustrate the use of this database and highlight the effects of CO upper limits on the inferred slopes of the local scaling relations between stellar mass, star formation rate (SFR), and H$_2$ surface densities. We find that the correlation between H$_2$ and SFR surface density is the tightest among the three relations.

astro-ph.GA

Comparative Analysis of Plastid Genomes Using Pangenome Research ToolKit (PGR-TK)

Plastid genomes (plastomes) of angiosperms are of great interest among biologists. High-throughput sequencing is making many such genomes accessible, increasing the need for tools to perform rapid comparative analysis. This exploratory analysis investigates whether the Pangenome Research Tool Kit (PGR-TK) is suitable for analyzing plastomes. After determining the optimal parameters for this tool on plastomes, we use it to compare sequences from each of the genera - Magnolia, Solanum, Fragaria and Cotoneaster, as well as a combined set from 20 rosid genera. PGR-TK recognizes large-scale plastome structures, such as the inverted repeats, among combined sequences from distant rosid families. If the plastid genomes are rotated to the same starting point, it also correctly groups different species from the same genus together in a generated cladogram. The visual approach of PGR-TK provides insights into genome evolution without requiring gene annotations.

q-bio.GN

Conversational Swarm Intelligence, a Pilot Study

Conversational Swarm Intelligence (CSI) is a new method for enabling large human groups to hold real-time networked conversations using a technique modeled on the dynamics of biological swarms. Through the novel use of conversational agents powered by Large Language Models (LLMs), the CSI structure simultaneously enables local dialog among small deliberative groups and global propagation of conversational content across a larger population. In this way, CSI combines the benefits of small-group deliberative reasoning and large-scale collective intelligence. In this pilot study, participants deliberating in conversational swarms (via text chat) (a) produced 30% more contributions (p<0.05) than participants deliberating in a standard centralized chat room and (b) demonstrated 7.2% less variance in contribution quantity. These results indicate that users contributed more content and participated more evenly when using the CSI structure.

cs.HC

Change-Point Analysis of Cyberbullying-Related Twitter Discussions During COVID-19

Due to the outbreak of COVID-19, users are increasingly turning to online services. An increase in social media usage has also been observed, leading to the suspicion that this has also raised cyberbullying. In this initial work, we explore the possibility of an increase in cyberbullying incidents due to the pandemic and high social media usage. To evaluate this trend, we collected 454,046 cyberbullying-related public tweets posted between January 1st, 2020 -- June 7th, 2020. We summarize the tweets containing multiple keywords into their daily counts. Our analysis showed the existence of at most one statistically significant changepoint for most of these keywords, which were primarily located around the end of March. Almost all these changepoint time-locations can be attributed to COVID-19, which substantiates our initial hypothesis of an increase in cyberbullying through analysis of discussions over Twitter.

cs.SI

Quantifying Susceptibility to Spear Phishing in a High School Environment Using Signal Detection Theory

Spear phishing is a deceptive attack that uses social engineering to obtain confidential information through targeted victimization. It is distinguished by its use of social cues and personalized information to target specific victims. Previous work on resilience to spear phishing has focused on convenience samples, with a disproportionate focus on students. In contrast, here, we report on an evaluation of a high school community. We engaged 57 high school students and faculty members (12 high school students, 45 staff members) as participants in research utilizing signal detection theory (SDT). Through scenario-based analysis, participants tasked with distinguishing phishing emails from authentic emails. The results revealed an overconfidence bias in self-detection from the participants, regardless of their technical background. These findings are critical for evaluating the decision-making of underrepresented populations and protecting people from potential spear phishing attacks by examining human susceptibility.

cs.CR

All About Phishing: Exploring User Research through a Systematic Literature Review

Phishing is a well-known cybersecurity attack that has rapidly increased in recent years. It poses legitimate risks to businesses, government agencies, and all users due to sensitive data breaches, subsequent financial and productivity losses, and social and personal inconvenience. Often, these attacks use social engineering techniques to deceive end-users, indicating the importance of user-focused studies to help prevent future attacks. We provide a detailed overview of phishing research that has focused on users by conducting a systematic literature review of peer-reviewed academic papers published in ACM Digital Library. Although published work on phishing appears in this data set as early as 2004, we found that of the total number of papers on phishing (N = 367) only 13.9% (n = 51) focus on users by employing user study methodologies such as interviews, surveys, and in-lab studies. Even within this small subset of papers, we note a striking lack of attention to reporting important information about methods and participants (e.g., the number and nature of participants), along with crucial recruitment biases in some of the research.

cs.CR