Searcharxiv⌕ Search

arXiv subjects

Boi Mai Quach

Publications and source records attributed to Boi Mai Quach.

4 recordsLinked to original sources

Pseudo-Incrementality Testing: Measuring Advertising Lift from Naturally Occurring Interventions

We develop a method for measuring the incremental effect of advertising when randomized exper- iments are unavailable. Firms generate abrupt interventions in their own marketing as a byproduct of operations: budgets are cut, channels launch, programs pause. We propose a two-stage proce- dure that treats these events as quasi-experiments. The first stage discovers and dates interventions by exact Bayesian run-length inference on marketing activity series. The second stage runs a causal impact analysis against a variance-constrained structural time series counterfactual, with simulation-based inference, five qualification conditions, and a closed-form power bound. We validate the procedure on the two experimental benchmarks that Meta released with its GeoLift framework. The method identifies the day of intervention exactly in both cases. On the experiment where advertising was removed, it recovers the effect within 6.2 percent of the experimental esti- mate (implied return of 1.50 versus 1.60). On the experiment where advertising was added, its 90 percent interval covers the experimental estimate and excludes zero. Of six estimators evaluated on the same released data, ours is the only one that requires no geographic panel and produces intervals that are both correct on both experiments and narrow enough to act on.

stat.ME↗

Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.

cs.CL↗

Decoding EEG Signals to Explore Next-Word Predictability in the Human Brain

Humans invented reading and have passed down this complex skill across generations through language. This study provides empirical evidence of the neural mechanisms underlying bottom-up (related to high-order linguistic structure) and top-down (related to next-word predictability) processes, which interact to guide comprehension during reading. While previous studies have focused on either the N400 effects of predictability or lexical categories, research on how predictability influences N400 responses across different lexical categories is limited, mainly due to constraints in publicly available datasets. Here, we examine how predictability influences brain responses, recorded at millisecond resolution using electroencephalography (EEG), with a focus on the N400 time window (300-500 ms post-stimulus) across different lexical and grammatical categories. Our results indicate that significant differences in N400 responses between high and low cloze probability levels were more pronounced for content words than function words. Among the two primary content categories, verbs exhibited greater N400 differences than nouns, while nouns carried more distinct information about their predictability than verbs. Moreover, we demonstrate that the decoding technique is more effective than the event-related potential (ERP) traditional analysis in capturing more detailed and distinct representations of cognitive processes over time.

cs.CL↗

Causal-driven attribution (CDA): Estimating channel influence without user-level data

Attribution modelling lies at the heart of marketing effectiveness, yet most existing approaches depend on user-level path data, which are increasingly inaccessible due to privacy regulations and platform restrictions. This paper introduces a Causal-Driven Attribution (CDA) framework that infers channel influence using only aggregated impression-level data, avoiding any reliance on user identifiers or click-path tracking. CDA integrates temporal causal discovery (using PCMCI) with causal effect estimation via a Structural Causal Model to recover directional channel relationships and quantify their contributions to conversions. Using large-scale synthetic data designed to replicate real marketing dynamics, we show that CDA achieves an average relative RMSE of 9.50% when given the true causal graph, and 24.23% when using the predicted graph, demonstrating strong accuracy under correct structure and meaningful signal recovery even under structural uncertainty. CDA captures cross-channel interdependencies while providing interpretable, privacy-preserving attribution insights, offering a scalable and future-proof alternative to traditional path-based models.

stat.ML↗