SearcharxivSearch

arXiv subjects

Kwan Soo Shin

Publications and source records attributed to Kwan Soo Shin.

8 recordsLinked to original sources

Beyond headcount and human capital: The Effective Cognitive Population as a decomposable capacity unit for AI-era planning

National planning counts population, human capital, and artificial-intelligence preparedness in separate ledgers. Demographic accounting has advanced from headcount to skills-adjusted stocks and still debates how much age structure retains once skills are modeled, yet no existing unit carries the conditions under which preparedness becomes productive capacity. This study introduces the Effective Cognitive Population (ECP), a decomposable unit that weights population by capability and by the conditions under which capability is deployed, anchored to the World Bank Human Capital Index Plus (HCI+) and the non-overlapping dimensions of the IMF AI Preparedness Index. The architecture is portable in principle; the case tested here is artificial intelligence, which has a published preparedness index. For 144 countries, HCI+ becomes a productivity level, AI opportunity uses digital infrastructure and innovation integration, conversion governance uses regulation and ethics, and the benchmark is ECP = N H(1 + AC). Against 2024 total output on identical population bases, ECP raises criterion R-squared from 0.849 for the HCI+-adjusted stock to 0.882 and lowers leave-one-country-out RMSE from 0.723 to 0.641, with the working-age comparison identical and bootstrap intervals excluding zero. Eighty-nine of 144 countries move at least ten rank positions from headcount, mostly through the human-capital adjustment itself. Results are stable across denominators, vintages, aggregation forms, and a 27-rule multiverse. The direct A by C interaction is not statistically supported, so the conjunction is a planning rule rather than causal complementarity. ECP is a diagnostic ledger whose scope excludes forecasts of population decline and estimates of AI's causal productivity effect.

econ.GN

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals

AI in radiology and other safety-critical workflows is evaluated on the hazards it is told to find, yet harm arises disproportionately from hazards no one specified. We show that conditioning a language or vision model on a narrow task suppresses its reporting of co-present, safety-critical signals it can otherwise report, a behavioral analogue of human inattentional blindness. Across radiology text scenarios and thoracic-image vision tasks, ordinary focused instructions suppressed reporting by up to 0.92; the gap ranged from minimal to complete across seven models, did not vary monotonically with scale, and persisted in a reasoning model, while one flagship model showed a robust safety-reporting override. We term this dissociation the Inattentional Gap: a system can score near-perfectly on specified hazards while omitting co-present safety-critical hazards. In a 24-scenario probe, an independent open-ended critic restored every omitted finding. We propose reporting-complete evaluation as an admission criterion for safety-critical deployment.

cs.CL

What Capital After Labor? Forecasting the Talent ROI Transition in the Human-AI Era

AI augmentation breaks the accounting link between labor time and productive contribution, yet firms continue to evaluate talent through time-based overhead bundles. This paper develops a forecasting framework for the transition from time-based talent accounting to output-based talent ROI in the human-AI era, organized around five theorems: Theorem 3 (ROI Inversion at τ*) carries the central transition claim, with overhead non-additivity, augmentation-saved-time pathways, innovation-premium amplification, and human-AI dyad attribution uncertainty as the mechanism architecture. Korea's staged 52-hour workweek mandate provides the early-warning case. In a DART panel of 365 firms (2,281 observations), the SG&A-to-revenue ratio rose from 18.26 percent (2018) to 20.06 percent (2020) and peaked at 20.10 percent (2024). Under the revenue-percentile cohort proxy, two-way fixed effects (+1.56 pp, p = 0.049), pooled event-study estimates (+4.21 pp at t = +3), and Callaway-Sant'Anna estimates (+4.51 pp at t = +4) converge on a positive overhead-pressure pattern. Institutional cohort evidence separates the two readings: under the statutory employee-size cohort the coefficient is indistinguishable from zero, weighing against a pure 52-hour-law interpretation and supporting the secular regime reading; a 2015-2017 backward extension (224 firms) argues against pre-existing trends. We read the Korean evidence as, to our knowledge, the first publicly documented signature of a secular pre-τ overhead-pressure regime in which time-based accounting still dominates while AI augmentation raises firm-internal overhead. Output-based firms are forecast to outperform time-based peers by 1.5-2.0 percentage points in TFP growth by 2032. The contribution is a forecasting model and planning tool for AI-augmented talent ROI accounting.

econ.GN

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation). Scoring a compact 146-million-parameter auditor's frozen-representation read-out and a frontier judge against each label on the identical 720 replies, the gap between the instruments moves by roughly 0.2 AUROC when the target changes. Under the judge's deployed interface, a single verdict, the ranking reverses: the auditor leads on exposure, 0.804 against 0.718, and trails on manifestation, 0.690 against 0.811. Matching the output resolution from either direction, by asking the judge a target-specific question answered with a continuous confidence score or by thresholding the auditor's read-out, removes the reversal but not the interaction, which excludes zero at all three resolutions (0.207, 0.237 and 0.169). The target governs how far apart the instruments are; the interface governs whether that distance changes their order. The auditor's hyperbolic geometry confers no advantage here. A single behavioural-detection AUROC is under-specified: such claims are comparable only when they state the estimand, the evaluator, and its output interface.

cs.CL

After the Euclidean Highway: Hyperbolic Expert AI as the Next Innovation

Expert domains are trees; the Euclidean transformer is not, diluting parent-child structure exponentially at depth. The hyperbolic turn left one question unasked: not how much of a network to curve, but where curvature may touch the gradient. Placement is a law, not a knob: the same geometry on a trainable adapter collapses training (seventeen training collapses, ~220 GPU-hours), yet at the loss layer alone it trains without one -- this is HySAT (Hyperbolic Structure-Aware Training), hyperbolic losses at the loss layer only. Across six expert SLMs we constructed and deployed (Llama 3.1 and EXAONE 3.5; four adapter strategies; 18.0M-sample corpus; zero NaN over ~317K optimizer steps), a matched four-arm ablation isolates the preserved manifold invariant, and three propositions and a lemma prove why loss-only placement is stable where adapter-on-manifold is not. Four models are operationally deployed (one live, consumer-facing), two open-weight, with per-step traces and a seventeen-incident failure ledger on Zenodo (CC-BY-4.0).

cs.LG

Forecasting AI-Era Productivity: The Intellectually Converged Human Framework and a Missing Cognitive Mediator in Production Function Theory

Why does massive AI investment fail to generate commensurate productivity gains? We argue the paradox is theoretically generated: prevailing production function frameworks encounter a structural boundary by treating AI as a separable factor of production without modeling the cognitive mediation through which AI generates productive value. This directs investment toward deployment when productivity requires prior development of what we term convergence capacity (C). We propose the Intellectually Converged Human (ICH) framework, a fifth-stage framework for production function theory: H-hat = H[1 + phi(A,C)], where effective productive capacity equals human capital (H) scaled by an augmentation factor [1 + phi], with phi jointly determined by AI utilization intensity (A) and convergence capacity (C), a four-dimensional cognitive construct encompassing embodied understanding, metacognition, temporal integration, and integrative thinking. The production function Y = F(K, H-hat) provides a human-centered mechanism for Solow's TFP residual: A_Solow = [1 + phi(A,C)]^(1-alpha). The framework predicts three augmentation regimes with distinct policy implications. Descriptive cross-national analysis of 20 OECD economies shows the AIxC interaction is associated with 86% of TFP variance versus 31% for AI alone, a pattern-consistent finding in the small-n theoretical tradition. South Korea exemplifies national-scale under-augmentation: high H, substantial A, low C produce phi = 0. We distinguish convergence capacity from adjacent constructs, absorptive capacity, dynamic capability, and human capital, and demonstrate that C constitutes the specific cognitive mediator that prior frameworks have left implicit. We derive C-first policy prescriptions and offer three empirically testable propositions with a falsifiable 10-year forecast.

econ.GN

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning

When copies of the same language model are prompted to debate, they produce diverse phrasings of one perspective rather than diverse perspectives. Multi-agent debate (MAD), and more broadly closed-system reasoning where agents iteratively transform each other's outputs, tends to preserve answer accuracy while degrading the reasoning behind those answers. We name the multi-agent case the Debate Trap and the broader phenomenon the Reasoning Trap, offering a programmatic theory of evidence-grounded reasoning failure.The framework has three parts: (i) SFS (Supported Faithfulness Score), a claim-level metric verifying decomposed atomic claims against provided evidence (decomposer-invariant rankings: Spearman rho=1.0); (ii) EGSR (Evidence-Grounded Socratic Reasoning), replacing adversarial argumentation with evidence-grounded inquiry; (iii) Theorem 1 (DPI Bound): under standard MAD, the chain E -> O^0 -> O^1 -> ... is Markov, and the Data Processing Inequality implies E[I(E;O^{t+1})] <= E[I(E;O^t)]. Three companion results -- open-system recovery (Theorem 2), EGSR accumulation (Lemma 2), and vote-aggregation floor (Proposition 1) -- partition multi-step LLM reasoning by its information-theoretic relationship to E. Across 16 conditions on SciFact (300 claims) and FEVER (1,000 claims), DebateCV (C13) preserves 88% of baseline accuracy while SFS drops 43%; majority-vote MAD (C15) reduces SFS to 1.7% of baseline (p < 10^{-6}, d = -0.96); EGSR recovers 98%. An R6 cohort study (Korean n=10x30 FEVER; English n=3x200 SciFact) finds inter-rater Fleiss kappa <= +0.018 with 0.8-1.4 Likert intra-rater shifts across language and domain -- the human agreement that faithfulness metrics have been calibrated against is not itself stable. We offer one falsifiable conjecture: any closed-system reasoning protocol preserving Theorem 1's Markov structure is, in expectation, subject to the same DPI bound.

cs.CL

The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't

An auditor instructs an AI assistant: "open each file individually using the Read tool -- no scripts, no agents." The AI replies "Yes" -- then issues a single batched call summarizing all fifty files at once. We call this the Compliance Gap: a third, orthogonal axis of AI honesty distinct from factual truthfulness and rhetorical substance. Three questions: does this verbal-behavioral disconnect exist (existence); can any text-only observer recover it (detectability); what infrastructure does AI deployment need (remedy)? Some 75 benchmarks (IFEval, SWE-bench, BFCL, COMPASS, SpecEval) measure outcome fidelity; none measures process fidelity. Theorem 1 shows the gap is structurally inevitable under RL that rewards text without observing behavior. Theorem 2, via the Data Processing Inequality, shows it is undetectable from text alone -- by any human or LLM observer, present or future. Thirteen experiments and 2,031 sessions on six frontier models confirm both predictions. Under default framing, all six exhibit instruction compliance rates of 0% -- Claude Sonnet 4 verbally agrees ten out of ten times then bypasses in all ten. The gap is selective: 97% compliance where rationale is rewarded (audit trails), 0-4% where it is not (file reading, privacy masking); removing delegation tools raises compliance to 75% (Cohen's d = 2.47), confirming environmental affordance rather than weight-encoded failure. Nine blinded human raters achieve Fleiss' kappa = 0.130 and correctly identify zero of fifteen compliant sessions, exactly as Theorem 2 predicts. Where humans show 47% intention-behavior gaps in psychology and 96.5pp gaps in surgical audits, RLHF-trained models approach 100% under default conditions -- a regime warranting its own measurement infrastructure. We release BS-Bench: the first open benchmark for process compliance, with seven tool-call-log audit metrics and a public leaderboard.

cs.CL