Searcharxiv⌕ Search

arXiv subjects

Dong Hyeon Jeon

Publications and source records attributed to Dong Hyeon Jeon.

3 recordsLinked to original sources

Silent Success: A Release Gate That Passed on Checks It Never Ran, and Eight More

A blocking quality gate in a production release pipeline reported PASS on a run in which one subgate had executed neither of its two checks and another had run six of its eight. Both keys that decided the run asked whether a violation had been observed, and both computed that from a population already stripped of the cases that failed to run, so absent data answered "no" and a run that checked almost nothing scored perfectly, for two weeks of green builds. Introducing a third value, pass, violate, and unable to determine, turned those silent passes into failures and put the shortfall into the exit status; separate work on the same specifications then found a detector firing with a margin of 0.000177 percentage points, its firing floor holding at the sixth decimal place. Eight further instances of the same form followed: five more from the same engagement, two in open-source projects, a gateway where a configured cache TTL was declared but never applied on the write path, and an inference server whose count of free cache slots was credited on every step for releases that returned nothing, and one committed by the author while writing this paper, using tooling built to prevent exactly it. What the nine cases are offered for is not the novelty of the category but the convergence of the remedy, nine different missed questions, one kind of act, seven of the nine answerable by a single query, command, or comparison, and that convergence is falsifiable, which is what this paper asks to be judged on.

cs.SE↗

Empty Intersection: Provenance Coverage Rose to 98% and Neither Verification Decision Moved

Two structural defenses for provenance, a grade on every row, so that a verification routine cannot mistake the system's own output for an observation, and a single write ingress, so that the grade is enforced rather than merely conventional, were measured against the production deployment that motivated them, over a frozen snapshot of 194,620 rows and the two verification decisions the snapshot supports. Neither reaches either decision. Both were prescribed by a companion paper, which diagnosed that deployment: its verification routines decided outcomes using values the system itself had written. Neither prescription is new: both are established practice in fields that do not cite one another, and no prior work measuring whether either changes a verdict was found, so what is offered here is the measurement and not the prescriptions. Filtering the verification queries by grade turns both decisions from pass to undetermined; widening the grade vocabulary raises classified coverage from 36.1% to 98.4%; a single ingress requiring a grade refuses 3,070 writes. None of the three gives either decision admissible input. The prescriptions do not fail at what they specify. Each is stated over the population and makes no reference to any decision, so neither says which rows a decision will read, and the rows each intervention repairs and the 32 rows the decisions read do not intersect. The intervention that changed the most rows shows the reach most plainly: all 121,296 rows it moved from unnameable to named fall outside both query windows. This paper reports the conditions, measured rather than designed, under which the decisions would have admissible input at all, and notes that the two decisions are blocked for different reasons.

cs.SE↗

What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers

GPT-3 shows remarkable in-context learning ability of large-scale language models (LMs) trained on hundreds of billion scale data. Here we address some remaining issues less reported by the GPT-3 paper, such as a non-English LM, the performances of different sized models, and the effect of recently introduced prompt optimization on in-context learning. To achieve this, we introduce HyperCLOVA, a Korean variant of 82B GPT-3 trained on a Korean-centric corpus of 560B tokens. Enhanced by our Korean-specific tokenization, HyperCLOVA with our training configuration shows state-of-the-art in-context zero-shot and few-shot learning performances on various downstream tasks in Korean. Also, we show the performance benefits of prompt-based learning and demonstrate how it can be integrated into the prompt engineering pipeline. Then we discuss the possibility of materializing the No Code AI paradigm by providing AI prototyping capabilities to non-experts of ML by introducing HyperCLOVA studio, an interactive prompt engineering interface. Lastly, we demonstrate the potential of our methods with three successful in-house applications.

cs.CL↗