SearcharxivSearch

arXiv subjects

Young Hyun Cho

Publications and source records attributed to Young Hyun Cho.

9 recordsLinked to original sources

Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations

Unexpected shocks recur in global operations, requiring decision rules that adapt as market and operating conditions change. Many operational systems also have hierarchical structures in which long-term and short-term decisions pursue a shared objective. We study how hierarchical reinforcement learning can strengthen resilience by adapting these interdependent rules jointly. We develop a two-timescale hierarchical reinforcement learning framework that adapts long-term and short-term policies at their respective time scales. Because the policies are interdependent, we synchronize their updates and prove, to our knowledge, the first convergence guarantees for coupled two-timescale learning. Over $T$ periods, our policies' average gap from an optimal policy pair is $O(T^{-1/2})$, improving to $O(\log T/T)$ when poor decisions produce clearer profit losses. In a used-car case study, inventory replenishment is the long-term decision and customer-arrival pricing the short-term decision. Relative to the strongest partially adaptive benchmark, the framework increases mean profit by $9.2\%$ under joint demand-supply shocks and by $11.8\%$ under a prolonged shock scenario, while maintaining a more stable profit trajectory over time. Short-term adaptation addresses routine seasonality and one-sided disruptions by responding immediately to changing conditions. Under joint demand-supply shocks, however, it is insufficient alone; long-term adaptation is also needed to create favorable conditions for short-term decisions. Joint adaptation thus yields higher and more stable profits through disruption and recovery. Because many organizations already use hierarchical planning, the framework strengthens operational resilience without altering existing decision structures.

stat.ML

When Should an AI Workflow Release? Always-Valid Inference for Black-Box Generate-Verify Systems

LLM-enabled AI workflows increasingly produce outputs through iterative generate-evaluate-revise loops. Each iteration can improve the candidate, but it also creates a release decision: when to stop and output the current result? This raises a statistical challenge because deployment-time evaluator scores are adaptively generated and repeatedly monitored, yet the likelihood models or exchangeability assumptions typically used for calibration are unavailable. We propose an always-valid release wrapper for existing generator-evaluator pipelines. The wrapper builds a hard-negative reference pool of high-scoring failures, calibrates deployment-time evaluator scores against this pool, and accumulates the resulting evidence with an e-process. This separates two roles: the reference pool turns black-box scores into conservative evidence, while the e-process provides validity under optional stopping. In theory, we show that a conservative reference pool yields finite-sample control of the probability of releasing on infeasible tasks, that is, tasks for which the given workflow is not capable of producing a reliable solution. We also characterize conditions under which the same conservative rule still achieves nontrivial release on feasible tasks. In an MBPP+ coding-agent case study, the wrapper reduces premature incorrect release relative to baseline stopping rules while still releasing on tasks for which the workflow repeatedly accumulates moderate supporting evidence.

stat.ML

Privacy-Preserving Reinforcement Learning from Human Feedback via Decoupled Reward Modeling

Preference-based fine-tuning has become an important component in training large language models, and the data used at this stage may contain sensitive user information. A central question is how to design a differentially private pipeline that is well suited to the distinct structure of reinforcement learning from human feedback. We propose a privacy-preserving framework that imposes differential privacy only on reward learning and derives the final policy from the resulting private reward model. Theoretically, we study the suboptimality gap and show that privacy contributes an additional additive term beyond the usual non-private statistical error. We also establish a minimax lower bound and show that the dominant term changes with sample size and privacy level, which in turn characterizes regimes in which the upper bound is rate-optimal up to logarithmic factors. Empirically, synthetic experiments confirm the scaling predicted by the theory, and experiments on the Anthropic HH-RLHF dataset using the Gemma-2B-IT model show stronger private alignment performance than existing differentially private baseline methods across privacy budgets.

stat.ML

Beyond Data Splitting: Full-Data Conformal Prediction by Differential Privacy

Privacy protection and uncertainty quantification are increasingly important in data-driven decision making. Conformal prediction provides finite-sample marginal coverage, but existing private approaches often rely on data splitting, reducing the effective sample size. We propose a full-data privacy-preserving conformal prediction framework that avoids splitting. Our framework leverages stability induced by differential privacy to control the gap between in-sample and out-of-sample conformal scores, and pairs this with a conservative private quantile routine designed to prevent under-coverage. We show that a generic differential privacy guarantee yields a universal coverage floor, yet cannot generally recover the nominal $1-α$ level. We then provide a refined, mechanism-specific stability analysis and yields asymptotic recovery of the nominal level. Experiments demonstrate sharper prediction sets than the split-based private baseline.

stat.ML

Privacy-Preserving Dynamic Assortment Selection

With the growing demand for personalized assortment recommendations, concerns over data privacy have intensified, highlighting the urgent need for effective privacy-preserving strategies. This paper presents a novel framework for privacy-preserving dynamic assortment selection using the multinomial logit (MNL) bandits model. Our approach employs a perturbed upper confidence bound method, integrating calibrated noise into user utility estimates to balance between exploration and exploitation while ensuring robust privacy protection. We rigorously prove that our policy satisfies Joint Differential Privacy (JDP), which better suits dynamic environments than traditional differential privacy, effectively mitigating inference attack risks. This analysis is built upon a novel objective perturbation technique tailored for MNL bandits, which is also of independent interest. Theoretically, we derive a near-optimal regret bound of $\tilde{O}(\sqrt{T})$ for our policy and explicitly quantify how privacy protection impacts regret. Through extensive simulations and an application to the Expedia hotel dataset, we demonstrate substantial performance enhancements over the benchmark method.

stat.ML

Formal Privacy Guarantees with Invariant Statistics

Motivated by the 2020 US Census products, this paper extends differential privacy (DP) to address the joint release of DP outputs and nonprivate statistics, referred to as invariant. Our framework, Semi-DP, redefines adjacency by focusing on datasets that conform to the given invariant, ensuring indistinguishability between adjacent datasets within invariant-conforming datasets. We further develop customized mechanisms that satisfy Semi-DP, including the Gaussian mechanism and the optimal $K$-norm mechanism for rank-deficient sensitivity spaces. Our framework is applied to contingency table analysis which is relevant to the 2020 US Census, illustrating how Semi-DP enables the release of private outputs given the one-way margins as the invariant. Additionally, we provide a privacy analysis of the 2020 US Decennial Census using the Semi-DP framework, revealing that the effective privacy guarantees are weaker than advertised.

cs.CR

Inverse Systems of Zero-dimensional Schemes in P^n

The authors construct the global Macaulay inverse system for a zero-dimensional subscheme Z of projective n-space P^n, from the local inverse systems of the irreducible components of Z. They show that when Z is locally Gorenstein a generic homogeneous form F of degree d apolar to Z determines Z when d is larger than an invariant b(Z). They also show that a natural upper bound for the Hiilbert function of Gorenstein Artin quotient of the coordinate ring is achieved for large socle degree. They show the uniqueness of generalized additive decompositions of a homogeneous form F into powers of linear forms, under suitable hypotheses. They include many examples.

math.AG

Conditions for Generic Initial Ideals to be Almost Reverse Lexicographic

Let $I$ be a homogeneous Artinian ideal in a polynomial ring $R=k[x_1,...,x_n]$ over a field $k$ of characteristic 0. We study an equivalent condition for the generic initial ideal $\gin(I)$ with respect to reverse lexicographic order to be almost reverse lexicographic. As a result, we show that Moreno-Socias conjecture implies Fröberg conjecture. And for the case $\Codim I \le 3$, we show that $R/I$ has the strong Lefschetz property if and only if $\gin(I)$ is almost reverse lexicographic. Finally for a monomial complete intersection Artinian ideal $I=(x_1^{d_1},...,x_n^{d_n})$, we prove that $\gin(I)$ is almost reverse lexicographic if $d_i > \sum_{j=1}^{i-1} d_j - i + 1$ for each $i \ge 4$. Using this, we give a positive partial answer to Moreno-Socias conjecture, and to Fröberg conjecture.

math.AC

Generic Initial Ideals of Artinian Ideals Having Lefschetz Properties or The Strong Stanley Property

For a standard Artinian $k$-algebra $A=R/I$, we give equivalent conditions for $A$ to have the weak (or strong) Lefschetz property or the strong Stanley property in terms of the minimal system of generators of the generic initial ideal $\mathrm{gin}(I)$ of $I$ under the reverse lexicographic order. Using the equivalent condition for the weak Lefschetz property, we show that some graded Betti numbers of $\mathrm{gin}(I)$ are determined just by the the Hilbert function of $I$ if $A$ has the weak Lefschetz property. Furthermore, for the case that $A$ is a standard Artinian $k$-algebra of codimension 3, we show that every graded Betti numbers of $\mathrm{gin}(I)$ are determined by the graded Betti numbers of $I$ if $A$ has the weak Lefschetz property. And if $A$ has the strong Lefschetz (resp. Stanley) property, then we show that the minimal system of generators of $\mathrm{gin}(I)$ is determined by the graded Betti numbers (resp. by the Hilbert function) of $I$.

math.AC