SearcharxivSearch

arXiv subjects

Junwei Lu

Publications and source records attributed to Junwei Lu.

At least 19 recordsLinked to original sources

Vector Balancing via Directional Total Variation

Our main result is a $3\sqrt{2\pi}$ bound for the Koml\'os signing problem: every finite family of real vectors of Euclidean norm at most one admits a signed sum of $\ell_\infty$-norm less than this constant, independently of the dimension and the family size. For any $\kappa\ge0$, if a bounded open convex set supports a probability density with directional total variation at most $\kappa$ in every unit direction, then its open-set Banaszczyk transform supports another such density with the same $\kappa$, provided the translation vector $v$ satisfies $\kappa\|v\|_2\le1/3$. As a consequence, every finite set system in which each element belongs to at most $t$ sets, where $t\ge1$ is an integer, admits a two-coloring whose imbalance in each set is less than $3\sqrt{2\pi t}$. This gives the square-root dependence predicted by the Beck-Fiala conjecture. The proof was discovered by the Odin Automatic AI Research Agent.

math.CO

Optimal Value Inference for Reinforcement Learning

We study offline inference for the optimal value in reinforcement learning. Two new nuisances are derived as fixed points of a self-induced Bellman equation, in which we approximate the maximum Bellman operator by its softmax correspondence. We propose a debiased estimator through the Neyman orthogonality and establish its asymptotic normality under diverging horizons even when the behavior policy changes with time, as long as the nuisances have the statistical rates that can be achieved by many machine learning methods. We provide a concrete estimating procedure for these nuisances and show they can lead to valid inference. Synthetic experiments validate the numerical performance of our inference method, and we implement it in real-life decision-making problems, including bike repositioning and AI agentic tool use.

stat.ML

An Improved Lower Bound for the Complex Grothendieck Constant

We prove $K_G^{\mathbb{C}}>1.35584631827168$ for the classical complex Grothendieck constant, closing more than one quarter of the gap between Davie's lower bound and Haagerup's upper bound. The numerical part of the proof is rigorously verified by interval arithmetic. The lower bound and the proof were discovered by the Odin Automatic AI Research Agent.

math.FA

On Unavoidable Faces of High-Dimensional Polytopes

Kalai's cube--simplex conjecture asserts that for all positive integers $\ell,k$, there is an integer $f(\ell,k)$ such that every polytope of dimension at least $f(\ell,k)$ has either a simplex $\ell$-face or a cube $k$-face; let $f_s(\ell,k)$ denote the threshold restricted to simple polytopes. Finiteness of $f(\ell,k)$ is known only for $\ell,k \leq 2$. In addition, Kalai proved that $f_s(2,k) \leq 2k^2$. Here we prove that $f_s(\ell,k)$ is finite for all $\ell \geq 2$ and $k \geq 3$, the first such result beyond $\ell = 2$, with $f_s(2,k) \leq 2k^2-1$ and $f_s(\ell,k) \leq \tfrac{1}{2}k^2\ell\,2^k$ for $\ell \geq 3$. In the opposite direction, we obtain the lower bounds $f(\ell,k) \geq (5\lfloor \ell/2 \rfloor + (\ell \bmod 2) - 1)(k-1)+1$ and $f_s(\ell,k) \geq \max\{4,\,2(\ell-1)\}(k-1)+1$. A companion question asks for the minimum possible size of a 3-face within a higher-dimensional polytope. Meisinger, Kleinschmidt and Kalai proved that every rational $d$-polytope with $d \geq 9$ has a $3$-face with fewer than $78$ vertices or fewer than $78$ facets. Here we improve their bound: every convex polytope of dimension at least $15$ has a $3$-face with at most $13$ facets. One step of our proof requires an explicit exact rational certificate or identity on flag numbers. This certificate is computed using linear programming.

math.CO

A Complex Structure on $S^2\times S^4$

We show that $S^2\times S^4$ admits a complex structure. Starting from a modular family of complex two-tori associated with the $(3,4,\infty)$ triangle group and the compactification constructed in [Alp\"oge 2026], we replace the period lattice by its unique monodromy-invariant index-two superlattice and compactify the resulting family. In the integral Mayer--Vietoris calculation, a primitive local class whose double is the class of a cusp component and the unique nonzero torsion class contributed by the multiplicity-four fibre restrict to the same class of order two on the common boundary. We then prove that the resulting compact complex threefold is diffeomorphic to $S^2\times S^4$. The proof is discovered by the Odin Automatic AI Research Agent.

math.GM

A Metric with Positive Sectional Curvature on $S^2\times S^3$

We prove that $S^2\times S^3$ admits a Riemannian metric with positive sectional curvature. We view it as a principal circle bundle over $S^2\times S^2$. A diagonal Cheeger deformation of the base and a connection whose curvature form vanishes on the remaining flat tori yield a nonnegatively curved connection metric whose zero-curvature planes are the horizontal lifts of the tangent planes to those tori. We then perturb this metric by the real part of a global complex-valued symmetric $2$-tensor. Differentiation along the circle fibers produces a trace-free first variation of the second fundamental form on local horizontal lifts of the flat tori. The Gauss equation converts this into a positive second-order curvature term that dominates as the fibers shrink. A quantitative lower bound for the Hessian in directions normal to the set of zero-curvature planes extends this positivity to nearby planes. The metric and the proof are discovered by the Odin Automatic AI Research Agent.

math.DG

Weak-Type Bounds for Convolution on the Boolean Hypercube

Let $G$ be the Boolean hypercube which carries uniform measure $\lambda$, and let $T_\mu$ denote convolution by a finite positive measure $\mu$ on $G$. For $\psi_\mu(u)=\sup\{u\lambda(\{T_\mu f\geq u\}):f\geq 0,\|f\|_1=1\},$ we prove Talagrand's convolution conjecture (Talagrand, 1989): if $\mu_a=((1+a)\delta_1/2+(1-a)\delta_{-1}/2)^{\otimes n}$ and $0 1$ and $n\geq1$, where $C_a$ depends only on $a$. The proof utilizes the reverse-heat and Boolean-bridge framework of Chen (2025) and the localized terminal-discrepancy method of Xiang and Zhang (2026). We introduce a new power coupling: each reverse edge ratio is split into two geometric powers. This choice produces a switched exponential weight which restores the exact reverse jump rate of the perturbed coordinate. The resulting endpoint comparison yields an anti-concentration profile estimate without the iterated-logarithmic factor. The proof was discovered by the Odin Automatic AI Research Agent.

math.PR

Sharp Bounds on the Independence Number of Simplicial Spheres

We study the maximum size of an independent set in the graph of a simplicial sphere. Let $\beta(d,n)$ denote this maximum over all simplicial $(d-1)$-spheres on $n$ vertices, and let $\alpha(d,n)$ denote the maximum restricted to flag $(d-1)$-spheres. For every fixed $d\geq4$, we prove $\beta(d,n)=n-\Theta(n^{1/\lfloor d/2\rfloor})$. For flag spheres, we show $\alpha(d,n)\geq n-4\sqrt n+O(1)$ for all $d\geq4$ and determine the correct asymptotic order $\alpha(d,n)=n-\Theta(\sqrt n)$ for dimensions $d=4,5$. We also investigate the independence sets of Bier spheres and show that, in contrast to our other results, for this very large family of spheres, the independence number cannot be larger than $\left\lfloor\frac{n}{2}\right\rfloor.$

math.CO

Hypothesis-Disciplined Multi-Agent Automated Formalization of Asymptotic Statistical Theory

Asymptotic statistical theory is a challenging domain for AI-assisted formalization: its central results mix convergence statements, asymptotic expansions, functional analysis, and regularity conditions that have a large gap from existing infrastructure in Lean 4 formalization. To address these challenges, we propose a hypothesis-disciplined Lean 4 formalization pipeline built from multiple agents: a manager that coordinates seven specialist roles for proof planning, skeleton scaffolding, Mathlib reconnaissance, proof construction, integration, independent review, and audit. The main methodological discipline is the hypothesis-disciplined audit, implemented by the Auditor agent: every main-theorem hypothesis and concept-layer field must be anchored in the source mathematical prose, justified as a Lean encoding adapter, marked as source-implied, or rejected as an unsupported strengthening. Using this workflow, we build a systematic formalization of asymptotic statistical theory, especially the parametric and semi-parametric models' asymptotic distribution and efficiency results. The resulting Lean development is axiom-clean and source-faithful, with Lean-checked and human-audited proofs of core parametric and semi-parametric theorems organized so that theorem-agnostic infrastructure and statistical concept definitions are separated from theorem-specific assembly. The formalization results are available at https://github.com/junwei-lu/Lean-Asymptotic-Statistical-Theory.

cs.AI

Channel-Oriented Design for EEG-to-Music Reconstruction

Brain-computer interfaces aim to decode naturalistic stimuli from neural signals, yet most progress to date has focused on vision and language. In this article, we study a more challenging but far less explored setting, EEG-to-music reconstruction, where signals are weak, distributed, and highly susceptible to noise and channel variability. Our central finding is that early channel mixing destroys weak but discriminative EEG signals. To address this, we propose a channel-oriented design with three key components. Specifically, channel-wise tokenization treats each electrode as an explicit token to retain spatially localized neural evidence, channel-wise multi-view self-distillation enforces consistency across temporal crops and random channel subsets to learn robust and distributed representations, and channel-wise data augmentation introduces structured channel dropout to improve invariance to noise, artifacts, and missing electrodes. Together, these components preserve weak yet informative signals across channels and enable stable alignment to a semantic music representation space. We integrate this channel-oriented design within an encoding-alignment-decoding pipeline for EEG-to-music reconstruction. Theoretically, we characterize when preserving channel-level structure leads to improved alignment. Empirically, we compare with a range of state-of-the-art baselines and demonstrate consistent and significant performance gains.

cs.SD

Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback

We study estimation and statistical inference for reward models used in aligning large language models (LLMs). A key component of LLM alignment is reinforcement learning from human feedback (RLHF), where humans compare pairs of model-generated answers and their preferences are used to train a reward model. However, human feedback is inherently heterogeneous, creating significant challenges for reliable reward learning. To address this, we adopt a heterogeneous preference framework that jointly models the latent reward of answers and human rationality. This leads to a challenging biconvex optimization problem, which we solve via an alternating gradient descent algorithm. We establish theoretical guarantees for the resulting estimator, including its convergence and asymptotic distribution. These results enable the construction of confidence intervals for reward estimates. Leveraging these uncertainty quantification results, we conduct valid statistical comparisons between rewards and incorporate uncertainty into the best-of-$N$ (BoN) policy framework. Extensive simulations demonstrate the effectiveness of our method, and applications to real LLM data highlight the practical value of accounting for uncertainty in reward modeling for LLM alignment.

stat.ML

DANIEL: A Distributed and Scalable Approach for Global Representation Learning with EHR Applications

Classical probabilistic graphical models face fundamental challenges in modern data environments, which are characterized by high dimensionality, source heterogeneity, and stringent data-sharing constraints. In this work, we revisit the Ising model, a well-established member of the Markov Random Field (MRF) family, and develop a distributed framework that enables scalable and privacy-preserving representation learning from large-scale binary data with inherent low-rank structure. Our approach optimizes a non-convex surrogate loss function via bi-factored gradient descent, offering substantial computational and communication advantages over conventional convex approaches. We evaluate our algorithm on multi-institutional electronic health record (EHR) datasets from 58,248 patients across the University of Pittsburgh Medical Center (UPMC) and Mass General Brigham (MGB), demonstrating superior performance in global representation learning and downstream clinical tasks, including relationship detection, patient phenotyping, and patient clustering. These results highlight a broader potential for statistical inference in federated, high-dimensional settings while addressing the practical challenges of data complexity and multi-institutional integration.

stat.ME

MOTION: ML-Assisted On-Device Low-Latency Motion Recognition

The use of tiny devices capable of low-latency gesture recognition is gaining momentum in everyday human-computer interaction and especially in medical monitoring fields. Embedded solutions such as fall detection, rehabilitation tracking, and patient supervision require fast and efficient tracking of movements while avoiding unwanted false alarms. This study presents an efficient solution on how to build very efficient motion-based models only using triaxial accelerometer sensors. We explore the capability of the AutoML pipelines to extract the most important features from the data segments. This approach also involves training multiple lightweight machine learning algorithms using the extracted features. We use WeBe Band, a multi-sensor wearable device that is equipped with a powerful enough MCU to effectively perform gesture recognition entirely on the device. Of the models explored, we found that the neural network provided the best balance between accuracy, latency, and memory use. Our results also demonstrate that reliable real-time gesture recognition can be achieved in WeBe Band, with great potential for real-time medical monitoring solutions that require a secure and fast response time.

cs.CV

Automated Hierarchical Graph Construction for Multi-source Electronic Health Records

Electronic Health Records (EHRs), comprising diverse clinical data such as diagnoses, medications, and laboratory results, hold great promise for translational research. EHR-derived data have advanced disease prevention, improved clinical trial recruitment, and generated real-world evidence. Synthesizing EHRs across institutions enables large-scale, generalizable studies that capture rare diseases and population diversity, but remains hindered by the heterogeneity of medical codes, institution-specific terminologies, and the absence of standardized data structures. These barriers limit the interpretability, comparability, and scalability of EHR-based analyses, underscoring the need for robust methods to harmonize and extract meaningful insights from distributed, heterogeneous data. To address this, we propose MASH (Multi-source Automated Structured Hierarchy), a fully automated framework that aligns medical codes across institutions using neural optimal transport and constructs hierarchical graphs with learned hyperbolic embeddings. During training, MASH integrates information from pre-trained language models, co-occurrence patterns, textual descriptions, and supervised labels to capture semantic and hierarchical relationships among medical concepts more effectively. Applied to real-world EHR data, including diagnosis, medication, and laboratory codes, MASH produces interpretable hierarchical graphs that facilitate the navigation and understanding of heterogeneous clinical data. Notably, it generates the first automated hierarchies for unstructured local laboratory codes, establishing foundational references for downstream applications.

stat.ML

Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation

Motivated by the need for rigorous and scalable evaluation of large language models, we study contextual preference inference for pairwise comparison functionals of context-dependent preference score functions across domains. Focusing on the contextual Bradley-Terry-Luce model, we develop a semiparametric efficient estimator that automates the debiased estimation through aggregating weighted residual balancing terms across the comparison graph. We show that the efficiency is achieved when the weights are derived from a novel strategy called Fisher random walk. We also propose a computationally feasible method to compute the weights by a potential representation of nuisance weight functions. We show our inference procedure is valid for general score function estimators accommodating the practitioners' need to implement flexible deep learning methods. We extend the procedure to multiple hypothesis testing using a Gaussian multiplier bootstrap that controls familywise error and to distributional shift via a cross-fitted importance-sampling adjustment for target-domain inference. Numerical studies, including language model evaluations under diverse contexts, corroborate the accuracy, efficiency, and practical utility of our method.

stat.ML

Latent Factor Point Processes for Patient Representation in Electronic Health Records

Electronic health records (EHR) contain valuable longitudinal patient-level information, yet most statistical methods reduce the irregular timing of EHR codes into simple counts, thereby discarding rich temporal structure. Existing temporal models often impose restrictive parametric assumptions or are tailored to code level rather than patient-level tasks. We propose the latent factor point process model, which represents code occurrences as a high-dimensional point process whose conditional intensity is driven by a low dimensional latent Poisson process. This low-rank structure reflects the clinical reality that thousands of codes are governed by a small number of underlying disease processes, while enabling statistically efficient estimation in high dimensions. Building on this model, we introduce the Fourier-Eigen embedding, a patient representation constructed from the spectral density matrix of the observed process. We establish theoretical guarantees showing that these embeddings efficiently capture subgroup-specific temporal patterns for downstream classification and clustering. Simulations and an application to an Alzheimer's disease EHR cohort demonstrate the practical advantages of our approach in uncovering clinically meaningful heterogeneity.

stat.ME

Time-Aware Attention for Enhanced Electronic Health Records Modeling

Electronic Health Records (EHR) contain valuable clinical information for predicting patient outcomes and guiding healthcare decisions. However, effectively modeling Electronic Health Records (EHRs) requires addressing data heterogeneity and complex temporal patterns. Standard approaches often struggle with irregular time intervals between clinical events. We propose TALE-EHR, a Transformer-based framework featuring a novel time-aware attention mechanism that explicitly models continuous temporal gaps to capture fine-grained sequence dynamics. To complement this temporal modeling with robust semantics, TALE-EHR leverages embeddings derived from standardized code descriptions using a pre-trained Large Language Model (LLM), providing a strong foundation for understanding clinical concepts. Experiments on the MIMIC-IV and PIC dataset demonstrate that our approach outperforms state-of-the-art baselines on tasks such as disease progression forecasting. TALE-EHR underscores the benefit of integrating explicit, continuous temporal modeling with strong semantic representations provides a powerful solution for advancing EHR analysis.

cs.LG

Contextual Online Uncertainty-Aware Preference Learning for Human Feedback

Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm in artificial intelligence to align large models with human preferences. In this paper, we propose a novel statistical framework to simultaneously conduct the online decision-making and statistical inference on the optimal model using human preference data based on dynamic contextual information. Our approach introduces an efficient decision strategy that achieves both the optimal regret bound and the asymptotic distribution of the estimators. A key challenge in RLHF is handling the dependent online human preference outcomes with dynamic contexts. To address this, in the methodological aspect, we propose a two-stage algorithm starting with $\epsilon$-greedy followed by exploitations; in the theoretical aspect, we tailor anti-concentration inequalities and matrix martingale concentration techniques to derive the uniform estimation rate and asymptotic normality of the estimators using dependent samples from both stages. Extensive simulation results demonstrate that our method outperforms state-of-the-art strategies. We apply the proposed framework to analyze the human preference data for ranking large language models on the Massive Multitask Language Understanding dataset, yielding insightful results on the performance of different large language models for medical anatomy knowledge.

stat.ML