SearcharxivSearch

arXiv subjects

Jing Lei

Publications and source records attributed to Jing Lei.

At least 19 recordsLinked to original sources

Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models

Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each token while leaving the Transformer backbone fixed. Depth-recurrent (looped) Transformers pursue this goal but are hard to scale, because looped computation does not fit naturally with the pipeline parallelism used to train the largest models. We add computation along the sequence-length dimension, where the extra computation is simply a longer input and stays compatible with standard large-model training. We propose Hidden Decoding, a sequence-length scaling method applied during continued pretraining (CPT). It expands each token into n streams with independent embedding tables and keeps the intermediate streams' key-value cache as context, so each token performs more internal computation without adding or widening Transformer layers. To keep this affordable at scale, we introduce Stream-Factorized Attention, in which most layers attend only within each stream and only a few layers mix across streams, reducing the attention cost from quadratic to roughly linear in n. Experiments support two scaling results. At frontier scale, we train WeLM-HD4-80B and WeLM-HD4-617B at n=4 and improve their matched non-HD baselines, making Hidden Decoding the first demonstrated sequence-length scaling method at the 100B+ MoE scale. Across expansion factors, the gains grow as n increases, showing that sequence-length expansion is a practical fixed-backbone scaling path for frontier-scale LLMs.

cs.CL

Cross-channel Specific Emitter Identification and Verification via Signal Envelope

Specific emitter identification (SEI) determines which known emitter a received signal originates from, while specific emitter verification (SEV) determines whether the received signal genuinely comes from its claimed emitter. In this paper, we consider the effect of wireless fading channels on SEI and SEV. When the Rician $K$-factor varies, the resulting distribution shift induced by the channel degrades both identification and verification performance. To address this issue, we first theoretically prove that the coefficient of variation of the signal envelope is strictly monotonic with respect to the Rician $K$-factor. Motivated by this property, we propose an envelope-guided adaptive feature modulation (EAFM) identifier for SEI and an EAFM with Mahalanobis distance metric learning (EAFM-MD) verifier for SEV. Specifically, the proposed EAFM identifier adopts a dual-branch neural network to extract device-oriented features from the IQ-domain input and channel-conditioning features from the normalized signal envelope, and adaptively modulates the former via feature-wise linear modulation. Then, we extend the EAFM identifier to an EAFM-MD verifier. The device-fingerprint library is constructed by storing the feature centroid and covariance for each enrolled device, along with the within-device Mahalanobis distances of training signals. For verification, the Mahalanobis distance between the extracted test features and each stored centroid is computed using the stored covariance matrix, and the minimum distance is compared to the corresponding device threshold to make a decision. Finally, numerical results show that the proposed EAFM identifier improves cross-channel identification performance, while the proposed EAFM-MD verifier achieves superior detection performance against unknown spoofing attacks.

eess.SP

A Survey of Physical-layer Authentication Enhanced by Emerging Spatial Domain Technologies

This article surveys spatial-domain-enhanced Physical-layer Authentication (PLA), with Dual-polarized Antennas (DPA), Massive Multiple-Input Multiple-Output (MIMO), and Reconfigurable Intelligent Surfaces (RIS) as the primary focus. With the rapid growth of wireless deployments, authentication mechanisms face stringent requirements for high security, low overhead, and low latency. PLA offers lightweight identity verification by exploiting physical-layer characteristics. However, the effectiveness of PLA critically depends on how physical observations are constructed and validated under wireless channels. Unlike existing surveys that mainly organize PLA by authentication modality, feature source, and evaluation metrics, this work emphasizes the connection between spatial-domain enhancement mechanisms, the resulting feature representation, and the authentication procedure. We review how DPA, Massive MIMO, and RIS reshape PLA feature representation, and we summarize newly introduced security threats along with representative defense strategies. Case studies further illustrate the practical impact, such as representative detection-probability trends across Signal-to-Noise Ratio regimes and quantitative comparisons among representative schemes. Finally, we outline promising future opportunities enabled by Dynamic Metasurface Antennas, Extra-large MIMO, and spatial configuration with artificial intelligence.

eess.SP

Probabilistic Win Ratio Method For Hierarchical Composite Endpoints With Coarsened Outcomes

The win ratio is increasingly used to analyze prioritized composite endpoints in clinical trials, but standard implementations rely on deterministic pairwise comparisons and can perform poorly in the presence of censoring and endpoint-specific missingness. In such settings, unresolved comparisons are often treated as ties, leading to loss of efficiency and potentially biased inference, particularly when lower-priority outcomes are incompletely observed. We propose the probabilistic win ratio (PWR), a framework for estimating the classical win ratio under coarsened observation. The PWR replaces deterministic pairwise decisions with conditional probabilities of win, loss, or tie given the observed data, allowing partially observed comparisons to contribute fractionally while being explicitly penalized according to their uncertainty. Comparisons with greater coarsening receive smaller effective weight, whereas fully observed comparisons contribute as in the classical analysis, preserving the clinical priority structure. When outcomes are fully observed, the PWR reduces exactly to the standard win ratio estimator. Simulation studies show that the PWR maintains low bias and mean squared error across a range of censoring and missingness scenarios. Two clinical trial case studies illustrate complementary data regimes, demonstrating calibration in near-complete data and stability under substantial right censoring.

stat.ME

Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization. We argue the root cause is not memorization itself, but its target: vanilla SFT drives models to exploit and memorize spurious surface correlations in problem-solution pairs, leaving them brittle to superficial input variations. To address this, we propose Theorem-SFT, which reorients supervision toward explicit theorem application by teaching models how rules are invoked rather than what answers look like. Theorem-SFT yields consistent gains across benchmarks and model families: +8.8% on MATH (LLaMA3.2-3B-Instruct) and +20.27% on GeoQA (Qwen2.5-VL-7B-Instruct) without modality-specific re-training. Fine-tuning MLP layers alone matches full-layers performance, implicating feed-forward components as the primary locus of reasoning rules. Our findings reframe the debate: Generalization failures stem not from memorization as a mechanism, but from memorizing the wrong inductive targets.

cs.LG

Socially Fluent, Socially Awkward: Artificial Intelligence Relational Talk Backfires in Commercial Interactions

Advancements in Artificial Intelligence (AI) technologies' social fluency are being integrated into commercial interactions. As tools such as OpenAI's assistant are integrated into platforms such as Shopify, Klarna, and Visa, understanding consumer responses to AI social features become essential. One such feature is relational talk, an informal and non-obligatory social communication embedded in transactional exchanges. Across four experiments, we find: 1) a negative main effect of AI relational talk on satisfaction, mediated by expectancy violation and perceived interaction awkwardness, and 2) goal-relevant relational talk to attenuate this effect. This paper extends the literature by challenging the assumption that increased social fluency will improve satisfaction, and highlights the complexity of integrating social features into AI systems. It also identifies awkwardness as a key emotional response and barrier to effective human-AI interaction, showing that even in the absence of real social repercussions, perceived awkwardness in AI-led commercial interactions can elicit negative responses.

cs.HC

Evaluating Black-Box Classifiers via Stable Adaptive Two-Sample Inference

We consider the problem of evaluating black-box multi-class classifiers. In the standard setup, we observe class labels $Y\in \{0,1,\ldots,M-1\}$ generated according to the conditional distribution $ Y|X \sim \text{ Multinom}\big(\eta(X)\big), $ where $X$ denotes the features and $\eta$ maps from the feature space to the $(M-1)$-dimensional simplex. A black-box classifier is an estimate $\hat{\eta}$ for which we make no assumptions about the training algorithm. Given holdout data, our goal is to evaluate the performance of the classifier $\hat{\eta}$. Recent work suggests treating this as a goodness-of-fit problem by testing the hypothesis $H_0: \rho((X,Y),(X',Y')) \le \delta$, where $\rho$ is some metric between two distributions, and $(X',Y')\sim P_X\times \text{ Multinom}(\hat\eta(X))$. Combining ideas from algorithmic fairness, Neyman-Pearson lemma, and conformal p-values, we propose a new methodology for this testing problem. The key idea is to generate a second sample $(X',Y') \sim P_X \times \text{ Multinom}\big(\hat\eta(X)\big)$ allowing us to reduce the task to two-sample conditional distribution testing. Using part of the data, we train an auxiliary binary classifier called a distinguisher to attempt to distinguish between the two samples. The distinguisher's ability to differentiate samples, measured using a rank-sum statistic, is then used to assess the difference between $\hat{\eta}$ and $\eta$ . Using techniques from cross-validation central limit theorems, we derive an asymptotically rigorous test under suitable stability conditions of the distinguisher.

stat.ME

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models

Multimodal large language models (MLLMs) often suffer from perceptual impairments under extended reasoning modes, particularly in visual question answering (VQA) tasks. We identify attention dispersion as the underlying cause: during multi-step reasoning, the model's visual attention becomes scattered and drifts away from question-relevant regions, effectively "losing focus" on the visual input. To better understand this phenomenon, we analyze the attention maps of MLLMs and observe that reasoning prompts significantly reduce attention to regions critical for answering the question. We further find a strong correlation between the model's overall attention on image tokens and the spatial dispersiveness of its attention within the image. Leveraging this insight, we propose a training-free Visual Region-Guided Attention (VRGA) framework that selects visual heads based on an entropy-focus criterion and reweights their attention, effectively guiding the model to focus on question-relevant regions during reasoning. Extensive experiments on vision-language benchmarks demonstrate that our method effectively alleviates perceptual degradation, leading to improvements in visual grounding and reasoning accuracy while providing interpretable insights into how MLLMs process visual information.

cs.CV

Sparse group principal component analysis via double thresholding with application to multi-cellular programs

Multi-cellular programs (MCPs) are coordinated patterns of gene expression across interacting cell types that collectively drive complex biological processes such as tissue development and immune responses. While MCPs are typically estimated from high-dimensional gene expression data using methods like sparse principal component analysis or latent factor models, these approaches often suffer from high computational costs and limited statistical power. In this work, we propose Sparse Group Principal Component Analysis (SGPCA) to estimate MCPs by leveraging their inherent group and individual sparsity. We introduce an efficient double-thresholding algorithm based on power iteration. In each iteration, a group thresholding step first identifies relevant gene groups, followed by an individual thresholding step to select active cell types. This algorithm achieves a linear computational complexity of $O(np)$, making it highly efficient and scalable for large-scale genomic analyses. We establish theoretical guarantees for SGPCA, including statistical consistency and a convergence rate that surpasses competing methods. Through extensive simulations, we demonstrate that SGPCA achieves superior estimation accuracy and improved statistical power for signal detection. Furthermore, We apply SGPCA to a Lupus study, discovering differentially expressed MCPs distinguishing Lupus patients from normal subjects.

stat.ME

A Low-Complexity Joint Fractional Delay and Doppler Frequency Estimator for AFDM-Enabled Vehicular LEO-ICAN Systems

Low-Earth-orbit (LEO) satellites and vehicle-to-everything (V2X) networks are driving integrated communication and navigation (ICAN) toward next-generation intelligent transportation. Affine frequency division multiplexing (AFDM) is a promising waveform for high-mobility LEO scenarios owing to its Doppler robustness, simple modulation, and low pilot overhead. However, applying existing high-accuracy AFDM fractional delay-Doppler estimators to LEO-ICAN entails substantial search or inference complexity, while the spectrum-wrapping-induced envelope structure in line-of-sight (LOS)-dominated channels remains underexploited. This paper analyzes and exploits the spectrum-wrapping-induced envelope structure of the fractional AFDM response, and proposes a low-complexity joint estimator that combines minimum-entropy fractional Doppler estimation with closed-form fractional delay estimation. Simulation results show that the proposed estimator approaches the root Cram\'er--Rao lower bound (RCRLB) and achieves root-mean-square error (RMSE) performance comparable to that of matched filtering (MF), matched filtering with generalized Fibonacci search (MF-GFS), and off-grid sparse Bayesian learning (OG-SBL), while requiring substantially lower computational complexity and runtime. This favorable accuracy-complexity profile highlights the potential of the proposed estimator for real-time ICAN processing in high-mobility LEO-assisted vehicular networks.

eess.SP

Specific Multi-emitter Identification: Theoretical Limits and Low-complexity Design

Specific emitter identification (SEI) distinguishes emitters by utilizing hardware-induced signal imperfections. However, conventional SEI techniques are primarily designed for single-emitter scenarios. This poses a fundamental limitation in distributed wireless networks, where simultaneous transmissions from multiple emitters result in overlapping signals that conventional single-emitter identification methods cannot effectively handle. To overcome this limitation, we present a specific multi-emitter identification (SMEI) framework via multi-label learning, treating identification as a problem of directly decoding emitter states from overlapping signals. Theoretically, we establish performance bounds using Fano's inequality. Methodologically, the multi-label formulation reduces output dimensionality from exponential to linear scale, thereby substantially decreasing computational complexity. Additionally, we propose an improved SMEI (I-SMEI), which incorporates multi-head attention to effectively capture features in correlated signal combinations. Experimental results demonstrate that SMEI achieves high identification accuracy with a linear computational complexity. Furthermore, the proposed I-SMEI scheme significantly improves identification accuracy across various overlapping scenarios compared to the proposed SMEI and other advanced methods.

eess.SP

Towards Efficient Inference under Nonmonotone Missingness with General Imputation

Missing data are ubiquitous in classical survey and longitudinal studies as well as modern multi-modality data analysis. A longstanding challenge arises under nonmonotone missingness, where different units may observe arbitrary subsets of all variables. We study parameter estimation and inference problem under this setting. Semiparametric efficiency theory characterizes the efficient estimator through inversion of an operator constructed from pattern-specific conditional expectations. However, this estimator is generally not tractable due to compositions of conditional expectations across patterns. We introduce the Restricted ANOVA hierarchY (RAY), a functional decomposition that reveals an almost-eigen structure of the operator under missing completely at random. This structure yields a closed-form, computable approximation to the efficient estimator. RAY estimator is applicable to general Z-estimation problems, and it remains unbiased for arbitrary independent imputation functions. In theory, we establish verifiable sufficient conditions where RAY attains the efficiency lower bound, and offer a general bound for the efficiency gap otherwise. We further develop adaptive RAY estimator, which attains the minimal asymptotic variance within a broader class containing RAY and other existing estimators. Finally, we investigate the extension of RAY under missing at random mechanism. Simulations and a single-cell multi-omics application demonstrate the efficiency gains of the proposed estimators.

stat.ME

Specific multi-emitter identification via multi-label learning

Specific emitter identification leverages hardware-induced impairments to uniquely determine a specific transmitter. However, existing approaches fail to address scenarios where signals from multiple emitters overlap. In this paper, we propose a specific multi-emitter identification (SMEI) method via multi-label learning to determine multiple transmitters. Specifically, the multi-emitter fingerprint extractor is designed to mitigate the mutual interference among overlapping signals. Then, the multi-emitter decision maker is proposed to assign the all emitter identification using the previous extracted fingerprint. Experimental results demonstrate that, compared with baseline approach, the proposed SMEI scheme achieves comparable identification accuracy under various overlapping conditions, while operating at significantly lower complexity. The significance of this paper is to identify multiple emitters from overlapped signal with a low complexity.

eess.SP

WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks

Graph Neural Networks (GNNs) are increasingly deployed in real-world applications, making ownership verification critical to protect their intellectual property against model theft. Fingerprinting and black-box watermarking are two main methods. However, the former relies on determining model similarity, which is computationally expensive and prone to ownership collisions after model post-processing. The latter embeds backdoors, exposing watermarked models to the risk of backdoor attacks. Moreover, both previous methods enable ownership verification but do not convey additional information about the copy model. If the owner has multiple models, each model requires a distinct trigger graph. To address these challenges, this paper proposes WGLE, a novel black-box watermarking paradigm for GNNs that enables embedding the multi-bit string in GNN models without using backdoors. WGLE builds on a key insight we term Layer-wise Distance Difference on an Edge (LDDE), which quantifies the difference between the feature distance and the prediction distance of two connected nodes in a graph. By assigning unique LDDE values to the edges and employing the LDDE sequence as the watermark, WGLE supports multi-bit capacity without relying on backdoor mechanisms. We evaluate WGLE on six public datasets across six mainstream GNN architectures, and compare WGLE with state-of-the-art GNN watermarking and fingerprinting methods. WGLE achieves 100% ownership verification accuracy, with an average fidelity degradation of only 1.41%. Additionally, WGLE exhibits robust resilience against potential attacks. The code is available in the repository.

cs.CR

A Modern Theory of Cross-Validation through the Lens of Stability

Modern data analysis and statistical learning are marked by complex data structures and black-box algorithms. Data complexity stems from technologies such as imaging, remote sensing, wearable devices, and genomic sequencing. At the same time, black-box models, especially deep neural networks, have achieved impressive results. This combination raises new challenges for uncertainty quantification and statistical inference, which we refer to as ``black-box inference.'' Black-box inference is difficult due to the lack of traditional modeling assumptions and the opaque behavior of modern estimators. These factors make it hard to characterize the distribution of estimation errors. A popular solution is post-hoc randomization, which, under mild assumptions such as exchangeability, can yield valid uncertainty quantification. Such methods range from classical techniques like permutation tests, the jackknife, and the bootstrap to more recent innovations like conformal inference. These approaches typically require little knowledge of data distributions or the internal workings of estimators. Many rely on the idea that estimators behave similarly under small perturbations of the data -- a concept formalized as stability. Over time, stability has become a key principle in data science, influencing research on generalization error, privacy, and adaptive inference. This article investigates cross-validation (CV) -- a widely used resampling method -- through the lens of stability. We first review recent theoretical results on CV for estimating generalization error and model selection under stability assumptions. We then examine uncertainty quantification for CV-based risk estimates. Together, these insights yield new theory and tools, which we apply to topics including model selection, selective inference, and conformal prediction.

math.ST

Semiparametric semi-supervised learning for general targets under distribution shift and decaying overlap

In modern scientific applications, large volumes of covariate data are readily available, while outcome labels are costly, sparse, and often subject to distribution shift. This asymmetry has spurred interest in semi-supervised (SS) learning, but most existing approaches rely on strong assumptions -- such as missing completely at random (MCAR) labeling or strict positivity -- that put substantial limitations on their practical usefulness. In this work, we introduce a general semiparametric framework for estimation, inference, and efficiency benchmarking in SS settings where labels are missing at random (MAR) and the overlap may vanish as sample size increases. Our framework, that we label D2S3, accommodates a wide range of smooth statistical targets -- including means, linear regression coefficients, quantiles, and causal effects -- and remains valid under high-dimensional nuisance estimation and distributional shift between labeled and unlabeled samples. We extend the theoretical guarantees of augmented inverse probability weighting estimators to preserve double robustness, asymptotic normality, and semiparametric efficiency under this challenging D2S3 regime. A key insight is that classical root-n convergence fails under vanishing overlap; we instead provide corrected asymptotic rates that capture the impact of the decay in overlap. We validate our theory through simulations and demonstrate practical utility in real-world applications on the internet of things and public health where labeled data are scarce.

math.ST

StablePCA: Distributionally Robust Learning of Shared Representations from Multi-Source Data

When synthesizing multi-source high-dimensional data, a key objective is to extract low-dimensional representations that effectively approximate the original features across different sources. Such representations facilitate the discovery of transferable structures and help mitigate systematic biases such as batch effects. We introduce Stable Principal Component Analysis (StablePCA), a distributionally robust framework for constructing stable latent representations by maximizing the worst-case explained variance over multiple sources. A primary challenge in extending classical PCA to the multi-source setting lies in the nonconvex rank constraint, which renders the StablePCA formulation a nonconvex optimization problem. To overcome this challenge, we conduct a convex relaxation of StablePCA and develop an efficient Mirror-Prox algorithm to solve the relaxed problem, with global convergence guarantees. Since the relaxed problem generally differs from the original formulation, we further introduce a data-dependent certificate to assess how well the algorithm solves the original nonconvex problem and establish the condition under which the relaxation is tight. Finally, we explore alternative distributionally robust formulations of multi-source PCA based on different loss functions.

cs.LG

Minimax Optimal Probability Matrix Estimation For Graphon With Spectral Decay

We study the optimal estimation of probability matrices of random graph models generated from graphons. This problem has been extensively studied in the case of step-graphons and H\"older smooth graphons. In this work, we characterize the regularity of graphons based on the decay rates of their eigenvalues. Our results show that for such classes of graphons, the minimax upper bound is achieved by a spectral thresholding algorithm and matches an information-theoretic lower bound up to a log factor. We provide insights on potential sources of this extra logarithm factor and discuss scenarios where exactly matching bounds can be obtained. This marks a difference from the step-graphon and H\"older smooth settings, because in those settings, there is a known computational-statistical gap where no polynomial time algorithm can achieve the statistical minimax rate. This contrast reflects a deeper observation that the spectral decay is an intrinsic feature of a graphon while smoothness is not.

math.ST