SearcharxivSearch

arXiv subjects

Wenjie Hu

Publications and source records attributed to Wenjie Hu.

At least 19 recordsLinked to original sources

LIGE-GR: A Smooth Leap from Ranking to Generative Recommendation in the LLM Era

The remarkable success of large language models (LLMs) has provided important inspiration for the next generation of recommender systems. Structurally, recommendation and language generation share a similarity: both aim to produce an ordered sequence that optimizes the user's experience. However, how to precisely absorb the essence of the LLM paradigm into mature industrial recommender systems remains an open problem. There are two challenges. First, it is unclear how to incorporate sequence-level generation and optimization from the LLM paradigm into recommendation. Second, real-world recommender systems are mature systems that have been iteratively customized for years around specific products, business constraints, serving infrastructure, and organizational ownership. Replacing such systems wholesale is often technically risky and organizationally disruptive. In this paper, we propose LIGE-GR, a listwise generation and evaluation recommendation framework that upgrades from a traditional ranking system based on itemwise recommendation toward a generative recommendation paradigm. Instead of rebuilding the entire recommendation stack from scratch, LIGE-GR generalizes the existing pointwise recommendation system into a listwise generation system. This allows mature recommender systems to benefit from listwise optimization while preserving compatibility with existing models, value functions, and serving infrastructure. We validate LIGE-GR in short-video recommendation on Instagram Reels and Facebook Video. On these recommendation surfaces, LIGE-GR improves time spent by 1.14 percent on Instagram Reels and 0.72 percent on Facebook Video, while requiring only modest additional inference resources.

cs.LG

Fractional-Quantum Ferroelectrics: A Route to High-Mobility Ferroelectric Semiconductors

Ferroelectric semiconductors are promising for multifunctional electronics, yet their typically low carrier mobilities remain a major limitation. Using first-principles phonon-limited transport calculations for monolayer In$_2$Se$_3$, we show that this limitation depends critically on the microscopic origin of ferroelectricity. In displacive $β'$-In$_2$Se$_3$, low-frequency ferroelectric modes dominate carrier scattering, with additional contributions from longitudinal-optical (LO) phonons, limiting the room-temperature electron mobility to a few cm$^2$/(V s). By contrast, in fractional-quantum ferroelectric $α$-In$_2$Se$_3$, ferroelectric-mode scattering is absent because polarization arises from discrete lattice-scale atomic displacements rather than soft-mode condensation. Transport is therefore dominated by LO phonons, yielding a room-temperature mobility above 70 cm$^2$/(V s). Carrier doping further screens long-range electron-LO-phonon interactions and raises the mobility beyond 300 cm$^2$/(V s) at experimentally accessible densities. These results establish that fractional-quantum ferroelectricity can decouple robust polarization from strong intrinsic carrier scattering, offering a route toward high-mobility ferroelectric semiconductors.

cond-mat.mtrl-sci

Simultaneous false discovery rate control in location families

When testing a number of statistical hypotheses using data from location families, it is often useful to control the false discovery rate (FDR) not just for hypotheses of the null values but also of other parameter values that are deemed practically insignificant. Here we consider FDR as a curve indexed by the location parameter and suggest a simple generalization of the Benjamini-Hochberg procedure that controls the FDR curve below any user-specified level. As a corollary of our main result, we show that the standard Benjamini-Hochberg procedure -- designed to control the FDR at the null -- also provides simultaneous control of the whole FDR curve for free. We further demonstrate the implications of our results and some practical considerations with a numerical example.

stat.ME

The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break

Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, interdependent action sequences. Despite rapid progress in agentic systems, these long-horizon failures remain poorly characterized, hindering principled diagnosis and comparison across domains. To address this gap, we introduce HORIZON, an initial cross-domain diagnostic benchmark for systematically constructing tasks and analyzing long-horizon failure behaviors in LLM-based agents. Using HORIZON, we evaluate state-of-the-art (SOTA) agents from multiple model families (GPT-5 variants and Claude models), collecting 3100+ trajectories across four representative agentic domains to study horizon-dependent degradation patterns. We further propose a trajectory-grounded LLM-as-a-Judge pipeline for scalable and reproducible failure attribution, and validate it with human annotation on trajectories, achieving strong agreement (inter-annotator κ=0.61; human-judge κ=0.84). Our findings offer an initial methodological step toward systematic, cross-domain analysis of long-horizon agent failures and offer practical guidance for building more reliable long-horizon agents. We release our project website at \href{https://xwang2775.github.io/horizon-leaderboard/}{HORIZON Leaderboard} and welcome contributions from the community.

cs.AI

Semiparametric Efficient Fusion of Individual Data and Summary Statistics

Suppose we have individual data from an internal study and various summary statistics from relevant external studies. External summary statistics have the potential to improve statistical inference for the internal population; however, it may lead to efficiency loss or bias if not used properly. We study the fusion of individual data and summary statistics in a semiparametric framework to investigate the efficient use of external summary statistics. Under a weak transportability assumption, we establish the semiparametric efficiency bound for estimating a general functional of the internal data distribution, which is no larger than that using only internal data and underpins the potential efficiency gain of integrating individual data and summary statistics. We propose a data-fused efficient estimator that achieves this efficiency bound. In addition, an adaptive fusion estimator is proposed to eliminate the bias of the original data-fused estimator when the transportability assumption fails. We establish the asymptotic oracle property of the adaptive fusion estimator. Simulations and application to a Helicobacter pylori infection dataset demonstrate the promising numerical performance of the proposed method.

stat.ME

Frugal coloring of graphs revisited

Given a graph $G$ and a positive integer $t$, an independent set $S\subseteq V(G)$ is $t$-frugal if every vertex has at most $t$ neighbors in $S$. A $t$-frugal coloring of $G$ is a partition of its vertex set into $t$-frugal independent sets. The maximum cardinality of a $t$-frugal independent set in $G$ is denoted by $α_t^f(G)$, while the minimum cardinality of a $t$-frugal coloring of $G$, $χ_t^f(G)$, is called the $t$-frugal chromatic number of $G$. Frugal colorings were introduced in 1998 and studied later in just a handful of papers. In this paper, we revisit this concept. While the NP-hardness of frugal coloring is known, we prove that the decision version of $α_t^f$ is NP-complete even for bipartite graphs, and present a linear-time algorithm to determine its value for trees. We prove a general sharp lower bound on $χ_{t}^{f}(G)$ expressed in terms of $α_{t}^{f}(G)$ and size of $G$. We also give a sharp upper bound on the $α_2^f$ of any graph $G$, which in the case of graphs with minimum degree $δ\geq2$ simplifies to $α_2^f(G)\le 2n/(δ+2)$. We prove that $3\leχ_2^f(G)\le 5$ holds for any graph $G$ with $Δ(G)=3$. For several classes of graphs such as block graphs, the Cartesian and strong products of multiple two-way infinite paths, we determine the exact values of $α_2^f$. We provide sharp bounds on the $α_2^f$ in all four standard graph products, which are expressed as different invariants of their factors. Finally, we obtain Nordhaus-Gaddum type inequalities for the sum of the $2$-frugal chromatic numbers of $G$ and its complement from below and from above by functions of the order of $G$. For the upper bound $χ_{2}^{f}(G)+χ_{2}^{f}(\overline{G})\leq 3n/2$, we characterize the family of extremal graphs $G$.

math.CO

Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention

Recent advances in Transformer-based Neural Operators have enabled significant progress in data-driven solvers for Partial Differential Equations (PDEs). Most current research has focused on reducing the quadratic complexity of attention to address the resulting low training and inference efficiency. Among these works, Transolver stands out as a representative method that introduces Physics-Attention to reduce computational costs. Physics-Attention projects grid points into slices for slice attention, then maps them back through deslicing. However, we observe that Physics-Attention can be reformulated as a special case of linear attention, and that the slice attention may even hurt the model performance. Based on these observations, we argue that its effectiveness primarily arises from the slice and deslice operations rather than interactions between slices. Building on this insight, we propose a two-step transformation to redesign Physics-Attention into a canonical linear attention, which we call Linear Attention Neural Operator (LinearNO). Our method achieves state-of-the-art performance on six standard PDE benchmarks, while reducing the number of parameters by an average of 40.0% and computational cost by 36.2%. Additionally, it delivers superior performance on two challenging, industrial-level datasets: AirfRANS and Shape-Net Car.

cs.LG

How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns

Large Language Models (LLMs) display strikingly different generalization behaviors: supervised fine-tuning (SFT) often narrows capability, whereas reinforcement-learning (RL) tuning tends to preserve it. The reasons behind this divergence remain unclear, as prior studies have largely relied on coarse accuracy metrics. We address this gap by introducing a novel benchmark that decomposes reasoning into atomic core skills such as calculation, fact retrieval, simulation, enumeration, and diagnostic, providing a concrete framework for addressing the fundamental question of what constitutes reasoning in LLMs. By isolating and measuring these core skills, the benchmark offers a more granular view of how specific cognitive abilities emerge, transfer, and sometimes collapse during post-training. Combined with analyses of low-level statistical patterns such as distributional divergence and parameter statistics, it enables a fine-grained study of how generalization evolves under SFT and RL across mathematical, scientific reasoning, and non-reasoning tasks. Our meta-probing framework tracks model behavior at different training stages and reveals that RL-tuned models maintain more stable behavioral profiles and resist collapse in reasoning skills, whereas SFT models exhibit sharper drift and overfit to surface patterns. This work provides new insights into the nature of reasoning in LLMs and points toward principles for designing training strategies that foster broad, robust generalization.

cs.LG

Free-carrier screening unlocks high electron mobility in ultrawide bandgap semiconductor CaSnO$_3$

Alkaline earth stannates have emerged as promising transparent conducting oxides due to their wide band gaps and high room-temperature electron mobilities. Among them, CaSnO$_3$ possesses the widest band gap, yet reported mobilities vary widely and are highly sample-dependent, leaving its intrinsic limit unclear. Here, we present ab initio calculations of electron mobility in CaSnO$_3$ across a range of temperatures and doping levels, using state-of-the-art methods that explicitly account for free-carrier screening in electron-phonon interactions. We identify the dominant limiting mechanism to be the long-range longitudinal optical phonon scattering, which is significantly suppressed at high doping due to free-carrier screening, leading to enhanced phonon-limited mobility. While ionized impurity scattering emerges as a competing mechanism at carrier concentrations up to ~10$^{20}$ cm$^{-3}$, the phonon scattering reduction dominates, yielding a net mobility increase with predicted room-temperature values reaching about twice the highest experimental report. Our work highlights the substantial untapped conductivity in CaSnO$_3$, establishing it as a compelling ultrawide bandgap semiconductor for transparent and high-power electronic applications.

cond-mat.mtrl-sci

Phonons Drive the Topological Phase Transition in Quasi-One-Dimensional Bi$_4$I$_4$

Quasi-one-dimensional bismuth halides offer an exceptional platform for exploring diverse topological phases, yet the nature of the room-temperature topological phase transition in Bi$_4$I$_4$ remains unresolved. While theory predicts the high-temperature $β$-phase to be a strong topological insulator (TI), experiments observe a weak TI. Here we resolve this discrepancy by revealing the critical but previously overlooked role of electron-phonon coupling in driving the topological phase transition. Using our newly developed ab initio framework for phonon-induced band renormalization, we show that thermal phonons alone drive $β$-Bi$_4$I$_4$ from the strong TI predicted by static-lattice calculations to a weak TI above ~180 K. At temperatures where $β$-Bi$_4$I$_4$ is stable, it is a weak TI with calculated surface states closely match experimental results, thereby reconciling theory with experiment. Our work establishes electron-phonon renormalization as essential for determining topological phases and provides a broadly applicable approach for predicting topological materials at finite temperatures.

cond-mat.mtrl-sci

CSTEapp: An interactive R-Shiny application of the covariate-specific treatment effect curve for visualizing individualized treatment rule

In precision medicine, deriving the individualized treatment rule (ITR) is crucial for recommending the optimal treatment based on patients' baseline covariates. The covariate-specific treatment effect (CSTE) curve presents a graphical method to visualize an ITR within a causal inference framework. Recent advancements have enhanced the causal interpretation of the CSTE curves and provided methods for deriving simultaneous confidence bands for various study types. To facilitate the implementation of these methods and make ITR estimation more accessible, we developed CSTEapp, a web-based application built on the R Shiny framework. CSTEapp allows users to upload data and create CSTE curves through simple point and click operations, making it the first application for estimating the ITRs. CSTEapp simplifies the analytical process by providing interactive graphical user interfaces with dynamic results, enabling users to easily report optimal treatments for individual patients based on their covariates information. Currently, CSTEapp is applicable to studies with binary and time-to-event outcomes, and we continually expand its capabilities to accommodate other outcome types as new methods emerge. We demonstrate the utility of CSTEapp using real-world examples and simulation datasets. By making advanced statistical methods more accessible, CSTEapp empowers researchers and practitioners across various fields to advance precision medicine and improve patient outcomes.

stat.CO

Hierarchical Deep Deterministic Policy Gradient for Autonomous Maze Navigation of Mobile Robots

Maze navigation is a fundamental challenge in robotics, requiring agents to traverse complex environments efficiently. While the Deep Deterministic Policy Gradient (DDPG) algorithm excels in control tasks, its performance in maze navigation suffers from sparse rewards, inefficient exploration, and long-horizon planning difficulties, often leading to low success rates and average rewards, sometimes even failing to achieve effective navigation. To address these limitations, this paper proposes an efficient Hierarchical DDPG (HDDPG) algorithm, which includes high-level and low-level policies. The high-level policy employs an advanced DDPG framework to generate intermediate subgoals from a long-term perspective and on a higher temporal scale. The low-level policy, also powered by the improved DDPG algorithm, generates primitive actions by observing current states and following the subgoal assigned by the high-level policy. The proposed method enhances stability with off-policy correction, refining subgoal assignments by relabeling historical experiences. Additionally, adaptive parameter space noise is utilized to improve exploration, and a reshaped intrinsic-extrinsic reward function is employed to boost learning efficiency. Further optimizations, including gradient clipping and Xavier initialization, are employed to improve robustness. The proposed algorithm is rigorously evaluated through numerical simulation experiments executed using the Robot Operating System (ROS) and Gazebo. Regarding the three distinct final targets in autonomous maze navigation tasks, HDDPG significantly overcomes the limitations of standard DDPG and its variants, improving the success rate by at least 56.59% and boosting the average reward by a minimum of 519.03 compared to baseline algorithms.

cs.RO

Marlin: Efficient Coordination for Autoscaling Cloud DBMS (Extended Version)

Modern cloud databases are shifting from converged architectures to storage disaggregation, enabling independent scaling and billing of compute and storage. However, cloud databases still rely on external, converged coordination services (e.g., ZooKeeper) for their control planes. These services are effectively lightweight databases optimized for low-volume metadata. As the control plane scales in the cloud, this approach faces similar limitations as converged databases did before storage disaggregation: scalability bottlenecks, low cost efficiency, and increased operational burden. We propose to disaggregate the cluster coordination to achieve the same benefits that storage disaggregation brought to modern cloud DBMSs. We present Marlin, a cloud-native coordination mechanism that fully embraces storage disaggregation. Marlin eliminates the need for external coordination services by consolidating coordination functionality into the existing cloud-native database it manages. To achieve failover without an external coordination service, Marlin allows cross-node modifications on coordination states. To ensure data consistency, Marlin employs transactions to manage both coordination and application states and introduces MarlinCommit, an optimized commit protocol that ensures strong transactional guarantees even under cross-node modifications. Our evaluations demonstrate that Marlin improves cost efficiency by up to 4.4x and reduces reconfiguration duration by up to 4.9x compared to converged coordination solutions.

cs.DB

Closing the Evaluation Gap: Developing a Behavior-Oriented Framework for Assessing Virtual Teamwork Competency

The growing reliance on remote work and digital collaboration has made virtual teamwork competencies essential for professional and academic success. However, the evaluation of such competencies remains a significant challenge. Existing assessment methods, predominantly based on self-reports and peer evaluations, often focus on short-term results or subjective perceptions rather than systematically examining observable teamwork behaviors. These limitations hinder the identification of specific areas for improvement and fail to support meaningful progress in skill development. Informed by group dynamic theory, this study developed a behavior-oriented framework for assessing virtual teamwork competencies among engineering students. Using focus group interviews combined with the Critical Incident Technique, the study identified three key dimensions - Group Task Dimension, Individual Task Dimension and Social Dimension - along with their behavioral indicators and student-perceived relationships between these components. The resulting framework provides a foundation for more effective assessment practices and supports the development of virtual teamwork competency essential for success in increasingly digital and globalized professional environments.

cs.CY

A self-learning magnetic Hopfield neural network with intrinsic gradient descent adaption

Physical neural networks using physical materials and devices to mimic synapses and neurons offer an energy-efficient way to implement artificial neural networks. Yet, training physical neural networks are difficult and heavily relies on external computing resources. An emerging concept to solve this issue is called physical self-learning that uses intrinsic physical parameters as trainable weights. Under external inputs (i.e. training data), training is achieved by the natural evolution of physical parameters that intrinsically adapt modern learning rules via autonomous physical process, eliminating the requirements on external computation resources.Here, we demonstrate a real spintronic system that mimics Hopfield neural networks (HNN) and unsupervised learning is intrinsically performed via the evolution of physical process. Using magnetic texture defined conductance matrix as trainable weights, we illustrate that under external voltage inputs, the conductance matrix naturally evolves and adapts Oja's learning algorithm in a gradient descent manner. The self-learning HNN is scalable and can achieve associative memories on patterns with high similarities. The fast spin dynamics and reconfigurability of magnetic textures offer an advantageous platform towards efficient autonomous training directly in materials.

cond-mat.dis-nn

Identifying average causal effect in regression discontinuity design with auxiliary data

Regression discontinuity designs are widely used when treatment assignment is determined by whether a running variable exceeds a predefined threshold. However, most research focuses on estimating local causal effects at the threshold, leaving the challenge of identifying treatment effects away from the cutoff largely unaddressed. The primary difficulty in this context is that the treatment assignment is deterministically defined by the running variable, violating the commonly assumed positivity assumption. In this paper, we introduce a novel framework for identifying the average causal effect in regression discontinuity designs. Our approach assumes the existence of an auxiliary variable for which the running variable can be seen as a surrogate, and an additional dataset that consists of the running variable and the auxiliary variable alongside the traditional regression discontinuity design setup. Under this framework, we propose three estimation methods for the ATE, which resembles the outcome regression, inverse propensity weighted and doubly robust estimators in classical causal inference literature. Asymptotically valid inference procedures are also provided. To demonstrate the practical application of our method, simulations are conducted to show the good performance of our methods; besides, we use the proposed methods to assess the causal effects of vitamin A supplementation on the severity of autism spectrum disorders in children, where a positive effect is found but with no statistical significance.

stat.ME

Improve Sensitivity Analysis Synthesizing Randomized Clinical Trials With Limited Overlap

Randomized clinical trials are the gold standard when estimating the average treatment effect. However, they are usually not a random sample from the real-world population because of the inclusion/exclusion rules. Meanwhile, observational studies typically consist of representative samples from the real-world population. However, due to unmeasured confounding, sensitivity analysis is often used to estimate bounds for the average treatment effect without relying on stringent assumptions of other existing methods. This article introduces a synthesis estimator that improves sensitivity analysis in observational studies by incorporating randomized clinical trial data, even when overlap in covariate distribution is limited due to inclusion/exclusion criteria. We show that the proposed estimator will give a tighter bound when a "separability" condition holds for the sensitivity parameter. Theoretical proofs and simulations show that this method provides a tighter bound than the sensitivity analysis using only observational study. We apply this method to combine an observational study on drug effectiveness with a partially overlapping RCT dataset, yielding improved average treatment effect bounds.

stat.ME

A constructive approach to selective risk control

Many modern applications require using data to select the statistical tasks and make valid inference after selection. In this article, we provide a unifying approach to control for a class of selective risks. Our method is motivated by a reformulation of the celebrated Benjamini-Hochberg (BH) procedure for multiple hypothesis testing as the fixed point iteration of the Benjamini-Yekutieli (BY) procedure for constructing post-selection confidence intervals. Building on this observation, we propose a constructive approach to control extra-selection risk (where selection is made after decision) by iterating decision strategies that control the post-selection risk (where decision is made after selection). We show that many previous methods and results are special cases of this general framework, and we further extend this approach to problems with multiple selective risks. Our development leads to two surprising results about the BH procedure: (1) in the context of one-sided location testing, the BH procedure not only controls the false discovery rate at the null but also at other locations for free; (2) in the context of permutation tests, the BH procedure with exact permutation p-values can be well approximated by a procedure which only requires a total number of permutations that is almost linear in the total number of hypotheses.

stat.ME