SearcharxivSearch

arXiv subjects

Mengyan Li

Publications and source records attributed to Mengyan Li.

At least 19 recordsLinked to original sources

RecGPT-V3 Technical Report

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior modeling, where each request reprocesses full user history, wasting computation and discarding prior analysis; (2) a tag-to-item information bottleneck, where natural-language tags form a lossy channel between user understanding and item grounding; and (3) inefficient explicit reasoning, whose lengthy chain-of-thought incurs untenable latency and compute overhead. We present RecGPT-V3, a stateful, hybrid-modal recommender that reasons over natural language for open-world knowledge and Semantic IDs (SIDs) for concrete item grounding. A Memory Hub maintains structured, continually evolving user memory that distills long-horizon behavior into condensed units, cutting user-modeling computation by 55.8%. A Hybrid-modal Foundation Model allows the LLM jointly reason over text tags and SIDs, opening a high-bandwidth channel into the item space. Latent Intent Reasoning internalizes verbose rationales into compact learnable latent tokens that remain decodable into readable explanations, lowering output token cost by 200x. Deployed in Taobao's "Guess What You Like" feed, RecGPT-V3 achieves consistent gains in large-scale online A/B tests: IPV +1.28%, CTR +1.00%, TC +1.97%, GMV +3.97%, while cutting end-to-end serving resource consumption by 52.4%.

cs.IR

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping

The wave of AI-native applications is moving shopping beyond page- and feed-based browsing toward intent-driven experiences orchestrated by LLM agents. A common design wraps an LLM around existing search and recommendation pipelines, forcing complex intents through low-bandwidth retrieval or ranking interfaces and leaving a gap between language understanding and item-space fulfillment. Generative recommendation gives LLMs a direct item-space interface through semantic IDs (SIDs), but existing models mainly generate candidates for retrieval rather than translate flexible intents into item-space outcomes. We propose ShopX to address this bottleneck by unifying intent understanding, execution planning, and flexible SID-native item-space operations into a single foundation model. We deploy ShopX in agentic shopping workflows through a model-native item-fulfillment framework with a serving harness that defines a model-facing action protocol and exposes support surfaces for context access, catalog grounding, and state management. Within this framework, ShopX plans and composes SID-based item-space operations such as SID beam-search retrieval, listwise ranking, or product bundling. This model-centric design reduces lossy hand-offs between agent orchestration and item-space execution. To build ShopX, we design semantically recoverable, LLM-operable SIDs and a training recipe that equips a general LLM for flexible multi-turn item-space fulfillment while retaining the knowledge and instruction-following abilities needed by a shopping agent. We evaluate the ShopX framework against tool-mediated agentic systems on single- and multi-turn fulfillment tasks derived from anonymized Taobao production logs, showing that model-native fulfillment improves overall framework behavior, especially on complex or ambiguous requests.

cs.IR

Semidefinite-programming hierarchies for classically simulable state families

Identifying whether a state family admits an irreducible quantum advantage is a fundamental task in quantum resource theory and quantum information processing. Here we study classically simulable state families, namely those residing within the convex hull of pairwise commuting families and therefore admitting a classical explanation. We develop a complete semidefinite-programming (SDP) hierarchy characterizing the set of classically simulable state families in arbitrary finite dimension. The key step is to reformulate classical simulability as a feasibility problem over deterministic response functions and auxiliary positive-operator-valued measures (POVMs) simulable by rank-one projective measurements. We establish a complete SDP hierarchy for rank-one projectively simulable POVMs and transfer the resulting characterization to state families, yielding both primal feasibility tests and dual affine witnesses certifying failure of classical simulability. Applying the hierarchy to state families mixed with depolarizing noise gives computable upper bounds on the critical classical visibility, which are matched by explicit classical simulations in several symmetric examples. These results provide a systematic convex-optimization framework for certifying classical simulability of quantum state families.

quant-ph

Geometric Construction of Optimal Teleportation Witnesses

Not all entangled states are useful for quantum teleportation. We present a geometric method to construct optimal teleportation witnesses, which provide operational necessary and sufficient criteria for identifying the teleportation usefulness of arbitrary two-qudit entangled states. Specifically, by developing a two-layer iterative cutting-plane algorithm to solve the shortest distance problem from the target state $\rho$ to the convex set $S$ of useless states, we obtain the projection point $\sigma^* \in S$ and then construct the optimal teleportation witness from the projection geometry. Moreover, the shortest distance $D(\rho)$ obtained during this construction also serves as a necessary and sufficient criterion for usefulness. We apply our method to identify the teleportation usefulness of three classes of entangled states.

quant-ph

Transfer Learning with Network Embeddings under Structured Missingness

Modern data-driven applications increasingly rely on large, heterogeneous datasets collected across multiple sites. Differences in data availability, feature representation, and underlying populations often induce structured missingness, complicating efforts to transfer information from data-rich settings to those with limited data. Many transfer learning methods overlook this structure, limiting their ability to capture meaningful relationships across sites. We propose TransNEST (Transfer learning with Network Embeddings under STructured missingness), a framework that integrates graphical data from source and target sites with prior group structure to construct and refine network embeddings. TransNEST accommodates site-specific features, captures within-group heterogeneity and between-site differences adaptively, and improves embedding estimation under partial feature overlap. We establish the convergence rate for the TransNEST estimator and demonstrate strong finite-sample performance in simulations. We apply TransNEST to a multi-site electronic health record study, transferring feature embeddings from a general hospital system to a pediatric hospital system. Using a hierarchical ontology structure, TransNEST improves pediatric embeddings and supports more accurate pediatric knowledge extraction, achieving the best accuracy for identifying pediatric-specific relational feature pairs compared with benchmark methods.

stat.ME

Semi-device-independent certification of high-dimensional quantum channels

Certifying high-dimensional quantum channels is essential for ensuring the reliability of quantum communication protocols. Existing certification schemes often rely on fully trusted internal devices, which is difficult to achieve in realistic scenarios. Here, we propose a semi-device-independent framework for certifying channel properties directly from observed statistics, assuming only that the system dimension is known. By explicitly incorporating the full set of structural constraints inherent to Choi states, our approach exploits the Choi-Jamio{\l}kowski isomorphism for rigorous certification of quantum channels. The entanglement dimensionality of quantum channels is first certified by introducing a witness and numerically determining its Schmidt-number-dependent bounds. This certification method reproduces known analytical benchmarks and is applied to dephasing and depolarizing noise channels, thereby confirming its validity. To provide a more complete assessment of channel performance, the entanglement fidelity of quantum channels is also certified using a hierarchy of semidefinite programming relaxations based on localizing matrices. Lower bounds on the entanglement fidelity are obtained that are compatible with either the full set of observed statistics or a single witness value.

quant-ph

HiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression

Generating structured narrations for real-world e-commerce videos requires models to perceive fine-grained visual details and organize them into coherent, high-level stories--capabilities that existing approaches struggle to unify. We introduce the E-commerce Hierarchical Video Captioning (E-HVC) dataset with dual-granularity, temporally grounded annotations: a Temporal Chain-of-Thought that anchors event-level observations and Chapter Summary that compose them into concise, story-centric summaries. Rather than directly prompting chapters, we adopt a staged construction that first gathers reliable linguistic and visual evidence via curated ASR and frame-level descriptions, then refines coarse annotations into precise chapter boundaries and titles conditioned on the Temporal Chain-of-Thought, yielding fact-grounded, time-aligned narratives. We also observe that e-commerce videos are fast-paced and information-dense, with visual tokens dominating the input sequence. To enable efficient training while reducing input tokens, we propose the Scene-Primed ASR-anchored Compressor (SPA-Compressor), which compresses multimodal tokens into hierarchical scene and event representations guided by ASR semantic cues. Built upon these designs, our HiVid-Narrator framework achieves superior narrative quality with fewer input tokens compared to existing methods.

cs.CV

A General Framework for Constructing Local Hidden-state Models to Determine the Steerability

Not all entangled states can exhibit quantum steering, and determining whether a given entangled state is steerable is a crucial problem in quantum information theory. The main challenge lies in verifying the existence of a local hidden-state (LHS) model capable of reproducing all post-measurement assemblages generated by arbitrary measurements. To address this, we propose a machine learning-based framework that employs batch sampling of measurements and gradient-based optimization to construct an optimal LHS model. We validate our method by analyzing the steerability of two-qubit Werner and two-qutrit isotropic states. For Werner states, our approach saturates the analytical visibility bounds under three Pauli measurements, arbitrary projective measurements (PVMs), and arbitrary positive operator-valued measurements (POVMs). For isotropic states, we achieve the known analytical bounds under arbitrary PVMs. We further investigate the steerability of this class of states under arbitrary POVMs, and our results suggest that POVMs can offer an advantage over PVMs in revealing the steerability of such states.

quant-ph

Advanced spectral clustering for heterogeneous data in credit risk monitoring systems

Heterogeneous data, which encompass both numerical financial variables and textual records, present substantial challenges for credit monitoring. To address this issue, we propose Advanced Spectral Clustering (ASC), a method that integrates financial and textual similarities through an optimized weight parameter and selects eigenvectors using a novel eigenvalue-silhouette optimization approach. Evaluated on a dataset comprising 1,428 small and medium-sized enterprises (SMEs), ASC achieves a Silhouette score that is 18% higher than that of a single-type data baseline method. Furthermore, the resulting clusters offer actionable insights; for instance, 51% of low-risk firms are found to include the term 'social recruitment' in their textual records. The robustness of ASC is confirmed across multiple clustering algorithms, including k-means, k-medians, and k-medoids, with {\Delta}Intra/Inter < 0.13 and {\Delta}Silhouette Coefficient < 0.02. By bridging spectral clustering theory with heterogeneous data applications, ASC enables the identification of meaningful clusters, such as recruitment-focused SMEs exhibiting a 30% lower default risk, thereby supporting more targeted and effective credit interventions.

cs.LG

A Weakly Supervised Transformer for Rare Disease Diagnosis and Subphenotyping from EHRs with Pulmonary Case Studies

Rare diseases affect an estimated 300-400 million people worldwide, yet individual conditions remain underdiagnosed and poorly characterized due to their low prevalence and limited clinician familiarity. Computational phenotyping offers a scalable approach to improving rare disease detection, but algorithm development is hindered by the scarcity of high-quality labeled data for training. Expert-labeled datasets from chart reviews and registries are clinically accurate but limited in scope and availability, whereas labels derived from electronic health records (EHRs) provide broader coverage but are often noisy or incomplete. To address these challenges, we propose WEST (WEakly Supervised Transformer for rare disease phenotyping and subphenotyping from EHRs), a framework that combines routinely collected EHR data with a limited set of expert-validated cases and controls to enable large-scale phenotyping. At its core, WEST employs a weakly supervised transformer model trained on extensive probabilistic silver-standard labels - derived from both structured and unstructured EHR features - that are iteratively refined during training to improve model calibration. We evaluate WEST on two rare pulmonary diseases using EHR data from Boston Children's Hospital and show that it outperforms existing methods in phenotype classification, identification of clinically meaningful subphenotypes, and prediction of disease progression. By reducing reliance on manual annotation, WEST enables data-efficient rare disease phenotyping that improves cohort definition, supports earlier and more accurate diagnosis, and accelerates data-driven discovery for the rare disease community.

cs.LG

Semi-supervised Clustering Through Representation Learning of Large-scale EHR Data

Electronic Health Records (EHR) offer rich real-world data for personalized medicine, providing insights into disease progression, treatment responses, and patient outcomes. However, their sparsity, heterogeneity, and high dimensionality make them difficult to model, while the lack of standardized ground truth further complicates predictive modeling. To address these challenges, we propose SCORE, a semi-supervised representation learning framework that captures multi-domain disease profiles through patient embeddings. SCORE employs a Poisson-Adapted Latent factor Mixture (PALM) Model with pre-trained code embeddings to characterize codified features and extract meaningful patient phenotypes and embeddings. To handle the computational challenges of large-scale data, it introduces a hybrid Expectation-Maximization (EM) and Gaussian Variational Approximation (GVA) algorithm, leveraging limited labeled data to refine estimates on a vast pool of unlabeled samples. We theoretically establish the convergence of this hybrid approach, quantify GVA errors, and derive SCORE's error rate under diverging embedding dimensions. Our analysis shows that incorporating unlabeled data enhances accuracy and reduces sensitivity to label scarcity. Extensive simulations confirm SCORE's superior finite-sample performance over existing methods. Finally, we apply SCORE to predict disability status for patients with multiple sclerosis (MS) using partially labeled EHR data, demonstrating that it produces more informative and predictive patient embeddings for multiple MS-related conditions compared to existing approaches.

stat.ME

Measuring network quantum steerability utilizing artificial neural networks

Network quantum steering plays a pivotal role in quantum information science, enabling robust certification of quantum correlations in scenarios with asymmetric trust assumptions among network parties. The intricate nature of quantum networks, however, poses significant challenges for the detection and quantification of steering. In this work, we develop a neural network-based method for measuring network quantum steerability, which can be generalized to arbitrary quantum networks and naturally applied to standard steering scenarios. Our method provides an effective framework for steerability analysis, demonstrating remarkable accuracy and efficiency in standard bipartite and multipartite steering scenarios. Numerical simulations involving isotropic states and noisy GHZ states yield results that are consistent with established findings in these respective scenarios. Furthermore, we demonstrate its utility in the bilocal network steering scenario, where an untrusted central party shares two-qubit isotropic states of different visibilities, $ν$ and $ω$, with trusted endpoint parties and performs a single Bell state measurement. Through explicit construction of a network local hidden state model derived from numerical results and incorporation of the entanglement properties of network assemblages, we analytically demonstrate that the network steering thresholds are determined by the curve $νω= {1}/{3}$ under the corresponding configuration.

quant-ph

Self-Testing Positive Operator-Valued Measurements and Certifying Randomness

In the device-independent scenario, positive operator-valued measurements (POVMs) can certify more randomness than projective measurements. This paper self-tests a three-outcome extremal qubit POVM in the X-Z plane of the Bloch sphere by achieving the maximal quantum violation of a newly constructed Bell expression C'3, adapted from the chained inequality C3. Using this POVM, approximately 1.58 bits of local randomness can be certified, which is the maximum amount of local randomness achievable by an extremal qubit POVM in this plane. Further modifications of C'3 produce C''3, enabling the self-testing of another three-outcome extremal qubit POVM. Together, these POVMs certify about 2.27 bits of global randomness. Both local and global randomness surpass the limitations certified from projective measurements. Additionally, the Navascués-Pironio-Acín hierarchy is employed to compare the lower bounds on global randomness certified by C3 and several other inequalities. As the extent of violation increases, C3 demonstrates superior performance in randomness certification.

quant-ph

Semi-supervised Triply Robust Inductive Transfer Learning

In this work, we propose a Semi-supervised Triply Robust Inductive transFer LEarning (STRIFLE) approach, which integrates heterogeneous data from a label-rich source population and a label-scarce target population and utilizes a large amount of unlabeled data simultaneously to improve the learning accuracy in the target population. Specifically, we consider a high dimensional covariate shift setting and employ two nuisance models, a density ratio model and an imputation model, to combine transfer learning and surrogate-assisted semi-supervised learning strategies effectively and achieve triple robustness. While the STRIFLE approach assumes the target and source populations to share the same conditional distribution of outcome Y given both the surrogate features S and predictors X, it allows the true underlying model of Y|X to differ between the two populations due to the potential covariate shift in S and X. Different from double robustness, even if both nuisance models are misspecified or the distribution of Y|(S, X) is not the same between the two populations, the triply robust STRIFLE estimator can still partially use the source population when the shifted source population and the target population share enough similarities. Moreover, it is guaranteed to be no worse than the target-only surrogate-assisted semi-supervised estimator with an additional error term from transferability detection. These desirable properties of our estimator are established theoretically and verified in finite samples via extensive simulation studies. We utilize the STRIFLE estimator to train a Type II diabetes polygenic risk prediction model for the African American target population by transferring knowledge from electronic health records linked genomic data observed in a larger European source population.

stat.ME

Domain Adaptation Targeting Heterogeneous and Imbalanced Subgroups

Domain adaptation enables generalizable and efficient data-driven research. However, existing work has largely focused on domain adaptation for some intrinsically homogeneous target cohort, overlooking inherent heterogeneity within the target, which can exacerbate biases and unfairness in the presence of subgroups with imbalanced sample sizes. We develop a novel domain adaptation framework that addresses more complicated target data that consists of heterogeneous and data-sparse subgroups and lacks gold-standard label observations. Our method simultaneously handles high-dimensionality, covariate shift, and outcome model heterogeneity by combining a model-assisted debiasing step used for covariate shift correction with an adaptive knowledge-guided sparsification procedure used to mitigate the issue of sample disparity. We also introduce a new model selection strategy to avoid negative knowledge transfer in the absence of labels in the target data. Our method is theoretically justified for being robust to nuisance model misspecification and adaptive to heterogeneity between the subgroups. Numerical experiments and two real-world applications, including genetic risk modeling of type II diabetes and prediction of mutation-induced protein stability changes, demonstrate the practical advantages of our method.

stat.ME

DATE: Domain Adaptive Product Seeker for E-commerce

Product Retrieval (PR) and Grounding (PG), aiming to seek image and object-level products respectively according to a textual query, have attracted great interest recently for better shopping experience. Owing to the lack of relevant datasets, we collect two large-scale benchmark datasets from Taobao Mall and Live domains with about 474k and 101k image-query pairs for PR, and manually annotate the object bounding boxes in each image for PG. As annotating boxes is expensive and time-consuming, we attempt to transfer knowledge from annotated domain to unannotated for PG to achieve un-supervised Domain Adaptation (PG-DA). We propose a {\bf D}omain {\bf A}daptive Produc{\bf t} S{\bf e}eker ({\bf DATE}) framework, regarding PR and PG as Product Seeking problem at different levels, to assist the query {\bf date} the product. Concretely, we first design a semantics-aggregated feature extractor for each modality to obtain concentrated and comprehensive features for following efficient retrieval and fine-grained grounding tasks. Then, we present two cooperative seekers to simultaneously search the image for PR and localize the product for PG. Besides, we devise a domain aligner for PG-DA to alleviate uni-modal marginal and multi-modal conditional distribution shift between source and target domains, and design a pseudo box generator to dynamically select reliable instances and generate bounding boxes for further knowledge transfer. Extensive experiments show that our DATE achieves satisfactory performance in fully-supervised PR, PG and un-supervised PG-DA. Our desensitized datasets will be publicly available here\footnote{\url{https://github.com/Taobao-live/Product-Seeking}}.

cs.CV

Doubly Robust Augmented Model Accuracy Transfer Inference with High Dimensional Features

Due to label scarcity and covariate shift happening frequently in real-world studies, transfer learning has become an essential technique to train models generalizable to some target populations using existing labeled source data. Most existing transfer learning research has been focused on model estimation, while there is a paucity of literature on transfer inference for model accuracy despite its importance. We propose a novel $\mathbf{D}$oubly $\mathbf{R}$obust $\mathbf{A}$ugmented $\mathbf{M}$odel $\mathbf{A}$ccuracy $\mathbf{T}$ransfer $\mathbf{I}$nferen$\mathbf{C}$e (DRAMATIC) method for point and interval estimation of commonly used classification performance measures in an unlabeled target population using labeled source data. Specifically, DRAMATIC derives and evaluates the risk model for a binary response $Y$ against some low dimensional predictors $\mathbf{A}$ on the target population, leveraging $Y$ from source data only and high dimensional adjustment features $\mathbf{X}$ from both the source and target data. The proposed estimators are doubly robust in the sense that they are $n^{1/2}$ consistent when at least one model is correctly specified and certain model sparsity assumptions hold. Simulation results demonstrate that the point estimation have negligible bias and the confidence intervals derived by DRAMATIC attain satisfactory empirical coverage levels. We further illustrate the utility of our method to transfer the genetic risk prediction model and its accuracy evaluation for type II diabetes across two patient cohorts in Mass General Brigham (MGB) collected using different sampling mechanisms and at different time points.

stat.ME

Deep Fusion Siamese Network for Automatic Kinship Verification

Automatic kinship verification aims to determine whether some individuals belong to the same family. It is of great research significance to help missing persons reunite with their families. In this work, the challenging problem is progressively addressed in two respects. First, we propose a deep siamese network to quantify the relative similarity between two individuals. When given two input face images, the deep siamese network extracts the features from them and fuses these features by combining and concatenating. Then, the fused features are fed into a fully-connected network to obtain the similarity score between two faces, which is used to verify the kinship. To improve the performance, a jury system is also employed for multi-model fusion. Second, two deep siamese networks are integrated into a deep triplet network for tri-subject (i.e., father, mother and child) kinship verification, which is intended to decide whether a child is related to a pair of parents or not. Specifically, the obtained similarity scores of father-child and mother-child are weighted to generate the parent-child similarity score for kinship verification. Recognizing Families In the Wild (RFIW) is a challenging kinship recognition task with multiple tracks, which is based on Families in the Wild (FIW), a large-scale and comprehensive image database for automatic kinship recognition. The Kinship Verification (track I) and Tri-Subject Verification (track II) are supported during the ongoing RFIW2020 Challenge. Our team (ustc-nelslip) ranked 1st in track II, and 3rd in track I. The code is available at https://github.com/gniknoil/FG2020-kinship.

cs.CV