SearcharxivSearch

arXiv subjects

Wuliang Huang

Publications and source records attributed to Wuliang Huang.

9 recordsLinked to original sources

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance exhibit diminishing performance gains with larger-scale raw text user behavioral input, which can be mitigated by tokenization. (ii) Lack of quantitative analysis of how tokenization configurations should scale with data size. In this report, we propose User Behavioral Densing Law for characterizing the quantitative relationship between data scale and the minimum sufficient tokenization capacity. Firstly, we conduct a pilot study on raw & tokenized scaling comparison on billion-scale Alipay dataset, revealing the raw data scaling bottleneck and the sustained gains enabled by tokenization. To derive the scaling pattern governing the minimum sufficient tokenization configuration at different data scales, theoretical analysis and systematic experiments are employed to summarize the quantitative scaling pattern. We find an approximately linear relationship between the logarithms of minimum sufficient tokenization capacity and input data size measured by tokens, and the scaling slope varies systematically with the tokenization method and data source, reflecting differences in representation-space redundancy and intra-source uniqueness. Guided by the proposed law, we further develop ALGN, an adaptive variable-length tokenization method that improves capacity allocation. Extensive experiments across diverse data sources, tokenization methods, and downstream tasks demonstrate the generalizability and reliability of the User Behavioral Densing Law, providing practical guidance for tokenization configuration selection in large-scale user representation learning. Moreover, ALGN outperforms existing baselines.

cs.IR

How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning

Decoder-only large language models are increasingly used as behavioral encoders for user representation learning, yet the impact of attention masking on the quality of user embeddings remains underexplored. In this work, we conduct a systematic study of causal, hybrid, and bidirectional attention masks within a unified contrastive learning framework trained on large-scale real-world Alipay data that integrates long-horizon heterogeneous user behaviors. To improve training dynamics when transitioning from causal to bidirectional attention, we propose Gradient-Guided Soft Masking, a gradient-based pre-warmup applied before a linear scheduler that gradually opens future attention during optimization. Evaluated on 9 industrial user cognition benchmarks covering prediction, preference, and marketing sensitivity tasks, our approach consistently yields more stable training and higher-quality bidirectional representations compared with causal, hybrid, and scheduler-only baselines, while remaining compatible with decoder pretraining. Overall, our findings highlight the importance of masking design and training transition in adapting decoder-only LLMs for effective user representation learning. Our code is available at https://github.com/JhCircle/Deepfind-GGSM.

cs.CL

FOUNDv2: Learning Unified User Quantized Tokenizers for User Representation

User representation learning serves as a fundamental pillar for personalized services on large-scale web platforms. Despite its importance, conventional continuous embedding methods face significant challenges, including the lack of a unified paradigm for multi-source data integration, prohibitive storage overhead due to low information density, and the lack of multi-scale modeling granularity. To overcome these limitations, we introduce FOUNDv2, a comprehensive user representation scheme centered on the Unified User Quantized Tokenizer U2QT) framework. FOUNDv2 transforms heterogeneous user data into a standardized discrete token space through a robust two-stage architecture. Specifically, the framework first extracts compact feature representations and subsequently employs a multi-view RQ-VAE to discretize them into storage-efficient tokens using shared and source-specific codebooks. To empower these representations with predictive intelligence, we further design multi-scale alignment objectives to capture both fine-grained behavioral dependencies and macro-temporal periodicity. Extensive experiments on various benchmarks demonstrate that FOUNDv2 consistently outperforms task-specific baselines while achieving substantial reductions in storage and computational costs. Finally, the large-scale deployment of FOUNDv2 on Alipay validates its practical scalability and efficiency across diverse industrial scenarios. The main code is available at: https://github.com/chuanhe1999/FOUNDv2.

cs.LG

Dark Matter, Quasars, and Superstructures in the Universe (with delta-particle search and spherical universe)

From the observed results of the space distribution of quasars we deduced that neutrino mass is about10^(-1) eV. The fourth stable elementary particle (delta particle) with mass about 10^(0) eV can help explain the energy resource mechanism in quasars, the peaks of Extra-galactic Background Light (EBL) etc. The key is the annihilation of dark matter particles under the super-strong gravitational/electromagnetic fields from supermassive compact celestial bodies (SCCB) such as galactic nuclei, quasars etc. Their origin is at the early era of the universe (EEU) as universe density lo large than lo(cr) and lo/c^(2)=const, and physical singularities (PSs) are produced from inflation-explosion stage. PSs can still hide in SCCB with FTL matter now. The energy scale at lo(cr) is the critical energy E(cr). At E(cr), inner symmetries are degenerated; below E(cr), it can hold at most 4-fold SM. Appendix is the theoretical/phenomenal part related to the mass scales sequence (MSS), inflation-explosion, spherical universe, BSM, new particles etc. The MSS, like a Mendeleev's periodic table, can be used to cosmology.

physics.gen-ph

Differentially Private Pre-Trained Model Fusion using Decentralized Federated Graph Matching

Model fusion is becoming a crucial component in the context of model-as-a-service scenarios, enabling the delivery of high-quality model services to local users. However, this approach introduces privacy risks and imposes certain limitations on its applications. Ensuring secure model exchange and knowledge fusion among users becomes a significant challenge in this setting. To tackle this issue, we propose PrivFusion, a novel architecture that preserves privacy while facilitating model fusion under the constraints of local differential privacy. PrivFusion leverages a graph-based structure, enabling the fusion of models from multiple parties without necessitating retraining. By employing randomized mechanisms, PrivFusion ensures privacy guarantees throughout the fusion process. To enhance model privacy, our approach incorporates a hybrid local differentially private mechanism and decentralized federated graph matching, effectively protecting both activation values and weights. Additionally, we introduce a perturbation filter adapter to alleviate the impact of randomized noise, thereby preserving the utility of the fused model. Through extensive experiments conducted on diverse image datasets and real-world healthcare applications, we provide empirical evidence showcasing the effectiveness of PrivFusion in maintaining model performance while preserving privacy. Our contributions offer valuable insights and practical solutions for secure and collaborative data analysis within the domain of privacy-preserving model fusion.

cs.LG

FedBone: Towards Large-Scale Federated Multi-Task Learning

Heterogeneous federated multi-task learning (HFMTL) is a federated learning technique that combines heterogeneous tasks of different clients to achieve more accurate, comprehensive predictions. In real-world applications, visual and natural language tasks typically require large-scale models to extract high-level abstract features. However, large-scale models cannot be directly applied to existing federated multi-task learning methods. Existing HFML methods also disregard the impact of gradient conflicts on multi-task optimization during the federated aggregation process. In this work, we propose an innovative framework called FedBone, which enables the construction of large-scale models with better generalization from the perspective of server-client split learning and gradient projection. We split the entire model into two components: a large-scale general model (referred to as the general model) on the cloud server and multiple task-specific models (referred to as the client model) on edge clients, solving the problem of insufficient computing power on edge clients. The conflicting gradient projection technique is used to enhance the generalization of the large-scale general model between different tasks. The proposed framework is evaluated on two benchmark datasets and a real ophthalmic dataset. Comprehensive results demonstrate that FedBone efficiently adapts to heterogeneous local tasks of each client and outperforms existing federated learning algorithms in most dense prediction and classification tasks with off-the-shelf computational resources on the client side.

cs.LG

Dark Matter, Mass Scales Sequence, and Superstructure in the Universe (with extension and summary)

We extend mass scale sequence to a mass tree. From mass tree, the evolution of the universe is described by three stages: chaos, inflation and expansion. The first two stages have c mutations and the inflation appears as a step by step fission process of black holes. The dark matter particles with low mass (neutrino and delta particle) are described in a dual SM or two-fold SM with new symmetry and new interaction, and delta-particle is like inert neutrino but has baryon number (L-B conservation). We emphasize how to search for delta-particle, how to research critical energy, critical density, background particles, and spherical universe. Critical density relates to a type of pseudo-balance black holes (or celestial bodies). Suppose the minimum black hole radius equal to proton radius means we live in a spherical universe, which belong to a big universe, mainly characterized by proton.

physics.gen-ph

Dark Matter Particles with Low Mass (and FTL)

From the observed results, we deduced that the mass of the neutrino is about 10^(-1) eV and the mass of the fourth stable elementary particle (delta) is about 10^(0) eV. While neutrino is related to electro-weak field, the fourth stable elementary particle delta is related to gravitation-"strong" field, and some new meta-stable baryons may appear near the TeV region. Therefore, a twofold standard model diagram is proposed, and involves some experiment phenomena: The new meta-stable baryons' decays produce delta particles, which are helpful in explaining the Dijet asymmetry phenomena at LHC of CERN, the different results for the Fermilab's data peak, etc; However, according to the (B-L) invariance, the sterile "neutrino" about the event excess in MiniBooNe is not the fourth neutrino but rather the delta particle; We think that the delta particles are related to the phenomenon about neutrinos FTL, and that anti-neutrinos are faster than neutrinos. FTL is also related to cosmic inflation, singular point disappearance, a finite universe, and abnormal red shift of SN Ia. Besides, the dark matter particles with low mass are helpful in explaining missing solar neutrinos, the CMB angular power spectrum measured by WMAP etc. Some experiments and observations are suggested, especially about the measurement for the speed of gravitational wave c'. c' and c, in physics, represent the limit speeds of moving particles made by different categories of matter with different Lorentz factors. Lorentz transformation is compatible with FTL. This will be helpful to look for new particles.

physics.gen-ph

Large Number, Dark Matter, Dark Energy, and the Superstructures in the Universe (with Extension)

Since there are dark matter particles (neutrino) with mass about 10^(-1)eV in the universe, the superstructures with a scale of 10^(19) solar mass [large number A is about 10^(19)] appeared around the era of the hydrogen recombination. The redshift z distributions of quasars support the existence of superstructures. Since there are superstructures in the universe, it is not necessary for the hypothesis of dark energy. While neutrino is related to electro-weak field, the fourth stable elementary particles (delta particle) with mass about 10^(0)eV to 10^(1)eV is related to gravitation-"strong" field, which suggests p + anti(p)--> n/anti(n) + anti(delta particle)/(delta particle) and that some new meta-stable baryons appeared near the TeV region. Therefore, a twofold standard model diagram is proposed, and related to many experiment phenomena: The new meta-stable baryons' decays produce delta particles, which are helpful to explain the Dijet asymmetry phenomena at LHC of CERN, the different results for the Fermilab's data peak, etc; However, according to the (B-L) invariance, the sterile "neutrino" from "the event excess in MiniBooNe" can not be the fourth neutrino but rather the delta particle; We think that the delta particles are related to the phenomenon about neutrinos FTL, and that anti-neutrinos are faster than neutrinos. FTL is also related to the cosmic inflation, singular point disappearance, and abnormal red shift of SN Ia. Some experiments and observations are suggested. In the Extension section, we clarify "mass tree", our finite universe, cosmic dual expansions, dual SM etc. And the LHC can look for new particles with decay products graviton/delta particle and new interaction indeed.

astro-ph