Searcharxiv⌕ Search

arXiv subjects

Yu Lu

Publications and source records attributed to Yu Lu.

At least 91 records · Page 5Linked to original sources

Hyperbolic Knowledge Transfer in Cross-Domain Recommendation System

Cross-Domain Recommendation (CDR) seeks to utilize knowledge from different domains to alleviate the problem of data sparsity in the target recommendation domain, and it has been gaining more attention in recent years. Although there have been notable advancements in this area, most current methods represent users and items in Euclidean space, which is not ideal for handling long-tail distributed data in recommendation systems. Additionally, adding data from other domains can worsen the long-tail characteristics of the entire dataset, making it harder to train CDR models effectively. Recent studies have shown that hyperbolic methods are particularly suitable for modeling long-tail distributions, which has led us to explore hyperbolic representations for users and items in CDR scenarios. However, due to the distinct characteristics of the different domains, applying hyperbolic representation learning to CDR tasks is quite challenging. In this paper, we introduce a new framework called Hyperbolic Contrastive Learning (HCTS), designed to capture the unique features of each domain while enabling efficient knowledge transfer between domains. We achieve this by embedding users and items from each domain separately and mapping them onto distinct hyperbolic manifolds with adjustable curvatures for prediction. To improve the representations of users and items in the target domain, we develop a hyperbolic contrastive learning module for knowledge transfer. Extensive experiments on real-world datasets demonstrate that hyperbolic manifolds are a promising alternative to Euclidean space for CDR tasks.

cs.IR↗

Retaining Key Information under High Compression Ratios: Query-Guided Compressor for LLMs

The growing popularity of Large Language Models has sparked interest in context compression for Large Language Models (LLMs). However, the performance of previous methods degrades dramatically as compression ratios increase, sometimes even falling to the closed-book level. This decline can be attributed to the loss of key information during the compression process. Our preliminary study supports this hypothesis, emphasizing the significance of retaining key information to maintain model performance under high compression ratios. As a result, we introduce Query-Guided Compressor (QGC), which leverages queries to guide the context compression process, effectively preserving key information within the compressed context. Additionally, we employ a dynamic compression strategy. We validate the effectiveness of our proposed QGC on the Question Answering task, including NaturalQuestions, TriviaQA, and HotpotQA datasets. Experimental results show that QGC can consistently perform well even at high compression ratios, which also offers significant benefits in terms of inference cost and throughput.

cs.CL↗

Automated Multi-level Preference for MLLMs

Current multimodal Large Language Models (MLLMs) suffer from ``hallucination'', occasionally generating responses that are not grounded in the input images. To tackle this challenge, one promising path is to utilize reinforcement learning from human feedback (RLHF), which steers MLLMs towards learning superior responses while avoiding inferior ones. We rethink the common practice of using binary preferences (i.e., superior, inferior), and find that adopting multi-level preferences (e.g., superior, medium, inferior) is better for two benefits: 1) It narrows the gap between adjacent levels, thereby encouraging MLLMs to discern subtle differences. 2) It further integrates cross-level comparisons (beyond adjacent-level comparisons), thus providing a broader range of comparisons with hallucination examples. To verify our viewpoint, we present the Automated Multi-level Preference (AMP) framework for MLLMs. To facilitate this framework, we first develop an automated dataset generation pipeline that provides high-quality multi-level preference datasets without any human annotators. Furthermore, we design the Multi-level Direct Preference Optimization (MDPO) algorithm to robustly conduct complex multi-level preference learning. Additionally, we propose a new hallucination benchmark, MRHal-Bench. Extensive experiments across public hallucination and general benchmarks, as well as our MRHal-Bench, demonstrate the effectiveness of our proposed method. Code is available at https://github.com/takomc/amp.

cs.CV↗

All-voltage control of Giant Magnetoresistance

The aim of voltage control of magnetism is to reduce the power consumption of spintronic devices. For a spin valve, the magnetization directions of two ferromagnetic layers determine the giant magnetoresistance magnitude. However, achieving all-voltage manipulation of the magnetization directions between parallel and antiparallel states is a significant challenge. Here, we demonstrate that by utilizing two exchange-biased Co/IrMn bilayers with opposite pinning directions and with ferromagnetic coupling through the Ruderman-Kittel-Kasuya-Yosida interaction between two Co layers, the magnetization directions of the two ferromagnetic layers of a spin valve can be switched between parallel and antiparallel states through allvoltage-induced strain control. The all-voltage controlled giant magnetoresistance is repeatable and nonvolatile. The rotation of magnetizations in the two Co layers under voltages, from antiparallel to parallel states, occurs in opposite directions as revealed through simulations utilizing the Landau-Lifshitz-Gilbert equation. This work can provide valuable reference for the development of low-power all-voltage-controlled spintronic devices.

cond-mat.mes-hall↗

Performance Analysis of RIS-aided MISO Systems with EMI and Channel Aging

In this paper, we investigate a reconfigurable intelligent surface (RIS)-aided multiple-input single-output (MISO) system in the presence of electromagnetic interference (EMI) and channel aging with a Rician fading channel model between the base station (BS) and user equipment (UE). Specifically, we derive the closed-form expression for downlink spectral efficiency (SE) with maximum ratio transmission (MRT) precoding. The Monte-Carlo simulation supports the theoretical results, demonstrating that amplifying the weight of the line-of-sight (LoS) component in Rician fading channels can boost SE, while EMI has a detrimental impact. Furthermore, continuously increasing the number of RIS elements is not an optimal choice when EMI exists. Nonetheless, RIS can be deployed to compensate for SE degradation caused by channel aging effects. Finally, enlarging the RIS elements size can significantly improve system performance.

cs.IT↗

Rediscovery of Numerical Lüscher's Formula from the Neural Network

We present that by predicting the spectrum in discrete space from the phase shift in continuous space, the neural network can remarkably reproduce the numerical Lüscher's formula to a high precision. The model-independent property of the Lüscher's formula is naturally realized by the generalizability of the neural network. This exhibits the great potential of the neural network to extract model-independent relation between model-dependent quantities, and this data-driven approach could greatly facilitate the discovery of the physical principles underneath the intricate data.

hep-lat↗

Unveiling the Secrets of Engaging Conversations: Factors that Keep Users Hooked on Role-Playing Dialog Agents

With the growing humanlike nature of dialog agents, people are now engaging in extended conversations that can stretch from brief moments to substantial periods of time. Understanding the factors that contribute to sustaining these interactions is crucial, yet existing studies primarily focusing on short-term simulations that rarely explore such prolonged and real conversations. In this paper, we investigate the factors influencing retention rates in real interactions with roleplaying models. By analyzing a large dataset of interactions between real users and thousands of characters, we systematically examine multiple factors and assess their impact on user retention rate. Surprisingly, we find that the degree to which the bot embodies the roles it plays has limited influence on retention rates, while the length of each turn it speaks significantly affects retention rates. This study sheds light on the critical aspects of user engagement with role-playing models and provides valuable insights for future improvements in the development of large language models for role-playing purposes.

cs.CL↗

A single-particle energy-conserving dissipative particle dynamics approach for simulating thermophoresis of nanoparticles in polymer networks

Thermophoresis is an effective method to drive the motion of nanoparticles in fluids. The transport of nanoparticles in polymer networks has significant fundamental and applied importance in biology and medicine, and can be described as Brownian particles crossing entropic barriers. This study proposes a novel extension of dissipative particle dynamics (DPD), called the single-particle energy-conserving dissipative particle dynamics (seDPD), which combines the features of single-particle dissipative particle dynamics (sDPD) and energy-conserving dissipative particle dynamics (eDPD) to simulate the thermophoresis of nanoparticles under temperature gradients. The reliability of the seDPD method is verified by considering the viscosity, thermal diffusivity, and hydrodynamic drag force on the nanoparticles. Using this method, the transport of nanoparticles driven by the thermophoretic force across the polymer network is simulated. The results show that the nanoparticles exhibit the phenomenon of giant acceleration of diffusion (GAD) in the polymer network, indicating that Brownian particles can exhibit GAD when crossing entropic barriers.

cond-mat.soft↗

ActiveDC: Distribution Calibration for Active Finetuning

The pretraining-finetuning paradigm has gained popularity in various computer vision tasks. In this paradigm, the emergence of active finetuning arises due to the abundance of large-scale data and costly annotation requirements. Active finetuning involves selecting a subset of data from an unlabeled pool for annotation, facilitating subsequent finetuning. However, the use of a limited number of training samples can lead to a biased distribution, potentially resulting in model overfitting. In this paper, we propose a new method called ActiveDC for the active finetuning tasks. Firstly, we select samples for annotation by optimizing the distribution similarity between the subset to be selected and the entire unlabeled pool in continuous space. Secondly, we calibrate the distribution of the selected samples by exploiting implicit category information in the unlabeled pool. The feature visualization provides an intuitive sense of the effectiveness of our approach to distribution calibration. We conducted extensive experiments on three image classification datasets with different sampling ratios. The results indicate that ActiveDC consistently outperforms the baseline performance in all image classification tasks. The improvement is particularly significant when the sampling ratio is low, with performance gains of up to 10%. Our code will be released.

cs.CV↗

FlowZero: Zero-Shot Text-to-Video Synthesis with LLM-Driven Dynamic Scene Syntax

Text-to-video (T2V) generation is a rapidly growing research area that aims to translate the scenes, objects, and actions within complex video text into a sequence of coherent visual frames. We present FlowZero, a novel framework that combines Large Language Models (LLMs) with image diffusion models to generate temporally-coherent videos. FlowZero uses LLMs to understand complex spatio-temporal dynamics from text, where LLMs can generate a comprehensive dynamic scene syntax (DSS) containing scene descriptions, object layouts, and background motion patterns. These elements in DSS are then used to guide the image diffusion model for video generation with smooth object motions and frame-to-frame coherence. Moreover, FlowZero incorporates an iterative self-refinement process, enhancing the alignment between the spatio-temporal layouts and the textual prompts for the videos. To enhance global coherence, we propose enriching the initial noise of each frame with motion dynamics to control the background movement and camera motion adaptively. By using spatio-temporal syntaxes to guide the diffusion process, FlowZero achieves improvement in zero-shot video synthesis, generating coherent videos with vivid motion.

cs.CV↗

Comparison of PM-HIP to Forged SA508 Pressure Vessel Steel Under High-Dose Neutron Irradiation

Powder metallurgy with hot isostatic pressing (PM-HIP) is an advanced manufacturing process that is envisioned to replace forging for heavy nuclear components, including the reactor pressure vessel (RPV). But PM-HIP products must at least demonstrate comparable irradiation tolerance than forgings in order to be qualified for nuclear applications. The objective of this study is to directly compare PM-HIP to forged SA508 Grade 3 Class 1 low-alloy RPV steel at two neutron irradiation conditions: ~0.5-1.0 displacements per atom (dpa) at ~270C and ~370C. PM-HIP SA508 experiences greater irradiation hardening and embrittlement (total elongation) than forged SA508. However, uniform elongation and approximate toughness are comparable across all irradiated materials, suggesting irradiated PM-HIP SA508 exhibits superior ductility at maximum load-bearing capacity. The irradiation hardening mechanism is linked to composition rather than fabrication method. Since PM-HIP SA508 has higher Mn and Ni concentration, it is more susceptible to irradiation-induced nucleation of Mn-Ni-Si-P (MNSP) nanoprecipitates and dislocation loops, which both contribute to hardening. Conversely, the forged material nucleates fewer MNSPs, causing dislocation loops to control irradiation hardening. These results show promise for the irradiation performance of PM-HIP SA508 and can motivate future nuclear code qualification of PM-HIP fabrication for RPVs.

cond-mat.mtrl-sci↗

RoFormer: Enhanced Transformer with Rotary Position Embedding

Position encoding recently has shown effective in the transformer architecture. It enables valuable supervision for dependency modeling between elements at different positions of the sequence. In this paper, we first investigate various methods to integrate positional information into the learning process of transformer-based language models. Then, we propose a novel method named Rotary Position Embedding(RoPE) to effectively leverage the positional information. Specifically, the proposed RoPE encodes the absolute position with a rotation matrix and meanwhile incorporates the explicit relative position dependency in self-attention formulation. Notably, RoPE enables valuable properties, including the flexibility of sequence length, decaying inter-token dependency with increasing relative distances, and the capability of equipping the linear self-attention with relative position encoding. Finally, we evaluate the enhanced transformer with rotary position embedding, also called RoFormer, on various long text classification benchmark datasets. Our experiments show that it consistently overcomes its alternatives. Furthermore, we provide a theoretical analysis to explain some experimental results. RoFormer is already integrated into Huggingface: \url{https://huggingface.co/docs/transformers/model_doc/roformer}.

cs.CL↗

Giant Acceleration of Diffusion in Soft Matter Potential

Diffusion of Brownian particles in the tilted periodic potential, usually referred to the washboard potential (WBP), is a well-known model to describe physical systems out of equilibrium. Considering that the biological medium is flexible and thermally fluctuating, a new model, namely the soft matter potential (SMP), is proposed to describe the biological medium. Compared to the washboard potential (WBP), SMP allows Brownian particles to actively modify the structure of the biological medium. Brenner's homogenization theory is applied to predict the diffusivity and velocity of Brownian particles driven by external forces in SMP. Thermodynamic uncertainty relation (TUR) is analyzed for Brownian particles in SMP. It is found that, compared to WBP, Brownian particles in SMP require a lower energy cost $\langle q \rangle$ to achieve accuracy $\mathcal{A}$, i.e. Brownian particles in SMP have higher transport efficiency when driven by external forces.

physics.chem-ph↗

Design of laser uniform illumination system based on aspheric lens and compound ellipsoidal cavity

In order to achieve uniform laser illumination with small aperture diameter and large field Angle,study laser active illumination system.An aspheric mirror combined with a composite ellipsoidal cavity is designed to achieve uniform illumination in this paper.Through an aspheric mirror,the fundamental mode of Gaussian beam is shaped into double Gaussian radiation and Flat-top radiation.The double Gaussian radiation rays are reflected again by the complex ellipsoidal cavity and decomposed into equal radiation flux,which is superimposed with the through Flat-top radiation rays to form a uniform distribution.The parameters of the complex ellipsoidal cavity are obtained by mapping equalization algorithm.After the superposition of the aspherical transmission Flat-top shaping and the composite ellipsoidal cavity secondary reflection shaping,the aperture is 29.7mm,whose aperture angle is 84.0 degrees,and the uniformity is 92.7% with 2m distance and 3.6m diameter.The optimization of uniformity is influenced by three factors:RMS,transmission and reflection power density ratio MT/R and transmission and reflection overlap degree.RMS and MT/R determine the design effect of the composite ellipsoidal cavity, which depends on the maximum reflection Angle and transmission Angle.MT/R is negatively correlated with the maximum reflection of Angle,and RMS is positively correlated with the transmission Angle.When the maximum reflection Angle is set to 32.0 degrees and the transmission Angle to 8.0 degrees,the minimum root-mean-square focusing radius is 108.6um,and the minimum effective transmission reflection power density ratio is 1.07.The degree overlap of transmission and reflection directly affects the uniformity of the target plane.The degree of transmission and reflection is adjusted by setting an adjustment factor.When the adjustment factor is 0.9,the uniformity of the target plane reaches the maximum.

physics.optics↗

Data-Driven Modeling of an Unsaturated Bentonite Buffer Model Test Under High Temperatures Using an Enhanced Axisymmetric Reproducing Kernel Particle Method

In deep geological repositories for high level nuclear waste with close canister spacings, bentonite buffers can experience temperatures higher than 100 °C. In this range of extreme temperatures, phenomenological constitutive laws face limitations in capturing the thermo-hydro-mechanical (THM) behavior of the bentonite, since the pre-defined functional constitutive laws often lack generality and flexibility to capture a wide range of complex coupling phenomena as well as the effects of stress state and path dependency. In this work, a deep neural network (DNN)-based soil-water retention curve (SWRC) of bentonite is introduced and integrated into a Reproducing Kernel Particle Method (RKPM) for conducting THM simulations of the bentonite buffer. The DNN-SWRC model incorporates temperature as an additional input variable, allowing it to learn the relationship between suction and degree of saturation under the general non-isothermal condition, which is difficult to represent using a phenomenological SWRC. For effective modeling of the tank-scale test, new axisymmetric Reproducing Kernel basis functions enriched with singular Dirichlet enforcement representing heater placement and an effective convective heat transfer coefficient representing thin-layer composite tank construction are developed. The proposed method is demonstrated through the modeling of a tank-scale experiment involving a cylindrical layer of MX-80 bentonite exposed to central heating.

cs.CE↗

NEOLAF, an LLM-powered neural-symbolic cognitive architecture

This paper presents the Never Ending Open Learning Adaptive Framework (NEOLAF), an integrated neural-symbolic cognitive architecture that models and constructs intelligent agents. The NEOLAF framework is a superior approach to constructing intelligent agents than both the pure connectionist and pure symbolic approaches due to its explainability, incremental learning, efficiency, collaborative and distributed learning, human-in-the-loop enablement, and self-improvement. The paper further presents a compelling experiment where a NEOLAF agent, built as a problem-solving agent, is fed with complex math problems from the open-source MATH dataset. The results demonstrate NEOLAF's superior learning capability and its potential to revolutionize the field of cognitive architectures and self-improving adaptive instructional systems.

cs.AI↗

Detection of Genuine Multipartite Entanglement in Arbitrary Multipartite systems

We study the genuine multipartite entanglement of arbitrary $n$-partite quantum states by representing the density matrices in terms of the generalized Pauli operators. We introduce a general framework for detecting genuine multipartite entanglement and non full-separability of multipartite quantum states with arbitrary dimensions based on correlation tensors. Effective criterion is derived to verify the genuine multipartite entanglement. Detailed examples are given to show that the criterion detects more genuine multipartite entanglement than the existing criteria.

quant-ph↗

Effect of Neutrinos on Angular Momentum of Dark Matter Halo

Massive neutrinos are expected to affect the large-scale structure formation, including the major component of solid substances, dark matter halos. How halos are influenced by neutrinos is vital and interesting, and angular momentum (AM) as a significant feature provides a statistical perspective for this issue. Exploring halos from TianNu N-body cosmological simulation with the co-evolving neutrino particles, we obtain some concrete conclusions. First, by comparing the same halos with and without neutrinos, in contrast to the neutrino-free case, over 89.71\% of halos have smaller halo moduli, over 71.06\% have smaller particle-mass-reduced (PMR) AM moduli, and over 95.44\% change their orientations of less than $0.65^\circ$. Moreover, the relative variation of PMR modulus is more visible for low-mass halos. Second, to explore the PMR moduli of halos in dense or sparse areas, we divide the whole box into big cubes, and search for halos within a small spherical cell in a single cube. From the two-level divisions, we discover that in denser cubes, the variation of PMR moduli with massive neutrinos decreases more significantly. This distinction suggests that neutrinos exert heavier influence on halos' moduli in compact regions. With massive neutrinos, most halos (86.60\%) have lower masses than without neutrinos.

astro-ph.CO↗