Searcharxiv⌕ Search

arXiv subjects

Yue Yu

Publications and source records attributed to Yue Yu.

At least 217 records · Page 12Linked to original sources

Optical multi-beam steering and communication using integrated acousto-optics arrays

Optical beam steering enables optical detection and imaging in macroscopic or microscopic scales and long-range communication over free space. It underpins numerous optical applications, including LiDAR, biomedical imaging, and remote sensing. Despite the inherent speed of light, advanced applications increasingly require the ability to steer multiple beams simultaneously to increase imaging throughput, boost communication bandwidth, and control arrays qubits for scalable quantum computing. Therefore, there is a significant demand for non-mechanical, integrated, and scalable multi-beam steering technology. Here, we report a scalable multi-beam steering system comprising an array of acousto-optic beam steering channels and photonic integrated circuits on a thin-film lithium niobate platform. Each channel generates tens of individually controllable beams of visible wavelength by exciting acoustic waves using digitally synthesized multi-tone microwave signals. We demonstrate the system's capabilities through multi-input, multi-output free-space communications, simultaneously transmitting to multiple receivers at megabits/sec data rates. This technology can be readily scaled up to steer hundreds of optical beams from a compact chip, potentially advancing many areas of optical technologies and enabling novel applications.

physics.optics↗

Constraint Preconditioning and Parameter Selection for a First-Order Primal-Dual Method applied to Model Predictive Control

Many techniques for real-time trajectory optimization and control require the solution of optimization problems at high frequencies. However, ill-conditioning in the optimization problem can significantly reduce the speed of first-order primal-dual optimization algorithms. We introduce a preconditioning technique and step-size heuristic for Proportional-Integral Projected Gradient (PIPG), a first-order primal-dual algorithm. The preconditioning technique, based on the QR factorization, aims to reduce the condition number of the KKT matrix associated with the optimization problem. Our step-size selection heuristic chooses step-sizes to minimize the upper bound on the convergence of the primal-dual gap for the optimization problem. These algorithms are tested on two model predictive control problem examples and show a solve-time reduction of at least 3.6x.

math.OC↗

Photorefractive and pyroelectric photonic memory and long-term stability in thin-film lithium niobate microresonators

The stability of the integrated photonic circuits is of critical importance for many applications that require high frequency precision or robust operation over time, such as optomechanical sensing, frequency conversion, optical communication, and quantum optics. Photonic memory is useful for low-energy optical computing and interconnects. Thin film lithium niobate (TFLN), as an emerging photonic platform, exhibits complex material properties including pyroelectric (PE) and photorefractive (PR) effects which could lead to intra-device drift and excess noise under different environmental or operating conditions as well as be utilized for building photonic memory. However, the long-term stability and memory effect of its optical properties has not been explored. In this paper, we discovered a long-lived change of optical refractive index as a result of light excitation and temporal temperature variation using Z-cut TFLN microresonators and reveal a strong dependence of the instability with the crystal orientation of the thin film form. The recovery time are measured to be over 10 hours. Leveraging the photonic memory with a long relaxation time, we realize optical trimming of the cavity resonance frequencies. Our result offers insights towards understanding the fundamental noise properties and dynamic behavior of the integrated TFLN material and devices.

physics.optics↗

Sensing Resource Allocation Against Data-Poisoning Attacks in Traffic Routing

Data-poisoning attacks can disrupt the efficient operations of transportation systems by misdirecting traffic flows via falsified data. One challenge in countering these attacks is to reduce the uncertainties on the types of attacks, such as the distribution of their targets and intensities. We introduce a resource allocation method in transportation networks to detect and distinguish different types of attacks and facilitate efficient traffic routing. The idea is to first cluster different types of attacks based on the corresponding optimal routing strategies, then allocate sensing resources to a subset of network links to distinguish attacks from different clusters via lexicographical mixed-integer programming. We illustrate the application of the proposed method using the Anaheim network, a benchmark model in traffic routing that contains more than 400 nodes and 900 links.

math.OC↗

Efficient Multi-task Prompt Tuning for Recommendation

With the expansion of business scenarios, real recommender systems are facing challenges in dealing with the constantly emerging new tasks in multi-task learning frameworks. In this paper, we attempt to improve the generalization ability of multi-task recommendations when dealing with new tasks. We find that joint training will enhance the performance of the new task but always negatively impact existing tasks in most multi-task learning methods. Besides, such a re-training mechanism with new tasks increases the training costs, limiting the generalization ability of multi-task recommendation models. Based on this consideration, we aim to design a suitable sharing mechanism among different tasks while maintaining joint optimization efficiency in new task learning. A novel two-stage prompt-tuning MTL framework (MPT-Rec) is proposed to address task irrelevance and training efficiency problems in multi-task recommender systems. Specifically, we disentangle the task-specific and task-sharing information in the multi-task pre-training stage, then use task-aware prompts to transfer knowledge from other tasks to the new task effectively. By freezing parameters in the pre-training tasks, MPT-Rec solves the negative impacts that may be brought by the new task and greatly reduces the training costs. Extensive experiments on three real-world datasets show the effectiveness of our proposed multi-task learning framework. MPT-Rec achieves the best performance compared to the SOTA multi-task learning method. Besides, it maintains comparable model performance but vastly improves the training efficiency (i.e., with up to 10% parameters in the full training way) in the new task learning.

cs.IR↗

Performance Metrics and Loss Mechanisms in Horticulture Luminescent Solar Concentrators

Horticulture Luminescent Solar Concentrators (HLSCs) represent an innovative concept developed in recent years to promote crop yields, building upon the foundation of traditional Luminescent Solar Concentrators (LSCs). HLSCs are characterized by two distinct properties: spectral conversion and light extraction. Unlike traditional LSCs, HLSCs focus on converting energy from one part of the solar spectrum (typically green) to a specific range (usually red) and aim for the converted photons to exit the device from the bottom surface rather than the edge surfaces. In this study, we start by examining the specific requirements of horticulture to clarify the motivation for using HLSCs. We re-evaluate and propose new optical metrics tailored to HLSCs. Additionally, we analyse potential loss channels for direct red emission and converted red emission. Utilizing Monte Carlo ray tracing method and experimental data, we further explore the factors influencing these loss channels. Our work provides a fundamental discussion on HLSCs and offers design guidelines for future HLSC research.

physics.optics↗

Flavor Nernst effects in quantum paramagnets

Recent advances in spin transport research have highlighted the potential of quantum paramagnets as platforms for exploring novel phenomena and developing next-generation technologies. In this paper, we investigate the flavor Nernst effect (FNE) in quantum paramagnets, focusing on the Hall-type thermal spin transport of crystal electric field (CEF) excitations with spin-orbit couplings. As a proof of principle, we investigate the quantum paramagnetic ground state in an effective spin-1 Hamiltonian with Dzyaloshinskii-Moriya interactions and a large hard-axis anisotropy. We employ linear flavor-wave theory to analyze the low-energy excitations, and obtain the flavor Nernst coefficients from the linear response theory. We demonstrate the FNE in a 2D pyrochlore thin film with an all-in-all-out Ising axis configuration, and investigate their dependence on temperature, anisotropy, DM interaction, and external fields. Our results reveal the connection between the FNE and the Berry curvature of the CEF excitations, suggesting potential applications in manipulating thermal spin currents and exploring topological spin transport phenomena in quantum paramagnets.

cond-mat.str-el↗

Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics Discovery

Despite the recent popularity of attention-based neural architectures in core AI fields like natural language processing (NLP) and computer vision (CV), their potential in modeling complex physical systems remains under-explored. Learning problems in physical systems are often characterized as discovering operators that map between function spaces based on a few instances of function pairs. This task frequently presents a severely ill-posed PDE inverse problem. In this work, we propose a novel neural operator architecture based on the attention mechanism, which we coin Nonlocal Attention Operator (NAO), and explore its capability towards developing a foundation physical model. In particular, we show that the attention mechanism is equivalent to a double integral operator that enables nonlocal interactions among spatial tokens, with a data-dependent kernel characterizing the inverse mapping from data to the hidden parameter field of the underlying operator. As such, the attention mechanism extracts global prior information from training data generated by multiple systems, and suggests the exploratory space in the form of a nonlinear kernel map. Consequently, NAO can address ill-posedness and rank deficiency in inverse PDE problems by encoding regularization and achieving generalizability. We empirically demonstrate the advantages of NAO over baseline neural models in terms of generalizability to unseen data resolutions and system states. Our work not only suggests a novel neural operator architecture for learning interpretable foundation models of physical systems, but also offers a new perspective towards understanding the attention mechanism.

cs.LG↗

Order by projection in single-band Hubbard model: a DMRG study

In a Fermi system near or at half-filling, a specific superconducting pairing channel, if not explicitly included in the Hamiltonian, can be boosted by suppressing a competing pairing channel; this is exemplified by the enhancement of extended $s$-wave correlations upon suppressing $s$-wave Cooper pairing. This phenomenon, originally found by the use of generalized uncertainty relations is referred to as \emph{order by projection}. The case of zero on-site Coulomb interaction in the thermodynamic limit, confirms this mechanism through the analytical solution. In this study, we go further and systematically investigate this mechanism for a strongly correlated fermionic Hubbard model, now with finite on-site interaction, on a square lattice with an extended set of hopping parameters. We explore the behaviors of different pairing channels when one of them is suppressed, utilizing density matrix renormalization group calculations. Our findings provide numerical evidence supporting the existence of \emph{order by projection} in the strongly correlated system we studied. We also investigate the effect of the strength of Hubbard $U$, next-nearest neighbor $t'$, hole-doping, as well as finite-size scaling approaching the thermodynamic limit.

cond-mat.str-el↗

Investigation on a quantum algorithm for linear differential equations

Ref.[BCOW17] introduced a pioneering quantum approach (coined BCOW algorithm) for solving linear differential equations with optimal error tolerance. Originally designed for a specific class of diagonalizable linear differential equations, the algorithm was extended by Krovi in [Kro23] to encompass broader classes, including non-diagonalizable and even singular matrices. Despite the common misconception, the original algorithm is indeed applicable to non-diagonalizable matrices, with diagonalisation primarily serving for theoretical analyses to establish bounds on condition number and solution error. By leveraging basic estimates from [Kro23], we derive bounds comparable to those outlined in the Krovi algorithm, thereby reinstating the advantages of the BCOW approach. Furthermore, we extend the BCOW algorithm to address time-dependent linear differential equations by transforming non-autonomous systems into higher-dimensional autonomous ones, a technique also applicable for the Krovi algorithm.

quant-ph↗

Dissipationless topological quantum computation for Majorana objects in sparse-dense mixed encoding process

Topological quantum computation based on Majorana objects is subject to a significant challenge because at least some of the two-qubit quantum gates rely on the fermion (either charge or spin) parity of the qubits. This dependency renders the quantum operations involving these gates probabilistic when attempting to advance quantum processes within the quantum circuit model. Such an approach leads to significant information loss whenever measurements yield the undesired fermion parity. To resolve the problem of wasting information, we devise topological operations that allow for the non-dissipative correction of information from undesired fermion parity to the desired one. We will use the sparse-dense mixed encoding process for the controlled-NOT gate as an example to explain how corrections can be implemented without affecting the quantum information carried by the computational qubits. This correction process can be applied {to} either the undesired input qubits or the fermion parity-dependent quantum gates, and it works for both Majorana-zero-mode-based and Majorana-edge-mode-based topological quantum computation.

quant-ph↗

Non-Fermi Liquid Behavior of the $t$-$J$ Model in the Strange Metal Phase: $U(1)$ Gauge Theory Consistent with Local Constraints

In the slave particle representation with $U(1)$ gauge symmetry, local constraints on physical states characterized by various mean field solutions belong to Dirac's second-class ones. Although constrained systems are extensively investigated, realistic methods to solve the gauge theory problem with second-class constraints are yet to be developed. We formulate a Becchi-Rouet-Stora-Tyutin (BRST) quantization theory, called consistent $U(1)$ gauge theory, that is consistent with both first- and second-class local constraints for strongly correlated condensed matter systems. In our consistent $U(1)$ gauge theory, the redundant gauge degrees of freedom are removed by proper gauge fixing conditions while the constraints are exactly retained and the gauge invariance is guaranteed by the BRST symmetry. Furthermore, the gauge fixing conditions endow the gauge field with dynamics. This turns the strongly correlated electron model into a weakly coupled slave boson model, so most of the system's physical properties can be calculated by the conventional quantum many-body perturbation method. We focus on the property of the strange metal phase in the $t$-$J$ model. The electron momentum distribution and the spectral function are calculated, and the non-Fermi liquid behavior agrees with the angle-resolved photoemission spectroscopy measurements for cuprate materials. We also study the electromagnetic responses of the strange metal state. The observed non-Fermi liquid anomalies are captured by our calculations. Especially, we find that the Hall resistivity decreases as temperature increases, and the sign of the Hall resistivity varies from negative to positive when the dopant concentration varies from optimal doping to underdoping in the strange metal regime.

cond-mat.str-el↗

RAM-EHR: Retrieval Augmentation Meets Clinical Predictions on Electronic Health Records

We present RAM-EHR, a Retrieval AugMentation pipeline to improve clinical predictions on Electronic Health Records (EHRs). RAM-EHR first collects multiple knowledge sources, converts them into text format, and uses dense retrieval to obtain information related to medical concepts. This strategy addresses the difficulties associated with complex names for the concepts. RAM-EHR then augments the local EHR predictive model co-trained with consistency regularization to capture complementary information from patient visits and summarized knowledge. Experiments on two EHR datasets show the efficacy of RAM-EHR over previous knowledge-enhanced baselines (3.4% gain in AUROC and 7.2% gain in AUPR), emphasizing the effectiveness of the summarized knowledge from RAM-EHR for clinical prediction tasks. The code will be published at \url{https://github.com/ritaranx/RAM-EHR}.

cs.CL↗

Heterogeneous Peridynamic Neural Operators: Discover Biotissue Constitutive Law and Microstructure From Digital Image Correlation Measurements

Human tissues are highly organized structures with collagen fiber arrangements varying from point to point. Anisotropy of the tissue arises from the natural orientation of the fibers, resulting in location-dependent anisotropy. Heterogeneity also plays an important role in tissue function. It is therefore critical to discover and understand the distribution of fiber orientations from experimental mechanical measurements such as digital image correlation (DIC) data. To this end, we introduce the Heterogeneous Peridynamic Neural Operator (HeteroPNO) approach for data-driven constitutive modeling of heterogeneous anisotropic materials. Our goal is to learn a nonlocal constitutive law together with the material microstructure, in the form of a heterogeneous fiber orientation field, from load-displacement field measurements. We propose a two-phase learning approach. Firstly, we learn a homogeneous constitutive law in the form of a neural network-based kernel function and a nonlocal bond force, to capture complex homogeneous material responses from data. Then, in the second phase we reinitialize the learnt bond force and the kernel function, and training them together with a fiber orientation field for each material point. Owing to the state-based peridynamic skeleton, our HeteroPNO-learned material models are objective and have the balance of linear and angular momentum guaranteed. Moreover, the effects from heterogeneity and nonlinear constitutive relationship are captured by the kernel function and the bond force respectively, enabling physical interpretability. As a result, our HeteroPNO architecture can learn a constitutive model for a biological tissue with anisotropic heterogeneous response undergoing large deformation regime. Moreover, the framework is capable to provide displacement and stress field predictions for new and unseen loading instances.

cond-mat.mtrl-sci↗

Model Provenance via Model DNA

Understanding the life cycle of the machine learning (ML) model is an intriguing area of research (e.g., understanding where the model comes from, how it is trained, and how it is used). This paper focuses on a novel problem within this field, namely Model Provenance (MP), which concerns the relationship between a target model and its pre-training model and aims to determine whether a source model serves as the provenance for a target model. This is an important problem that has significant implications for ensuring the security and intellectual property of machine learning models but has not received much attention in the literature. To fill in this gap, we introduce a novel concept of Model DNA which represents the unique characteristics of a machine learning model. We utilize a data-driven and model-driven representation learning method to encode the model's training data and input-output information as a compact and comprehensive representation (i.e., DNA) of the model. Using this model DNA, we develop an efficient framework for model provenance identification, which enables us to identify whether a source model is a pre-training model of a target model. We conduct evaluations on both computer vision and natural language processing tasks using various models, datasets, and scenarios to demonstrate the effectiveness of our approach in accurately identifying model provenance.

cs.LG↗

Nematic Bogoliubov Fermi surfaces from magnetic toroidal order in FeSe$_{1-x}$S$_x$

Recently it has been argued that the superconducting state of FeSe$_{1-x}$S$_x$ exhibits Bogoliubov Fermi surfaces for $x>0.17$. These Bogoliubov Fermi surfaces appear together with broken time-reversal symmetry and surprisingly demonstrate nematic behavior in a structurally tetragonal phase. Through a symmetry-based analysis of Bogoliubov Fermi surfaces that can arise from broken time-reversal symmetry, we argue that the likely origin of time-reversal symmetry breaking is due to magnetic toroidal order. We show that this magnetic toroidal order naturally appears as a consequence of either static Néel antiferromagnetic order or due to the formation of a spontaneous pair density wave superconducting order. Finally, we reveal that independent of the presence of Bogoliubov Fermi surfaces, supercurrents will induce Néel magnetic order in many Fe-based superconductors.

cond-mat.supr-con↗

Large language models, physics-based modeling, experimental measurements: the trinity of data-scarce learning of polymer properties

Large language models (LLMs) bear promise as a fast and accurate material modeling paradigm for evaluation, analysis, and design. Their vast number of trainable parameters necessitates a wealth of data to achieve accuracy and mitigate overfitting. However, experimental measurements are often limited and costly to obtain in sufficient quantities for finetuning. To this end, we present a physics-based training pipeline that tackles the pathology of data scarcity. The core enabler is a physics-based modeling framework that generates a multitude of synthetic data to align the LLM to a physically consistent initial state before finetuning. Our framework features a two-phase training strategy: (1) utilizing the large-in-amount while less accurate synthetic data for supervised pretraining, and (2) finetuning the phase-1 model with limited experimental data. We empirically demonstrate that supervised pretraining is vital to obtaining accurate finetuned LLMs, via the lens of learning polymer flammability metrics where cone calorimeter data is sparse.

cs.LG↗

RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs

Large language models (LLMs) typically utilize the top-k contexts from a retriever in retrieval-augmented generation (RAG). In this work, we propose a novel instruction fine-tuning framework RankRAG, which instruction-tunes a single LLM for the dual purpose of context ranking and answer generation in RAG. In particular, the instruction-tuned LLMs work surprisingly well by adding a small fraction of ranking data into the training blend, and outperform existing expert ranking models, including the same LLM exclusively fine-tuned on a large amount of ranking data. For generation, we compare our model with many strong baselines, including GPT-4-0613, GPT-4-turbo-2024-0409, and ChatQA-1.5, an open-sourced model with the state-of-the-art performance on RAG benchmarks. Specifically, our Llama3-RankRAG significantly outperforms Llama3-ChatQA-1.5 and GPT-4 models on nine knowledge-intensive benchmarks. In addition, it also performs comparably to GPT-4 on five RAG benchmarks in the biomedical domain without instruction fine-tuning on biomedical data, demonstrating its superb capability for generalization to new domains.

cs.CL↗