SearcharxivSearch

arXiv subjects

Tianchi Zhang

Publications and source records attributed to Tianchi Zhang.

16 recordsLinked to original sources

Wafer-scale monolithic integration of Ce:YIG films and magneto-optical isolators on silicon

Silicon integrated cerium doped yttrium iron garnet (Ce:YIG) thin films are promising candidates for integrated nonreciprocal photonic devices, cryogenic photonic modulators and optical computing applications. However, previously reported Ce:YIG thin film on silicon is limited to milimeter sizes. Wafer-scale integration and non-destructive characterization of high quality Ce:YIG thin films on silicon has been elusive. Here, we report growth of 4-inch wafer-scale Ce:YIG thin films on silicon substrates by radio-frequency magnetron sputtering. Strong Faraday effect of 2318 deg/cm, low propagation loss of 80 dB/cm and excellent thickness uniformity of 3.5% is demonstrated across the 4-inch silicon wafer. Furthermore, a custom designed wafer-scale, non-destructive magneto-ellipsometry was established to characterize the film thickness, optical constants and magneto-optical constants across the wafer. Wafer-scale integration of ring resonator type magneto-optical isolators are also demonstrated. Our work demonstrates a step forward toward wafer-scale heterogeneous integration and characterization of magneto-optical thin films on silicon, providing material candidates for non-reciprocal photonic device arrays, magneto-optical in-memory computing networks and integrated magneto-optic magnetometers.

physics.optics

Localized crystallization of Ce:YIG thin films on Si using CO2 laser annealing for integrated nonreciprocal photonic device applications

Laser annealing (LA) technique has emerged as an effective method for localized crystallization of magneto-optical (MO) garnet thin films on semiconductor substrates. However, no studies have explored the crystallization and magneto-optical (MO) properties of cerium-substituted yttrium iron garnet (Ce:YIG, Ce1Y2Fe5O12) thin films for integrated photonic device applications using LA technique. In this study, we provide a comprehensive investigation into the laser annealing of Ce:YIG films deposited on SiO2 substrates and silicon nitride photonic waveguides for integrated nonreciprocal photonic device applications. Garnet phase was successfully observed in films grown on SiO2 substrates, and SiN waveguides with laser annealing of sputtered Ce:YIG films on top of a laser annealed Y3Fe5O12 seed layer. The magneto-optical (MO) properties of Ce:YIG films on oxidized Si substrates were found to be comparable to those prepared by rapid thermal annealing (RTA). A Mach-Zehnder Interferometer (MZI) type optical isolator based on Ce:YIG film on SiN was fabricated, exhibiting a saturation Faraday rotation of -2317.7 deg/cm and propagation loss of 188.2 dB/cm. Isolation ratio of 27.1 dB and insertion loss of 10.1 dB were achieved at 1552.7 nm wavelength.

cond-mat.mtrl-sci

A Global-Local Graph Attention Network for Traffic Forecasting

Traffic forecasting is a significant part of intelligent transportation systems. One of the critical challenges of traffic forecasting is to find spatio-temporal correlations. In recent years, graph convolutional networks and graph attention networks have replaced traditional statistical models to predict future traffic. However, it is complicated for both of them to allow vertices to have far different characters. To address this, we propose the Global-Local Graph Attention Network (GLGAT) with pairwise encoding and the event-based adjacency matrix. The GLGAT allows vertices to have a global attention matrix set for the whole graph and assigns local attention matrix sets to each vertex. Experiments on two real-world traffic datasets show that GLGAT can effectively capture spatio-temporal correlations and has competitive performance against other state-of-the-art baselines.

cs.AI

Nonreciprocal optical circuit switching

Directly switching optical signals outperforms conventional optoelectronic hardware in terms of cost, latency, and energy efficiency, and is expected to address the growing demand for data node capacity driven by the development of machine learning and artificial intelligence (AI) technologies. Therefore, optical circuit switching (OCS) technology has piqued widespread research interest in various technical solutions, including silicon photonics. However, silicon-based integrated OCS remains constrained by challenges such as network performance and port scalability. Here we propose a magneto-optical heterogeneous integrated nonreciprocal OCS (NOCS) network based on a silicon photonics platform, achieving bidirectional full-duplex nonreciprocal transmission by programming reciprocal and nonreciprocal phase shifters. We demonstrate that compared with the existing OCS architecture, NOCS has the advantages of ultra-high reconfiguration speed, large-scale integration compatibility, and bidirectional channel isolation reducing the number of required ports. NOCS could meet the programming speed requirements of the AI backend network, or supports nonreciprocal optical switching applications without multiplexing technology.

physics.optics

Unlock giant nonreciprocity via multi-valued behavior of non-Hermitian zero-index materials

Although Einstein's field equations are time-independent, the multivalued feature of the horizon of a blackhole naturally enables the one-way transmission, leading to the strong arrow of time from the time-independent gravitational interaction. Here we experimentally demonstrate a photonic analogue of this principle and reveal the infinite nonreciprocity of the time-reversal-symmetric Maxwell equations. By designing a non-Hermitian zero-index magneto-optical metawaveguide, we introduce multivalued feature to this metawaveguide's complex eigenspace via an exceptional point with non-zero residue, bringing nonlocal, path-dependent historical memory to the system. Hence, a weak magneto-optical response can direct forward and backward waves to two photonic branches with largely distinct momenta and losses, leading to the optical nonreciprocity far beyond the limitation imposed by the magneto-optical material. We fabricated an a-Si/Ce:YIG metawaveguide, achieving nonreciprocal phase shift of 47.78 rad/mm and nonreciprocal loss of 53.9 dB/mm near 1575 nm, exceeding state-of-the-art nonreciprocal devices by an order of magnitude. Our principle universally applies from microwave to visible frequencies, leading to compact isolators, circulators, and sensors. Our principle can also be extended to nonreciprocal acoustic, elastic, and thermal systems. The proposed new paradigm -- geometry-based strong arrow of time in covariant and reversible physical systems -- has broad implications in many disciplines including string theory, cosmology, and astronomy.

physics.optics

Prediction of Individual Halo Concentrations Across Cosmic Time Using Neural Networks

The concentration of dark matter haloes is closely linked to their mass accretion history. We utilize the halo mass accretion histories from large cosmological N-body simulations as inputs for our neural networks, which we train to predict the concentration of individual haloes at a given redshift. The trained model performs effectively in other cosmological simulations, achieving the root mean square error between the actual and predicted concentrations that significantly lower than that of the model by Zhao et al. and Giocoli et al. at any redshift. This model serves as a valuable tool for rapidly predicting halo concentrations at specified redshifts in large cosmological simulations.

astro-ph.CO

Nonreciprocal Optical Routing in Multi-port Magneto-Optical Devices on Silicon

Nonreciprocal optical devices are key components in photonic integrated circuits for light reflection blocking and routing. Most reported silicon integrated nonreciprocal optical devices to date were unit devices. To allow complex signal routing between multi-ports in photonic networks, multi-port magneto-optical (MO) nonreciprocal photonic devices are desired. In this study, we report experimental demonstration of a silicon integrated 5*5 multiport nonreciprocal photonic device based on magneto-optical waveguides. By introducing different nonreciprocal phase shift effect to planar photonic waveguides, the device focuses light to different ports for both forward and backward propagation. The device shows designable nonreciprocal transmission between 5*5 ports, achieving 16 dB isolation ratio and -18 dB crosstalk.

physics.optics

UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications

Applications with irregular data structures, data-dependent control flows and fine-grained data transfers (e.g., real-world graph computations) perform poorly on cache-based systems. We propose the UpDown accelerator that supports fine-grained execution with novel architecture mechanisms - lightweight threading, event-driven scheduling, efficient ultra-short threads, and split-transaction DRAM access with software-controlled synchronization. These hardware primitives support software programmable events, enabling high performance on diverse data structures and algorithms. UpDown also supports scalable performance; hardware replication enables programs to scale up performance. Evaluation results show UpDown's flexibility and scalability enable it to outperform CPUs on graph mining and analytics computations by up to 116-195x geomean speedup and more than 4x speedup over prior accelerators. We show that UpDown generates high memory parallelism (~4.6x over CPU) required for memory intensive graph computations. We present measurements that attribute the performance of UpDown (23x architectural advantage) to its individual architectural mechanisms. Finally, we also analyze the area and power cost of UpDown's mechanisms for software programmability.

cs.AR

Combining Cloud and Mobile Computing for Machine Learning

Although the computing power of mobile devices is increasing, machine learning models are also growing in size. This trend creates problems for mobile devices due to limitations like their memory capacity and battery life. While many services, like ChatGPT and Midjourney, run all the inferences in the cloud, we believe a flexible and fine-grained task distribution is more desirable. In this work, we consider model segmentation as a solution to improving the user experience, dividing the computation between mobile devices and the cloud in a way that offloads the compute-heavy portion of the model while minimizing the data transfer required. We show that the division not only reduces the wait time for users but can also be fine-tuned to optimize the workloads of the cloud. To achieve that, we design a scheduler that collects information about network quality, client device capability, and job requirements, making decisions to achieve consistent performance across a range of devices while reducing the work the cloud needs to perform.

cs.DC

FP8-BERT: Post-Training Quantization for Transformer

Transformer-based models, such as BERT, have been widely applied in a wide range of natural language processing tasks. However, one inevitable side effect is that they require massive memory storage and inference cost when deployed in production. Quantization is one of the popularized ways to alleviate the cost. However, the previous 8-bit quantization strategy based on INT8 data format either suffers from the degradation of accuracy in a Post-Training Quantization (PTQ) fashion or requires an expensive Quantization-Aware Training (QAT) process. Recently, a new numeric format FP8 (i.e. floating-point of 8-bits) has been proposed and supported in commercial AI computing platforms such as H100. In this paper, we empirically validate the effectiveness of FP8 as a way to do Post-Training Quantization without significant loss of accuracy, with a simple calibration and format conversion process. We adopt the FP8 standard proposed by NVIDIA Corp. (2022) in our extensive experiments of BERT variants on GLUE and SQuAD v1.1 datasets, and show that PTQ with FP8 can significantly improve the accuracy upon that with INT8, to the extent of the full-precision model.

cs.AI

The halo concentration and mass relation traced by satellite galaxies

We study the relation between halo concentration and mass (c-M relation) using the Seventh and Eighth Data Release of the Sloan Digital Sky Survey (SDSS DR7 and DR8) galaxy catalogue. Assuming that the satellite galaxies follow the distribution of dark matter, we derive the halo concentration by fitting the satellite radial profile with a Nararro Frank and White (NFW) format. The derived c-M relation covers a wide halo mass range from $10^{11.6}$ to $10^{14.1} \rm\ M_\odot$. We confirm the anti-correlation between the halo mass and concentration as predicted in cosmological simulations. Our results are in good agreement with those derived using galaxy dynamics and gravitational lensing for halos of $10^{11.6}-10^{12.9} \rm\ M_\odot$, while they are slightly lower for halos of $10^{12.9}-10^{14.1}\rm\ M_\odot$. It is because blue satellite galaxies are less concentrated, especially in the inner regions. Instead of using all satellite galaxies, red satellites could be better tracers of the underlying dark matter distribution in galaxy groups.

astro-ph.GA

You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language Model

Large-scale Transformer models bring significant improvements for various downstream vision language tasks with a unified architecture. The performance improvements come with increasing model size, resulting in slow inference speed and increased cost for severing. While some certain predictions benefit from the full complexity of the large-scale model, not all of inputs need the same amount of computation to conduct, potentially leading to computation resource waste. To handle this challenge, early exiting is proposed to adaptively allocate computational power in term of input complexity to improve inference efficiency. The existing early exiting strategies usually adopt output confidence based on intermediate layers as a proxy of input complexity to incur the decision of skipping following layers. However, such strategies cannot apply to encoder in the widely-used unified architecture with both encoder and decoder due to difficulty of output confidence estimation in the encoder. It is suboptimal in term of saving computation power to ignore the early exiting in encoder component. To handle this challenge, we propose a novel early exiting strategy for unified visual language models, which allows dynamically skip the layers in encoder and decoder simultaneously in term of input layer-wise similarities with multiple times of early exiting, namely \textbf{MuE}. By decomposing the image and text modalities in the encoder, MuE is flexible and can skip different layers in term of modalities, advancing the inference efficiency while minimizing performance drop. Experiments on the SNLI-VE and MS COCO datasets show that the proposed approach MuE can reduce expected inference time by up to 50\% and 40\% while maintaining 99\% and 96\% performance respectively.

cs.CV

The spatial distribution of satellites in galaxy clusters

The planar distributions of satellite galaxies around the Milky Way and Andromeda have been extensively studied as potential challenges to the standard cosmological model. Using the Sloan Digital Sky Survey and the Millennium simulation we extend such studies to the satellite galaxies of massive galaxy clusters. We find that both observations and simulations of galaxy clusters show an excess of anisotropic satellite distributions. On average, satellites in clusters have a higher degree of anisotropy than their counterparts in Milky-Way-mass hosts once we account for the difference in their radial distributions. The normal vector of the plane of satellites is strongly aligned with the host halo's minor axis, while the alignment with the large-scale structure is weak. At fixed cluster mass, the degree of anisotropy is higher at higher redshift. This reflects the highly anisotropic nature of satellites accretion points, a feature that is partly erased by the subsequent orbital evolution of the satellites. We also find that satellite galaxies are mostly accreted singly so group accretion is not the explanation for the high flattening of the planes of satellites.

astro-ph.GA

Numerical convergence of pre-initial conditions on dark matter halo properties

Generating pre-initial conditions (or particle loads) is the very first step to set up a cosmological N-body simulation. In this work, we revisit the numerical convergence of pre-initial conditions on dark matter halo properties using a set of simulations which only differs in initial particle loads, i.e. grid, glass, and the newly introduced capacity constrained Voronoi tessellation (CCVT). We find that the median halo properties agree fairly well (i.e. within a convergence level of a few per cent) among simulations running from different initial loads. We also notice that for some individual haloes cross-matched among different simulations, the relative difference of their properties sometimes can be several tens of per cent. By looking at the evolution history of these poorly converged haloes, we find that they are usually merging haloes or haloes have experienced recent merger events, and their merging processes in different simulations are out-of-sync, making the convergence of halo properties become poor temporarily. We show that, comparing to the simulation starting with an anisotropic grid load, the simulation with an isotropic CCVT load converges slightly better to the simulation with a glass load, which is also isotropic. Among simulations with different pre-initial conditions, haloes in higher density environments tend to have their properties converged slightly better. Our results confirm that CCVT loads behave as well as the widely used grid and glass loads at small scales, and for the first time we quantify the convergence of two independent isotropic particle loads (i.e. glass and CCVT) on halo properties.

astro-ph.CO

Galaxy properties in the cosmic web of EAGLE simulation

We investigate the dependence of the galaxy properties on cosmic web environments using the most up-to-date hydrodynamic simulation: Evolution and Assembly of Galaxies and their Environments (EAGLE). The baryon fractions in haloes and the amplitudes of the galaxy luminosity function decrease going from knots to filaments to sheets to voids. Interestingly, the value of L$^*$ varies dramatically in different cosmic web environments. At z = 0, we find a characteristic halo mass of $10^{12} h^{-1}\rm M_{\odot}$, below which the stellar-to-halo mass ratio is higher in knots while above which it reverses. This particular halo mass corresponds to a characteristic stellar mass of $1.8\times 10^{10} h^{-1}\rm M_{\odot}$. Below the characteristic stellar mass central galaxies have redder colors, lower sSFRs and higher metallicities in knots than those in filaments, sheets and voids, while above this characteristic stellar mass, the cosmic web environmental dependences either reverse or vanish. Such dependences can be attributed to the fact that the active galaxy fraction decreases along voids, sheets, filaments and knots. The cosmic web dependences get weaker towards higher redshifts for most of the explored galaxy properties and scaling relations, except for the gas metallicity vs. stellar mass relation.

astro-ph.GA

The optimal gravitational softening length for cosmological N-body simulations

Gravitational softening length is one of the key parameters to properly set up a cosmological $N$-body simulation. In this paper, we perform a large suit of high-resolution $N$-body simulations to revise the optimal softening scheme proposed by Power et al. (P03). Our finding is that P03 optimal scheme works well but is over conservative. Using smaller softening lengths than that of P03 can achieve higher spatial resolution and numerically convergent results on both circular velocity and density profiles. However using an over small softening length overpredicts matter density at the inner most region of dark matter haloes. We empirically explore a better optimal softening scheme based on P03 form and find that a small modification works well. This work will be useful for setting up cosmological simulations.

astro-ph.CO