SearcharxivSearch

arXiv subjects

Y. Wu

Publications and source records attributed to Y. Wu.

At least 19 recordsLinked to original sources

Low-Variance Randomised Numerical Linear Algebra for Finite Element Simulation

We present a low-variance randomised numerical linear algebra approach for multi-query finite element systems arising from parametric elliptic partial differential equations with applications to digital twins and online model calibration. The method relies on Galerkin subspace projection for reducing the dimensionality, and then combines parameter-oblivious leverage-score Bernoulli sampling with a control variates scheme to yield a reduced-variance `forward' sketch and an invertible `inverse' sketch that are then fused to a single efficient regularised estimator. Effectively, this reduces the computational cost in computing the projected system of equations while preserving the structure, stability, and accuracy of the underlying FEM formulation. We derive probabilistic bounds for the sketching error, invertibility, and estimator variance, and then validate the method on large-scale example problems. The results show that when the parameter fields do not vary too sharply, the synergy of control variates together with the sketch fusion can largely offset the loss incurred by the sub-optimal parameter-oblivious sampling. In this regime, our method achieves substantial savings in time, memory, and communication while maintaining accuracy levels that are acceptable for scientific simulation.

math.NA

The Murchison Widefield Array Phase III upgrade: Sensitivity Doubled, Number of Baselines Quadrupled, Flexibility Enhanced, and EoR Observations Optimised

We describe the latest iteration of upgrades (designated Phase III) to the Murchison Widefield Array (MWA), in the fourth paper in a series that covers the evolution of the telescope from design concept to initial operational facility, and through two major upgrades. As part of the Phase III upgrade of the MWA, we report the completion of work to design, build, and deploy a new fleet of digital receivers that further optimise the MWA for Epoch of Reionisation observations. These receivers complement existing receivers, such that the MWA now supports the full correlation of all 256 antenna tiles currently in the array. This step releases the MWA from the prior constraint of having to correlate only 128 of the 256 tiles at any given time, which means that the maximum instantaneous sensitivity of the MWA is doubled and the maximum number of interferometric baselines is approximately quadrupled. The upgrade is fundamentally enabled by the new MWAX correlator and various other improvements to the MWA sub-systems. In this paper we describe the new digital receivers and the other improvements that result in the Phase III system. A range of operational benefits arise from the upgrade and scientific flexibility is increased. We also comment on the transition from the MWA to the SKA-Low facility near the end of the decade, including a description of some unique science opportunities utilising joint MWA/SKA-Low data during the Science Verification phase of the SKA-Low Array Assembly 2 (AA2) period.

astro-ph.IM

Inverse design of exceptional points in a single-resonance two-port network

Exceptional points (EPs) in non-Hermitian photonic systems enable unconventional control of wave amplitude and phase. However, identifying the EPs in a multidimensional parameter space of a system can be nontrivial and, in some cases, even infeasible. Here we propose an inverse-design method to efficiently locate the scattering EPs for a two-port resonant system supporting a single mode. The proposed method provides a direct way for tuning of geometric parameters to realize scattering EPs, as confirmed by both full-wave simulation and equivalent circuit model. In principle, our method is compatible with multi-mode system and applicable to a broad class of resonant systems.

physics.optics

Abrupt crystallization from shock-compressed CaSiO3 glass

We have performed in situ time-resolved X-ray diffraction at ~100 GPa on laser-shocked CaSiO3 glass to investigate the glass-to-crystal transition. At this extreme pressure, we observe the ultrafast crystallization of the CaSiO3 perovskite structure from the compressed amorphous phase, with a typical nucleation time of 1.69 +/- 0.10 ns and a final grainsize of ~20 nm. The grain size temporal evolution suggest a diffusion controlled transformation. Moreover, the observed concomitant explosive grain growth together with the release wave arrival into shocked CaSiO3 also suggests a role of the release in the nucleation process.

cond-mat.mtrl-sci

Spin-flop-like transition as quantum critical point in Cs$_2$RuO$_4$

We report thermodynamic, neutron diffraction, and inelastic neutron scattering measurements on Cs$_2$RuO$_4$, a member of the celebrated family of frustrated magnets Cs$_2$MX$_4$ (M = Cu, Co, X = Br, Cl). Unlike the previously studied members, it is based on $4d$ transition metal ions with $S=1$. Mapping out the $H-T$ magnetic phase diagram reveals an unusual continuous spin-flop-like phase transition associated with a quantum critical point within the antiferromagnetically ordered phase. A quantitative analysis of the complex magnetic excitation spectrum measured in zero field allows us to derive a model magnetic Hamiltonian for this compound. Its main feature is a frustration of magnetic anisotropy on a level that is much higher than in any of the previously studied species. This frustration naturally explains the peculiar phase transition observed.

cond-mat.str-el

The calibration house in JUNO

As an auxiliary system within the calibration system of the Jiangmen Underground Neutrino Observatory, a calibration house is designed to provide interfaces for connecting the central detector and accommodating various calibration sub-systems. Onsite installation has demonstrated that the calibration house interfaces are capable of effectively connecting to the central detector and supporting the installation of complex and sophisticated calibration sub-systems. Additionally, controlling the levels of radon and oxygen within the calibration house is critical. Radon can increase the experimental background, while oxygen can degrade the quality of the liquid scintillator. The oxygen concentration can be maintained at levels below 10 parts per million, and the radon concentration can be kept below 15 mBq/m$^{3}$. This paper will provide detailed information on the calibration house and its methods for radon and oxygen concentration control.

physics.ins-det

Study of atomic effects on electron spectrum in bound-muon decay process

For the bound-muon decay process, the study of atomic effects on the electron spectrum near its endpoint is performed within the framework of the Fermi effective theory. The analysis takes into account for corrections due to finite-nuclear-size, nuclear-deformation, electron-screening, and vacuum-polarization effects, all of which are incorporated self-consistently into the Dirac equation. Furthermore, the nuclear-recoil correction to the muon binding energy is included. Calculations are carried out for the isotopes of C, Al, and Si, which are of a particular importance for forthcoming experiments aimed at search for the charged-lepton flavor-violating process of muon-to-electron conversion in a nuclear field.

physics.atom-ph

Pressure tuning of Kitaev spin liquid candidate Na$_3$Co$_2$SbO$_6$

The search for Kitaev's quantum spin liquid (KQSL) state in real materials has recently expanded with the prediction that honeycomb lattices of divalent, high-spin cobalt ions could host the dominant bond-dependent exchange interactions required to stabilize the elusive entangled quantum state. The layered honeycomb Na$_3$Co$_2$SbO$_6$ has been singled out as a leading candidate provided that the trigonal crystal field acting on Co $3d$ orbitals, which enhances non-Kitaev exchange interactions between $J_{\rm eff}=\frac{1}{2}$ spin-orbital pseudospins, is reduced. We find that applied pressure leads to anisotropic compression of the layered structure, significantly reducing the trigonal distortion of CoO$_6$ octahedra. A strong enhancement of ferromagnetic correlations between pseudospins is observed in the spin-polarized (3 Tesla) phase up to about 60 GPa. Higher pressures drive a spin transition into a low-spin state destroying the $J_{\rm eff}=\frac{1}{2}$ local moments required to map the spin Hamiltonian into Kitaev's model. The spin transition strongly suppresses the low-temperature magnetic susceptibility and appears to stabilize a paramagnetic phase driven by frustration. Although applied pressure fails to realize a KQSL state, the possible emergence of frustrated magnetism of localized, low-spin $S=\frac{1}{2}$ moments opens the door for exploration of novel magnetic quantum states in compressed honeycomb lattices of divalent cobaltates.

cond-mat.str-el

Superconductivity in the parent infinite-layer nickelate NdNiO$_2$

We report evidence for superconductivity with onset temperatures up to 11 K in thin films of the infinite-layer nickelate parent compound NdNiO$_2$. A combination of oxide molecular-beam epitaxy and atomic hydrogen reduction yields samples with high crystallinity and low residual resistivities, a substantial fraction of which exhibit superconducting transitions. We survey a large series of samples with a variety of techniques, including electrical transport, scanning transmission electron microscopy, x-ray absorption spectroscopy, and resonant inelastic x-ray scattering, to investigate the possible origins of superconductivity. We propose that superconductivity could be intrinsic to the undoped infinite-layer nickelates but suppressed by disorder due to its nodal order parameter, a finding which would necessitate a reconsideration of the nickelate phase diagram. Another possible hypothesis is that the parent materials can be hole doped from randomly dispersed apical oxygen atoms, which would suggest an alternative pathway for achieving superconductivity.

cond-mat.supr-con

Anisotropic Spin Stripe Domains in Bilayer La$_3$Ni$_2$O$_7$

The discovery of superconductivity in La$_3$Ni$_2$O$_7$ under pressure has motivated the investigation of a parent spin density wave (SDW) state, which could provide the underlying pairing interaction. Here, we employ resonant soft x-ray scattering and polarimetry on thin films of bilayer La$_3$Ni$_2$O$_7$ to determine that the magnetic structure of the SDW forms unidirectional diagonal spin stripes with moments lying within the NiO$_2$ plane and perpendicular to $\mathbf{Q}_{SDW}$, but without evidence of the strong charge disproportionation typically associated with other nickelates. These stripes form anisotropic domains with shorter correlation lengths perpendicular versus parallel to $\mathbf{Q}_{SDW}$, revealing nanoscale rotational and translational symmetry breaking analogous to the cuprate and Fe-based superconductors, with possible Bloch-like antiferromagnetic domain walls separating orthogonal domains.

cond-mat.supr-con

Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models

Parameter-efficient fine-tuning (PEFT) is crucial for customizing Large Language Models (LLMs) with constrained resources. Although there have been various PEFT methods for dense-architecture LLMs, PEFT for sparse-architecture LLMs is still underexplored. In this work, we study the PEFT method for LLMs with the Mixture-of-Experts (MoE) architecture and the contents of this work are mainly threefold: (1) We investigate the dispersion degree of the activated experts in customized tasks, and found that the routing distribution for a specific task tends to be highly concentrated, while the distribution of activated experts varies significantly across different tasks. (2) We propose Expert-Specialized Fine-Tuning, or ESFT, which tunes the experts most relevant to downstream tasks while freezing the other experts and modules; experimental results demonstrate that our method not only improves the tuning efficiency, but also matches or even surpasses the performance of full-parameter fine-tuning. (3) We further analyze the impact of the MoE architecture on expert-specialized fine-tuning. We find that MoE models with finer-grained experts are more advantageous in selecting the combination of experts that are most relevant to downstream tasks, thereby enhancing both the training efficiency and effectiveness. Our code is available at https://github.com/deepseek-ai/ESFT.

cs.CL

Petty projection inequality on the sphere and on the hyperbolic space

Using gnomonic projection and Poincar\'e model, we first define the spherical projection body and hyperbolic projection body in spherical space $\mathbb{S}^n$ and hyperbolic space $\mathbb{H}^n$, then define the spherical Steiner symmetrization and hyperbolic Steiner symmetrization, finally prove the spherical projection inequality and hyperbolic projection inequality.

math.MG

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

We present DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks. Specifically, DeepSeek-Coder-V2 is further pre-trained from an intermediate checkpoint of DeepSeek-V2 with additional 6 trillion tokens. Through this continued pre-training, DeepSeek-Coder-V2 substantially enhances the coding and mathematical reasoning capabilities of DeepSeek-V2, while maintaining comparable performance in general language tasks. Compared to DeepSeek-Coder-33B, DeepSeek-Coder-V2 demonstrates significant advancements in various aspects of code-related tasks, as well as reasoning and general capabilities. Additionally, DeepSeek-Coder-V2 expands its support for programming languages from 86 to 338, while extending the context length from 16K to 128K. In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks.

cs.SE

Magnetic properties of the quasi-XY Shastry-Sutherland magnet Er$_2$Be$_2$SiO$_7$

Polycrystalline and single crystal samples of the insulating Shastry-Sutherland compound Er$_2$Be$_2$SiO$_7$ were synthesized via a solid-state reaction and the floating zone method respectively. The crystal structure, Er single ion anisotropy, zero-field magnetic ground state, and magnetic phase diagrams along high-symmetry crystallographic directions were investigated by bulk measurement techniques, x-ray and neutron diffraction, and neutron spectroscopy. We establish that Er$_2$Be$_2$SiO$_7$ crystallizes in a tetragonal space group with planes of orthogonal Er dimers and a strong preference for the Er moments to lie in the local plane perpendicular to each dimer bond. We also find that this system has a non-collinear ordered ground state in zero field with a transition temperature of 0.841 K consisting of antiferromagnetic dimers and in-plane moments. Finally, we mapped out the $H-T$ phase diagrams for Er$_2$Be$_2$SiO$_7$ along the directions $H \parallel$ [001], [100], and [110]. While an increasing in-plane field simply induces a phase transition to a field-polarized phase, we identify three metamagnetic transitions before the field-polarized phase is established in the $H \parallel$ [001] case. This complex behavior establishes insulating Er$_2$Be$_2$SiO$_7$ and other isostructural family members as promising candidates for uncovering exotic magnetic properties and phenomena that can be readily compared to theoretical predictions of the exactly soluble Shastry-Sutherland model.

cond-mat.str-el

A study of Galactic Plane Planck Galactic Cold Clumps observed by SCOPE and the JCMT Plane Survey

We have investigated the physical properties of Planck Galactic Cold Clumps (PGCCs) located in the Galactic Plane, using the JCMT Plane Survey (JPS) and the SCUBA-2 Continuum Observations of Pre-protostellar Evolution (SCOPE) survey. By utilising a suite of molecular-line surveys, velocities and distances were assigned to the compact sources within the PGCCs, placing them in a Galactic context. The properties of these compact sources show no large-scale variations with Galactic environment. Investigating the star-forming content of the sample, we find that the luminosity-to-mass ratio (L/M) is an order of magnitude lower than in other Galactic studies, indicating that these objects are hosting lower levels of star formation. Finally, by comparing ATLASGAL sources that are associated or are not associated with PGCCs, we find that those associated with PGCCs are typically colder, denser, and have a lower L/M ratio, hinting that PGCCs are a distinct population of Galactic Plane sources.

astro-ph.GA

Hyperbolic Secant representation of the logistic function: Application to probabilistic Multiple Instance Learning for CT intracranial hemorrhage detection

Multiple Instance Learning (MIL) is a weakly supervised paradigm that has been successfully applied to many different scientific areas and is particularly well suited to medical imaging. Probabilistic MIL methods, and more specifically Gaussian Processes (GPs), have achieved excellent results due to their high expressiveness and uncertainty quantification capabilities. One of the most successful GP-based MIL methods, VGPMIL, resorts to a variational bound to handle the intractability of the logistic function. Here, we formulate VGPMIL using P\'olya-Gamma random variables. This approach yields the same variational posterior approximations as the original VGPMIL, which is a consequence of the two representations that the Hyperbolic Secant distribution admits. This leads us to propose a general GP-based MIL method that takes different forms by simply leveraging distributions other than the Hyperbolic Secant one. Using the Gamma distribution we arrive at a new approach that obtains competitive or superior predictive performance and efficiency. This is validated in a comprehensive experimental study including one synthetic MIL dataset, two well-known MIL benchmarks, and a real-world medical problem. We expect that this work provides useful ideas beyond MIL that can foster further research in the field.

cs.LG

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.

cs.CL

DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

The rapid development of large language models has revolutionized code intelligence in software development. However, the predominance of closed-source models has restricted extensive research and development. To address this, we introduce the DeepSeek-Coder series, a range of open-source code models with sizes from 1.3B to 33B, trained from scratch on 2 trillion tokens. These models are pre-trained on a high-quality project-level code corpus and employ a fill-in-the-blank task with a 16K window to enhance code generation and infilling. Our extensive evaluations demonstrate that DeepSeek-Coder not only achieves state-of-the-art performance among open-source code models across multiple benchmarks but also surpasses existing closed-source models like Codex and GPT-3.5. Furthermore, DeepSeek-Coder models are under a permissive license that allows for both research and unrestricted commercial use.

cs.SE