SearcharxivSearch

arXiv subjects

Li

Publications and source records attributed to Li.

At least 37 records · Page 2Linked to original sources

Voltage Support Capability Analysis of Grid-Forming Inverters with Current-Limiting Control Under Asymmetrical Grid Faults

Voltage support capability is critical for grid-forming (GFM) inverters with current-limiting control (CLC) during grid faults. Despite the findings on the voltage support for symmetrical grid faults, its applicability to more common but complex asymmetrical grid faults has yet to be verified rigorously. This letter fills the gap in the voltage support capability analysis for asymmetrical grid faults by establishing and analyzing positive- and negative-sequence equivalent circuit models, where the virtual impedance is adopted to emulate various CLCs. It is discovered that matching the phase angle of the virtual impedance, emulated by the CLC, with that of the composed impedance from the capacitor to the fault location can maximize the voltage support capability of GFM inverters under asymmetrical grid faults. Rigorous theoretical analysis and experimental results verify this conclusion.

eess.SY

A Survey: Collaborative Hardware and Software Design in the Era of Large Language Models

The rapid development of large language models (LLMs) has significantly transformed the field of artificial intelligence, demonstrating remarkable capabilities in natural language processing and moving towards multi-modal functionality. These models are increasingly integrated into diverse applications, impacting both research and industry. However, their development and deployment present substantial challenges, including the need for extensive computational resources, high energy consumption, and complex software optimizations. Unlike traditional deep learning systems, LLMs require unique optimization strategies for training and inference, focusing on system-level efficiency. This paper surveys hardware and software co-design approaches specifically tailored to address the unique characteristics and constraints of large language models. This survey analyzes the challenges and impacts of LLMs on hardware and algorithm research, exploring algorithm optimization, hardware design, and system-level innovations. It aims to provide a comprehensive understanding of the trade-offs and considerations in LLM-centric computing systems, guiding future advancements in AI. Finally, we summarize the existing efforts in this space and outline future directions toward realizing production-grade co-design methodologies for the next generation of large language models and AI systems.

cs.AR

Distilling Text Style Transfer With Self-Explanation From LLMs

Text Style Transfer (TST) seeks to alter the style of text while retaining its core content. Given the constraints of limited parallel datasets for TST, we propose CoTeX, a framework that leverages large language models (LLMs) alongside chain-of-thought (CoT) prompting to facilitate TST. CoTeX distills the complex rewriting and reasoning capabilities of LLMs into more streamlined models capable of working with both non-parallel and parallel data. Through experimentation across four TST datasets, CoTeX is shown to surpass traditional supervised fine-tuning and knowledge distillation methods, particularly in low-resource settings. We conduct a comprehensive evaluation, comparing CoTeX against current unsupervised, supervised, in-context learning (ICL) techniques, and instruction-tuned LLMs. Furthermore, CoTeX distinguishes itself by offering transparent explanations for its style transfer process.

cs.CL

Fourier-Laplace transforms in polynomial Ornstein-Uhlenbeck volatility models

We consider the Fourier-Laplace transforms of a broad class of polynomial Ornstein-Uhlenbeck (OU) volatility models, including the well-known Stein-Stein, Schöbel-Zhu, one-factor Bergomi, and the recently introduced Quintic OU models motivated by the SPX-VIX joint calibration problem. We show the connection between the joint Fourier-Laplace functional of the log-price and the integrated variance, and the solution of an infinite dimensional Riccati equation. Next, under some non-vanishing conditions of the Fourier-Laplace transforms, we establish an existence result for such Riccati equation and we provide a discretized approximation of the joint characteristic functional that is exponentially entire. On the practical side, we develop a numerical scheme to solve the stiff infinite dimensional Riccati equations and demonstrate the efficiency and accuracy of the scheme for pricing SPX options and volatility swaps using Fourier and Laplace inversions, with specific examples of the Quintic OU and the one-factor Bergomi models and their calibration to real market data.

q-fin.MF

Self Expanding Convolutional Neural Networks

In this paper, we present a novel method for dynamically expanding Convolutional Neural Networks (CNNs) during training, aimed at meeting the increasing demand for efficient and sustainable deep learning models. Our approach, drawing from the seminal work on Self-Expanding Neural Networks (SENN), employs a natural expansion score as an expansion criteria to address the common issue of over-parameterization in deep convolutional neural networks, thereby ensuring that the model's complexity is finely tuned to the task's specific needs. A significant benefit of this method is its eco-friendly nature, as it obviates the necessity of training multiple models of different sizes. We employ a strategy where a single model is dynamically expanded, facilitating the extraction of checkpoints at various complexity levels, effectively reducing computational resource use and energy consumption while also expediting the development cycle by offering diverse model complexities from a single training session. We evaluate our method on the CIFAR-10 dataset and our experimental results validate this approach, demonstrating that dynamically adding layers not only maintains but also improves CNN performance, underscoring the effectiveness of our expansion criteria. This approach marks a considerable advancement in developing adaptive, scalable, and environmentally considerate neural network architectures, addressing key challenges in the field of deep learning.

cs.CV

Radio Jet Feedback on the Inner Disk of Virgo Spiral Galaxy Messier 58

Spitzer spectral maps reveal a disk of highly luminous, warm (>150 K) H2 in the center of the massive spiral galaxy Messier 58, which hosts a radio-loud AGN. The inner 2.6 kpc of the galaxy appears to be overrun by shocks from the radio jet cocoon. Gemini NIRI imaging of the H2 1-0 S(1) emission line, ALMA CO 2-1, and HST multiband imagery indicate that much of the molecular gas is shocked in-situ, corresponding to lanes of dusty molecular gas that spiral towards the galaxy nucleus. The CO 2-1 and ionized gas kinematics are highly disturbed, with velocity dispersion up to 300 km/s. Dissipation of the associated kinetic energy and turbulence, likely injected into the ISM by radio-jet driven outflows, may power the observed molecular and ionized gas emission from the inner disk. The PAH fraction and composition in the inner disk appear to be normal, in spite of the jet and AGN activity. The PAH ratios are consistent with excitation by the interstellar radiation field from old stars in the bulge, with no contribution from star formation. The phenomenon of jet-shocked H2 may substantially reduce star formation and help to regulate the stellar mass of the inner disk and supermassive black hole in this otherwise normal spiral galaxy. Similarly strong H2 emission is found at the centers of several nearby spiral and lenticular galaxies with massive bulges and radio-loud AGN.

astro-ph.GA

Towards Fast, Adaptive, and Hardware-Assisted User-Space Scheduling

Modern datacenter applications are prone to high tail latencies since their requests typically follow highly-dispersive distributions. Delivering fast interrupts is essential to reducing tail latency. Prior work has proposed both OS- and system-level solutions to reduce tail latencies for microsecond-scale workloads through better scheduling. Unfortunately, existing approaches like customized dataplane OSes, require significant OS changes, experience scalability limitations, or do not reach the full performance capabilities hardware offers. The emergence of new hardware features like UINTR exposed new opportunities to rethink the design paradigms and abstractions of traditional scheduling systems. We propose LibPreemptible, a preemptive user-level threading library that is flexible, lightweight, and adaptive. LibPreemptible was built with a set of optimizations like LibUtimer for scalability, and deadline-oriented API for flexible policies, time-quantum controller for adaptiveness. Compared to the prior state-of-the-art scheduling system Shinjuku, our system achieves significant tail latency and throughput improvements for various workloads without modifying the kernel. We also demonstrate the flexibility of LibPreemptible across scheduling policies for real applications experiencing varying load levels and characteristics.

cs.DC

Considerations for Master Protocols Using External Controls

There has been an increasing use of master protocols in oncology clinical trials because of its efficiency and flexibility to accelerate cancer drug development. Depending on the study objective and design, a master protocol trial can be a basket trial, an umbrella trial, a platform trial, or any other form of trials in which multiple investigational products and/or subpopulations are studied under a single protocol. Master protocols can use external data and evidence (e.g., external controls) for treatment effect estimation, which can further improve efficiency of master protocol trials. This paper provides an overview of different types of external controls and their unique features when used in master protocols. Some key considerations in master protocols with external controls are discussed including construction of estimands, assessment of fit-for-use real-world data, and considerations for different types of master protocols. Similarities and differences between regular randomized controlled trials and master protocols when using external controls are discussed. A targeted learning-based causal roadmap is presented which constitutes three key steps: (1) define a target statistical estimand that aligns with the causal estimand for the study objective, (2) use an efficient estimator to estimate the target statistical estimand and its uncertainty, and (3) evaluate the impact of causal assumptions on the study conclusion by performing sensitivity analyses. Two illustrative examples for master protocols using external controls are discussed for their merits and possible improvement in causal effect estimation.

stat.AP

Foundation Models for Generalist Geospatial Artificial Intelligence

Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled datasets through self-supervision, and then fine-tuned for various downstream tasks with small labeled datasets. This paper introduces a first-of-a-kind framework for the efficient pre-training and fine-tuning of foundational models on extensive geospatial data. We have utilized this framework to create Prithvi, a transformer-based geospatial foundational model pre-trained on more than 1TB of multispectral satellite imagery from the Harmonized Landsat-Sentinel 2 (HLS) dataset. Our study demonstrates the efficacy of our framework in successfully fine-tuning Prithvi to a range of Earth observation tasks that have not been tackled by previous work on foundation models involving multi-temporal cloud gap imputation, flood mapping, wildfire scar segmentation, and multi-temporal crop segmentation. Our experiments show that the pre-trained model accelerates the fine-tuning process compared to leveraging randomly initialized weights. In addition, pre-trained Prithvi compares well against the state-of-the-art, e.g., outperforming a conditional GAN model in multi-temporal cloud imputation by up to 5pp (or 5.7%) in the structural similarity index. Finally, due to the limited availability of labeled data in the field of Earth observation, we gradually reduce the quantity of available labeled data for refining the model to evaluate data efficiency and demonstrate that data can be decreased significantly without affecting the model's accuracy. The pre-trained 100 million parameter model and corresponding fine-tuning workflows have been released publicly as open source contributions to the global Earth sciences community through Hugging Face.

cs.CV

A Search for Technosignatures Around 11,680 Stars with the Green Bank Telescope at 1.15-1.73 GHz

We conducted a search for narrowband radio signals over four observing sessions in 2020-2023 with the L-band receiver (1.15-1.73 GHz) of the 100 m diameter Green Bank Telescope. We pointed the telescope in the directions of 62 TESS Objects of Interest, capturing radio emissions from a total of ~11,680 stars and planetary systems in the ~9 arcminute beam of the telescope. All detections were either automatically rejected or visually inspected and confirmed to be of anthropogenic nature. In this work, we also quantified the end-to-end efficiency of radio SETI pipelines with a signal injection and recovery analysis. The UCLA SETI pipeline recovers 94.0% of the injected signals over the usable frequency range of the receiver and 98.7% of the injections when regions of dense RFI are excluded. In another pipeline that uses incoherent sums of 51 consecutive spectra, the recovery rate is ~15 times smaller at ~6%. The pipeline efficiency affects calculations of transmitter prevalence and SETI search volume. Accordingly, we developed an improved Drake Figure of Merit and a formalism to place upper limits on transmitter prevalence that take the pipeline efficiency and transmitter duty cycle into account. Based on our observations, we can state at the 95% confidence level that fewer than 6.6% of stars within 100 pc host a transmitter that is detectable in our search (EIRP > 1e13 W). For stars within 20,000 ly, the fraction of stars with detectable transmitters (EIRP > 5e16 W) is at most 3e-4. Finally, we showed that the UCLA SETI pipeline natively detects the signals detected with AI techniques by Ma et al. (2023).

astro-ph.IM

LESS: Label-efficient Multi-scale Learning for Cytological Whole Slide Image Screening

In computational pathology, multiple instance learning (MIL) is widely used to circumvent the computational impasse in giga-pixel whole slide image (WSI) analysis. It usually consists of two stages: patch-level feature extraction and slide-level aggregation. Recently, pretrained models or self-supervised learning have been used to extract patch features, but they suffer from low effectiveness or inefficiency due to overlooking the task-specific supervision provided by slide labels. Here we propose a weakly-supervised Label-Efficient WSI Screening method, dubbed LESS, for cytological WSI analysis with only slide-level labels, which can be effectively applied to small datasets. First, we suggest using variational positive-unlabeled (VPU) learning to uncover hidden labels of both benign and malignant patches. We provide appropriate supervision by using slide-level labels to improve the learning of patch-level features. Next, we take into account the sparse and random arrangement of cells in cytological WSIs. To address this, we propose a strategy to crop patches at multiple scales and utilize a cross-attention vision transformer (CrossViT) to combine information from different scales for WSI classification. The combination of our two steps achieves task-alignment, improving effectiveness and efficiency. We validate the proposed label-efficient method on a urine cytology WSI dataset encompassing 130 samples (13,000 patches) and FNAC 2019 dataset with 212 samples (21,200 patches). The experiment shows that the proposed LESS reaches 84.79%, 85.43%, 91.79% and 78.30% on a urine cytology WSI dataset, and 96.88%, 96.86%, 98.95%, 97.06% on FNAC 2019 dataset in terms of accuracy, AUC, sensitivity and specificity. It outperforms state-of-the-art MIL methods on pathology WSIs and realizes automatic cytological WSI cancer screening.

eess.IV

The quintic Ornstein-Uhlenbeck volatility model that jointly calibrates SPX & VIX smiles

The quintic Ornstein-Uhlenbeck volatility model is a stochastic volatility model where the volatility process is a polynomial function of degree five of a single Ornstein-Uhlenbeck process with fast mean reversion and large vol-of-vol. The model is able to achieve remarkable joint fits of the SPX-VIX smiles with only 6 effective parameters and an input curve that allows to match certain term structures. We provide several practical specifications of the input curve, study their impact on the joint calibration problem and consider additionally time-dependent parameters to help achieve better fits for longer maturities going beyond 1 year. Even better, the model remains very simple and tractable for pricing and calibration: the VIX squared is again polynomial in the Ornstein-Uhlenbeck process, leading to efficient VIX derivative pricing by a simple integration against a Gaussian density; simulation of the volatility process is exact; and pricing SPX products derivatives can be done efficiently and accurately by standard Monte Carlo techniques with suitable antithetic and control variates.

q-fin.MF

A Cross-direction Task Decoupling Network for Small Logo Detection

Logo detection plays an integral role in many applications. However, handling small logos is still difficult since they occupy too few pixels in the image, which burdens the extraction of discriminative features. The aggregation of small logos also brings a great challenge to the classification and localization of logos. To solve these problems, we creatively propose Cross-direction Task Decoupling Network (CTDNet) for small logo detection. We first introduce Cross-direction Feature Pyramid (CFP) to realize cross-direction feature fusion by adopting horizontal transmission and vertical transmission. In addition, Multi-frequency Task Decoupling Head (MTDH) decouples the classification and localization tasks into two branches. A multi frequency attention convolution branch is designed to achieve more accurate regression by combining discrete cosine transform and convolution creatively. Comprehensive experiments on four logo datasets demonstrate the effectiveness and efficiency of the proposed method.

cs.CV

Chemo-dynamical substructure in the M31 inner halo globular clusters: Further evidence for a recent accretion event

Based upon a metallicity selection, we identify a significant sub-population of the inner halo globular clusters in the Andromeda Galaxy which we name the Dulais Structure. It is distinguished as a co-rotating group of 10-20 globular clusters which appear to be kinematically distinct from, and on average more metal-poor than, the majority of the inner halo population. Intriguingly, the orbital axis of this Dulais Structure is closely aligned with that of the younger accretion event recently identified using a sub-population of globular clusters in the outer halo of Andromeda, and this is strongly suggestive of a causal relationship between the two. If this connection is confirmed, a natural explanation for the kinematics of the globular clusters in the Dulais Structure is that they trace the accretion of a substantial progenitor (~10^11 Msun) into the halo of Andromeda during the last few billion years, that may have occurred as part of a larger group infall.

astro-ph.GA

High capacity topological coding based on nested vortex knots and links

Optical knots and links have attracted great attention because of their exotic topological characteristics. Recent investigations have shown that the information encoding based on optical knots could possess robust features against external perturbations. However, as a superior coding scheme, it is also necessary to achieve a high capacity, which is hard to be fulfilled by existing knot-carriers owing to the limit number of associated topological invariants. Thus, how to realize the knot-based information coding with a high capacity is a key problem to be solved. Here, we create a type of nested vortex knot, and show that it can be used to fulfill the robust information coding with a high capacity assisted by a large number of intrinsic topological invariants. In experiments, we design and fabricate metasurface holograms to generate light fields sustaining different kinds of nested vortex links. Furthermore, we verify the feasibility of the high-capacity coding scheme based on those topological optical knots. Our work opens another way to realize the robust and high capacity optical coding, which may have useful impacts on the field of information transfer and storage.

physics.optics

Closing the "Quantum Supremacy" Gap: Achieving Real-Time Simulation of a Random Quantum Circuit Using a New Sunway Supercomputer

We develop a high-performance tensor-based simulator for random quantum circuits(RQCs) on the new Sunway supercomputer. Our major innovations include: (1) a near-optimal slicing scheme, and a path-optimization strategy that considers both complexity and compute density; (2) a three-level parallelization scheme that scales to about 42 million cores; (3) a fused permutation and multiplication design that improves the compute efficiency for a wide range of tensor contraction scenarios; and (4) a mixed-precision scheme to further improve the performance. Our simulator effectively expands the scope of simulatable RQCs to include the 10*10(qubits)*(1+40+1)(depth) circuit, with a sustained performance of 1.2 Eflops (single-precision), or 4.4 Eflops (mixed-precision)as a new milestone for classical simulation of quantum circuits; and reduces the simulation sampling time of Google Sycamore to 304 seconds, from the previously claimed 10,000 years.

quant-ph

Directional dark-field implicit x-ray speckle tracking using an anisotropic-diffusion Fokker-Planck equation

When a macroscopic-sized non-crystalline sample is illuminated using coherent x-ray radiation, a bifurcation of photon energy flow may occur. The coarse-grained complex refractive index of the sample may be considered to attenuate and refract the incident coherent beam, leading to a coherent component of the transmitted beam. Spatially-unresolved sample microstructure, associated with the fine-grained components of the complex refractive index, introduces a diffuse component to the transmitted beam. This diffuse photon-scattering channel may be viewed in terms of position-dependent fans of ultra-small-angle x-ray scatter. These position-dependent fans, at the exit surface of the object, may under certain circumstances be approximated as having a locally-elliptical shape. By using an anisotropic-diffusion Fokker-Planck approach to model this bifurcated x-ray energy flow, we show how all three components (attenuation, refraction and locally-elliptical diffuse scatter) may be recovered. This is done via x-ray speckle tracking, in which the sample is illuminated with spatially-random x-ray fields generated by coherent illumination of a spatially-random membrane. The theory is developed, and then successfully applied to experimental x-ray data.

physics.med-ph

AlphaZero Based Post-Storm Repair Crew Dispatch for Distribution Grid Restoration

Natural disasters such as storms usually bring significant damages to distribution grids. This paper investigates the optimal routing of utility vehicles to restore outages in the distribution grid as fast as possible after a storm. First, the poststorm repair crew dispatch task with multiple utility vehicles is formulated as a sequential stochastic optimization problem. In the formulated optimization model, the belief state of the power grid is updated according to the phone calls from customers and the information collected by utility vehicles. Second, an AlphaZero[1] based utility vehicle routing (AlphaZero-UVR) approach is developed to achieve the real-time dispatching of the repair crews. The proposed AlphaZero-UVR approach combines deep neural networks with stochastic Monte-Carlo tree search (MCTS) to give a lookahead search decisions, which can learn to navigate repair crews without human guidance. Simulation results show that the proposed approach can efficiently navigate crews to repair all outages.

eess.SY