SearcharxivSearch

arXiv subjects

Vincent Oria

Publications and source records attributed to Vincent Oria.

13 recordsLinked to original sources

DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation

Prompting-based (i.e., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not changed for the task, face three problems: (i) relying on coarse-grained schema information that may not reveal the fine-grained relationships needed to distinguish ambiguous columns, (ii) failing to capture recurring SQL-generation failures, and (iii) suffering from omission or hallucination of components in complex questions. This paper develops DexterSQL, a prompting/non-fine-tuning-based Text-to-SQL system that improves SQL generation with three novel components: (i) deep schema explorator that identifies ambiguous columns, analyzes their individual and joint data distributions to uncover their relationships and the distinct role of each, (ii) database-agnostic rule creator that mines mismatches between generated and gold SQL only on the training database and converts them into database-agnostic corrective rules that capture recurring LLM failure patterns; and (iii) multi-path SQL generation that introduces a dependency-tree-based intermediate representation that uses the question's sentence structure to guide its decomposition into an SQL skeleton for final SQL generation. DexterSQL achieves a higher accuracy compared to the state-of-the-art using both open-source/weight and closed-source/weight models. Particularly, DexterSQL shows a high improvement of at least 5.5% using an open-weight model (GPT-OSS-120B) on BIRDDev, with total accuracy 70.4%. DexterSQL also shows better improvement of at least 1.4% using closed-weight models, with total accuracy 72.1% and 72.9% on BIRD-Dev with GPT-4o and GPT-5.2.

cs.DB

Daily Predictions of F10.7 and F30 Solar Indices with Deep Learning

The F10.7 and F30 solar indices are the solar radio fluxes measured at wavelengths of 10.7 cm and 30 cm, respectively, which are key indicators of solar activity. F10.7 is valuable for explaining the impact of solar ultraviolet (UV) radiation on the upper atmosphere of Earth, while F30 is more sensitive and could improve the reaction of thermospheric density to solar stimulation. In this study, we present a new deep learning model, named the Solar Index Network, or SINet for short, to predict daily values of the F10.7 and F30 solar indices. The SINet model is designed to make medium-term predictions of the index values (1-60 days in advance). The observed data used for SINet training were taken from the National Oceanic and Atmospheric Administration (NOAA) as well as Toyokawa and Nobeyama facilities. Our experimental results show that SINet performs better than five closely related statistical and deep learning methods for the prediction of F10.7. Furthermore, to our knowledge, this is the first time deep learning has been used to predict the F30 solar index.

astro-ph.SR

FedDAG: Clustered Federated Learning via Global Data and Gradient Integration for Heterogeneous Environments

Federated Learning (FL) enables a group of clients to collaboratively train a model without sharing individual data, but its performance drops when client data are heterogeneous. Clustered FL tackles this by grouping similar clients. However, existing clustered FL approaches rely solely on either data similarity or gradient similarity; however, this results in an incomplete assessment of client similarities. Prior clustered FL approaches also restrict knowledge and representation sharing to clients within the same cluster. This prevents cluster models from benefiting from the diverse client population across clusters. To address these limitations, FedDAG introduces a clustered FL framework, FedDAG, that employs a weighted, class-wise similarity metric that integrates both data and gradient information, providing a more holistic measure of similarity during clustering. In addition, FedDAG adopts a dual-encoder architecture for cluster models, comprising a primary encoder trained on its own clients' data and a secondary encoder refined using gradients from complementary clusters. This enables cross-cluster feature transfer while preserving cluster-specific specialization. Experiments on diverse benchmarks and data heterogeneity settings show that FedDAG consistently outperforms state-of-the-art clustered FL baselines in accuracy.

cs.LG

Time Series of Magnetic Field Parameters of Merged MDI and HMI Space-Weather Active Region Patches as Potential Tool for Solar Flare Forecasting

Solar flare prediction studies have been recently conducted with the use of Space-Weather MDI (Michelson Doppler Imager onboard Solar and Heliospheric Observatory) Active Region Patches (SMARP) and Space-Weather HMI (Helioseismic and Magnetic Imager onboard Solar Dynamics Observatory) Active Region Patches (SHARP), which are two currently available data products containing magnetic field characteristics of solar active regions. The present work is an effort to combine them into one data product, and perform some initial statistical analyses in order to further expand their application in space weather forecasting. The combined data are derived by filtering, rescaling, and merging the SMARP with SHARP parameters, which can then be spatially reduced to create uniform multivariate time series. The resulting combined MDI-HMI dataset currently spans the period between April 4, 1996, and December 13, 2022, and may be extended to a more recent date. This provides an opportunity to correlate and compare it with other space weather time series, such as the daily solar flare index or the statistical properties of the soft X-ray flux measured by the Geostationary Operational Environmental Satellites (GOES). Time-lagged cross-correlation indicates that a relationship may exist, where some magnetic field properties of active regions lead the flare index in time. Applying the rolling window technique makes it possible to see how this leader-follower dynamic varies with time. Preliminary results indicate that areas of high correlation generally correspond to increased flare activity during the peak solar cycle.

astro-ph.SR

Predicting Solar Proton Events of Solar Cycles 22-24 using GOES Proton & soft X-Ray flux features

Solar Energetic Particle (SEP) events and their major subclass, Solar Proton Events (SPEs), can have unfavorable consequences on numerous aspects of life and technology, making them one of the most harmful effects of solar activity. Garnering knowledge preceding such events by studying operational data flows is essential for their forecasting. Considering only Solar Cycle (SC) 24 in our previous study, Sadykov et al. 2021, we found that it may be sufficient to utilize only proton and soft X-ray (SXR) parameters for SPE forecasts. Here, we report a catalog recording $\geq$ 10 MeV $\geq$ 10 particle flux unit SPEs with their properties, spanning SCs 22-24, using NOAA's Geostationary Operational Environmental Satellite flux data. We report an additional catalog of daily proton and SXR flux statistics for this period, employing it to test the application of machine learning (ML) on the prediction of SPEs using a Support Vector Machine (SVM) and eXtreme Gradient Boosting (XGBoost). We explore the effects of training models with data from one and two SCs, evaluating how transferable a model can be across different time periods. XGBoost proved to be more accurate than SVMs for almost every test considered, while outperforming operational SWPC NOAA predictions and a persistence forecast. Interestingly, training done with SC 24 produces weaker TSS and HSS2, even when paired with SC 22 or SC 23, indicating transferability issues. This work contributes towards validating forecasts using long-spanning data -- an understudied area in SEP research that should be considered to verify the cross-cycle robustness of ML-driven forecasts.

astro-ph.SR

Statistical Study of the Correlation between Solar Energetic Particles and Properties of Active Regions

The flux of energetic particles originating from the Sun fluctuates during the solar cycles. It depends on the number and properties of Active Regions (ARs) present in a single day and associated solar activities, such as solar flares and coronal mass ejections (CMEs). Observational records of the Space Weather Prediction Center (SWPC NOAA) enable the creation of time-indexed databases containing information about ARs and particle flux enhancements, most widely known as Solar Energetic Particle events (SEPs). In this work, we utilize the data available for Solar Cycles 21-24, and the initial phase of Cycle 25 to perform a statistical analysis of the correlation between SEPs and properties of ARs inferred from the McIntosh and Hale classifications. We find that the complexity of the magnetic field, longitudinal location, area, and penumbra type of the largest sunspot of ARs are most correlated with the production of SEPs. It is found that most SEPs ($\approx$60\%, or 108 out of 181 considered events) were generated from an AR classified with the 'k' McIntosh subclass as the second component, and these ARs are more likely to produce SEPs if they fall in a Hale class containing $δ$ component. The resulting database containing information about SEP events and ARs is publicly available and can be used for the development of Machine Learning (ML) models to predict the occurrence of SEPs.

astro-ph.SR

The Random Hivemind: An Ensemble Deep Learner Application to Solar Energetic Particle Prediction Problem

The application of machine learning and deep learning, including the wide use of non-ensemble, conventional neural networks (CoNN), for predicting various phenomena has become very popular in recent years thanks to the efficiencies and the abilities of these techniques to find relationships in data without human intervention. However, certain CoNN setups may not work on some datasets, especially if the parameters passed to it, including model parameters and hyperparameters, are arguably arbitrary in nature and need to continuously be updated with the need to retrain the model. This concern can be partially alleviated by employing committees of neural networks that are identical in terms of input features and architectures, initialized randomly, and "vote" on the decisions made by the committees as a whole. Yet, it is possible for the committee members to "agree" on identical sets of weights and biases for all nodes and edges. Members of these committees also cannot be expanded to accommodate new features and entire committees must therefore be retrained in order to do so. We propose the Random Hivemind (RH) approach, which helps to alleviate this concern by having multiple neural network estimators make decisions based on random permutations of features and prescribing a method to determine the weight of the decision of each individual estimator. The effectiveness of RH is demonstrated through experimentation in the predictions of hazardous Solar Energetic Particle (SEP) events by comparing it to that of using both CoNNs and the aforementioned setup of committees. Our results demonstrate that RH, while having a comparable or better performance than the CoNN and a Committee-based approach, demonstrates a lesser score spread for the individual experiments, and shows promising results with respect to capturing almost every single flare instance leading to SEPs.

astro-ph.SR

Revisiting the Solar Research Cyberinfrastructure Needs: A White Paper of Findings and Recommendations

Solar and Heliosphere physics are areas of remarkable data-driven discoveries. Recent advances in high-cadence, high-resolution multiwavelength observations, growing amounts of data from realistic modeling, and operational needs for uninterrupted science-quality data coverage generate the demand for a solar metadata standardization and overall healthy data infrastructure. This white paper is prepared as an effort of the working group "Uniform Semantics and Syntax of Solar Observations and Events" created within the "Towards Integration of Heliophysics Data, Modeling, and Analysis Tools" EarthCube Research Coordination Network (@HDMIEC RCN), with primary objectives to discuss current advances and identify future needs for the solar research cyberinfrastructure. The white paper summarizes presentations and discussions held during the special working group session at the EarthCube Annual Meeting on June 19th, 2020, as well as community contribution gathered during a series of preceding workshops and subsequent RCN working group sessions. The authors provide examples of the current standing of the solar research cyberinfrastructure, and describe the problems related to current data handling approaches. The list of the top-level recommendations agreed by the authors of the current white paper is presented at the beginning of the paper.

astro-ph.IM

Prediction of Solar Proton Events with Machine Learning: Comparison with Operational Forecasts and "All-Clear" Perspectives

Solar Energetic Particle events (SEPs) are among the most dangerous transient phenomena of solar activity. As hazardous radiation, SEPs may affect the health of astronauts in outer space and adversely impact current and future space exploration. In this paper, we consider the problem of daily prediction of Solar Proton Events (SPEs) based on the characteristics of the magnetic fields in solar Active Regions (ARs), preceding soft X-ray and proton fluxes, and statistics of solar radio bursts. The machine learning (ML) algorithm uses an artificial neural network of custom architecture designed for whole-Sun input. The predictions of the ML model are compared with the SWPC NOAA operational forecasts of SPEs. Our preliminary results indicate that 1) for the AR-based predictions, it is necessary to take into account ARs at the western limb and on the far side of the Sun; 2) characteristics of the preceding proton flux represent the most valuable input for prediction; 3) daily median characteristics of ARs and the counts of type II, III, and IV radio bursts may be excluded from the forecast without performance loss; and 4) ML-based forecasts outperform SWPC NOAA forecasts in situations in which missing SPE events is very undesirable. The introduced approach indicates the possibility of developing robust "all-clear" SPE forecasts by employing machine learning methods.

astro-ph.SR

Compression of Solar Spectroscopic Observations: a Case Study of Mg II k Spectral Line Profiles Observed by NASA's IRIS Satellite

In this study we extract the deep features and investigate the compression of the Mg II k spectral line profiles observed in quiet Sun regions by NASA's IRIS satellite. The data set of line profiles used for the analysis was obtained on April 20th, 2020, at the center of the solar disc, and contains almost 300,000 individual Mg II k line profiles after data cleaning. The data are separated into train and test subsets. The train subset was used to train the autoencoder of the varying embedding layer size. The early stopping criterion was implemented on the test subset to prevent the model from overfitting. Our results indicate that it is possible to compress the spectral line profiles more than 27 times (which corresponds to the reduction of the data dimensionality from 110 to 4) while having a 4 DN average reconstruction error, which is comparable to the variations in the line continuum. The mean squared error and the reconstruction error of even statistical moments sharply decrease when the dimensionality of the embedding layer increases from 1 to 4 and almost stop decreasing for higher numbers. The observed occasional improvements in training for values higher than 4 indicate that a better compact embedding may potentially be obtained if other training strategies and longer training times are used. The features learned for the critical four-dimensional case can be interpreted. In particular, three of these four features mainly control the line width, line asymmetry, and line dip formation respectively. The presented results are the first attempt to obtain a compact embedding for spectroscopic line profiles and confirm the value of this approach, in particular for feature extraction, data compression, and denoising.

astro-ph.SR

Machine Learning in Heliophysics and Space Weather Forecasting: A White Paper of Findings and Recommendations

The authors of this white paper met on 16-17 January 2020 at the New Jersey Institute of Technology, Newark, NJ, for a 2-day workshop that brought together a group of heliophysicists, data providers, expert modelers, and computer/data scientists. Their objective was to discuss critical developments and prospects of the application of machine and/or deep learning techniques for data analysis, modeling and forecasting in Heliophysics, and to shape a strategy for further developments in the field. The workshop combined a set of plenary sessions featuring invited introductory talks interleaved with a set of open discussion sessions. The outcome of the discussion is encapsulated in this white paper that also features a top-level list of recommendations agreed by participants.

astro-ph.SR

Roadmap for Reliable Ensemble Forecasting of the Sun-Earth System

The authors of this report met on 28-30 March 2018 at the New Jersey Institute of Technology, Newark, New Jersey, for a 3-day workshop that brought together a group of data providers, expert modelers, and computer and data scientists, in the solar discipline. Their objective was to identify challenges in the path towards building an effective framework to achieve transformative advances in the understanding and forecasting of the Sun-Earth system from the upper convection zone of the Sun to the Earth's magnetosphere. The workshop aimed to develop a research roadmap that targets the scientific challenge of coupling observations and modeling with emerging data-science research to extract knowledge from the large volumes of data (observed and simulated) while stimulating computer science with new research applications. The desire among the attendees was to promote future trans-disciplinary collaborations and identify areas of convergence across disciplines. The workshop combined a set of plenary sessions featuring invited introductory talks and workshop progress reports, interleaved with a set of breakout sessions focused on specific topics of interest. Each breakout group generated short documents, listing the challenges identified during their discussions in addition to possible ways of attacking them collectively. These documents were combined into this report-wherein a list of prioritized activities have been collated, shared and endorsed.

astro-ph.SR

Interactive Multi-Instrument Database of Solar Flares

Solar flares are complicated physical phenomena that are observable in a broad range of the electromagnetic spectrum, from radiowaves to $γ$-rays. For a more comprehensive understanding of flares, it is necessary to perform a combined multi-wavelength analysis using observations from many satellites and ground-based observatories. For efficient data search, integration of different flare lists and representation of observational data, we have developed an Interactive Multi-Instrument Database of Solar Flares (https://solarflare.njit.edu/). The web accessible database is fully functional and allows the user to search for uniquely-identified flare events based on their physical descriptors and availability of observations by a particular set of instruments. Currently, the data from three primary flare lists (GOES, RHESSI and HEK) and a variety of other event catalogs (Hinode, Fermi GBM, Konus-Wind, OVSA flare catalogs, CACTus CME catalog, Filament eruption catalog) and observing logs (IRIS and Nobeyama coverage) are integrated, and an additional set of physical descriptors (temperature and emission measure) is provided along with an observing summary, data links, and multi-wavelength light curves for each flare event since January, 2002. We envision that this new tool will allow researchers to significantly speed up the search of events of interest for statistical and case studies.

astro-ph.SR