SearcharxivSearch

arXiv subjects

Ahsan Ali

Publications and source records attributed to Ahsan Ali.

15 recordsLinked to original sources

2D MoS$_2$/Au interfaces for enhanced opto-electronic response with sub-bandgap photons

Monolayer MoS$_2$ is a direct band gap semiconductor with potential applications in optoelectronics and photonics. MoS$_2$ also has a large optical nonlinearity. However, the atomic thickness of the monolayer limits the strength of the measured functional signals, such as the photocurrent or photoluminescence, in optoelectronic devices. Here, we show that photocurrent in monolayer MoS$_2$ can be induced by sub-band gap photons by depositing Au nanoparticles on it. In this system, the nonlinear light-matter interaction in Au nanoparticles enhanced by the localized surface plasmons results in the generation of supercontinuum, which is reabsorbed by MoS$_2$ due to efficient resonant energy transfer. Au nanoparticle assisted photocurrent is more than an order of magnitude larger than two-photon photocurrent in monolayer MoS$_2$. Optimization of the shape, size and composition of the nanoparticle has the potential to enhance the photocurrent significantly with the prospect of applications in the detection of NIR photons, and related technologies including optical telecommunication.

physics.optics

ND1 centers in diamond for long-term data storage in extreme conditions

Practically feasible long-term data storage under extreme conditions is an unsolved problem in modern data storage systems. This study introduces a novel approach using ND1 centers in diamonds for high-density, three-dimensional optical data storage. By employing near-infrared femtosecond laser pulses, we demonstrate the creation of sub-micron ND1 defect sites with precise spatial control, enabling efficient data encoding as luminescent ''pits." The ND1 centers exhibit robust photoluminescence in the UV spectrum, driven by three-photon absorption, which intrinsically provides a 3D reading of the data. Remarkably, these centers remain stable under extreme electric and magnetic fields, temperatures ranging from 4 K to 500 K, and corrosive chemical environments, with no degradation observed over extended periods. A reading speed of 500 MBits/s, limited by the lifetime of the photoluminescence, surpasses conventional Blu-ray technology while maintaining compatibility with existing optical data storage infrastructure. Our findings highlight diamond-based ND1 centers as a promising medium for durable, high-capacity data storage, capable of preserving critical information for millions of years, even under harsh conditions.

physics.optics

Optimal transfer operators in algebraic two-level methods for nonsymmetric and indefinite problems

Consider an algebraic two-level method applied to the $n$-dimensional linear system $A \mathbf{x} = \mathbf{b}$ using fine-space preconditioner (i.e., ``relaxation'' or ``smoother'') $M$, with $M \approx A$, restriction and interpolation $R$ and $P$, and algebraic coarse-space operator ${A_c := R^*AP}$. Then, what are the the best possible transfer operators $R$ and $P$ of a given dimension $n_c < n$? Brannick et al. (2018) showed that when $A$ and $M$ are Hermitian positive definite (HPD), the optimal interpolation is such that its range contains the $n_c$ smallest generalized eigenvectors of the matrix pencil $(A, M)$. Recently, in Ali et al. (2025) we generalized this framework to the non-HPD setting, by considering both right (interpolation) and left (restriction) generalized eigenvectors of $(A, M)$ and defining corresponding nonsymmetric transfer operators $\{R_\#,P_\#\}$. Tight convergence bounds for $\{R_\#,P_\#\}$ are derived in spectral radius, as well as a proof of pseudo-optimality. Note, $\{R_\#,P_\#\}$ are typically complex valued, which is not practical for real-valued problems. Here we build on Ali et al. (2025), first characterizing all inner products in which the coarse-space correction defined by $\{R_\#,P_\#\}$ is orthogonal. We then develop tight two-level convergence bounds in these norms, and prove that the underlying transfer operators $\{R_\#,P_\#\}$ are genuinely optimal. As a special case, our theory both recovers and extends the HPD results from Brannick et al. (2018). Finally, we show how to construct optimal, real-valued transfer operators in the case of that $A$ and $M$ are real valued, but are not HPD. Numerical examples arising from discretized advection and wave-equation problems are used to verify and illustrate the theory.

math.NA

Generalized Optimal AMG Convergence Theory for Stokes Equations Using Smooth Aggregation and Vanka Relaxation Strategies

This paper discusses our recent generalized optimal algebraic multigrid (AMG) convergence theory applied to the steady-state Stokes equations discretized using Taylor-Hood elements ($\pmb{ \mathbb{P}}_2/\mathbb{P}_{1}$). The generalized theory is founded on matrix-induced orthogonality of the left and right eigenvectors of a generalized eigenvalue problem involving the system matrix and relaxation operator. This framework establishes a rigorous lower bound on the spectral radius of the two-grid error-propagation operator, enabling precise predictions of the convergence rate for symmetric indefinite problems, such as those arising from saddle-point systems. We apply this theory to the recently developed monolithic smooth aggregation AMG (SA-AMG) solver for Stokes, constructed using evolution-based strength of connection, standard aggregation, and smoothed prolongation. The performance of these solvers is evaluated using additive and multiplicative Vanka relaxation strategies. Additive Vanka relaxation constructs patches algebraically on each level, resulting in a nonsymmetric relaxation operator due to the partition of unity being applied on one side of the block-diagonal matrix. Although symmetry can be restored by eliminating the partition of unity, this compromises convergence. Alternatively, multiplicative Vanka relaxation updates velocity and pressure sequentially within each patch, propagating updates multiplicatively across the domain and effectively addressing velocity-pressure coupling, ensuring a symmetric relaxation. We demonstrate that the generalized optimal AMG theory consistently provides accurate lower bounds on the convergence rate for SA-AMG applied to Stokes equations. These findings suggest potential avenues for further enhancement in AMG solver design for saddle-point systems.

math.NA

Constrained Local Approximate Ideal Restriction for Advection-Diffusion Problems

This paper focuses on developing a reduction-based algebraic multigrid method that is suitable for solving general (non)symmetric linear systems and is naturally robust from pure advection to pure diffusion. Initial motivation comes from a new reduction-based algebraic multigrid (AMG) approach, $\ell$AIR (local approximate ideal restriction), that was developed for solving advection-dominated problems. Though this new solver is very effective in the advection dominated regime, its performance degrades in cases where diffusion becomes dominant. This is consistent with the fact that in general, reduction-based AMG methods tend to suffer from growth in complexity and/or convergence rates as the problem size is increased, especially for diffusion dominated problems in two or three dimensions. Motivated by the success of $\ell$AIR in the advective regime, our aim in this paper is to generalize the AIR framework with the goal of improving the performance of the solver in diffusion dominated regimes. To do so, we propose a novel way to combine mode constraints as used commonly in energy minimization AMG methods with the local approximation of ideal operators used in $\ell$AIR. The resulting constrained $\ell$AIR (C$\ell$AIR) algorithm is able to achieve fast scalable convergence on advective and diffusive problems. In addition, it is able to achieve standard low complexity hierarchies in the diffusive regime through aggressive coarsening, something that has been previously difficult for reduction-based methods.

math.NA

Generalized Optimal AMG Convergence Theory for Nonsymmetric and Indefinite Problems

Algebraic multigrid (AMG) is known to be an effective solver for many sparse symmetric positive definite (SPD) linear systems. For SPD systems, the convergence theory of AMG is well-understood in terms of the $A$-norm, but in a nonsymmetric setting, such an energy norm is non-existent. For this reason, convergence of AMG for nonsymmetric systems of equations remains an open area of research. A particular aspect missing from theory of nonsymmetric and indefinite AMG is the incorporation of general relaxation schemes. In the SPD setting, the classical form of optimal AMG interpolation provides a useful insight in determining the best possible two-grid convergence rate of a method based on an arbitrary symmetrized relaxation scheme. In this work, we discuss a generalization of the optimal AMG convergence theory targeting nonsymmetric problems, using a certain matrix-induced orthogonality of the left and right eigenvectors of a generalized eigenvalue problem relating the system matrix and relaxation operator. We show that using this generalization of the optimal convergence theory, one can obtain a measure of the spectral radius of the two grid error transfer operator that is mathematically equivalent to the derivation in the SPD setting for optimal interpolation, which instead uses norms. In addition, this generalization of the optimal AMG convergence theory can be further extended for symmetric indefinite problems, such as those arising from saddle point systems so that one can obtain a precise convergence rate of the resulting two-grid method based on optimal interpolation. We provide supporting numerical examples of the convergence theory for nonsymmetric advection-diffusion problems, two-dimensional Dirac equation motivated by $\gamma_5$-symmetry, and the mixed Darcy flow problem corresponding to a saddle point system.

math.NA

Rapid detection of rare events from in situ X-ray diffraction data using machine learning

High-energy X-ray diffraction methods can non-destructively map the 3D microstructure and associated attributes of metallic polycrystalline engineering materials in their bulk form. These methods are often combined with external stimuli such as thermo-mechanical loading to take snapshots over time of the evolving microstructure and attributes. However, the extreme data volumes and the high costs of traditional data acquisition and reduction approaches pose a barrier to quickly extracting actionable insights and improving the temporal resolution of these snapshots. Here we present a fully automated technique capable of rapidly detecting the onset of plasticity in high-energy X-ray microscopy data. Our technique is computationally faster by at least 50 times than the traditional approaches and works for data sets that are up to 9 times sparser than a full data set. This new technique leverages self-supervised image representation learning and clustering to transform massive data into compact, semantic-rich representations of visually salient characteristics (e.g., peak shapes). These characteristics can be a rapid indicator of anomalous events such as changes in diffraction peak shapes. We anticipate that this technique will provide just-in-time actionable information to drive smarter experiments that effectively deploy multi-modal X-ray diffraction methods that span many decades of length scales.

cs.LG

Effects of impurity band on multiphoton photocurrent from InGaN and GaN photodetectors

Multiphoton absorption of wide band-gap semiconductors has shown great prospects in many fundamental researches and practical applications. With intensity-modulated femtosecond lasers by acousto-optic frequency shifters, photocurrents and yellow luminescence induced by two-photon absorption of InGaN and GaN photodetectors are investigated experimentally. Photocurrent from InGaN detector shows nearly perfect quadratic dependence on excitation intensity, while that in GaN detector shows cubic and higher order dependence. Yellow luminescence from both detectors show sub-quadratic dependence on excitation intensity. Highly nonlinear photocurrent from GaN is ascribed to absorption of additional photons by long-lived electrons in traps and impurity bands. Our investigation indicates that InGaN can serve as a superior detector for multiphoton absorption, absent of linear and higher order process, while GaN, which suffers from absorption by trapped electrons and impurity bands, must be used with caution.

physics.optics

Probing Silicon Carbide with Phase-Modulated Femtosecond Laser Pulses: Insights into Multiphoton Photocurrent

Wide bandgap semiconductors are widely used in photonic technologies due to their advantageous features, such as large optical bandgap, low losses, and fast operational speeds. Silicon carbide is a prototypical wide bandgap semiconductor with high optical nonlinearities, large electron transport, and a high breakdown threshold. Integration of silicon carbide in nonlinear photonics requires a systematic analysis of the multiphoton contribution to the device functionality. Here, multiphoton photocurrent in a silicon carbide photodetector is investigated using phase-modulated femtosecond pulses. Multiphoton absorption is quantified using a 1030 nm phase-modulated pulsed laser. Our measurements show that although the bandgap is less than the energy of three photons, only four-photon absorption has a significant contribution to the photocurrent. We interpret the four-photon absorption as a direct transition from the valance to the conduction band at the Γ point. More importantly, silicon carbide withstands higher excitation intensities compared to other wide bandgap semiconductors making it an ideal system for high-power nonlinear applications.

physics.optics

fairDMS: Rapid Model Training by Data and Model Reuse

Extracting actionable information rapidly from data produced by instruments such as the Linac Coherent Light Source (LCLS-II) and Advanced Photon Source Upgrade (APS-U) is becoming ever more challenging due to high (up to TB/s) data rates. Conventional physics-based information retrieval methods are hard-pressed to detect interesting events fast enough to enable timely focusing on a rare event or correction of an error. Machine learning~(ML) methods that learn cheap surrogate classifiers present a promising alternative, but can fail catastrophically when changes in instrument or sample result in degradation in ML performance. To overcome such difficulties, we present a new data storage and ML model training architecture designed to organize large volumes of data and models so that when model degradation is detected, prior models and/or data can be queried rapidly and a more suitable model retrieved and fine-tuned for new conditions. We show that our approach can achieve up to 100x data labelling speedup compared to the current state-of-the-art, 200x improvement in training speed, and 92x speedup in-terms of end-to-end model updating time.

cs.LG

SMLT: A Serverless Framework for Scalable and Adaptive Machine Learning Design and Training

In today's production machine learning (ML) systems, models are continuously trained, improved, and deployed. ML design and training are becoming a continuous workflow of various tasks that have dynamic resource demands. Serverless computing is an emerging cloud paradigm that provides transparent resource management and scaling for users and has the potential to revolutionize the routine of ML design and training. However, hosting modern ML workflows on existing serverless platforms has non-trivial challenges due to their intrinsic design limitations such as stateless nature, limited communication support across function instances, and limited function execution duration. These limitations result in a lack of an overarching view and adaptation mechanism for training dynamics and an amplification of existing problems in ML workflows. To address the above challenges, we propose SMLT, an automated, scalable, and adaptive serverless framework to enable efficient and user-centric ML design and training. SMLT employs an automated and adaptive scheduling mechanism to dynamically optimize the deployment and resource scaling for ML tasks during training. SMLT further enables user-centric ML workflow execution by supporting user-specified training deadlines and budget limits. In addition, by providing an end-to-end design, SMLT solves the intrinsic problems in serverless platforms such as the communication overhead, limited function execution duration, need for repeated initialization, and also provides explicit fault tolerance for ML training. SMLT is open-sourced and compatible with all major ML frameworks. Our experimental evaluation with large, sophisticated modern ML models demonstrate that SMLT outperforms the state-of-the-art VM based systems and existing serverless ML training frameworks in both training speed (up to 8X) and monetary cost (up to 3X)

cs.DC

Bridging Data Center AI Systems with Edge Computing for Actionable Information Retrieval

Extremely high data rates at modern synchrotron and X-ray free-electron laser light source beamlines motivate the use of machine learning methods for data reduction, feature detection, and other purposes. Regardless of the application, the basic concept is the same: data collected in early stages of an experiment, data from past similar experiments, and/or data simulated for the upcoming experiment are used to train machine learning models that, in effect, learn specific characteristics of those data; these models are then used to process subsequent data more efficiently than would general-purpose models that lack knowledge of the specific dataset or data class. Thus, a key challenge is to be able to train models with sufficient rapidity that they can be deployed and used within useful timescales. We describe here how specialized data center AI (DCAI) systems can be used for this purpose through a geographically distributed workflow. Experiments show that although there are data movement cost and service overhead to use remote DCAI systems for DNN training, the turnaround time is still less than 1/30 of using a locally deploy-able GPU.

cs.LG

Curse or Redemption? How Data Heterogeneity Affects the Robustness of Federated Learning

Data heterogeneity has been identified as one of the key features in federated learning but often overlooked in the lens of robustness to adversarial attacks. This paper focuses on characterizing and understanding its impact on backdooring attacks in federated learning through comprehensive experiments using synthetic and the LEAF benchmarks. The initial impression driven by our experimental results suggests that data heterogeneity is the dominant factor in the effectiveness of attacks and it may be a redemption for defending against backdooring as it makes the attack less efficient, more challenging to design effective attack strategies, and the attack result also becomes less predictable. However, with further investigations, we found data heterogeneity is more of a curse than a redemption as the attack effectiveness can be significantly boosted by simply adjusting the client-side backdooring timing. More importantly,data heterogeneity may result in overfitting at the local training of benign clients, which can be utilized by attackers to disguise themselves and fool skewed-feature based defenses. In addition, effective attack strategies can be made by adjusting attack data distribution. Finally, we discuss the potential directions of defending the curses brought by data heterogeneity. The results and lessons learned from our extensive experiments and analysis offer new insights for designing robust federated learning methods and systems

cs.LG

TiFL: A Tier-based Federated Learning System

Federated Learning (FL) enables learning a shared model across many clients without violating the privacy requirements. One of the key attributes in FL is the heterogeneity that exists in both resource and data due to the differences in computation and communication capacity, as well as the quantity and content of data among different clients. We conduct a case study to show that heterogeneity in resource and data has a significant impact on training time and model accuracy in conventional FL systems. To this end, we propose TiFL, a Tier-based Federated Learning System, which divides clients into tiers based on their training performance and selects clients from the same tier in each training round to mitigate the straggler problem caused by heterogeneity in resource and data quantity. To further tame the heterogeneity caused by non-IID (Independent and Identical Distribution) data and resources, TiFL employs an adaptive tier selection approach to update the tiering on-the-fly based on the observed training performance and accuracy overtime. We prototype TiFL in a FL testbed following Google's FL architecture and evaluate it using popular benchmarks and the state-of-the-art FL benchmark LEAF. Experimental evaluation shows that TiFL outperforms the conventional FL in various heterogeneous conditions. With the proposed adaptive tier selection policy, we demonstrate that TiFL achieves much faster training performance while keeping the same (and in some cases - better) test accuracy across the board.

cs.LG