SearcharxivSearch

arXiv subjects

Shijing Sun

Publications and source records attributed to Shijing Sun.

At least 19 recordsLinked to original sources

Conditional Diffusion Models for Energy-Efficient Driving

Electrification of commercial delivery fleets is shifting fleet routing from distance- and time-based optimization toward energy-aware decision-making. Existing sequence models primarily provide deterministic point estimates or limited uncertainty summaries, which do not capture the range of plausible energy-consumption trajectories required for operational decision-making. In this work, we introduce a conditional diffusion framework that generates EV battery-current profiles conditioned on route features such as vehicle velocity and ambient temperature. The model combines a latent conditioning encoder with a temporal 1D U-Net denoising backbone that enables trip-related conditions to be mapped into a shared representation and guides the reverse diffusion process. We evaluate the framework on an open-access commercial EV telemetry dataset containing 12k trips from 9 vehicles. The proposed latent-conditioned diffusion model generates realistic cur- rent trajectories that capture both the dominant temporal envelope and sharp transient events. The model achieves a Wasserstein distance of 0.0029 between generated and measured current distributions below the real vs real reference distance of 0.0085 indicating that generated samples lie within the empirical variability of the test set. We further demonstrate that learned latent conditioning substantially improves performance over direct condition injection, reducing the Wasserstein distance by 89.1% and MAE by 52.8%. This work demonstrates a generative modeling framework for characterizing EV energy consumption under real-world operating conditions, providing an essential foundation for uncertainty-aware fleet planning in large-scale operational settings.

cs.LG

The Rise of Generative AI for Metal-Organic Framework Design and Synthesis

Advances in generative artificial intelligence are transforming how metal-organic frameworks (MOFs) are designed and discovered. This Perspective introduces the shift from laborious enumeration of MOF candidates to generative approaches that can autonomously propose and synthesize in the laboratory new porous reticular structures on demand. We outline the progress of employing deep learning models, such as variational autoencoders, diffusion models, and large language model-based agents, that are fueled by the growing amount of available data from the MOF community and suggest novel crystalline materials designs. These generative tools can be combined with high-throughput computational screening and even automated experiments to form accelerated, closed-loop discovery pipelines. The result is a new paradigm for reticular chemistry in which AI algorithms more efficiently direct the search for high-performance MOF materials for clean air and energy applications. Finally, we highlight remaining challenges such as synthetic feasibility, dataset diversity, and the need for further integration of domain knowledge.

cond-mat.mtrl-sci

Lessons in Reproducibility: Insights from NLP Studies in Materials Science

Natural Language Processing (NLP), a cornerstone field within artificial intelligence, has been increasingly utilized in the field of materials science literature. Our study conducts a reproducibility analysis of two pioneering works within this domain: "Machine-learned and codified synthesis parameters of oxide materials" by Kim et al., and "Unsupervised word embeddings capture latent knowledge from materials science literature" by Tshitoyan et al. We aim to comprehend these studies from a reproducibility perspective, acknowledging their significant influence on the field of materials informatics, rather than critiquing them. Our study indicates that both papers offered thorough workflows, tidy and well-documented codebases, and clear guidance for model evaluation. This makes it easier to replicate their results successfully and partially reproduce their findings. In doing so, they set commendable standards for future materials science publications to aspire to. However, our analysis also highlights areas for improvement such as to provide access to training data where copyright restrictions permit, more transparency on model architecture and the training process, and specifications of software dependency versions. We also cross-compare the word embedding models between papers, and find that some key differences in reproducibility and cross-compatibility are attributable to design choices outside the bounds of the models themselves. In summary, our study appreciates the benchmark set by these seminal papers while advocating for further enhancements in research reproducibility practices in the field of NLP for materials science. This balance of understanding and continuous improvement will ultimately propel the intersecting domains of NLP and materials science literature into a future of exciting discoveries.

physics.chem-ph

What is missing in autonomous discovery: Open challenges for the community

Self-driving labs (SDLs) leverage combinations of artificial intelligence, automation, and advanced computing to accelerate scientific discovery. The promise of this field has given rise to a rich community of passionate scientists, engineers, and social scientists, as evidenced by the development of the Acceleration Consortium and recent Accelerate Conference. Despite its strengths, this rapidly developing field presents numerous opportunities for growth, challenges to overcome, and potential risks of which to remain aware. This community perspective builds on a discourse instantiated during the first Accelerate Conference, and looks to the future of self-driving labs with a tempered optimism. Incorporating input from academia, government, and industry, we briefly describe the current status of self-driving labs, then turn our attention to barriers, opportunities, and a vision for what is possible. Our field is delivering solutions in technology and infrastructure, artificial intelligence and knowledge generation, and education and workforce development. In the spirit of community, we intend for this work to foster discussion and drive best practices as our field grows.

cond-mat.mtrl-sci

Introducing flexible perovskites to the IoT world using photovoltaic-powered wireless tags

Billions of everyday objects could become part of the Internet of Things (IoT) by augmentation with low-cost, long-range, maintenance-free wireless sensors. Radio Frequency Identification (RFID) is a low-cost wireless technology that could enable this vision, but it is constrained by short communication range and lack of sufficient energy available to power auxiliary electronics and sensors. Here, we explore the use of flexible perovskite photovoltaic cells to provide external power to semi-passive RFID tags to increase range and energy availability for external electronics such as microcontrollers and digital sensors. Perovskites are intriguing materials that hold the possibility to develop high-performance, low-cost, optically tunable (to absorb different light spectra), and flexible light energy harvesters. Our prototype perovskite photovoltaic cells on plastic substrates have an efficiency of 13% and a voltage of 0.88 V at maximum power under standard testing conditions. We built prototypes of RFID sensors powered with these flexible photovoltaic cells to demonstrate real-world applications. Our evaluation of the prototypes suggests that: i) flexible PV cells are durable up to a bending radius of 5 mm with only a 20 % drop in relative efficiency; ii) RFID communication range increased by 5x, and meets the energy needs (10-350 microwatt) to enable self-powered wireless sensors; iii) perovskite powered wireless sensors enable many battery-less sensing applications (e.g., perishable good monitoring, warehouse automation)

eess.SP

Processing Induced Distinct Charge Carrier Dynamics of Bulky Organic Halide Treated Perovskites

State-of-the-art metal halide perovskite-based photovoltaics often employ organic ammonium salts, AX, as a surface passivator, where A is a large organic cation and X is a halide. These surface treatments passivate the perovskite by forming layered perovskites (e.g., A2PbX4) or by AX itself serving as a surface passivation agent on the perovskite photoactive film. It remains unclear whether layered perovskites or AX is the ideal passivator due to an incomplete understanding of the interfacial impact and resulting photoexcited carrier dynamics of AX treatment. In the present study, we use TRPL measurements to selectively probe the different interfaces of glass/perovskite/AX to demonstrate the vastly distinct interfacial photoexcited state dynamics with the presence of A2PbX4 or AX. Coupling the TRPL results with X-ray diffraction and nanoscale microscopy measurements, we find that the presence of AX not only passivates the traps at the surface and the grain boundaries, but also induces an α/δ-FAPbI3 phase mixing that alters the carrier dynamics near the glass/perovskite interface and enhances the photoluminescence quantum yield. In contrast, the passivation with A2PbI4 is mostly localized to the surface and grain boundaries near the top surface where the availability of PbI2 directly determines the formation of A2PbI4. Such distinct mechanisms significantly impact the corresponding solar cell performance, and we find AX passivation that has not been converted to a layered perovskite allows for a much larger processing window (e.g., larger allowed variance of AX concentration which is critical for improving the eventual manufacturing yield) and more reproducible condition to realize device performance improvements, while A2PbI4 as a passivator yields a much narrower processing window. We expect these results to enable a more rational route for developing AX for perovskite.

cond-mat.mtrl-sci

An invertible crystallographic representation for general inverse design of inorganic crystals with targeted properties

Realizing general inverse design could greatly accelerate the discovery of new materials with user-defined properties. However, state-of-the-art generative models tend to be limited to a specific composition or crystal structure. Herein, we present a framework capable of general inverse design (not limited to a given set of elements or crystal structures), featuring a generalized invertible representation that encodes crystals in both real and reciprocal space, and a property-structured latent space from a variational autoencoder (VAE). In three design cases, the framework generates 142 new crystals with user-defined formation energies, bandgap, thermoelectric (TE) power factor, and combinations thereof. These generated crystals, absent in the training database, are validated by first-principles calculations. The success rates (number of first-principles-validated target-satisfying crystals/number of designed crystals) ranges between 7.1% and 38.9%. These results represent a significant step toward property-driven general inverse design using generative models, although practical challenges remain when coupled with experimental synthesis.

physics.comp-ph

Predicting antimicrobial activity of conjugated oligoelectrolyte molecules via machine learning

New antibiotics are needed to battle growing antibiotic resistance, but the development process from hit, to lead, and ultimately to a useful drug, takes decades. Although progress in molecular property prediction using machine-learning methods has opened up new pathways for aiding the antibiotics development process, many existing solutions rely on large datasets and finding structural similarities to existing antibiotics. Challenges remain in modelling of unconventional antibiotics classes that are drawing increasing research attention. In response, we developed an antimicrobial activity prediction model for conjugated oligoelectrolyte molecules, a new class of antibiotics that lacks extensive prior structure-activity relationship studies. Our approach enables us to predict minimum inhibitory concentration for E. coli K12, with 21 molecular descriptors selected by recursive elimination from a set of 5,305 descriptors. This predictive model achieves an R2 of 0.65 with no prior knowledge of the underlying mechanism. We find the molecular representation optimum for the domain is the key to good predictions of antimicrobial activity. In the case of conjugated oligoelectrolytes, a representation reflecting the 3-dimensional shape of the molecules is most critical. Although it is demonstrated with a specific example of conjugated oligoelectrolytes, our proposed approach for creating the predictive model can be readily adapted to other novel antibiotic candidate domains.

physics.app-ph

Opportunities for Machine Learning to Accelerate Halide Perovskite Commercialization and Scale-Up

While halide perovskites attract significant academic attention, examples of at-scale industrial production are still sparse. In this perspective, we review practical challenges hindering the commercialization of halide perovskites, and discuss how machine-learning (ML) tools could help: (1) active-learning algorithms that blend institutional knowledge and human expertise could help stabilize and rapidly update baseline manufacturing processes; (2) ML-powered metrology, including computer imaging, could help narrow the performance gap between large- and small-area devices; and (3) inference methods could help accelerate root-cause analysis by reconciling multiple data streams and simulations, focusing research effort on areas with highest probability for improvement. We conclude that to satisfy many of these challenges, incremental -- not radical -- adaptations of existing ML and statistical methods are needed. We identify resources to help develop in-house data-science talent, and propose how industry-academic partnerships could help adapt "ready-now" ML tools to specific industry needs, further improve process control by revealing underlying mechanisms, and develop "gamechanger" discovery-oriented algorithms to better navigate vast materials combination spaces and the literature.

cond-mat.mtrl-sci

Discovering Equations that Govern Experimental Materials Stability under Environmental Stress using Scientific Machine Learning

While machine learning (ML) in experimental research has demonstrated impressive predictive capabilities, inductive reasoning and knowledge extraction remain elusive tasks, in part because of the difficulty extracting fungible knowledge representations from experimental data. In this manuscript, we use ML to infer the underlying dynamical differential equation (DE) from experimental data of degrading organic-inorganic methylammonium lead iodide (MAPI) perovskite thin films under environmental stressors (elevated temperature, humidity, and light). We apply a sparse regression algorithm that automatically identifies the differential equation describing the dynamics from time-series data. We find that the underlying DE governing MAPI degradation across a broad temperature range of 35 to 85°C is described minimally with three terms (specifically, a second-order polynomial), and not a simple single-order reaction (i.e. 0th, 1st, or 2nd-order reaction). We demonstrate how computer-derived results can aid the researcher to develop profound mechanistic insights. This DE corresponds to the Verhulst logistic function, which describes reaction kinetics analogous in functional form to autocatalytic or self-propagating reactions, suggesting future strategies to suppress MAPI degradation. We examine the robustness of our conclusions to experimental luck-of-the-draw variance and Gaussian noise using a combination of experiment and simulation, and describe the experimental limits within which this methodology can be applied. Our study demonstrates the application of scientific ML in experimental chemical and materials systems, highlighting the promise and challenges associated with ML-aided scientific discovery.

cond-mat.mtrl-sci

Benchmarking the Performance of Bayesian Optimization across Multiple Experimental Materials Science Domains

In the field of machine learning (ML) for materials optimization, active learning algorithms, such as Bayesian Optimization (BO), have been leveraged for guiding autonomous and high-throughput experimentation systems. However, very few studies have evaluated the efficiency of BO as a general optimization algorithm across a broad range of experimental materials science domains. In this work, we evaluate the performance of BO algorithms with a collection of surrogate model and acquisition function pairs across five diverse experimental materials systems, namely carbon nanotube polymer blends, silver nanoparticles, lead-halide perovskites, as well as additively manufactured polymer structures and shapes. By defining acceleration and enhancement metrics for general materials optimization objectives, we find that for surrogate model selection, Gaussian Process (GP) with anisotropic kernels (automatic relevance detection, ARD) and Random Forests (RF) have comparable performance and both outperform the commonly used GP without ARD. We discuss the implicit distributional assumptions of RF and GP, and the benefits of using GP with anisotropic kernels in detail. We provide practical insights for experimentalists on surrogate model selection of BO during materials optimization campaigns.

cond-mat.mtrl-sci

Bridging the gap between photovoltaics R&D and manufacturing with data-driven optimization

Novel photovoltaics, such as perovskites and perovskite-inspired materials, have shown great promise due to high efficiency and potentially low manufacturing cost. So far, solar cell R&D has mostly focused on achieving record efficiencies, a process that often results in small batches, large variance, and limited understanding of the physical causes of underperformance. This approach is intensive in time and resources, and ignores many relevant factors for industrial production, particularly the need for high reproducibility and high manufacturing yield, and the accompanying need of physical insights. The record-efficiency paradigm is effective in early-stage R&D, but becomes unsuitable for industrial translation, requiring a repetition of the optimization procedure in the industrial setting. This mismatch between optimization objectives, combined with the complexity of physical root-cause analysis, contributes to decade-long timelines to transfer new technologies into the market. Based on recent machine learning and technoeconomic advances, our perspective articulates a data-driven optimization framework to bridge R&D and manufacturing optimization approaches. We extend the maximum-efficiency optimization paradigm by considering two additional dimensions: a technoeconomic figure of merit and scalable physical inference. Our framework naturally aligns different stages of technology development with shared optimization objectives, and accelerates the optimization process by providing physical insights.

physics.app-ph

Embedding Physics Domain Knowledge into a Bayesian Network Enables Layer-by-Layer Process Innovation for Photovoltaics

Process optimization of photovoltaic devices is a time-intensive, trial and error endeavor, without full transparency of the underlying physics, and with user-imposed constraints that may or may not lead to a global optimum. Herein, we demonstrate that embedding physics domain knowledge into a Bayesian network enables an optimization approach that identifies the root cause(s) of underperformance with layer by-layer resolution and reveals alternative optimal process windows beyond global black-box optimization. Our Bayesian-network approach links process conditions to materials descriptors (bulk and interface properties, e.g., bulk lifetime, doping, and surface recombination) and device performance parameters (e.g., cell efficiency), using a Bayesian inference framework with an autoencoder-based surrogate device-physics model that is 100x faster than numerical solvers. With the trained surrogate model, our approach is robust and reduces significantly the time consuming experimentalist intervention, even with small numbers of fabricated samples. To demonstrate our method, we perform layer-by-layer optimization of GaAs solar cells. In a single cycle of learning, we find an improved growth temperature for the GaAs solar cells without any secondary measurements, and demonstrate a 6.5% relative AM1.5G efficiency improvement above baseline and traditional black-box optimization methods.

physics.app-ph

Perovskite PV-powered RFID: enabling low-cost self-powered IoT sensors

Photovoltaic (PV) cells have the potential to serve as on-board power sources for low-power IoT devices. Here, we explore the use of perovskite solar cells to power Radio Frequency (RF) backscatter-based IoT devices with a few μW power demand. Perovskites are suitable for low-cost, high-performance, low-temperature processing, and flexible light energy harvesting that hold the possibility to significantly extend the range and lifetime of current backscatter techniques such as Radio Frequency Identification (RFID). For these reasons, perovskite solar cells are prominent candidates for future low-power wireless applications. We report on realizing a functional perovskite-powered wireless temperature sensor with 4 m communication range. We use a 10.1% efficient perovskite PV module generating an output voltage of 4.3 V with an active area of 1.06 cm2 under 1 sun illumination, with AM 1.5G spectrum, to power a commercial off-the-shelf RFID IC, requiring 10 - 45 μW of power. Having an on-board energy harvester provides extra-energy to boost the range of the sensor (5x) in addition to providing energy to carry out high-volume sensor measurements (hundreds of measurements per min). Our evaluation of the prototype suggests that perovskite photovoltaic cells are able to meet the energy needs to enable fully autonomous low-power RF backscatter applications of the future. We conclude with an outlook into a range of applications that we envision to leverage the synergies offered by combining perovskite photovoltaics and RFID.

eess.SP

Self-powered sensors enabled by wide-bandgap perovskite indoor photovoltaic cells

We present a new approach to ubiquitous sensing for indoor applications, using high-efficiency and low-cost indoor perovksite photovoltaic cells as external power sources for backscatter sensors. We demonstrate wide-bandgap perovskite photovoltaic cells for indoor light energy harvesting with the 1.63eV and 1.84 eV devices demonstrate efficiencies of 21% and 18.5% respectively under indoor compact fluorescent lighting, with a champion open-circuit voltage of 0.95 V in a 1.84 eV cell under a light intensity of 0.16 mW/cm2. Subsequently, we demonstrate a wireless temperature sensor self-powered by a perovskite indoor light-harvesting module. We connect three perovskite photovoltaic cells in series to create a module that produces 14.5 uW output power under 0.16 mW/cm2 of compact fluorescent illumination with an efficiency of 13.2%. We use this module as an external power source for a battery-assisted RFID temperature sensor and demonstrate a read range by of 5.1 meters while maintaining very high frequency measurements every 1.24 seconds. Our combined indoor perovskite photovoltaic modules and backscatter radio-frequency sensors are further discussed as a route to ubiquitous sensing in buildings given their potential to be manufactured in an integrated manner at very low-cost, their lack of a need for battery replacement and the high frequency data collection possible.

physics.app-ph

Fast and interpretable classification of small X-ray diffraction datasets using data augmentation and deep neural networks

X-ray diffraction (XRD) data acquisition and analysis is among the most time-consuming steps in the development cycle of novel thin-film materials. We propose a machine-learning-enabled approach to predict crystallographic dimensionality and space group from a limited number of thin-film XRD patterns. We overcome the scarce-data problem intrinsic to novel materials development by coupling a supervised machine learning approach with a model agnostic, physics-informed data augmentation strategy using simulated data from the Inorganic Crystal Structure Database (ICSD) and experimental data. As a test case, 115 thin-film metal halides spanning 3 dimensionalities and 7 space-groups are synthesized and classified. After testing various algorithms, we develop and implement an all convolutional neural network, with cross validated accuracies for dimensionality and space-group classification of 93% and 89%, respectively. We propose average class activation maps, computed from a global average pooling layer, to allow high model interpretability by human experimentalists, elucidating the root causes of misclassification. Finally, we systematically evaluate the maximum XRD pattern step size (data acquisition rate) before loss of predictive accuracy occurs, and determine it to be 0.16°, which enables an XRD pattern to be obtained and classified in 5.5 minutes or less.

physics.data-an

Accelerating Photovoltaic Materials Development via High-Throughput Experiments and Machine-Learning-Assisted Diagnosis

Accelerating the experimental cycle for new materials development is vital for addressing the grand energy challenges of the 21st century. We fabricate and characterize 75 unique halide perovskite-inspired solution-based thin-film materials within a two-month period, with 87% exhibiting band gaps between 1.2 eV and 2.4 eV that are of interest for energy-harvesting applications. This increased throughput is enabled by streamlining experimental workflows, developing a set of precursors amenable to high-throughput synthesis, and developing machine-learning assisted diagnosis. We utilize a deep neural network to classify compounds based on experimental X-ray diffraction data into 0D, 2D, and 3D structures more than 10 times faster than human analysis and with 90% accuracy. We validate our methods using lead-halide perovskites and extend the application to novel lead-free compositions. The wider synthesis window and faster cycle of learning enables three noteworthy scientific findings: (1) we realize four inorganic layered perovskites, A3B2Br9 (A = Cs, Rb; B = Bi, Sb) in thin-film form via one-step liquid deposition; (2) we report a multi-site lead-free alloy series that was not previously described in literature, Cs3(Bi1-xSbx)2(I1-xBrx)9; and (3) we reveal the effect on bandgap (reduction to <2 eV) and structure upon simultaneous alloying on the B-site and X-site of Cs3Bi2I9 with Sb and Br. This study demonstrates that combining an accelerated experimental cycle of learning and machine-learning based diagnosis represents an important step toward realizing fully-automated laboratories for materials discovery and development.

physics.app-ph

Homogenization of Halide Distribution and Carrier Dynamics in Alloyed Organic-Inorganic Perovskites

Perovskite solar cells have shown remarkable efficiencies beyond 22%, through organic and inorganic cation alloying. However, the role of alkali-metal cations is not well-understood. By using synchrotron-based nano-X-ray fluorescence and complementary measurements, we show that when adding RbI and/or CsI the halide distribution becomes homogenous. This homogenization translates into long-lived charge carrier decays, spatially homogenous carrier dynamics visualized by ultrafast microscopy, as well as improved photovoltaic device performance. We find that Rb and K phase-segregate in highly concentrated aggregates. Synchrotron-based X-ray-beam-induced current and electron-beam-induced current of solar cells show that Rb clusters do not contribute to the current and are recombination active. Our findings bring light to the beneficial effects of alkali metal halides in perovskites, and point at areas of weakness in the elemental composition of these complex perovskites, paving the way to improved performance in this rapidly growing family of materials for solar cell applications.

cond-mat.mtrl-sci