SearcharxivSearch

arXiv subjects

Olga Kovaleva

Publications and source records attributed to Olga Kovaleva.

6 recordsLinked to original sources

Energy-saving technologies and energy efficiency in the post-pandemic world

This paper explores the role of energy-saving technologies and energy efficiency in the post-COVID era. The pandemic meant major rethinking of the entrenched patterns in energy saving and efficiency. It also provided opportunities for reevaluating energy consumption for households and industries. In addition, it highlighted the importance of employing digital tools and technologies in energy networks and smart grids (e.g. Internet of Energy (IoE), peer-to-peer (P2P) prosumer networks, or the AI-powered autonomous power systems (APS)). In addition, the pandemic added novel legal aspects to the energy efficiency and energy saving and enhanced inter-national collaborations and partnerships. The paper highlights the importance of energy efficiency measures and examines various technologies that can contribute to a sustainable and resilient energy future. Using the bibliometric network analysis of 12960 publications indexed in Web of Science databases, it demonstrates the potential benefits and challenges associated with implementing energy-saving technologies and autonomic power systems in a post-COVID world. Our findings emphasize the need for robust policies, technological advancements, and public engagement to foster energy efficiency and mitigate the environmental impacts of energy consumption.

econ.GN

Down and Across: Introducing Crossword-Solving as a New NLP Benchmark

Solving crossword puzzles requires diverse reasoning capabilities, access to a vast amount of knowledge about language and the world, and the ability to satisfy the constraints imposed by the structure of the puzzle. In this work, we introduce solving crossword puzzles as a new natural language understanding task. We release the specification of a corpus of crossword puzzles collected from the New York Times daily crossword spanning 25 years and comprised of a total of around nine thousand puzzles. These puzzles include a diverse set of clues: historic, factual, word meaning, synonyms/antonyms, fill-in-the-blank, abbreviations, prefixes/suffixes, wordplay, and cross-lingual, as well as clues that depend on the answers to other clues. We separately release the clue-answer pairs from these puzzles as an open-domain question answering dataset containing over half a million unique clue-answer pairs. For the question answering task, our baselines include several sequence-to-sequence and retrieval-based generative models. We also introduce a non-parametric constraint satisfaction baseline for solving the entire crossword puzzle. Finally, we propose an evaluation framework which consists of several complementary performance metrics.

cs.CL

BERT Busters: Outlier Dimensions that Disrupt Transformers

Multiple studies have shown that Transformers are remarkably robust to pruning. Contrary to this received wisdom, we demonstrate that pre-trained Transformer encoders are surprisingly fragile to the removal of a very small number of features in the layer outputs (<0.0001% of model weights). In case of BERT and other pre-trained encoder Transformers, the affected component is the scaling factors and biases in the LayerNorm. The outliers are high-magnitude normalization parameters that emerge early in pre-training and show up consistently in the same dimensional position throughout the model. We show that disabling them significantly degrades both the MLM loss and the downstream task performance. This effect is observed across several BERT-family models and other popular pre-trained Transformer architectures, including BART, XLNet and ELECTRA; we also show a similar effect in GPT-2.

cs.CL

A Primer in BERTology: What we know about how BERT works

Transformer-based models have pushed state of the art in many areas of NLP, but our understanding of what is behind their success is still limited. This paper is the first survey of over 150 studies of the popular BERT model. We review the current state of knowledge about how BERT works, what kind of information it learns and how it is represented, common modifications to its training objectives and architecture, the overparameterization issue and approaches to compression. We then outline directions for future research.

cs.CL

Revealing the Dark Secrets of BERT

BERT-based architectures currently give state-of-the-art performance on many NLP tasks, but little is known about the exact mechanisms that contribute to its success. In the current work, we focus on the interpretation of self-attention, which is one of the fundamental underlying components of BERT. Using a subset of GLUE tasks and a set of handcrafted features-of-interest, we propose the methodology and carry out a qualitative and quantitative analysis of the information encoded by the individual BERT's heads. Our findings suggest that there is a limited set of attention patterns that are repeated across different heads, indicating the overall model overparametrization. While different heads consistently use the same attention patterns, they have varying impact on performance across different tasks. We show that manually disabling attention in certain heads leads to a performance improvement over the regular fine-tuned BERT models.

cs.CL

Direct observation of faulting by means of a rotary shear test under X-ray micro-computed tomography

Friction and fault surface evolution are critical aspects in earthquake studies. We present the preliminary result from a novel experimental approach that combines rotary shear testing with X-ray micro-computed tomography ($μ$CT) technology. An artificial fault was sheared at small incremental rotational steps under the normal stress of 2.5 MPa. During shearing, mechanical data including normal force and torque were measured and used to calculate the friction coefficient. After each rotation increment, a $μ$CT scan was conducted to observe the sample structure. The careful and quantitative $μ$CT image analysis allowed for direct and continuous observation of the fault evolution. We observed that fracturing due to asperity interlocking and breakage dominated the initial phase of slipping. The frictional behavior stabilized after ~1 mm slip distance, which inferred the critical slip distance. We developed a novel approach to estimate the real contact area on the fault surface by means of $μ$CT image analysis. Real contact area varied with increased shear distances as the contacts between asperities changed, and it eventually stabilized at approximately 12% of the nominal fault area. The dimension of the largest contact patch on the surface was close to observed critical slip distance, suggesting that the frictional behavior may be controlled by contacting large asperities. These observations improved our understanding of fault evolution and associated friction variation. Moreover, this work demonstrates that the $μ$CT technology is a powerful tool for the study of earthquake physics.

physics.geo-ph