SearcharxivSearch

arXiv subjects

Yukio Ohsawa

Publications and source records attributed to Yukio Ohsawa.

At least 19 recordsLinked to original sources

Extracting and Validating Explanatory Word Archipelagoes using Dual Entropy

The logical connectivity of text is represented by the connectivity of words that form archipelagoes. Here, each archipelago is a sequence of islands of the occurrences of a certain word. An island here means the local sequence of sentences where the word is emphasized, and an archipelago of a length comparable to the target text is extracted using the co-variation of entropy A (the window-based entropy) on the distribution of the word's occurrences with the width of each time window. Then, the logical connectivity of text is evaluated on entropy B (the graph-based entropy) computed on the distribution of sentences to connected word-clusters obtained on the co-occurrence of words. The results show the parts of the target text with words forming archipelagoes extracted on entropy A, without learned or prepared knowledge, form an explanatory part of the text that is of smaller entropy B than the parts extracted by the baseline methods.

cs.CL

Analyzing Polysemy Evolution Using Semantic Cells

The senses of words evolve. The sense of the same word may change from today to tomorrow, and multiple senses of the same word may be the result of the evolution of each other, that is, they may be parents and children. If we view Juba as an evolving ecosystem, the paradigm of learning the correct answer, which does not move with the sense of a word, is no longer valid. This paper is a case study that shows that word polysemy is an evolutionary consequence of the modification of Semantic Cells, which has al-ready been presented by the author, by introducing a small amount of diversity in its initial state as an example of analyzing the current set of short sentences. In particular, the analysis of a sentence sequence of 1000 sentences in some order for each of the four senses of the word Spring, collected using Chat GPT, shows that the word acquires the most polysemy monotonically in the analysis when the senses are arranged in the order in which they have evolved. In other words, we present a method for analyzing the dynamism of a word's acquiring polysemy with evolution and, at the same time, a methodology for viewing polysemy from an evolutionary framework rather than a learning-based one.

cs.CL

Semantic Cells: Evolutional Process to Acquire Sense Diversity of Items

Previous models for learning the semantic vectors of items and their groups, such as words, sentences, nodes, and graphs, using distributed representation have been based on the assumption that the basic sense of an item corresponds to one vector composed of dimensions corresponding to hidden contexts in the target real world, from which multiple senses of the item are obtained by conforming to lexical databases or adapting to the context. However, there may be multiple senses of an item, which are hardly assimilated and change or evolve dynamically following the contextual shift even within a document or a restricted period. This is a process similar to the evolution or adaptation of a living entity with/to environmental shifts. Setting the scope of disambiguation of items for sensemaking, the author presents a method in which a word or item in the data embraces multiple semantic vectors that evolve via interaction with others, similar to a cell embracing chromosomes crossing over with each other. We obtained two preliminary results: (1) the role of a word that evolves to acquire the largest or lower-middle variance of semantic vectors tends to be explainable by the author of the text; (2) the epicenters of earthquakes that acquire larger variance via crossover, corresponding to the interaction with diverse areas of land crust, are likely to correspond to the epicenters of forthcoming large earthquakes.

cs.LG

Collect and Connect Data Leaves to Feature Concepts: Interactive Graph Generation Toward Well-being

Feature concepts and data leaves have been invented using datasets to foster creative thoughts for creating well-being in daily life. The idea, simply put, is to attach selected and collected data leaves that are summaries of event flows to be discovered from corresponding datasets, on the target feature concept representing the well-being aimed. A graph of existing or expected datasets to be attached to a feature concept is generated semi-automatically. Rather than sheer automated generative AI, our work addresses the process of generative artificial and natural intelligence to create the basis for data use and reuse.

cs.LG

Generating a Map of Well-being Regions using Multiscale Moving Direction Entropy on Mobile Sensors

The well-being of individuals in a crowd is interpreted as a product of the crossover of individuals from heterogeneous communities, which may occur via interactions with other crowds. The index moving-direction entropy corresponding to the diversity of the moving directions of individuals is introduced to represent such an inter-community crossover. Multi-scale moving direction entropies, composed of various geographical mesh sizes to compute the index values, are used to capture the information flow owing to human movements from/to various crowds. The generated map of high values of multiscale moving direction entropy is shown to coincide significantly with the preference of people to live in each region.

econ.GN

Data Leaves: Scenario-oriented Metadata for Data Federative Innovation

A method for representing the digest information of each dataset is proposed, oriented to the aid of innovative thoughts and the communication of data users who attempt to create valuable products, services, and business models using or combining datasets. Compared with methods for connecting datasets via shared attributes (i.e., variables), this method connects datasets via events, situations, or actions in a scenario that is supposed to be active in the real world. This method reflects the consideration of the fitness of each metadata to the feature concept, which is an abstract of the information or knowledge expected to be acquired from data; thus, the users of the data acquire practical knowledge that fits the requirements of real businesses and real life, as well as grounds for realistic application of AI technologies to data.

cs.DB

Improving Blockchain scalability based on one-time cross-chain contract and gossip network

This study proposes a novel solution that provides secure interoperability for blockchains, which improves the overall scalability of the whole blockchain network. In our solution, a cross-chain task will build a one-time cross-blockchain contract. Each blockchain system can follow the contract to complete or this task. The result of tasks is bound with the system, hence can be anchored to all other blockchain systems through the gossip network. This work shows our result can provide linear scalability for the whole system and achieve consistency among honest systems.

cs.NI

Feature Concepts for Data Federative Innovations

A feature concept, the essence of the data-federative innovation process, is presented as a model of the concept to be acquired from data. A feature concept may be a simple feature, such as a single variable, but is more likely to be a conceptual illustration of the abstract information to be obtained from the data. For example, trees and clusters are feature concepts for decision tree learning and clustering, respectively. Useful feature concepts for satis-fying the requirements of users of data have been elicited so far via creative communication among stakeholders in the market of data. In this short paper, such a creative communication is reviewed, showing a couple of appli-cations, for example, change explanation in markets and earthquakes, and highlight the feature concepts elicited in these cases.

cs.LG

Effects of Interregional Travels and Vaccination in Infection Spreads Simulated by Lattice of SEIRS Circuits

The SEIRS model, an extension of the SEIR model for analyzing and predicting the spread of virus infection, was further extended to consider the movement of people across regions. In contrast to previous models that con-sider the risk of travelers from/to other regions, we consider two factors. First, we consider the movements of susceptible (S), exposed (E), and recovered (R) individuals who may get infected and infect others in the destination region, as well as infected (I) individuals. Second, people living in a region and moving from other regions are dealt as separate but interacting groups with respect to their states, S, E, R, or I. This enables us to consider the potential influence of movements before individuals become infected, difficult to detect by testing at the time of immigration, on the spread of infection. In this paper, we show the results of the simulation where individuals travel across regions, which means prefectures here, and the government chooses regions to vaccinate with priority. We found a general law that a quantity of vaccines can be used efficiently by maximizing an index value, the conditional entropy Hc, when we distribute vaccines to regions. The efficiency of this strategy, which maximizes Hc, was found to outperform that of vaccinating regions with a larger effective re-generation number. This law also explains the surprising result that travel activities across regional borders may suppress the spread if vaccination is processed at a sufficiently high pace, introducing the concept of social muddling.

physics.med-ph

Hierarchical entropy and domain interaction to understand the structure in an image

In this study, we devise a model that introduces two hierarchies into information entropy. The two hierarchies are the size of the region for which entropy is calculated and the size of the component that determines whether the structures in the image are integrated or not. And this model uses two indicators, hierarchical entropy and domain interaction. Both indicators increase or decrease due to the integration or fragmentation of the structure in the image. It aims to help people interpret and explain what the structure in an image looks like from two indicators that change with the size of the region and the component. First, we conduct experiments using images and qualitatively evaluate how the two indicators change. Next, we explain the relationship with the hidden structure of Vermeer's girl with a pearl earring using the change of hierarchical entropy. Finally, we clarify the relationship between the change of domain interaction and the appropriate segment result of the image by an experiment using a questionnaire.

cs.CV

Data Combination for Problem-solving: A Case of an Open Data Exchange Platform

In recent years, rather than enclosing data within a single organization, exchanging and combining data from different domains has become an emerging practice. Many studies have discussed the economic and utility value of data and data exchange, but the characteristics of data that contribute to problem solving through data combination have not been fully understood. In big data and interdisciplinary data combinations, large-scale data with many variables are expected to be used, and value is expected to be created by combining data as much as possible. In this study, we conduct three experiments to investigate the characteristics of data, focusing on the relationships between data combinations and variables in each dataset, using empirical data shared by the local government. The results indicate that even datasets that have a few variables are frequently used to propose solutions for problem solving. Moreover, we found that even if the datasets in the solution do not have common variables, there are some well-established solutions to the problems. The findings of this study shed light on mechanisms behind data combination for problem-solving involving multiple datasets and variables.

cs.CY

Detecting and explaining changes in various assets' relationships in financial markets

We study the method for detecting relationship changes in financial markets and providing human-interpretable network visualization to support the decision-making of fund managers dealing with multi-assets. First, we construct co-occurrence networks with each asset as a node and a pair with a strong relationship in price change as an edge at each time step. Second, we calculate Graph-Based Entropy to represent the variety of price changes based on the network. Third, we apply the Differential Network to finance, which is traditionally used in the field of bioinformatics. By the method described above, we can visualize when and what kind of changes are occurring in the financial market, and which assets play a central role in changes in financial markets. Experiments with multi-asset time-series data showed results that were well fit with actual events while maintaining high interpretability. It is suggested that this approach is useful for fund managers to use as a new option for decision making.

q-fin.GN

Data Requests and Scenarios for Data Design of Unobserved Events in Corona-related Confusion Using TEEDA

Due to the global violence of the novel coronavirus, various industries have been affected and the breakdown between systems has been apparent. To understand and overcome the phenomenon related to this unprecedented crisis caused by the coronavirus infectious disease (COVID-19), the importance of data exchange and sharing across fields has gained social attention. In this study, we use the interactive platform called treasuring every encounter of data affairs (TEEDA) to externalize data requests from data users, which is a tool to exchange not only the information on data that can be provided but also the call for data, what data users want and for what purpose. Further, we analyze the characteristics of missing data in the corona-related confusion stemming from both the data requests and the providable data obtained in the workshop. We also create three scenarios for the data design of unobserved events focusing on variables.

cs.HC

Stay with Your Community: Bridges between Clusters Trigger Expansion of COVID-19

The spreading of virus infection is here simulated over artificial human networks. The real-space urban life of people is modeled as a modified scale-free network with constraints. A scale-free network has been adopted in several studies for modeling on-line communities so far but is modified here for the aim to represent peoples' social behaviors where the generated communities are restricted reflecting the spatiotemporal constraints in the real life. Furthermore, the networks have been extended by introducing multiple cliques in the initial step of network construction and enabling people to zero-degree people as well as popular (large degree) people. As a result, four findings and a policy proposal have been obtained. First, the "second waves" occur without external influence or constraints on contacts or the releasing of the constraints. These second waves, mostly lower than the first wave, implies the bridges between infected and fresh clusters may trigger new expansions of spreading. Second, if the network changes the structure on the way of infection spreading or after its suppression, the peak of the second wave can be larger than the first. Third, the peak height in the time series depends on the difference between the upper bound of the number of people each member accepts to meet and the number of people one chooses to meet. This tendency is observed for two kinds of artificial networks and implies the impact of the bridges between communities on the virus spreading. Fourth, the release of once given constraint may trigger a second wave higher than the peak of the time series without introducing any constraint from the beginning, if the release is introduced at a time close to the peak. Thus, both governments and individuals should be careful in returning to human society with inter-community contacts.

cs.SI

Modeling Stakeholder-centric Value Chain of Data to Understand Data Exchange Ecosystem

In recent years, the expectation that new businesses and economic value can be created by combining/exchanging data from different fields has risen. However, value creation by data exchange involves not only data, but also technologies and a variety of stakeholders that are integrated and in competition with one another. This makes the data exchange ecosystem a challenging subject to study. In this paper, we propose a model describing the stakeholder-centric value chain (SVC) of data by focusing on the relationships among stakeholders in data businesses and discussing creative ways to use them. The SVC model enables the analysis and understanding of the structural characteristics of the data exchange ecosystem. We identified stakeholders who carry potential risk, those who play central roles in the ecosystem, and the distribution of profit among them using business models collected by the SVC.

cs.CY

COVID-19 Should be Suppressed by Mixed Constraints -- from Simulations on Constrained Scale-Free Networks

The spreading of virus infection is here simulated over artificial human networks. Here, the real-space urban life of people is modeled as a scale-free network with constraints. A scale-free network has been adopted for modeling on-line communities so far but is employed here for the aim to represent peoples' social behaviors where the generated communities are restricted reflecting the spatiotemporal constraints in the real life. As a result, three findings and a policy proposal have been obtained. First, the height of the peaks in the time sequence of the number of infection cases tends to get reduced corresponding to the upper bound of the size of groups where all members meet. Second, if we adopt the constraint on m0, the number of all other people one meets separately each at a time, to the range between 2 and 8, its effect on the suppression of infections may be weak as far as we allow group meetings of size W of 8 or larger. Third, such a moderate constraint may temporarily seem to work for the reduction of infections in the early stage but it may turn out to be just a delay of peaks. Based on these results, a policy is proposed here: for quickly suppressing the number of infections, restrict W to less than 4 if the constraint to make m0 at most 1 is too strict. If W is set to less than 4, setting m0 to 4 or less works for quick reduction of infections according to the result.

physics.soc-ph

Variable-Based Network Analysis of Datasets on Data Exchange Platforms

Recently, data exchange platforms have emerged in the digital economy to enable better resource allocation in a data-driven society, which requires cross-organizational data collaborations. Understanding the characteristics of the data on these platforms is important for their application; however, the structures of such platforms have not been extensively investigated. In this study, we apply a network approach with a novel variable-based structural analysis to the metadata of datasets on two data platform services. It was noted that the structures of the data networks are locally dense and highly assortative, similar to human-related net-works. Even though the data on these platforms are designed and collected differently, depending on the use objectives, the variables of heterogeneous data exhibit a power distribution, and the data networks exhibit multi-scaling behavior. Furthermore, we found that the data collection strategies of the platforms are related to the variety of variables, density of the networks, and their robustness from the viewpoint of sustainability and social acceptability of the data platforms.

cs.SI

Tangled String for Multi-Scale Explanation of Contextual Shifts in Stock Market

The original research question here is given by marketers in general, i.e., how to explain the changes in the desired timescale of the market. Tangled String, a sequence visualization tool based on the metaphor where contexts in a sequence are compared to tangled pills in a string, is here extended and diverted to detecting stocks that trigger changes in the market and to explaining the scenario of contextual shifts in the market. Here, the sequential data on the stocks of top 10 weekly increase rates in the First Section of the Tokyo Stock Exchange for 12 years are visualized by Tangled String. The changing in the prices of stocks is a mixture of various timescales and can be explained in the time-scale set as desired by using TS. Also, it is found that the change points found by TS coincided by high precision with the real changes in each stock price. As TS has been created from the data-driven innovation platform called Innovators Marketplace on Data Jackets and is extended to satisfy data users, this paper is as evidence of the contribution of the market of data to data-driven innovations.

cs.CE