SearcharxivSearch

arXiv subjects

Sneha Jha

Publications and source records attributed to Sneha Jha.

4 recordsLinked to original sources

LGDE: Local Graph-based Dictionary Expansion

We present Local Graph-based Dictionary Expansion (LGDE), a method for data-driven discovery of the semantic neighbourhood of words using tools from manifold learning and network science. At the heart of LGDE lies the creation of a word similarity graph from the geometry of word embeddings followed by local community detection based on graph diffusion. The diffusion in the local graph manifold allows the exploration of the complex nonlinear geometry of word embeddings to capture word similarities based on paths of semantic association, over and above direct pairwise similarities. Exploiting such semantic neighbourhoods enables the expansion of dictionaries of pre-selected keywords, an important step for tasks in information retrieval, such as database queries and online data collection. We validate LGDE on two user-generated English-language corpora and show that LGDE enriches the list of keywords with improved performance relative to methods based on direct word similarities or co-occurrences. We further demonstrate our method through a real-world use case from communication science, where LGDE is evaluated quantitatively on the expansion of a conspiracy-related dictionary from online data collected and analysed by domain experts. Our empirical results and expert user assessment indicate that LGDE expands the seed dictionary with more useful keywords due to the manifold-learning-based similarity network.

cs.CL

A Web-Based Application Leveraging Geospatial Information to Automate On-Farm Trial Design

On-farm sensor data have allowed farmers to implement field management techniques and intensively track the corresponding responses. These data combined with historical records open the door for real-time field management improvements with the help of current advancements in computing power. However, despite these advances, the statistical design of experiments is rarely used to evaluate the performance of field management techniques accurately. Traditionally, randomized block design is prevalent in statistical designs of field trials, but in practice it is limited in dealing with large variations in soil classes, management practices, and crop varieties. More specifically, although this experimental design is suited for most trial types, it is not the optimal choice when multiple factors are tested over multifarious natural variations in farms, due to the economic constraints caused by the sheer number of variables involved. Experimental refinement is required to better estimate the effects of the primary factor in the presence of auxiliary factors. In this way, farmers can better understand the characteristics and limitations of the primary factor. This work presents a framework for automating the analysis of local field variations by fusing soil classification data and lidar topography data with historical yield. This framework will be leveraged to automate the designing of field experiments based on multiple topographic features

stat.ME

Analyzing trends for agricultural decision support system using twitter data

The trends and reactions of the general public towards global events can be analyzed using data from social platforms, including Twitter. The number of tweets has been reported to help detect variations in communication traffic within subsets like countries, age groups and industries. Similarly, publicly accessible data and (in particular) data from social media about agricultural issues provide a great opportunity for obtaining instantaneous snapshots of farmer opinions and a method to track changes in opinion through temporal analysis. In this paper we hypothesize that the presence of keywords like precision agriculture, digital agriculture, Internet of Things (IoT), BigData, remote sensing, GPS, etc., in tweets could serve as an indicator of discussions centered around interest in modern farming practices. We extracted relevant tweets using keywords such as IoT, BigData and Geographical Information System (GIS), and then analyzed their geographical origin and frequency of their mention. We analyzed the Twitter data for the period of 1st -11th January 2018 to understand these trends and the factors affecting them. These factors, such as special events, projects, biogeography, etc., were further analyzed using tweet sources and trending hashtags from the database. The regions with the highest interest in the keywords were United States, Egypt, Brazil, Japan and China. A comparison of frequency of keywords revealed IoT as the most tweeted word (77.6%) in the downloaded data. The most used language was English followed by Spanish, Japanese and French. Periodical tweets on IoT from an account handled by IoT project on Twitter and Seminars on IoT in January in Santa Catarina (Brazil) were found to be the underlying factors for the observed trends.

stat.AP

Data-Driven Web-Based Patching Management Tool Using Multi-Sensor Pavement Structure Measurements

Automating pavement maintenance suggestions is challenging,especially for actionable recommendations such as patching location,depth and priority.It is common practice among State agencies to manually inspect road segments of interest and decide maintenance requirements based on the pavement condition index (PCI).However,standalone PCI only evaluates the pavement surface condition and coupled with the variability in human perception of pavement distress,limits the accuracy and quality of current pavement maintenance practices.Here,a need for multi-sensor data integrated with standardized pavement distress condition ratings is required.This study explores the possibility of estimating the appropriate pavement patching strategy (i.e.,patching location,depth,and quantity) by integrating pavement structural and surface condition assessment with pavement specific ratings of distress.Especially,it combines pavement structural condition assessment parameter;falling weight deflectometer deflections along with surface condition assessment parameters;international roughness index,and cracking density for a better representation of overall pavement distress condition.Then,a pavement specific threshold-based patching suggestion algorithm is implemented to evaluate the pavement overall distress condition into a priority-based patching suggestion.The novelty in the use of pavement specific thresholds is placed on its data-driven ability to determine threshold values from current road condition measurements using a reliability concept validated by the theoretical pavement condition rating,pavement structural number.A web-based patching manager tool (PMT) was implemented to automate the patching suggestion procedure and visualize the results.Validated with road surface images obtained from three-dimensional laser sensors,PMT could successfully capture localized distresses in existing pavements.

eess.SP