SearcharxivSearch

arXiv subjects

Vasco Furtado

Publications and source records attributed to Vasco Furtado.

18 recordsLinked to original sources

CDJUR-BR -- A Golden Collection of Legal Document from Brazilian Justice with Fine-Grained Named Entities

A basic task for most Legal Artificial Intelligence (Legal AI) applications is Named Entity Recognition (NER). However, texts produced in the context of legal practice make references to entities that are not trivially recognized by the currently available NERs. There is a lack of categorization of legislation, jurisprudence, evidence, penalties, the roles of people in a legal process (judge, lawyer, victim, defendant, witness), types of locations (crime location, defendant's address), etc. In this sense, there is still a need for a robust golden collection, annotated with fine-grained entities of the legal domain, and which covers various documents of a legal process, such as petitions, inquiries, complaints, decisions and sentences. In this article, we describe the development of the Golden Collection of the Brazilian Judiciary (CDJUR-BR) contemplating a set of fine-grained named entities that have been annotated by experts in legal documents. The creation of CDJUR-BR followed its own methodology that aimed to attribute a character of comprehensiveness and robustness. Together with the CDJUR-BR repository we provided a NER based on the BERT model and trained with the CDJUR-BR, whose results indicated the prevalence of the CDJUR-BR.

cs.CL

The ubiquitous efficiency of going further: how street networks affect travel speed

As cities struggle to adapt to more ``people-centered'' urbanism, transportation planning and engineering must innovate to expand the street network strategically in order to ensure efficiency but also to deter sprawl. Here, we conducted a study of over 200 cities around the world to understand the impact that the patterns of deceleration points in streets due to traffic signs has in trajectories done from motorized vehicles. We demonstrate that there is a ubiquitous nonlinear relationship between time and distance in the optimal trajectories within each city. More precisely, given a specific period of time $τ$, without any traffic, one can move on average up to the distance $\left \langle D \right \rangle \simτ^β$. We found a super-linear relationship for almost all cities in which $β>1.0$. This points to an efficiency of scale when traveling large distances, meaning the average speed will be higher for longer trips when compared to shorter trips. We demonstrate that this efficiency is a consequence of the spatial distribution of large segments of streets without deceleration points, favoring access to routes in which a vehicle can cross large distances without stops. These findings show that cities must consider how their street morphology can affect travel speed.

physics.soc-ph

Tracing contacts to evaluate the transmission of COVID-19 from highly exposed individuals in public transportation

We investigate, through a data-driven contact tracing model, the transmission of COVID-19 inside buses during distinct phases of the pandemic in a large Brazilian city. From this microscopic approach, we recover the networks of close contacts within consecutive time windows. A longitudinal comparison is then performed by upscaling the traced contacts with the transmission computed from a mean-field compartmental model for the entire city. Our results show that the effective reproduction numbers inside the buses, $Re^{bus}$, and in the city, $Re^{city}$, followed a compatible behavior during the first wave of the local outbreak. Moreover, by distinguishing the close contacts of healthcare workers in the buses, we discovered that their transmission, $Re^{health}$, during the same period, was systematically higher than $Re^{bus}$. This result reinforces the need for special public transportation policies for highly exposed groups of people.

physics.soc-ph

Understanding the impact of the alphabetical ordering of names in user interfaces: a gender bias analysis

Listing people alphabetically on an electronic output device is a traditional technique, since alphabetical order is easily perceived by users and facilitates access to information. However, this apparently harmless technique, especially when the list is ordered by first name, needs to be used with caution by designers and programmers. We show, via empirical data analysis, that when an interface displays people's first name in alphabetical order in several pages/screens, each page/screen may have imbalances in respect to gender of its Top-k individuals.k represents the size of the list of names visualized first, which may be the number of names that fits in a screen page of a certain device.The research work was carried out with the analysis of actual datasets of names of five different countries. Each dataset has a person name and the frequency of adoption of the name in the country.Our analysis shows that, even though all countries have exhibit imbalance problems, the samples of individuals with Brazilian and Spanish first names are more prone to gender imbalance among their Top-k individuals. These results can be useful for designers and engineers to construct information systems that avoid gender bias induction.

cs.HC

A Temporal Clustering Algorithm for Achieving the trade-off between the User Experience and the Equipment Economy in the Context of IoT

We present here the Temporal Clustering Algorithm (TCA), an incremental learning algorithm applicable to problems of anticipatory computing in the context of the Internet of Things. This algorithm was tested in a specific prediction scenario of consumption of an electric water dispenser typically used in tropical countries, in which the ambient temperature is around 30-degree Celsius. In this context, the user typically wants to drinking iced water therefore uses the cooler function of the dispenser. Real and synthetic water consumption data was used to test a forecasting capacity on how much energy can be saved by predicting the pattern of use of the equipment. In addition to using a small constant amount of memory, which allows the algorithm to be implemented at the lowest cost, while using microcontrollers with a small amount of memory (less than 1Kbyte) available on the market. The algorithm can also be configured according to user preference, prioritizing comfort, keeping the water at the desired temperature longer, or prioritizing energy savings. The main result is that the TCA achieved energy savings of up to 40% compared to the conventional mode of operation of the dispenser with an average success rate higher than 90% in its times of use.

cs.LG

A universal approach for drainage basins

Drainage basins are essential to Geohydrology and Biodiversity. Defining those regions in a simple, robust and efficient way is a constant challenge in Earth Science. Here, we introduce a model to delineate multiple drainage basins through an extension of the Invasion Percolation-Based Algorithm (IPBA). In order to prove the potential of our approach, we apply it to real and artificial datasets. We observe that the perimeter and area distributions of basins and anti-basins display long tails extending over several orders of magnitude and following approximately power-law behaviors. Moreover, the exponents of these power laws depend on spatial correlations and are invariant under the landscape orientation, not only for terrestrial, but lunar and martian landscapes. The terrestrial and martian results are statistically identical, which suggests that a hypothetical martian river would present similarity to the terrestrial rivers. Finally, we propose a theoretical value for the Hack's exponent based on the fractal dimension of watersheds, $γ=D/2$. We measure $γ=0.54 \pm 0.01$ for Earth, which is close to our estimation of $γ\approx 0.55$. Our study suggests that Hack's law can have its origin purely in the maximum and minimum lines of the landscapes.

physics.comp-ph

A worldwide model for boundaries of urban settlements

The shape of urban settlements plays a fundamental role in their sustainable planning. Properly defining the boundaries of cities is challenging and remains an open problem in the Science of Cities. Here, we propose a worldwide model to define urban settlements beyond their administrative boundaries through a bottom-up approach that takes into account geographical biases intrinsically associated with most societies around the world, and reflected in their different regional growing dynamics. The generality of the model allows to study the scaling laws of cities at all geographical levels: countries, continents, and the entire world. Our definition of cities is robust and holds to one of the most famous results in Social Sciences: Zipf's law. According to our results, the largest cities in the world are not in line with what was recently reported by the United Nations. For example, we find that the largest city in the world is an agglomeration of several small settlements close to each other, connecting three large settlements: Alexandria, Cairo, and Luxor. Our definition of cities opens the doors to the study of the economy of cities in a systematic way independently of arbitrary definitions that employ administrative boundaries.

physics.soc-ph

Towards Understanding the Impact of Crime in a Choice of a Route by a Bus Passenger

In this paper we describe a simulation platform that supports studies on the impact of crime on urban mobility. We present an example of how this can be achieved by seeking to understand the effect, on the transport system, if users of this system decide to choose optimal routes of time between origins and destinations that normally follow. Based on real data from a large Brazilian metropolis, we found that the percentage of users who follow this policy is small. Most prefer to follow less efficient routes by making bus exchanges at terminals. This can be understood as an indication that the users of the transport system privilege the security factor.

cs.CY

Increasing the Likelihood of Finding Public Transport Riders that Face Problems Through a Data-Driven approach

The maintenance of big cities public transport service quality requires constant monitoring, which may become an expensive and time-consuming practice. The perception of quality, from the users point of view is an important aspect of quality monitoring. In this sense, we proposed a methodology for data analysis and visualization, supported by software, which allows for the structuring of estimates and assumptions of where and who seems to be having unsatisfactory experiences while making use of the public transportation in populous metropolitan areas. Moreover, it provides support in setting up a plan for on-site quality surveys. The proposed methodology increases the likelihood that, with the on-site visits, the interviewer finds users who suffer inconveniences, which influence their behavior. Simulation comparison and a small-scale pilot survey stand for the validity of the proposed method.

cs.CY

Towards Understanding the Impact of Human Mobility on Police Allocation

Motivated by recent findings that human mobility is proxy for crime behavior in big cities and that there is a superlinear relationship between the people's movement and crime, this article aims to evaluate the impact of how these findings influence police allocation. More precisely, we shed light on the differences between an allocation strategy, in which the resources are distributed by clusters of floating population, and conventional allocation strategies, in which the police resources are distributed by an Administrative Area (typically based on resident population). We observed a substantial difference in the distributions of police resources allocated following these strategies, what evidences the imprecision of conventional police allocation methods.

physics.soc-ph

Contextual Data Collection for Smart Cities

As part of Smart Cities initiatives, national, regional and local governments all over the globe are under the mandate of being more open regarding how they share their data. Under this mandate, many of these governments are publishing data under the umbrella of open government data, which includes measurement data from city-wide sensor networks. Furthermore, many of these data are published in so-called data portals as documents that may be spreadsheets, comma-separated value (CSV) data files, or plain documents in PDF or Word documents. The sharing of these documents may be a convenient way for the data provider to convey and publish data but it is not the ideal way for data consumers to reuse the data. For example, the problems of reusing the data may range from difficulty opening a document that is provided in any format that is not plain text, to the actual problem of understanding the meaning of each piece of knowledge inside of the document. Our proposal tackles those challenges by identifying metadata that has been regarded to be relevant for measurement data and providing a schema for this metadata. We further leverage the Human-Aware Sensor Network Ontology (HASNetO) to build an architecture for data collected in urban environments. We discuss the use of HASNetO and the supporting infrastructure to manage both data and metadata in support of the City of Fortaleza, a large metropolitan area in Brazil.

cs.AI

From Data to City Indicators: A Knowledge Graph for Supporting Automatic Generation of Dashboards

In the context of Smart Cities, indicator definitions have been used to calculate values that enable the comparison among different cities. The calculation of an indicator values has challenges as the calculation may need to combine some aspects of quality while addressing different levels of abstraction. Knowledge graphs (KGs) have been used successfully to support flexible representation, which can support improved understanding and data analysis in similar settings. This paper presents an operational description for a city KG, an indicator ontology that support indicator discovery and data visualization and an application capable of performing metadata analysis to automatically build and display dashboards according to discovered indicators. We describe our implementation in an urban mobility setting.

cs.AI

A Service-Oriented Architecture for Assisting the Authoring of Semantic Crowd Maps

Although there are increasingly more initiatives for the generation of semantic knowledge based on user participation, there is still a shortage of platforms for regular users to create applications on which semantic data can be exploited and generated automatically. We propose an architecture, called Semantic Maps (SeMaps), for assisting the authoring and hosting of applications in which the maps combine the aggregation of a Geographic Information System and crowd-generated content (called here crowd maps). In these systems, the digital map works as a blackboard for accommodating stories told by people about events they want to share with others typically participating in their social networks. SeMaps offers an environment for the creation and maintenance of sites based on crowd maps with the possibility for the user to characterize semantically that which s/he intends to mark on the map. The designer of a crowd map, by informing a linguistic expression that designates what has to be marked on the maps, is guided in a process that aims to associate a concept from a common-sense base to this linguistic expression. Thus, the crowd maps start to have dominion over common-sense inferential relations that define the meaning of the marker, and are able to make inferences about the network of linked data. This makes it possible to generate maps that have the power to perform inferences and access external sources (such as DBpedia) that constitute information that is useful and appropriate to the context of the map. In this paper we describe the architecture of SeMaps and how it was applied in a crowd map authoring tool.

cs.AI

Geracao Automatica de Paineis de Controle para Analise de Mobilidade Urbana Utilizando Redes Complexas

In this paper we describe an automatic generator to support the data scientist to construct, in a user-friendly way, dashboards from data represented as networks. The generator called SBINet (Semantic for Business Intelligence from Networks) has a semantic layer that, through ontologies, describes the data that represents a network as well as the possible metrics to be calculated in the network. Thus, with SBINet, the stages of the dashboard constructing process that uses complex network metrics are facilitated and can be done by users who do not necessarily know about complex networks.

cs.AI

Detecção de comunidades em redes complexas para identificar gargalos e desperdício de recursos em sistemas de ônibus

We propose here a methodology to help to understand the shortcomings of public transportation in a city via the mining of complex networks representing the supply and demand of public transport. We show how to build these networks based upon data on smart card use in buses via the application of algorithms that estimate an OD and reconstruct the complete itinerary of the passengers. The overlapping of the two networks sheds light in potential overload and waste in the offer of resources that can be mitigated with strategies for balancing supply and demand.

cs.AI

Análise comparativa de pesquisas de origens e destinos: uma abordagem baseada em Redes Complexas

In this paper, a comparative study was conducted between complex networks representing origin and destination survey data. Similarities were found between the characteristics of the networks of Brazilian cities with networks of foreign cities. Power laws were found in the distributions of edge weights and this scale - free behavior can occur due to the economic characteristics of the cities.

cs.SI

Human Mobility in Large Cities as a Proxy for Crime

We investigate at the subscale of the neighborhoods of a highly populated city the incidence of property crimes in terms of both the resident and the floating population. Our results show that a relevant allometric relation could only be observed between property crimes and floating population. More precisely, the evidence of a superlinear behavior indicates that a disproportional number of property crimes occurs in regions where an increased flow of people takes place in the city. For comparison, we also found that the number of crimes of peace disturbance only correlates well, and in a superlinear fashion too, with the resident population. Our study raises the interesting possibility that the superlinearity observed in previous studies [Bettencourt et al., Proc. Natl. Acad. Sci. USA 104, 7301 (2007) and Melo et al., Sci. Rep. 4, 6239 (2014)] for homicides versus population at the city scale could have its origin in the fact that the floating population, and not the resident one, should be taken as the relevant variable determining the intrinsic microdynamical behavior of the system.

physics.soc-ph

Micro-interventions in urban transport from pattern discovery on the flow of passengers and on the bus network

In this paper, we describe a case study in a big metropolis, in which from data collected by digital sensors, we tried to understand mobility patterns of persons using buses and how this can generate knowledge to suggest interventions that are applied incrementally into the transportation network in use. We have first estimated an Origin-Destination matrix of buses users from datasets about the ticket validation and GPS positioning of buses. Then we represent the supply of buses with their routes through bus stops as a complex network, which allowed us to understand the bottlenecks of the current scenario and, in particular, applying community discovery techniques, to identify clusters that the service supply infrastructure has. Finally, from the superimposing of the flow of people represented in the OriginDestination matrix in the supply network, we exemplify how micro-interventions can be prospected by means of an example of the introduction of express routes.

cs.AI