SearcharxivSearch

arXiv subjects

Valmir C. Barbosa

Publications and source records attributed to Valmir C. Barbosa.

At least 19 recordsLinked to original sources

Tilings of a bounded region of the plane by maximal one-dimensional tiles

We study the tiling of a two-dimensional region of the plane by $K$-cell one-dimensional tiles, or $K$-mers. Unlike previous studies, which typically allowed for one single value of $K$ or sometimes a small assortment of fixed values, here a tiling may concomitantly employ $K$-mers comprising any number $K$ of cells, provided a maximality constraint is satisfied. In essence, this constraint requires each of the $K$-mers in use to be as lengthy as possible, given its surroundings in the resulting tiling. Maximality aims to limit the variety of possible tilings while allowing for interesting behavior in terms of the statistical physical observables of interest. In fact, by introducing an energy function based on cell contacts and parameterizing it appropriately, we have been able to observe relatively unexpected behavior, including the suggestion of phase transitions as the system's temperature evolves.

cond-mat.stat-mech

Near-optimal pilot assignment in cell-free massive MIMO

Cell-free massive MIMO systems are currently being considered as potential enablers of future (6G) technologies for wireless communications. By combining distributed processing and massive MIMO, they are expected to deliver improved user coverage and efficiency. A possible source of performance degradation in such systems is pilot contamination, which contributes to causing interference during uplink training and affects channel estimation negatively. Contamination occurs when the same pilot sequence is assigned to more than one user. This is in general inevitable, as the number of mutually orthogonal pilot sequences corresponds to only a fraction of the coherence interval. We introduce a new algorithm for pilot assignment and analyze its performance both from a theoretical perspective and in computational experiments. We show that it has an approximation ratio close to 1 for a plausibly large number of orthogonal pilot sequences, as well as low computational complexity under massive parallelism. We also show that, on average, it outperforms other methods in terms of per-user SINR and throughput on the uplink.

cs.IT

Exact solution of the full RMSA problem in elastic optical networks

Exact solutions of the Routing, Modulation, and Spectrum Allocation (RMSA) problem in Elastic Optical Networks (EONs), so that the number of admitted demands is maximized while those of regenerators and frequency slots used are minimized, require a complex ILP formulation taking into account frequency-slot continuity and contiguity. We introduce the first such formulation, ending a hiatus of some years since the last ILP formulation for a much simpler RMSA variation was introduced. By exploiting a number of problem and solver specificities, we use the NSFNET topology to illustrate the practicality and importance of obtaining exact solutions.

cs.NI

Configuration space partitioning in tilings of a bounded region of the plane

Given a finite collection of two-dimensional tile types, the field of study concerned with covering the plane with tiles of these types exclusively has a long history, having enjoyed great prominence in the last six to seven decades. Much of this interest has revolved around fundamental geometrical problems such as minimizing the variety of tile types to be used, and also around important applications in areas such as crystallography as well as others. All these applications are of course confined to finite spatial regions, but in many cases they refer back directly to progress in tiling the whole, unbounded plane. Tilings of bounded regions of the plane have also been actively studied, but in general the additional complications imposed by the boundary conditions tend to constrain progress to mostly indirect results, such as recurrence relations. Here we study the tiling of rectangular regions of the plane by rectangular tiles. The tile types we use are squares, dominoes, and straight tetraminoes. For this set of tile types, not even recurrence relations seem to be available. Our approach is to seek to characterize this complex system through some fundamental physical quantities. We do this on two parallel tracks, one analytical for what seems to be the most complex special case still amenable to such approach, the other based on the Wang-Landau method for state-density estimation. Given a simple energy function based solely on tile contacts, we have found either approach to lead to illuminating depictions of entropy, temperature, and above all partitions of the configuration space. The notion of a configuration, in this context, refers to how many tiles of each type are used. We have found that certain partitions help bind together different aspects of the system in question and conjecture that future applications will benefit from the possibilities they afford.

cond-mat.stat-mech

A flexible algorithm to offload DAG applications for edge computing

Multi-access Edge Computing (MEC) is an enabling technology to leverage new network applications, such as virtual/augmented reality, by providing faster task processing at the network edge. This is done by deploying servers closer to the end users to run the network applications. These applications are often intensive in terms of task processing, memory usage, and communication; thus mobile devices may take a long time or even not be able to run them efficiently. By transferring (offloading) the execution of these applications to the servers at the network edge, it is possible to achieve a lower completion time (makespan) and meet application requirements. However, offloading multiple entire applications to the edge server can overwhelm its hardware and communication channel, as well as underutilize the mobile devices' hardware. In this paper, network applications are modeled as Directed Acyclic Graphs (DAGs) and partitioned into tasks, and only part of these tasks are offloaded to the edge server. This is the DAG application partitioning and offloading problem, which is known to be NP-hard. To approximate its solution, this paper proposes the FlexDO algorithm. FlexDO combines a greedy phase with a permutation phase to find a set of offloading decisions, and then chooses the one that achieves the shortest makespan. FlexDO is compared with a proposal from the literature and two baseline decisions, considering realistic DAG applications extracted from the Alibaba Cluster Trace Program. Results show that FlexDO is consistently only 3.9% to 8.9% above the optimal makespan in all test scenarios, which include different levels of CPU availability, a multi-user case, and different communication channel transmission rates. FlexDO outperforms both baseline solutions by a wide margin, and is three times closer to the optimal makespan than its competitor.

cs.NI

Shape complexity in cluster analysis

In cluster analysis, a common first step is to scale the data aiming to better partition them into clusters. Even though many different techniques have throughout many years been introduced to this end, it is probably fair to say that the workhorse in this preprocessing phase has been to divide the data by the standard deviation along each dimension. Like division by the standard deviation, the great majority of scaling techniques can be said to have roots in some sort of statistical take on the data. Here we explore the use of multidimensional shapes of data, aiming to obtain scaling factors for use prior to clustering by some method, like k-means, that makes explicit use of distances between samples. We borrow from the field of cosmology and related areas the recently introduced notion of shape complexity, which in the variant we use is a relatively simple, data-dependent nonlinear function that we show can be used to help with the determination of appropriate scaling factors. Focusing on what might be called "midrange" distances, we formulate a constrained nonlinear programming problem and use it to produce candidate scaling-factor sets that can be sifted on the basis of further considerations of the data, say via expert knowledge. We give results on some iconic data sets, highlighting the strengths and potential weaknesses of the new approach. These results are generally positive across all the data sets used.

cs.LG

A simple linear model to aid in analyses of the Beta Pictoris moving group

We build a four-dimensional linear model of object membership in the Beta Pictoris moving group (BPMG), using two nested applications of Principal Component Analysis (PCA) to high-quality data on about 1.5 million objects. These data contain the objects' galactic space velocities and also their Gaia $G$ magnitudes. Through PCA, they ultimately result in a four-dimensional straight line, referred to as PC $1'$, about which both the bona fide members used to obtain the straight line and the candidate members used to test the model congregate at generally small distances. Our bona fide members come from a recent, Gaia DR2-based compilation. Most candidate members are from a compilation from 2017. Using a standard procedure to flag groups of outliers in data sets, we argue that flagging the few possible outliers we identified on account of distances to PC $1'$ is consistent with the nature of the candidate members in use. The spatial and kinematic measurements that backed their inclusion in the 2017 compilation were of course from before the availability of data from the Gaia mission. Moreover, their radial velocities at the time were either unknown or estimated somewhat unreliably. We propose that PC $1'$ be added to the tool set for BPMG analyses and potentially extended to other young stellar moving groups.

astro-ph.SR

Integrated optimization of heterogeneous-network management and the elusive role of macrocells

We consider heterogeneous wireless networks in the physical interference model and introduce a new formulation of the mixed-integer nonlinear programming problem that addresses base-station activation and many-to-many associations while minimizing power consumption. We also introduce HetNetGA, a genetic algorithm that can tackle the problem without any approximations. Though unsuitable for practical deployment, HetNetGA enables the investigation of such networks' true possibilities. Results for scenarios involving both macrocells and picocells often align with what is expected, but sometimes are unexpected and essentially point to the need to better understand the role of macrocells in helping provide capacity while remaining energetically advantageous.

cs.NI

Interspecies evolutionary dynamics mediated by public goods in bacterial quorum sensing

Bacterial quorum sensing is the communication that takes place between bacteria as they secrete certain molecules into the intercellular medium that later get absorbed by the secreting cells themselves and by others. Depending on cell density, this uptake has the potential to alter gene expression and thereby affect global properties of the community. We consider the case of multiple bacterial species coexisting, referring to each one of them as a genotype and adopting the usual denomination of the molecules they collectively secrete as public goods. A crucial problem in this setting is characterizing the coevolution of genotypes as some of them secrete public goods (and pay the associated metabolic costs) while others do not but may nevertheless benefit from the available public goods. We introduce a network model to describe genotype interaction and evolution when genotype fitness depends on the production and uptake of public goods. The model comprises a random graph to summarize the possible evolutionary pathways the genotypes may take as they interact genetically with one another, and a system of coupled differential equations to characterize the behavior of genotype abundance in time. We study some simple variations of the model analytically and more complex variations computationally. Our results point to a simple trade-off affecting the long-term survival of those genotypes that do produce public goods. This trade-off involves, on the producer side, the impact of producing and that of absorbing the public good. On the non-producer side, it involves the impact of absorbing the public good as well, now compounded by the molecular compatibility between the producer and the non-producer. Depending on how these factors turn out, producers may or may not survive.

q-bio.PE

Scheduling wireless links in the physical interference model by fractional edge coloring

We consider the problem of scheduling the links of wireless mesh networks for capacity maximization in the physical interference model. We represent such a network by an undirected graph $G$, with vertices standing for network nodes and edges for links. We define network capacity to be $1/χ'^*_\mathrm{phys}(G)$, where $χ'^*_\mathrm{phys}(G)$ is a novel edge-chromatic indicator of $G$, one that modifies the notion of $G$'s fractional chromatic index. This index asks that the edges of $G$ be covered by matchings in a certain optimal way. The new indicator does the same, but requires additionally that the matchings used be all feasible in the sense of the physical interference model. Sometimes the resulting optimal covering of $G$'s edge set by feasible matchings is simply a partition of the edge set. In such cases, the index $χ'^*_\mathrm{phys}(G)$ becomes the particular case that we denote by $χ'_\mathrm{phys}(G)$, a similar modification of $G$'s well-known chromatic index. We formulate the exact computation of $χ'^*_\mathrm{phys}(G)$ as a linear programming problem, which we solve for an extensive collection of random geometric graphs used to instantiate networks in the physical interference model. We have found that, depending on node density (number of nodes per unit deployment area), often $G$ is such that $χ'^*_\mathrm{phys}(G)<χ'_\mathrm{phys}(G)$. This bespeaks the possibility of increased network capacity by virtue of simply defining it so that edges are colored in the fractional, rather than the integer, sense.

cs.NI

On the mediation of program allocation in high-demand environments

In this paper we challenge the widely accepted premise that, in order to carry out a distributed computation, say on the cloud, users have to inform, along with all the inputs that the algorithm in use requires, the number of processors to be used. We discuss the complicated nature of deciding the value of such parameter, should it be chosen optimally, and propose the alternative scenario in which this choice is passed on to the server side for automatic determination. We show that the allocation problem arising from this alternative is NP-hard only weakly, being therefore solvable in pseudo-polynomial time. In our proposal, one key component on which the automatic determination of the number of processors is based is the cost model. The one we use, which is being increasingly adopted in the wake of the cloud-computing movement, posits that each single execution of a program is to be subject to current circumstances on both user and server side, and as such be priced independently of all others. Running through our proposal is thus a critique of the established common sense that sizing a set of processors to handle a submission to some provider is entirely up to the user.

cs.DC

Co-evolution of the mitotic and meiotic modes of eukaryotic cellular division

The genetic material of a eukaryotic cell comprises both nuclear DNA (ncDNA) and mitochondrial DNA (mtDNA). These differ markedly in several aspects but nevertheless must encode proteins that are compatible with one another. Here we introduce a network model of the hypothetical co-evolution of the two most common modes of cellular division for reproduction: by mitosis (supporting asexual reproduction) and by meiosis (supporting sexual reproduction). Our model is based on a random hypergraph, with two nodes for each possible genotype, each encompassing both ncDNA and mtDNA. One of the nodes is necessarily generated by mitosis occurring at a parent genotype, the other by meiosis occurring at two parent genotypes. A genotype's fitness depends on the compatibility of its ncDNA and mtDNA. The model has two probability parameters, $p$ and $r$, the former accounting for the diversification of ncDNA during meiosis, the latter for the diversification of mtDNA accompanying both meiosis and mitosis. Another parameter, $λ$, is used to regulate the relative rate at which mitosis- and meiosis-generated genotypes are produced. We have found that, even though $p$ and $r$ do affect the existence of evolutionary pathways in the network, the crucial parameter regulating the coexistence of the two modes of cellular division is $λ$. Depending on genotype size, $λ$ can be valued so that either mode of cellular division prevails. Our study is closely related to a recent hypothesis that views the appearance of cellular division by meiosis, as opposed to division by mitosis, as an evolutionary strategy for boosting ncDNA diversification to keep up with that of mtDNA. Our results indicate that this may well have been the case, thus lending support to the first hypothesis in the field to take into account the role of such ubiquitous and essential organelles as mitochondria.

q-bio.PE

Information-theoretic signatures of biodiversity in the barcoding gene

The COI mitochondrial gene is present in all animal phyla and in a few others, and is the leading candidate for species identification through DNA barcoding. Calculating a generalized form of total correlation on publicly available data on the gene yields distinctive information-theoretic descriptors of the phyla represented in the data. Moreover, performing principal component analysis on standardized versions of these descriptors reveals a strong correlation between the first principal component and the natural logarithm of the number of known living species. The descriptors thus constitute clear information-theoretic signatures of the processes whereby evolution has given rise to current biodiversity.

q-bio.PE

A distributed system for SearchOnMath based on the Microsoft BizSpark program

Mathematical information retrieval is a relatively new area, so the first search tools capable of retrieving mathematical formulas began to appear only a few years ago. The proposals made public so far mostly implement searches on internal university databases, small sets of scientific papers, or Wikipedia in English. As such, only modest computing power is required. In this context, SearchOnMath has emerged as a pioneering tool in that it indexes several different databases and is compatible with several mathematical representation languages. Given the significantly greater number of formulas it handles, a distributed system becomes necessary to support it. The present study is based on the Microsoft BizSpark program and has aimed, for 38 different distributed-system scenarios, to pinpoint the one affording the best response times when searching the SearchOnMath databases for a collection of 120 formulas.

cs.IR

Power-law decay of the degree-sequence probabilities of multiple random graphs with application to graph isomorphism

We consider events over the probability space generated by the degree sequences of multiple independent Erdős-Rényi random graphs, and consider an approximation probability space where such degree sequences are deemed to be sequences of i.i.d. random variables. We show that, for any sequence of events with probabilities asymptotically smaller than some power law in the approximation model, the same upper bound also holds in the original model. We accomplish this by extending an approximation framework proposed in a seminal paper by McKay and Wormald. Finally, as an example, we apply the developed framework to bound the probability of isomorphism-related events over multiple independent random graphs.

math.PR

Information integration from distributed threshold-based interactions

We consider a collection of distributed units that interact with one another through the sending of messages. Each message carries a positive ($+1$) or negative ($-1$) tag and causes the receiving unit to send out messages as a function of the tags it has received and a threshold. This simple model abstracts some of the essential characteristics of several systems used in the field of artificial intelligence, and also of biological systems epitomized by the brain. We study the integration of information inside a temporal window as the model's dynamics unfolds. We quantify information integration by the total correlation, relative to the window's duration ($w$), of a set of random variables valued as a function of message arrival. Total correlation refers to the rise of information gain above and beyond that which the units already achieve individually, being therefore related to consciousness studies in some models. We report on extensive computational experiments that explore the interrelations of the model's parameters (two probabilities and the threshold), highlighting relevant scenarios of message traffic and how they impact the behavior of total correlation as a function of $w$. We find that total correlation can occur at significant fractions of the maximum possible value and provide semi-analytical results on the message-traffic characteristics associated with values of $w$ for which it peaks. We then reinterpret the model's parameters in terms of the current best estimates of some quantities pertaining to cortical structure and dynamics. We find the resulting possibilities for best values of $w$ to be well aligned with the time frames within which percepts are thought to be processed and eventually rendered conscious.

q-bio.NC

Local symmetry in random graphs

Quite often real-world networks can be thought of as being symmetric, in the abstract sense that vertices can be found to have similar or equivalent structural roles. However, traditional measures of symmetry in graphs are based on their automorphism groups, which do not account for the similarity of local structures. We introduce the concept of local symmetry, which reflects the structural equivalence of the vertices' egonets. We study the emergence of asymmetry in the Erdős-Rényi random graph model and identify regimes of both asymptotic local symmetry and asymptotic local asymmetry. We find that local symmetry persists at least to an average degree of $n^{1/3}$ and local asymmetry emerges at an average degree not greater than $n^{1/2}$, which are regimes of much larger average degree than for traditional, global asymmetry.

math.PR

Counting independent terms in big-oh notation

The field of computational complexity is concerned both with the intrinsic hardness of computational problems and with the efficiency of algorithms to solve them. Given such a problem, normally one designs an algorithm to solve it and sets about establishing bounds on its performance as functions of the algorithm's variables, particularly upper bounds expressed via the big-oh notation. But if we were given some inscrutable code and were asked to figure out its big-oh profile from performance data on a given set of inputs, how hard would we have to grapple with the various possibilities before zooming in on a reasonably small set of candidates? Here we show that, even if we restricted our search to upper bounds given by polynomials, the number of possibilities could be arbitrarily large for two or more variables. This is unexpected, given the available body of examples on algorithmic efficiency, and serves to illustrate the many facets of the big-oh notation, as well as its counter-intuitive twists.

cs.CC