SearcharxivSearch

arXiv subjects

Daniel Fraiman

Publications and source records attributed to Daniel Fraiman.

16 recordsLinked to original sources

Imbalanced Classification under Capacity Constraints

Detecting observations from a minority class under severe class imbalance is a central challenge in applications such as fraud detection, medical screening, and industrial quality control. In these settings, each positive prediction triggers a costly follow-up action, an MRI scan, a transaction audit, whose execution is subject to real operational constraints. This paper proposes a formal classification framework under capacity constraints: given a user-defined bound limit $b$ on the proportion of observations that can be labeled as belonging to the minority class, the goal is to find the classifier that maximizes sensitivity on that class. We characterize the optimal classifier under this constraint and establish its equivalence with the classical Bayes classifier under a reweighting of the prior probabilities. We also introduce a capacity-adjusted performance metric $M$ that accounts for the effective detection rate when the capacity constraint is binding. The framework is implemented on top of standard learning methods, k-NN, SVM, random forests, and neural networks, and statistical consistency is established for each. We further show that these methods reduce to post-hoc thresholding when no hyperparameters are oriented toward the capacity-constrained objective, and introduce a capacity-aware support vector machine that exploits the constraint during training and achieves the strongest empirical performance. Experiments on the Taiwanese credit card default dataset confirm that capacity-constrained classifiers substantially outperform both classical approaches and SMOTE under high imbalance regimes. The framework extends naturally to multiclass settings and online environments.

stat.ML

An Out-of-Equilibrium 1D Particle System Undergoing Perfectly Plastic Collisions

At time zero, there are $N$ identical point particles in the line (1D) which are characterized by their positions and velocities. Both values are given randomly and independently from each other, with arbitrary probability densities. Each particle evolves at constant velocity until eventually they meet. When this happens, a perfectly-plastic collision is produced, resulting in a new particle composed by the sum of their masses and the weighted average velocity. The merged particles evolve indistinguishably from the non-merged ones, i.e. they move at constant velocity until a new plastic collision eventually happens. As in any open system, the particles are not confined to any region or reservoir, so as time progresses, they go on to infinity. From this non-equilibrium process, the number of (now, non-identical) final particles, $\tilde{X}_N$, the distribution of masses of these final particles and the kinetic energy loss from all plastic collisions, is studied. The principal findings shown in this paper are outlined as follows: (1) A method has been developed to determine the number and mass of the final particles based solely on the initial conditions, eliminating the need to evolve the particle system. (2) A similar model of merging particles, with a universal number of final particles, $\tilde{Z}_N$, is introduced. (3) Strong evidence that $\tilde{X}_N$ is also universal and has the same law of probability as $\tilde{Z}_N$ is presented. (4) An accurate approximation of the energy loss is presented. (5) Results for $\tilde{X}_N$ for an explosive-like initial condition are analyzed.

cond-mat.stat-mech

Group testing with nested pools

In order to identify the infected individuals of a population, their samples are divided in equally sized groups called pools and a single laboratory test is applied to each pool. Individuals whose samples belong to pools that test negative are declared healthy, while each pool that tests positive is divided into smaller, equally sized pools which are tested in the next stage. In the $(k+1)$-th stage all remaining samples are tested. If $p<1-3^{-1/3}$, we minimize the expected number of tests per individual as a function of the number $k+1$ of stages, and of the pool sizes in the first $k$ stages. We show that for each $p\in (0, 1-3^{-1/3})$ the optimal choice is one of four possible schemes, which are explicitly described. We conjecture that for each $p$, the optimal choice is one of the two sequences of pool sizes $(3^k\text{ or }3^{k-1}4,3^{k-1},\dots,3^2,3 )$, with a precise description of the range of $p$'s where each is optimal. The conjecture is supported by overwhelming numerical evidence for $p>2^{-51}$. We also show that the cost of the best among the schemes $(3^k,\dots,3)$ is of order $O\big(p\log(1/p)\big)$, comparable to the information theoretical lower bound $p\log_2(1/p)+(1-p)\log_2(1/(1-p))$, the entropy of a Bernoulli$(p)$ random variable.

math.ST

A self-organized criticality participative pricing mechanism for selling zero-marginal cost products

In today's economy, selling a new zero-marginal cost product is a real challenge, as it is difficult to determine a product's "correct" sales price based on its profit and dissemination. As an example, think of the price of a new app or video game. New sales mechanisms for selling this type of product need to be designed, in particular ones that consider consumer preferences and reality. Current auction mechanisms establish a time deadline for the auction to take place. This deadline is set to increase the number of bidders and thus the final offering price. Consumers want to obtain the product as quickly as possible from the moment they become interested in it, and this time does not always coincide with the seller's deadline. Naturally, consumers also want to pay a price they consider "fair". Here we introduce an auction model where buyers continuously place bids and the challenge is to decide quickly whether or not to accept them. The model does not include a deadline for placing bids, and exhibits self-organized criticality; it presents a critical price from which a bid is accepted with probability one, and avalanches of sales above this value are observed. This model is of particular interest for startup companies interested in profit as well as making the product known on the market.

q-fin.TR

Critical values in Bak-Sneppen type models

In the Bak-Sneppen model, the lowest fitness particle and its two nearest neighbors are renewed at each temporal step with a uniform (0,1) fitness distribution. The model presents a critical value that depends on the interaction criteria (two nearest neighbors) and on the update procedure (uniform). Here we calculate the critical value for models where one or both properties are changed. We study models with non-uniform updates, models with random neighbors and models with binary fitness and obtain exact results for the average fitness and for $p_c$.

nlin.AO

Local equilibrium in the Bak-Sneppen model

The Bak Sneppen (BS) model is a very simple model that exhibits all the richness of self-organized criticality theory. At the thermodynamic limit, the BS model converges to a situation where all particles have a fitness that is uniformly distributed between a critical value $p_c$ and 1. The $p_c$ value is unknown, as are the variables that influence and determine this value. Here, we study the Bak Sneppen model in the case in which the lowest fitness particle interacts with an arbitrary even number of $m$ nearest neighbors. We show that $p_{c,m}$ verifies a simple local equilibrium relationship. Based on this relationship, we can determine bounds for $p_{c,m}$.

nlin.AO

Statistical comparison of (brain) networks

The study of random networks in a neuroscientific context has developed extensively over the last couple of decades. By contrast, techniques for the statistical analysis of these networks are less developed. In this paper, we focus on the statistical comparison of brain networks in a nonparametric framework and discuss the associated detection and identification problems. We tested network differences between groups with an analysis of variance (ANOVA) test we developed specifically for networks. We also propose and analyse the behaviour of a new statistical procedure designed to identify different subnetworks. As an example, we show the application of this tool in resting-state fMRI data obtained from the Human Connectome Project. Finally, we discuss the potential bias in neuroimaging findings that is generated by some behavioural and brain structure variables. Our method can also be applied to other kind of networks such as protein interaction networks, gene networks or social networks.

q-bio.NC

Banking Networks and Leverage Dependence: Evidence from Selected Emerging Countries

We use bank-level balance sheet data from 2005 to 2010 to study interactions within the banking system of five emerging countries: Argentina, Brazil, Mexico, South Africa, and Taiwan. For each country we construct a financial network based on the leverage ratio dependence between each pair of banks, and find results that are comparable across countries. Banks present a variety of leverage ratio behaviors. This leverage diversity produces financial networks that exhibit a modular structure characterized by one large bank community, some small ones and isolated banks. There exist compact structures that have synchronized dynamics. Many groups of banks merge together creating a financial network topology that converges to a unique big cluster at a relatively low leverage dependence level. Finally, we propose a model that includes corporate and interbank loans for studying the banking system. This model generates networks similar to the empirical ones. Moreover, we find that faster-growing banks tend to be more highly interconnected between them, and this is also observed in empirical data.

q-fin.ST

A test of hypotheses for random graph distributions built from EEG data

The theory of random graphs is being applied in recent years to model neural interactions in the brain. While the probabilistic properties of random graphs has been extensively studied in the literature, the development of statistical inference methods for this class of objects has received less attention. In this work we propose a non-parametric test of hypotheses to test if two samples of random graphs were originated from the same probability distribution. We show how to compute efficiently the test statistic and we study its performance on simulated data. We apply the test to compare graphs of brain functional network interactions built from electroencephalographic (EEG) data collected during the visualization of point light displays depicting human locomotion.

stat.AP

Non Parametric Statistics of Dynamic Networks with distinguishable nodes

The study of random graphs and networks had an explosive development in the last couple of decades. Meanwhile, techniques for the statistical analysis of sequences of networks were less developed. In this paper we focus on networks sequences with a fixed number of labeled nodes and study some statistical problems in a nonparametric framework. We introduce natural notions of center and a depth function for networks that evolve in time. We develop several statistical techniques including testing, supervised and unsupervised classification, and some notions of principal component sets in the space of networks. Some examples and asymptotic results are given, as well as two real data examples.

cond-mat.dis-nn

What kind of noise is brain noise: anomalous scaling behavior of the resting brain activity fluctuations

The continuous interaction between brain regions "at rest" defines the so-called resting state networks (RSN) which can be reconstructed from the analysis of functional magnetic resonance imaging (fMRI) data. What dynamical mechanism allows for a flexible large-scale organization of the RSN still remains an important challenge. Here, three key novel properties of the RSN are uncovered. First, the correlation length (i.e., the length at which correlation between two regions vanishes) diverges with the cluster's size considered. Second, this divergence it is observed also for measures of mutual information. Third, the variance of the fMRI mean signal remains constant across the entire range of observed clusters sizes, in contrast with naive expectations. The unveiled scale invariance exposes the RSN optimal information-sharing properties across very diverse networks sizes, architectures and functions, which can be an important marker of healthy brain dynamics.

q-bio.NC

Point process analysis of large-scale brain fMRI dynamics

Functional magnetic resonance imaging (fMRI) techniques have contributed significantly to our understanding of brain function. Current methods are based on the analysis of \emph{gradual and continuous} changes in the brain blood oxygenated level dependent (BOLD) signal. Departing from that approach, recent work has shown that equivalent results can be obtained by inspecting only the relatively large amplitude BOLD signal peaks, suggesting that relevant information can be condensed in \emph{discrete} events. This idea is further explored here to demonstrate how brain dynamics at resting state can be captured just by the timing and location of such events, i.e., in terms of a spatiotemporal point process. As a proof of principle, we show that the resting state networks (RSN) maps can be extracted from such point processes. Furthermore, the analysis uncovers avalanches of activity which are ruled by the same dynamical and statistical properties described previously for neuronal events at smaller scales. Given the demonstrated functional relevance of the resting state brain dynamics, its representation as a discrete process might facilitate large scale analysis of brain function both in health and disease.

q-bio.NC

Ising-like dynamics in large-scale functional brain networks

Brain "rest" is defined -more or less unsuccessfully- as the state in which there is no explicit brain input or output. This work focuss on the question of whether such state can be comparable to any known \emph{dynamical} state. For that purpose, correlation networks from human brain Functional Magnetic Resonance Imaging (fMRI) are constrasted with correlation networks extracted from numerical simulations of the Ising model in 2D, at different temperatures. For the critical temperature $T_c$, striking similarities appear in the most relevant statistical properties, making the two networks indistinguishable from each other. These results are interpreted here as lending support to the conjecture that the dynamics of the functioning brain is near a critical point.

cond-mat.dis-nn

The brain: What is critical about it?

We review the recent proposal that the most fascinating brain properties are related to the fact that it always stays close to a second order phase transition. In such conditions, the collective of neuronal groups can reliably generate robust and flexible behavior, because it is known that at the critical point there is the largest abundance of metastable states to choose from. Here we review the motivation, arguments and recent results, as well as further implications of this view of the functioning brain.

cond-mat.dis-nn

Growing Networks: Limit in-degree distribution for arbitrary out-degree one

We compute the stationary in-degree probability, $P_{in}(k)$, for a growing network model with directed edges and arbitrary out-degree probability. In particular, under preferential linking, we find that if the nodes have a light tail (finite variance) out-degree distribution, then the corresponding in-degree one behaves as $k^{-3}$. Moreover, for an out-degree distribution with a scale invariant tail, $P_{out}(k)\sim k^{-α}$, the corresponding in-degree distribution has exactly the same asymptotic behavior only if $2<α<3$ (infinite variance). Similar results are obtained when attractiveness is included. We also present some results on descriptive statistics measures %descriptive statistics such as the correlation between the number of in-going links, $D_{in}$, and outgoing links, $D_{out}$, and the conditional expectation of $D_{in}$ given $D_{out}$, and we calculate these measures for the WWW network. Finally, we present an application to the scientific publications network. The results presented here can explain the tail behavior of in/out-degree distribution observed in many real networks.

physics.soc-ph

Growing Directed Networks: Estimation and Hypothesis Testing

Based only on the information gathered in a snapshot of a directed network, we present a formal way of checking if the proposed model is correct for the empirical growing network under study. In particular, we show how to estimate the attractiveness, and present an application of the model presented in [arxiv:0704.1847] to the scientific publications network from the ISI dataset.

physics.soc-ph