Searcharxiv⌕ Search

arXiv subjects

Ondřej Sokol

Publications and source records attributed to Ondřej Sokol.

6 recordsLinked to original sources

Product Relation Correlation and Its Use in Product Clustering

This paper introduces product relation correlation, a measure of product relatedness that assesses the extent to which products may function as substitutes or complements through analysis of shared purchasing patterns. Product relation correlation can be used for tasks such as product clustering and shelf space optimization, enabling retailers to arrange items in ways that enhance customer experience. Applied to data from a retail drugstore chain, the measure demonstrates an alignment with cross-price elasticity, increasing as products diverge from independence. With computational simplicity, requirement for only commonly available data, and a robust theoretical interpretation, product relation correlation serves as a practical and efficient tool for deriving useful product insights.

stat.AP↗

Clustering Retail Products Based on Customer Behaviour

The categorization of retail products is essential for the business decision-making process. It is a common practice to classify products based on their quantitative and qualitative characteristics. In this paper we use a purely data-driven approach. Our clustering of products is based exclusively on the customer behaviour. We propose a method for clustering retail products using market basket data. Our model is formulated as an optimization problem which is solved by a genetic algorithm. It is demonstrated on simulated data how our method behaves in different settings. The application using real data from a Czech drugstore company shows that our method leads to similar results in comparison with the classification by experts. The number of clusters is a parameter of our algorithm. We demonstrate that if more clusters are allowed than the original number of categories is, the method yields additional information about the structure of the product categorization.

stat.AP↗

The NP-hard problem of computing the maximal sample variance over interval data is solvable in almost linear time with high probability

We consider the algorithm by Ferson et al. (Reliable computing 11(3), p. 207-233, 2005) designed for solving the NP-hard problem of computing the maximal sample variance over interval data, motivated by robust statistics (in fact, the formulation can be written as a nonconvex quadratic program with a specific structure). First, we propose a new version of the algorithm improving its original time bound $O(n^2 2^ω)$ to $O(n \log n+n\cdot 2^ω)$, where $n$ is number of input data and $ω$ is the clique number in a certain intersection graph. Then we treat input data as random variables as it is usual in statistics) and introduce a natural probabilistic data generating model. We get $2^ω= O(n^{1/\log\log n})$ and $ω= O(\log n / \log\log n)$ on average. This results in average computing time $O(n^{1+ε})$ for $ε> 0$ arbitrarily small, which may be considered as "surprisingly good" average time complexity for solving an NP-hard problem. Moreover, we prove the following tail bound on the distribution of computation time: hard instances, forcing the algorithm to compute in time $2^{Ω(n)}$, occur rarely, with probability tending to zero at the rate $e^{-n\log\log n}$.

math.OC↗

Clustering with Penalty for Joint Occurrence of Objects: Computational Aspects

The method of Holý, Sokol and Černý (Applied Soft Computing, 2017, Vol. 60, p. 752-762) clusters objects based on their incidence in a large number of given sets. The idea is to minimize the occurrence of multiple objects from the same cluster in the same set. In the current paper, we study computational aspects of the method. First, we prove that the problem of finding the optimal clustering is NP-hard. Second, to numerically find a suitable clustering, we propose to use the genetic algorithm augmented by a renumbering procedure, a fast task-specific local search heuristic and an initial solution based on a simplified model. Third, in a simulation study, we demonstrate that our improvements of the standard genetic algorithm significantly enhance its computational performance.

cs.AI↗

How Many Customers Does a Retail Store Have?

The knowledge of the number of customers is the pillar of retail business analytics. In our setting, we assume that a portion of customers is monitored and easily counted due to the loyalty program while the rest is not monitored. The behavior of customers in both groups may significantly differ making the estimation of the number of unmonitored customers a non-trivial task. We identify shopping patterns of several customer segments which allows us to estimate the distribution of customers without the loyalty card using the maximum likelihood method. In a simulation study, we find that the proposed approach is quite precise even when the data sample is very small and its assumptions are violated to a certain degree. In an empirical study of a drugstore chain, we validate and illustrate the proposed approach in practice. The actual number of customers estimated by the proposed method is much higher than the number suggested by the naive estimate assuming the constant customer distribution. The proposed method can also be utilized to determine penetration of the loyalty program in the individual customer segments.

stat.AP↗

The Role of Shopping Mission in Retail Customer Segmentation

In retailing, it is important to understand customer behavior and determine customer value. A useful tool to achieve such goals is the cluster analysis of transaction data. Typically, a customer segmentation is based on the recency, frequency and monetary value of shopping or the structure of purchased products. We take a different approach and base our segmentation on the shopping mission - a reason why a customer visits the shop. Shopping missions include focused purchases of specific product categories and general purchases of various sizes. In an application to a Czech drugstore chain, we show that the proposed segmentation brings unique information about customers and should be used alongside the traditional methods.

stat.AP↗