Searcharxiv⌕ Search

arXiv subjects

John R. J. Thompson

Publications and source records attributed to John R. J. Thompson.

6 recordsLinked to original sources

From Corridor Selection to Earthwork: A Multi-Stage Framework for Automated Road Design via Steiner Trees and Convex Optimization

Designing road networks for wind farms in complex terrain is a challenging task, especially under tight construction budgets. Traditional manual methods are time-consuming and may yield suboptimal results. We propose a structured three-phase optimization framework, TriPhase, to automate and minimize road construction costs from corridor selection to earthwork. Phase one formulates corridor selection as a Steiner minimum tree problem over a terrain-aware graph, incorporating slope and curvature constraints. In phase two, each road segment undergoes horizontal alignment optimization using a bilevel model, where a mixed-integer program evaluates vertical alignment costs. To ensure solver compatibility and performance, the model is reformulated explicitly for Gurobi. The final phase applies a network-wide convex optimization model to refine vertical alignment. Numerical experiments on real-world sites demonstrate up to 14% cost savings compared to industry-standard manual designs, validating the framework's effectiveness and practical relevance.

cs.CE↗

Spectrally Tuned Bandwidth Selection for Kernel Fuzzy Relational Clustering

Fuzzy clustering is used to identify overlapping geometric cluster structures through partial memberships. However, classical methods are limited by the assumption of equal variable importance and by sensitivity to the fuzzifier parameter. These limitations may yield equal cluster membership probabilities, which we refer to as the uniform solution. To address these issues, we propose Kernel Fuzzy Relational Clustering (KFRC) equipped with a bandwidth selection algorithm tuned via the spectral properties of the induced kernel Gram matrix. The KFRC framework implicitly performs unsupervised kernel metric learning by controlling the geometric embedding of the data through adjustable bandwidth parameters. We conduct a formal stability analysis to identify the exact theoretical conditions under which relational clustering collapses, thereby ensuring the stable performance of KFRC. We find that our two-stage bandwidth selection procedure adapts to the data structure while actively avoiding the uniform solution. Furthermore, this theoretical analysis leads to the proposal of a novel fuzzifier function that presents distinct advantages over the power fuzzifier function. We conduct experiments on several synthetic and publicly available data sets to demonstrate that the proposed framework consistently recovers complex structures that traditional methods fail to resolve, while ensuring a purely fuzzy solution.

stat.ME↗

Mixed-type Distance Shrinkage and Selection for Clustering via Kernel Metric Learning

Distance-based clustering and classification are widely used in various fields to group mixed numeric and categorical data. In many algorithms, a predefined distance measurement is used to cluster data points based on their dissimilarity. While there exist numerous distance-based measures for data with pure numerical attributes and several ordered and unordered categorical metrics, an efficient and accurate distance for mixed-type data that utilizes the continuous and discrete properties simulatenously is an open problem. Many metrics convert numerical attributes to categorical ones or vice versa. They handle the data points as a single attribute type or calculate a distance between each attribute separately and add them up. We propose a metric called KDSUM that uses mixed kernels to measure dissimilarity, with cross-validated optimal bandwidth selection. We demonstrate that KDSUM is a shrinkage method from existing mixed-type metrics to a uniform dissimilarity metric, and improves clustering accuracy when utilized in existing distance-based clustering algorithms on simulated and real-world datasets containing continuous-only, categorical-only, and mixed-type data.

cs.LG↗

Anisotropic local constant smoothing for change-point regression function estimation

Understanding forest fire spread in any region of Canada is critical to promoting forest health, and protecting human life and infrastructure. Quantifying fire spread from noisy images, where regions of a fire are separated by change-point boundaries, is critical to faithfully estimating fire spread rates. In this research, we develop a statistically consistent smooth estimator that allows us to denoise fire spread imagery from micro-fire experiments. We develop an anisotropic smoothing method for change-point data that uses estimates of the underlying data generating process to inform smoothing. We show that the anisotropic local constant regression estimator is consistent with convergence rate $O\left(n^{-1/{(q+2)}}\right)$. We demonstrate its effectiveness on simulated one- and two-dimensional change-point data and fire spread imagery from micro-fire experiments.

stat.ME↗

Measuring Financial Advice: aligning client elicited and revealed risk

Financial advisors use questionnaires and discussions with clients to determine a suitable portfolio of assets that will allow clients to reach their investment objectives. Financial institutions assign risk ratings to each security they offer, and those ratings are used to guide clients and advisors to choose an investment portfolio risk that suits their stated risk tolerance. This paper compares client Know Your Client (KYC) profile risk allocations to their investment portfolio risk selections using a value-at-risk discrepancy methodology. Value-at-risk is used to measure elicited and revealed risk to show whether clients are over-risked or under-risked, changes in KYC risk lead to changes in portfolio configuration, and cash flow affects a client's portfolio risk. We demonstrate the effectiveness of value-at-risk at measuring clients' elicited and revealed risk on a dataset provided by a private Canadian financial dealership of over $50,000$ accounts for over $27,000$ clients and $300$ advisors. By measuring both elicited and revealed risk using the same measure, we can determine how well a client's portfolio aligns with their stated goals. We believe that using value-at-risk to measure client risk provides valuable insight to advisors to ensure that their practice is KYC compliant, to better tailor their client portfolios to stated goals, communicate advice to clients to either align their portfolios to stated goals or refresh their goals, and to monitor changes to the clients' risk positions across their practice.

econ.EM↗

Know Your Clients' behaviours: a cluster analysis of financial transactions

In Canada, financial advisors and dealers are required by provincial securities commissions and self-regulatory organizations--charged with direct regulation over investment dealers and mutual fund dealers--to respectively collect and maintain Know Your Client (KYC) information, such as their age or risk tolerance, for investor accounts. With this information, investors, under their advisor's guidance, make decisions on their investments which are presumed to be beneficial to their investment goals. Our unique dataset is provided by a financial investment dealer with over 50,000 accounts for over 23,000 clients. We use a modified behavioural finance recency, frequency, monetary model for engineering features that quantify investor behaviours, and machine learning clustering algorithms to find groups of investors that behave similarly. We show that the KYC information collected does not explain client behaviours, whereas trade and transaction frequency and volume are most informative. We believe the results shown herein encourage financial regulators and advisors to use more advanced metrics to better understand and predict investor behaviours.

econ.EM↗