SearcharxivSearch

arXiv subjects

James Martin

Publications and source records attributed to James Martin.

At least 19 recordsLinked to original sources

Two-stage Adaptive Design Cluster Randomised Trials

Adaptive sample size re-estimation, early stopping, and trial re-design at interim analyses can reduce expected sample sizes in randomised trials. Cluster randomised trials, in which groups of participants are randomly allocated to treatment status, may particularly benefit as they can be costly and their required sample sizes depend on one or more auxiliary parameters governing correlations within and between clusters, which are often estimated with high uncertainty. We adapt a combination test approach to the cluster trial setting allowing for early stopping for futility or efficacy and accounting for correlations between trial stages and other nuisance parameters. We consider design decisions for multi-dimensional sample sizes involving clusters, participants, and time and allowing for modifications to intervention roll-out patterns. We use a Pareto optimality approach to balance objectives relating to different components of the sample size and costs. We also examine the interim estimation of auxiliary parameters and trial re-design for efficiency. We illustrate the methods including an example of a parallel cluster trial re-design and a re-analysis of the large cluster randomised trial E-MOTIVE.

stat.ME

RealSynCol: a high-fidelity synthetic colon dataset for 3D reconstruction applications

Deep learning has the potential to improve colonoscopy by enabling 3D reconstruction of the colon, providing a comprehensive view of mucosal surfaces and lesions, and facilitating the identification of unexplored areas. However, the development of robust methods is limited by the scarcity of large-scale ground truth data. We propose RealSynCol, a highly realistic synthetic dataset designed to replicate the endoscopic environment. Colon geometries extracted from 10 CT scans were imported into a virtual environment that closely mimics intraoperative conditions and rendered with realistic vascular textures. The resulting dataset comprises 28\,130 frames, paired with ground truth depth maps, optical flow, 3D meshes, and camera trajectories. A benchmark study was conducted to evaluate the available synthetic colon datasets for the tasks of depth and pose estimation. Results demonstrate that the high realism and variability of RealSynCol significantly enhance generalization performance on clinical images, proving it to be a powerful tool for developing deep learning algorithms to support endoscopic diagnosis.

cs.CV

Permutation invariance in last-passage percolation and the distribution of the Busemann process

In i.i.d. exponential last-passage percolation, we describe the joint distribution of Busemann functions, over all edges and over all directions, in terms of a joint last-passage problem in a finite inhomogeneous environment. More specifically, the Busemann increments within a $k\times\ell$ grid, and associated to $d$ different directions, are equal in distribution to a particular collection of last-passage increments inside a $(k+d-1)\times(\ell+d-1)$ grid. The joint Busemann distribution was previously described along a horizontal line by Fan and the fourth author, using certain queuing maps. By contrast, our new description explicitly gives the joint distribution for any collection of edges (not just along a horizontal line) using only finitely many random variables. Our result thus provides an exact and accessible way to sample from the joint distribution. In the proof, we rely on one-directional marginal distributions of the inhomogeneous Busemann functions recently studied by Janjigian and the second and fourth authors. The second ingredient of our proof is a novel joint invariance of inhomogeneous last-passage times under permutations of the inhomogeneity parameters. Our proof of the invariance is different from earlier proofs of such results, using the Burke property instead of the RSK correspondence, and leading to an explicit coupling of the weights before and after the permutation of the parameters.

math.PR

Towards Actionable Pedagogical Feedback: A Multi-Perspective Analysis of Mathematics Teaching and Tutoring Dialogue

Effective feedback is essential for refining instructional practices in mathematics education, and researchers often turn to advanced natural language processing (NLP) models to analyze classroom dialogues from multiple perspectives. However, utterance-level discourse analysis encounters two primary challenges: (1) multifunctionality, where a single utterance may serve multiple purposes that a single tag cannot capture, and (2) the exclusion of many utterances from domain-specific discourse move classifications, leading to their omission in feedback. To address these challenges, we proposed a multi-perspective discourse analysis that integrates domain-specific talk moves with dialogue act (using the flattened multi-functional SWBD-MASL schema with 43 tags) and discourse relation (applying Segmented Discourse Representation Theory with 16 relations). Our top-down analysis framework enables a comprehensive understanding of utterances that contain talk moves, as well as utterances that do not contain talk moves. This is applied to two mathematics education datasets: TalkMoves (teaching) and SAGA22 (tutoring). Through distributional unigram analysis, sequential talk move analysis, and multi-view deep dive, we discovered meaningful discourse patterns, and revealed the vital role of utterances without talk moves, demonstrating that these utterances, far from being mere fillers, serve crucial functions in guiding, acknowledging, and structuring classroom discourse. These insights underscore the importance of incorporating discourse relations and dialogue acts into AI-assisted education systems to enhance feedback and create more responsive learning environments. Our framework may prove helpful for providing human educator feedback, but also aiding in the development of AI agents that can effectively emulate the roles of both educators and students.

cs.CL

Wireless Network Topology Inference: A Markov Chains Approach

We address the problem of inferring the topology of a wireless network using limited observational data. Specifically, we assume that we can detect when a node is transmitting, but no further information regarding the transmission is available. We propose a novel network estimation procedure grounded in the following abstract problem: estimating the parameters of a finite discrete-time Markov chain by observing, at each time step, which states are visited by multiple ``anonymous'' copies of the chain. We develop a consistent estimator that approximates the transition matrix of the chain in the operator norm, with the number of required samples scaling roughly linearly with the size of the state space. Applying this estimation procedure to wireless networks, our numerical experiments demonstrate that the proposed method accurately infers network topology across a wide range of parameters, consistently outperforming transfer entropy, particularly under conditions of high network congestion.

cs.NI

Percolation and localisation: Sub-leading eigenvalues of the nonbacktracking matrix

The spectrum of the nonbacktracking matrix associated to a network is known to contain fundamental information regarding percolation properties of the network. Indeed, the inverse of its leading eigenvalue is often used as an estimate for the percolation threshold. However, for many networks with nonbacktracking centrality localised on a few nodes, such as networks with a core-periphery structure, this spectral approach badly underestimates the threshold. In this work, we study networks that exhibit this localisation effect by looking beyond the leading eigenvalue and searching deeper into the spectrum of the nonbacktracking matrix. We identify that, when localisation is present, the threshold often more closely aligns with the inverse of one of the sub-leading real eigenvalues: the largest real eigenvalue with a "delocalised" corresponding eigenvector. We investigate a core-periphery network model and determine, both theoretically and experimentally, a regime of parameters for which our approach closely approximates the threshold, while the estimate derived using the leading eigenvalue does not. We further present experimental results on large scale real-world networks that showcase the usefulness of our approach.

physics.soc-ph

An iterative spectral algorithm for digraph clustering

Graph clustering is a fundamental technique in data analysis with applications in many different fields. While there is a large body of work on clustering undirected graphs, the problem of clustering directed graphs is much less understood. The analysis is more complex in the directed graph case for two reasons: the clustering must preserve directional information in the relationships between clusters, and directed graphs have non-Hermitian adjacency matrices whose properties are less conducive to traditional spectral methods. Here we consider the problem of partitioning the vertex set of a directed graph into $k\ge 2$ clusters so that edges between different clusters tend to follow the same direction. We present an iterative algorithm based on spectral methods applied to new Hermitian representations of directed graphs. Our algorithm performs favourably against the state-of-the-art, both on synthetic and real-world data sets. Additionally, it is able to identify a "meta-graph" of $k$ vertices that represents the higher-order relations between clusters in a directed graph. We showcase this capability on data sets pertaining food webs, biological neural networks, and the online card game Hearthstone.

physics.soc-ph

Adaptive Data Transport Mechanism for UAV Surveillance Missions in Lossy Environments

Unmanned Aerial Vehicles (UAVs) play an increasingly critical role in Intelligence, Surveillance, and Reconnaissance (ISR) missions such as border patrolling and criminal detection, thanks to their ability to access remote areas and transmit real-time imagery to processing servers. However, UAVs are highly constrained by payload size, power limits, and communication bandwidth, necessitating the development of highly selective and efficient data transmission strategies. This has driven the development of various compression and optimal transmission technologies for UAVs. Nevertheless, most methods strive to preserve maximal information in transferred video frames, missing the fact that only certain parts of images/video frames might offer meaningful contributions to the ultimate mission objectives in the ISR scenarios involving moving object detection and tracking (OD/OT). This paper adopts a different perspective, and offers an alternative AI-driven scheduling policy that prioritizes selecting regions of the image that significantly contributes to the mission objective. The key idea is tiling the image into small patches and developing a deep reinforcement learning (DRL) framework that assigns higher transmission probabilities to patches that present higher overlaps with the detected object of interest, while penalizing sharp transitions over consecutive frames to promote smooth scheduling shifts. Although we used Yolov-8 object detection and UDP transmission protocols as a benchmark testing scenario the idea is general and applicable to different transmission protocols and OD/OT methods. To further boost the system's performance and avoid OD errors for cluttered image patches, we integrate it with interframe interpolations.

eess.IV

The inhomogeneous $t$-PushTASEP and Macdonald polynomials

We study a multispecies $t$-PushTASEP system on a finite ring of $n$ sites with site-dependent rates $x_1,\dots,x_n$. Let $\lambda=(\lambda_1,\dots,\lambda_n)$ be a partition whose parts represent the species of the $n$ particles on the ring. We show that for each composition $\eta$ obtained by permuting the parts of $\lambda$, the stationary probability of being in state $\eta$ is proportional to the ASEP polynomial $F_{\eta}(x_1,\dots,x_n; q,t)$ at $q=1$; the normalizing constant (or partition function) is the Macdonald polynomial $P_{\lambda}(x_1,\dots,x_n;q,t)$ at $q=1$. Our approach involves new relations between the families of ASEP polynomials and of non-symmetric Macdonald polynomials at $q=1$. We also use multiline diagrams, showing that a single jump of the PushTASEP system is closely related to the operation of moving from one line to the next in a multiline diagram. We derive symmetry properties for the system under permutation of its jump rates, as well as a formula for the current of a single-species system.

math.CO

Deep API Learning Revisited

Understanding the correct API usage sequences is one of the most important tasks for programmers when they work with unfamiliar libraries. However, programmers often encounter obstacles to finding the appropriate information due to either poor quality of API documentation or ineffective query-based searching strategy. To help solve this issue, researchers have proposed various methods to suggest the sequence of APIs given natural language queries representing the information needs from programmers. Among such efforts, Gu et al. adopted a deep learning method, in particular an RNN Encoder-Decoder architecture, to perform this task and obtained promising results on common APIs in Java. In this work, we aim to reproduce their results and apply the same methods for APIs in Python. Additionally, we compare the performance with a more recent Transformer-based method, i.e., CodeBERT, for the same task. Our experiment reveals a clear drop in performance measures when careful data cleaning is performed. Owing to the pretraining from a large number of source code files and effective encoding technique, CodeBERT outperforms the method by Gu et al., to a large extent.

cs.SE

The Foata-Fuchs proof of Cayley's formula, and its probabilistic uses

We present a very simple bijective proof of Cayley's formula due to Foata and Fuchs (1970). This bijection turns out to be very useful when seen through a probabilistic lens; we explain some of the ways in which it can be used to derive probabilistic identities, bounds, and growth procedures for random trees with given degrees, including random d-ary trees. We also introduce a partial order on the degree sequences of rooted trees, and conjecture that it induces a stochastic partial order on heights of random rooted trees with given degrees.

math.CO

Lessons Learned from the Real-world Deployment of a Connected Vehicle Testbed

The connected vehicle (CV) system promises unprecedented safety, mobility, environmental, economic and social benefits, which can be unlocked using the enormous amount of data shared between vehicles and infrastructure (e.g., traffic signals, centers). Real world CV deployments including pilot deployments help solve technical issues and observe potential benefits, both of which support the broader adoption of the CV system. This study focused on the Clemson University Connected Vehicle Testbed (CUCVT) with the goal of sharing the lessons learned from the CUCVT deployment. The motivation of this study was to enhance early CV deployments with the objective of depicting the lessons learned from the CUCVT testbed, which includes unique features to support multiple CV applications running simultaneously. The lessons learned in the CUCVT testbed are described at three different levels: i) the development of system architecture and prototyping in a controlled environment, ii) the deployment of the CUCVT testbed, and iii) the validation of the CV application experiments in the CUCVT. Our field experiments with a CV application validated the functionalities needed for running multiple diverse CV applications simultaneously under heterogeneous wireless networking, and realtime and non real time data analytics requirements. The unique deployment experiences, related to heterogeneous wireless networks, real time data aggregation, a distribution using broker system and data archiving with big data management tools, gained from the CUCVT testbed, could be used to advance CV research and guide public and private agencies for the deployment of CVs in the real world.

cs.CY

Adaptive Queue Prediction Algorithm for an Edge Centric Cyber Physical System Platform in a Connected Vehicle Environment

In the early days of connected vehicles (CVs), data will be collected only from a limited number of CVs (i.e., low CV penetration rate) and not from other vehicles (i.e., non-connected vehicles). Moreover, the data loss rate in the wireless CV environment contributes to the unavailability of data from the limited number of CVs. Thus, it is very challenging to predict traffic behavior, which changes dynamically over time, with the limited CV data. The primary objective of this study was to develop an adaptive queue prediction algorithm to predict real-time queue status in the CV environment in an edge-centric cyber-physical system (CPS), which is a relatively new CPS concept. The adaptive queue prediction algorithm was developed using a machine learning algorithm with a real-time feedback system. The algorithm was evaluated using SUMO (i.e., Simulation of Urban Mobility) and ns3 (Network Simulator 3) simulation platforms to illustrate the efficacy of the algorithm on a roadway network in Clemson, South Carolina, USA. The performance of the adaptive queue prediction application was measured in terms of queue detection accuracy with varying CV penetration levels and data loss rates. The analyses revealed that the adaptive queue prediction algorithm with feedback system outperforms without feedback system algorithm.

cs.CY

Critical random forests

Let $F(N,m)$ denote a random forest on a set of $N$ vertices, chosen uniformly from all forests with $m$ edges. Let $F(N,p)$ denote the forest obtained by conditioning the Erdos-Renyi graph $G(N,p)$ to be acyclic. We describe scaling limits for the largest components of $F(N,p)$ and $F(N,m)$, in the critical window $p=N^{-1}+O(N^{-4/3})$ or $m=N/2+O(N^{2/3})$. Aldous described a scaling limit for the largest components of $G(N,p)$ within the critical window in terms of the excursion lengths of a reflected Brownian motion with time-dependent drift. Our scaling limit for critical random forests is of a similar nature, but now based on a reflected diffusion whose drift depends on space as well as on time.

math.PR

Stability of Service under Time-of-Use Pricing

We consider "time-of-use" pricing as a technique for matching supply and demand of temporal resources with the goal of maximizing social welfare. Relevant examples include energy, computing resources on a cloud computing platform, and charging stations for electric vehicles, among many others. A client/job in this setting has a window of time during which he needs service, and a particular value for obtaining it. We assume a stochastic model for demand, where each job materializes with some probability via an independent Bernoulli trial. Given a per-time-unit pricing of resources, any realized job will first try to get served by the cheapest available resource in its window and, failing that, will try to find service at the next cheapest available resource, and so on. Thus, the natural stochastic fluctuations in demand have the potential to lead to cascading overload events. Our main result shows that setting prices so as to optimally handle the {\em expected} demand works well: with high probability, when the actual demand is instantiated, the system is stable and the expected value of the jobs served is very close to that of the optimal offline algorithm.

cs.GT

Optimal low-rank approximations of Bayesian linear inverse problems

In the Bayesian approach to inverse problems, data are often informative, relative to the prior, only on a low-dimensional subspace of the parameter space. Significant computational savings can be achieved by using this subspace to characterize and approximate the posterior distribution of the parameters. We first investigate approximation of the posterior covariance matrix as a low-rank update of the prior covariance matrix. We prove optimality of a particular update, based on the leading eigendirections of the matrix pencil defined by the Hessian of the negative log-likelihood and the prior precision, for a broad class of loss functions. This class includes the F\"{o}rstner metric for symmetric positive definite matrices, as well as the Kullback-Leibler divergence and the Hellinger distance between the associated distributions. We also propose two fast approximations of the posterior mean and prove their optimality with respect to a weighted Bayes risk under squared-error loss. These approximations are deployed in an offline-online manner, where a more costly but data-independent offline calculation is followed by fast online evaluations. As a result, these approximations are particularly useful when repeated posterior mean evaluations are required for multiple data sets. We demonstrate our theoretical results with several numerical examples, including high-dimensional X-ray tomography and an inverse heat conduction problem. In both of these examples, the intrinsic low-dimensional structure of the inference problem can be exploited while producing results that are essentially indistinguishable from solutions computed in the full space.

math.NA

Likelihood-informed dimension reduction for nonlinear inverse problems

The intrinsic dimensionality of an inverse problem is affected by prior information, the accuracy and number of observations, and the smoothing properties of the forward operator. From a Bayesian perspective, changes from the prior to the posterior may, in many problems, be confined to a relatively low-dimensional subspace of the parameter space. We present a dimension reduction approach that defines and identifies such a subspace, called the "likelihood-informed subspace" (LIS), by characterizing the relative influences of the prior and the likelihood over the support of the posterior distribution. This identification enables new and more efficient computational methods for Bayesian inference with nonlinear forward models and Gaussian priors. In particular, we approximate the posterior distribution as the product of a lower-dimensional posterior defined on the LIS and the prior distribution marginalized onto the complementary subspace. Markov chain Monte Carlo sampling can then proceed in lower dimensions, with significant gains in computational efficiency. We also introduce a Rao-Blackwellization strategy that de-randomizes Monte Carlo estimates of posterior expectations for additional variance reduction. We demonstrate the efficiency of our methods using two numerical examples: inference of permeability in a groundwater system governed by an elliptic PDE, and an atmospheric remote sensing problem based on Global Ozone Monitoring System (GOMOS) observations.

stat.CO

Long-range last-passage percolation on the line

We consider directed last-passage percolation on the random graph G = (V,E) where V = Z and each edge (i,j), for i < j, is present in E independently with some probability 0 < p <= 1. To every present edge (i,j) we attach i.i.d. random weights v_{i,j} > 0. We are interested in the behaviour of w_{0,n}, which is the maximum weight of all directed paths from 0 to n, as n tends to infinity. We see two very different types of behaviour, depending on whether E[v_{i,j}^2] is finite or infinite. In the case where E[v_{i,j}^2] is finite we show that the process has a certain regenerative structure, and prove a strong law of large numbers and, under an extra assumption, a functional central limit theorem. In the situation where E[v_{i,j}^2] is infinite we obtain scaling laws and asymptotic distributions expressed in terms of a "continuous last-passage percolation" model on [0,1]; these are related to corresponding results for two-dimensional last-passage percolation with heavy-tailed weights obtained by Hambly and Martin.

math.PR