SearcharxivSearch

arXiv subjects

Daniel Lawson

Publications and source records attributed to Daniel Lawson.

13 recordsLinked to original sources

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

Categorising invoices into the correct General Ledger (GL) code underpins financial reporting and tax compliance. This is a skilled accounting judgement rather than a routine task: the correct category depends subtly on the nature of the purchasing business, the vendor and the invoice text. Whilst AI is increasingly being adopted across industries to automate tasks, including invoice categorisation, implementations built on in-house small language models (SLMs) can simultaneously reduce cost and improve data security, confidentiality, and interpretability. We investigate this approach by first analysing the pre-trained embedding geometry of a small sentence transformer (SBERT) and classic SLM (DeBERTa). The sentence-embedding space of this financial corpus is globally anisotropic but composed of locally isotropic clusters, extending prior token-level findings to sentence embeddings in a financial setting, and these clusters are strongly correlated with the vendor identity. SBERT fine-tuned on a single GPU reaches 0.96 accuracy on invoice classification, above both a zero-shot LLM and a vendor identity baseline, increasing performance for smaller, challenging categories and new clients. For this important generalisation problem, SBERT reaches 0.9 F1 with roughly 100 client-specific invoices, showing that an in-house SLM implementation is promising. Combining these results with geometric analysis shows that pre-trained embedding geometry is associated with classification performance and reveals a counterintuitive finding that a structured input that would help a human reader does not improve the SLM performance.

stat.ML

Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning

While goal-conditioned behavior cloning (GCBC) methods can perform well on in-distribution training tasks, they do not necessarily generalize zero-shot to tasks that require conditioning on novel state-goal pairs, i.e. combinatorial generalization. In part, this limitation can be attributed to a lack of temporal consistency in the state representation learned by BC; if temporally correlated states are properly encoded to similar latent representations, then the out-of-distribution gap for novel state-goal pairs would be reduced. We formalize this notion by demonstrating how encouraging long-range temporal consistency via successor representations (SR) can facilitate generalization. We then propose a simple yet effective representation learning objective, $\text{BYOL-}\gamma$ for GCBC, which theoretically approximates the successor representation in the finite MDP case through self-predictive representations, and achieves competitive empirical performance across a suite of challenging tasks requiring combinatorial generalization.

cs.LG

Unsupervised Attributed Dynamic Network Embedding with Stability Guarantees

Stability for dynamic network embeddings ensures that nodes behaving the same at different times receive the same embedding, allowing comparison of nodes in the network across time. We present attributed unfolded adjacency spectral embedding (AUASE), a stable unsupervised representation learning framework for dynamic networks in which nodes are attributed with time-varying covariate information. To establish stability, we prove uniform convergence to an associated latent position model. We quantify the benefits of our dynamic embedding by comparing with state-of-the-art network representation learning methods on four real attributed networks. To the best of our knowledge, AUASE is the only attributed dynamic embedding that satisfies stability guarantees without the need for ground truth labels, which we demonstrate provides significant improvements for link prediction and node classification.

stat.ML

Differentiable Composite Neural Signed Distance Fields for Robot Navigation in Dynamic Indoor Environments

Neural Signed Distance Fields (SDFs) provide a differentiable environment representation to readily obtain collision checks and well-defined gradients for robot navigation tasks. However, updating neural SDFs as the scene evolves entails re-training, which is tedious, time consuming, and inefficient, making it unsuitable for robot navigation with limited field-of-view in dynamic environments. Towards this objective, we propose a compositional framework of neural SDFs to solve robot navigation in indoor environments using only an onboard RGB-D sensor. Our framework embodies a dual mode procedure for trajectory optimization, with different modes using complementary methods of modeling collision costs and collision avoidance gradients. The primary stage queries the robot body's SDF, swept along the route to goal, at the obstacle point cloud, enabling swift local optimization of trajectories. The secondary stage infers the visible scene's SDF by aligning and composing the SDF representations of its constituents, providing better informed costs and gradients for trajectory optimization. The dual mode procedure combines the best of both stages, achieving a success rate of 98%, 14.4% higher than baseline with comparable amortized plan time on iGibson 2.0. We also demonstrate its effectiveness in adapting to real-world indoor scenarios.

cs.RO

EarthquakeNPP: A Benchmark for Earthquake Forecasting with Neural Point Processes

For decades, classical point process models, such as the epidemic-type aftershock sequence (ETAS) model, have been widely used for forecasting the event times and locations of earthquakes. Recent advances have led to Neural Point Processes (NPPs), which promise greater flexibility and improvements over such classical models. However, the currently-used benchmark for NPPs does not represent an up-to-date challenge in the seismological community, since it contains data leakage and omits the largest earthquake sequence from the region. Additionally, initial earthquake forecasting benchmarks fail to compare NPPs with state-of-the-art forecasting models commonly used in seismology. To address these gaps, we introduce EarthquakeNPP: a benchmarking platform that curates and standardizes existing public resources: globally available earthquake catalogs, the ETAS model, and evaluation protocols from the seismology community. The datasets cover a range of small to large target regions within California, dating from 1971 to 2021, and include different methodologies for dataset generation. Benchmarking experiments, using both log-likelihood and generative evaluation metrics widely recognised in seismology, show that none of the five NPPs tested outperform ETAS. These findings suggest that current NPP implementations are not yet suitable for practical earthquake forecasting. Nonetheless, EarthquakeNPP provides a platform to foster future collaboration between the seismology and machine learning communities.

physics.geo-ph

Ultracompact programmable silicon photonics using layers of low-loss phase-change material Sb$_2$Se$_3$ of increasing thickness

High-performance programmable silicon photonic circuits are considered to be a critical part of next generation architectures for optical processing, photonic quantum circuits and neural networks. Low-loss optical phase change materials (PCMs) offer a promising route towards non-volatile free-form control of light. Here, we exploit direct-write digital patterning of waveguides using layers of the PCM Sb$_2$Se$_3$ with a thickness of up to 100 nm, demonstrating the ability to strongly increase the effect per pixel compared to previous implementations where much thinner PCM layers were used. We exploit the excellent refractive index matching between Sb$_2$Se$_3$ and silicon to achieve a low-loss hybrid platform for programmable photonics. A five-fold reduction in modulation length of a Mach-Zehnder interferometer is achieved compared to previous work using thin-film Sb$_2$Se$_3$ devices, decreased to 5 $\mu$m in this work. Application of the thicker PCM layers in direct-write digital programming of a multimode interferometer (MMI) shows a three-fold reduction of the number of programmed pixels to below 10 pixels per device. The demonstrated scaling of performance with PCM layer thickness is important for establishing the optimum working range for hybrid silicon-PCM devices and holds promise for achieving ultracompact programmable photonic circuits.

physics.optics

Optical switching beyond a million cycles of low-loss phase change material Sb$_2$Se$_3$

The development of the next generation of optical phase change technologies for integrated photonic and free-space platforms relies on the availability of materials that can be switched repeatedly over large volumes and with low optical losses. In recent years, the antimony-based chalcogenide phase-change material Sb$_2$Se$_3$ has been identified as particularly promising for a number of applications owing to good optical transparency in the near-infrared part of the spectrum and a high refractive index close to silicon. The crystallization temperature of Sb$_2$Se$_3$ of around 460 K allows switching to be achieved at moderate energies using optical or electrical control signals while providing sufficient data retention time for non-volatile storage. Here, we investigate the parameter space for optical switching of films of Sb$_2$Se$_3$ for a range of film thicknesses relevant for optical applications. By identifying optimal switching conditions, we demonstrate endurance of up to 10$^7$ cycles at reversible switching rates of 20 kHz. Our work demonstrates that the combination of intrinsic film parameters with pumping conditions is particularly critical for achieving high endurance in optical phase change applications.

physics.optics

Merging Decision Transformers: Weight Averaging for Forming Multi-Task Policies

Recent work has shown the promise of creating generalist, transformer-based, models for language, vision, and sequential decision-making problems. To create such models, we generally require centralized training objectives, data, and compute. It is of interest if we can more flexibly create generalist policies by merging together multiple, task-specific, individually trained policies. In this work, we take a preliminary step in this direction through merging, or averaging, subsets of Decision Transformers in parameter space trained on different MuJoCo locomotion problems, forming multi-task models without centralized training. We also demonstrate the importance of various methodological choices when merging policies, such as utilizing common pre-trained initializations, increasing model capacity, and utilizing Fisher information for weighting parameter importance. In general, we believe research in this direction could help democratize and distribute the process that forms multi-task robotics policies. Our implementation is available at https://github.com/daniellawson9999/merging-decision-transformers.

cs.LG

Co-learning Planning and Control Policies Constrained by Differentiable Logic Specifications

Synthesizing planning and control policies in robotics is a fundamental task, further complicated by factors such as complex logic specifications and high-dimensional robot dynamics. This paper presents a novel reinforcement learning approach to solving high-dimensional robot navigation tasks with complex logic specifications by co-learning planning and control policies. Notably, this approach significantly reduces the sample complexity in training, allowing us to train high-quality policies with much fewer samples compared to existing reinforcement learning algorithms. In addition, our methodology streamlines complex specification extraction from map images and enables the efficient generation of long-horizon robot motion paths across different map layouts. Moreover, our approach also demonstrates capabilities for high-dimensional control and avoiding suboptimal policies via policy alignment. The efficacy of our approach is demonstrated through experiments involving simulated high-dimensional quadruped robot dynamics and a real-world differential drive robot (TurtleBot3) under different types of task specifications.

cs.RO

Control Transformer: Robot Navigation in Unknown Environments through PRM-Guided Return-Conditioned Sequence Modeling

Learning long-horizon tasks such as navigation has presented difficult challenges for successfully applying reinforcement learning to robotics. From another perspective, under known environments, sampling-based planning can robustly find collision-free paths in environments without learning. In this work, we propose Control Transformer that models return-conditioned sequences from low-level policies guided by a sampling-based Probabilistic Roadmap (PRM) planner. We demonstrate that our framework can solve long-horizon navigation tasks using only local information. We evaluate our approach on partially-observed maze navigation with MuJoCo robots, including Ant, Point, and Humanoid. We show that Control Transformer can successfully navigate through mazes and transfer to unknown environments. Additionally, we apply our method to a differential drive robot (Turtlebot3) and show zero-shot sim2real transfer under noisy observations.

cs.RO

Time-resolved reversible optical switching of the ultralow-loss phase change material Sb2Se3

The antimony-based chalcogenide Sb2Se3 is a rapidly emerging material for photonic phase change applications owing to its ultra-low optical losses at telecommunication wavelengths in both crystalline and amorphous phases. Here, we investigate the dynamical response of these materials from nanoseconds to milliseconds under optical pumping conditions. We apply bichromatic pump-probe transient reflectance spectroscopy which is a widely used method to study the optical performance of optical phase change materials. Amorphous regions of several hundreds of nanometers in diameter are induced by pulsed excitation of the material using a wavelength of 488 nm above the absorption edge, while the transient reflectance is probed using a continuous wave 980 nm laser, well below the absorption edge of the material. We find vitrification dynamics in the nanosecond range and observe crystallization on millisecond time scales. These results show a large five-orders of magnitude difference in time scales between crystallization and vitrification dynamics in this material. The insights provided in this work are fundamental for the optimisation of the material family and its employment in photonic applications.

physics.optics

The species-area relationship and evolution

Models relating to the Species-Area curve are usually defined at the species level, and concerned only with ecological timescales. We examine an individual-based model of co-evolution on a spatial lattice based on the Tangled Nature model, and show that reproduction, mutation and dispersion by diffusion in an interacting system produces power-law Species-Area Relations as observed in ecological measurements at medium scales. We find that co-evolutionary habitats form, allowing high diversity levels in a spatially homogenous system, and these are maintained for exponentially increasing time when increasing system size.

q-bio.PE

Diversity as a product of interspecial interactions

We demonstrate diversification rather than optimisation for highly interacting organisms in a well mixed biological system by means of a simple model and reference to experiment, and find the cause to be the complex network of interactions formed, allowing species less well adapted to an environment to flourish by co-interaction over the `best' species. This diversification can be considered as the construction of many co-evolutionary niches by the network of interactions between species. Evidence for this comes from work with the bacteria Escherichia coli, which may coexist with their own mutants under certain conditions. Diversification only occurs above a certain threshold interaction strength, below which competitive exclusion occurs.

q-bio.PE