SearcharxivSearch

arXiv subjects

John Scoville

Publications and source records attributed to John Scoville.

13 recordsLinked to original sources

Thomson: Continual Learning of Frontier Models for SovereignAI

The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an organisation's capability to independently build, deploy and govern AI use), but offers little concrete advice on how this can be achieved in the short term under a diversity of funding settings. We argue that frontier performance is achievable by a wide range of institutions through Continual Learning on readily available open-weight models. Unlike limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation of a frozen model, our approach exploits a modern mid- & post-training stack while introducing safeguards that preserve both plasticity and stability at each stage, making the minimal number of high-impact interventions on the parameters. This yields gains comparable to those typically seen across multiple successive model generations, at compute and personnel budgets substantially lower than commonly thought, making ownership of large parts of the SovereignAI stack (model, tool infrastructure, values & data privacy) viable for far more actors. We demonstrate this with Thomson, a general-purpose frontier model trained with an enhanced focus on high-stakes professional work. Thomson performs competitively with recent frontier models across agentic tasks, safety, legal, tax & multilingualism, and large-scale Deep Research. Evaluations show a distinctive $\pi$-shaped pattern: distinct improvements across a wide range of capabilities, including those not explicitly targeted, while almost completely eliminating the forgetting problem common to narrow domain adaptation.

cs.AI

Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models

Recent meta-reasoning frameworks improve LLM reasoning by wrapping chain-of-thought generation in an iterative control loop, allowing more effective backtracking, termination of reasoning loops, and injection of promising reasoning patterns, among other strategy adjustments. Despite promising results, methods often rely on backward-looking reward functions, utilize coarse search actions, or require additional reasoning controller training requiring many-shot supervision. We introduce Cognitive Demand Steering (CDS), a training-free meta-reasoning framework equipped with residual demand assessment: at each step, an LLM-based progress evaluator characterizes the residual reasoning required to arrive at a solution rather than merely evaluating the previous step. This allows a meta-controller to select reasoning interventions comprising both general-purpose exemplars and actions (e.g., general guidance for quantitative reasoning) that directly tackle this forward-looking demand signal. This shift eliminates the need for any trained component while enabling zero-shot transfer across models and tasks with no adaptation. Rather than relying on coarse characterizations, we employ cognitive scales to both design interventions as well as profile initial problem complexity and residual demand signal over 16 dimensions motivated by cognitive science (e.g., attention and scan, learning and abstraction, spatio-physical reasoning), giving the controller a fine-grained vocabulary for diagnosing. Averaged across three frontier LLMs and six reasoning benchmarks, CDS improves accuracy by $21.9\%$ over direct calls and $9\%$ over standard CoT reasoning, with the largest gains on difficult mathematics and coding tasks.

cs.AI

A Little Confidence Goes a Long Way

We introduce a group of related methods for binary classification tasks using probes of the hidden state activations in large language models (LLMs). Performance is on par with the largest and most advanced LLMs currently available, but requiring orders of magnitude fewer computational resources and not requiring labeled data. This approach involves translating class labels into a semantically rich description, spontaneous symmetry breaking of multilayer perceptron probes for unsupervised learning and inference, training probes to generate confidence scores (prior probabilities) from hidden state activations subject to known constraints via entropy maximization, and selecting the most confident probe model from an ensemble for prediction. These techniques are evaluated on four datasets using five base LLMs.

cs.LG

Earthquake precursors in the light of peroxy defects theory: critical review of systematic observations

The starting point of the present review is to acknowledge that there are innumerable reports of non-seismic types of earthquake precursory phenomena that are intermittent and seem not to occur systematically, while associated reports are not widely accepted by the geoscience community at large because no one could explain their origins. We review a unifying theory for a solid-state mechanism, based on decades of research bridging semi-conductor physics, chemistry and rock physics. A synthesis has emerged that all pre-earthquake phenomena could trace back to one fundamental physical process: the activation of electronic charges (electrons and positive holes) in rocks subjected to ever-increasing tectonic stresses prior to any major seismic activity, via the rupture of peroxy bonds. In the second part of the review, we critically examine satellite and ground station data, recorded before past large earthquakes, as they have been claimed to provide evidence that precursory signals tend to become measurable days, sometimes weeks before the disasters. We review some of the various phenomena that can be directly predicted by the peroxy defect theory , namely, radon gas emanations, corona discharges, thermal infrared emissions, air ionization, ion and electron content in the ionosphere, and electro-magnetic anomalies. Our analysis demonstrates the need for further systematic investigations, in particular with strong continuous statistical testing of the relevance and confidence of the precursors. Only then, the scientific community will be able to assess and improve the performance of earthquake forecasts.

physics.geo-ph

Stochastic Time-Series Spectroscopy

Spectroscopically measuring low levels of non-equilibrium phenomena (e.g. emission in the presence of a large thermal background) can be problematic due to an unfavorable signal-to-noise ratio. An approach is presented to use time-series spectroscopy to separate non-equilibrium quantities from slowly varying equilibria. A stochastic process associated with the non-equilibrium part of the spectrum is characterized in terms of its central moments or cumulants, which may vary over time. This parameterization encodes information about the non-equilibrium behavior of the system. Stochastic time-series spectroscopy (STSS) can be implemented at very little expense in many settings since a series of scans are typically recorded in order to generate a low-noise averaged spectrum. Higher moments or cumulants may be readily calculated from this series, enabling the observation of quantities that would be difficult or impossible to determine from an average spectrum or from prinicipal components analysis (PCA). This method is more scalable than PCA, having linear time complexity, yet it can produce comparable or superior results, as shown in example applications. One example compares an STSS-derived CO$_2$ bending mode to a standard reference spectrum and the result of PCA. A second example shows that STSS can reveal conditions of stress in rocks, a scenario where traditional methods such as PCA are inadequate. This allows spectral lines and non-equilibrium behavior to be precisely resolved. A relationship between 2nd order STSS and a time-varying form of PCA is considered. Although the possible applications of STSS have not been fully explored, it promises to reveal information that previously could not be easily measured, possibly enabling new domains of spectroscopy and remote sensing.

physics.geo-ph

Paradox of Peroxy Defects and Positive Holes in Rocks Part II: Outflow of Electric Currents from Stressed Rocks

Understanding the electrical properties of rocks is of fundamental interest. We report on currents generated when stresses are applied. Loading the center of gabbro tiles, 30x30x0.9 cm$^3$, across a 5 cm diameter piston, leads to positive currents flowing from the center to the unstressed edges. Changing the constant rate of loading over 5 orders of magnitude from 0.2 kPa/s to 20 MPa/s produces positive currents, which start to flow already at low stress levels, <5 MPa. The currents increase as long as stresses increase. At constant load they flow for hours, days, even weeks and months, slowly decreasing with time. When stresses are removed, they rapidly disappear but can be made to reappear upon reloading. These currents are consistent with the stress-activation of peroxy defects, such as O$_3$Si-OO-SiO$_3$, in the matrix of rock-forming minerals. The peroxy break-up leads to positive holes h$^{\bullet}$, i.e. electronic states associated with O$^-$ in a matrix of O$^{2-}$, plus electrons, e'. Propagating along the upper edge of the valence band, the holes are able to flow from stressed to unstressed rock, traveling fast and far by way of a phonon-assisted electron hopping mechanism using energy levels at the upper edge of the valence band. Impacting the tile center leads to h$^{\bullet}$ pulses, 4-6 ms long, flowing outward at ~100 m/sec at a current equivalent to 1-2 x 10$^9$ A/km$^3$. Electrons, trapped in the broken peroxy bonds, are also mobile, but only within the stressed volume.

physics.geo-ph

A Distributed Magnetometer Network

Various possiblities for a distributed magnetometer network are considered. We discuss strategies such as croudsourcing smartphone magnetometer data, the use of trees as magnetometers, and performing interferometry using magnetometer arrays to synthesize the magnetometers into a large-scale low-frequency radio telescope. Geophysical and other applications of such a network are discussed.

physics.geo-ph

Pre-earthquake Magnetic Pulses

A semiconductor model of rocks is shown to describe unipolar magnetic pulses, a phenomenon that has been observed prior to earthquakes. These pulses are observable because their extremely long wavelength allows them to pass through the Earth's crust. Interestingly, the source of these pulses may be triangulated to pinpoint locations where stress is building deep within the crust. We couple a semiconductor drift-diffusion model to a magnetic field in order to describe the electromagnetic effects associated with electrical currents flowing within rocks. The resulting system of equations is solved numerically and it is seen that a volume of rock may act as a diode that produces transient currents when it switches bias. These unidirectional currents are expected to produce transient unipolar magnetic pulses similar in form, amplitude, and duration to those observed before earthquakes, and this suggests that the pulses could be the result of geophysical semiconductor processes.

physics.geo-ph

Fast Autocorrelated Context Models for Data Compression

A method is presented to automatically generate context models of data by calculating the data's autocorrelation function. The largest values of the autocorrelation function occur at the offsets or lags in the bitstream which tend to be the most highly correlated to any particular location. These offsets are ideal for use in predictive coding, such as predictive partial match (PPM) or context-mixing algorithms for data compression, making such algorithms more efficient and more general by reducing or eliminating the need for ad-hoc models based on particular types of data. Instead of using the definition of the autocorrelation function, which considers the pairwise correlations of data requiring O(n^2) time, the Weiner-Khinchin theorem is applied, quickly obtaining the autocorrelation as the inverse Fast Fourier transform of the data's power spectrum in O(n log n) time, making the technique practical for the compression of large data objects. The method is shown to produce the highest levels of performance obtained to date on a lossless image compression benchmark.

cs.IT

Bounding Lossy Compression using Lossless Codes at Reduced Precision

An alternative approach to two-part 'critical compression' is presented. Whereas previous results were based on summing a lossless code at reduced precision with a lossy-compressed error or noise term, the present approach uses a similar lossless code at reduced precision to establish absolute bounds which constrain an arbitrary lossy data compression algorithm applied to the original data.

cs.MM

Critical Data Compression

A new approach to data compression is developed and applied to multimedia content. This method separates messages into components suitable for both lossless coding and 'lossy' or statistical coding techniques, compressing complex objects by separately encoding signals and noise. This is demonstrated by compressing the most significant bits of data exactly, since they are typically redundant and compressible, and either fitting a maximally likely noise function to the residual bits or compressing them using lossy methods. Upon decompression, the significant bits are decoded and added to a noise function, whether sampled from a noise model or decompressed from a lossy code. This results in compressed data similar to the original. For many test images, a two-part image code using JPEG2000 for lossy coding and PAQ8l for lossless coding produces less mean-squared error than an equal length of JPEG2000. Computer-generated images typically compress better using this method than through direct lossy coding, as do many black and white photographs and most color photographs at sufficiently high quality levels. Examples applying the method to audio and video coding are also demonstrated. Since two-part codes are efficient for both periodic and chaotic data, concatenations of roughly similar objects may be encoded efficiently, which leads to improved inference. Applications to artificial intelligence are demonstrated, showing that signals using an economical lossless code have a critical level of redundancy which leads to better description-based inference than signals which encode either insufficient data or too much detail.

cs.IT

On Universal Complexity Measures

We relate the computational complexity of finite strings to universal representations of their underlying symmetries. First, Boolean functions are classified using the universal covering topologies of the circuits which enumerate them. A binary string is classified as a fixed point of its automorphism group; the irreducible representation of this group is the string's universal covering group. Such a measure may be used to test the quasi-randomness of binary sequences with regard to first-order set membership. Next, strings over general alphabets are considered. The complexity of a general string is given by a universal representation which recursively factors the codeword number associated with a string. This is the complexity of the representation recursively decoding a Godel number having the value of the string; the result is a tree of prime numbers which forms a universal representation of the string's group symmetries.

cs.IT

On Macroscopic Complexity and Perceptual Coding

The theoretical limits of 'lossy' data compression algorithms are considered. The complexity of an object as seen by a macroscopic observer is the size of the perceptual code which discards all information that can be lost without altering the perception of the specified observer. The complexity of this macroscopically observed state is the simplest description of any microstate comprising that macrostate. Inference and pattern recognition based on macrostate rather than microstate complexities will take advantage of the complexity of the macroscopic observer to ignore irrelevant noise.

cs.IT