SearcharxivSearch

arXiv subjects

John Davis

Publications and source records attributed to John Davis.

8 recordsLinked to original sources

CogAdapt: Adapting Clinical ECG Foundation Models for Wearable Cognitive Load Assessment

Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects. Recent ECG foundation models, pre-trained on millions of clinical diagnostic ECG recordings, yet they do not apply directly to wearable devices when the sensor configuration and the task both differ. We present CogAdapt, a framework that adapts a clinical ECG foundation model to wearable cognitive load assessment. CogAdapt has two parts. LeadBridge is a learnable adapter that maps 3-lead wearable signals to a 12-lead-compatible representation. ProFine is a progressive fine-tuning strategy that unfreezes encoder layers in stages while limiting representational drift in the pre-trained model. On two public datasets (CLARE and CL-Drive) under leave-one-subject-out cross-validation, CogAdapt reaches macro-F1 of 0.626 and 0.768, improving over from-scratch baselines by 11.2 and 16.1 percentage points. The results show that a clinical ECG pretraining can support subject-independent cognitive load assessment from wearable sensors.

cs.LG

MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data

Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such as driver vigilance monitoring or automated flight deck assistance, yet two challenges persist: handling frequent data missingness from blinks and tracking failures, and efficiently modeling long-range temporal dependencies. We propose MambaGaze (Bi-Mamba), a framework that addresses these challenges through (1) XMD encoding, which augments raw features with observation masks and time-deltas to explicitly model data uncertainty, and (2) bidirectional Mamba-2, which captures temporal dependencies with linear computational complexity. Experiments on CLARE and CL-Drive datasets under leave-one-subject-out evaluation show that MambaGaze achieves 77.1% accuracy and 59.2% macro-F1 on CLARE, and 69.4% accuracy and 51.5% macro-F1 on CL-Drive, attaining the highest average LOSO macro-F1 (55.3%) across all ten compared models. Input-stream ablation indicates that log-scaled time-deltas are the strongest single channel in our setting, and combining all three XMD streams provides consistent gains of 5-20 pp macro-F1. Edge deployment benchmarks on three NVIDIA Jetson Orin platforms show real-time inference at 27-36 FPS with power consumption below 6.6 W, supporting feasibility for embedded cognitive load monitoring.

cs.LG

Google COVID-19 Vaccination Search Insights: Anonymization Process Description

This report describes the aggregation and anonymization process applied to the COVID-19 Vaccination Search Insights (published at http://goo.gle/covid19vaccinationinsights), a publicly available dataset showing aggregated and anonymized trends in Google searches related to COVID-19 vaccination. The applied anonymization techniques protect every user's daily search activity related to COVID-19 vaccinations with $(\varepsilon, \delta)$-differential privacy for $\varepsilon = 2.19$ and $\delta = 10^{-5}$.

cs.CR

Google COVID-19 Search Trends Symptoms Dataset: Anonymization Process Description (version 1.0)

This report describes the aggregation and anonymization process applied to the initial version of COVID-19 Search Trends symptoms dataset (published at https://goo.gle/covid19symptomdataset on September 2, 2020), a publicly available dataset that shows aggregated, anonymized trends in Google searches for symptoms (and some related topics). The anonymization process is designed to protect the daily symptom search activity of every user with $\varepsilon$-differential privacy for $\varepsilon$ = 1.68.

cs.CR

Google COVID-19 Community Mobility Reports: Anonymization Process Description (version 1.1)

This document describes the aggregation and anonymization process applied to the initial version of Google COVID-19 Community Mobility Reports (published at http://google.com/covid19/mobility on April 2, 2020), a publicly available resource intended to help public health authorities understand what has changed in response to work-from-home, shelter-in-place, and other recommended policies aimed at flattening the curve of the COVID-19 pandemic. Our anonymization process is designed to ensure that no personal data, including an individual's location, movement, or contacts, can be derived from the resulting metrics. The high-level description of the procedure is as follows: we first generate a set of anonymized metrics from the data of Google users who opted in to Location History. Then, we compute percentage changes of these metrics from a baseline based on the historical part of the anonymized metrics. We then discard a subset which does not meet our bar for statistical reliability, and release the rest publicly in a format that compares the result to the private baseline.

cs.CR

Learning Over Dirty Data Without Cleaning

Real-world datasets are dirty and contain many errors. Examples of these issues are violations of integrity constraints, duplicates, and inconsistencies in representing data values and entities. Learning over dirty databases may result in inaccurate models. Users have to spend a great deal of time and effort to repair data errors and create a clean database for learning. Moreover, as the information required to repair these errors is not often available, there may be numerous possible clean versions for a dirty database. We propose DLearn, a novel relational learning system that learns directly over dirty databases effectively and efficiently without any preprocessing. DLearn leverages database constraints to learn accurate relational models over inconsistent and heterogeneous data. Its learned models represent patterns over all possible clean instances of the data in a usable form. Our empirical study indicates that DLearn learns accurate models over large real-world databases efficiently.

cs.DB

Distribution and Structure of Matter in and around Galaxies

Understanding the origins and distribution of matter in the Universe is one of the most important quests in physics and astronomy. Themes range from astro-particle physics to chemical evolution in the Galaxy to cosmic nucleosynthesis and chemistry in an anticipation of a full account of matter in the Universe. Studies of chemical evolution in the early Universe will answer questions about when and where the majority of metals were formed, how they spread and why they appar today as they are. The evolution of matter in our Universe cannot be characterized as a simple path of development. In fact the state of matter today tells us that mass and matter is under constant reformation through on-going star formation, nucleosynthesis and mass loss on stellar and galactic scales. X-ray absorption studies have evolved in recent years into powerful means to probe the various phases of interstellar and intergalactic media. Future observatories such as IXO and Gen-X will provide vast new opportunities to study structure and distribution of matter with high resolution X-ray spectra. Specifically the capabilities of the soft energy gratings with a resolution of R=3000 onboard IXO will provide ground breaking determinations of element abundance, ionization structure, and dispersion velocities of the interstellar and intergalactic media of our Galaxy and the Local Group

astro-ph.HE

Structure and Evolution of Pre-Main Sequence Stars

Low-mass pre-main sequence (PMS) stars are strong and variable X-ray emitters, as has been well established by EINSTEIN and ROSAT observatories. It was originally believed that this emission was of thermal nature and primarily originated from coronal activity (magnetically confined loops, in analogy with Solar activity) on contracting young stars. Broadband spectral analysis showed that the emission was not isothermal and that elemental abundances were non-Solar. The resolving power of the Chandra and XMM X-ray gratings spectrometers have provided the first, tantalizing details concerning the physical conditions such as temperatures, densities, and abundances that characterize the X-ray emitting regions of young star. These existing high resolution spectrometers, however, simply do not have the effective area to measure diagnostic lines for a large number of PMS stars over required to answer global questions such as: how does magnetic activity in PMS stars differ from that of main sequence stars, how do they evolve, what determines the population structure and activity in stellar clusters, and how does the activity influence the evolution of protostellar disks. Highly resolved (R>3000) X-ray spectroscopy at orders of magnitude greater efficiency than currently available will provide major advances in answering these questions. This requires the ability to resolve the key diagnostic emission lines with a precision of better than 100 km/s.

astro-ph.HE