SearcharxivSearch

arXiv subjects

Dennis Wang

Publications and source records attributed to Dennis Wang.

11 recordsLinked to original sources

What Makes an Initial Reaction Ready for Discussion?: Multi-Persona AI Support for Stance Reflection and Writing

An initial reaction to a social or community issue can feel meaningful before it is ready to become a message: people still need to clarify the claim, anticipate audience risks, and decide how much reasoning should become visible to others. We present StanceLab, a prototype for preparing a stance before entering a discussion. The prototype compares a three-persona mode, where an Interviewer, Mentor, and Opponent respond in parallel to help users diagnose and revise a stance, with a standalone LLM mode. In a formative within-subject pilot with six participants and 12 task sessions, every session produced a short final message in the notepad. The pilot revealed two design requirements: persona roles should diagnose useful blind spots or objections, and parallel responses need coordination support. We propose a future diagnosis-and-writing workflow that turns persona-based reflection into selective, audience-aware final messages.

cs.HC

OpenMHC: Accelerating the Science of Wearable Foundation Models

Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. However, the largest wearable datasets are not publicly available for research, and leading wearable foundation models trained on such datasets are rarely open-weight or come with reproducible training code. To accelerate open science in wearable health, we release OpenMyHeartCounts (OpenMHC), the largest and most comprehensive broadly accessible wearable health dataset to date, released to qualified researchers, alongside open-source implementations of recent wearable foundation models. OpenMHC, derived from over a decade of data collected through the My Heart Counts study app, includes >60 million hours of wearable data across 19 sensor channels (e.g., step count, heart rate, sleep, workouts) and up to 169 linked variables, including health, lifestyle, mood, and behavior from 11,894 consenting participants. Furthermore, we introduce a unified, open benchmark that enables standardized comparison of wearable health models across three tracks: health and behavior downstream prediction, multivariate data imputation, and time-series forecasting. We benchmark classical methods alongside recent wearable and multivariate time series foundation models. By releasing data under broad research access, alongside open-source code and model weights, at this unprecedented scale, we aim to democratize wearable health AI research and enable the community to drive open progress in this domain.

cs.LG

Structured Transfer Learning for Survival Risk Stratification in Data-Sparse Clinical Cohorts

Background: Survival prediction models are often less reliable in clinical groups with limited sample sizes or few outcome events. Target-only models may be unstable, whereas models from larger cohorts may transfer poorly when risk-factor effects differ across populations. We evaluated whether structured transfer learning can improve survival risk stratification in data-sparse cohorts while allowing cohort-specific adaptation. Methods: We developed the COhort-shared Rank-rEduced Cox model (CORE-Cox), a two-stage framework for multi-outcome survival prediction. CORE-Cox learns shared risk-factor patterns across related outcomes in a larger source cohort via a low-rank Cox coefficient structure, then adapts these patterns to a smaller target cohort through regularized residual correction. We evaluated CORE-Cox in UK Biobank (White source, n=150,093; Asian target, n=2,534) and MIMIC-IV (White ICU source, n=15,997; Asian ICU target, n=672), comparing against target-only Cox, penalized Cox, low-rank multi-task, naive pooling, direct transfer, and single-outcome residual transfer under repeated nested cross-validation. Results: CORE-Cox achieved best or near-best discrimination across most outcomes. Mean C-index improved from 0.733 to 0.766 in UK Biobank and from 0.628 to 0.658 in MIMIC-IV, with gains in eight of nine outcomes. CORE-Cox also improved top-15% risk enrichment, with hazard-ratio estimates typically intermediate between source-only and target-only models. Discussion: CORE-Cox offers an interpretable transfer-learning framework for survival risk stratification in data-sparse cohorts, combining shared cross-outcome structure with cohort-specific adaptation. Further validation is needed before use in calibrated absolute-risk prediction or clinical decision-making.

stat.ME

The Capacity to Care: Designing Social Technology for Sustained Engagement With Societal Challenges

People care about climate change, injustice, and humanitarian crises. The challenge is not apathy but capacity: sustained engagement with large-scale problems is psychologically costly, and social media architecture often amplifies awareness while providing few pathways to meaningful action. The result is rising distress, overwhelm, and disengagement -- particularly among young people who encounter global suffering through platforms designed for attention capture rather than constructive response. This workshop examines how social technology design shapes the conditions for sustained engagement with societal challenges. Drawing on Tronto's care ethics framework and research in moral psychology and platform studies, we ask why caring at scale is difficult and how social media can both exacerbate and potentially mitigate this difficulty. Tronto's framework shows that good care requires more than awareness: it demands responsibility, competence, and community. Dominant social media architectures stall the caring process at its earliest phase. We invite researchers and designers to identify platform designs that deplete or support the capacity to care, and to develop design directions for sustainable care: engagement that people can maintain over time without burning out.

cs.HC

Are Large Language Models Effective Knowledge Graph Constructors?

Knowledge graphs (KGs) are widely used in knowledge-intensive applications, yet it remains unclear how effectively current large language models (LLMs) can construct document-grounded KGs in a zero-shot, schema-free setting without relying on complex task-specific frameworks. We introduce Detail-to-Abstract Hierarchical Knowledge Graph (D2A-HKG) construction framework, which decomposes KG construction into three stages: initial extraction, splitting, and abstraction, and evaluates the resulting graphs from both semantic and structural perspectives. Using seven frontier LLMs, we benchmark zero-shot KG construction on CMW-Lit, a dataset derived from published paediatric research articles on children's mental well-being. CMW-Lit provides a challenging test bed due to its heterogeneous evidence, interconnected factors, and complex, statistically qualified relationships. Our results show that state-of-the-art LLMs can generally produce relevant and document-faithful triples with limited hallucination, while exhibiting substantially different extraction behaviors across the construction stages. These findings provide empirical insight into the strengths and limitations of frontier LLMs for direct knowledge graph construction. We further release CMW-Lit and the resulting knowledge graphs as resources for future research, with the generated graphs providing a strong foundation for expert refinement and downstream knowledge-intensive applications.

cs.CL

Prospective Prediction of Body Mass Index Trajectories using Multi-task Gaussian Processes

Clinicians often investigate the body mass index (BMI) trajectories of children to assess their growth with respect to their peers, as well as to anticipate future growth and disease risk. While retrospective modelling of BMI trajectories has been an active area of research, prospective prediction of continuous BMI trajectories from historical growth data has not been well investigated. Using weight and height measurements from birth to age 10 years from a longitudinal mother-offspring cohort, we leveraged a multi-task Gaussian processes model, called MagmaClust, to derive probabilistic predictions for BMI trajectories over various forecasting periods. Experiments were conducted to evaluate the accuracy, sensitivity to missing values, and number of clusters. The results were compared with cubic B-spline regression and a parametric Jenss-Bayley mixed effects model. A downstream tool computing individual overweight probabilities was also proposed and evaluated. In all experiments, MagmaClust outperformed conventional models in prediction accuracy while correctly calibrating uncertainty regardless of the missing data amount (up to 90\% missing) or the forecasting period (from 2 to 8 years in the future). Moreover, the overweight probabilities computed from MagmaClust's uncertainty quantification exhibited high specificity ($0.94$ to $0.96$) and accuracy ($0.86$ to $0.94$) in predicting the 10-year overweight status even from age 2 years. MagmaClust provides a probabilistic non-parametric framework to prospectively predict BMI trajectories, which is robust to missing values and outperforms conventional BMI trajectory modelling approaches. It also clusters individuals to identify typical BMI patterns (early peak, adiposity rebounds) during childhood. Overall, we demonstrated its potential to anticipate BMI evolution throughout childhood, allowing clinicians to implement prevention strategies.

stat.AP

Longitudinal prediction of DNA methylation to forecast epigenetic outcomes

Interrogating the evolution of biological changes at early stages of life requires longitudinal profiling of molecules, such as DNA methylation, which can be challenging with children. We introduce a probabilistic and longitudinal machine learning framework based on multi-mean Gaussian processes (GPs), accounting for individual and gene correlations across time. This method provides future predictions of DNA methylation status at different individual ages while accounting for uncertainty. Our model is trained on a birth cohort of children with methylation profiled at ages 0-4, and we demonstrated that the status of methylation sites for each child can be accurately predicted at ages 5-7. We show that methylation profiles predicted by multi-mean GPs can be used to estimate other phenotypes, such as epigenetic age, and enable comparison to other health measures of interest. This approach encourages epigenetic studies to move towards longitudinal design for investigating epigenetic changes during development, ageing and disease progression.

q-bio.GN

Fast Parallel Hypertree Decompositions in Logarithmic Recursion Depth

Modern trends in data collection are bringing current mainstream techniques for database query processing to their limits. Consequently, various novel approaches for efficient query processing are being actively studied. One such approach is based on hypertree decompositions (HDs), which have been shown to carry great potential to process complex queries more efficiently and with stronger theoretical guarantees. However, using HDs for query execution relies on the difficult task of computing decompositions of the query structure, which guides the efficient execution of the query. From theoretical results we know that the performance of purely sequential methods is inherently limited, yet the problem is susceptible to parallelisation. In this paper we propose the first algorithm for computing hypertree decompositions that is well-suited for parallelisation. The proposed algorithm log-k-decomp requires only a logarithmic number of recursion levels and additionally allows for highly parallelised pruning of the search space by restriction to balanced separators. We provide detailed experimental evaluation over the HyperBench benchmark and demonstrate that our approach is highly effective especially for complex queries.

cs.DB

Complete Strain Mapping of Nanosheets of Tantalum Disulfide

Quasi-two-dimensional (quasi-2D) materials hold promise for future electronics because of their unique band structures that result in electronic and mechanical properties sensitive to crystal strains in all three dimensions. Quantifying crystal strain is a prerequisite to correlating it with the performance of the device, and calls for high resolution but spatially resolved rapid characterization methods. Here we show that using fly-scan nano X-ray diffraction we can accomplish a tensile strain sensitivity below 0.001% with a spatial resolution of better than 80 nm over a spatial extent of 100 $μ$m on quasi 2D flakes of 1T-TaS2. Coherent diffraction patterns were collected from a $\sim$ 100 nm thick sheet of 1T-TaS2 by scanning 12keV focused X-ray beam across and rotating the sample. We demonstrate that the strain distribution around micron and sub-micron sized 'bubbles' that are present in the sample may be reconstructed from these images. The experiments use state of the art synchrotron instrumentation, and will allow rapid and non-intrusive strain mapping of thin film samples and electronic devices based on quasi 2D materials.

cond-mat.mtrl-sci

Atomic scale characterization of graphene p-n junctions for electron-optical applications

Graphene p-n junctions offer a potentially powerful approach towards controlling electron trajectories via collimation and focusing in ballistic solid-state devices. The ability of p-n junctions to control electron trajectories depends crucially on the doping profile and roughness of the junction. Here, we use four-probe scanning tunneling microscopy and spectroscopy (STM/STS) to characterize two state-of-the-art graphene p-n junction geometries at the atomic scale, one with CMOS polySi gates and another with naturally cleaved graphite gates. Using spectroscopic imaging, we characterize the local doping profile across and along the p-n junctions. We find that realistic junctions exhibit non-ideality both in their geometry as well as in the doping profile across the junction. We show that the geometry of the junction can be improved by using the cleaved edge of van der Waals metals such as graphite to define the junction. We quantify the geometric roughness and doping profiles of junctions experimentally and use these parameters in Nonequilibrium Green's Function based simulations of focusing and collimation in these realistic junctions. We find that for realizing Veselago focusing, it is crucial to minimize lateral interface roughness which only natural graphite gates achieve, and to reduce junction width, in which both devices under investigation underperform. We also find that carrier collimation is currently limited by the non-linearity of the doping profile across the junction. Our work provides benchmarks of the current graphene p-n junction quality and provides guidance for future improvements.

cond-mat.mes-hall

Absence of a Band Gap at Interface of a Metal and Highly Doped Monolayer $MoS_2$

High quality electrical contact to semiconducting transition metal dichalcogenides (TMDCs) such as $MoS_2$ is key to unlocking their unique electronic and optoelectronic properties for fundamental research and device applications. Despite extensive experimental and theoretical efforts reliable ohmic contact to doped TMDCs remains elusive and would benefit from a better understanding of the underlying physics of the metal-TMDC interface. Here we present measurements of the atomic-scale energy band diagram of junctions between various metals and heavily doped monolayer $MoS_2$ using ultra-high vacuum scanning tunneling microscopy (UHV-STM). Our measurements reveal that the electronic properties of these junctions are dominated by 2D metal induced gap states (MIGS). These MIGS are characterized by a spatially growing measured gap in the local density of states (L-DOS) of the $MoS_2$ within 2 nm of the metal-semiconductor interface. Their decay lengths extend from a minimum of ~0.55 nm near mid gap to as long as 2 nm near the band edges and are nearly identical for Au, Pd and graphite contacts, indicating that it is a universal property of the monolayer semiconductor. Our findings indicate that even in heavily doped semiconductors, the presence of MIGS sets the ultimate limit for electrical contact.

cond-mat.mtrl-sci