SearcharxivSearch

arXiv subjects

Jia Tang

Publications and source records attributed to Jia Tang.

9 recordsLinked to original sources

PocketPPD: Screening for Postpartum Depression Risk Using Passive Smartphone Sensing

Postpartum depression (PPD) is a serious perinatal mental health condition affecting approximately 20% of new mothers worldwide. Common screening approaches for PPD, such as self-report questionnaires and active digital logs, rely heavily on user input and thus impose a substantial burden on participants, limiting their feasibility for long-term use. Recent passive mobile sensing (PMS) approaches have enabled low-burden detection of depressive symptoms using machine learning methods with multi-modal sensor data from off-the-shelf mobile devices including smartphones. However, the postpartum period entails distinct behavioral patterns, raising uncertainty about whether sensing-based indicators for general depression and mental disorders generalize to PPD. To address this gap, we propose PocketPPD, a PMS-based PPD screening method that detects PPD risk using maternal contextual features, such as disruptions in behavioral rhythms and shifts in stability, collected through a smartphone. In our exploratory four-week feasibility study with 61 postpartum women, the PMS-only model achieved an AUC of 0.75, while the best-performing model, integrating PMS-oriented data and self-report features, achieved an AUC of 0.83. Moreover, we find that morning and late-night routine volatility ranks among the top digital biomarkers, dynamically moderated by maternal contexts such as infant developmental stage and employment status. This work provides empirical evidence for low-burden PPD risk screening and our findings lay the groundwork for continuous perinatal mental health monitoring.

cs.HC

PriorZero: Bridging Language Priors and World Models for Decision Making

Leveraging the rich world knowledge of Large Language Models (LLMs) to enhance Reinforcement Learning (RL) agents offers a promising path toward general intelligence. However, a fundamental prior-dynamics mismatch hinders existing approaches: static LLM knowledge cannot directly adapt to the complex transition dynamics of long-horizon tasks. Using LLM priors as fixed policies limits exploration diversity, as the prior is blind to environment-specific dynamics; while end-to-end fine-tuning suffers from optimization instability and credit assignment issues. To bridge this gap, we propose PriorZero, a unified framework that integrates LLM-derived conceptual priors into world-model-based planning through a decoupled rollout-training design. During rollout, a novel root-prior injection mechanism incorporates LLM priors exclusively at the root node of Monte Carlo Tree Search (MCTS), focusing search on semantically promising actions while preserving the world model's deep lookahead capability. During training, PriorZero decouples world-model learning from LLM adaptation: the world model is continuously refined on interaction data to jointly improve its dynamics, policy, and value predictions, its value estimates are then leveraged to provide fine-grained credit assignment signals for stable LLM fine-tuning via alternating optimization. Experiments across diverse benchmarks, including text-based adventure games in Jericho and instruction-following gridworld tasks in BabyAI, demonstrate that PriorZero consistently improves both exploration efficiency and asymptotic performance, establishing a promising framework for LLM-empowered decision-making. Our code is available at https://github.com/opendilab/LightZero.

cs.LG

OSI: One-step Inversion Excels in Extracting Diffusion Watermarks

Watermarking is an important mechanism for provenance and copyright protection of diffusion-generated images. Training-free methods, exemplified by Gaussian Shading, embed watermarks into the initial noise of diffusion models with negligible impact on the quality of generated images. However, extracting this type of watermark typically requires multi-step diffusion inversion to obtain precise initial noise, which is computationally expensive and time-consuming. To address this issue, we propose One-step Inversion (OSI), a significantly faster and more accurate method for extracting Gaussian Shading style watermarks. OSI reformulates watermark extraction as a learnable sign classification problem, which eliminates the need for precise regression of the initial noise. Then, we initialize the OSI model from the diffusion backbone and finetune it on synthesized noise-image pairs with a sign classification objective. In this manner, the OSI model is able to accomplish the watermark extraction efficiently in only one step. Our OSI substantially outperforms the multi-step diffusion inversion method: it is 20x faster, achieves higher extraction accuracy, and doubles the watermark payload capacity. Extensive experiments across diverse schedulers, diffusion backbones, and cryptographic schemes consistently show improvements, demonstrating the generality of our OSI framework.

cs.CV

Global Pre-fixing, Local Adjusting: A Simple yet Effective Contrastive Strategy for Continual Learning

Continual learning (CL) involves acquiring and accumulating knowledge from evolving tasks while alleviating catastrophic forgetting. Recently, leveraging contrastive loss to construct more transferable and less forgetful representations has been a promising direction in CL. Despite advancements, their performance is still limited due to confusion arising from both inter-task and intra-task features. To address the problem, we propose a simple yet effective contrastive strategy named \textbf{G}lobal \textbf{P}re-fixing, \textbf{L}ocal \textbf{A}djusting for \textbf{S}upervised \textbf{C}ontrastive learning (GPLASC). Specifically, to avoid task-level confusion, we divide the entire unit hypersphere of representations into non-overlapping regions, with the centers of the regions forming an inter-task pre-fixed \textbf{E}quiangular \textbf{T}ight \textbf{F}rame (ETF). Meanwhile, for individual tasks, our method helps regulate the feature structure and form intra-task adjustable ETFs within their respective allocated regions. As a result, our method \textit{simultaneously} ensures discriminative feature structures both between tasks and within tasks and can be seamlessly integrated into any existing contrastive continual learning framework. Extensive experiments validate its effectiveness.

cs.LG

One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning

In heterogeneous multi-task decision-making, tasks not only exhibit diverse observation and action spaces but also vary substantially in their underlying complexities. While conventional multi-task world models like UniZero excel in single-task settings, we find that when handling a broad and diverse suite of tasks, gradient conflicts and the loss of model plasticity often constrain their sample efficiency. In this work, we address these challenges from two complementary perspectives: the single learning iteration and the overall learning process. First, to mitigate the gradient conflicts, we systematically investigate key architectural designs for extending UniZero. Our investigation identifies a Mixture-of-Experts (MoE) architecture as the most effective approach. We demonstrate, both theoretically and empirically, that this architecture alleviates gradient conflicts by routing task-specific representations to specialized sub-networks. This finding leads to our proposed model, \textit{ScaleZero}. Second, to dynamically allocate model capacity throughout the learning process, we introduce an online Dynamic Parameter Scaling (DPS) strategy. This strategy progressively integrates LoRA adapters in response to task-specific progress, enabling adaptive knowledge retention and parameter expansion. Evaluations on a diverse set of standard benchmarks (Atari, DMC, Jericho) demonstrate that ScaleZero, utilizing solely online reinforcement learning with one model, performs on par with specialized single-task agents. With the DPS strategy, it remains competitive while using just 71.5% of the environment interactions. These findings underscore the potential of ScaleZero for effective multi-task planning. Our code is available at \textcolor{magenta}{https://github.com/opendilab/LightZero}.

cs.LG

Frequency conversion between optical and microwave photons in non-Markovian environments

In this paper, we propose a scheme for frequency conversion between optical photons and microwave photons in non-Markovian environments using both magnetic and mechanical excitations as intermediate media. When the frequencies of optical photons, magnons, phonons, and microwave photons resonance, the conversion efficiency can be made close to reach 98.76$\%$ by adjusting the defined complex cooperativities, while in the case of Markovian, the conversion efficiency is 90.44$\%$. By controlling the environmental spectral widths, the efficiency of frequency conversion exhibits a transition from Markovian regimes to non-Markovian regimes. This transformation simultaneously improves frequency conversion efficiency and conversion bandwidth, which is due to the excitation backflow generated by the interaction between the system and the non-Markovian environments. In the case, when the optical pump power in the non-Markovian regimes are of a large order of magnitude, the conversion bandwidth can be increased, but at the cost of reduced conversion efficiency. Our scheme improves the frequency conversion efficiency and bandwidth between optical photons and microwave photons, breaking the limitations of frequency conversion in Markovian environments and providing a new approach for long-distance quantum communication research of other non-Markovian quantum systems in quantum optics.

physics.optics

Design and Implementation of a Psychiatry Resident Training System Based on Large Language Models

Mental disorders have become a significant global public health issue, while the shortage of psychiatrists and inefficient training systems severely hinder the accessibility of mental health services. This paper designs and implements an artificial intelligence-based training system for psychiatrists. By integrating technologies such as large language models, knowledge graphs, and expert systems, the system constructs an intelligent and standardized training platform. It includes six functional modules: case generation, consultation dialogue, examination prescription, diagnostic decision-making, integrated traditional Chinese and Western medicine prescription, and expert evaluation, providing comprehensive support from clinical skill training to professional level assessment.The system adopts a B/S architecture, developed using the Vue.js and Node.js technology stack, and innovatively applies deep learning algorithms for case generation and doctor-patient dialogue. In a clinical trial involving 60 psychiatrists at different levels, the system demonstrated excellent performance and training outcomes: system stability reached 99.95%, AI dialogue accuracy achieved 96.5%, diagnostic accuracy reached 92.5%, and user satisfaction scored 92.3%. Experimental data showed that doctors using the system improved their knowledge mastery, clinical thinking, and diagnostic skills by 35.6%, 28.4%, and 23.7%, respectively.The research results provide an innovative solution for improving the efficiency of psychiatrist training and hold significant importance for promoting the standardization and scalability of mental health professional development.

cs.CY

Spin Josephson effects of spin-orbit-coupled Bose-Einstein condensates in a non-Hermitian double well

In this paper, we investigate the spin and tunneling dynamics of a spin-orbit-coupled noninteracting Bose-Einstein condensate in a periodically driven non-Hermitian double-well potential. Under high-frequency driving, we obtain the effective time-averaged Hamiltonian by using the standard time-averaging method, and analytically calculate the Floquet quasienergies, revealing that the parity-time (PT)-breaking phase transition appears even for arbitrarily small non-Hermitian parameters when the spin-orbit coupling strength takes half-integer value, irrespective of the values of other parameters used. When the system is PT-symmetric with balanced gain and loss, we find numerically and analytically that in the broken PT-symmetric regions, there will exist the net spin current together with a vanishing atomic current, if we drop the contribution of the exponential growth of the norm to the current behaviors. When the system is non-PT-symmetric, though the quasienergies are partial complex, a stable net spin current can be generated by controlling the periodic driving field, which is accompanied by a spatial localization of the condensate in the well with gain. The results deepen the understanding of non-Hermitian physics and could be useful for engineering a variety of devices for spintronics.

cond-mat.quant-gas

On finite termination of the generalized Newton method for solving absolute value equations

Motivated by the framework constructed by Brugnano and Casulli $[$SIAM J. Sci. Comput. 30: 463--472, 2008$]$, we analyze the finite termination property of the generalized Netwon method (GNM) for solving the absolute value equation (AVE). More precisely, for some special matrices, GNM is terminated in at most $2n + 2$ iterations. A new result for the unique solvability and unsolvability of the AVE is obtained. Numerical experiments are given to demonstrate the theoretical analysis.

math.OC