SearcharxivSearch

arXiv subjects

Haoyu Ji

Publications and source records attributed to Haoyu Ji.

12 recordsLinked to original sources

Hourly U.S.-wide flood simulation beyond the limits of traditional and data-driven models

As increasingly-damaging floods can strike within hours of a storm and in ungauged reaches, hourly network-wide simulation has become critical societal infrastructure. Here we demonstrate a multi-timescale physics-embedded learning model which outperforms the United States' operational system and surpasses AI-based systems at flood peaks. Covering more than 800,000 river reaches of the conterminous U.S., the model elevates median hourly Nash-Sutcliffe efficiency at 2,831 gauges to 0.683 from 0.461 for the operational National Water Model v3.0, and narrows flood-peak timing errors from 7-8 hours to 4.5-6 hours. dHBV2.0MTS-MC captures 33% more >=50-year floods than NWM3.0 and 159% more than an operational LSTM baseline. Against recent AI models, overall hourly skill is comparable while rare-flood accuracy is distinctly higher, with relative peak-magnitude error reduced by 34% for >=100-year floods. It combines long-term hydrologic context, short-term shocks, and infiltration excess to resolve extraordinary hourly peaks not visible on a daily plot. Process-model parameters and hourly discharge are produced for every reach, seamlessly covering the continent at 7.2 km2 median resolution. This candidate for the next-generation National Water Model sets a new operational accuracy level for national-scale flood prediction.

physics.geo-ph

AI-Driven SERS for Non-invasive and Label-Free Extracellular Vesicle Detection Across Cellular Origins in Tears and Sweat

Wearable sensing technology capable of point-of-care, continuous and non-invasive analysis of exosomes in biofluid such as tears and sweat is an essential part for future personalized medicine. Major detection and identification methods of cell secreted Extracellular Vesicles (EVs) often require labeling and are time-consuming, resulting in low efficiency in EV mechanism research and disease diagnosis. While the label-free Surface-enhanced Raman spectroscopy (SERS) has been combined with deep learning model for EV identification in blood, their application to non-invasive detection of EVs in tears and sweat are missing. Here, we filled this gap by developing an artificial intelligence (AI)-assisted Surface-enhanced Raman spectroscopy (SERS) method based on salt-induced nanoparticle aggregation for fast EV identification in tears and sweat with high accuracy. Significantly, our label-free detection and AI differentiation of EVs from 6 cell lines (HepG2, Hela, 143B, LO-2, BMSC, H8) achieved the identification of EVs in tear fluids from 7 different disease sources with accuracies >92%. Our results showed that this platform can not only distinguish EVs from multiple cell sources but also generate highly reproducible and selective EV signals in tear fluids without a need for chemical labeling or separation steps. Molecular dynamics simulations revealed that silver atoms (Ag) form electrostatic interactions with oxygen atoms of multiple amino acid residues in proteins, suggesting a high affinity. This strategy realizes ultra-sensitive and anti-interference detection of EVs, providing a new idea for the rapid diagnosis of clinical diseases.

cond-mat.mes-hall

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling

Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generation paradigms often struggle with a lack of spatial persistence and insufficient visual realism, making it difficult to support seamless navigation in complex environments. To address these challenges, we propose INSPATIO-WORLD, a novel real-time framework capable of recovering and generating high-fidelity, dynamic interactive scenes from a single reference video. At the core of our approach is a Spatiotemporal Autoregressive (STAR) architecture, which enables consistent and controllable scene evolution through two tightly coupled components: Implicit Spatiotemporal Cache aggregates reference and historical observations into a latent world representation, ensuring global consistency during long-horizon navigation; Explicit Spatial Constraint Module enforces geometric structure and translates user interactions into precise and physically plausible camera trajectories. Furthermore, we introduce Joint Distribution Matching Distillation (JDMD). By using real-world data distributions as a regularizing guide, JDMD effectively overcomes the fidelity degradation typically caused by over-reliance on synthetic data. Extensive experiments demonstrate that INSPATIO-WORLD significantly outperforms existing state-of-the-art (SOTA) models in spatial consistency and interaction precision, ranking first among real-time interactive methods on the WorldScore-Dynamic benchmark, and establishing a practical pipeline for navigating 4D environments reconstructed from monocular videos.

cs.CV

LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation

Skeleton-based Temporal Action Segmentation (STAS) aims to densely parse untrimmed skeletal sequences into frame-level action categories. However, existing methods, while proficient at capturing spatio-temporal kinematics, neglect the underlying physical dynamics that govern human motion. This oversight limits inter-class discriminability between actions with similar kinematics but distinct dynamic intents, and hinders precise boundary localization where dynamic force profiles shift. To address these, we propose the Lagrangian-Dynamic Informed Network (LaDy), a framework integrating principles of Lagrangian dynamics into the segmentation process. Specifically, LaDy first computes generalized coordinates from joint positions and then estimates Lagrangian terms under physical constraints to explicitly synthesize the generalized forces. To further ensure physical coherence, our Energy Consistency Loss enforces the work-energy theorem, aligning kinetic energy change with the work done by the net force. The learned dynamics then drive a Spatio-Temporal Modulation module: Spatially, generalized forces are fused with spatial representations to provide more discriminative semantics. Temporally, salient dynamic signals are constructed for temporal gating, thereby significantly enhancing boundary awareness. Experiments on challenging datasets show that LaDy achieves state-of-the-art performance, validating the integration of physical dynamics for action segmentation. Code is available at https://github.com/HaoyuJi/LaDy.

cs.CV

Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation

Skeleton-based Temporal Action Segmentation (STAS) seeks to densely segment and classify diverse actions within long, untrimmed skeletal motion sequences. However, existing STAS methodologies face challenges of limited inter-class discriminability and blurred segmentation boundaries, primarily due to insufficient distinction of spatio-temporal patterns between adjacent actions. To address these limitations, we propose Spectral Scalpel, a frequency-selective filtering framework aimed at suppressing shared frequency components between adjacent distinct actions while amplifying their action-specific frequencies, thereby enhancing inter-action discrepancies and sharpening transition boundaries. Specifically, Spectral Scalpel employs adaptive multi-scale spectral filters as scalpels to edit frequency spectra, coupled with a discrepancy loss between adjacent actions serving as the surgical objective. This design amplifies representational disparities between neighboring actions, effectively mitigating boundary localization ambiguities and inter-class confusion. Furthermore, complementing long-term temporal modeling, we introduce a frequency-aware channel mixer to strengthen channel evolution by aggregating spectra across channels. This work presents a novel paradigm for STAS that extends conventional spatio-temporal modeling by incorporating frequency-domain analysis. Extensive experiments on five public datasets demonstrate that Spectral Scalpel achieves state-of-the-art performance. Code is available at https://github.com/HaoyuJi/SpecScalpel.

cs.CV

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential frame generation and incur substantial latency due to window-level processing, InSpatio-WorldFM adopts a frame-based paradigm that generates each frame independently, enabling low-latency real-time spatial inference. By enforcing multi-view spatial consistency through explicit 3D anchors and implicit spatial memory, the model preserves global scene geometry while maintaining fine-grained visual details across viewpoint changes. We further introduce a progressive three-stage training pipeline that transforms a pretrained image diffusion model into a controllable frame model and finally into a real-time generator through few-step distillation. Experimental results show that InSpatio-WorldFM achieves strong multi-view consistency while supporting interactive exploration on consumer-grade GPUs, providing an efficient alternative to traditional video-based world models for real-time world simulation.

cs.CV

Progressive Mixture-of-Experts with autoencoder routing for continual RANS turbulence modelling

Developing Reynolds-averaged Navier-Stokes (RANS) turbulence models that remain accurate across diverse flow regimes is a long-standing challenge. In this work, we propose a novel framework, termed the progressive mixture-of-experts (PMoE), designed to enable continual learning for RANS turbulence modelling. The framework employs a modular autoencoder-based router to associate each flow scenario with a specialised turbulence model, referred to as an expert. When a new flow regime cannot be adequately represented by the existing router and expert set, a new expert together with its routing component can be introduced at low cost, without modifying or degrading previously trained ones, thereby naturally avoiding catastrophic forgetting. The framework is applied to a range of flows with distinct physical characteristics, including airfoil wake, channel, periodic hill, and square duct flows. The resulting PMoE model effectively integrates multiple experts and achieves improved predictive accuracy across both seen and unseen test cases that differ in operating conditions or configurations. Owing to sparse activation, model expansion does not incur additional computational cost during inference. The proposed framework therefore provides a scalable pathway towards lifelong-learning turbulence models for industrial computational fluid dynamics.

physics.flu-dyn

Diffusion-Based Probabilistic Modeling for Hourly Streamflow Prediction and Assimilation

Hourly predictions are critical for issuing flood warnings as the flood peaks on the hourly scale can be distinctly higher than the corresponding daily ones. Currently a popular hourly data-driven prediction scheme is multi-time-scale long short-term memory (MTS-LSTM), yet such models face challenges in probabilistic forecasts or integrating observations when available. Diffusion artificial intelligence (AI) models represent a promising method to predict high-resolution information, e.g., hourly streamflow. Here we develop a denoising diffusion probabilistic model (h-Diffusion) for hourly streamflow prediction that conditions on either observed or simulated daily discharge from hydrologic models to generate hourly hydrographs. The model is benchmarked on the CAMELS hourly dataset against record-holding MTS-LSTM and multi-frequency LSTM (MF-LSTM) baselines. Results show that h-Diffusion outperforms baselines in terms of general performance and extreme metrics. Furthermore, the h-Diffusion model can utilize the inpainting technique and recent observations to accomplish data assimilation that largely improves flood forecasting performance. These advances can greatly reduce flood forecasting uncertainty and provide a unified probabilistic framework for downscaling, prediction, and data assimilation at the hourly scale, representing risks where daily models cannot.

physics.geo-ph

Distinct hydrologic response patterns and trends worldwide revealed by physics-embedded learning

To track rapid changes within our water sector, Global Water Models (GWMs) need to realistically represent hydrologic systems' response patterns - such as baseflow fraction - but are hindered by their limited ability to learn from data. Here we introduce a high-resolution physics-embedded big-data-trained model as a breakthrough in reliably capturing characteristic hydrologic response patterns ('signatures') and their shifts. By realistically representing the long-term water balance, the model revealed widespread shifts - up to ~20% over 20 years - in fundamental green-blue-water partitioning and baseflow ratios worldwide. Shifts in these response patterns, previously considered static, contributed to increasing flood risks in northern mid-latitudes, heightening water supply stresses in southern subtropical regions, and declining freshwater inputs to many European estuaries, all with ecological implications. With more accurate simulations at monthly and daily scales than current operational systems, this next-generation model resolves large, nonlinear seasonal runoff responses to rainfall ('elasticity') and streamflow flashiness in semi-arid and arid regions. These metrics highlight regions with management challenges due to large water supply variability and high climate sensitivity, but also provide tools to forecast seasonal water availability. This capability newly enables global-scale models to deliver reliable and locally relevant insights for water management.

physics.geo-ph

Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation

Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling to establish dependencies among joints as well as frames, and utilize one-hot encoding with cross-entropy loss for frame-wise classification supervision. However, these methods overlook the intrinsic correlations among joints and actions within skeletal features, leading to a limited understanding of human movements. To address this, we propose a Text-Derived Relational Graph-Enhanced Network (TRG-Net) that leverages prior graphs generated by Large Language Models (LLM) to enhance both modeling and supervision. For modeling, the Dynamic Spatio-Temporal Fusion Modeling (DSFM) method incorporates Text-Derived Joint Graphs (TJG) with channel- and frame-level dynamic adaptation to effectively model spatial relations, while integrating spatio-temporal core features during temporal modeling. For supervision, the Absolute-Relative Inter-Class Supervision (ARIS) method employs contrastive learning between action features and text embeddings to regularize the absolute class distributions, and utilizes Text-Derived Action Graphs (TAG) to capture the relative inter-class relationships among action features. Additionally, we propose a Spatial-Aware Enhancement Processing (SAEP) method, which incorporates random joint occlusion and axial rotation to enhance spatial generalization. Performance evaluations on four public datasets demonstrate that TRG-Net achieves state-of-the-art results.

cs.CV

Language-Assisted Human Part Motion Learning for Skeleton-Based Temporal Action Segmentation

Skeleton-based Temporal Action Segmentation involves the dense action classification of variable-length skeleton sequences. Current approaches primarily apply graph-based networks to extract framewise, whole-body-level motion representations, and use one-hot encoded labels for model optimization. However, whole-body motion representations do not capture fine-grained part-level motion representations and the one-hot encoded labels neglect the intrinsic semantic relationships within the language-based action definitions. To address these limitations, we propose a novel method named Language-assisted Human Part Motion Representation Learning (LPL), which contains a Disentangled Part Motion Encoder (DPE) to extract dual-level (i.e., part and whole-body) motion representations and a Language-assisted Distribution Alignment (LDA) strategy for optimizing spatial relations within representations. Specifically, after part-aware skeleton encoding via DPE, LDA generates dual-level action descriptions to construct a textual embedding space with the help of a large-scale language model. Then, LDA motivates the alignment of the embedding space between text descriptions and motions. This alignment allows LDA not only to enhance intra-class compactness but also to transfer the language-encoded semantic correlations among actions to skeleton-based motion learning. Moreover, we propose a simple yet efficient Semantic Offset Adapter to smooth the cross-domain misalignment. Our experiments indicate that LPL achieves state-of-the-art performance across various datasets (e.g., +4.4\% Accuracy, +5.6\% F1 on the PKU-MMD dataset). Moreover, LDA is compatible with existing methods and improves their performance (e.g., +4.8\% Accuracy, +4.3\% F1 on the LARa dataset) without additional inference costs.

cs.CV

Label-free detection of exosomes from different cellular sources based on surface-enhanced Raman spectroscopy combined with machine learning models

Exosomes are significant facilitators of inter-cellular communication that can unveil cell-cell interactions, signaling pathways, regulatory mechanisms and disease diagnostics. Nonetheless, current analysis required large amount of data for exosome identification that it hampers efficient and timely mechanism study and diagnostics. Here, we used a machine-learning assisted Surface-enhanced Raman spectroscopy (SERS) method to detect exosomes derived from six distinct cell lines (HepG2, Hela, 143B, LO-2, BMSC, and H8) with small amount of data. By employing sodium borohydride-reduced silver nanoparticles and sodium borohydride solution as an aggregating agent, 100 SERS spectra of the each types of exosomes were collected and then subjected to multivariate and machine learning analysis. By integrating Principal Component Analysis with Support Vector Machine (PCA-SVM) models, our analysis achieved a high accuracy rate of 94.4% in predicting exosomes originating from various cellular sources. In comparison to other machine learning analysis, our method used small amount of SERS data to allow a simple and rapid exosome detection, which enables a timely subsequent study of cell-cell interactions, communication mechanisms, and disease mechanisms in life sciences.

q-bio.BM