SearcharxivSearch

arXiv subjects

Ke Yu

Publications and source records attributed to Ke Yu.

At least 19 recordsLinked to original sources

HELIX: Hybrid Encoding with Learnable Identity and Cross-dimensional Synthesis for Time Series Imputation

Time series imputation benefits from leveraging cross-feature correlations, yet existing attention-based methods re-discover feature relationships at each layer, lacking persistent anchors to maintain consistent representations. To address this, we propose HELIX, which assigns each feature a learnable feature identity, a persistent embedding that captures intrinsic semantic properties throughout the network. Unlike graph-based methods that rely on predefined topology and assume homogeneous spatial relationships, HELIX learns arbitrary feature dependencies end-to-end from temporal co-variation, naturally handling datasets where features mix spatial locations with semantic variables. Integrated with hybrid temporal-feature attention, HELIX achieves the state-of-the-art performance, surpassing all 16 baselines on 5 public datasets across 21 experimental settings in our evaluation. Furthermore, our mechanistic analysis reveals that HELIX aligns learned feature identities and dependencies with latent physical and semantic structure progressively across layers, demonstrating that it more effectively translates cross-feature structure into imputation accuracy.

cs.LG

Lattice Discrete Particle Model (LDPM): Comparison of Various Time Integration Solvers and Implementations

This article presents a comparison of various implementations of the Lattice Discrete Particle Model (LDPM) for the numerical simulation of concrete and other heterogeneous quasibrittle materials. The comparison involves the use of transient implicit and explicit solvers and steady-state (static) solvers and implementations for Central Processing Unit (CPU) as well as Graphics Processing Unit (GPU). The various implementations are compared on the basis of a set of benchmarks tests describing behaviors of increasing computational complexity. They include elastic vibrations, confined strain-hardening compressive response, tensile fracture, and unconfined strain-softening compressive response. Metrics of interest extracted from the simulations include macroscopic stress versus strain responses, computational times, number of iterations, and energy balance error. Pairwise comparison of final crack patterns is provided through the correlation coefficient and normalized root mean square error of the crack opening vectors. Moreover, for the most numerically challenging case of unconfined compression with sliding boundary conditions, the stability of the strain-softening response is tested by perturbing the solutions as well as changing the convergence criteria and time step size. Attached to this paper is the complete input data of the benchmark tests; this will allow researchers to run the examples and compare them with their own implementations. In addition, most of the reported implementations are publicly available in open source packages.

cs.CE

Pavement Missing Condition Data Imputation through Collective Learning-Based Graph Neural Networks

Pavement condition data is important in providing information regarding the current state of the road network and in determining the needs of maintenance and rehabilitation treatments. However, the condition data is often incomplete due to various reasons such as sensor errors and non-periodic inspection schedules. Missing data, especially data missing systematically, presents loss of information, reduces statistical power, and introduces biased assessment. Existing methods in dealing with missing data usually discard entire data points with missing values or impute through data correlation. In this paper, we used a collective learning-based Graph Convolutional Networks, which integrates both features of adjacent sections and dependencies between observed section conditions to learn missing condition values. Unlike other variants of graph neural networks, the proposed approach is able to capture dependent relationship between the conditions of adjacent pavement sections. In the case study, pavement condition data collected from Texas Department of Transportation Austin District were used. Experiments show that the proposed model was able to produce promising results in imputing the missing data.

cs.LG

A Cautionary Tale of Self-Supervised Learning for Imaging Biomarkers: Alzheimer's Disease Case Study

Discovery of sensitive and biologically grounded biomarkers is essential for early detection and monitoring of Alzheimer's disease (AD). Structural MRI is widely available but typically relies on hand-crafted features such as cortical thickness or volume. We ask whether self-supervised learning (SSL) can uncover more powerful biomarkers from the same data. Existing SSL methods underperform FreeSurfer-derived features in disease classification, conversion prediction, and amyloid status prediction. We introduce Residual Noise Contrastive Estimation (R-NCE), a new SSL framework that integrates auxiliary FreeSurfer features while maximizing additional augmentation-invariant information. R-NCE outperforms traditional features and existing SSL methods across multiple benchmarks, including AD conversion prediction. To assess biological relevance, we derive Brain Age Gap (BAG) measures and perform genome-wide association studies. R-NCE-BAG shows high heritability and associations with MAPT and IRAG1, with enrichment in astrocytes and oligodendrocytes, indicating sensitivity to neurodegenerative and cerebrovascular processes.

cs.LG

Multiscale Cross-Modal Mapping of Molecular, Pathologic, and Radiologic Phenotypes in Lipid-Deficient Clear Cell Renal CellCarcinoma

Clear cell renal cell carcinoma (ccRCC) exhibits extensive intratumoral heterogeneity on multiple biological scales, contributing to variable clinical outcomes and limiting the effectiveness of conventional TNM staging, which highlights the urgent need for multiscale integrative analytic frameworks. The lipid-deficient de-clear cell differentiated (DCCD) ccRCC subtype, defined by multi-omics analyses, is associated with adverse outcomes even in early-stage disease. Here, we establish a hierarchical cross-scale framework for the preoperative identification of DCCD-ccRCC. At the highest layer, cross-modal mapping transferred molecular signatures to histological and CT phenotypes, establishing a molecular-to-pathology-to-radiology supervisory bridge. Within this framework, each modality-specific model is designed to mirror the inherent hierarchical structure of tumor biology. PathoDCCD captured multi-scale microscopic features, from cellular morphology and tissue architecture to meso-regional organization. RadioDCCD integrated complementary macroscopic information by combining whole-tumor and its habitat-subregions radiomics with a 2D maximal-section heterogeneity metric. These nested models enabled integrated molecular subtype prediction and clinical risk stratification. Across five cohorts totaling 1,659 patients, PathoDCCD reliably recapitulated molecular subtypes, while RadioDCCD provided reliable preoperative prediction. The consistent predictions identified patients with the poorest clinical outcomes. This cross-scale paradigm unifies molecular biology, computational pathology, and quantitative radiology into a biologically grounded strategy for preoperative noninvasive molecular phenotyping of ccRCC.

q-bio.QM

SRLR: Symbolic Regression based Logic Recovery to Counter Programmable Logic Controller Attacks

Programmable Logic Controllers (PLCs) are critical components in Industrial Control Systems (ICSs). Their potential exposure to external world makes them susceptible to cyber-attacks. Existing detection methods against controller logic attacks use either specification-based or learnt models. However, specification-based models require experts' manual efforts or access to PLC's source code, while machine learning-based models often fall short of providing explanation for their decisions. We design SRLR -- a it Symbolic Regression based Logic Recovery} solution to identify the logic of a PLC based only on its inputs and outputs. The recovered logic is used to generate explainable rules for detecting controller logic attacks. SRLR enhances the latest deep symbolic regression methods using the following ICS-specific properties: (1) some important ICS control logic is best represented in frequency domain rather than time domain; (2) an ICS controller can operate in multiple modes, each using different logic, where mode switches usually do not happen frequently; (3) a robust controller usually filters out outlier inputs as ICS sensor data can be noisy; and (4) with the above factors captured, the degree of complexity of the formulas is reduced, making effective search possible. Thanks to these enhancements, SRLR consistently outperforms all existing methods in a variety of ICS settings that we evaluate. In terms of the recovery accuracy, SRLR's gain can be as high as 39% in some challenging environment. We also evaluate SRLR on a distribution grid containing hundreds of voltage regulators, demonstrating its stability in handling large-scale, complex systems with varied configurations.

cs.LG

Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning

Recent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational resources for fine-tuning in new domains. To address this issue, we develop a new adaptation method for large vision-language models, called \textit{Training-free Dual Hyperbolic Adapters} (T-DHA). We characterize the vision-language relationship between semantic concepts, which typically has a hierarchical tree structure, in the hyperbolic space instead of the traditional Euclidean space. Hyperbolic spaces exhibit exponential volume growth with radius, unlike the polynomial growth in Euclidean space. We find that this unique property is particularly effective for embedding hierarchical data structures using the Poincar\'e ball model, achieving significantly improved representation and discrimination power. Coupled with negative learning, it provides more accurate and robust classifications with fewer feature dimensions. Our extensive experimental results on various datasets demonstrate that the T-DHA method significantly outperforms existing state-of-the-art methods in few-shot image recognition and domain generalization tasks.

cs.CV

Two-stage primary acceleration in filament initial eruption under a fan-spine magnetic configuration

Understaning the filament rising process is crucial for unveiling the triggering mechanisms of the coronal mass ejections and forecasting the space weather. In this paper, we present a detailed study on the filament initial eruption under a fan-spine structure. It was found that the filament underwent two distinct acceleration stages corresponding to a calss M1.0 and M4.6 flare event, respectively. The first acceleration stage commenced with the filament splitting, after which the upper portion was subsequently heated being a hot channel and slow rose at an average speed of 22 km/s. A set of hot reverse C-shaped loops appeared repeatedly during the filament splitting and a hook structure was recognized at this phase, suggesting ongoing growth of the magnetic flux rope (MFR). When it reached a certain altitude, the hot channel appeared to get into a quasi-static phase with its upper edge seriously decelerated and lower edge expanding downward. Approximately 30 minutes later, as a distinct annular ribbon appeared outside the hook structure, the hot channel rose again at a velocity over 50 km/s accompanied with rapid footpoints drifting, and experienced the second acceleration stage with its axial flux increased to 1.1 X 10^{21} Mx. It is deduced that the filament initial eruption under a magnetic dome possess multi kinetic process. We suggest that the magnetic reconnection taken place within and beneath the filament continues to trigger the growth of pre-eruptive MFR and the first acceleration, when the magnetic reconnection above the filament plays a key role in the second acceleration.

astro-ph.SR

SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data

Synthetic tabular data, which are widely used in domains such as healthcare, enterprise operations, and customer analytics, are increasingly evaluated to ensure that they preserve both privacy and utility. While existing evaluation practices typically focus on distributional similarity (e.g., the Kullback-Leibler divergence) or predictive performance (e.g., Train-on-Synthetic-Test-on-Real (TSTR) accuracy), these approaches fail to assess semantic fidelity, that is, whether models trained on synthetic data follow reasoning patterns consistent with those trained on real data. To address this gap, we introduce the SHapley Additive exPlanations (SHAP) Distance, a novel explainability-aware metric that is defined as the cosine distance between the global SHAP attribution vectors derived from classifiers trained on real versus synthetic datasets. By analyzing datasets that span clinical health records with physiological features, enterprise invoice transactions with heterogeneous scales, and telecom churn logs with mixed categorical-numerical attributes, we demonstrate that the SHAP Distance reliably identifies semantic discrepancies that are overlooked by standard statistical and predictive measures. In particular, our results show that the SHAP Distance captures feature importance shifts and underrepresented tail effects that the Kullback-Leibler divergence and Train-on-Synthetic-Test-on-Real accuracy fail to detect. This study positions the SHAP Distance as a practical and discriminative tool for auditing the semantic fidelity of synthetic tabular data, and offers practical guidelines for integrating attribution-based evaluation into future benchmarking pipelines.

cs.LG

Two-sided-loop jet originates from the filament internal reconnection

Magnetic reconnection driving two-sided-loop jet is typically associated with interactions between an emerging bipole and the overlying horizontal magnetic field, or between filaments from separate magnetic systems. Leveraging high temporal and spatial resolution observations from ground-based and space-borne instruments, we have identified a two-sided-loop jet originating from magnetic reconnection between threads within a single filament. Our observations show that as two initially crossing filamentary threads within the filament converge, reconnection takes place at their intersection. In the Doppler images, distinct redshift and blueshift signals are observed at the locations where the filament threads intersected. This process generates a two-sided-loop jet with outflow speeds of \speed{22.2} and \speed{62.5}. Following reconnection, the original crossing threads transform into two parallel threads that subsequently separate at speeds of \speed{2.8} and \speed{8.3}. This observation offers a new perspective on the mechanisms responsible for jet formation.

astro-ph.SR

GraphFedMIG: Tackling Class Imbalance in Federated Graph Learning via Mutual Information-Guided Generation

Federated graph learning (FGL) enables multiple clients to collaboratively train powerful graph neural networks without sharing their private, decentralized graph data. Inherited from generic federated learning, FGL is critically challenged by statistical heterogeneity, where non-IID data distributions across clients can severely impair model performance. A particularly destructive form of this is class imbalance, which causes the global model to become biased towards majority classes and fail at identifying rare but critical events. This issue is exacerbated in FGL, as nodes from a minority class are often surrounded by biased neighborhood information, hindering the learning of expressive embeddings. To grapple with this challenge, we propose GraphFedMIG, a novel FGL framework that reframes the problem as a federated generative data augmentation task. GraphFedMIG employs a hierarchical generative adversarial network where each client trains a local generator to synthesize high-fidelity feature representations. To provide tailored supervision, clients are grouped into clusters, each sharing a dedicated discriminator. Crucially, the framework designs a mutual information-guided mechanism to steer the evolution of these client generators. By calculating each client's unique informational value, this mechanism corrects the local generator parameters, ensuring that subsequent rounds of mutual information-guided generation are focused on producing high-value, minority-class features. We conduct extensive experiments on four real-world datasets, and the results demonstrate the superiority of the proposed GraphFedMIG compared with other baselines.

cs.LG

Considering Spatial Structure of the Road Network in Pavement Deterioration Modeling

Pavement deterioration modeling is important in providing information regarding the future state of the road network and in determining the needs of preventive maintenance or rehabilitation treatments. This research incorporated spatial dependence of road network into pavement deterioration modeling through a graph neural network (GNN). The key motivation of using a GNN for pavement performance modeling is the ability to easily and directly exploit the rich structural information in the network. This paper explored if considering spatial structure of the road network will improve the prediction performance of the deterioration models. The data used in this research comprises a large pavement condition data set with more than a half million observations taken from the Pavement Management Information System (PMIS) maintained by the Texas Department of Transportation. The promising comparison results indicates that pavement deterioration prediction models perform better when spatial relationship is considered.

cs.LG

3D fast-mode Wave Propagation from Corona to Chromosphere: Triggering Mechanism for 3D Oscillations of filaments

Moreton waves are widely regarded as the chromospheric counterpart of extreme ultraviolet (EUV) waves propagating in the corona. However, direct observational evidence confirming their simultaneous propagation across multiple atmospheric layers from the corona through the transition region to the chromosphere has been lacking. In this study, we present comprehensive observational evidence of a three-dimensional (3D) fast-mode wave propagating from the corona through the transition region into the chromosphere, exhibiting a gradual deceleration. Additionally, this wave interacts with three filaments (F1, F2, and F3) along its path, inducing oscillation with multiple amplitudes: Filaments F1 and F2 exhibit simultaneous horizontal and vertical large-scale oscillations ($\sim$\speed{20}), while Filament F3 only exhibits vertical small-scale oscillation ($\sim$\speed{4}). Interestingly, F1 displays a similar oscillation period of about 500\,s in both horizontal and vertical directions, whereas F2 shows significantly different periods in these two dimensions (1100\,s and 750\,s), and F3 exhibits only a vertical oscillation with a period of about 450\,s. Based on this kinematic behavior, we propose that their oscillations were likely triggered by compression from the flanks of the dome-shaped wavefront. We further estimate the magnetic fields of the filaments. The radial (axial) magnetic fields for F1 and F2 are estimated to be 14.9\,G (28.6\,G) and 9.9\,G (18.6\,G), respectively. For F3, we estimate its radial magnetic field to be 16.6\,G.

astro-ph.SR

Temperature dependence of quasi-localized phonons-mediated non-Markovianity dynamics of SiV^- centers in diamond

Here we investigate the temperature-dependent non-Markovian dynamics of the SiV^- center in diamond, focusing on the roles of low- and high-frequency quasi-localized phonon modes. Low-frequency phonons exhibit stronger electron-phonon coupling, leading to long-lived dephasing rate, while high-frequency phonons induce rapid attenuation of oscillatory dephasing rate facilitating a persistent memory effect. The non-Markovianity measure N_C shows memory effects persisting at low temperatures but diminishing at high temperatures due to enhanced damping. The temperature dependence of N_C follows a monotonic decay, from which a transition temperature T_NM=110 K is determined. These results highlight the interplay between phonon activation and damping in shaping quantum coherence, offering insights for optimizing solid-state quantum systems.

physics.optics

Rethinking Text-based Protein Understanding: Retrieval or LLM?

In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding. Current approaches focus on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment, enabling simultaneous comprehension of textual descriptions and protein sequences. Through a thorough analysis of existing model architectures and text-based protein understanding benchmarks, we identify significant data leakage issues present in current benchmarks. Moreover, conventional metrics derived from natural language processing fail to accurately assess the model's performance in this domain. To address these limitations, we reorganize existing datasets and introduce a novel evaluation framework based on biological entities. Motivated by our observation, we propose a retrieval-enhanced method, which significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios. Our code and data can be seen at https://github.com/IDEA-XL/RAPM.

cs.CL

ACPs: Agent Collaboration Protocols for the Internet of Agents

With the rapid advancement of artificial intelligence, the proliferation of autonomous agents has introduced new challenges in interoperability, scalability, and coordination. The Internet of Agents (IoA) aims to interconnect heterogeneous agents through standardized communication protocols, enabling seamless collaboration and intelligent task execution. However, existing agent communication protocols such as MCP, A2A, and ANP remain fragmented and scenario-specific. To address this gap, we propose Agent Collaboration Protocols (ACPs), a comprehensive protocol suite for the IoA. ACPs include registration, discovery, interaction, and tooling protocols to support trustable access, capability orchestration, and workflow construction. We present the architecture, key technologies, and application workflows of ACPs, and demonstrate its effectiveness in a collaborative restaurant booking scenario. ACPs lay the foundation for building a secure, open, and scalable agent internet infrastructure.

cs.MA

Learning and Current Prediction of PMSM Drive via Differential Neural Networks

Learning models for dynamical systems in continuous time is significant for understanding complex phenomena and making accurate predictions. This study presents a novel approach utilizing differential neural networks (DNNs) to model nonlinear systems, specifically permanent magnet synchronous motors (PMSMs), and to predict their current trajectories. The efficacy of our approach is validated through experiments conducted under various load disturbances and no-load conditions. The results demonstrate that our method effectively and accurately reconstructs the original systems, showcasing strong short-term and long-term prediction capabilities and robustness. This study provides valuable insights into learning the inherent dynamics of complex dynamical data and holds potential for further applications in fields such as weather forecasting, robotics, and collective behavior analysis.

cs.LG

ICODE: Modeling Dynamical Systems with Extrinsic Input Information

Learning models of dynamical systems with external inputs, which may be, for example, nonsmooth or piecewise, is crucial for studying complex phenomena and predicting future state evolution, which is essential for applications such as safety guarantees and decision-making. In this work, we introduce \emph{Input Concomitant Neural ODEs (ICODEs)}, which incorporate precise real-time input information into the learning process of the models, rather than treating the inputs as hidden parameters to be learned. The sufficient conditions to ensure the model's contraction property are provided to guarantee that system trajectories of the trained model converge to a fixed point, regardless of initial conditions across different training processes. We validate our method through experiments on several representative real dynamics: Single-link robot, DC-to-DC converter, motion dynamics of a rigid body, Rabinovich-Fabrikant equation, Glycolytic-glycogenolytic pathway model, and heat conduction equation. The experimental results demonstrate that our proposed ICODEs efficiently learn the ground truth systems, achieving superior prediction performance under both typical and atypical inputs. This work offers a valuable class of neural ODE models for understanding physical systems with explicit external input information, with potentially promising applications in fields such as physics and robotics. Our code is available online at https://github.com/EEE-ai59/ICODE.git.

cs.LG