SearcharxivSearch

arXiv subjects

Yong Sun

Publications and source records attributed to Yong Sun.

At least 19 recordsLinked to original sources

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses. While LLM-based approaches show promise in code analysis, they struggle with deep protocol-level logic bugs involving complex state-dependent behaviors across multiple execution stages. We present Agora, a domain-aware multi-agent framework that integrates hypothesis-driven testing with LLM capabilities for systematic protocol verification. Agora employs specialized agents that collaboratively explore protocol state spaces, synthesize attack scenarios using domain-specific constraints, and validate findings through iterative refinement. This explicit role separation enables reasoning about global protocol invariants beyond single-function code analysis. We evaluate Agora on four consensus implementations (Raft, EPaxos, HotStuff, BullShark) using four state-of-the-art LLMs. Agora discovers 15 previously unknown protocol-level logic bugs that violate safety properties, while existing LLM-based agents fail to detect any such protocol-level logic bugs. Our results demonstrate that domain-aware multi-agent collaboration is essential for detecting deep logic bugs in complex protocols.

cs.SE

Wave-number-dependent closure condition for fluid moment equations

Fluid models offer crucial computational efficiency for plasma simulations, yet accurately capturing kinetic effects like Landau damping remains a fundamental challenge. While conventional closures (e.g., Hammett-Perkins and Hunana) are widely used, their fidelity relative to exact kinetic response degrades significantly depending on the perturbation wave number. Here, we propose a novel wave-number-dependent closure condition for the three-moment fluid equations that explicitly preserves the primary dispersion relation. By mapping Pad\'e approximant coefficients directly to the kinetic roots of the collisionless Vlasov-Poisson system, we derive an analytical closure that rigorously embeds exact kinetic scaling across all spatial scales. We further demonstrate that this framework readily extends to collisional plasmas via the BGK model. This deterministic approach precisely captures the long-term macroscopic evolution of fluid moments and field energy, offering a rigorous foundation for high-fidelity fluid modeling.

physics.plasm-ph

Personalized Multi-Interest Modeling for Cross-Domain Recommendation to Cold-Start Users

Cross-domain recommendation (CDR) has demonstrated to be an effective solution for alleviating the user cold-start issue. By leveraging rich user-item interactions available in a richly informative source domain, CDR could improve the recommendation performance for cold-start users in the target domain. Previous CDR approaches mostly adhere the Embedding and Mapping (EMCDR) paradigm, which learns a user-shared mapping function to transfer users' preference from the source domain to the target domain, neglecting users' personalized preference. Recent CDR approaches further leverage the meta-learning paradigm, considering the CDR task for each user independently and learning user-specific mapping functions for each user. However, they mostly learn representations for each user individually, which ignores the common preference between different users, neglecting valuable information for CDR. In addition, all these approaches usually summarize the user's preference into an overall representation, which can hardly capture the user's multi-interest preference. To this end, we propose a personalized multi-interest modeling framework for CDR to cold-start users, termed as NF-NPCDR. Specifically, we propose a personalized preference encoder that enhances the neural process (NP) with the normalizing flow (NF) to convert the Gaussian (unimodal) distribution to a multimodal distribution, providing a novel way to capture the user's personalized multi-interest preference. Then, we propose a common preference encoder with a preference pool to capture the common preference between different users. Furthermore, we introduce a stochastic adaptive decoder to incorporate both the personalized and common preference for cold-start users, adaptively modulating both preference for better recommendation.

cs.IR

Anomalous Non-Hermitian Topological Anderson Insulator

Strong disorder drives conventional Hermitian systems into Anderson insulating states, suppressing all topological phases. Here, we unveil symmetry-protected, anomalous topological phases in the strong disorder limit of a non-Hermitian system, characterized by a scale-invariant merging of zero-energy modes. Using the maximally symmetric Jx lattice as an ideal platform and introducing specifically engineered (ABBA-type) symmetry-preserving non-Hermitian disorder, we observe a sequence of disorder-induced phase transitions: from a trivial insulator into and through a non-Hermitian topological Anderson insulator (TAI) phase, culminating in a stable anomalous non-Hermitian TAI phase characterized by a quantized polarization P_x \approx 0.25. Within this anomalous phase protected by the mobility gap, the zero-energy modes exhibit a distinct (N/2)-mode coalescence that scales with system size. Our findings demonstrate that non-Hermitian disorder engineered to preserve symmetry can induce and protect novel topological order inaccessible to conventional Hermitian disorder, thereby advancing the fundamental understanding of topological phenomena mediated by the interplay of disorder and non-Hermiticity.

cond-mat.mes-hall

Realization of staircase topological Anderson phase transitions

One-dimensional topological Anderson insulators provide a paradigm for disorder-induced topological phases in which the underlying system turns from a trivial to a topological phase. It is widely recognized that the latter vanishes at large disorder amplitude. Here, and contrary to the general belief, we provide evidence for a successive disorder-driven topological transitions in a single-wall nanotube, culminating in a topological Anderson phase that remains unexpectedly robust at strong disorder. This phenomenon is confirmed by analysis of the corresponding topological invariant, which increases stepwise as disorder increases, giving evidence for the emergence of edge states. We experimentally implement these topological Anderson staircase phase transitions in a one-dimensional topolectrical circuit, where the persistence of edge states is revealed by node-voltage measurements. The robustness of the edge states is corroborated by numerical calculations of their localization properties. Our work opens the road to topological disordertronics, where topological phases can be tuned by disorder.

cond-mat.mes-hall

EmbryoDiff: A Conditional Diffusion Framework with Multi-Focal Feature Fusion for Fine-Grained Embryo Developmental Stage Recognition

Identification of fine-grained embryo developmental stages during In Vitro Fertilization (IVF) is crucial for assessing embryo viability. Although recent deep learning methods have achieved promising accuracy, existing discriminative models fail to utilize the distributional prior of embryonic development to improve accuracy. Moreover, their reliance on single-focal information leads to incomplete embryonic representations, making them susceptible to feature ambiguity under cell occlusions. To address these limitations, we propose EmbryoDiff, a two-stage diffusion-based framework that formulates the task as a conditional sequence denoising process. Specifically, we first train and freeze a frame-level encoder to extract robust multi-focal features. In the second stage, we introduce a Multi-Focal Feature Fusion Strategy that aggregates information across focal planes to construct a 3D-aware morphological representation, effectively alleviating ambiguities arising from cell occlusions. Building on this fused representation, we derive complementary semantic and boundary cues and design a Hybrid Semantic-Boundary Condition Block to inject them into the diffusion-based denoising process, enabling accurate embryonic stage classification. Extensive experiments on two benchmark datasets show that our method achieves state-of-the-art results. Notably, with only a single denoising step, our model obtains the best average test performance, reaching 82.8% and 81.3% accuracy on the two datasets, respectively.

cs.CV

MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context

Existing text-driven motion generation methods often treat synthesis as a bidirectional mapping between language and motion, but remain limited in capturing the causal logic of action execution and the human intentions that drive behavior. The absence of visual grounding further restricts precision and personalization, as language alone cannot specify fine-grained spatiotemporal details. We propose MoGIC, a unified framework that integrates intention modeling and visual priors into multimodal motion synthesis. By jointly optimizing multimodal-conditioned motion generation and intention prediction, MoGIC uncovers latent human goals, leverages visual priors to enhance generation, and exhibits versatile multimodal generative capability. We further introduce a mixture-of-attention mechanism with adaptive scope to enable effective local alignment between conditional tokens and motion subsequences. To support this paradigm, we curate Mo440H, a 440-hour benchmark from 21 high-quality motion datasets. Experiments show that after finetuning, MoGIC reduces FID by 38.6\% on HumanML3D and 34.6\% on Mo440H, surpasses LLM-based methods in motion captioning with a lightweight text head, and further enables intention prediction and vision-conditioned generation, advancing controllable motion synthesis and intention understanding. The code is available at https://github.com/JunyuShi02/MoGIC

cs.CV

Ergodicity and Measurements In Static and Dynamic Light Scattering

Static Light Scattering (SLS) and Dynamic Light Scattering (DLS) are very important techniques to study the characteristics of nano-particles in dispersion. The data of SLS is determined by the optical characteristic and the measured values of DLS are determined by optical and hydrodynamic characteristics of different size nano-particles. In general, the nano-particles investigated are considered to be cross-linked soft particles or hyper-branched chains with a three dimensional network structure. Therefore the density of these nano-particles is not homogeneous and the different parts have different optical characteristics. However our experiments reveal that the long time average data of scattered intensity can be perfectly described by homogeneous spherical model. Based on the size distribution obtained using the SLS or TEM technique and the relation between the static and hydrodynamic radii, all the calculated and measured values of $g^{\left( 2\right) }\left( \tau \right) $ investigated are also consistent very well. Since the long time average scattered intensity of non-homogeneous spherical nano-particles can be perfectly described by homogeneous spherical model, the phenomenon of ergodicity happen in our experiments. The results also reveal that the root mean-square radius of gyration $\left\langle R_{g}^{2}\right\rangle ^{1/2}$ obtained using the Zimm plot, Berry plot or Guinier plot is an optical weight size. Due to the different optical average methods of the root mean-square radius of gyration $\left\langle R_{g}^{2}\right\rangle ^{1/2}$ and apparent hydrodynamic radius $R_{app,h}$ and the complex hydrodynamic structures, the dimensionless shape parameter $\rho =\left\langle R_{g}^{2}\right\rangle ^{1/2}/R_{app,h}$ has lost the physical significance to judge the shapes of nano-particles.

physics.chem-ph

Time-Lapse Video-Based Embryo Grading via Complementary Spatial-Temporal Pattern Mining

Artificial intelligence has recently shown promise in automated embryo selection for In-Vitro Fertilization (IVF). However, current approaches either address partial embryo evaluation lacking holistic quality assessment or target clinical outcomes inevitably confounded by extra-embryonic factors, both limiting clinical utility. To bridge this gap, we propose a new task called Video-Based Embryo Grading - the first paradigm that directly utilizes full-length time-lapse monitoring (TLM) videos to predict embryologists' overall quality assessments. To support this task, we curate a real-world clinical dataset comprising over 2,500 TLM videos, each annotated with a grading label indicating the overall quality of embryos. Grounded in clinical decision-making principles, we propose a Complementary Spatial-Temporal Pattern Mining (CoSTeM) framework that conceptually replicates embryologists' evaluation process. The CoSTeM comprises two branches: (1) a morphological branch using a Mixture of Cross-Attentive Experts layer and a Temporal Selection Block to select discriminative local structural features, and (2) a morphokinetic branch employing a Temporal Transformer to model global developmental trajectories, synergistically integrating static and dynamic determinants for grading embryos. Extensive experimental results demonstrate the superiority of our design. This work provides a valuable methodological framework for AI-assisted embryo selection. The dataset and source code will be publicly available upon acceptance.

cs.CV

ExoGait-MS: Learning Periodic Dynamics with Multi-Scale Graph Network for Exoskeleton Gait Recognition

Current exoskeleton control methods often face challenges in delivering personalized treatment. Standardized walking gaits can lead to patient discomfort or even injury. Therefore, personalized gait is essential for the effectiveness of exoskeleton robots, as it directly impacts their adaptability, comfort, and rehabilitation outcomes for individual users. To enable personalized treatment in exoskeleton-assisted therapy and related applications, accurate recognition of personal gait is crucial for implementing tailored gait control. The key challenge in gait recognition lies in effectively capturing individual differences in subtle gait features caused by joint synergy, such as step frequency and step length. To tackle this issue, we propose a novel approach, which uses Multi-Scale Global Dense Graph Convolutional Networks (GCN) in the spatial domain to identify latent joint synergy patterns. Moreover, we propose a Gait Non-linear Periodic Dynamics Learning module to effectively capture the periodic characteristics of gait in the temporal domain. To support our individual gait recognition task, we have constructed a comprehensive gait dataset that ensures both completeness and reliability. Our experimental results demonstrate that our method achieves an impressive accuracy of 94.34% on this dataset, surpassing the current state-of-the-art (SOTA) by 3.77%. This advancement underscores the potential of our approach to enhance personalized gait control in exoskeleton-assisted therapy.

cs.RO

RoboAct-CLIP: Video-Driven Pre-training of Atomic Action Understanding for Robotics

Visual Language Models (VLMs) have emerged as pivotal tools for robotic systems, enabling cross-task generalization, dynamic environmental interaction, and long-horizon planning through multimodal perception and semantic reasoning. However, existing open-source VLMs predominantly trained for generic vision-language alignment tasks fail to model temporally correlated action semantics that are crucial for robotic manipulation effectively. While current image-based fine-tuning methods partially adapt VLMs to robotic applications, they fundamentally disregard temporal evolution patterns in video sequences and suffer from visual feature entanglement between robotic agents, manipulated objects, and environmental contexts, thereby limiting semantic decoupling capability for atomic actions and compromising model generalizability.To overcome these challenges, this work presents RoboAct-CLIP with dual technical contributions: 1) A dataset reconstruction framework that performs semantic-constrained action unit segmentation and re-annotation on open-source robotic videos, constructing purified training sets containing singular atomic actions (e.g., "grasp"); 2) A temporal-decoupling fine-tuning strategy based on Contrastive Language-Image Pretraining (CLIP) architecture, which disentangles temporal action features across video frames from object-centric characteristics to achieve hierarchical representation learning of robotic atomic actions.Experimental results in simulated environments demonstrate that the RoboAct-CLIP pretrained model achieves a 12% higher success rate than baseline VLMs, along with superior generalization in multi-object manipulation tasks.

cs.RO

GenM$^3$: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation

Scaling up motion datasets is crucial to enhance motion generation capabilities. However, training on large-scale multi-source datasets introduces data heterogeneity challenges due to variations in motion content. To address this, we propose Generative Pretrained Multi-path Motion Model (GenM\(^3\)), a comprehensive framework designed to learn unified motion representations. GenM\(^3\) comprises two components: 1) a Multi-Expert VQ-VAE (MEVQ-VAE) that adapts to different dataset distributions to learn a unified discrete motion representation, and 2) a Multi-path Motion Transformer (MMT) that improves intra-modal representations by using separate modality-specific pathways, each with densely activated experts to accommodate variations within that modality, and improves inter-modal alignment by the text-motion shared pathway. To enable large-scale training, we integrate and unify 11 high-quality motion datasets (approximately 220 hours of motion data) and augment it with textual annotations (nearly 10,000 motion sequences labeled by a large language model and 300+ by human experts). After training on our integrated dataset, GenM\(^3\) achieves a state-of-the-art FID of 0.035 on the HumanML3D benchmark, surpassing state-of-the-art methods by a large margin. It also demonstrates strong zero-shot generalization on IDEA400 dataset, highlighting its effectiveness and adaptability across diverse motion scenarios.

cs.CV

TeraSim: Uncovering Unknown Unsafe Events for Autonomous Vehicles through Generative Simulation

Traffic simulation is essential for autonomous vehicle (AV) development, enabling comprehensive safety evaluation across diverse driving conditions. However, traditional rule-based simulators struggle to capture complex human interactions, while data-driven approaches often fail to maintain long-term behavioral realism or generate diverse safety-critical events. To address these challenges, we propose TeraSim, an open-source, high-fidelity traffic simulation platform designed to uncover unknown unsafe events and efficiently estimate AV statistical performance metrics, such as crash rates. TeraSim is designed for seamless integration with third-party physics simulators and standalone AV stacks, to construct a complete AV simulation system. Experimental results demonstrate its effectiveness in generating diverse safety-critical events involving both static and dynamic agents, identifying hidden deficiencies in AV systems, and enabling statistical performance evaluation. These findings highlight TeraSim's potential as a practical tool for AV safety assessment, benefiting researchers, developers, and policymakers. The code is available at https://github.com/mcity/TeraSim.

cs.RO

MoReFun: Past-Movement Guided Motion Representation Learning for Future Motion Prediction and Understanding

3D human motion prediction aims to generate coherent future motions from observed sequences, yet existing end-to-end regression frameworks often fail to capture complex dynamics and tend to produce temporally inconsistent or static predictions-a limitation rooted in representation shortcutting, where models rely on superficial cues rather than learning meaningful motion structure. We propose a two-stage self-supervised framework that decouples representation learning from prediction. In the pretraining stage, the model performs unified past-future self-reconstruction, reconstructing the past sequence while recovering masked joints in the future sequence under full historical guidance. A velocity-based masking strategy selects highly dynamic joints, forcing the model to focus on informative motion components and internalize the statistical dependencies between past and future states without regression interference. In the fine-tuning stage, the pretrained model predicts the entire future sequence, now treated as fully masked, and is further equipped with a lightweight future-text prediction head for joint optimization of low-level motion prediction and high-level motion understanding. Experiments on Human3.6M, 3DPW, and AMASS show that our method reduces average prediction errors by 8.8% over state-of-the-art methods while achieving competitive future-motion understanding performance compared to LLM-based models. Code is available at: https://github.com/JunyuShi02/MoReFun

cs.CV

Fusion Makes Perfection: An Efficient Multi-Grained Matching Approach for Zero-Shot Relation Extraction

Predicting unseen relations that cannot be observed during the training phase is a challenging task in relation extraction. Previous works have made progress by matching the semantics between input instances and label descriptions. However, fine-grained matching often requires laborious manual annotation, and rich interactions between instances and label descriptions come with significant computational overhead. In this work, we propose an efficient multi-grained matching approach that uses virtual entity matching to reduce manual annotation cost, and fuses coarse-grained recall and fine-grained classification for rich interactions with guaranteed inference speed. Experimental results show that our approach outperforms the previous State Of The Art (SOTA) methods, and achieves a balance between inference efficiency and prediction accuracy in zero-shot relation extraction tasks. Our code is available at https://github.com/longls777/EMMA.

cs.CL

My Understanding for Static and Dynamic Light Scattering

The results obtained using my computing program are consistent with the values obtained twenty years ago. It also makes me believe that how to obtain particle size information using the Light Scattering technique needs to be reconsidered. Static Light Scattering (SLS) and Dynamic Light Scattering (DLS) are very important techniques to study the characteristics of nano-particles in dispersion. The data of SLS is determined by the optical characteristic and the measured values of DLS are determined by optical and hydrodynamic characteristics of different size nano-particles in dispersion. Then considering the optical characteristic of nano-particles and using the SLS technique further, the size distribution can be measured accurately and is also consistent with the results measured using the TEM technique. Based on the size distribution obtained using the SLS or TEM technique and the relation between the static and hydrodynamic radii, all the expected and measured values of $g^{\left( 2\right) }\left( \tau \right) $ investigated are very well consistent. Since the data measured using the DLS technique contains the information of the optical and hydrodynamic properties of nano-particles together, therefore the accurate size distribution cannot be obtained from the experimental data of $g^{\left( 2\right) }\left( \tau \right) $ for an unknown sample. The traditional particle information: apparent hydrodynamic radius and polydispersity index measured using the DLS technique are determined by the optical and hydrodynamic characteristics and size distribution together. They cannot represent a number distribution of nano-particles in dispersion. Using the light scattering technique not only can measure the size distribution accurately but also can provide a method to understand the optical and hydrodynamic characteristics of nano-particles.

physics.chem-ph

Aggregate Model of District Heating Network for Integrated Energy Dispatch: A Physically Informed Data-Driven Approach

The district heating network (DHN) is essential in enhancing the operational flexibility of integrated energy systems (IES). Yet, it is hard to obtain an accurate and concise DHN model for the operation owing to complicated network features and imperfect measurements. Considering this, this paper proposes a physical-ly informed data-driven aggregate model (AGM) for the DHN, providing a concise description of the source-load relationship of DHN without exposing network details. First, we derive the analytical relationship between the state variables of the source and load nodes of the DHN, offering a physical fundament for the AGM. Second, we propose a physics-informed estimator for the AGM that is robust to low-quality measurements, in which the physical constraints associated with the parameter normalization and sparsity are embedded to improve the accuracy and robustness. Finally, we propose a physics-enhanced algorithm to solve the nonlinear estimator with non-closed constraints efficiently. Simulation results verify the effectiveness of the proposed method.

eess.SY

Inelastic electron transfer in olfaction: multiphonons processes

Inelastic electron transfer being regarded as one of the potential mechanisms to explain the odorant recognition in the atomic-scale processes is still a matter of intense debate. Here, we propose multiphonon processes of electrons transfer using the Markvart model and calculate their lifetimes with values of key parameters widely adopted in olfactory systems. We find that these multiphonon processes are as quickly as the single phonon process, which suggest that contributions from different phonon modes of odorant molecule for electrons transfer in olfaction should be included. Meanwhile, temperature dependence of electron transfer could be analyzed effectively based on the reorganization energy is expanded into the linewidth of multiphonon processes. Our theoretical results not only enrich the knowledge of the mechanism of the olfaction recognition, but also provide insights for quantum processes in the biological system.

physics.bio-ph