Searcharxiv⌕ Search

arXiv subjects

Rui Sun

Publications and source records attributed to Rui Sun.

At least 55 records · Page 3Linked to original sources

Advancing Opinion Dynamics Modeling with Neural Diffusion-Convection-Reaction Equation

Advanced opinion dynamics modeling is vital for deciphering social behavior, emphasizing its role in mitigating polarization and securing cyberspace. To synergize mechanistic interpretability with data-driven flexibility, recent studies have explored the integration of Physics-Informed Neural Networks (PINNs) for opinion modeling. Despite this promise, existing methods are tailored to incomplete priors, lacking a comprehensive physical system to integrate dynamics from local, global, and endogenous levels. Moreover, penalty-based constraints adopted in existing methods struggle to deeply encode physical priors, leading to optimization pathologies and discrepancy between latent representations and physical transparency. To this end, we offer a physical view to interpret opinion dynamics via Diffusion-Convection-Reaction (DCR) system inspired by interacting particle theory. Building upon the Neural ODEs, we define the neural opinion dynamics to coordinate neural networks with physical priors, and further present the OPINN, a physics-informed neural framework for opinion dynamics modeling. Evaluated on real-world and synthetic datasets, OPINN achieves state-of-the-art performance in opinion evolution forecasting, offering a promising paradigm for the nexus of cyber, physical, and social systems.

cs.AI↗

Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring

Accurate cell instance segmentation is foundational for digital pathology analysis. Existing methods based on contour detection and distance mapping still face significant challenges in processing complex and dense cellular regions. Graph coloring-based methods provide a new paradigm for this task, yet the effectiveness of this paradigm in real-world scenarios with dense overlaps and complex topologies has not been verified. Addressing this issue, we release a large-scale dataset GBC-FS 2025, which contains highly complex and dense sub-cellular nuclear arrangements. We conduct the first systematic analysis of the chromatic properties of cell adjacency graphs across four diverse datasets and reveal an important discovery: most real-world cell graphs are non-bipartite, with a high prevalence of odd-length cycles (predominantly triangles). This makes simple 2-coloring theory insufficient for handling complex tissues, while higher-chromaticity models would cause representational redundancy and optimization difficulties. Building on this observation of complex real-world contexts, we propose Disco (Densely-overlapping Cell Instance Segmentation via Adjacency-aware COllaborative Coloring), an adjacency-aware framework based on the "divide and conquer" principle. It uniquely combines a data-driven topological labeling strategy with a constrained deep learning system to resolve complex adjacency conflicts. First, "Explicit Marking" strategy transforms the topological challenge into a learnable classification task by recursively decomposing the cell graph and isolating a "conflict set." Second, "Implicit Disambiguation" mechanism resolves ambiguities in conflict regions by enforcing feature dissimilarity between different instances, enabling the model to learn separable feature representations.

cs.CV↗

MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition

Capsule Network (CapsNet) has demonstrated significant potential in visual recognition by capturing spatial relationships and part-whole hierarchies for learning equivariant feature representations. However, existing CapsNet and variants often rely on a single high-level feature map, overlooking the rich complementary information from multi-scale features. Furthermore, conventional feature fusion strategies (e.g., addition and concatenation) struggle to reconcile multi-scale feature discrepancies, leading to suboptimal classification performance. To address these limitations, we propose the Multi-Scale Patchify Capsule Network (MSPCaps), a novel architecture that integrates multi-scale feature learning and efficient capsule routing. Specifically, MSPCaps consists of three key components: a Multi-Scale ResNet Backbone (MSRB), a Patchify Capsule Layer (PatchifyCaps), and Cross-Agreement Routing (CAR) blocks. First, the MSRB extracts diverse multi-scale feature representations from input images, preserving both fine-grained details and global contextual information. Second, the PatchifyCaps partitions these multi-scale features into primary capsules using a uniform patch size, equipping the model with the ability to learn from diverse receptive fields. Finally, the CAR block adaptively routes the multi-scale capsules by identifying cross-scale prediction pairs with maximum agreement. Unlike the simple concatenation of multiple self-routing blocks, CAR ensures that only the most coherent capsules contribute to the final voting. Our proposed MSPCaps achieves remarkable scalability and superior robustness, consistently surpassing multiple baseline methods in terms of classification accuracy, with configurations ranging from a highly efficient Tiny model (344.3K parameters) to a powerful Large model (10.9M parameters), highlighting its potential in advancing feature representation learning.

cs.CV↗

AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning

Evaluating the clinical correctness and reasoning fidelity of automatically generated medical imaging reports remains a critical yet unresolved challenge. Existing evaluation methods often fail to capture the structured diagnostic logic that underlies radiological interpretation, resulting in unreliable judgments and limited clinical relevance. We introduce AgentsEval, a multi-agent stream reasoning framework that emulates the collaborative diagnostic workflow of radiologists. By dividing the evaluation process into interpretable steps including criteria definition, evidence extraction, alignment, and consistency scoring, AgentsEval provides explicit reasoning traces and structured clinical feedback. We also construct a multi-domain perturbation-based benchmark covering five medical report datasets with diverse imaging modalities and controlled semantic variations. Experimental results demonstrate that AgentsEval delivers clinically aligned, semantically faithful, and interpretable evaluations that remain robust under paraphrastic, semantic, and stylistic perturbations. This framework represents a step toward transparent and clinically grounded assessment of medical report generation systems, fostering trustworthy integration of large language models into clinical practice.

cs.AI↗

A Control Theoretic Approach to Decentralized AI Economy Stabilization via Dynamic Buyback-and-Burn Mechanisms

The democratization of artificial intelligence through decentralized networks represents a paradigm shift in computational provisioning, yet the long-term viability of these ecosystems is critically endangered by the extreme volatility of their native economic layers. Current tokenomic models, which predominantly rely on static or threshold-based buyback heuristics, are ill-equipped to handle complex system dynamics and often function pro-cyclically, exacerbating instability during market downturns. To bridge this gap, we propose the Dynamic-Control Buyback Mechanism (DCBM), a formalized control-theoretic framework that utilizes a Proportional-Integral-Derivative (PID) controller with strict solvency constraints to regulate the token economy as a dynamical system. Extensive agent-based simulations utilizing Jump-Diffusion processes demonstrate that DCBM fundamentally outperforms static baselines, reducing token price volatility by approximately 66% and lowering operator churn from 19.5% to 8.1% in high-volatility regimes. These findings establish that converting tokenomics from static rules into continuous, structurally constrained control loops is a necessary condition for secure and sustainable decentralized intelligence networks.

cs.GT↗

Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective

Pseudo-label learning is widely used in semantic segmentation, particularly in label-scarce scenarios such as unsupervised domain adaptation (UDA) and semisupervised learning (SSL). Despite its success, this paradigm can generate erroneous pseudo-labels, which are further amplified during training due to utilization of one-hot encoding. To address this issue, we propose ECOCSeg, a novel perspective for segmentation models that utilizes error-correcting output codes (ECOC) to create a fine-grained encoding for each class. ECOCSeg offers several advantages. First, an ECOC-based classifier is introduced, enabling model to disentangle classes into attributes and handle partial inaccurate bits, improving stability and generalization in pseudo-label learning. Second, a bit-level label denoising mechanism is developed to generate higher-quality pseudo-labels, providing adequate and robust supervision for unlabeled images. ECOCSeg can be easily integrated with existing methods and consistently demonstrates significant improvements on multiple UDA and SSL benchmarks across different segmentation architectures. Code is available at https://github.com/Woof6/ECOCSeg.

cs.CV↗

The ${\cal N}=1$ supersymmetric Pati-Salam models with extra $SU(2)_{L_2/R_2}$ gauge symmetry from intersecting D6-branes

By introducing an extra stack of D6-branes to standard ${\cal N}=1$ supersymmetric Pati-Salam models, we extend the landscape of its complete search. In this construction, the $d$-stack of D6-branes is introduced besides the standard $a,~b,~c$-stacks. More intersections from the extra stacks of D6-branes appear, and thus Higgs/Higgs-like particles arise from more origins. Among these models, we find eight new classes of ${\cal N}=1$ supersymmetric Pati-Salam models with gauge symmetries $SU(4)_C\times SU(2)_L\times SU(2)_{R_1}\times SU(2)_{R_2}$ and $SU(4)_C\times SU(2)_{L_1}\times SU(2)_{R}\times SU(2)_{L_2}$, where $d$-stack of D6-branes carries the gauge symmetries $SU(2)_{R_2}$ and $SU(2)_{L_2}$, respectively. The $SU(2)_{L_1/R_1} \times SU(2)_{L_2/R_2}$ can be broken down to the diagonal $SU(2)_{L/R}$ gauge symmetry via bifundamental Higgs fields. In such a way, we for the first time successfully constructed three-family supersymmetric Pati-Salam models from non-rigid D6-branes with extra $d$-stacks of D6-branes as visible sectors. Interestingly, by introducing extra stack of D6-branes to the standard supersymmetric Pati-Salam models, the number of filler brane reduces in general, and eventually the models without any $USp(N)$ gauge symmetry present. This reduces the exotic particles from filler brane intersection yet provides more vector-like particles from ${\cal N}=2$ subsector that are useful in renormalization group equation evolution as an advantage. Moreover, interesting degeneracy behavior with the same gauge coupling ratio exists in certain class of models.

hep-th↗

Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification

Reinforcement Learning (RL) has enabled Large Language Models (LLMs) to achieve remarkable reasoning in domains like mathematics and coding, where verifiable rewards provide clear signals. However, extending this paradigm to financial decision is challenged by the market's stochastic nature: rewards are verifiable but inherently noisy, causing standard RL to degenerate into reward hacking. To address this, we propose Trade-R1, a model training framework that bridges verifiable rewards to stochastic environments via process-level reasoning verification. Our key innovation is a verification method that transforms the problem of evaluating reasoning over lengthy financial documents into a structured Retrieval-Augmented Generation (RAG) task. We construct a triangular consistency metric, assessing pairwise alignment between retrieved evidence, reasoning chains, and decisions to serve as a validity filter for noisy market returns. We explore two reward integration strategies: Fixed-effect Semantic Reward (FSR) for stable alignment signals, and Dynamic-effect Semantic Reward (DSR) for coupled magnitude optimization. Experiments on different country asset selection demonstrate that our paradigm reduces reward hacking, with DSR achieving superior cross-market generalization while maintaining the highest reasoning consistency.

cs.AI↗

Tracing the Heart's Pathways: ECG Representation Learning from a Cardiac Conduction Perspective

The multi-lead electrocardiogram (ECG) stands as a cornerstone of cardiac diagnosis. Recent strides in electrocardiogram self-supervised learning (eSSL) have brightened prospects for enhancing representation learning without relying on high-quality annotations. Yet earlier eSSL methods suffer a key limitation: they focus on consistent patterns across leads and beats, overlooking the inherent differences in heartbeats rooted in cardiac conduction processes, while subtle but significant variations carry unique physiological signatures. Moreover, representation learning for ECG analysis should align with ECG diagnostic guidelines, which progress from individual heartbeats to single leads and ultimately to lead combinations. This sequential logic, however, is often neglected when applying pre-trained models to downstream tasks. To address these gaps, we propose CLEAR-HUG, a two-stage framework designed to capture subtle variations in cardiac conduction across leads while adhering to ECG diagnostic guidelines. In the first stage, we introduce an eSSL model termed Conduction-LEAd Reconstructor (CLEAR), which captures both specific variations and general commonalities across heartbeats. Treating each heartbeat as a distinct entity, CLEAR employs a simple yet effective sparse attention mechanism to reconstruct signals without interference from other heartbeats. In the second stage, we implement a Hierarchical lead-Unified Group head (HUG) for disease diagnosis, mirroring clinical workflow. Experimental results across six tasks show a 6.84% improvement, validating the effectiveness of CLEAR-HUG. This highlights its ability to enhance representations of cardiac conduction and align patterns with expert diagnostic guidelines.

cs.LG↗

Generalized Three-Family Supersymmetric Pati-Salam Models from Type IIA Intersecting D6-Branes

Generalizing three-family chiral fermion conditions to $I_{ac}=-(3+h)$ and $I_{ac'}=h$, with positive integer $h$, we extend the landscape of three-family ${\cal N}=1$ supersymmetric Pati-Salam models in a broader region. Differing from the former investigation with $I_{ac}=-3$ and $I_{ac'}=0$, we do not restrict that the $a$ stack of D6-branes must be parallel to the orientifold image of the $c$-stack along one of the three two-tori. In this investigation, without the simple parallel construction, we find four new classes of supersymmetric Pati-Salam models that are allowed by the extended three generation condition with $I_{ac}=3, I_{ac'}=-6$ and $I_{ac}=-1, I_{ac'}=-2$ through the intersections of $a$- and $c/c'$-branes. Moreover, with the $SU(2)_{L'}$ gauge coupling realized from $SU(2)_{L_1}\times SU(2)_{L_2}$ symmetry breaking, the canonical normalization requirement of the gauge kinetic term provides an alternative approach that can be imposed before the renormalization group equation evolution for $SU(2)_{L'}$ gauge coupling. This turns out to be an effective mechanism to realize the string-scale gauge coupling relation, especially for the new supersymmetric Pati-Salam models with large $g_b/g_a$ ratio. We show that this symmetry-breaking modified renormalization group evolution can highly suppress $g_b/g_a$, and finally realizes string-scale gauge coupling relations for the extended supersymmetric Pati-Salam models as well.

hep-th↗

From Pretraining to Privacy: Federated Ultrasound Foundation Model with Self-Supervised Learning

Ultrasound imaging is widely used in clinical diagnosis due to its non-invasive nature and real-time capabilities. However, traditional ultrasound diagnostics relies heavily on physician expertise and is often hampered by suboptimal image quality, leading to potential diagnostic errors. While artificial intelligence (AI) offers a promising solution to enhance clinical diagnosis by detecting abnormalities across various imaging modalities, existing AI methods for ultrasound face two major challenges. First, they typically require vast amounts of labeled medical data, raising serious concerns regarding patient privacy. Second, most models are designed for specific tasks, which restricts their broader clinical utility. To overcome these challenges, we present UltraFedFM, an innovative privacy-preserving ultrasound foundation model. UltraFedFM is collaboratively pre-trained using federated learning across 16 distributed medical institutions in 9 countries, leveraging a dataset of over 1 million ultrasound images covering 19 organs and 10 ultrasound modalities. This extensive and diverse data, combined with a secure training framework, enables UltraFedFM to exhibit strong generalization and diagnostic capabilities. It achieves an average area under the receiver operating characteristic curve (AUROC) of 0.927 for disease diagnosis and a dice similarity coefficient (DSC) of 0.878 for lesion segmentation. Notably, UltraFedFM surpasses the diagnostic accuracy of mid-level ultrasonographers (4-8 years of experience) and matches the performance of expert-level sonographers (10+ years of experience) in the joint diagnosis of 8 common systemic diseases.c These findings indicate that UltraFedFM can significantly enhance clinical diagnostics while safeguarding patient privacy, marking a significant advancement in AI-driven ultrasound imaging for future clinical applications.

eess.IV↗

VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel

Accurate vessel segmentation is critical for clinical applications such as disease diagnosis and surgical planning, yet remains challenging due to thin, branching structures and low texture contrast. While foundation models like the Segment Anything Model (SAM) have shown promise in generic segmentation, they perform sub-optimally on vascular structures. In this work, we present VesSAM, a powerful and efficient framework tailored for 2D vessel segmentation. VesSAM integrates (1) a convolutional adapter to enhance local texture features, (2) a multi-prompt encoder that fuses anatomical prompts, including skeletons, bifurcation points, and segment midpoints, via hierarchical cross-attention, and (3) a lightweight mask decoder to reduce jagged artifacts. We also introduce an automated pipeline to generate structured multi-prompt annotations, and curate a diverse benchmark dataset spanning 8 datasets across 5 imaging modalities. Experimental results demonstrate that VesSAM consistently outperforms state-of-the-art PEFT-based SAM variants by over 10% Dice and 13% IoU, and achieves competitive performance compared to fully fine-tuned methods, with significantly fewer parameters. VesSAM also generalizes well to out-of-distribution (OoD) settings, outperforming all baselines in average OoD Dice and IoU.

cs.CV↗

Towards Unsupervised Domain Bridging via Image Degradation in Semantic Segmentation

Semantic segmentation suffers from significant performance degradation when the trained network is applied to a different domain. To address this issue, unsupervised domain adaptation (UDA) has been extensively studied. Despite the effectiveness of selftraining techniques in UDA, they still overlook the explicit modeling of domain-shared feature extraction. In this paper, we propose DiDA, an unsupervised domain bridging approach for semantic segmentation. DiDA consists of two key modules: (1) Degradation-based Intermediate Domain Construction, which creates continuous intermediate domains through simple image degradation operations to encourage learning domain-invariant features as domain differences gradually diminish; (2) Semantic Shift Compensation, which leverages a diffusion encoder to disentangle and compensate for semantic shift information with degraded timesteps, preserving discriminative representations in the intermediate domains. As a plug-and-play solution, DiDA supports various degradation operations and seamlessly integrates with existing UDA methods. Extensive experiments on multiple domain adaptive semantic segmentation benchmarks demonstrate that DiDA consistently achieves significant performance improvements across all settings. Code is available at https://github.com/Woof6/DiDA.

cs.CV↗

Balanced Learning for Domain Adaptive Semantic Segmentation

Unsupervised domain adaptation (UDA) for semantic segmentation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Despite the effectiveness of self-training techniques in UDA, they struggle to learn each class in a balanced manner due to inherent class imbalance and distribution shift in both data and label space between domains. To address this issue, we propose Balanced Learning for Domain Adaptation (BLDA), a novel approach to directly assess and alleviate class bias without requiring prior knowledge about the distribution shift. First, we identify over-predicted and under-predicted classes by analyzing the distribution of predicted logits. Subsequently, we introduce a post-hoc approach to align the logits distributions across different classes using shared anchor distributions. To further consider the network's need to generate unbiased pseudo-labels during self-training, we estimate logits distributions online and incorporate logits correction terms into the loss function. Moreover, we leverage the resulting cumulative density as domain-shared structural knowledge to connect the source and target domains. Extensive experiments on two standard UDA semantic segmentation benchmarks demonstrate that BLDA consistently improves performance, especially for under-predicted classes, when integrated into various existing methods. Code is available at https://github.com/Woof6/BLDA.

cs.CV↗

Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs

Depth-wise pruning accelerates LLM inference in resource-constrained scenarios but suffers from performance degradation due to direct removal of entire Transformer layers. This paper reveals ``Patch-like'' redundancy across layers via correlation analysis of the outputs of different layers in reproducing kernel Hilbert space, demonstrating consecutive layers exhibit high functional similarity. Building on this observation, this paper proposes Sliding-Window Merging (SWM) - a dynamic compression method that selects consecutive layers from top to bottom using a pre-defined similarity threshold, and compacts patch-redundant layers through a parameter consolidation, thereby simplifying the model structure while maintaining its performance. Extensive experiments on LLMs with various architectures and different parameter scales show that our method outperforms existing pruning techniques in both zero-shot inference performance and retraining recovery quality after pruning. In particular, in the experiment with 35% pruning on the Vicuna-7B model, our method achieved a 1.654% improvement in average performance on zero-shot tasks compared to the existing method. Moreover, we further reveal the potential of combining depth pruning with width pruning to enhance the pruning effect. Our codes are available at https://github.com/920927/SLM-a-sliding-layer-merging-method.

cs.CV↗

The Global Minimum Tax, Investment Incentives and Asymmetric Tax Competition

This paper investigates the OECD's global minimum tax (GMT) in a formal model of tax competition between asymmetric countries. We consider both profit shifting and real responses of multinational enterprises, and highlight the role of the substance-based income exclusion (SBIE) in investment incentives and tax rate setting. The GMT reduces the true tax rate differential and benefits the large country, while the revenue effect is generally ambiguous for the small country. In the short run where tax rates are fixed, the GMT reduces the small country's revenue if profit shifting costs are low and increases it otherwise. In the long run where countries adjust tax rates, the GMT reshapes the tax game and the competition pattern. We reveal that the minimum rate binds the small country only if it is low. With the rise of the GMT rate, countries will set tax rates below the minimum to boost capital investments and collect top-up taxes. Simulations show that a moderate GMT rate can raise both countries' revenues and the large country's welfare in the long run. However, it may reduce the small country's welfare if the welfare weight of private income is high.

econ.GN↗

The purely leptonic and semileptonic decays of $D^{*}_{s}$ meson

With the potential prospects of the $D^{*}_{s}$ at high-luminosity heavy-flavor experiments in the future, we investigated the CKM-favored and tree-dominated leptonic $D^{*}_{s}\to\ell\barν_{\ell}$ and semileptonic $D^{*}_{s}\to M\ell\barν_{\ell}$ ($M=ϕ, η^{(\prime)}$ and $\ell=e, μ$) weak decays in the Standard Model (SM). The theoretical predictions and some discussions for the observable quantities including the total width of $D^{*}_{s}$ mesons, the branching fractions of leptonic $D^{*}_{s}\to\ell\barν_{\ell}$ and semileptonic $D^{*}_{s}\to M\ell\barν_{\ell}$ weak decays, the lepton spin asymmetry and forward-backward asymmetry are presented. Numerically, the weak decays of $D^{*}_{s}\to\ell\barν_{\ell}$ and $D^{*}_{s} \to M\ell\barν_{\ell}$ have relatively large branching fractions of the order $\mathcal{O}(10^{-5})$ and $\mathcal{O}(10^{-7})$ respectively, which are expected to be observed in future experiments.

hep-ph↗

High thermal conductivity of rutile-GeO$_2$ films grown by MOCVD: $52.9~\mathrm{W\,m^{-1}\,K^{-1}}$

Rutile germanium dioxide (r-GeO2) has recently emerged as a promising ultrawide-bandgap (UWBG) semiconductor owing to its wide bandgap (~4.4-5.1 eV), ambipolar doping potential, and high theoretical thermal conductivity. However, experimental data on the thermal conductivity of r-GeO2 epitaxial layers have not been reported, primarily due to challenges in phase control and surface roughness. Here, we report a high thermal conductivity of 52.9 +/- 6.6 W m^-1 K^-1 for high-quality (002) r-GeO2 films grown by metal-organic chemical vapor deposition (MOCVD) and characterized using time-domain thermoreflectance (TDTR). The phase control was achieved through a seed-driven stepwise crystallization (SDSC) approach, and the surface roughness was significantly reduced from 76 nm to 16 nm (locally as low as 1 A) via chemical mechanical polishing (CMP). These results highlight the promise of r-GeO2 as a UWBG oxide platform for power electronics applications.

cond-mat.mtrl-sci↗