SearcharxivSearch

arXiv subjects

Rui Shi

Publications and source records attributed to Rui Shi.

At least 19 recordsLinked to original sources

RailSyn: Diagnosis-Guided Image Generation for Traceable Data Completion in Railway Foreign Object Detection

Railway foreign object detection (RFOD) is critical to safe railway operation, yet scarce real positive samples incompletely represent task-relevant variations in object scale, intrusion relation, railway scene, illumination, and adverse weather. Existing synthetic augmentation can improve RFOD detection, but its gains lack an explicit account of the task-relevant deficiencies complemented by the generated data. We therefore introduce RailSyn, a diagnosis-guided framework comprising a real-referenced Inspector and a requirement-aligned Generator. The Inspector constructs a variable-radius empirical cover from finite real observations to localize candidate completion regions and profile synthetic pools. The resulting audit identifies railway-context, intrusion-semantic, and visual-consistency requirements; the Generator addresses them through domain adaptation, agent-planned placement and physical contact relations, and plan-consistent conditional refinement. Using the Inspector, we further trace representation-space changes across generation variants; the complete system attains a local-shell occupation of $C_{gap}$ to 13.64%, which measures generated coverage of real-derived completion regions. Extensive experiments show AP50--95 gains of up to 4.9 points and consistent improvements across nine mainstream detectors, demonstrating broad cross-architecture utility.

cs.CV

RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation

Small-object detection under long-tailed data distributions is a fundamental yet challenging problem in multimedia. Railway Foreign Object Detection (RFOD) epitomizes this challenge with easily confused small intrusions and scarce samples. To address these issues, we propose a generative-augmented detection paradigm that leverages multimodal image generation to enrich the feature space of rare and small objects. We first construct RailGen, a multimodal image generation agent based on large models. Under semantic constraints, RailGen automatically invokes tools to generate railway scenes, calibrate intrusion positions, extract foreign objects, and fuse them into realistic intrusion effects. This process produces high-quality synthetic samples that effectively densify the feature representations of tail classes and complete the small-object feature space. Within this paradigm, we further propose FocalDEIM, a detection framework designed to enhance training with generated data. FocalDEIM improves dense matching with Focal Modulation for better small-object discrimination and adopts Focal Loss to emphasize hard samples, thereby alleviating blurred inter-class boundaries in complex railway scenes. Experimental results demonstrate that RailGen can generate high-quality small-scale foreign objects, reducing the object pixel area by up to 58x and 13.85x on average. Equipped with these challenging samples, our paradigm surpasses the baseline DEIM by 5.6% and 7.5% in mAP@50 and mAP@(50-95), respectively, and outperforms existing state-of-the-art methods. Ablation studies verify RailGen's feature-space enrichment and FocalDEIM's boundary discrimination. The paradigm provides an effective multimodal generative solution for long-tailed small-object detection in safety-critical applications.

cs.CV

Kaplansky's second test problem in operator algebras

Kaplansky's second test problem on similarity asks: if $T$ and $S$ are elements in a unital Banach algebra $\mathcal{B}$ and $T\oplus T$ is similar to $S\oplus S$ in $\mathbb{M}_2(\mathcal{B})$, is $T$ similar to $S$ in $\mathcal{B}$? We answer this problem affirmatively if $T$ is an operator with property $(J)$ in a type $\mathrm{I}_n$ von Neumann algebra $\mathcal{M}$, i.e., $\{T\}'\cap\mathcal{M}$ contains a bounded maximal abelian family of idempotents. Moreover, the condition of property $(J)$ can be removed for $1\leqslant n\leqslant 3$. A similar result is proved if $T$ is an element in a unital Banach algebra $\mathcal{B}$ with essentially finite-dimensional commutant, i.e., the relative commutant of $T$ in $\mathcal{B}$ is finite-dimensional modulo its Jacobson radical. Finally, we point out that one of our main results can be applied to the implementation of local unitary (LU) equivalence of quantum states.

math.FA

Bargmann invariants and local unitary equivalence

In this paper, we study the local unitary equivalence of quantum states on $\mathbb{C}^n\otimes\mathbb{C}^n$, which is an important notion in quantum information theory. For the case $n=2$, the local unitary orbits of two-qubit states are completely determined by their local unitary Bargmann invariants. We show that the local unitary Bargmann invariants do not form a complete set of invariants of local unitary orbits for $n\geqslant 3$, which negatively answers a problem proposed by L. Zhang and B. Xie.

quant-ph

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.

cs.CV

Detectors for CLASS-W2: The second 90 GHz telescope of the Cosmology Large Angular Scale Surveyor

The Cosmology Large Angular Scale Surveyor (CLASS) is measuring the Cosmic Microwave Background (CMB) polarization anisotropy on the largest angular scales (>1 degree) to probe the epochs of inflation and reionization. To enhance the CMB mapping speed, we have built, tested, and commissioned in August 2025 a second 90 GHz receiver (CLASS-W2) with a detector focal plane composed of feedhorn-coupled Transition Edge Sensor (TES) bolometers fabricated at NIST-Boulder. The focal plane consists of four modules, each containing 37 feedhorns coupling orthogonal polarizations onto two TES bolometers, for a total of 296 optically sensitive detectors. Laboratory tests show highly uniform TES properties with an array average critical temperature of 184+-3 mK, a thermal conductance of 460+-47 pW/K, and a normal resistance of 7.8+-0.3 mOhms. The detector array has an average band center frequency of 95.2 GHz with 28.3 GHz bandwidth, and achieves a detector yield of 94%. On-sky measurements indicate a mean detector optical load of 3.3 pW, corresponding to an antenna temperature of ~23 K. The array's average beam solid angle is 124 $\mu$sr, with a full width at half maximum of 0.592 degrees, and the end-to-end average optical efficiency is 0.37. We find that high-frequency 'blue-leak' radiation couples directly to the TES bolometer islands; adding a metal-mesh low-pass filter with cutoff frequency of 157 GHz in front of the focal plane suppresses the 'blue-leak' power by 0.9 pW. The four-module array achieves a noise-equivalent temperature of NET= 16 uKrtS. Adding this array has boosted the CLASS 90 GHz mapping speed by 41%.

astro-ph.IM

PriSrv: Privacy-Enhanced and Highly Usable Service Discovery in Wireless Communications

Service discovery is essential in wireless communications. However, existing protocols provide limited privacy protection, leaking sensitive device information and opening routes to network attacks. This paper proposes a private service discovery protocol, called PriSrv, which enables both service providers and clients to specify fine-grained authentication policies before establishing connections. PriSrv achieves this via a dual-layer matching architecture: an outer layer filters mismatched entities using public attributes, while an inner layer handles mutual authentication using selectively disclosed private attributes. As a core component, we introduce the primitive of anonymous credential-based matchmaking encryption (ACME), which enables dual-layer matching in a single step to achieve bilateral policy control, selective attribute disclosure, and multi-show unlinkability. To instantiate ACME, we design a fast anonymous credential (FAC) scheme providing constant-size credentials and efficient verification. We demonstrate PriSrv's interoperability by integrating it with popular wireless frameworks including EAP, mDNS, BLE, and AirDrop. Detailed formal security proofs and extensive performance evaluations across desktop, laptop, smartphone, and Raspberry Pi platforms demonstrate that PriSrv provides enhanced privacy guarantees with high usability, achieving secure discovery in less than one second on mainstream mobile devices.

cs.CR

PriSrv+: Privacy and Usability-Enhanced Wireless Service Discovery with Fast and Expressive Matchmaking Encryption

Service discovery is a fundamental process in wireless networks, enabling devices to find and communicate with services dynamically, and is critical for the seamless operation of modern systems like 5G and IoT. This paper introduces PriSrv+, an advanced privacy and usability-enhanced service discovery protocol for modern wireless networks and resource-constrained environments. PriSrv+ builds upon PriSrv (NDSS'24), by addressing critical limitations in expressiveness, privacy, scalability, and efficiency, while maintaining compatibility with widely-used wireless protocols such as mDNS, BLE, and Wi-Fi. A key innovation in PriSrv+ is the development of Fast and Expressive Matchmaking Encryption (FEME), the first matchmaking encryption scheme capable of supporting expressive access control policies with an unbounded attribute universe, allowing any arbitrary string to be used as an attribute. FEME significantly enhances the flexibility of service discovery while ensuring robust message and attribute privacy. Compared to PriSrv, PriSrv+ optimizes cryptographic operations, achieving 7.62* faster for encryption and 6.23* faster for decryption, and dramatically reduces ciphertext sizes by 87.33%. In addition, PriSrv+ reduces communication costs by 87.33% for service broadcast and 86.64% for anonymous mutual authentication compared with PriSrv. Formal security proofs confirm the security of FEME and PriSrv+. Extensive evaluations on multiple platforms demonstrate that PriSrv+ achieves superior performance, scalability, and efficiency compared to existing state-of-the-art protocols.

cs.CR

Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference

Quantization is essential for efficient large language model (LLM) inference, yet the dequantization step-converting low-bit weights back to high-precision for matrix multiplication has become a critical bottleneck on modern AI accelerators. On architectures with decoupled compute units (e.g., Ascend NPUs), dequantization operations can consume more cycles than the matrix multiplication itself, leaving the high-throughput tensor cores underutilized. This paper presents Multi-Scale Dequant (MSD), a quantization framework that removes weight/KV dequantization from the GEMM critical path. Instead of lifting low-bit weights to BF16 precision, MSD decomposes high-precision BF16 activations into multiple low-precision components, each of which can be multiplied directly with quantized weights via native hardware-accelerated GEMM. This approach shifts the computational paradigm from precision conversion to multi-scale approximation, avoiding INT8-to-BF16 weight conversion before GEMM. We instantiate MSD for two weight formats and derive tight error bounds for each. For INT8 weights (W4A16), two-pass INT8 decomposition achieves near 16 effective bits. For MXFP4 weights (W4A16), two-pass MXFP4 decomposition yields near 6.6 effective bits with error bound 1/64 per block surpassing single-pass MXFP8(5.24 bits) while maintaining the same effective GEMM compute time. We further derive closed-form latency and HBM traffic models showing that MSD avoids the Vector-Cube pipeline stall caused by dequantization and reduces KV cache HBM traffic by up to 2.5 times in attention. Numerical simulations on matrix multiplication and Flash Attention kernels confirm that MSD does not degrade accuracy compared to dequantization baselines, and in many settings achieves lower L2 error.

stat.ML

The similarity of irreducible operators in factors

An operator $T$ in a separable factor $\mathcal{M}$ is said to be irreducible in $\mathcal{M}$ if the von Neumann subalgebra $W^*(T)$ generated by $T$ is an irreducible subfactor of $\mathcal{M}$, i.e., $W^*(T)'\cap\mathcal{M}=\mathbb{C}I$. We say that $T$ is a single generator of $\mathcal{M}$ if $W^*(T)=\mathcal{M}$. In this paper, we study generators of separable factors related to maximal abelian self-adjoint subalgebras. As an application, we obtain a complete characterization of normal operators in separable factors which are similar to irreducible operators.

math.OA

Segmentation of Gray Matters and White Matters from Brain MRI data

Accurate segmentation of brain tissues such as gray matter and white matter from magnetic resonance imaging is essential for studying brain anatomy, diagnosing neurological disorders, and monitoring disease progression. Traditional methods, such as FSL FAST, produce tissue probability maps but often require task-specific adjustments and face challenges with diverse imaging conditions. Recent foundation models, such as MedSAM, offer a prompt-based approach that leverages large-scale pretraining. In this paper, we propose a modified MedSAM model designed for multi-class brain tissue segmentation. Our preprocessing pipeline includes skull stripping with FSL BET, tissue probability mapping with FSL FAST, and converting these into 2D axial, sagittal, coronal slices with multi-class labels (background, gray matter, and white matter). We extend MedSAM's mask decoder to three classes, freezing the pre-trained image encoder and fine-tuning the prompt encoder and decoder. Experiments on the IXI dataset achieve Dice scores up to 0.8751. This work demonstrates that foundation models like MedSAM can be adapted for multi-class medical image segmentation with minimal architectural modifications. Our findings suggest that such models can be extended to more diverse medical imaging scenarios in future work.

cs.CV

Can Large Language Models be a Cardinality Estimator? An Empirical study

Cardinality estimation (CardEst) still remains a challenging problem for DBMS. Recent years have witnessed the success of ML-based cardinality estimators in outperforming traditional methods. However, these solutions suffer from poor generalizability to new data or query distribution, inability to handle complex queries, and substantial data preparation overhead, thus preventing their wide adoption in the real-world DBMS. Some recent efforts have been dedicated to addressing some but not all of these issues. We notice that the recent emerging Large Language Models (LLMs) have shown their remarkable generalizability to unseen tasks, capabilities to understand complex programs, and power to perform data-efficient fine-tuning. In light of this, we propose to leverage LLMs to mitigate the above issues. Specifically, we carefully craft prompts, and subsequently perform fine-tuning and self-correction during inference with LLMs for CardEst task. We then extensively evaluate LLMs' in-distribution and out-of-distribution generalizability, feasibility to support complex queries, and training data efficiency during fine-tuning LLMs on pre-training datasets. The results suggest that LLMs outperform the state-of-the-art in almost all settings, thus indicating their potential for the CardEst task. We further measure the end-to-end query execution time in DBMS by using the estimated cardinalities of LLMs in some practical settings, which suggests that the inference overhead of LLMs can be outweighed by the benefits brought by LLMs for CardEst.

cs.DB

A $\Gamma$-valley Moir\'e Platform for Tunable Square Lattice Hubbard Model

Moir\'e superlattices have emerged as a premier platform for simulating the Hubbard model, yet achieving high tunability in square-lattice systems remains a key challenge. We demonstrate that $\Gamma$-valley twisted square homobilayers provide a faithful and highly tunable realization of $t-t'-U$ Hubbard model, extending the recent proposal in M-valley systems. We show that at small twist angles, an emergent layer-exchange symmetry decouples electronic states into flat bands residing on two nested square sublattices. An interlayer displacement field breaks this symmetry to induce controllable inter-sublattice hybridization, enabling wide-range experimental tuning of the effective hopping ratio $t'/t$. By establishing a direct correspondence between $\Gamma$- and M-valley systems, we provide a unified framework for understanding displacement-field tunability in square moir\'e physics. These findings establish $\Gamma$-valley twisted bilayers as a versatile platform for simulating the square-lattice Hubbard model and exploring its rich landscape of correlated phenomena.

cond-mat.mes-hall

Moir\'e Ferroelectricity-Driven Band Engineering in Twisted Square Bilayers

We develop the moir\'e band theory for M-valley twisted square homobilayers with layer groups $P$-$42m$ and $P$-$4m2$, and propose candidate material realizations. We show that moir\'e ferroelectricity-originating from sliding ferroelectricity in the untwisted bilayers-provides an independent control knob for miniband engineering in addition to interlayer tunneling. The competition between these two effects enables controlled switching between layer-resolved bilayer minibands and an effective single isolated miniband. Remarkably, these systems exhibit an emergent momentum-space nonsymmorphic symmetry in the absence of external magnetic fields. Large-scale \emph{ab initio} calculations identify Cu$_2$WS$_4$ and GeCl$_2$ as representative materials realizing the ferroelectricity- and tunneling-dominated regimes, respectively. Our results establish twisted square homobilayers as a promising platform for correlated band engineering beyond moir\'e hexagonal systems.

cond-mat.mes-hall

CEI-3D: Collaborative Explicit-Implicit 3D Reconstruction for Realistic and Fine-Grained Object Editing

Existing 3D editing methods often produce unrealistic and unrefined results due to the deeply integrated nature of their reconstruction networks. To address the challenge, this paper introduces CEI-3D, an editing-oriented reconstruction pipeline designed to facilitate realistic and fine-grained editing. Specifically, we propose a collaborative explicit-implicit reconstruction approach, which represents the target object using an implicit SDF network and a differentially sampled, locally controllable set of handler points. The implicit network provides a smooth and continuous geometry prior, while the explicit handler points offer localized control, enabling mutual guidance between the global 3D structure and user-specified local editing regions. To independently control each attribute of the handler points, we design a physical properties disentangling module to decouple the color of the handler points into separate physical properties. We also propose a dual-diffuse-albedo network in this module to process the edited and non-edited regions through separate branches, thereby preventing undesired interference from editing operations. Building on the reconstructed collaborative explicit-implicit representation with disentangled properties, we introduce a spatial-aware editing module that enables part-wise adjustment of relevant handler points. This module employs a cross-view propagation-based 3D segmentation strategy, which helps users to edit the specified physical attributes of a target part efficiently. Extensive experiments on both real and synthetic datasets demonstrate that our approach achieves more realistic and fine-grained editing results than the state-of-the-art (SOTA) methods while requiring less editing time. Our code is available on https://github.com/shiyue001/CEI-3D.

cs.CV

CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling

Large language models (LLMs) training or inference across multiple nodes introduces significant pressure on GPU memory and interconnect bandwidth. The Compute Express Link (CXL) shared memory pool offers a scalable solution by enabling memory sharing across nodes, reducing over-provisioning and improving resource utilization. We propose \name, a collective communication library, leveraging the CXL shared memory pool to support cross-node GPU operations without relying on traditional RDMA-based networking. Our design addresses the challenges on synchronization, data interleaving, and communication parallelization faced by using the CXL shared memory pool for collective communications. Evaluating on multiple nodes with a TITAN-II CXL switch and six Micron CZ120 memory cards, we show that \name achieves highly efficient collective operations across hosts, demonstrating CXL's potential for scalable, memory-centric GPU communication. Our evaluation demonstrates that \name achieves average performance improvements of 1.34$\times$ for AllGather, 1.84$\times$ for Broadcast, 1.94$\times$ for Gather, and 1.04$\times$ for Scatter, compared to the original RDMA-based implementation over 200 Gbps InfiniBand. \textcolor{dong}{In addition, the evaluation with a case of LLM training shows 1.11$\times$ speedup compared with the InfiniBand while saving production cost by $2.75\times$ in hardware.}

cs.DC

Morphis: SLO-Aware Resource Scheduling for Microservices with Time-Varying Call Graphs

Modern microservice systems exhibit continuous structural evolution in their runtime call graphs due to workload fluctuations, fault responses, and deployment activities. Despite this complexity, our analysis of over 500,000 production traces from ByteDance reveals a latent regularity: execution paths concentrate around a small set of recurring invocation patterns. However, existing resource management approaches fail to exploit this structure. Industrial autoscalers like Kubernetes HPA ignore inter-service dependencies, while recent academic methods often assume static topologies, rendering them ineffective under dynamic execution contexts. In this work, we propose Morphis, a dependency-aware provisioning framework that unifies pattern-aware trace analysis with global optimization. It introduces structural fingerprinting that decomposes traces into a stable execution backbone and interpretable deviation subgraphs. Then, resource allocation is formulated as a constrained optimization problem over predicted pattern distributions, jointly minimizing aggregate CPU usage while satisfying end-to-end tail-latency SLOs. Our extensive evaluations on the TrainTicket benchmark demonstrate that Morphis reduces CPU consumption by 35-38% compared to state-of-the-art baselines while maintaining 98.8% SLO compliance.

cs.SE

Operators with disconnected spectrum in von Neumann algebras

Let $\mathcal{M}$ be a von Neumann algebra, $\mathcal{I}$ a weak-operator dense ideal in $\mathcal{M}$, and $\Phi$ a unitarily invariant $\|\cdot\|$-dominating norm on $\mathcal{I}$. In this paper, we provide a necessary and sufficient condition on $\Phi$ such that every operator in $\mathcal{M}$ can be expressed as the sum of an operator in $\mathcal{M}$ with disconnected spectrum and an operator in $\mathcal{I}$ whose $\Phi$-norm is arbitrarily small. Similarly, if $\mathcal{A}$ is a unital $C^*$-algebra of real rank zero with dimension greater than one and $\mathcal{I}$ is an essential ideal in $\mathcal{A}$, then every element in $\mathcal{A}$ can be written as the sum of an operator in $\mathcal{A}$ with disconnected spectrum and an operator in $\mathcal{I}$ whose norm is arbitrarily small.

math.OA