SearcharxivSearch

arXiv subjects

Gang Li

Publications and source records attributed to Gang Li.

At least 19 recordsLinked to original sources

When Fusion Fails: Corruption-Aware Rebalanced Fusion for Multi-Modal Medical Image Segmentation

Multi-modal medical image segmentation leverages complementary diagnostic information, yet fusion can underperform single-modality baselines when spatially aligned inputs differ in quality. Here, "corruption" primarily denotes resolution-induced degradation rather than misalignment or complete modality absence, while synthetic noise is evaluated only as an auxiliary setting. We identify a critical optimization-inference inconsistency: degraded modalities can receive weak training updates yet substantially affect predictions, indicating active interference with fusion. We attribute this failure to resampling-induced feature corruption and optimization bias, where noisy features propagate through skip connections and encourage unreliable modality selection. We therefore propose CoReFuse-Med, a Corruption-aware Rebalanced Fusion framework that suppresses corruption during feature transmission and rebalances modality contributions during high-level fusion. Experiments on EPVS, BraTS, and WMH, including multiple Z-axis slice-retention ratios and an auxiliary noise test, demonstrate improved accuracy and robustness under modality-quality discrepancies. Our code is available at https://github.com/lrever/CoReFuse.

cs.CV

From Splats to Silicon: Rethinking Computational Efficiency of 3DGS

3D Gaussian splatting (3DGS) represents scenes with explicit primitives and supports real-time novel-view synthesis, yet its system efficiency varies substantially across scenes, viewpoints, rendering paths, and platform constraints. Existing studies pursue efficiency through representation and algorithm design, GPU runtime optimization, and architectural support, but their reported gains correspond to different points along the rendering and update paths. Connecting these indicators to end-to-end system benefit requires tracing how each optimization changes Gaussian selection, screen-space work, data movement, and stage or frame time. We therefore use a workload-centric framework to connect representation and algorithm research, GPU runtimes, and hardware architectures and to identify recurring workload patterns. We complement literature analysis with reproduced measurements and controlled GPU profiling of selected implementations, relating workload counts to stage time and memory traffic. Together, these comparisons show that system gains depend on workload reductions reaching downstream execution, granularity matching each stage, and the cost of data transfers, synchronization, and cached results, gradients, and optimizer data. Building on these findings, we discuss more consistent evaluation under rendering-quality constraints and identify key directions for future system design.

cs.AR

Radiative decays of the isoscalar $S$-wave $D\bar D$ molecule

In the hadronic molecular picture, an isoscalar $S$-wave $D\bar{D}$ molecule is naturally expected as the spin-0 partner of $X(3872)$. Assuming this state, denoted by $X_0$, to be a pure $D\bar D$ molecule with quantum numbers $J^{PC}=0^{++}$, we investigate its radiative decays $X_0\to\gamma V (V=\rho^0,\omega)$ within an effective Lagrangian framework. The decay amplitudes are generated by intermediate charmed-meson loops, with electromagnetic gauge invariance consistently maintained throughout the calculation. We evaluate the partial decay widths and examine their dependence on the $X_0$ mass and the cutoff. Although the individual decay widths exhibit a sizable dependence on the cutoff, their ratio is remarkably stable. In particular, we obtain that the ratio for the decays $X_0\to\gamma\rho^0$ and $X_0\to\gamma\omega$ is approximately 2.26, which is insensitive to the variation of the cutoff. This robust ratio provides a useful model-insensitive signature of the molecular structure of $X_0$.

hep-ph

From citation intent to knowledge contribution: Classifying what cited papers actually contribute

Understanding the flow and evolution of scientific knowledge is essential for assessing research impact. Existing citation analysis methods mainly focus on citing authors' subjective intents, failing to consistently characterize cited papers' knowledge contributions. This study proposes the Knowledge Contribution Taxonomy (KCT), derived from the Scientific Research Logic Model, which identifies the type of knowledge a cited paper contributes based on the citation context. KCT classifies citations into Method, Resource Tool, Empirical Finding, and Background, further distinguishing core from non-core contributions. We propose a Dual-Path Fusion model for the classification task, which achieves an accuracy of 85.5%, outperforming mainstream large language models. An analysis of 802,202 citations from the ACL Anthology reveals that core knowledge contributions account for only 39.09% of all citations. The core knowledge contribution citation count achieves higher hit rates for award-winning papers than the traditional citation count at all ranking cutoffs, reflecting the value of differentiating knowledge contributions for research evaluation and impact prediction. In dissemination prediction experiments, KCT outperforms citation intent classification, demonstrating its stronger predictive validity for scholarly dissemination. By focusing on the knowledge contributions of cited papers, the KCT can support differentiated research evaluation.

cs.DL

A large-scale dataset of sub-institution name disambiguation and hierarchical structures from OpenAlex

Accurate attribution of scholarly work to specific sub-institutional units, such as schools or departments of a university, is crucial for granular research assessment and policymaking. While robust identifiers exist for top-level institutions, standardized data for sub-level units remains scarce due to the linguistic and structural variability of affiliation strings. In this study, we introduce OpenSubAffil, a large-scale dataset mapping raw affiliation strings from OpenAlex to disambiguated sub-institutional entities and their hierarchical structures. We developed a pipeline integrating named entity recognition (NER) with embedding-based clustering. Furthermore, we proposed a multi-signal scoring function that synthesizes lexical and co-occurrence evidence to reconstruct the sub-institutional hierarchy. OpenSubAffil comprises mappings for 40 million affiliation strings to 638,843 disambiguated sub-units across 18,635 top-level institutions, together with their hierarchical relationships. Validation against Wikidata benchmarks and manual investigation show that our method achieves promising performance. Overall, this dataset bridges the granularity gap between individual researchers and top-level institutions, enabling high-resolution analyses of scholarly output and communication at the sub-institutional level. The OpenSubAffil dataset is publicly available at https://doi.org/10.5281/zenodo.19602782.

cs.DL

Hundred-hertz quantum circuit iteration rate in a reusable neutral-atom array

Neutral-atom quantum processors have rapidly advanced in scale and coherence, yet their practical performance remains constrained by limited quantum circuit iteration rates (qCIRs) and information throughput. Here we experimentally demonstrate a high-throughput neutral-atom system based on non-destructive readout and atom reuse. By integrating a chip-based photonic interface with a 10-qubit array, we implement non-destructive readout with a retention probability of 99.7%, and further achieve a raw qCIR of 101Hz and a post-selected qCIR of 74.8Hz. More importantly, we verify a general throughput optimization methodology and obtain a normalized Fisher information rate of 57.7Hz, improving the achievable throughput by more than one order of magnitude compared with conventional methods. Our results establish a practical route toward high-throughput neutral-atom quantum processors.

quant-ph

A scalable chip-integrated single-photon source array based on 50 individually addressable neutral atoms

Scalable arrays of identical single-photon sources are a central resource for photonic quantum information processing, quantum networks and quantum metrology. Neutral atoms provide intrinsically identical emitters that can be assembled and rearranged in optical tweezers, but a many-channel fiber interface to individually trapped atoms has remained a major technical challenge. Here we demonstrate a chip-interfaced single-photon source array based on 50 individually addressable $^{87}\mathrm{Rb}$ atoms. A glass waveguide fan-out converts the \SI{5}{\micro m} pitch of the optical-tweezer array to the \SI{127}{\micro m} pitch of a commercial fiber array, mapping each atom to its own waveguide, fiber and single-photon detector. We resolve all 50 channels with an average nearest-neighbor cross-talk of $0.4\%$ and a uniform insertion loss of \SI{2.9}{dB}, and verify single-photon emission with $g^{(2)}(0)=0.29$, presently limited by detector dark counts and residual cooling-light scattering. Combining per-channel atom discrimination, rearrangement and reservoir replenishment, we prepare source subarrays of up to 24 atoms with a $93\%$ fill fraction. For small target numbers, atom loss is repaired from the reservoir at the detection-limited rate of \SI{118}{Hz}. We further fabricate a 784-channel waveguide chip, showing that the photonic interface can be extended well beyond the present number. This architecture establishes a fiber-native neutral-atom platform for larger arrays of identical single-photon sources.

physics.atom-ph

FastJM: An R Package for Efficient Implementation of Semiparametric Joint Models for Longitudinal and Survival Data

Joint models provide a flexible framework for characterizing the association between longitudinal and time-to-event processes and have been widely applied in biomedical research. However, fitting joint models can be computationally challenging for large-scale and complex biomedical data. This paper introduces the \proglang{R} package \pkg{FastJM}, which provides computationally efficient frequentist estimation for three classes of semiparametric joint models: joint models with a single longitudinal biomarker, joint models with multiple longitudinal biomarkers, and joint models with a single longitudinal biomarker with heterogeneous within-subject (WS) variability. Within an expectation--maximization framework, \pkg{FastJM} employs customized linear-scan algorithms to efficiently update the nonparametric baseline hazards, thereby addressing a major computational bottleneck in semiparametric joint modeling. The package also supports commonly used time-dependent latent association structures by integrating these algorithms with a landmark multivariate joint modeling framework. \pkg{FastJM} provides a unified interface for model specification, estimation, inference, visualization, dynamic prediction, and prediction performance assessment, including cross-validated time-dependent accuracy measures and time-independent concordance statistics. We describe the underlying methodology and software implementation and demonstrate the main functionality of \pkg{FastJM} through reproducible examples.

stat.ME

A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era

The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping manufacturing faster than engineering curricula can adapt, widening the gap between the competencies required on the shop floor and those delivered by traditional engineering and technology education. This paper proposes a Workforce Readiness Level (WRL) framework, which adapts the Technology Readiness Level scale into nine progressive competency stages and a four-pillar rubric, digital and AI literacy, cyber-physical systems fluency, human-machine collaboration, and data-driven decision making, aggregated through a composite stage score and a cohort-level workforce-readiness index under a ``no-thin-pillar'' rule. The framework is instantiated at a university smart-manufacturing teaching laboratory and draws on 89 sponsored capstone projects delivered over four semesters, four of which are analyzed in depth. Four pillars jointly span the relevant ABET student outcomes. Across the highlighted cohorts the workforce-readiness index ranged from 5.2 to 6.4, and the no-thin-pillar rule was diagnostically informative in three of the four cases and the binding certification constraint in one, repeatedly surfacing cyber-physical and data-driven-decision gaps concealed behind strong analytics profiles; advancement to the highest stages was gated by industry-embedded experience rather than additional coursework. WRL offers educators, accreditation bodies, and regional workforce systems a common, evidence-based instrument for diagnosing and advancing workforce readiness; future work will calibrate pillar weights and test reliability and predictive validity.

eess.SY

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross-modal optimization and leading to poor speech intelligibility. To overcome these limitations, we introduce MiDashengLM-Gen, an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. MiDashengLM-Gen represents a first approach for general text-to-audio generation with one end-to-end trained model. Empirical evaluations demonstrate that MiDashengLM-Gen drastically improves speech intelligibility over existing unified models. On the Seed-TTS benchmark, English Word Error Rate (WER) drops from 12.15% to 2.79%, approaching the performance of dedicated Text-to-Speech (TTS) systems (1.24%). Furthermore, the framework extends effectively to multilingual settings, yielding highly competitive multilingual WERs compared to existing baselines. Lastly, the model maintains competitive mixed-audio generation quality on the MECAT benchmark. Code and checkpoints are available at https://github.com/xiaomi-research/midashenglm-gen and https://huggingface.co/mispeech/midashenglm-gen, and the demo page is available at https://xingws.github.io/midashenglm-gen-demo/.

eess.AS

Debiased Machine Learning for Partially Linear Accelerated Failure Time Models

The Cox model remains the default for survival analysis, but the proportional hazards assumption is often violated and hazard ratios can be difficult to interpret. Accelerated failure time (AFT) models provide an intuitive time-scale alternative, yet flexible covariate adjustment while preserving valid inference on a target exposure remains challenging. For the partially linear AFT model under right censoring, a rank-based debiased machine learning (DML) framework remains undeveloped: the rank-based pairwise moment is not Neyman orthogonal and standard cross-fitting does not directly apply to U-statistics. We develop the first such framework by combining an orthogonalized rank-based U-statistic, a censoring-corrected influence function, and block-pairwise cross-fitting, yielding valid inference under flexible nuisance estimation. Simulations and an application to All of Us electronic health record data demonstrate finite-sample performance and practical utility.

stat.ME

DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization

3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of PEs leads to marginal performance improvement during rendering. We identify that the root cause is the tightly coupled ``checking-while-blending'' dataflow, which exacerbates PE underutilization caused by spatial redundancy from irregular Gaussian coverage and temporal redundancy from asynchronous pixel-wise termination under parallel execution. To address this issue, we propose DeGS, a scalable architecture for efficient 3DGS inference. To systematically eliminate the redundancies inherent in rendering, DeGS exploits a decoupled dataflow, restructuring the coupled $\alpha$-checking, transmittance checking, and $\alpha$-blending of the standard rendering process into consecutive workload parsing, reorganization, and blending stages. This allows the fragmented, length-variable, and temporal-dependent workloads to be reorganized into compact, conflict-free, and dense workloads prior to blending, thereby significantly improving PE utilization during parallel blending. Implemented in 28 nm technology, DeGS achieves 2.36$\times$--7.25$\times$ throughput, 1.82$\times$--6.02$\times$ end-to-end speedup, and 1.59$\times$--4.42$\times$ energy efficiency over state-of-the-art 3DGS accelerators (GSCore, GBU, GCC) across diverse scenes and resolutions (720p to 8K). Moreover, scaling from 16 to 1024 PEs, DeGS maintains over 80\% PE utilization at high resolutions, significantly outperforming existing accelerators.

cs.AR

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot's execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation.

cs.AI

LATTICE: Constraint-Directed Scheduling, Memory Planning, and Pipeline Refinement for NPUs

General-purpose NPUs execute fine-grained command DAGs across heterogeneous compute and memory-transfer engines backed by finite, explicitly managed on-chip memories. This execution model creates a directed dependency between scheduling and memory planning: different legal topological orders induce different lifetime overlap, placement opportunities, and spill behavior, while a materialized layout introduces physical-address reuse constraints absent from the input precedence DAG. Command order therefore shapes the feasible memory plan, and the realized plan in turn defines the legal space for subsequent timing refinement. We present LATTICE, a deterministic constraint-directed compiler pipeline. Memory-Pressure-Aware Topological Scheduling reshapes lifetime geometry before address binding; Deterministic Linear Repackaging materializes tiered placement, spill/reload events, and plan-induced reuse constraints; and Critical Path Enhancement recovers pipeline parallelism while preserving the selected memory plan. Every accepted schedule passes independent memory and timing verification. Across six artifact-provided command traces labeled as derived from a Da Vinci NPU flow, LATTICE achieves the best or tied-best result in all 24 evaluated workload-metric comparisons. Relative to the best evaluated baseline for each workload and metric, it reduces peak memory, extra DDR traffic, spill count, and modeled makespan by 18.3% lower, 20.4% lower, 14.1% lower, and 16.3% lower, respectively. Plan-preserving CPE further reduces makespan by 12.1% over Freeze while leaving placement and memory traffic unchanged, establishing the static memory plan as a verifiable scheduling contract between memory planning and pipeline optimization.

cs.NI

Measurement of $\Xi^-/\bar{\Xi}^{+}$ production in jets from $Z$ boson decays with the DELPHI open data

The production rates of $\Xi^{-}/\bar{\Xi}^{+}$ baryons in energy-ranked jets produced in $Z\to\text{hadrons}$ decays are measured using $3.2$ million hadronic $Z$ events recorded by the DELPHI experiment. Jets are reconstructed using the Durham algorithm with $y_{\text{cut}}=0.005$. Quark- and gluon-enriched jet samples are obtained by ranking the jet energies in three-jet events. The softest jet are found to produce fewer $\Xi^{-}/\bar{\Xi}^{+}$ and less energetic baryons than the other jets. The ratio of $\Xi^{-}/\bar{\Xi}^{+}$ production rates in gluon and quark jets, each normalized to the corresponding mean charged-particle multiplicity, is measured to be $1.21 \pm 0.18~\mathrm{(stat.)} \pm 0.26~\mathrm{(syst.)}$. The result is consistent with the JETSET expectation and the OPAL measurements of $K_S^0$ and $\Lambda$ productions in $Z$ decays. This study presents the first measurement of the gluon-to-quark production ratio for baryons containing two $s$ quarks, providing new insights into strange-quark production and hadronization. Future $e^{+}e^{-}$ colliders such as CEPC and FCCee will provide much larger $Z$-boson samples and will allow far more precise studies of the subject.

hep-ex

Production of hidden-charm molecular candidates in $\psi(4660)$ decays

We investigate the production of several hidden-charm exotic candidates, including $Z_c(3900)$, $Z_c(4020)$, $Z_{cs}(3985)$, and $Z_2(4250)$, in $\psi(4660)$ decays under the assumption that these states are predominantly hadronic molecules. Treating $\psi(4660)$ as a conventional $\psi(5S)$ charmonium state, the production mechanisms are described through intermediate charmed-meson triangle loops, with its couplings to charmed-meson pairs estimated within the quark model. A systematic analysis of the processes $\psi(4660)\to Z_c(3900)\pi$, $\psi(4660)\to Z_c(4020)\pi$, $\psi(4660)\to Z_{cs}(3985)K$, and $\psi(4660)\to Z_2(4250)\pi$ is performed within a unified framework. The predicted branching fractions are found to be of the order of $10^{-2}$, $10^{-4}$, $10^{-3}$, and $10^{-6}$, respectively, exhibiting only a mild dependence on the cutoff parameter. We further find that the contributions from the $SHH$ intermediate loops dominate over those from the $THH$ and $HHH$ loops in most channels. The sizable production rates obtained in this work indicate that $\psi(4660)$ decays provide a promising platform for probing the molecular nature of charged hidden-charm exotic states and testing their underlying production mechanisms.

hep-ph

Study of exotic hadron states in the $DD^{*}$ system via the complex momentum representation and Green's function method

In this paper, we propose a novel approach to investigate exotic hadronic states. For the $DD^{*}$ system, we employ the projection operator method to derive the momentum-space interaction potential. Subsequently, the complex momentum representation (CMR) method is adopted to realize a unified description of bound states, resonant states, and the continuum. By combining the Green's function and the CMR, the scattering phase shifts and cross sections are determined. This integrated approach provides a comprehensive framework for analyzing the scattering dynamics of the $DD^{*}$ system. In the hadronic molecular state framework, the $X(3872)$, $T_{cc}^+$, and $Z_c(3900)$ states can be consistently explained as bound states, while the $G(3900)$ can be interpreted as a $P$-wave resonant state. The decomposition of the scattering phase shifts and cross sections facilitates understanding the roles of resonant and continuum spectrum.

hep-ph

Low-latency FPGA-based electronic control system for fast preparation of defect-free atom arrays

The scalability of neutral atom quantum computing demands integrated electronic control systems with low latency, modular architecture, and real-time feedback capability. Here, we present an FPGA-based electronic control system that eliminates the PC from the feedback loop, integrating photon counting, real-time decision-making, and waveform generation within a unified PXIe architecture. The system achieves a total feedback latency of $282\,\mathrm{\mu s}$ and is validated in practical experiments by assembling defect-free atom arrays from 24 stochastically loaded optical tweezers. A single-round rearrangement achieves a filling fraction of $\sim96\%$, while feedback-controlled iterative rearrangement over five rounds boosts the success probability for generating a 10-atom defect-free array from $65.7\%$ to $95.4\%$. This system establishes the electronic infrastructure necessary for mid-circuit measurement and real-time quantum error correction on neutral-atom platforms.

quant-ph