Searcharxiv⌕ Search

arXiv subjects

Dan Song

Publications and source records attributed to Dan Song.

At least 37 records · Page 2Linked to original sources

Unified Multi-Modal Image Synthesis for Missing Modality Imputation

Multi-modal medical images provide complementary soft-tissue characteristics that aid in the screening and diagnosis of diseases. However, limited scanning time, image corruption and various imaging protocols often result in incomplete multi-modal images, thus limiting the usage of multi-modal data for clinical purposes. To address this issue, in this paper, we propose a novel unified multi-modal image synthesis method for missing modality imputation. Our method overall takes a generative adversarial architecture, which aims to synthesize missing modalities from any combination of available ones with a single model. To this end, we specifically design a Commonality- and Discrepancy-Sensitive Encoder for the generator to exploit both modality-invariant and specific information contained in input modalities. The incorporation of both types of information facilitates the generation of images with consistent anatomy and realistic details of the desired distribution. Besides, we propose a Dynamic Feature Unification Module to integrate information from a varying number of available modalities, which enables the network to be robust to random missing modalities. The module performs both hard integration and soft integration, ensuring the effectiveness of feature combination while avoiding information loss. Verified on two public multi-modal magnetic resonance datasets, the proposed method is effective in handling various synthesis tasks and shows superior performance compared to previous methods.

cs.CV↗

The Rise of Artificial Intelligence in Educational Measurement: Opportunities and Ethical Challenges

The integration of artificial intelligence (AI) in educational measurement has revolutionized assessment methods, enabling automated scoring, rapid content analysis, and personalized feedback through machine learning and natural language processing. These advancements provide timely, consistent feedback and valuable insights into student performance, thereby enhancing the assessment experience. However, the deployment of AI in education also raises significant ethical concerns regarding validity, reliability, transparency, fairness, and equity. Issues such as algorithmic bias and the opacity of AI decision-making processes pose risks of perpetuating inequalities and affecting assessment outcomes. Responding to these concerns, various stakeholders, including educators, policymakers, and organizations, have developed guidelines to ensure ethical AI use in education. The National Council of Measurement in Education's Special Interest Group on AI in Measurement and Education (AIME) also focuses on establishing ethical standards and advancing research in this area. In this paper, a diverse group of AIME members examines the ethical implications of AI-powered tools in educational measurement, explores significant challenges such as automation bias and environmental impact, and proposes solutions to ensure AI's responsible and effective use in education.

cs.CY↗

CAT-DM: Controllable Accelerated Virtual Try-on with Diffusion Model

Generative Adversarial Networks (GANs) dominate the research field in image-based virtual try-on, but have not resolved problems such as unnatural deformation of garments and the blurry generation quality. While the generative quality of diffusion models is impressive, achieving controllability poses a significant challenge when applying it to virtual try-on and multiple denoising iterations limit its potential for real-time applications. In this paper, we propose Controllable Accelerated virtual Try-on with Diffusion Model (CAT-DM). To enhance the controllability, a basic diffusion-based virtual try-on network is designed, which utilizes ControlNet to introduce additional control conditions and improves the feature extraction of garment images. In terms of acceleration, CAT-DM initiates a reverse denoising process with an implicit distribution generated by a pre-trained GAN-based model. Compared with previous try-on methods based on diffusion models, CAT-DM not only retains the pattern and texture details of the inshop garment but also reduces the sampling steps without compromising generation quality. Extensive experiments demonstrate the superiority of CAT-DM against both GANbased and diffusion-based methods in producing more realistic images and accurately reproducing garment patterns.

cs.CV↗

Possible molecular states from interactions of charmed strange baryons

In this work, we perform an investigation of possible molecular states composed of two charmed strange baryons from the $Ξ_c^{(',*)}Ξ_c^{(',*)}$ interaction, and their hidden-charm hidden-strange partners from the $Ξ_c^{(',*)}\barΞ_c^{(',*)}$ interaction. With the help of the heavy quark chiral effective Lagrangians, the interactions of charmed strange baryons are described with light meson exchanges. The potential kernels are constructed, and inserted into the quasipotential Bethe-Salpeter equation. The bound states are produced from most interactions considered, which suggests that strong attractions exist widely between the charmed strange baryons. Experimental searching for such molecular states is suggested in future high-precision measurements.

hep-ph↗

Temporal-spatial Correlation Attention Network for Clinical Data Analysis in Intensive Care Unit

In recent years, medical information technology has made it possible for electronic health record (EHR) to store fairly complete clinical data. This has brought health care into the era of "big data". However, medical data are often sparse and strongly correlated, which means that medical problems cannot be solved effectively. With the rapid development of deep learning in recent years, it has provided opportunities for the use of big data in healthcare. In this paper, we propose a temporal-saptial correlation attention network (TSCAN) to handle some clinical characteristic prediction problems, such as predicting death, predicting length of stay, detecting physiologic decline, and classifying phenotypes. Based on the design of the attention mechanism model, our approach can effectively remove irrelevant items in clinical data and irrelevant nodes in time according to different tasks, so as to obtain more accurate prediction results. Our method can also find key clinical indicators of important outcomes that can be used to improve treatment options. Our experiments use information from the Medical Information Mart for Intensive Care (MIMIC-IV) database, which is open to the public. Finally, we have achieved significant performance benefits of 2.0\% (metric) compared to other SOTA prediction methods. We achieved a staggering 90.7\% on mortality rate, 45.1\% on length of stay. The source code can be find: \url{https://github.com/yuyuheintju/TSCAN}.

cs.LG↗

Deep Reinforcement Learning Framework for Thoracic Diseases Classification via Prior Knowledge Guidance

The chest X-ray is often utilized for diagnosing common thoracic diseases. In recent years, many approaches have been proposed to handle the problem of automatic diagnosis based on chest X-rays. However, the scarcity of labeled data for related diseases still poses a huge challenge to an accurate diagnosis. In this paper, we focus on the thorax disease diagnostic problem and propose a novel deep reinforcement learning framework, which introduces prior knowledge to direct the learning of diagnostic agents and the model parameters can also be continuously updated as the data increases, like a person's learning process. Especially, 1) prior knowledge can be learned from the pre-trained model based on old data or other domains' similar data, which can effectively reduce the dependence on target domain data, and 2) the framework of reinforcement learning can make the diagnostic agent as exploratory as a human being and improve the accuracy of diagnosis through continuous exploration. The method can also effectively solve the model learning problem in the case of few-shot data and improve the generalization ability of the model. Finally, our approach's performance was demonstrated using the well-known NIH ChestX-ray 14 and CheXpert datasets, and we achieved competitive results. The source code can be found here: \url{https://github.com/NeaseZ/MARL}.

eess.IV↗

Chest X-ray Image Classification: A Causal Perspective

The chest X-ray (CXR) is one of the most common and easy-to-get medical tests used to diagnose common diseases of the chest. Recently, many deep learning-based methods have been proposed that are capable of effectively classifying CXRs. Even though these techniques have worked quite well, it is difficult to establish whether what these algorithms actually learn is the cause-and-effect link between diseases and their causes or just how to map labels to photos.In this paper, we propose a causal approach to address the CXR classification problem, which constructs a structural causal model (SCM) and uses the backdoor adjustment to select effective visual information for CXR classification. Specially, we design different probability optimization functions to eliminate the influence of confounders on the learning of real causality. Experimental results demonstrate that our proposed method outperforms the open-source NIH ChestX-ray14 in terms of classification performance.

eess.IV↗

Privacy Amplification via Compression: Achieving the Optimal Privacy-Accuracy-Communication Trade-off in Distributed Mean Estimation

Privacy and communication constraints are two major bottlenecks in federated learning (FL) and analytics (FA). We study the optimal accuracy of mean and frequency estimation (canonical models for FL and FA respectively) under joint communication and $(\varepsilon, δ)$-differential privacy (DP) constraints. We show that in order to achieve the optimal error under $(\varepsilon, δ)$-DP, it is sufficient for each client to send $Θ\left( n \min\left(\varepsilon, \varepsilon^2\right)\right)$ bits for FL and $Θ\left(\log\left( n\min\left(\varepsilon, \varepsilon^2\right) \right)\right)$ bits for FA to the server, where $n$ is the number of participating clients. Without compression, each client needs $O(d)$ bits and $\log d$ bits for the mean and frequency estimation problems respectively (where $d$ corresponds to the number of trainable parameters in FL or the domain size in FA), which means that we can get significant savings in the regime $ n \min\left(\varepsilon, \varepsilon^2\right) = o(d)$, which is often the relevant regime in practice. Our algorithms leverage compression for privacy amplification: when each client communicates only partial information about its sample, we show that privacy can be amplified by randomly selecting the part contributed by each client.

stat.ML↗

Possible $Λ_c\barΛ_c$ molecular states and their productions in nulceon-antinulceon collision

In this work, a study of possible molecular states from the $Λ_c\barΛ_c$ interaction and their productions in nucleon-antinucleon collision is performed in a quasipotential Bethe-Salpeter equation approach. Two bound states with quantum numbers $J^{PC}=0^{-+}$ and $1^{--}$ are produced with almost the same binding energy from the $Λ_c\barΛ_c$ interaction which is described by the light meson exchanges. However, the result does not support the assignment of experimentally observed $Y(4630)$ as a $Λ_c\barΛ_c$ molecular state because it is hard to obtain a peak near experimental mass of the $Y(4630)$ which is far above the $Λ_c\barΛ_c$ threshold. The possibility to search these states in nucleon-antinucleon collision is studied by including couplings to $N\bar{N}$ and $D^{(*)}\bar{D}^{(*)}$ channels. The peaks can be found obviously near the $Λ_c\barΛ_c$ threshold in the $D^*\bar{D}^*$ channel at an order of amplitude of 10 $μ$b. Too small width of state with $0^{-+}$ may lead to the difficulty to be observed in experiment. Based on the results in the current work, search for the $Λ_c\barΛ_c$ molecular state with $1^{--}$ is suggested in process $N\bar{N}\to D^*\bar{D}^*$, which is accessible at $\rm \bar{P}ANDA$.

hep-ph↗

Possible molecular states from interactions of charmed baryons

In this work, we perform a systematic study of possible molecular states composed of two charmed baryons including hidden-charm systems $Λ_c\barΛ_c$, $Σ_c^{(*)}\barΣ_c^{(*)}$, and $Λ_c\barΣ_c^{(*)}$, and corresponding double-charm systems $Λ_cΛ_c$, $Σ_c^{(*)}Σ_c^{(*)}$, and $Λ_cΣ_c^{(*)}$. With the help of the heavy quark chiral effective Lagrangians, the interactions are described with $π$, $ρ$, $η$, $ω$, $ϕ$, and $σ$ exchanges. The potential kernels are constructed, and inserted into the quasipotential Bethe-Salpeter equation. The bound states from the interactions considered is studied by searching for the poles of the scattering amplitude. The results suggest that strong attractions exist in both hidden-charm and double-charm systems considered in the current work, and bound states can be produced in most of the systems. More experiment studies about these molecular states are suggested though the nucleon-nucleon collison at LHC and nucleon-antinucleon collison at $\rm \bar{P}ANDA$.

hep-ph↗

CTooth+: A Large-scale Dental Cone Beam Computed Tomography Dataset and Benchmark for Tooth Volume Segmentation

Accurate tooth volume segmentation is a prerequisite for computer-aided dental analysis. Deep learning-based tooth segmentation methods have achieved satisfying performances but require a large quantity of tooth data with ground truth. The dental data publicly available is limited meaning the existing methods can not be reproduced, evaluated and applied in clinical practice. In this paper, we establish a 3D dental CBCT dataset CTooth+, with 22 fully annotated volumes and 146 unlabeled volumes. We further evaluate several state-of-the-art tooth volume segmentation strategies based on fully-supervised learning, semi-supervised learning and active learning, and define the performance principles. This work provides a new benchmark for the tooth volume segmentation task, and the experiment can serve as the baseline for future AI-based dental imaging research and clinical application development.

eess.IV↗

CTooth: A Fully Annotated 3D Dataset and Benchmark for Tooth Volume Segmentation on Cone Beam Computed Tomography Images

3D tooth segmentation is a prerequisite for computer-aided dental diagnosis and treatment. However, segmenting all tooth regions manually is subjective and time-consuming. Recently, deep learning-based segmentation methods produce convincing results and reduce manual annotation efforts, but it requires a large quantity of ground truth for training. To our knowledge, there are few tooth data available for the 3D segmentation study. In this paper, we establish a fully annotated cone beam computed tomography dataset CTooth with tooth gold standard. This dataset contains 22 volumes (7363 slices) with fine tooth labels annotated by experienced radiographic interpreters. To ensure a relative even data sampling distribution, data variance is included in the CTooth including missing teeth and dental restoration. Several state-of-the-art segmentation methods are evaluated on this dataset. Afterwards, we further summarise and apply a series of 3D attention-based Unet variants for segmenting tooth volumes. This work provides a new benchmark for the tooth volume segmentation task. Experimental evidence proves that attention modules of the 3D UNet structure boost responses in tooth areas and inhibit the influence of background and noise. The best performance is achieved by 3D Unet with SKNet attention module, of 88.04 \% Dice and 78.71 \% IOU, respectively. The attention-based Unet framework outperforms other state-of-the-art methods on the CTooth dataset. The codebase and dataset are released.

cs.CV↗

Efficiently Computable Converses for Finite-Blocklength Communication

This paper presents a method for computing a finite-blocklength converse for the rate of fixed-length codes with feedback used on discrete memoryless channels (DMCs). The new converse is expressed in terms of a stochastic control problem whose solution can be efficiently computed using dynamic programming and Fourier methods. For channels such as the binary symmetric channel (BSC) and binary erasure channel (BEC), the accuracy of the proposed converse is similar to that of existing special-purpose converse bounds, but the new converse technique can be applied to arbitrary DMCs. We provide example applications of the new converse technique to the binary asymmetric channel (BAC) and the quantized amplitude-constrained AWGN channel.

cs.IT↗

Achieving Short-Blocklength RCU bound via CRC List Decoding of TCM with Probabilistic Shaping

This paper applies probabilistic amplitude shaping (PAS) to a cyclic redundancy check (CRC) aided trellis coded modulation (TCM) to achieve the short-blocklength random coding union (RCU) bound. In the transmitter, the equally likely message bits are first encoded by distribution matcher to generate amplitude symbols with the desired distribution. The binary representations of the distribution matcher outputs are then encoded by a CRC. Finally, the CRC-encoded bits are encoded and modulated by Ungerboeck's TCM scheme, which consists of a $\frac{k_0}{k_0+1}$ systematic tail-biting convolutional code and a mapping function that maps coded bits to channel signals with capacity-achieving distribution. This paper proves that, for the proposed transmitter, the CRC bits have uniform distribution and that the channel signals have symmetric distribution. In the receiver, the serial list Viterbi decoding (S-LVD) is used to estimate the information bits. Simulation results show that, for the proposed CRC-TCM-PAS system with 87 input bits and 65-67 8-AM coded output symbols, the decoding performance under additive white Gaussian noise channel achieves the RCU bound with properly designed CRC and convolutional codes.

cs.IT↗

Heavy-strange meson molecules and possible candidates $D^*_{s0}(2317)$, $D_{s1}(2460)$, and $X_0(2900)$

In this work, we systematically investigate the heavy-strange meson systems, $D^{(*)}K^{(*)}/\bar{B}^{(*)}K^{(*)}$ and $\bar{D}^{(*)}K^{(*)}/B^{(*)}K^{(*)}$, to study possible molecules in a quasipotenial Bethe-Salpter equation approach together with the one-boson exchange model. The potential is achieved with the help of the hidden-gauge Lagrangians. Molecular states are found from all six $S$-wave isoscalar interactions of $D^{(*)}K^{(*)}$ or $\bar{B}^{(*)}K^{(*)}$. The charmed-strange mesons $D^*_{s0}(2317)$ and $D_{s1}(2460)$ can be related to the ${D}K$ and $D^*K$ states with spin parities $0^+$ and $1^+$, respectively. In the current model, the $\bar{B}K^*$ molecular state with $1^+$ is the best candidate of the recent observed $B_{sJ}(6158)$. Four molecular states are produced from the interactions of $\bar{D}^{(*)}K^{(*)}$ or $B^{(*)}K^{(*)}$. The relation between the $\bar{D}^*{K}^*$ molecular state with $0^+$ and the $X_0(2900)$ is also discussed. No isovector molecular states are found from the interactions considered. The current results are helpful to understand the internal structure of $D^*_{s0}(2317)$, $D_{s1}(2460)$, $X_0(2900)$, and new $B_{sJ}$ states. The experimental research for more heavy-strange meson molecules is suggested.

hep-ph↗

Hidden and doubly heavy molecular states from interactions $D^{(*)}_{(s)}{\bar{D}}^{(*)}_{s}$/$B^{(*)}_{(s)}{\bar{B}}^{(*)}_{s}$ and ${D}^{(*)}_{(s)}D_{s}^{(*)}$/${B}^{(*)}_{(s)}B_{s}^{(*)}$

In this work, we perform a systematical investigation about the possible hidden and doubly heavy molecular states with open and hidden strangeness from interactions of $D^{(*)}{\bar{D}}^{(*)}_{s}$/$B^{(*)}{\bar{B}}^{(*)}_{s}$, ${D}^{(*)}_{s}{\bar{D}}^{(*)}_{s}$/${B}^{(*)}_{s}{\bar{B}}^{(*)}_{s}$, ${D}^{(*)}D_{s}^{(*)}$/${B}^{(*)}B_{s}^{(*)}$, and $D_{s}^{(*)}D_{s}^{(*)}$/$B_{s}^{(*)}B_{s}^{(*)}$ in a quasipotential Bethe-Salpeter equation approach. The interactions of the systems considered are described within the one-boson-exchange model, which includes exchanges of light mesons and $J/ψ/Υ$ meson. Possible molecular states are searched for as poles of scattering amplitudes of the interactions considered. The results suggest that recently observed $Z_{cs}(3985)$ can be assigned as a molecular state of $D^*\bar{D}_s+D\bar{D}^*_s$, which is a partner of $Z_c(3900)$ state as a $D\bar{D}^*$ molecular state. The calculation also favors the existence of hidden heavy states $D_s\bar{D}_s/B_s\bar{B}_s$ with spin parity $J^P=0^+$, $D_s\bar{D}^*_s/B_s\bar{B}^*_s$ with $1^{+}$, and $D^*_s\bar{D}^*_s/B^*_s\bar{B}^*_s$ with $0^+$, $1^+$, and $2^+$. In the doubly heavy sector, the bound states can be found from the interactions $(D^*D_s+DD^*_s)/(B^*B_s+BB^*_s)$ with $1^+$, $D_s\bar{D}_s^*/B_s\bar{B}_s^*$ with $1^+$, $D^*D^*_s/B^*B^*_s$ with $1^+$ and $2^+$, and $D^*_sD^*_s/B^*_sB^*_s$ with $1^+$ and $2^+$. Some other interactions are also found attractive, but may be not strong enough to produce a bound state. The results in this work are helpful for understanding the $Z_{cs}(3985)$, and future experimental search for the new molecular states.

hep-ph↗

Illumination-aware Faster R-CNN for Robust Multispectral Pedestrian Detection

Multispectral images of color-thermal pairs have shown more effective than a single color channel for pedestrian detection, especially under challenging illumination conditions. However, there is still a lack of studies on how to fuse the two modalities effectively. In this paper, we deeply compare six different convolutional network fusion architectures and analyse their adaptations, enabling a vanilla architecture to obtain detection performances comparable to the state-of-the-art results. Further, we discover that pedestrian detection confidences from color or thermal images are correlated with illumination conditions. With this in mind, we propose an Illumination-aware Faster R-CNN (IAF RCNN). Specifically, an Illumination-aware Network is introduced to give an illumination measure of the input image. Then we adaptively merge color and thermal sub-networks via a gate function defined over the illumination value. The experimental results on KAIST Multispectral Pedestrian Benchmark validate the effectiveness of the proposed IAF R-CNN.

cs.CV↗

Multispectral Pedestrian Detection via Simultaneous Detection and Segmentation

Multispectral pedestrian detection has attracted increasing attention from the research community due to its crucial competence for many around-the-clock applications (e.g., video surveillance and autonomous driving), especially under insufficient illumination conditions. We create a human baseline over the KAIST dataset and reveal that there is still a large gap between current top detectors and human performance. To narrow this gap, we propose a network fusion architecture, which consists of a multispectral proposal network to generate pedestrian proposals, and a subsequent multispectral classification network to distinguish pedestrian instances from hard negatives. The unified network is learned by jointly optimizing pedestrian detection and semantic segmentation tasks. The final detections are obtained by integrating the outputs from different modalities as well as the two stages. The approach significantly outperforms state-of-the-art methods on the KAIST dataset while remain fast. Additionally, we contribute a sanitized version of training annotations for the KAIST dataset, and examine the effects caused by different kinds of annotation errors. Future research of this problem will benefit from the sanitized version which eliminates the interference of annotation errors.

cs.CV↗