SearcharxivSearch

arXiv subjects

Akihiro Toyoda

Publications and source records attributed to Akihiro Toyoda.

6 recordsLinked to original sources

Single-Round Clustered Federated Learning via Data Collaboration Analysis for Non-IID Data

Federated Learning (FL) enables distributed learning across multiple clients without sharing raw data. When statistical heterogeneity across clients is severe, Clustered Federated Learning (CFL) can im-prove performance by grouping similar clients and training cluster-wise models. However, most CFL approaches rely on multiple communication rounds for cluster estimation and model updates, which limits their practicality under tight constraints on communication rounds. We propose Data Collaboration-based Clustered Federated Learning (DC-CFL), a single-round framework that completes both client clustering and cluster-wise learning, using only the information shared in DC analysis. DC-CFL quantifies inter-client similarity via total variation distance between label distributions, estimates clusters using hierarchical clustering, and performs cluster-wise learning via DC analysis. Experiments on multiple open datasets under representative non-IID conditions show that DC-CFL achieves accuracy comparable to multi-round baselines while requiring only one communication round. These results indicate that DC-CFL is a practical alternative for collaborative AI model development when multiple communication rounds are impractical. Our source code is publicly available at https://github.com/souta-suga/DC-CFL.

cs.LG

A new type of federated clustering: A non-model-sharing approach

In recent years, the growing need to leverage sensitive data across institutions has led to increased attention on federated learning (FL), a decentralized machine learning paradigm that enables model training without sharing raw data. However, existing FL-based clustering methods, known as federated clustering, typically assume simple data partitioning scenarios such as horizontal or vertical splits, and cannot handle more complex distributed structures. This study proposes data collaboration clustering (DC-Clustering), a novel federated clustering method that supports clustering over complex data partitioning scenarios where horizontal and vertical splits coexist. In DC-Clustering, each institution shares only intermediate representations instead of raw data, ensuring privacy preservation while enabling collaborative clustering. The method allows flexible selection between k-means and spectral clustering, and achieves final results with a single round of communication with the central server. We conducted extensive experiments using synthetic and open benchmark datasets. The results show that our method achieves clustering performance comparable to centralized clustering where all data are pooled. DC-Clustering addresses an important gap in current FL research by enabling effective knowledge discovery from distributed heterogeneous data. Its practical properties -- privacy preservation, communication efficiency, and flexibility -- make it a promising tool for privacy-sensitive domains such as healthcare and finance.

cs.LG

Estimating Covariate-balanced Survival Curve in Distributed Data Environment using Data Collaboration Quasi-Experiment

The sharing of patient-level data necessary for covariate-adjusted survival analysis between medical institutions is difficult due to privacy protection restrictions. We propose a privacy-preserving framework that estimates balanced Kaplan-Meier curves from distributed observational data without exchanging raw data. Each institution sends only the low-dimensional representation obtained through dimensionality reduction of the covariate matrix. Analysts reconstruct the aggregated dataset, perform propensity score matching, and estimate survival curves. Experiments using simulation datasets and five publicly available medical datasets showed that the proposed method consistently outperformed single-site analyses. This method can handle both horizontal and vertical data distribution scenarios and enables the collaborative acquisition of reliable survival curves with minimal communication and no disclosure of raw data.

stat.ME

Data collaboration for causal inference from limited medical testing and medication data

Observational studies enable causal inferences when randomized controlled trials (RCTs) are not feasible. However, integrating sensitive medical data across multiple institutions introduces significant privacy challenges. The data collaboration quasi-experiment (DC-QE) framework addresses these concerns by sharing "intermediate representations" -- dimensionality-reduced data derived from raw data -- instead of the raw data. While the DC-QE can estimate treatment effects, its application to medical data remains unexplored. This study applied the DC-QE framework to medical data from a single institution to simulate distributed data environments under independent and identically distributed (IID) and non-IID conditions. We propose a novel method for generating intermediate representations within the DC-QE framework. Experimental results demonstrated that DC-QE consistently outperformed individual analyses across various accuracy metrics, closely approximating the performance of centralized analysis. The proposed method further improved performance, particularly under non-IID conditions. These outcomes highlight the potential of the DC-QE framework as a robust approach for privacy-preserving causal inferences in healthcare. Broader adoption of this framework and increased use of intermediate representations could grant researchers access to larger, more diverse datasets while safeguarding patient confidentiality. This approach may ultimately aid in identifying previously unrecognized causal relationships, support drug repurposing efforts, and enhance therapeutic interventions for rare diseases.

stat.ME

Stable beam operation of approximately 1 mA beam under highly efficient energy recovery conditions at compact energy-recovery linac

A compact energy-recovery linac (cERL) has been un-der construction at KEK since 2009 to develop key technologies for the energy-recovery linac. The cERL began operating in 2013 to create a high-current beam with a low-emittance beam with stable continuous wave (CW) superconducting cavities. Owing to the development of critical components, such as the DC gun, superconducting cavities, and the design of ideal beam transport optics, we have successfully established approximately 1 mA stable CW operation with a small beam emittance and extremely small beam loss. This study presents the details of our key technologies and experimental results for achieving 100% energy recovery operation with extremely small beam loss during a stable, approximately 1 mA CW beam operation.

physics.acc-ph

Radionuclides in the Cooling Water Systems for the NuMi Beamline and the Antiproton Production Target Station at Fermilab

At the 120-GeV proton accelerator facilities of Fermilab, USA, water samples were collected from the cooling water systems for the target, magnetic horn1, magnetic horn2, decay pipe, and hadron absorber at the NuMI beamline as well as from the cooling water systems for the collection lens, pulse magnet and collimator, and beam absorber at the antiproton production target station, just after the shutdown of the accelerators for a maintenance period. Specific activities of γ -emitting radionuclides and 3H in these samples were determined using high-purity germanium detectors and a liquid scintillation counter. The cooling water contained various radionuclides depending on both major and minor materials in contact with the water. The activity of the radionuclides depended on the presence of a deionizer. Specific activities of 3H were used to estimate the residual rates of 7Be. The estimated residual rates of 7Be in the cooling water were approximately 5% for systems without deionizers and less than 0.1% for systems with deionizers, although the deionizers function to remove 7Be from the cooling water.

physics.acc-ph