SearcharxivSearch

arXiv subjects

Yuta Nakamura

Publications and source records attributed to Yuta Nakamura.

13 recordsLinked to original sources

Passivity-based Semi-autonomous Rotational Motion Navigation for Rigid-body Networks: Stability and Human Passivity Analysis

This paper presents a novel passivity-based semi-autonomous attitude control framework, with a particular focus on attitude kinematics defined on the special orthogonal group $SO(3)$. While human-robot interaction facilitates the successful execution of complex tasks, ensuring stability of human-in-the-loop systems on the $SO(3)$ manifold remains a largely unsolved challenge. We first propose a new control architecture in which a multi-robot system preserves invariance of the average information fed back to the human operator through so-called stealthy control, and the human intervention is mediated through a virtual leader, which is coupled with the robots via a passivity-based attitude synchronization law. We then rigorously prove closed-loop stability of the proposed human-in-the-loop system under the assumption that the human behaves as a passive system. To support this analysis, simulation studies are conducted to identify the human operator as a dynamical system, and to examine passivity properties of the identified model.

eess.SY

Stealthy Coverage Control for Human-enabled Real-Time 3D Reconstruction

In this paper, we propose a novel semi-autonomous image sampling strategy, called stealthy coverage control, for human-enabled 3D structure reconstruction. The present mission involves a fundamental problem: while the number of images required to accurately reconstruct a 3D model depends on the structural complexity of the target scene to be reconstructed, it is not realistic to assume prior knowledge of the spatially non-uniform structural complexity. We approach this issue by leveraging human flexible reasoning and situational recognition capabilities. Specifically, we design a semi-autonomous system that leaves identification of regions that need more images and navigation of the drones to such regions to a human operator. To this end, we first present a way to reflect the human intention in autonomous coverage control. Subsequently, in order to avoid operational conflicts between manual control and autonomous coverage control, we develop the stealthy coverage control that decouples the drone motion for efficient image sampling from navigation by the human. Simulation studies on a Unity/ROS2-based simulator demonstrate that the present semi-autonomous system outperforms the one without human interventions in the sense of the reconstructed model quality.

eess.SY

Simultaneous power generation and cooling using semiconductor-sensitized thermal cells

This manuscript reports a semiconductor-sensitized thermal cell (STC) that converts ambient heat into electrical power while simultaneously reducing its own temperature under isothermal conditions. Using a printable semiconductor--electrolyte architecture, we fabricate $4\,\mathrm{cm} \times 4\,\mathrm{cm}$ devices that generate up to approximately $0.2\,\mathrm{mW}$ at temperatures of $40$--$55\,^\circ\mathrm{C}$. During continuous discharge, the STC exhibits a transient temperature decrease followed by thermal equilibration with the environment. In contrast, periodic on--off discharge produces sustained cooling of approximately $1\,^\circ\mathrm{C}$ relative to a non-discharging reference. Notably, parallel integration of four STCs yields a nonlinear enhancement of cooling (approximately $5\,^\circ\mathrm{C}$) without a corresponding increase in electrical output. The observed behavior can be understood within a macroscopic energy-balance framework, in which time modulation of electrochemical heat consumption prevents the establishment of thermal steady state. These results demonstrate sustained isothermal cooling induced by heat-to-electricity conversion at practical device scales, and highlight semiconductor-sensitized thermal cells as a platform for coupled energy harvesting and thermal management.

cond-mat.mtrl-sci

Observation of Galactic center in the sub-MeV gamma-ray band with electron-tracking Compton camera

We report the direct detection of gamma-ray emission from the Galactic center in the 150-600 keV band using the electron-tracking Compton camera (ETCC), which has a wide field of view of 3.1 sr. This represents the first application of this linear, imaging-spectroscopy method to observations of the Galactic center. Measurements in a one-day flight over Australia yielded significant gamma-ray detection in the light curve and revealed a $7.9\sigma$ excess over the background in the image map from the Galactic center region. These results, obtained through a simple and unambiguous analysis, demonstrate the high reliability and sensitivity of the ETCC and establish its potential for future high-precision MeV gamma-ray observations. The measured intensity and spatial distribution were tested against three emission models: a single point-like source, a multi-component structure, and a symmetric two-dimensional Gaussian. All three were found to be statistically consistent with the data. The positronium-related flux provided by the multi-component model is $(3.2~\pm~1.4) \times 10^{-2}$ photons cm$^{-2}$s$^{-1}$, consistent with the value reported by INTEGRAL within $1\sigma$. These results establish the potential of the ETCC for future high-precision MeV gamma-ray surveys.

astro-ph.HE

Zero-shot 3D Segmentation of Abdominal Organs in CT Scans Using Segment Anything Model 2: Adapting Video Tracking Capabilities for 3D Medical Imaging

Objectives: To evaluate the zero-shot performance of Segment Anything Model 2 (SAM 2) in 3D segmentation of abdominal organs in CT scans, and to investigate the effects of prompt settings on segmentation results. Materials and Methods: In this retrospective study, we used a subset of the TotalSegmentator CT dataset from eight institutions to assess SAM 2's ability to segment eight abdominal organs. Segmentation was initiated from three different z-coordinate levels (caudal, mid, and cranial levels) of each organ. Performance was measured using the Dice similarity coefficient (DSC). We also analyzed the impact of "negative prompts," which explicitly exclude certain regions from the segmentation process, on accuracy. Results: 123 patients (mean age, 60.7 \pm 15.5 years; 63 men, 60 women) were evaluated. As a zero-shot approach, larger organs with clear boundaries demonstrated high segmentation performance, with mean DSCs as follows: liver 0.821 \pm 0.192, right kidney 0.862 \pm 0.212, left kidney 0.870 \pm 0.154, and spleen 0.891 \pm 0.131. Smaller organs showed lower performance: gallbladder 0.531 \pm 0.291, pancreas 0.361 \pm 0.197, and adrenal glands, right 0.203 \pm 0.222, left 0.308 \pm 0.234. The initial slice for segmentation and the use of negative prompts significantly influenced the results. By removing negative prompts from the input, the DSCs significantly decreased for six organs. Conclusion: SAM 2 demonstrated promising zero-shot performance in segmenting certain abdominal organs in CT scans, particularly larger organs. Performance was significantly influenced by input negative prompts and initial slice selection, highlighting the importance of optimizing these factors.

eess.IV

Development of a Large-scale Dataset of Chest Computed Tomography Reports in Japanese and a High-performance Finding Classification Model

Background: Recent advances in large language models highlight the need for high-quality multilingual medical datasets. While Japan leads globally in CT scanner deployment and utilization, the lack of large-scale Japanese radiology datasets has hindered the development of specialized language models for medical imaging analysis. Objective: To develop a comprehensive Japanese CT report dataset through machine translation and establish a specialized language model for structured finding classification. Additionally, to create a rigorously validated evaluation dataset through expert radiologist review. Methods: We translated the CT-RATE dataset (24,283 CT reports from 21,304 patients) into Japanese using GPT-4o mini. The training dataset consisted of 22,778 machine-translated reports, while the validation dataset included 150 radiologist-revised reports. We developed CT-BERT-JPN based on "tohoku-nlp/bert-base-japanese-v3" architecture for extracting 18 structured findings from Japanese radiology reports. Results: Translation metrics showed strong performance with BLEU scores of 0.731 and 0.690, and ROUGE scores ranging from 0.770 to 0.876 for Findings and from 0.748 to 0.857 for Impression sections. CT-BERT-JPN demonstrated superior performance compared to GPT-4o in 11 out of 18 conditions, including lymphadenopathy (+14.2%), interlobular septal thickening (+10.9%), and atelectasis (+7.4%). The model maintained F1 scores exceeding 0.95 in 14 out of 18 conditions and achieved perfect scores in four conditions. Conclusions: Our study establishes a robust Japanese CT report dataset and demonstrates the effectiveness of a specialized language model for structured finding classification. The hybrid approach of machine translation and expert validation enables the creation of large-scale medical datasets while maintaining high quality.

cs.CL

High-energy extension of the gamma-ray band observable with an electron-tracking Compton camera

Although the MeV gamma-ray band is a promising energy-band window in astrophysics, the current situation of MeV gamma-ray astronomy significantly lags behind those of the other energy bands in angular resolution and sensitivity. An electron-tracking Compton camera (ETCC), a next-generation MeV detector, is expected to revolutionize the situation. An ETCC tracks each Compton-recoil electron with a gaseous electron tracker and determines the incoming direction of each gamma-ray photon; thus, it has a strong background rejection power and yields a better angular resolution than classical Compton cameras. Here, we study ETCC events in which the Compton-recoil electrons do not deposit all energies to the electron tracker but escape and hit the surrounding pixel scintillator array (PSA). We developed an analysis method for this untapped class of events and applied it to laboratory and simulation data. We found that the energy spectrum obtained from the simulation agreed with that of the actual data within a factor of 1.2. We then evaluated the detector performance using the simulation data. The angular resolution for the new-class events was found to be twice as good as in the previous study at the energy range 1.0--2.0~MeV, where both analyses overlap. We also found that the total effective area is dominated by the contribution of the double-hit events above an energy of 1.5~MeV. Notably, applying this new method extends the sensitive energy range with the ETCC from 0.2--2.1 MeV in the previous studies to up to 3.5~MeV. Adjusting the PSA dynamic range should improve the sensitivity in even higher energy gamma-rays. The development of this new analysis method would pave the way for future observations by ETCC to fill the MeV-band sensitivity gap in astronomy.

astro-ph.HE

Playing the Werewolf game with artificial intelligence for language understanding

The Werewolf game is a social deduction game based on free natural language communication, in which players try to deceive others in order to survive. An important feature of this game is that a large portion of the conversations are false information, and the behavior of artificial intelligence (AI) in such a situation has not been widely investigated. The purpose of this study is to develop an AI agent that can play Werewolf through natural language conversations. First, we collected game logs from 15 human players. Next, we fine-tuned a Transformer-based pretrained language model to construct a value network that can predict a posterior probability of winning a game at any given phase of the game and given a candidate for the next action. We then developed an AI agent that can interact with humans and choose the best voting target on the basis of its probability from the value network. Lastly, we evaluated the performance of the agent by having it actually play the game with human players. We found that our AI agent, Deep Wolf, could play Werewolf as competitively as average human players in a villager or a betrayer role, whereas Deep Wolf was inferior to human players in a werewolf or a seer role. These results suggest that current language models have the capability to suspect what others are saying, tell a lie, or detect lies in conversations.

cs.AI

Statistical mechanics analysis of general multi-dimensional knapsack problems

Knapsack problem (KP) is a representative combinatorial optimization problem that aims to maximize the total profit by selecting a subset of items under given constraints on the total weights. In this study, we analyze a generalized version of KP, which is termed the generalized multidimensional knapsack problem (GMDKP). As opposed to the basic KP, GMDKP allows multiple choices per item type under multiple weight constraints. Although several efficient algorithms are known and the properties of their solutions have been examined to a significant extent for basic KPs, there is a paucity of known algorithms and studies on the solution properties of GMDKP. To gain insight into the problem, we assess the typical achievable limit of the total profit for a random ensemble of GMDKP using the replica method. Our findings are summarized as follows: (1) When the profits of item types are normally distributed, the total profit grows in the leading order with respect to the number of item types as the maximum number of choices per item type $x^{\rm max}$ increases while it depends on $x^{\rm max}$ only in a sub-leading order if the profits are constant among the item types. (2) A greedy-type heuristic can find a nearly optimal solution whose total profit is lower than the optimal value only by a sub-leading order with a low computational cost. (3) The sub-leading difference from the optimal total profit can be improved by a heuristic algorithm based on the cavity method. Extensive numerical experiments support these findings.

math.OC

First observation of MeV gamma-ray universe with bijective imaging spectroscopy using the Electron-Tracking Compton Telescope aboard SMILE-2+

MeV gamma-rays provide a unique window for the direct measurement of line emissions from radioisotopes, but observations have made little significant progress after COMPTEL/{\it CGRO}. To observe celestial objects in this band, we are developing an electron-tracking Compton camera (ETCC), which realizes both bijective imaging spectroscopy and efficient background reduction gleaned from the recoil electron track information. The energy spectrum of the observation target can then be obtained by a simple ON-OFF method using a correctly defined point spread function on the celestial sphere. The performance of celestial object observations was validated on the second balloon SMILE-2+ installed with an ETCC having a gaseous electron tracker with a volume of 30$\times$30$\times$30 cm$^3$. Gamma-rays from the Crab nebula were detected with a significance of 4.0$σ$ in the energy range 0.15--2.1 MeV with a live time of 5.1 h, as expected before launching. Additionally, the light curve clarified an enhancement of gamma-ray events generated in the Galactic center region, indicating that a significant proportion of the final remaining events are cosmic gamma rays. Independently, the observed intensity and time variation were consistent with the pre-launch estimates except in the Galactic center region. The estimates were based on the total background of extragalactic diffuse, atmospheric, and instrumental gamma-rays after accounting for the variations in the atmospheric depth and rigidity during the level flight. The Crab results and light curve strongly support our understanding of both the detection sensitivity and the background in real observations. This work promises significant advances in MeV gamma-ray astronomy.

astro-ph.HE

KART: Parameterization of Privacy Leakage Scenarios from Pre-trained Language Models

For the safe sharing pre-trained language models, no guidelines exist at present owing to the difficulty in estimating the upper bound of the risk of privacy leakage. One problem is that previous studies have assessed the risk for different real-world privacy leakage scenarios and attack methods, which reduces the portability of the findings. To tackle this problem, we represent complex real-world privacy leakage scenarios under a universal parameterization, \textit{Knowledge, Anonymization, Resource, and Target} (KART). KART parameterization has two merits: (i) it clarifies the definition of privacy leakage in each experiment and (ii) it improves the comparability of the findings of risk assessments. We show that previous studies can be simply reviewed by parameterizing the scenarios with KART. We also demonstrate privacy risk assessments in different scenarios under the same attack method, which suggests that KART helps approximate the upper bound of risk under a specific attack or scenario. We believe that KART helps integrate past and future findings on privacy risk and will contribute to a standard for sharing language models.

cs.CL

Twofold Multiprior Preferences and Failures of Contingent Reasoning

We propose a model of incomplete \textit{twofold multiprior preferences}, in which an act $f$ is ranked above an act $g$ only when $f$ provides higher utility in a worst-case scenario than what $g$ provides in a best-case scenario. The model explains failures of contingent reasoning, captured through a weakening of the state-by-state monotonicity (or dominance) axiom. Our model gives rise to rich comparative statics results, as well as extension exercises, and connections to choice theory. We present an application to second-price auctions.

econ.TH

Content-defined Merkle Trees for Efficient Container Delivery

Containerization simplifies the sharing and deployment of applications when environments change in the software delivery chain. To deploy an application, container delivery methods push and pull container images. These methods operate on file and layer (set of files) granularity, and introduce redundant data within a container. Several container operations such as upgrading, installing, and maintaining become inefficient, because of copying and provisioning of redundant data. In this paper, we reestablish recent results that block-level deduplication reduces the size of individual containers, by verifying the result using content-defined chunking. Block-level deduplication, however, does not improve the efficiency of push/pull operations which must determine the specific blocks to transfer. We introduce a content-defined Merkle Tree (\CDMT{}) over deduplicated storage in a container. \CDMT{} indexes deduplicated blocks and determines changes to blocks in logarithmic time on the client. \CDMT{} efficiently pushes and pulls container images from a registry, especially as containers are upgraded and (re-)provisioned on a client. We also describe how a registry can efficiently maintain the \CDMT{} index as new image versions are pushed. We show the scalability of \CDMT{} over Merkle Trees in terms of disk and network I/O savings using 15 container images and 233 image versions from Docker Hub.

cs.DB