SearcharxivSearch

arXiv subjects

Satoshi Tsutsui

Publications and source records attributed to Satoshi Tsutsui.

At least 19 recordsLinked to original sources

WBCAtt+: Fine-Grained Pixel-Level Morphological Annotations for White Blood Cell Images

The microscopic examination of white blood cells (WBCs) plays a fundamental role in pathology and is essential for diagnosing blood disorders such as leukemia and anemia. To support further research on WBC images, multiple datasets have been proposed. However, they mainly annotate cell categories, and lack detailed morphological characteristics that pathologists use to explain their interpretations of cells. To address this gap, we introduce WBCAtt+, a novel dataset of WBC images densely annotated with 11 morphological attributes and five pixel-level cell components. With 113k image-level labels and 10k segmentation maps, WBCAtt+ is the first to provide comprehensive annotations for WBC images. Leveraging this dataset, we provide baseline models for attribute recognition and semantic segmentation. We also design an attribute recognition model to incorporate compositional structure of cells, further improving the recognition performance. Lastly, we showcase various applications enabled by our dataset, such as explainable AI models, including counterfactual example generation. \revision{The dataset and code are publicly available\footnote{https://doi.org/10.57967/hf/8143}}.

cs.CV

Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models

Infants learn not only object categories but also fine-grained visual attributes such as color, size, and texture from limited experience. Prior infant-scale vision--language models have mainly been evaluated on object recognition, leaving open whether they support within-class attribute discrimination. We introduce a controlled benchmark that varies color, size, and texture across 67 everyday object classes using synthetic rendering to decouple attribute values from object identity. We evaluate infant-trained models (CVCL and an infant-trained DINO baseline) against web-scale and ImageNet models (CLIP, SigLIP, ResNeXt) under two complementary settings: an image-only prototype test and a text--vision test with attribute--object prompts. We find a dissociation between visual and linguistic attribute information: infant-trained models form strong visual representations for size and discriminate texture comparably to other models, but perform poorly on visual color discrimination, and in the text--vision setting they struggle to ground color and show only modest size grounding. In contrast, web-trained vision--language models strongly ground color from text while exhibiting weaker visual size discrimination.

cs.LG

Acoustic phonon softening and lattice instability driven by on-site $f$-$d$ hybridization in CeCoSi

Soft phonon modes in tetragonal CeCoSi, which undergoes a structural transition at $T_0=12$ K followed by antiferromagnetic order at $T_{\text{N}}=9.5$ K, have been investigated using high-resolution inelastic x-ray scattering. Pronounced softening was detected in the transverse acoustic modes corresponding to the $(yz+zx)$-type monoclinic distortion, consistent with the experimentally determined triclinic structure. Remarkably, the softening persists up to the zone boundary along (0, 0, $q$), indicating a short correlation length of the lattice instability. This instability, characterized by a Curie-type strain susceptibility, is interpreted as a consequence of the on-site $4f$-$5d$ hybridization, which is intrinsic to this crystal structure due to the lack of inversion symmetry at the two Ce sites.

cond-mat.str-el

Digital Staining with Knowledge Distillation: A Unified Framework for Unpaired and Paired-But-Misaligned Data

Staining is essential in cell imaging and medical diagnostics but poses significant challenges, including high cost, time consumption, labor intensity, and irreversible tissue alterations. Recent advances in deep learning have enabled digital staining through supervised model training. However, collecting large-scale, perfectly aligned pairs of stained and unstained images remains difficult. In this work, we propose a novel unsupervised deep learning framework for digital cell staining that reduces the need for extensive paired data using knowledge distillation. We explore two training schemes: (1) unpaired and (2) paired-but-misaligned settings. For the unpaired case, we introduce a two-stage pipeline, comprising light enhancement followed by colorization, as a teacher model. Subsequently, we obtain a student staining generator through knowledge distillation with hybrid non-reference losses. To leverage the pixel-wise information between adjacent sections, we further extend to the paired-but-misaligned setting, adding the Learning to Align module to utilize pixel-level information. Experiment results on our dataset demonstrate that our proposed unsupervised deep staining method can generate stained images with more accurate positions and shapes of the cell targets in both settings. Compared with competing methods, our method achieves improved results both qualitatively and quantitatively (e.g., NIQE and PSNR).We applied our digital staining method to the White Blood Cell (WBC) dataset, investigating its potential for medical applications.

cs.CV

Towards Robust and Reliable Concept Representations: Reliability-Enhanced Concept Embedding Model

Concept Bottleneck Models (CBMs) aim to enhance interpretability by predicting human-understandable concepts as intermediates for decision-making. However, these models often face challenges in ensuring reliable concept representations, which can propagate to downstream tasks and undermine robustness, especially under distribution shifts. Two inherent issues contribute to concept unreliability: sensitivity to concept-irrelevant features (e.g., background variations) and lack of semantic consistency for the same concept across different samples. To address these limitations, we propose the Reliability-Enhanced Concept Embedding Model (RECEM), which introduces a two-fold strategy: Concept-Level Disentanglement to separate irrelevant features from concept-relevant information and a Concept Mixup mechanism to ensure semantic alignment across samples. These mechanisms work together to improve concept reliability, enabling the model to focus on meaningful object attributes and generate faithful concept representations. Experimental results demonstrate that RECEM consistently outperforms existing baselines across multiple datasets, showing superior performance under background and domain shifts. These findings highlight the effectiveness of disentanglement and alignment strategies in enhancing both reliability and robustness in CBMs.

cs.CV

Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning

Infants develop complex visual understanding rapidly, even preceding the acquisition of linguistic skills. As computer vision seeks to replicate the human vision system, understanding infant visual development may offer valuable insights. In this paper, we present an interdisciplinary study exploring this question: can a computational model that imitates the infant learning process develop broader visual concepts that extend beyond the vocabulary it has heard, similar to how infants naturally learn? To investigate this, we analyze a recently published model in Science by Vong et al., which is trained on longitudinal, egocentric images of a single child paired with transcribed parental speech. We perform neuron labeling to identify visual concept neurons hidden in the model's internal representations. We then demonstrate that these neurons can recognize objects beyond the model's original vocabulary. Furthermore, we compare the differences in representation between infant models and those in modern computer vision models, such as CLIP and ImageNet pre-trained model. Ultimately, our work bridges cognitive science and computer vision by analyzing the internal representations of a computational model trained on an infant visual and linguistic inputs. Project page is available at https://kexueyi.github.io/webpage-discover-hidden-visual-concepts.

cs.CV

Integrating Clinical Knowledge into Concept Bottleneck Models

Concept bottleneck models (CBMs), which predict human-interpretable concepts (e.g., nucleus shapes in cell images) before predicting the final output (e.g., cell type), provide insights into the decision-making processes of the model. However, training CBMs solely in a data-driven manner can introduce undesirable biases, which may compromise prediction performance, especially when the trained models are evaluated on out-of-domain images (e.g., those acquired using different devices). To mitigate this challenge, we propose integrating clinical knowledge to refine CBMs, better aligning them with clinicians' decision-making processes. Specifically, we guide the model to prioritize the concepts that clinicians also prioritize. We validate our approach on two datasets of medical images: white blood cell and skin images. Empirical validation demonstrates that incorporating medical guidance enhances the model's classification performance on unseen datasets with varying preparation methods, thereby increasing its real-world applicability.

cs.CV

Evolving Storytelling: Benchmarks and Methods for New Character Customization with Diffusion Models

Diffusion-based models for story visualization have shown promise in generating content-coherent images for storytelling tasks. However, how to effectively integrate new characters into existing narratives while maintaining character consistency remains an open problem, particularly with limited data. Two major limitations hinder the progress: (1) the absence of a suitable benchmark due to potential character leakage and inconsistent text labeling, and (2) the challenge of distinguishing between new and old characters, leading to ambiguous results. To address these challenges, we introduce the NewEpisode benchmark, comprising refined datasets designed to evaluate generative models' adaptability in generating new stories with fresh characters using just a single example story. The refined dataset involves refined text prompts and eliminates character leakage. Additionally, to mitigate the character confusion of generated results, we propose EpicEvo, a method that customizes a diffusion-based visual story generation model with a single story featuring the new characters seamlessly integrating them into established character dynamics. EpicEvo introduces a novel adversarial character alignment module to align the generated images progressively in the diffusive process, with exemplar images of new characters, while applying knowledge distillation to prevent forgetting of characters and background details. Our evaluation quantitatively demonstrates that EpicEvo outperforms existing baselines on the NewEpisode benchmark, and qualitative studies confirm its superior customization of visual story generation in diffusion models. In summary, EpicEvo provides an effective way to incorporate new characters using only one example story, unlocking new possibilities for applications such as serialized cartoons.

cs.CV

Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces

Deepfake videos are becoming increasingly realistic, showing few tampering traces on facial areasthat vary between frames. Consequently, existing Deepfake detection methods struggle to detect unknown domain Deepfake videos while accurately locating the tampered region. To address thislimitation, we propose Delocate, a novel Deepfake detection model that can both recognize andlocalize unknown domain Deepfake videos. Ourmethod consists of two stages named recoveringand localization. In the recovering stage, the modelrandomly masks regions of interest (ROIs) and reconstructs real faces without tampering traces, leading to a relatively good recovery effect for realfaces and a poor recovery effect for fake faces. Inthe localization stage, the output of the recoveryphase and the forgery ground truth mask serve assupervision to guide the forgery localization process. This process strategically emphasizes the recovery phase of fake faces with poor recovery, facilitating the localization of tampered regions. Ourextensive experiments on four widely used benchmark datasets demonstrate that Delocate not onlyexcels in localizing tampered areas but also enhances cross-domain detection performance.

cs.CV

Basis Function Dependence of Estimation Precision for Synchrotron-Radiation-Based M\"ossbauer Spectroscopy

M\"ossbauer spectroscopy is a technique employed to investigate the microscopic properties of materials using transitions between energy levels in the nuclei. Conventionally, in synchrotron-radiation-based M\"ossbauer spectroscopy, the measurement window is decided by the researcher heuristically, although this decision has a significant impact on the shape of the measurement spectra. In this paper, we propose a method for evaluating the precision of the spectral position by introducing Bayesian estimation. The proposed method makes it possible to select the best measurement window by calculating the precision of M\"ossbauer spectroscopy from the data. Based on the results, the precision of the M\"ossbauer center shifts improved by more than three times compared with the results achieved with the conventional simple fitting method using the Lorentzian function.

physics.comp-ph

WBCAtt: A White Blood Cell Dataset Annotated with Detailed Morphological Attributes

The examination of blood samples at a microscopic level plays a fundamental role in clinical diagnostics, influencing a wide range of medical conditions. For instance, an in-depth study of White Blood Cells (WBCs), a crucial component of our blood, is essential for diagnosing blood-related diseases such as leukemia and anemia. While multiple datasets containing WBC images have been proposed, they mostly focus on cell categorization, often lacking the necessary morphological details to explain such categorizations, despite the importance of explainable artificial intelligence (XAI) in medical domains. This paper seeks to address this limitation by introducing comprehensive annotations for WBC images. Through collaboration with pathologists, a thorough literature review, and manual inspection of microscopic images, we have identified 11 morphological attributes associated with the cell and its components (nucleus, cytoplasm, and granules). We then annotated ten thousand WBC images with these attributes. Moreover, we conduct experiments to predict these attributes from images, providing insights beyond basic WBC classification. As the first public dataset to offer such extensive annotations, we also illustrate specific applications that can benefit from our attribute annotations. Overall, our dataset paves the way for interpreting WBC recognition models, further advancing XAI in the fields of pathology and hematology.

cs.CV

Structural and Dynamical Changes in a Gd-Co Metallic Glass by Cryogenic Rejuvenation

To experimentally clarify the changes in structural and dynamic heterogeneities in a metallic glass (MG), Gd65Co35, by rejuvenation with a temperature cycling (cryogenic rejuvenation), high-energy x-ray diffraction (HEXRD), anomalous x-ray scattering (AXS), and inelastic x-ray scattering (IXS) experiments were carried out. By a repeated temperature change between liquid N2 and room temperatures 40 times, tiny but clear structural changes are observed by HEXRD even in the first neighboring range. Partial structural information obtained by AXS reveals that slight movements of the Gd and Co atoms occur in the first- and second-neighboring shells around the central Gd atom. The concentration inhomogeneity in the nm size drastically increases for the Gd atoms by the temperature cycling, while the other heterogeneities are negligible. A distinct change was detected in a microscopic elastic property by IXS: The width of longitudinal acoustic excitation broadens by about 20%, indicating an increase of the elastic heterogeneity of this MG by the thermal treatments. These static and dynamic results explicitly clarify the features of the cryogenic rejuvenation effect experimentally.

cond-mat.mtrl-sci

Recap: Detecting Deepfake Video with Unpredictable Tampered Traces via Recovering Faces and Mapping Recovered Faces

The exploitation of Deepfake techniques for malicious intentions has driven significant research interest in Deepfake detection. Deepfake manipulations frequently introduce random tampered traces, leading to unpredictable outcomes in different facial regions. However, existing detection methods heavily rely on specific forgery indicators, and as the forgery mode improves, these traces become increasingly randomized, resulting in a decline in the detection performance of methods reliant on specific forgery traces. To address the limitation, we propose Recap, a novel Deepfake detection model that exposes unspecific facial part inconsistencies by recovering faces and enlarges the differences between real and fake by mapping recovered faces. In the recovering stage, the model focuses on randomly masking regions of interest (ROIs) and reconstructing real faces without unpredictable tampered traces, resulting in a relatively good recovery effect for real faces while a poor recovery effect for fake faces. In the mapping stage, the output of the recovery phase serves as supervision to guide the facial mapping process. This mapping process strategically emphasizes the mapping of fake faces with poor recovery, leading to a further deterioration in their representation, while enhancing and refining the mapping of real faces with good representation. As a result, this approach significantly amplifies the discrepancies between real and fake videos. Our extensive experiments on standard benchmarks demonstrate that Recap is effective in multiple scenarios.

cs.CV

Benchmarking White Blood Cell Classification Under Domain Shift

Recognizing the types of white blood cells (WBCs) in microscopic images of human blood smears is a fundamental task in the fields of pathology and hematology. Although previous studies have made significant contributions to the development of methods and datasets, few papers have investigated benchmarks or baselines that others can easily refer to. For instance, we observed notable variations in the reported accuracies of the same Convolutional Neural Network (CNN) model across different studies, yet no public implementation exists to reproduce these results. In this paper, we establish a benchmark for WBC recognition. Our results indicate that CNN-based models achieve high accuracy when trained and tested under similar imaging conditions. However, their performance drops significantly when tested under different conditions. Moreover, the ResNet classifier, which has been widely employed in previous work, exhibits an unreasonably poor generalization ability under domain shifts due to batch normalization. We investigate this issue and suggest some alternative normalization techniques that can mitigate it. We make fully-reproducible code publicly available\footnote{\url{https://github.com/apple2373/wbc-benchmark}}.

eess.IV

Mover: Mask and Recovery based Facial Part Consistency Aware Method for Deepfake Video Detection

Deepfake techniques have been widely used for malicious purposes, prompting extensive research interest in developing Deepfake detection methods. Deepfake manipulations typically involve tampering with facial parts, which can result in inconsistencies across different parts of the face. For instance, Deepfake techniques may change smiling lips to an upset lip, while the eyes remain smiling. Existing detection methods depend on specific indicators of forgery, which tend to disappear as the forgery patterns are improved. To address the limitation, we propose Mover, a new Deepfake detection model that exploits unspecific facial part inconsistencies, which are inevitable weaknesses of Deepfake videos. Mover randomly masks regions of interest (ROIs) and recovers faces to learn unspecific features, which makes it difficult for fake faces to be recovered, while real faces can be easily recovered. Specifically, given a real face image, we first pretrain a masked autoencoder to learn facial part consistency by dividing faces into three parts and randomly masking ROIs, which are then recovered based on the unmasked facial parts. Furthermore, to maximize the discrepancy between real and fake videos, we propose a novel model with dual networks that utilize the pretrained encoder and masked autoencoder, respectively. 1) The pretrained encoder is finetuned for capturing the encoding of inconsistent information in the given video. 2) The pretrained masked autoencoder is utilized for mapping faces and distinguishing real and fake videos. Our extensive experiments on standard benchmarks demonstrate that Mover is highly effective.

cs.MM

Mover: Mask and Recovery based Facial Part Consistency Aware Method for Deepfake Video Detection

Deepfake techniques have been widely used for malicious purposes, prompting extensive research interest in developing Deepfake detection methods. Deepfake manipulations typically involve tampering with facial parts, which can result in inconsistencies across different parts of the face. For instance, Deepfake techniques may change smiling lips to an upset lip, while the eyes remain smiling. Existing detection methods depend on specific indicators of forgery, which tend to disappear as the forgery patterns are improved. To address the limitation, we propose Mover, a new Deepfake detection model that exploits unspecific facial part inconsistencies, which are inevitable weaknesses of Deepfake videos. Mover randomly masks regions of interest (ROIs) and recovers faces to learn unspecific features, which makes it difficult for fake faces to be recovered, while real faces can be easily recovered. Specifically, given a real face image, we first pretrain a masked autoencoder to learn facial part consistency by dividing faces into three parts and randomly masking ROIs, which are then recovered based on the unmasked facial parts. Furthermore, to maximize the discrepancy between real and fake videos, we propose a novel model with dual networks that utilize the pretrained encoder and masked autoencoder, respectively. 1) The pretrained encoder is finetuned for capturing the encoding of inconsistent information in the given video. 2) The pretrained masked autoencoder is utilized for mapping faces and distinguishing real and fake videos. Our extensive experiments on standard benchmarks demonstrate that Mover is highly effective.

cs.CV

Cluster Toroidal Multipoles Formed by Electric-Quadrupole and Magnetic-Octupole Trimers: A Possible Scenario for Hidden Orders in Ca$_5$Ir$_3$O$_{12}$

Cluster multipole orderings composed of atomic high-rank multipole moments are theoretically investigated with a 5$d$-electron compound Ca$_5$Ir$_3$O$_{12}$ in mind. Ca$_5$Ir$_3$O$_{12}$ exhibits two hidden orders: One is an intermediate-temperature phase with time-reversal symmetry and the other is a low-temperature phase without time-reversal symmetry. By performing the symmetry and augmented multipole analyses for a $d$-orbital model under the hexagonal point group $D_{\rm 3h}$, we find that the 120$^{\circ}$-type ordering of the electric quadrupole corresponds to cluster electric toroidal dipole ordering with the electric ferroaxial moment, which can become the microscopic origin of the intermediate-temperature phase in Ca$_5$Ir$_3$O$_{12}$. Furthermore, based on ${}^{193}$Ir synchrotron-radiation-based Mössbauer spectroscopy, we propose that the low-temperature phase in Ca$_5$Ir$_3$O$_{12}$ is regarded as a coexisting state with cluster electric toroidal dipole and cluster magnetic toroidal quadrupole, the latter of which is formed by the 120$^{\circ}$-type ordering of the magnetic octupole and accompanies a small uniform magnetization as a secondary effect. Our results provide a clue to two hidden phases in Ca$_5$Ir$_3$O$_{12}$.

cond-mat.str-el

Experimental Observation of Mesoscopic Fluctuations to Identify Origin of Thermodynamic Anomalies of Ambient Liquid Water

We report a new experimental approach for observing mesoscopic fluctuations underlying the thermodynamic anomalies of ambient liquid water. In this approach, two sound velocity measurements with different frequencies, namely inelastic X-ray scattering (IXS) in THz band and ultrasonic (US) in MHz band, are required to investigate the relaxation phenomenon with the characteristic frequency between the two aforementioned frequencies. We performed IXS measurements to obtain the IXS sound velocity of liquid water from the ambient conditions to the supercritical region of liquid-gas phase transition (LGT) and compared the results with the US sound velocity in the literature. We found that the ratio of the two sound velocities, Sf, which corresponds to the relaxation intensity, exhibits a simple but significant change. Two distinct rises were observed in the high-temperature and low-temperature regions, implying that two relaxation phenomena exist: in the high-temperature region, a peak was observed near the LGT critical ridge line, which was linked with changes in the density fluctuation and isochoric and isobaric specific heat capacities; in the low-temperature region, Sf increased toward the low-temperature region, which was linked with the change in the isochoric heat capacity. We concluded that these two relaxation phenomena are originated from critical fluctuations of liquid-gas phase transition (LGT) and liquid-liquid phase transition, respectively. The linkage between Sf and isochoric heat capacity in the low-temperature region proves that the relaxation is the cause of the well-known heat capacity anomaly of ambient liquid water. In this study, both LGT and LLT critical fluctuations were observed, and the relationship between thermodynamics and the critical fluctuations was comprehensively discussed.

cond-mat.soft