SearcharxivSearch

arXiv subjects

Min Jung Kim

Publications and source records attributed to Min Jung Kim.

6 recordsLinked to original sources

Ablation-Corrected Evaluation of Attribution Maps in Echocardiographic Ejection-Fraction Models

Attribution maps for echocardiographic ejection-fraction models are evaluated by their overlap with an expert left-ventricular annotation, compared against a chance level that is computed from an area ratio rather than measured. We measure it. Two architectures trained on the same task attain overlap at 3.55 and 4.20 times measured chance, an eighteen percent difference a reader would take as the size of the gap between them. It is not. Replacing the annotated ventricle with a composition-matched surrogate changes one model's prediction six times more than replacing an equal-area control region and the other's twice, a factor of three; on sixty-two percent of cases for the second, neither the ventricle nor the attributed region moves the prediction at all. The reliability of this measurement, from two independent intervention batches, is 0.90 to 0.98. Overlap does not track the difference: scored against an ablation-derived label of which cases the model depends on the ventricle for, it reaches an area under the curve of 0.52 and 0.67. The design of the ablation also determines its answer, since ablating only the two annotated frames leaves the ventricle indistinguishable from a control region while ablating across all frames does not. A pediatric cohort reproduces all of these. We give the measured chance level, the symmetric-ablation protocol, the reliability estimate, and the ablation-prediction score as things to report alongside overlap.

cs.CV

Recovery Rates Are Not Comparable Across Transcription Factors: Chance Correction for Attribution Evaluation

Attribution methods for genomic sequence models are commonly evaluated by how much of a known motif they recover, or by how a prediction degrades as evidence is deleted. Neither score is interpretable without the value it would take by chance, and neither is routinely reported against one. We show that this omission is not a matter of precision but of validity. The uniform chance level for contiguous motif overlap is \(L/(N-L+1)\); across 268 transcription factors in UniBind it ranges from 0.0118 to 0.0427, a 3.6-fold spread determined by motif length and window size alone. For two factors the bootstrap intervals of the chance levels themselves do not overlap, so their raw recovery rates are not comparable quantities. Correcting for this dissolves a published three-way classification of five factors: a factor reported as a resolution failure attains the second-highest corrected value, ahead of one of the two positive controls, and two reported as complete failures fall at or below chance. We further show that perturbation-based evaluation can fail its own precondition: for one factor a fully masked input still scores above the decision boundary, and the curve is not monotone in the number of masked positions, so the area under it is not a measure of faithfulness. We provide chance levels in closed form, a chance-corrected score, and two screens that run before any attribution is computed.

q-bio.OT

Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

Objective: Concept bottleneck models route prediction through interpretable intermediate variables, and their validity is normally judged by how accurately those variables are predicted. We ask whether that judgement is sufficient, using left ventricular volumes as the concepts underlying ejection fraction estimation from echocardiographic video. Methods: A video transformer encoder was trained on a publicly available echocardiography dataset. End-systolic and end-diastolic volumes formed a concept layer from which ejection fraction was computed analytically, with no residual path to the output. We compared training under an ejection fraction objective alone against training with additional supervision of the volumes in millilitres, and evaluated both on 1276 held-out studies. Results: The concept bottleneck did not increase ejection fraction error relative to direct regression, at 6.89 against 7.13 mean absolute error. Without volume supervision, however, the spread of predicted volumes collapsed to 0.1 millilitres against reference spreads of 35.7 and 45.7 millilitres, while correlation was partly preserved. We show that this follows from an invariance property of the objective: ejection fraction is a ratio and is unchanged when both volumes are rescaled, so the loss determines the concept layer only up to scale. Supervision in absolute units reduced volume error from 89.8 to 25.8 millilitres at a cost of 0.4 in ejection fraction error. Conclusion: Concept accuracy alone can conceal a concept layer that carries no physical scale. Significance: Interpretable intermediate variables in clinical models should be validated against the invariance structure of the training objective, not only against prediction accuracy.

cs.AI

MC-GenRef: Annotation-free mammography microcalcification segmentation with generative posterior refinement

Microcalcification (MC) analysis is clinically important in screening mammography because clustered puncta can be an early sign of malignancy, yet dense MC segmentation remains challenging: targets are extremely small and sparse, dense pixel-level labels are expensive and ambiguous, and cross-site shift often induces texture-driven false positives and missed puncta in dense tissue. We propose MC-GenRef, a real dense-label-free framework that combines high-fidelity synthetic supervision with test-time generative posterior refinement (TT-GPR). During training, real negative mammogram patches are used as backgrounds, and physically plausible MC patterns are injected through a lightweight image formation model with local contrast modulation and blur, yielding exact image-mask pairs without real dense annotation. Using only these synthetic labeled pairs, MC-GenRef trains a base segmentor and a seed-conditioned rectified-flow (RF) generator that serves as a controllable generative prior. During inference, TT-GPR treats segmentation as approximate posterior inference: it derives a sparse seed from the current prediction, forms seed-consistent RF projections, converts them into case-specific surrogate targets through the frozen segmentor, and iteratively refines the logits with overlap-consistent and edge-aware regularization. On INbreast, the synthetic-only initializer achieved the best Dice without real dense annotations, while TT-GPR improved miss-sensitive performance to Recall and FNR, with strong class-balanced behavior (Bal.Acc., G-Mean). On an external private Yonsei cohort ( n=50 ), TT-GPR consistently improved the synthetic-only initializer under cross-site shift, increasing Dice and Recall while reducing FNR. These results suggest that test-time generative posterior refinement is a practical route to reduce MC misses and improve robustness without additional real dense labeling.

eess.IV

Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children

This study presents a model of automatic speech recognition (ASR) designed to diagnose pronunciation issues in children with speech sound disorders (SSDs) to replace manual transcriptions in clinical procedures. Since ASR models trained for general purposes primarily predict input speech into real words, employing a well-known high-performance ASR model for evaluating pronunciation in children with SSDs is impractical. We fine-tuned the wav2vec 2.0 XLS-R model to recognize speech as pronounced rather than as existing words. The model was fine-tuned with a speech dataset from 137 children with inadequate speech production pronouncing 73 Korean words selected for actual clinical diagnosis. The model's predictions of the pronunciations of the words matched the human annotations with about 90% accuracy. While the model still requires improvement in recognizing unclear pronunciation, this study demonstrates that ASR models can streamline complex pronunciation error diagnostic procedures in clinical fields.

cs.CL

On the maximum area of inscribed polygons

Given a convex $n$-gon $P$ and a positive integer $m$ such that $3\le m\le n-1$, let $Q$ denote the largest area convex $m$-gon contained in $P$. We are interested in the minimum value of $Δ(Q)/Δ(P)$, the ratio of the areas of these two polygons. More precisely, given positive integers $n$ and $m$, with $3 \le m \le n-1$, define \begin{equation*} f_n(m)=\min_{P\in \mathcal {P}_n} \max_{Q \subset P,|Q|=m} \frac{Δ(Q)}{Δ(P)} \end{equation*} where the maximum is taken over all $m$-gons contained in $P$, and the minimum is taken over $\mathcal{P}_n$, the entire class of convex $n$-gons. The values of $f_4(3)$, $f_5(4)$ and $f_6(3)$ are known. In this paper we compute the values of $f_5(3)$, $f_6(5)$ and $f_6(4)$. In addition, we prove that for all $n\ge 6$ we have \begin{equation*} \frac{4}{n}\cdot\sin^2\left(\fracπ{n}\right)\le 1-f_n(n-1)\le \min\left(\frac{1}{n}, \frac{4}{n}\cdot\sin^2\left(\frac{2π}{n}\right)\right). \end{equation*} These bounds can be used to improve the known estimates for $f_n(m)$.

math.CO