SearcharxivSearch

arXiv subjects

Seunghyun Kim

Publications and source records attributed to Seunghyun Kim.

12 recordsLinked to original sources

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry

This paper describes the participation of team "Go-To-Germany" in the ImageCLEF 2026 Audio Deepfake Detection and Generation task. Our detection system, built on a four-backbone self-supervised learning (SSL) ensemble combining WavLM-Large, Wav2Vec2-XLS-R-300M, ECAPA-TDNN, and x-vector representations, achieved a final score of 0.9522 on the official ImageCLEF 2026 evaluation, with perfect accuracy (1.0000) on participant-generated deepfakes and 0.8875 on the held-out organizer ground-truth real data. For the Generation sub-task, our official team submission, an F5-TTS v1 baseline processed with a uniform reverberation pass and submitted as a deliberate anti-forensic probe, ranked first with a final score of 0.4304 (word error rate (WER) 4.99%, character error rate (CER) 2.07%); details of our four-model program (GLM-TTS, F5-TTS, XTTS v2, CosyVoice3), from which the official entry was drawn, appear in the paper. We present a cross-track analysis revealing a pronounced asymmetry: our detection system identifies 100% of participant-generated deepfakes, while our official generation entry, despite ranking first in the Audio Generation sub-task and evading 61.4% and 56.2% of participant and organizer detectors, attains a Final Score of 0.4304 against 0.9522 on the Detection side. We further report falsification-based ablation experiments (LOSO 56-speaker cross-validation, three-region backbone geometry, bootstrap confidence intervals, and PCA analysis) that motivate our architectural-insurance hypothesis for multi-backbone SSL ensembling. We complement these results with five cross-track insights and five pre-registered falsification experiments connecting generation-side evasion to detection-side design decisions, and we openly report an 11.25% false-positive gap on held-out organizer real recordings as the principal open challenge for deployment.

cs.SD

Adversarial Deepfake Generation and an Investigation of Purification-Based Adversarial Detection

This paper describes the participation of team "Go To Germany" in the ImageCLEF 2026 Deepfake Detection and Generation Task. For the image generation task, we employ FLUX.1-dev with PuLID for identity-preserving face synthesis, combined with a multi-model PGD adversarial attack targeting 12 detectors simultaneously (DiffJPEG-in-loop, MI/DI/EoT, adaptive weighting, two-stage warm-start). Our approach achieved 90% evasion against organizer detectors and 57.6% against participant detectors, with a final generation score of 0.4170. For the image detection task, we combine two complementary detectors - SigLIP+DINOv2 for AI-generated images and GenD-DINOv3 for face manipulations - in a max-probability ensemble, achieving 99.4% accuracy on baseline deepfakes but suffering from high false-positive rates on real images, resulting in a final detection score of 0.6986. Beyond the official submission, we conducted a self-initiated investigation of purification-based adversarial detection, comparing three families of detection signals across six detectors that share a CLIP ViT-L/14 backbone. We find that raw $|Δ\text{logit}|$ under median-3 purification, applied through the EFFORT detector, separates adversarial inputs from clean inputs with AUROC 0.81-0.98 across four adversarial source types - a finding that refutes the simple backbone-preservation hypothesis and exposes a sharp JPEG-quality cliff at Q70 where the signal collapses.

cs.CV

Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory Protection

As DRAM scales to higher density and I/O speeds, ensuring data correctness becomes increasingly difficult. Industry has responded with a three-layer stack: on-die ECC (O-ECC), link ECC (L-ECC), and system ECC (S-ECC). However, these layers have evolved independently, often duplicating redundancy, leaving coverage gaps, and occasionally interfering. We propose Cerberus, a cross-layer ECC co-design that unifies protection across device, link, and system while preserving the native role of each layer. At its core is an Encode-Once, Decode-Many (EODM) architecture: the controller performs a single encoding whose redundancy is reused by L-ECC for immediate write-path detection and retry, by O-ECC for in-device repair on reads, and by S-ECC for strong end-to-end recovery. Cerberus jointly designs complementary parity and syndrome structures, orders decoders, and allocates the correction budget to prevent miscorrection amplification and enable selective correction under tight redundancy constraints. Our evaluations show improved resilience to clustered and peripheral faults while reducing redundant overhead, underscoring the importance of coordinated cross-layer protection for next-generation memory systems, such as custom HBMs.

cs.AR

RLDX-1 Technical Report

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs (e.g. $π_{0.5}$ and GR00T N1.6) across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8% while $π_{0.5}$ and GR00T N1.6 achieve around 40%, highlighting the ability of RLDX-1 to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

cs.RO

HFI: A unified framework for training-free detection and implicit watermarking of latent diffusion model generated images

Dramatic advances in the quality of the latent diffusion models (LDMs) also led to the malicious use of AI-generated images. While current AI-generated image detection methods assume the availability of real/AI-generated images for training, this is practically limited given the vast expressibility of LDMs. This motivates the training-free detection setup where no related data are available in advance. The existing LDM-generated image detection method assumes that images generated by LDM are easier to reconstruct using an autoencoder than real images. However, we observe that this reconstruction distance is overfitted to background information, leading the current method to underperform in detecting images with simple backgrounds. To address this, we propose a novel method called HFI. Specifically, by viewing the autoencoder of LDM as a downsampling-upsampling kernel, HFI measures the extent of aliasing, a distortion of high-frequency information that appears in the reconstructed image. HFI is training-free, efficient, and consistently outperforms other training-free methods in detecting challenging images generated by various generative models. We also show that HFI can successfully detect the images generated from the specified LDM as a means of implicit watermarking. HFI outperforms the best baseline method while achieving magnitudes of

cs.CV

Well-posedness issues for the generalized Benjamin--Bona--Mahony equation

In this paper, we consider the one-dimensional generalized Benjamin--Bona--Mahony (gBBM) equation \[(1-\partial_x^2)u_t+(u+u^p)_x=0,\qquad p=2,3,4,\dots,\] posed either on the real line $\mathbb R$ or on the torus $\mathbb T$. This equation may be viewed as a regularized model for the propagation of long-crested surface water waves. The main results of this work are threefold: \medskip First, we establish \emph{unconditional local well-posedness} in the class $C([0,T];H^s)$ without imposing any auxiliary spaces for \[s\ge \frac{p-2}{2p},\] which is \emph{sharp} in the sense that the multilinear estimate in $H^s$ is optimal. In addition, we prove \emph{unconditional uniqueness} for all distributional solutions in $L^\infty((0,T);H^s)$. \medskip Second, we show that below this regularity threshold, the flow map cannot be of class $C^p$. Precisely, if the flow map is well-defined and continuous near the origin from $H^s$ to $C([0,T];H^s)$ for every $s<\frac{p-2}{2p}$, then it cannot be of class $C^p$ at the origin. The proof is based on a high-to-low frequency interaction, implemented differently on $\mathbb R$ and $\mathbb T$. \medskip Third, in the odd-power case, we prove \emph{global well-posedness} below $H^1$ in the following cases: $p=3$ with $s\ge \frac14$, and $p=5$ with $s>\frac12$. To the best of our knowledge, these are the first global well-posedness results in the Sobolev framework for the generalized BBM equation below $H^1$. The argument is based on the Bona--Tzvetkov approach \cite{BT}, while being initially inspired by Bourgain's high--low method \cite{Bourgain1998, Bourgain1999}. A key new ingredient is the use of a Hamiltonian conservation law below the $H^1$ energy level. This allows us to control the higher-degree nonlinear contributions in the energy estimate, thereby preventing the Grönwall iteration from blowing up.

math.AP

Sparsification of the Generalized Persistence Diagrams for Scalability through Gradient Descent

The generalized persistence diagram (GPD) is a natural extension of the classical persistence barcode to the setting of multi-parameter persistence and beyond. The GPD is defined as an integer-valued function whose domain is the set of intervals in the indexing poset of a persistence module, and is known to be able to capture richer topological information than its single-parameter counterpart. However, computing the GPD is computationally prohibitive due to the sheer size of the interval set. Restricting the GPD to a subset of intervals provides a way to manage this complexity, compromising discriminating power to some extent. However, identifying and computing an effective restriction of the domain that minimizes the loss of discriminating power remains an open challenge. In this work, we introduce a novel method for optimizing the domain of the GPD through gradient descent optimization. To achieve this, we introduce a loss function tailored to optimize the selection of intervals, balancing computational efficiency and discriminative accuracy. The design of the loss function is based on the known erosion stability property of the GPD. We showcase the efficiency of our sparsification method for dataset classification in supervised machine learning. Experimental results demonstrate that our sparsification method significantly reduces the time required for computing the GPDs associated to several datasets, while maintaining classification accuracies comparable to those achieved using full GPDs. Our method thus opens the way for the use of GPD-based methods to applications at an unprecedented scale.

math.AT

Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments

In embodied instruction-following (EIF), the integration of pretrained language models (LMs) as task planners emerges as a significant branch, where tasks are planned at the skill level by prompting LMs with pretrained skills and user instructions. However, grounding these pretrained skills in different domains remains challenging due to their intricate entanglement with the domain-specific knowledge. To address this challenge, we present a semantic skill grounding (SemGro) framework that leverages the hierarchical nature of semantic skills. SemGro recognizes the broad spectrum of these skills, ranging from short-horizon low-semantic skills that are universally applicable across domains to long-horizon rich-semantic skills that are highly specialized and tailored for particular domains. The framework employs an iterative skill decomposition approach, starting from the higher levels of semantic skill hierarchy and then moving downwards, so as to ground each planned skill to an executable level within the target domain. To do so, we use the reasoning capabilities of LMs for composing and decomposing semantic skills, as well as their multi-modal extension for assessing the skill feasibility in the target domain. Our experiments in the VirtualHome benchmark show the efficacy of SemGro in 300 cross-domain EIF scenarios.

cs.AI

An Experimental Investigation of Cavitation Bulk Nanobubbles Characteristics: Effects of pH and Surface-active Agents

Understanding the behavior of nanobubbles (NBs) in various aqueous solutions is a challenging task. The present work investigates the effects of various surfactants (i.e., anionic, cationic, and nonionic) and pH medium on bulk NBs formation, size, concentration, bubble size distribution (BSD), zeta potential, and stability. The effect of surfactant was investigated at various concentrations above and below critical micelle concentrations. NBs were created in DI water using a piezoelectric transducer. The stability of NBs was assessed by tracking the change in size and concentration over time. NBs size is small in the neutral medium compared to the other surfactant or pH mediums. The size, concentration, BSD, and stability of NBs are strongly influenced by the zeta potential rather than the solution medium. BSD curve shifts to lower bubble sizes when the magnitude of zeta potential is high in any solution. NBs were observed to exist for a long time, either in pure water, surfactant, or pH solutions. The longevity of NBs is shortened in environments with pH less than 3. Surfactant adsorption on the NBs surface increases with surfactant concentration up to a certain limit, beyond which it declines considerably. The Derjaguin-LandauVerwey-Overbeek (DLVO) theory was used to interpret the NBs stability, which resulted in a total potential energy barrier that is positive and greater than 43.90kBT for pH ranging from 6.0 to 11.0, whereas, for pH below 6, the potential energy barrier essentially vanishes. Moreover, an effort has also been made to elucidate the plausible prospect of ion distribution and its alignment surrounding NBs in cationic and anionic surfactants. The present research will extend the in-depth investigation of NBs for industrial applications involving NBs.

physics.flu-dyn

Pressure Test: Quantifying the impact of positive stress on companies from online employee reviews

Workplace stress is often considered to be negative, yet lab studies on individuals suggest that not all stress is bad. There are two types of stress: distress refers to harmful stimuli, while eustress refers to healthy, euphoric stimuli that create a sense of fulfillment and achievement. Telling the two types of stress apart is challenging, let alone quantifying their impact across corporations. By leveraging a dataset of 440K reviews about S&P 500 companies published during twelve successive years, we developed a deep learning framework to extract stress mentions from these reviews. We proposed a new methodology that places each company on a stress-by-rating quadrant (based on its overall stress score and overall rating on the site), and accordingly scores the company to be, on average, either a low stress}, passive, negative stress, or positive stress company. We found that (former) employees of positive stress companies tended to describe high-growth and collaborative workplaces in their reviews, and that such companies' stock evaluations grew, on average, 5.1 times in 10 years (2009-2019) as opposed to the companies of the other three stress types that grew, on average, 3.7 times in the same time period. We also found that the four stress scores aggregated every year -- from 2008 to 2020 -- closely followed the unemployment rate in the U.S.: a year of positive stress (2008) was rapidly followed by several years of negative stress (2009-2015), which peaked during the Great Recession (2009-2011). These results suggest that automated analyses of the language used by employees on corporate social-networking tools offer yet another way of tracking workplace stress, allowing quantification of its impact on corporations.

cs.SI

Online Stochastic Gradient Methods Under Sub-Weibull Noise and the Polyak-Łojasiewicz Condition

This paper focuses on the online gradient and proximal-gradient methods with stochastic gradient errors. In particular, we examine the performance of the online gradient descent method when the cost satisfies the Polyak-Łojasiewicz (PL) inequality. We provide bounds in expectation and in high probability (that hold iteration-wise), with the latter derived by leveraging a sub-Weibull model for the errors affecting the gradient. The convergence results show that the instantaneous regret converges linearly up to an error that depends on the variability of the problem and the statistics of the sub-Weibull gradient error. Similar convergence results are then provided for the online proximal-gradient method, under the assumption that the composite cost satisfies the proximal-PL condition. In the case of static costs, we provide new bounds for the regret incurred by these methods when the gradient errors are modeled as sub-Weibull random variables. Illustrative simulations are provided to corroborate the technical findings.

math.OC