SearcharxivSearch

arXiv subjects

Sung Park

Publications and source records attributed to Sung Park.

13 recordsLinked to original sources

iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding

Brain-computer interfaces (BCIs) for speech restoration hold transformative potential for the approximately 173,000--232,500 individuals worldwide with ALS-related dysarthria. Despite recent progress, high-performance speech BCIs have been demonstrated in only 22--31 patients globally, largely due to limitations in neural decoding accuracy and practical input interfaces. We present iPhoneme, a brain-to-text communication system that jointly addresses these challenges through integrated modeling and interaction design. The system combines a deep learning phoneme decoder based on a modified Conformer architecture (ConformerXL, 192.9M parameters) with a gaze-assisted phoneme input interface that mitigates the Midas touch problem in eye-tracking systems. The acoustic model incorporates a temporal prenet with multi-scale dilated convolutions and bidirectional GRU for neural jitter correction, temporal subsampling for CTC stability, and Pre-RMSNorm stabilization across 12 encoder blocks, trained with AdamW and cosine scheduling. On the interaction side, iPhoneme introduces a chorded gaze-plus-silent-speech paradigm that replaces dwell-time selection, enabling more efficient input. We evaluate the system on the T15 dataset (45 sessions, 8,071 trials) of 256-channel intracranial EEG from speech motor cortex regions. A 6-gram phoneme language model trained on 3.1M sequences, combined with WFST beam search (beam=128), achieves 92.14% phoneme accuracy (7.86% PER) and 73.39% word accuracy (26.61% WER), approximately 3% above prior state-of-the-art. The system operates on CPU with 180 ms latency, demonstrating real-time, high-accuracy brain-to-text communication for ALS.

cs.SD

The Empty Quadrant: AI Teammates for Embodied Field Learning

For four decades, AIED research has rested on what we term the Sedentary Assumption: the unexamined design commitment to a stationary learner seated before a screen. Mobile learning and museum guides have moved learners into physical space, and context-aware systems have delivered location-triggered content -- yet these efforts predominantly cast AI in the role of information-de-livery tool rather than epistemic partner. We map this gap through a 2 x 2 matrix (AI Role x Learning Environment) and identify an undertheorized intersection: the configuration in which AI serves as an epistemic teammate during unstruc-tured, place-bound field inquiry and learning is assessed through trajectory rather than product. To fill it, we propose Field Atlas, a framework grounded in embod-ied, embedded, enactive, and extended (4E) cognition, active inference, and dual coding theory that shifts AIED's guiding metaphor from instruction to sensemak-ing. The architecture pairs volitional photography with immediate voice reflec-tion, constrains AI to Socratic provocation rather than answer delivery, and ap-plies Epistemic Trajectory Modeling (ETM) to represent field learning as a con-tinuous trajectory through conjoined physical-epistemic space. We demonstrate the framework through a museum scenario and argue that the resulting trajecto-ries -- bound to a specific body, place, and time -- constitute process-based evi-dence structurally resistant to AI fabrication, offering a new assessment paradigm and reorienting AIED toward embodied, dialogic human-AI sensemaking in the wild.

cs.HC

When Should an AI Act? A Human-Centered Model of Scene, Context, and Behavior for Agentic AI Design

Agentic AI increasingly intervenes proactively by inferring users' situations from contextual data yet often fails for lack of principled judgment about when, why, and whether to act. We address this gap by proposing a conceptual model that reframes behavior as an interpretive outcome integrating Scene (observable situation), Context (user-constructed meaning), and Human Behavior Factors (determinants shaping behavioral likelihood). Grounded in multidisciplinary perspectives across the humanities, social sciences, HCI, and engineering, the model separates what is observable from what is meaningful to the user and explains how the same scene can yield different behavioral meanings and outcomes. To translate this lens into design action, we derive five agent design principles (behavioral alignment, contextual sensitivity, temporal appropriateness, motivational calibration, and agency preservation) that guide intervention depth, timing, intensity, and restraint. Together, the model and principles provide a foundation for designing agentic AI systems that act with contextual sensitivity and judgment in interactions.

cs.AI

Exploring the Effects of Generative AI Assistance on Writing Self-Efficacy

Generative AI (GenAI) is increasingly used in academic writing, yet its effects on students' writing self-efficacy remain contingent on how assistance is configured. This pilot study investigates how ideation-level, sentence-level, full-process, and no AI support differentially shape undergraduate writers' self-efficacy using a 2 by 2 experimental design with Korean undergraduates completing argumentative writing tasks. Results indicate that AI assistance does not uniformly enhance self-efficacy full AI support produced high but stable self-efficacy alongside signs of reduced ownership, sentence-level AI support led to consistent self-efficacy decline, and ideation-level AI support was associated with both high self-efficacy and positive longitudinal change. These findings suggest that the locus of AI intervention, rather than the amount of assistance, is critical in fostering writing self-efficacy while preserving learner agency.

cs.HC

The Effect of Empathic Expression Levels in Virtual Human Interaction: A Controlled Experiment

As artificial intelligence (AI) systems become increasingly embedded in everyday life, the ability of interactive agents to express empathy has become critical for effective human-AI interaction, particularly in emotionally sensitive contexts. Rather than treating empathy as a binary capability, this study examines how different levels of empathic expression in virtual human interaction influence user experience. We conducted a between-subject experiment (n = 70) in a counseling-style interaction context, comparing three virtual human conditions: a neutral dialogue-based agent, a dialogue-based empathic agent, and a video-based empathic agent that incorporates users' facial cues. Participants engaged in a 15-minute interaction and subsequently evaluated their experience using subjective measures of empathy and interaction quality. Results from analysis of variance (ANOVA) revealed significant differences across conditions in affective empathy, perceived naturalness of facial movement, and appropriateness of facial expression. The video-based empathic expression condition elicited significantly higher affective empathy than the neutral baseline (p < .001) and marginally higher levels than the dialogue-based condition (p < .10). In contrast, cognitive empathy did not differ significantly across conditions. These findings indicate that empathic expression in virtual humans should be conceptualized as a graded design variable, rather than a binary capability, with visually grounded cues playing a decisive role in shaping affective user experience.

cs.HC

Significant Other AI: Identity, Memory, and Emotional Regulation as Long-Term Relational Intelligence

Significant Others (SOs) stabilize identity, regulate emotion, and support narrative meaning-making, yet many people today lack access to such relational anchors. Recent advances in large language models and memory-augmented AI raise the question of whether artificial systems could support some of these functions. Existing empathic AIs, however, remain reactive and short-term, lacking autobiographical memory, identity modeling, predictive emotional regulation, and narrative coherence. This manuscript introduces Significant Other Artificial Intelligence (SO-AI) as a new domain of relational AI. It synthesizes psychological and sociological theory to define SO functions and derives requirements for SO-AI, including identity awareness, long-term memory, proactive support, narrative co-construction, and ethical boundary enforcement. A conceptual architecture is proposed, comprising an anthropomorphic interface, a relational cognition layer, and a governance layer. A research agenda outlines methods for evaluating identity stability, longitudinal interaction patterns, narrative development, and sociocultural impact. SO-AI reframes AI-human relationships as long-term, identity-bearing partnerships and provides a foundational blueprint for investigating whether AI can responsibly augment the relational stability many individuals lack today.

cs.HC

Design Framework for Conversational Agent in Couple relationships: A Systematic Review

The development of conversational agents (CAs) has shown strong potential in supporting mental health through dialogue. While many studies focus on CAs for individual psychological care, research on agents designed for couples facing relational or emotional challenges remains limited. This study aims to identify design considerations for CAs that address the relational context of couples and support their well-being. Following PRISMA guidelines, a systematic review was conducted across seven databases: CINAHL, Embase, PubMed, PsycINFO, Scopus, Web of Science, and the ACM Digital Library. Peer-reviewed empirical studies were screened, duplicates removed, and selection criteria applied, resulting in twelve studies for analysis. Thematic analysis was conducted across three dimensions: AI interaction design, relational framing, and technical limitations. Three key themes emerged: (1) the need for a relational expert persona, (2) technological directions leveraging state-of-the-art AI for relational specificity and emotional competence, and (3) a shift from content-centered to relationship-centered design. Based on these insights, eight design considerations are proposed for couple-oriented CAs: (1) agent persona, (2) individual mode, (3) concurrent mode, (4) conjoint mode, (5) ethics, (6) data and privacy, (7) interaction pattern, and (8) safety mechanism. These principles guide CAs as relational mediators capable of maintaining multiple alliances, respecting cultural and ethical boundaries, and ensuring fairness and emotional safety between partners. Ultimately, this review introduces a design framework that integrates relational theory with advanced AI technologies to inform future development of CAs for couple-based mental health interventions.

cs.HC

Persode: Personalized Visual Journaling with Episodic Memory-Aware AI Agent

Reflective journaling often lacks personalization and fails to engage Generation Alpha and Z, who prefer visually immersive and fast-paced interactions over traditional text-heavy methods. Visual storytelling enhances emotional recall and offers an engaging way to process personal expe- riences. Designed with these digital-native generations in mind, this paper introduces Persode, a journaling system that integrates personalized onboarding, memory-aware conversational agents, and automated visual storytelling. Persode captures user demographics and stylistic preferences through a tailored onboarding process, ensuring outputs resonate with individual identities. Using a Retrieval-Augmented Generation (RAG) framework, it prioritizes emotionally significant memories to provide meaningful, context-rich interactions. Additionally, Persode dynamically transforms user experiences into visually engaging narratives by generating prompts for advanced text-to-image models, adapting characters, backgrounds, and styles to user preferences. By addressing the need for personalization, visual engagement, and responsiveness, Persode bridges the gap between traditional journaling and the evolving preferences of Gen Alpha and Z.

cs.HC

Bandits for Online Calibration: An Application to Content Moderation on Social Media Platforms

We describe the current content moderation strategy employed by Meta to remove policy-violating content from its platforms. Meta relies on both handcrafted and learned risk models to flag potentially violating content for human review. Our approach aggregates these risk models into a single ranking score, calibrating them to prioritize more reliable risk models. A key challenge is that violation trends change over time, affecting which risk models are most reliable. Our system additionally handles production challenges such as changing risk models and novel risk models. We use a contextual bandit to update the calibration in response to such trends. Our approach increases Meta's top-line metric for measuring the effectiveness of its content moderation strategy by 13%.

cs.LG

Nanoscale spectroscopic studies of two different physical origins of the tip-enhanced force: dipole and thermal

When light illuminates the junction formed between a sharp metal tip and a sample, different mechanisms can con-tribute to the measured photo-induced force simultaneously. Of particular interest are the instantaneous force be-tween the induced dipoles in the tip and in the sample and the force related to thermal heating of the junction. A key difference between these two force mechanisms is their spectral behaviors. The magnitude of the thermal response follows a dissipative Lorentzian lineshape, which measures the heat exchange between light and matter, while the induced dipole response exhibits a dispersive spectrum and relates to the real part of the material polarizability. Be-cause the two interactions are sometimes comparable in magnitude, the origin of the nanoscale chemical selectivity in the recently developed photo-induced force microscopy (PiFM) is often unclear. Here, we demonstrate theoretically and experimentally how light absorption followed by nanoscale thermal expansion generates a photo-induced force in PiFM. Furthermore, we explain how this thermal force can be distinguished from the induced dipole force by tuning the relaxation time of samples. Our analysis presented here helps the interpretation of nanoscale chemical measure-ments of heterogeneous materials and sheds light on the nature of light-matter coupling in van der Waals materials.

physics.optics

Eigenmodes of a quartz tuning fork and their application to photo-induced force microscopy

We examine the mechanical eigenmodes of a quartz tuning fork (QTF) for the purpose of facilitat- ing its use as a probe for multi-frequency atomic force microscopy (AFM). We perform simulations based on the three-dimensional finite element method (FEM) and compare the observed motions of the beams with experimentally measured resonance frequencies of two QTF systems. The com- parison enabled us to assign the first seven asymmetric eigenmodes of the QTF. We also find that a modified version of single beam theory can be used to guide the assignment of mechanical eigen- modes of QTFs. The usefulness of the QTF for multi-frequency AFM measurements is demonstrated through photo-induced force microscopy (PiFM) measurements. By using the QTF in different con- figurations, we show that the vectorial components of the photo-induced force can be independently assessed, and that lateral forces can be probed in true non-contact mode.

physics.ins-det

Berkeley Supernova Ia Program I: Observations, Data Reduction, and Spectroscopic Sample of 582 Low-Redshift Type Ia Supernovae

In this first paper in a series we present 1298 low-redshift (z\leq0.2) optical spectra of 582 Type Ia supernovae (SNe Ia) observed from 1989 through 2008 as part of the Berkeley SN Ia Program (BSNIP). 584 spectra of 199 SNe Ia have well-calibrated light curves with measured distance moduli, and many of the spectra have been corrected for host-galaxy contamination. Most of the data were obtained using the Kast double spectrograph mounted on the Shane 3 m telescope at Lick Observatory and have a typical wavelength range of 3300-10,400 Ang., roughly twice as wide as spectra from most previously published datasets. We present our observing and reduction procedures, and we describe the resulting SN Database (SNDB), which will be an online, public, searchable database containing all of our fully reduced spectra and companion photometry. In addition, we discuss our spectral classification scheme (using the SuperNova IDentification code, SNID; Blondin & Tonry 2007), utilising our newly constructed set of SNID spectral templates. These templates allow us to accurately classify our entire dataset, and by doing so we are able to reclassify a handful of objects as bona fide SNe Ia and a few other objects as members of some of the peculiar SN Ia subtypes. In fact, our dataset includes spectra of nearly 90 spectroscopically peculiar SNe Ia. We also present spectroscopic host-galaxy redshifts of some SNe Ia where these values were previously unknown. [Abridged]

astro-ph.CO

A non-spherical core in the explosion of supernova SN 2004dj

An important and perhaps critical clue to the mechanism driving the explosion of massive stars as supernovae is provided by the accumulating evidence for asymmetry in the explosion. Indirect evidence comes from high pulsar velocities, associations of supernovae with long-soft gamma-ray bursts, and asymmetries in late-time emission-line profiles. Spectropolarimetry provides a direct probe of young supernova geometry, with higher polarization generally indicating a greater departure from spherical symmetry. Large polarizations have been measured for 'stripped-envelope' (that is, type Ic) supernovae, which confirms their non-spherical morphology; but the explosions of massive stars with intact hydrogen envelopes (type II-P supernovae) have shown only weak polarizations at the early times observed. Here we report multi-epoch spectropolarimetry of a classic type II-P supernova that reveals the abrupt appearance of significant polarization when the inner core is first exposed in the thinning ejecta (~90 days after explosion). We infer a departure from spherical symmetry of at least 30 per cent for the inner ejecta. Combined with earlier results, this suggests that a strongly non-spherical explosion may be a generic feature of core-collapse supernovae of all types, where the asphericity in type II-P supernovae is cloaked at early times by the massive, opaque, hydrogen envelope.

astro-ph