SearcharxivSearch

arXiv subjects

Hyunju Kim

Publications and source records attributed to Hyunju Kim.

At least 19 recordsLinked to original sources

SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange

Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new event vocabularies by matching frozen audio and visual features with text-encoded event names. However, related labels share evidence. An incorrect label can then score at least as high as a correct one. We call this a false co-activation (FCA). No scalar cutoff can reject the incorrect label while keeping every correct one. Class-specific thresholds may prevent that label from becoming a final prediction, but the FCA remains in the underlying score vector. We introduce SCoPE, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other. We derive an exact condition for when this competition removes an FCA in a two-label fit. With identical frozen CLIP+CLAP backbones on LLP, SCoPE improves Type@seg by 7.45 points and Event@seg by 5.04 points compared with the reported AV$^2$A values. The same fixed configuration transfers unchanged to OV-AVEBench and VGGSound-AVEL100k.

cs.CV

Longitudinal wearable monitoring and polygenic risk for incident major depressive disorder in the All of Us Research Program

Major depressive disorder (MDD) risk reflects both stable inherited liability and dynamic behavioral patterns, yet these dimensions are rarely examined together using long-term objective data in real-world settings. Here, we integrated genomic, electronic health record (EHR), and longitudinal Fitbit wearable data from 3,030 adults of genetically inferred European ancestry in the All of Us Research Program, 284 of whom developed EHR-recorded incident MDD after a 180-day baseline period. Time-varying Cox models examined associations of MDD polygenic risk scores (PRS), monthly wearable-derived physical activity and sleep features, and interactions between wearable features and MDD PRS with incident MDD. Higher MDD PRS, lower daily steps, lower light and vigorous physical activity, lower sleep efficiency, and greater sleep duration variability were associated with higher risk of EHR-recorded incident MDD. The associations of sedentary time and sleep duration variability with incident MDD differed across MDD PRS levels, with stronger risk associations among participants with higher MDD PRS. PRS-stratified hazard ratio curves further indicated that comparable estimated risk corresponded to more favorable behavioral levels (such as higher daily step counts and more stable sleep) among participants with higher MDD PRS than among those with lower MDD PRS. Sequentially integrating MDD PRS, baseline wearable features, monthly wearable features, and selected interactions between wearable features and MDD PRS increased model discrimination, with the C-index increasing from 0.637 to 0.705. These findings support the complementary value of inherited liability and longitudinal real-world behavioral monitoring for incident MDD risk characterization and may inform future work on genetically informed digital phenotyping for personalized risk monitoring and prevention.

stat.AP

Mechanics-trained neural coordinate mapping for B-spline analysis of crack-tip and corner singularities

Near a crack tip or re-entrant corner, fractional radial powers can have unbounded derivatives and slow the convergence of high-order splines. A singular mapping grades the computational radius so that the pulled-back field is smoother without changing the physical domain. Classical maps require the singular exponent and grading in advance. Here the radial coordinate is trained from the mechanics problem. It is the normalized integral of a positive neural density, which fixes both radial boundaries and ensures r'(s)>0 away from the collapsed tip. The radial grading exponent and density-correction weights are inferred from Galerkin energies evaluated at discrete equilibrium, without exponent labels or exact interior fields. We test the mapping in scalar and plane-strain B-spline formulations and compare it with the identity map, radially graded knot vectors, prescribed power maps, and an adaptive enriched B-spline method. At 156 vector degrees of freedom, the mechanics-trained map reduces the relative energy-norm error by a factor of 33.42 compared with the identity map. The maximum error in the mixed-mode stress intensity factors recovered on 3 contours is 1.923 x 10^(-5). For the straight crack, the learned density correction vanishes, and the prescribed r=s^2 map gives the same improvement. For a nonlinear Robin family with test parameters outside the training interval, the density correction remains nonzero and gives a fixed-q incremental energy gain of at least 4.839. Thus, for the problems considered here, the neural correction is useful when the required coordinate is not represented by a prescribed power map.

math.NA

Data-efficient reconstruction of critical quantum dynamics via blind fractional-envelope extrapolation

Simulating real-time dynamics of quantum systems is often limited to short times by entanglement growth. Finite-pole reconstructions such as linear prediction and related machineries extrapolate such data reliably when the spectrum is a finite set of excitations, but at criticality the low-energy spectrum is a power-law continuum $A(\omega)\sim|\omega|^{\alpha-1}$ -- a branch cut whose real-time tail $G(t)\sim t^{-\alpha}$ finitely many poles cannot represent. Here we develop a fractional-calculus-motivated envelope extrapolation for such data. Its structure is motivated by a fractional form of Schwinger--Dyson (fSD) equation, in which the Laplace symbol $s^{\alpha}$ carries the branch cut analytically while the residual self-energy remains meromorphic. On real data we employ the corresponding operational alternative -- the exponent $\alpha$ is selected blindly inside the fit window, the signal is detrended by $t^{\alpha}$, the residual is fitted by a stabilized finite-pole model, and the algebraic envelope is restored. On the critical XXZ chain this blind fractional-envelope method (fSD for short) extrapolates short-time data typically several-fold more accurately than finite-pole methods, with the exponent $\alpha$ identified blindly from the fit window alone and bracketing the closed-form Luttinger value at weak coupling. The same blind search finds the $z=2$ dilute-magnon exponent $\alpha=1/2$ at the $\Delta=1$ saturation transition, and the advantage persists in the gapped free-magnon phase with its sharp band edges. On noncritical dynamical mean-field spectra, whose low-frequency response is regular, fSD by contrast fails to select any stable fractional envelope and reduces to the standard pole result rather than manufacturing a spurious power law, making it an efficient and accurate route to quantum critical dynamics when only short simulation times are accessible.

cond-mat.str-el

Fast and Scalable Caputo Fractional Gradient Descent via Perturbation-Preserving Memory Compression

Fractional gradient descent (FGD) incorporates long-range memory through Caputo-type operators and has been shown to improve stability in ill-conditioned and nonconvex optimization problems. Despite these advantages, its practical use remains limited, mainly due to the high computational cost of evaluating history-dependent convolutions, which scales quadratically with the number of iterations. In this paper, we focus on making Caputo-based optimization computationally viable without sacrificing its intrinsic memory structure. We begin by expressing the fractional descent direction as a discrete convolution over past gradients, which provides a unified view of the method. Based on this formulation, we introduce two complementary mechanisms to reduce the cost of the memory term. The first uses a sum-of-exponentials (SOE) approximation of the power-law kernel, leading to efficient recursive updates. The second approach, newly proposed in this paper as dyadic hierarchical discrete convolution (DHDC), compresses the gradient history through a multiscale aggregation strategy. Rather than treating these approximations as purely numerical accelerations, we interpret them as perturbations of the ideal Caputo operator. This viewpoint allows us to analyze how the compressed memory affects the optimization dynamics. Under standard $\mu$-strong convexity and $L$-smoothness assumptions, we show that the resulting method still exhibits monotone descent and linear convergence, provided that the approximation error remains controlled.

math.OC

Preoperative Decline and Postoperative Recovery of Wearable-Derived Physical Activity Over a Four-Year Perioperative Period in Total Knee and Hip Arthroplasty: Evidence from the All of Us Research Program

Total knee arthroplasty (TKA) and total hip arthroplasty (THA) improve symptoms in end-stage osteoarthritis, yet long-term objective characterization of perioperative physical activity trajectories remains limited. We conducted a longitudinal observational study within the All of Us Research Program dataset, linking electronic health records with continuous Fitbit-derived step count data over a four-year perioperative window (two years before and two years after arthroplasty). Piecewise linear mixed-effects models characterized preoperative declines and postoperative recovery trajectories, and time-to-recovery was evaluated using Kaplan-Meier curves and Cox proportional hazards models under remote and immediate preoperative physical activity baseline definitions. Among 238 participants (147 TKA; 91 THA), both procedures exhibited progressive preoperative decline with distinct procedure-specific patterns and staged postoperative recovery: rapid improvement during weeks 1-6, decelerating gains through weeks 7-19/20, and subsequent stabilization through week 104. Recovery to remote and immediate baselines differed in timing (median 22 vs 13 weeks) and associated predictors. Higher immediate preoperative activity was associated with greater likelihood of recovery to habitual activity levels, underscoring the relevance of preoperative functional reserve and surgical timing. These findings demonstrate the potential of long-term wearable monitoring to refine assessment of functional outcomes, guide recovery expectations, and support perioperative management.

stat.AP

Physical Activity Trajectories Preceding Incident Major Depressive Disorder Diagnosis Using Consumer Wearable Devices in the All of Us Research Program: Case-Control Study

Low physical activity is a known risk factor for major depressive disorder (MDD), but changes in activity before a first clinical diagnosis remain unclear, especially using long-term objective measurements. This study characterized trajectories of wearable-measured physical activity during the year preceding incident MDD diagnosis. We conducted a retrospective nested case-control study using linked electronic health record and Fitbit data from the All of Us Research Program. Adults with at least 6 months of valid wearable data in the year before diagnosis were eligible. Incident MDD cases were matched to controls on age, sex, body mass index, and index time (up to four controls per case). Daily step counts and moderate-to-vigorous physical activity (MVPA) were aggregated into monthly averages. Linear mixed-effects models compared trajectories from 12 months before diagnosis to diagnosis. Within cases, contrasts identified when activity first significantly deviated from levels 12 months prior. The cohort included 4,104 participants (829 cases and 3,275 controls; 81.7% women; median age 48.4 years). Compared with controls, cases showed consistently lower activity and significant downward trajectories in both step counts and MVPA during the year before diagnosis (P < 0.001). Significant declines appeared about 4 months before diagnosis for step counts and 5 months for MVPA. Exploratory analyses suggested subgroup differences, including steeper declines in men, greater intensity reductions at older ages, and persistently low activity among individuals with obesity. Sustained within-person declines in physical activity emerged months before incident MDD diagnosis. Longitudinal wearable monitoring may provide early signals to support risk stratification and earlier intervention.

stat.AP

DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning

Aligning egocentric video with wearable sensors have shown promise for human action recognition, but face practical limitations in user discomfort, privacy concerns, and scalability. We explore exocentric video with ambient sensors as a non-intrusive, scalable alternative. While prior egocentric-wearable works predominantly adopt Global Alignment by encoding entire sequences into unified representations, this approach fails in exocentric-ambient settings due to two problems: (P1) inability to capture local details such as subtle motions, and (P2) over-reliance on modality-invariant temporal patterns, causing misalignment between actions sharing similar temporal patterns with different spatio-semantic contexts. To resolve these problems, we propose DETACH, a decomposed spatio-temporal framework. This explicit decomposition preserves local details, while our novel sensor-spatial features discovered via online clustering provide semantic grounding for context-aware alignment. To align the decomposed features, our two-stage approach establishes spatial correspondence through mutual supervision, then performs temporal alignment via a spatial-temporal weighted contrastive loss that adaptively handles easy negatives, hard negatives, and false negatives. Comprehensive experiments with downstream tasks on Opportunity++ and HWU-USP datasets demonstrate substantial improvements over adapted egocentric-wearable baselines.

cs.CV

Beyond Gaussian Initializations: Signal Preserving Weight Initialization for Odd-Sigmoid Activations

Activation functions critically influence trainability and expressivity, and recent work has therefore explored a broad range of nonlinearities. However, widely used Gaussian i.i.d. initializations are designed to preserve activation variance under wide or infinite width assumptions. In deep and relatively narrow networks with sigmoidal nonlinearities, these schemes often drive preactivations into saturation, and collapse gradients. To address this, we introduce an odd-sigmoid activations and propose an activation aware initialization tailored to any function in this class. Our method remains robust over a wide band of variance scales, preserving both forward signal variance and backpropagated gradient norms even in very deep and narrow networks. Empirically, across standard image benchmarks we find that the proposed initialization is substantially less sensitive to depth, width, and activation scale than Gaussian initializations. In physics informed neural networks (PINNs), scaled odd-sigmoid activations combined with our initialization achieve lower losses than Gaussian based setups, suggesting that diagonal-plus-noise weights provide a practical alternative when Gaussian initialization breaks down.

cs.LG

Self-supervised New Activity Detection in Sensor-based Smart Environments

With the rapid advancement of ubiquitous computing technology, human activity analysis based on time series data from a diverse range of sensors enables the delivery of more intelligent services. Despite the importance of exploring new activities in real-world scenarios, existing human activity recognition studies generally rely on predefined known activities and often overlook detecting new patterns (novelties) that have not been previously observed during training. Novelty detection in human activities becomes even more challenging due to (1) diversity of patterns within the same known activity, (2) shared patterns between known and new activities, and (3) differences in sensor properties of each activity dataset. We introduce CLAN, a two-tower model that leverages Contrastive Learning with diverse data Augmentation for New activity detection in sensor-based environments. CLAN simultaneously and explicitly utilizes multiple types of strongly shifted data as negative samples in contrastive learning, effectively learning invariant representations that adapt to various pattern variations within the same activity. To enhance the ability to distinguish between known and new activities that share common features, CLAN incorporates both time and frequency domains, enabling the learning of multi-faceted discriminative representations. Additionally, we design an automatic selection mechanism of data augmentation methods tailored to each dataset's properties, generating appropriate positive and negative pairs for contrastive learning. Comprehensive experiments on real-world datasets show that CLAN achieves a 9.24% improvement in AUROC compared to the best-performing baseline model.

cs.LG

Robust Weight Initialization for Tanh Neural Networks with Fixed Point Analysis

As a neural network's depth increases, it can improve generalization performance. However, training deep networks is challenging due to gradient and signal propagation issues. To address these challenges, extensive theoretical research and various methods have been introduced. Despite these advances, effective weight initialization methods for tanh neural networks remain insufficiently investigated. This paper presents a novel weight initialization method for neural networks with tanh activation function. Based on an analysis of the fixed points of the function $\tanh(ax)$, the proposed method aims to determine values of $a$ that mitigate activation saturation. A series of experiments on various classification datasets and physics-informed neural networks demonstrates that the proposed method outperforms Xavier initialization methods~(with or without normalization) in terms of robustness across different network sizes, data efficiency, and convergence speed. Code is available at https://github.com/1HyunwooLee/Tanh-Init

cs.LG

DiffIM: Differentiable Influence Minimization with Surrogate Modeling and Continuous Relaxation

In social networks, people influence each other through social links, which can be represented as propagation among nodes in graphs. Influence minimization (IMIN) is the problem of manipulating the structures of an input graph (e.g., removing edges) to reduce the propagation among nodes. IMIN can represent time-critical real-world applications, such as rumor blocking, but IMIN is theoretically difficult and computationally expensive. Moreover, the discrete nature of IMIN hinders the usage of powerful machine learning techniques, which requires differentiable computation. In this work, we propose DiffIM, a novel method for IMIN with two differentiable schemes for acceleration: (1) surrogate modeling for efficient influence estimation, which avoids time-consuming simulations (e.g., Monte Carlo), and (2) the continuous relaxation of decisions, which avoids the evaluation of individual discrete decisions (e.g., removing an edge). We further propose a third accelerating scheme, gradient-driven selection, that chooses edges instantly based on gradients without optimization (spec., gradient descent iterations) on each test instance. Through extensive experiments on real-world graphs, we show that each proposed scheme significantly improves speed with little (or even no) IMIN performance degradation. Our method is Pareto-optimal (i.e., no baseline is faster and more effective than it) and typically several orders of magnitude (spec., up to 15,160X) faster than the most effective baseline while being more effective.

cs.LG

FlowerFormer: Empowering Neural Architecture Encoding using a Flow-aware Graph Transformer

The success of a specific neural network architecture is closely tied to the dataset and task it tackles; there is no one-size-fits-all solution. Thus, considerable efforts have been made to quickly and accurately estimate the performances of neural architectures, without full training or evaluation, for given tasks and datasets. Neural architecture encoding has played a crucial role in the estimation, and graphbased methods, which treat an architecture as a graph, have shown prominent performance. For enhanced representation learning of neural architectures, we introduce FlowerFormer, a powerful graph transformer that incorporates the information flows within a neural architecture. FlowerFormer consists of two key components: (a) bidirectional asynchronous message passing, inspired by the flows; (b) global attention built on flow-based masking. Our extensive experiments demonstrate the superiority of FlowerFormer over existing neural encoding methods, and its effectiveness extends beyond computer vision models to include graph neural networks and auto speech recognition models. Our code is available at http://github.com/y0ngjaenius/CVPR2024_FLOWERFormer.

cs.LG

DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition

With the advancement of IoT technology, recognizing user activities with machine learning methods is a promising way to provide various smart services to users. High-quality data with privacy protection is essential for deploying such services in the real world. Data streams from surrounding ambient sensors are well suited to the requirement. Existing ambient sensor datasets only support constrained private spaces and those for public spaces have yet to be explored despite growing interest in research on them. To meet this need, we build a dataset collected from a meeting room equipped with ambient sensors. The dataset, DOO-RE, includes data streams from various ambient sensor types such as Sound and Projector. Each sensor data stream is segmented into activity units and multiple annotators provide activity labels through a cross-validation annotation process to improve annotation quality. We finally obtain 9 types of activities. To our best knowledge, DOO-RE is the first dataset to support the recognition of both single and group activities in a real meeting room with reliable annotations.

cs.HC

A Causality-Aware Pattern Mining Scheme for Group Activity Recognition in a Pervasive Sensor Space

Human activity recognition (HAR) is a key challenge in pervasive computing and its solutions have been presented based on various disciplines. Specifically, for HAR in a smart space without privacy and accessibility issues, data streams generated by deployed pervasive sensors are leveraged. In this paper, we focus on a group activity by which a group of users perform a collaborative task without user identification and propose an efficient group activity recognition scheme which extracts causality patterns from pervasive sensor event sequences generated by a group of users to support as good recognition accuracy as the state-of-the-art graphical model. To filter out irrelevant noise events from a given data stream, a set of rules is leveraged to highlight causally related events. Then, a pattern-tree algorithm extracts frequent causal patterns by means of a growing tree structure. Based on the extracted patterns, a weighted sum-based pattern matching algorithm computes the likelihoods of stored group activities to the given test event sequence by means of matched event pattern counts for group activity recognition. We evaluate the proposed scheme using the data collected from our testbed and CASAS datasets where users perform their tasks on a daily basis and validate its effectiveness in a real environment. Experiment results show that the proposed scheme performs higher recognition accuracy and with a small amount of runtime overhead than the existing schemes.

cs.LG

Four-set Hypergraphlets for Characterization of Directed Hypergraphs

A directed hypergraph, which consists of nodes and hyperarcs, is a higher-order data structure that naturally models directional group interactions (e.g., chemical reactions of molecules). Although there have been extensive studies on local structures of (directed) graphs in the real world, those of directed hypergraphs remain unexplored. In this work, we focus on measurements, findings, and applications related to local structures of directed hypergraphs, and they together contribute to a systematic understanding of various real-world systems interconnected by directed group interactions. Our first contribution is to define 91 directed hypergraphlets (DHGs), which disjointly categorize directed connections and overlaps among four node sets that compose two incident hyperarcs. Our second contribution is to develop exact and approximate algorithms for counting the occurrences of each DHG. Our last contribution is to characterize 11 real-world directed hypergraphs and individual hyperarcs in them using the occurrences of DHGs, which reveals clear domain-based local structural patterns. Our experiments demonstrate that our DHG-based characterization gives up to 12% and 33% better performances on hypergraph clustering and hyperarc prediction, respectively, than baseline characterization methods. Moreover, we show that CODA-A, which is our proposed approximate algorithm, is up to 32X faster than its competitors with similar characterization quality.

cs.DS

Hypergraph Motifs and Their Extensions Beyond Binary

Hypergraphs naturally represent group interactions, which are omnipresent in many domains: collaborations of researchers, co-purchases of items, and joint interactions of proteins, to name a few. In this work, we propose tools for answering the following questions: (Q1) what are the structural design principles of real-world hypergraphs? (Q2) how can we compare local structures of hypergraphs of different sizes? (Q3) how can we identify domains from which hypergraphs are? We first define hypergraph motifs (h-motifs), which describe the overlapping patterns of three connected hyperedges. Then, we define the significance of each h-motif in a hypergraph as its occurrences relative to those in properly randomized hypergraphs. Lastly, we define the characteristic profile (CP) as the vector of the normalized significance of every h-motif. Regarding Q1, we find that h-motifs' occurrences in 11 real-world hypergraphs from 5 domains are clearly distinguished from those of randomized hypergraphs. Then, we demonstrate that CPs capture local structural patterns unique to each domain, and thus comparing CPs of hypergraphs addresses Q2 and Q3. The concept of CP is extended to represent the connectivity pattern of each node or hyperedge as a vector, which proves useful in node classification and hyperedge prediction. Our algorithmic contribution is to propose MoCHy, a family of parallel algorithms for counting h-motifs' occurrences in a hypergraph. We theoretically analyze their speed and accuracy and show empirically that the advanced approximate version MoCHy-A+ is more accurate and faster than the basic approximate and exact versions, respectively. Furthermore, we explore ternary hypergraph motifs that extends h-motifs by taking into account not only the presence but also the cardinality of intersections among hyperedges. This extension proves beneficial for all previously mentioned applications.

cs.SI

Characterization of Simplicial Complexes by Counting Simplets Beyond Four Nodes

Simplicial complexes are higher-order combinatorial structures which have been used to represent real-world complex systems. In this paper, we concentrate on the local patterns in simplicial complexes called simplets, a generalization of graphlets. We formulate the problem of counting simplets of a given size in a given simplicial complex. For this problem, we extend a sampling algorithm based on color coding from graphs to simplicial complexes, with essential technical novelty. We theoretically analyze our proposed algorithm named SC3, showing its correctness, unbiasedness, convergence, and time/space complexity. Through the extensive experiments on sixteen real-world datasets, we show the superiority of SC3 in terms of accuracy, speed, and scalability, compared to the baseline methods. Finally, we use the counts given by SC3 for simplicial complex analysis, especially for characterization, which is further used for simplicial complex clustering, where SC3 shows a strong ability of characterization with domain-based similarity.

cs.SI