SearcharxivSearch

arXiv subjects

Daeun Kim

Publications and source records attributed to Daeun Kim.

14 recordsLinked to original sources

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Data augmentation is fundamental to training modern deep vision and multimodal models. While individual methods, such as RandAug, CutMix, Mixup, RandErase, and DropPath, offer strong regularization effects, their combined use has saturated in performance due to overlapping functionalities, and aggressive pixel-level manipulations may disrupt delicate cross-modal alignment. This saturation motivates the search for a new augmentation axis within the embedding space rather than the input space. We introduce Aether, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent. Inspired by feature-space perturbations in language models and image degradation in generative pretraining, Aether induces mild yet effective perturbations that smooth the representations without compromising the fine-grained structural information required for strong vision-language encoders. Across diverse architectures and across multiple recognition tasks, Aether delivers consistent gains over the advanced recipe combining CutMix, Mixup, DropPath, and RandAug---a level of improvement rarely observed with modern augmentation alternatives. Notably, Aether demonstrates superior effectiveness in multi-modal alignment, succeeding where traditional pixel-space augmentations fail by providing a stable, isotropic regularization signal that respects the integrity of the high-dimensional feature space.

cs.CV

Comparison of Image Processing Models in Quark Gluon Jet Classification

We present a comprehensive comparison of convolutional and transformer-based models for distinguishing quark and gluon jets using simulated jet images from Pythia 8. By encoding jet substructure into a three-channel representation of particle kinematics, we evaluate the performance of convolutional neural networks (CNNs), Vision Transformers (ViTs), and Swin Transformers (Swin-Tiny) under both supervised and self-supervised learning setups. Our results show that fine-tuning only the final two transformer blocks of the Swin-Tiny model achieves the best trade-off between efficiency and accuracy, reaching 81.4% accuracy and an AUC (area under the ROC curve) of 88.9%. Self-supervised pretraining with Momentum Contrast (MoCo) further enhances feature robustness and reduces the number of trainable parameters. These findings highlight the potential of hierarchical attention-based models for jet substructure studies and for domain transfer to real collision data.

physics.data-an

D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse

Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces D\'ej\`a Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, D\'ej\`a Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that D\'ej\`a Vu accelerates embedding generation by up to a 2.64x within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.

cs.DC

MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization

Diffusion Transformer (DiT) has driven significant progress in image generation tasks. However, DiT inferencing is notoriously compute-intensive and incurs long latency even on datacenter-scale GPUs, primarily due to its iterative nature and heavy reliance on GEMM operations inherent to its encoder-based structure. To address the challenge, prior work has explored quantization, but achieving low-precision quantization for DiT inferencing with both high accuracy and substantial speedup remains an open problem. To this end, this paper proposes MixDiT, an algorithm-hardware co-designed acceleration solution that exploits mixed Microscaling (MX) formats to quantize DiT activation values. MixDiT quantizes the DiT activation tensors by selectively applying higher precision to magnitude-based outliers, which produce mixed-precision GEMM operations. To achieve tangible speedup from the mixed-precision arithmetic, we design a MixDiT accelerator that enables precision-flexible multiplications and efficient MX precision conversions. Our experimental results show that MixDiT delivers a speedup of 2.10-5.32 times over RTX 3090, with no loss in FID.

cs.AR

Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models

In this work, we dive deep into the impact of additive noise in pre-training deep networks. While various methods have attempted to use additive noise inspired by the success of latent denoising diffusion models, when used in combination with masked image modeling, their gains have been marginal when it comes to recognition tasks. We thus investigate why this would be the case, in an attempt to find effective ways to combine the two ideas. Specifically, we find three critical conditions: corruption and restoration must be applied within the encoder, noise must be introduced in the feature space, and an explicit disentanglement between noised and masked tokens is necessary. By implementing these findings, we demonstrate improved pre-training performance for a wide range of recognition tasks, including those that require fine-grained, high-frequency information to solve.

cs.CV

Spectrum Sharing Between Low Earth Orbit Satellite and Terrestrial Networks: A Stochastic Geometry Perspective Analysis

Low Earth orbit (LEO) satellite networks with mega constellations have the potential to provide 5G and beyond services ubiquitously. However, these networks may introduce mutual interference to both satellite and terrestrial networks, particularly when sharing spectrum resources. In this paper, we present a system-level performance analysis to address these interference issues using the tool of stochastic geometry. We model the spatial distributions of satellites, satellite users, terrestrial base stations (BSs), and terrestrial users using independent Poisson point processes on the surfaces of concentric spheres. Under these spatial models, we derive analytical expressions for the ergodic spectral efficiency of uplink (UL) and downlink (DL) satellite networks when they share spectrum with both UL and DL terrestrial networks. These derived ergodic expressions capture comprehensive network parameters, including the densities of satellite and terrestrial networks, the path-loss exponent, and fading. From our analysis, we determine the conditions under which spectrum sharing with UL terrestrial networks is advantageous for both UL and DL satellite networks. Our key finding is that the optimal spectrum sharing configuration among the four possible configurations depends on the density ratio between terrestrial BSs and users, providing a design guideline for spectrum management. Simulation results confirm the accuracy of our derived expressions.

eess.SP

Ergodic Secrecy Rate Analysis for LEO Satellite Downlink Networks

Satellite networks are recognized as an effective solution to ensure seamless connectivity worldwide, catering to a diverse range of applications. However, the broad coverage and broadcasting nature of satellite networks also expose them to security challenges. Despite these challenges, there is a lack of analytical understanding addressing the secrecy performance of these networks. This paper presents a secrecy rate analysis for downlink low Earth orbit (LEO) satellite networks by modeling the spatial distribution of satellites, users, and potential eavesdroppers as homogeneous Poisson point processes on concentric spheres. Specifically, we provide an analytical expression for the ergodic secrecy rate of a typical downlink user in terms of the satellite network parameters, fading parameters, and path-loss exponent. Simulation results show the exactness of the provided expressions and we find that optimal satellite altitude increases with eavesdropper density.

eess.SP

Coverage Analysis of Dynamic Coordinated Beamforming for LEO Satellite Downlink Networks

In this paper, we investigate the coverage performance of downlink satellite networks employing dynamic coordinated beamforming. Our approach involves modeling the spatial arrangement of satellites and users using Poisson point processes situated on concentric spheres. We derive analytical expressions for the coverage probability, which take into account the in-cluster geometry of the coordinated satellite set. These expressions are formulated in terms of various parameters, including the number of antennas per satellite, satellite density, fading characteristics, and path-loss exponent. To offer a more intuitive understanding, we also develop an approximation for the coverage probability. Furthermore, by considering the distribution of normalized distances, we derive the spatially averaged coverage probability, thereby validating the advantages of coordinated beamforming from a spatial average perspective. Our primary finding is that dynamic coordinated beamforming significantly improves coverage compared to the absence of satellite coordination, in direct proportion to the number of antennas on each satellite. Moreover, we observe that the optimal cluster size, which maximizes the ergodic spectral efficiency, increases with higher satellite density, provided that the number of antennas on the satellites is sufficiently large. Our findings are corroborated by simulation results, confirming the accuracy of the derived expressions.

eess.SP

BanditLinQ: A Scalable Link Scheduling for Dense D2D Networks with One-Bit Feedback

This paper addresses cooperative link scheduling problems for base station (BS) aided device-to-device (D2D) communications using limited channel state information (CSI) at BS. We first derive the analytical form of ergodic sum-spectral efficiency as a function of network parameters, assuming statistical CSI at the BS. However, the optimal link scheduling, which maximizes the ergodic sum-spectral efficiency, becomes computationally infeasible when network density increases. To overcome this challenge, we present a low-complexity link scheduling algorithm that divides the D2D network into sub-networks and identifies the optimal link scheduling strategy per sub-network. Furthermore, we consider the scenario when the statistical CSI is not available to the BS. In such cases, we propose a quasi-optimal scalable link scheduling algorithm that utilizes one-bit feedback information from D2D receivers. The algorithm clusters the links and applies the UCB algorithm per cluster using the collected one-bit feedback information. We highlight that even with reduced scheduling complexity, the proposed algorithm identifies a link scheduling action that ensures optimality within a constant throughput gap. We also demonstrate through simulations that the proposed algorithm achieves higher sum-spectral efficiency than the existing link scheduling algorithms, even without explicit CSI or network parameters knowledge.

eess.SP

CoVA: Exploiting Compressed-Domain Analysis to Accelerate Video Analytics

Modern retrospective analytics systems leverage cascade architecture to mitigate bottleneck for computing deep neural networks (DNNs). However, the existing cascades suffer two limitations: (1) decoding bottleneck is either neglected or circumvented, paying significant compute and storage cost for pre-processing; and (2) the systems are specialized for temporal queries and lack spatial query support. This paper presents CoVA, a novel cascade architecture that splits the cascade computation between compressed domain and pixel domain to address the decoding bottleneck, supporting both temporal and spatial queries. CoVA cascades analysis into three major stages where the first two stages are performed in compressed domain while the last one in pixel domain. First, CoVA detects occurrences of moving objects (called blobs) over a set of compressed frames (called tracks). Then, using the track results, CoVA prudently selects a minimal set of frames to obtain the label information and only decode them to compute the full DNNs, alleviating the decoding bottleneck. Lastly, CoVA associates tracks with labels to produce the final analysis results on which users can process both temporal and spatial queries. Our experiments demonstrate that CoVA offers 4.8x throughput improvement over modern cascade systems, while imposing modest accuracy loss.

cs.CV

Groove-Assisted Global Spontaneous Alignment of Carbon Nanotubes in Vacuum Filtration

Ever since the discovery of carbon nanotubes (CNTs), it has long been a challenging goal to create macroscopically ordered assemblies, or crystals, of CNTs that preserve the one-dimensional quantum properties of individual CNTs on a macroscopic scale. Recently, a simple and well-controlled method was reported for producing wafer-scale crystalline films of highly aligned and densely packed CNTs through spontaneous global alignment that occurs during vacuum filtration [\textit{Nat.\ Nanotechnol}.\ \textbf{11}, 633 (2016)]. However, a full understanding of the mechanism of such global alignment has not been achieved. Here, we report results of a series of systematic experiments that demonstrate that the CNT alignment direction can be controlled by the surface morphology of the filter membrane used in the vacuum filtration process. More specifically, we found that the direction of parallel grooves pre-existing on the surface of the filter membrane dictates the direction of the resulting CNT alignment. Furthermore, we intentionally imprinted periodically spaced parallel grooves on a filter membranes using a diffraction grating, which successfully defined the direction of the global alignment of CNTs in a precise and reproducible manner.

physics.app-ph

Multidimensional Correlation Spectroscopic Imaging of Exponential Decays: From Theoretical Principles to In Vivo Human Applications

Multiexponential modeling of relaxation or diffusion MR signal decays is a popular approach for estimating and spatially mapping different microstructural tissue compartments. While this approach can be quite powerful, it is also limited by the fact that one-dimensional multiexponential modeling is an ill-posed inverse problem with substantial ambiguities. In this paper, we present an overview of a recent multidimensional correlation spectroscopic imaging approach to this problem. This approach helps to alleviate ill-posedness by leveraging multidimensional contrast encoding (e.g., 2D diffusion-relaxation encoding or 2D relaxation-relaxation encoding) combined with a regularized spatial-spectral estimation procedure. Theoretical calculations, simulations, and experimental results are used to illustrate the benefits of this approach relative to classical methods. In addition, we demonstrate an initial proof-of-principle application of this kind of approach to in vivo human MRI experiments.

eess.IV

Supervised-Learning for Multi-Hop MU-MIMO Communications with One-Bit Transceivers

This paper considers a nonlinear multi-hop multi-user multiple-input multiple-output (MU-MIMO) relay channel, in which multiple users send information symbols to a multi-antenna base station (BS) with one-bit analog-to-digital converters via intermediate relays, each with one-bit transceiver. To understand the fundamental limit of the detection performance, the optimal maximum-likelihood (ML) detector is proposed with the assumption of perfect and global channel state information (CSI) at the BS. This multi-user detector, however, is not practical due to the unrealistic CSI assumption and the overwhelming detection complexity. These limitations are addressed by presenting a novel detection framework inspired by supervised-learning. The key idea is to model the complicated multihop MU-MIMO channel as a simplified channel with much fewer and learnable parameters. One major finding is that, even using the simplified channel model, a near ML detection performance is achievable with a reasonable amount of pilot overheads in a certain condition. In addition, an online supervised-learning detector is proposed, which adaptively tracks channel variations. The idea is to update the model parameters with a reliably detected data symbol by treating it as a new training (labelled) data. Lastly, a multi-user detector using a deep neural network is proposed. Unlike the model-based approaches, this model-free approach enables to remove the errors in the simplified channel model, while increasing the computational complexity for parameter learning. Via simulations, the detection performances of classical, model-based, and model-free detectors are thoroughly compared to demonstrate the effectiveness of the supervised-learning approaches in this channel.

cs.IT

OEDIPUS: An Experiment Design Framework for Sparsity-Constrained MRI

This paper introduces a new estimation-theoretic framework for experiment design in the context of MR image reconstruction under sparsity constraints. The new framework is called OEDIPUS (Oracle-based Experiment Design for Imaging Parsimoniously Under Sparsity constraints), and is based on combining the constrained Cramér-Rao bound with classical experiment design techniques. Compared to popular random sampling approaches, OEDIPUS is fully deterministic and automatically tailors the sampling pattern to the specific imaging context of interest (i.e., accounting for coil geometry, anatomy, image contrast, etc.). OEDIPUS-based experiment designs are evaluated using retrospectively subsampled in vivo MRI data in several different contexts. Results demonstrate that OEDIPUS-based experiment designs have some desirable characteristics relative to conventional MRI sampling approaches.

eess.SP