SearcharxivSearch

arXiv subjects

Daehee Kim

Publications and source records attributed to Daehee Kim.

At least 19 recordsLinked to original sources

Microwave-Free $^{13}$C Hyperpolarization of Diamond Particles Enabled by Magic Angle Spinning and NV Centers

Nuclear hyperpolarization from optically pumped color centers in solids offers an alternative to conventional microwave-driven dynamic nuclear polarization (DNP). Diamond can host the nitrogen vacancy (NV) center, whose ground spin state can be readily polarized by light at room temperature, making diamond a candidate platform for nuclear hyperpolarization. We report $^{13}{\rm C}$ nuclear hyperpolarization in randomly oriented diamond particles with sizes ranging from 0.2 to 2 $\mu$m, both at natural $^{13}{\rm C}$ abundance (1.1 %) and at 20 % isotopic enrichment, at magnetic fields of 7.1 T and 9.4 T. The protocol combines optical illumination with magic angle spinning (MAS) and does not require microwave irradiation. By investigating the nuclear polarization as a function of the MAS frequency between 0 and 6 kHz at the magnetic field of 7.1 T, we find maximum light-induced polarization enhancements of $280$-fold for the isotopically enriched sample and $411$-fold for the natural abundance sample. Under continuous illumination, steady-state absolute $^{13}{\rm C}$ polarization levels above 0.1 % are reached. A model involving optical pumping of NV centers and spin dynamics near level anticrossings (LACs) in three-spin clusters formed by NV, a substitutional nitrogen (P1) and $^{13}{\rm C}$ is used to describe these findings. The protocol strongly mitigates the effect of the anisotropy of the NV spin Hamiltonian, allowing more than $99.9\%$ of NV orientations to participate in the polarization transfer process. These results represent a first step toward transferring nuclear polarization from diamond particles to external nuclei, with potential applications in sensitive and high-resolution NMR at room temperature.

quant-ph

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control loop, with its entire action space declared solely in the prompt. Recent systems approach this setting but increasingly narrow the model's decision-making. We widen it back. We introduce DroneCATS-Agent, an architecture where the MLLM is a swappable component, and DroneCATS, a benchmark treating the model as the independent variable. Beyond merely flying toward a pixel, our agent entrusts the model to yaw and search, deliberate when unsure, and self-declare arrival---all without fine-tuning or function-calling schemas. Evaluating frontier and open models across four core capabilities---approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet---reveals that even the simplest embodied settings are far from solved. Crucially, to identify what breaks first at the edge, our roster scales down to 2B parameters. The findings expose a stark paradox: it is not the flying that fails. Small open models often navigate into the success radius more reliably than frontier models, yet lose the episode by declaring arrival prematurely or not at all. Multi-drone commanding amplifies this divide, with small models failing by blindly copying a single coordinate across distinct views. Viewed as vision-language-action agents, the models' spatial perception holds up, but their action protocol does not. What separates a deployable edge model from a frontier model is not navigation, but the discipline to sustain a declared protocol and emit the correct terminating action. The open problem is closing this gap at onboard compute costs---yielding a fast model that plans persistently and knows exactly when it is done---and DroneCATS is built to measure that distance.

cs.RO

Spin-force from a Nitrogen-Vacancy ensemble drives a 100 mg levitated resonator

The force experienced by a spin in a magnetic field gradient underlies many proposals for hybrid quantum systems. These include schemes for mechanically mediated quantum gates, spin squeezing, searches for exotic forces, and motional superpositions for probing the interface between quantum and gravity. Yet, experimentally observing this spin-force for anything larger than atomic scales has proved challenging. In our work, we demonstrate controllable Center-of-Mass motion of a $128 \rm\: mg$ diamagnetically levitated oscillator due to force from an ensemble of Nitrogen-Vacancy (NV) defects in diamond. We induce coherent motion in the oscillator by periodic optical initialisation of the NV spin states, achieving motional amplitudes exceeding $100 \rm\:nm$. Our results mark a key milestone towards spin-based engineering of motional states deep in the high-mass regime.

quant-ph

Efficient Knowledge Tracing Leveraging Higher-Order Information in Integrated Graphs

The rise of online learning has led to the development of various knowledge tracing (KT) methods. However, existing methods have overlooked the problem of increasing computational cost when utilizing large graphs and long learning sequences. To address this issue, we introduce Dual Graph Attention-based Knowledge Tracing (DGAKT), a graph neural network model designed to leverage high-order information from subgraphs representing student-exercise-KC relationships. DGAKT incorporates a subgraph-based approach to enhance computational efficiency. By processing only relevant subgraphs for each target interaction, DGAKT significantly reduces memory and computational requirements compared to full global graph models. Extensive experimental results demonstrate that DGAKT not only outperforms existing KT models but also sets a new standard in resource efficiency, addressing a critical need that has been largely overlooked by prior KT approaches.

cs.LG

GuRE:Generative Query REwriter for Legal Passage Retrieval

Legal Passage Retrieval (LPR) systems are crucial as they help practitioners save time when drafting legal arguments. However, it remains an underexplored avenue. One primary reason is the significant vocabulary mismatch between the query and the target passage. To address this, we propose a simple yet effective method, the Generative query REwriter (GuRE). We leverage the generative capabilities of Large Language Models (LLMs) by training the LLM for query rewriting. "Rewritten queries" help retrievers to retrieve target passages by mitigating vocabulary mismatch. Experimental results show that GuRE significantly improves performance in a retriever-agnostic manner, outperforming all baseline methods. Further analysis reveals that different training objectives lead to distinct retrieval behaviors, making GuRE more suitable than direct retriever fine-tuning for real-world applications. Codes are avaiable at github.com/daehuikim/GuRE.

cs.CL

A magnetically levitated conducting rotor with ultra-low rotational damping circumventing eddy loss

Levitation of macroscopic objects in a vacuum is key towards the development of high-precision inertial sensors and pressure sensors, as well as towards the fundamental studies of quantum mechanics and its relation to gravity. Diamagnetic levitation offers a passive method at room temperature to isolate macroscopic objects in vacuum environments, yet eddy current damping remains a critical limitation for electrically conductive materials. We show that there are situations where the motion of conductors in magnetic fields does not, in principle, produce eddy damping, and demonstrate an electrically conducting rotor diamagnetically levitated in an axially symmetric magnetic field in a high vacuum. Experimental measurements and finite-element simulations reveal gas collision damping as the dominant loss mechanism at high pressures, while residual eddy damping, which arises from symmetry-breaking factors such as platform tilt or material imperfections, dominates at low pressures. The conclusion is supported by an analytic proof and an analytic example of zero steady current density for a rotating conductor in a magnetic field. This demonstrates a macroscopic levitated rotor with extremely low rotational damping and paves the way to fully suppress rotor damping, enabling ultra-low-loss rotors for gyroscopes, pressure sensing, and fundamental physics tests.

quant-ph

Closer through commonality: Enhancing hypergraph contrastive learning with shared groups

Hypergraphs provide a superior modeling framework for representing complex multidimensional relationships in the context of real-world interactions that often occur in groups, overcoming the limitations of traditional homogeneous graphs. However, there have been few studies on hypergraphbased contrastive learning, and existing graph-based contrastive learning methods have not been able to fully exploit the highorder correlation information in hypergraphs. Here, we propose a Hypergraph Fine-grained contrastive learning (HyFi) method designed to exploit the complex high-dimensional information inherent in hypergraphs. While avoiding traditional graph augmentation methods that corrupt the hypergraph topology, the proposed method provides a simple and efficient learning augmentation function by adding noise to node features. Furthermore, we expands beyond the traditional dichotomous relationship between positive and negative samples in contrastive learning by introducing a new relationship of weak positives. It demonstrates the importance of fine-graining positive samples in contrastive learning. Therefore, HyFi is able to produce highquality embeddings, and outperforms both supervised and unsupervised baselines in average rank on node classification across 10 datasets. Our approach effectively exploits high-dimensional hypergraph information, shows significant improvement over existing graph-based contrastive learning methods, and is efficient in terms of training speed and GPU memory cost. The source code is available at https://github.com/Noverse0/HyFi.git.

cs.LG

Exploring Iterative Controllable Summarization with Large Language Models

Large language models (LLMs) have demonstrated remarkable performance in abstractive summarization tasks. However, their ability to precisely control summary attributes (e.g., length or topic) remains underexplored, limiting their adaptability to specific user preferences. In this paper, we systematically explore the controllability of LLMs. To this end, we revisit summary attribute measurements and introduce iterative evaluation metrics, failure rate and average iteration count to precisely evaluate controllability of LLMs, rather than merely assessing errors. Our findings show that LLMs struggle more with numerical attributes than with linguistic attributes. To address this challenge, we propose a guide-to-explain framework (GTE) for controllable summarization. Our GTE framework enables the model to identify misaligned attributes in the initial draft and guides it in self-explaining errors in the previous output. By allowing the model to reflect on its misalignment, GTE generates well-adjusted summaries that satisfy the desired attributes with robust effectiveness, requiring surprisingly fewer iterations than other iterative approaches.

cs.CL

Ontology-Free General-Domain Knowledge Graph-to-Text Generation Dataset Synthesis using Large Language Model

Knowledge Graph-to-Text (G2T) generation involves verbalizing structured knowledge graphs into natural language text. Recent advancements in Pretrained Language Models (PLMs) have improved G2T performance, but their effectiveness depends on datasets with precise graph-text alignment. However, the scarcity of high-quality, general-domain G2T generation datasets restricts progress in the general-domain G2T generation research. To address this issue, we introduce Wikipedia Ontology-Free Graph-text dataset (WikiOFGraph), a new large-scale G2T dataset generated using a novel method that leverages Large Language Model (LLM) and Data-QuestEval. Our new dataset, which contains 5.85M general-domain graph-text pairs, offers high graph-text consistency without relying on external ontologies. Experimental results demonstrate that PLM fine-tuned on WikiOFGraph outperforms those trained on other datasets across various evaluation metrics. Our method proves to be a scalable and effective solution for generating high-quality G2T data, significantly advancing the field of G2T generation.

cs.CL

Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models

Recent advances in Large Language Models (LLMs) have stimulated a surge of research aimed at extending their applications to the visual domain. While these models exhibit promise in generating abstract image captions and facilitating natural conversations, their performance on text-rich images still requires improvement. In this paper, we introduce Contrastive Reading Model (Cream), a novel neural architecture designed to enhance the language-image understanding capability of LLMs by capturing intricate details that are often overlooked in existing methods. Cream combines vision and auxiliary encoders, fortified by a contrastive feature alignment technique, to achieve a more effective comprehension of language information in visually situated contexts within the images. Our approach bridges the gap between vision and language understanding, paving the way for the development of more sophisticated Document Intelligence Assistants. Through rigorous evaluations across diverse visually-situated language understanding tasks that demand reasoning capabilities, we demonstrate the compelling performance of Cream, positioning it as a prominent model in the field of visual document understanding. We provide our codebase and newly-generated datasets at https://github.com/naver-ai/cream .

cs.CL

SCOB: Universal Text Understanding via Character-wise Supervised Contrastive Learning with Online Text Rendering for Bridging Domain Gap

Inspired by the great success of language model (LM)-based pre-training, recent studies in visual document understanding have explored LM-based pre-training methods for modeling text within document images. Among them, pre-training that reads all text from an image has shown promise, but often exhibits instability and even fails when applied to broader domains, such as those involving both visual documents and scene text images. This is a substantial limitation for real-world scenarios, where the processing of text image inputs in diverse domains is essential. In this paper, we investigate effective pre-training tasks in the broader domains and also propose a novel pre-training method called SCOB that leverages character-wise supervised contrastive learning with online text rendering to effectively pre-train document and scene text domains by bridging the domain gap. Moreover, SCOB enables weakly supervised learning, significantly reducing annotation costs. Extensive benchmarks demonstrate that SCOB generally improves vanilla pre-training methods and achieves comparable performance to state-of-the-art methods. Our findings suggest that SCOB can be served generally and effectively for read-type pre-training methods. The code will be available at https://github.com/naver-ai/scob.

cs.CV

Fully Decentralized Peer-to-Peer Community Grid with Dynamic and Congestion Pricing

Peer-to-peer (P2P) electricity markets enable prosumers to minimize their costs, which has been extensively studied in recent research. However, there are several challenges with P2P trading when physical network constraints are also included. Moreover, most studies use fixed prices for grid power prices without considering dynamic grid pricing, and equity for all participants. This policy may negatively affect the long-term development of the market if prosumers with low demand are not treated fairly. An initial step towards addressing these problems is the design of a new decentralized P2P electricity market with two dynamic grid pricing schemes that are determined by consumer demand. Futhermore, we consider a decentralized system with physical constraints for optimizing power flow in networks without compromising privacy. We propose a dynamic congestion price to effectively address congestion and then prove the convergence and global optimality of the proposed method. Our experiments show that P2P energy trade decreases generation cost of main grid by 56.9% compared with previous works. Consumers reduce grid trading by 57.3% while the social welfare of consumers is barely affected by the increase of grid price.

eess.SY

Towards Unified Scene Text Spotting based on Sequence Generation

Sequence generation models have recently made significant progress in unifying various vision tasks. Although some auto-regressive models have demonstrated promising results in end-to-end text spotting, they use specific detection formats while ignoring various text shapes and are limited in the maximum number of text instances that can be detected. To overcome these limitations, we propose a UNIfied scene Text Spotter, called UNITS. Our model unifies various detection formats, including quadrilaterals and polygons, allowing it to detect text in arbitrary shapes. Additionally, we apply starting-point prompting to enable the model to extract texts from an arbitrary starting point, thereby extracting more texts beyond the number of instances it was trained on. Experimental results demonstrate that our method achieves competitive performance compared to state-of-the-art methods. Further analysis shows that UNITS can extract a larger number of texts than it was trained on. We provide the code for our method at https://github.com/clovaai/units.

cs.CV

Supervised Contrastive ResNet and Transfer Learning for the In-vehicle Intrusion Detection System

High-end vehicles have been furnished with a number of electronic control units (ECUs), which provide upgrading functions to enhance the driving experience. The controller area network (CAN) is a well-known protocol that connects these ECUs because of its modesty and efficiency. However, the CAN bus is vulnerable to various types of attacks. Although the intrusion detection system (IDS) is proposed to address the security problem of the CAN bus, most previous studies only provide alerts when attacks occur without knowing the specific type of attack. Moreover, an IDS is designed for a specific car model due to diverse car manufacturers. In this study, we proposed a novel deep learning model called supervised contrastive (SupCon) ResNet, which can handle multiple attack identification on the CAN bus. Furthermore, the model can be used to improve the performance of a limited-size dataset using a transfer learning technique. The capability of the proposed model is evaluated on two real car datasets. When tested with the car hacking dataset, the experiment results show that the SupCon ResNet model improves the overall false-negative rates of four types of attack by four times on average, compared to other models. In addition, the model achieves the highest F1 score at 0.9994 on the survival dataset by utilizing transfer learning. Finally, the model can adapt to hardware constraints in terms of memory size and running time.

cs.CR

Detecting In-vehicle Intrusion via Semi-supervised Learning-based Convolutional Adversarial Autoencoders

With the development of autonomous vehicle technology, the controller area network (CAN) bus has become the de facto standard for an in-vehicle communication system because of its simplicity and efficiency. However, without any encryption and authentication mechanisms, the in-vehicle network using the CAN protocol is susceptible to a wide range of attacks. Many studies, which are mostly based on machine learning, have proposed installing an intrusion detection system (IDS) for anomaly detection in the CAN bus system. Although machine learning methods have many advantages for IDS, previous models usually require a large amount of labeled data, which results in high time and labor costs. To handle this problem, we propose a novel semi-supervised learning-based convolutional adversarial autoencoder model in this paper. The proposed model combines two popular deep learning models: autoencoder and generative adversarial networks. First, the model is trained with unlabeled data to learn the manifolds of normal and attack patterns. Then, only a small number of labeled samples are used in supervised training. The proposed model can detect various kinds of message injection attacks, such as DoS, fuzzy, and spoofing, as well as unknown attacks. The experimental results show that the proposed model achieves the highest F1 score of 0.99 and a low error rate of 0.1\% with limited labeled data compared to other supervised methods. In addition, we show that the model can meet the real-time requirement by analyzing the model complexity in terms of the number of trainable parameters and inference time. This study successfully reduced the number of model parameters by five times and the inference time by eight times, compared to a state-of-the-art model.

cs.CR

MILP-based optimal day-ahead scheduling for system-centric CEMS supporting different types of homes and energy trading

Optimal day-ahead scheduling for a system-centric community energy management system (CEMS) is proposed to provide economic benefits and user comfort of energy management at the community level. Our proposed community includes different types of homes and allows prosumers to trade energy locally using mid-market rate pricing. A mathematical model of the community is constructed and the optimization problem of this model is transformed into an MILP problem that can be solved in a short time. By solving this MILP problem, the optimization of the overall energy cost of the community and satisfaction of the thermal comfort at every home are achieved. For comparison, we also establish two different scenarios for the same community: a prosumer-centric CEMS and no CEMS. The simulation results demonstrate that the overall energy cost of the community with the system-centric CEMS is the smallest among the three scenarios and is only half that of the community with the prosumer-centric CEMS. Moreover, by using linear transformation, the computational time of the optimization problem of the proposed system-centric CEMS is only 118.2 s for a 500-home community, which is a short time for day-ahead scheduling of a community.

eess.SY

A Supervised-Learning based Hour-Ahead Demand Response of a Behavior-based HEMS approximating MILP Optimization

The demand response (DR) program of a traditional HEMS usually intervenes appliances by controlling or scheduling them to achieve multiple objectives such as minimizing energy cost and maximizing user comfort. In this study, instead of intervening appliances and changing resident behavior, our proposed strategy for hour-ahead DR firstly learns appliance use behavior of residents and then silently controls ESS and RES to minimize daily energy cost based on its knowledge. To accomplish the goal, our proposed deep neural networks (DNNs) models approximate MILP optimization by using supervised learning. The datasets for training DNNs are created from optimal outputs of a MILP solver with historical data. After training, at each time slot, these DNNs are used to control ESS and RES with real-time data of the surrounding environment. For comparison, we develop two different strategies named multi-agent reinforcement learning-based strategy, a kind of hour-ahead strategy and forecast-based MILP strategy, a kind of day-ahead strategy. For evaluation and verification, our proposed strategies are applied at three different real-world homes with real-world real-time global horizontal irradiation and real-world real-time prices. Numerical results verify that the proposed MILP-based supervised learning strategy is effective in term of daily energy cost and is the best one among three proposed strategies

eess.SY

SelfReg: Self-supervised Contrastive Regularization for Domain Generalization

In general, an experimental environment for deep learning assumes that the training and the test dataset are sampled from the same distribution. However, in real-world situations, a difference in the distribution between two datasets, domain shift, may occur, which becomes a major factor impeding the generalization performance of the model. The research field to solve this problem is called domain generalization, and it alleviates the domain shift problem by extracting domain-invariant features explicitly or implicitly. In recent studies, contrastive learning-based domain generalization approaches have been proposed and achieved high performance. These approaches require sampling of the negative data pair. However, the performance of contrastive learning fundamentally depends on quality and quantity of negative data pairs. To address this issue, we propose a new regularization method for domain generalization based on contrastive learning, self-supervised contrastive regularization (SelfReg). The proposed approach use only positive data pairs, thus it resolves various problems caused by negative pair sampling. Moreover, we propose a class-specific domain perturbation layer (CDPL), which makes it possible to effectively apply mixup augmentation even when only positive data pairs are used. The experimental results show that the techniques incorporated by SelfReg contributed to the performance in a compatible manner. In the recent benchmark, DomainBed, the proposed method shows comparable performance to the conventional state-of-the-art alternatives. Codes are available at https://github.com/dnap512/SelfReg.

cs.CV