SearcharxivSearch

arXiv subjects

Hojin Lee

Publications and source records attributed to Hojin Lee.

At least 19 recordsLinked to original sources

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and research, yet conventional tabular metrics often overlook temporal structure. Existing single-table and relational evaluation protocols largely collapse records into static distributions, leaving key temporal properties insufficiently evaluated. We introduce Seq2Synth, a unified benchmark for assessing these properties. Its taxonomy characterizes temporal and schema properties to determine applicable evaluations, covering timestamp, cross-sectional, longitudinal, and structural fidelity, alongside trajectory-aware utility and privacy. Across seven core datasets from a 13-dataset benchmark and eight generators, models with near-perfect static fidelity still violate basic temporal constraints, producing duplicate timestamps, irregular intervals, and incomplete observation grids. Moreover, static and temporal-aware rankings diverge substantially, showing that temporal fidelity must be evaluated directly rather than inferred from static or relational scores. Project page and online appendices are available at: https://seq2synth.github.io/.

cs.LG

Gravitational Metric of a Star

Solving the classical equations of motion in general relativity recursively, we consider the metric of a spatially localized and stationary source of matter. Having in mind a star of general composition, we characterize it by means of its infinite set of mass and current multipoles. Specializing to de Donder gauge we set up the recursive equations that produce the metric outside the star to any desired order in perturbation theory, expanded both in Newton's constant and in the order of multipoles. Up to second post-Minkowskian order we express the result to any order in the multipole expansion in terms of generalized (tensor) bubble integrals in momentum space and a corresponding simple expansion in inverse distances. In a special corner of the space of multipoles we recover the Kerr black hole solution to the given order. By tweaking just slightly the multipoles away from the Kerr limit the metric will describe stars that are Kerr-like and yet are not black holes. A subtlety with respect to the gauge ambiguity of de Donder gauge is also pointed out.

hep-th

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---particularly the learning rate---at extreme scales of both model size and token budget via sweeping remains computationally prohibitive. In this paper, we propose a compute-efficient, two-step hyperparameter transfer framework that estimates optimal learning rates for training large MoE models by transferring them across scaling model widths, and subsequently extrapolating to trillion-token horizons. First, we formulate a Maximal Update Parameterization ($μ$P) adaptation for MoE architectures utilizing Multi-head Latent Attention (MLA) and the Muon optimizer, demonstrating that optimal learning rates transfer consistently across width-scaled models. Second, we extend this transferability along the token dimension by establishing a predictive scaling law. By applying linear regression to the optimal values derived from small proxy models on limited budgets, we successfully extrapolate the ideal learning rate to massive training horizons (e.g., 10 trillion tokens) with high fidelity ($R^2=0.95$). Consequently, this indicates that proxy training on small models is sufficient to determine the optimal learning rate for the extensive training of large-scale MoEs. We apply the proposed methodology to pretrain our foundation model (155B total, 17B active parameters) from scratch, and the stable training and evaluation results validate that optimal configurations for full-scale target models can be accurately predicted with minimal ablation costs.

cs.LG

Iterative Solution of the Kerr Black Hole Metric

Using a recursive solution of the Einstein equations, we consider the perturbative expansion of the metric corresponding to a Kerr black hole. Because the metric is a function of two parameters, Newton's constant G and the Kerr spin parameter a, the perturbation theory naturally becomes a double expansion. In harmonic gauge the recursion relations can be solved to arbitrarily high orders in these two expansion parameters but to re-sum the series into the closed-form harmonic gauge metric requires the introduction of terms that are redundant and correspond to the addition of harmonic functions to the coordinates. Issues related to dimensional regularization of Fourier transforms are explained in detail.

hep-th

UrbanFlow-3K: A Dataset of 3,000 Lattice-Boltzmann Simulations of Random Building Layouts

The analysis of flow around buildings has gained significant research interest across various domains, including pedestrian safety, pollutant dispersion, natural ventilation, and building energy efficiency. While these domains frequently include high-resolution computational fluid dynamics (CFD) data, predicting urban flow fields with machine learning (ML) models has emerged as a promising approach to overcome the prohibitive costs of CFD simulations. However, the availability of open-source datasets for training such ML models remains scarce. In particular, publicly available two-dimensional datasets of urban flow fields are nearly non-existent, despite their potential value for early development and debugging stages of data-driven models, before scaling to computationally expensive three-dimensional datasets. To bridge this gap, this study presents a comprehensive dataset consisting of 3,000 two-dimensional urban flow simulations conducted using a lattice-Boltzmann method across three distinct Reynolds numbers. The dataset contains the time-averaged velocity fields. A key feature of this dataset is its high geometric diversity: each layout incorporates between three and six buildings with randomized sizes, positions, and rotation angles ranging from 0° to 90°. This extensive variability enables the dataset to capture several critical flow characteristics, including wake formation, flow acceleration, shielding effects, and recirculation zones, across a wide range of orchestrated urban canopies. The large sample size and consistent simulation setup make the dataset particularly suitable for developing and benchmarking ML architectures. In addition, the dataset can support transfer-learning strategies in which models trained on large two-dimensional datasets are adapted to smaller and more computationally expensive three-dimensional datasets.

physics.flu-dyn

Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs

The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional hardware is challenging due to limited DRAM capacity and high GPU costs. Thus, in this work, we propose offloading the single-batch token generation to a 3D NAND flash processing-in-memory (PIM) device, leveraging its high storage density to overcome the DRAM capacity wall. We explore 3D NAND flash configurations and present a re-architected PIM array with an H-tree network for optimal latency and cell density. Along with the well-chosen PIM array size, we develop operation tiling and mapping methods for LLM layers, achieving a 2.4x speedup over four RTX4090 with vLLM and comparable performance to four A100 with only 4.9% latency overhead. Our detailed area analysis reveals that the proposed 3D NAND flash PIM architecture can be integrated within a 4.98mm2 die area under the memory array, without extra area overhead.

cs.AR

Classical eikonal in relativistic scattering

The classical eikonal is defined to be the generator of all scattering observables in a scattering problem in classical mechanics. It was originally introduced as the log of the quantum S-matrix in the classical limit. But its classical nature calls for a definition and computational methods independent of quantum mechanics. In this paper, we formulate a classical interaction picture which serves as the foundation of the classical eikonal. Our emphasis is on generality. In perturbation theories, both Hamiltonian deformation and symplectic deformation are considered. Particles and fields are treated on a similar footing. The causality prescription of the propagator is essentially the same for non-relativistic and relativistic kinematics. For a probe particle in electromagnetic or gravitational background, we present all order formulas for the perturbative eikonal. In the electromagnetic setting, we also illustrate how the eikonal encodes the information on radiation of external fields.

hep-th

Cognitive Weave: Synthesizing Abstracted Knowledge with a Spatio-Temporal Resonance Graph

The emergence of capable large language model (LLM) based agents necessitates memory architectures that transcend mere data storage, enabling continuous learning, nuanced reasoning, and dynamic adaptation. Current memory systems often grapple with fundamental limitations in structural flexibility, temporal awareness, and the ability to synthesize higher-level insights from raw interaction data. This paper introduces Cognitive Weave, a novel memory framework centered around a multi-layered spatio-temporal resonance graph (STRG). This graph manages information as semantically rich insight particles (IPs), which are dynamically enriched with resonance keys, signifiers, and situational imprints via a dedicated semantic oracle interface (SOI). These IPs are interconnected through typed relational strands, forming an evolving knowledge tapestry. A key component of Cognitive Weave is the cognitive refinement process, an autonomous mechanism that includes the synthesis of insight aggregates (IAs) condensed, higher-level knowledge structures derived from identified clusters of related IPs. We present comprehensive experimental results demonstrating Cognitive Weave's marked enhancement over existing approaches in long-horizon planning tasks, evolving question-answering scenarios, and multi-session dialogue coherence. The system achieves a notable 34% average improvement in task completion rates and a 42% reduction in mean query latency when compared to state-of-the-art baselines. Furthermore, this paper explores the ethical considerations inherent in such advanced memory systems, discusses the implications for long-term memory in LLMs, and outlines promising future research trajectories.

cs.AI

Kanana: Compute-efficient Bilingual Language Models

We introduce Kanana, a series of bilingual language models that demonstrate exceeding performance in Korean and competitive performance in English. The computational cost of Kanana is significantly lower than that of state-of-the-art models of similar size. The report details the techniques employed during pre-training to achieve compute-efficient yet competitive models, including high quality data filtering, staged pre-training, depth up-scaling, and pruning and distillation. Furthermore, the report outlines the methodologies utilized during the post-training of the Kanana models, encompassing supervised fine-tuning and preference optimization, aimed at enhancing their capability for seamless interaction with users. Lastly, the report elaborates on plausible approaches used for language model adaptation to specific scenarios, such as embedding, retrieval augmented generation, and function calling. The Kanana model series spans from 2.1B to 32.5B parameters with 2.1B models (base, instruct, embedding) publicly released to promote research on Korean language models.

cs.CL

Recursion for Differential Cross-Section from the Optical Theorem

We present a novel framework for computing differential cross-sections in quantum field theory using the optical theorem and loop amplitudes, circumventing the traditional method of squaring scattering amplitudes. This approach addresses two major computational challenges in high-multiplicity processes: complexity from amplitude squaring and the extensive summations over color and helicity. Our method employs quantum off-shell recursion, a loop-level generalization of Berends--Giele recursion, combined with Veltman's largest time equation (LTE) through a doubling prescription of fields. By deriving Dyson--Schwinger equations within this doubled framework and constructing quantum perturbiner expansions, we develop recursive relations for generating LTEs. We validate our method by successfully reproducing the differential cross-section for tree-level $2 \to 2$ and $2 \to 4$ scalar scattering for $ϕ^{4}$ theory through one-loop and three-loop amplitude calculation respectively. This framework offers an efficient alternative to conventional methods and can be broadly applied to theories with color charges, such as QCD and the Standard Model.

hep-ph

Controlled Text Generation for Black-box Language Models via Score-based Progressive Editor

Controlled text generation is very important for the practical use of language models because it ensures that the produced text includes only the desired attributes from a specific domain or dataset. Existing methods, however, are inapplicable to black-box models or suffer a significant trade-off between controlling the generated text and maintaining its fluency. This paper introduces the Score-based Progressive Editor (ScoPE), a novel approach designed to overcome these issues. ScoPE modifies the context at the token level during the generation process of a backbone language model. This modification guides the subsequent text to naturally include the target attributes. To facilitate this process, ScoPE employs a training objective that maximizes a target score, thoroughly considering both the ability to guide the text and its fluency. Experimental results on diverse controlled generation tasks demonstrate that ScoPE can effectively regulate the attributes of the generated text while fully utilizing the capability of the backbone large language models. Our codes are available at \url{https://github.com/ysw1021/ScoPE}.

cs.CL

Learning Terrain-Aware Kinodynamic Model for Autonomous Off-Road Rally Driving With Model Predictive Path Integral Control

High-speed autonomous driving in off-road environments has immense potential for various applications, but it also presents challenges due to the complexity of vehicle-terrain interactions. In such environments, it is crucial for the vehicle to predict its motion and adjust its controls proactively in response to environmental changes, such as variations in terrain elevation. To this end, we propose a method for learning terrain-aware kinodynamic model which is conditioned on both proprioceptive and exteroceptive information. The proposed model generates reliable predictions of 6-degree-of-freedom motion and can even estimate contact interactions without requiring ground truth force data during training. This enables the design of a safe and robust model predictive controller through appropriate cost function design which penalizes sampled trajectories with unstable motion, unsafe interactions, and high levels of uncertainty derived from the model. We demonstrate the effectiveness of our approach through experiments on a simulated off-road track, showing that our proposed model-controller pair outperforms the baseline and ensures robust high-speed driving performance without control failure.

cs.RO

Poincaré generators at second post-Minkowskian order

We verify the global Poincaré invariance of the Hamiltonian mechanics of gravitating binary dynamics at the second post Minkowskian (2PM) order. For spinless point particles, based on the known 2PM Hamiltonian in the center of momentum frame, we compute the general 2PM Hamiltonian valid in an arbitrary reference frame. An off-shell extension of the 1PM Hamiltonian, which contributes at the 2PM order through an iteration process, plays a crucial role. We then construct the 2PM boost generator that uniquely satisfies all the conditions imposed by the Poincaré algebra.

hep-th

Classical observables from partial wave amplitudes

We study the formalism of Kosower-Maybee-O'Connell (KMOC) to extract classical impulse from quantum amplitude in the context of the partial wave expansion of a 2-to-2 elastic scattering. We take two complementary approaches to establish the connection. The first one takes advantage of Clebsch-Gordan relations for the base amplitudes of the partial wave expansion. The second one is a novel adaptation of the traditional saddle point approximation in the semi-classical limit. In the former, an interference between the S-matrix and its conjugate leads to a large degree of cancellation such that the saddle point approximation to handle a rapidly oscillating integral is no longer needed. As an example with a non-orbital angular momentum, we apply our methods to the charge-monopole scattering problem in the probe limit and reproduce both of the two angles characterizing the classical scattering. A spinor basis for the partial wave expansion, a non-relativistic avatar of the spinor-helicity variables, plays a crucial role throughout our computations.

hep-th

Poincaré invariance of spinning binary dynamics in the post-Minkowskian Hamiltonian approach

We initiate the construction of the global Poincaré algebra generators in the context of the post-Minkowskian Hamiltonian formulation of gravitating binary dynamics in isotropic coordinates that is partly inspired by scattering amplitudes. At the first post-Minkowskian (1PM) order, we write down the Hamiltonian in a form valid in an arbitrary inertial frame. Then we construct the boost generator at the same order which uniquely solves all the equations required by the Poincaré algebra. Our results are linear in Newton's constant but exact in velocities and spins, including all spin multiple moments. We also compute the generators of canonical transformations that proves the equivalence between our new generators and the corresponding generators in the ADM coordinates up to the second post-Newtonian (2PN) order.

gr-qc

Learning-based Uncertainty-aware Navigation in 3D Off-Road Terrains

This paper presents a safe, efficient, and agile ground vehicle navigation algorithm for 3D off-road terrain environments. Off-road navigation is subject to uncertain vehicle-terrain interactions caused by different terrain conditions on top of 3D terrain topology. The existing works are limited to adopt overly simplified vehicle-terrain models. The proposed algorithm learns the terrain-induced uncertainties from driving data and encodes the learned uncertainty distribution into the traversability cost for path evaluation. The navigation path is then designed to optimize the uncertainty-aware traversability cost, resulting in a safe and agile vehicle maneuver. Assuring real-time execution, the algorithm is further implemented within parallel computation architecture running on Graphics Processing Units (GPU).

cs.RO

Physics Embedded Neural Network Vehicle Model and Applications in Risk-Aware Autonomous Driving Using Latent Features

Non-holonomic vehicle motion has been studied extensively using physics-based models. Common approaches when using these models interpret the wheel/ground interactions using a linear tire model and thus may not fully capture the nonlinear and complex dynamics under various environments. On the other hand, neural network models have been widely employed in this domain, demonstrating powerful function approximation capabilities. However, these black-box learning strategies completely abandon the existing knowledge of well-known physics. In this paper, we seamlessly combine deep learning with a fully differentiable physics model to endow the neural network with available prior knowledge. The proposed model shows better generalization performance than the vanilla neural network model by a large margin. We also show that the latent features of our model can accurately represent lateral tire forces without the need for any additional training. Lastly, We develop a risk-aware model predictive controller using proprioceptive information derived from the latent features. We validate our idea in two autonomous driving tasks under unknown friction, outperforming the baseline control framework.

cs.RO

TOAST: Trajectory Optimization and Simultaneous Tracking using Shared Neural Network Dynamics

Neural networks have been increasingly employed in Model Predictive Controller (MPC) to control nonlinear dynamic systems. However, MPC still poses a problem that an achievable update rate is insufficient to cope with model uncertainty and external disturbances. In this paper, we present a novel control scheme that can design an optimal tracking controller using the neural network dynamics of the MPC, making it possible to be applied as a plug-and-play extension for any existing model-based feedforward controller. We also describe how our method handles a neural network containing history information, which does not follow a general form of dynamics. The proposed method is evaluated by its performance in classical control benchmarks with external disturbances. We also extend our control framework to be applied in an aggressive autonomous driving task with unknown friction. In all experiments, our method outperformed the compared methods by a large margin. Our controller also showed low control chattering levels, demonstrating that our feedback controller does not interfere with the optimal command of MPC.

cs.RO