SearcharxivSearch

arXiv subjects

Jinwoo Park

Publications and source records attributed to Jinwoo Park.

At least 19 recordsLinked to original sources

Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting

In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. This averaging is convenient, but the optimizer sees only the aggregated gradient, which does not reveal whether the variable-wise contributions align or oppose one another. To quantify how often this disagreement arises, we measure the variable-wise gradients directly and find that 30.6% of their pairwise cosine similarities are negative on average across seven datasets. However, conflict and harm are not the same thing. Under shared training 35 of the 64 variables do worse than a full-input single-target oracle, and the harmed fraction is not reliably predicted by how often gradients conflict. We propose Per-Variable Surgery (PV-Surgery), an optimizer-side training strategy for backbones with cache-compatible layers. One backward pass builds variable-wise gradient proxies from output-side signals and keeps the pointwise forecasting loss. Reliability-aware selection targets layers whose proxy sums closely approximate their shared-gradient slices. Conditional pooling forms anchor and conflict pools without dropping variables. Common-direction surgery aligns variable or pooled gradients with their normalized mean and restores input norms to avoid reweighting. In experiments across five backbones, seven datasets, and four horizons, PV-Surgery lowers MSE by 3.61% and MAE by 2.93% on average. For multivariate forecasting, this indicates that the variable-wise structure hidden by mean-loss training is a usable optimization signal.

cs.LG

Band-like Carriers in a Soft, Anharmonic Lattice: Lead-Halide Perovskites

Lead-halide perovskites APbX3 present a striking dichotomy: they exhibit band-like electronic transport despite their exceptionally soft and anharmonic lattices. Optically and electrically, they resemble conventional direct-gap semiconductors, exhibiting light carrier masses and steep absorption onsets, whereas their lattices display liquid-like dynamics, including overdamped octahedral motions, quasielastic Raman central peaks, and exceptionally low thermal conductivities. Solution-processed films nevertheless sustain micrometre-scale carrier diffusion at defect densities that would severely suppress transport in conventional semiconductors. Here, we argue that these apparently disparate properties emerge from a common microscopic framework: a soft, strongly anharmonic, and polar [PbX3]- framework, strongly influenced by the Pb 6s2 lone pair, which simultaneously shapes the antibonding orbital character of the band edges, the magnitude and multiple timescales of the dielectric response, and the slow relaxational dynamics that dress every charge carrier. We therefore invert the conventional order and develop the lattice before the electronic structure, because the nominally cubic phase is better viewed as a thermally fluctuating ensemble of locally symmetry-broken configurations rather than a single geometry. Within this framework, we discuss excitons, Frohlich large polarons in the intermediate-coupling regime, carrier transport, defect tolerance, dimensional reduction in two-dimensional and nanocrystalline derivatives, and symmetry-breaking phenomena. We critically assess three contested issues - defect tolerance, ferroelectricity, and the interpretation of the T-3/2 mobility law - and identify seven open questions together with the key measurements needed to resolve them

cond-mat.mtrl-sci

Terahertz Time-Domain Spectroscopy as a Universal Defect Fingerprinting Tool for Organic Halide Perovskite Solar Cells

Organic-inorganic hybrid perovskites (OHPs) deliver certified single-junction power conversion efficiencies (PCEs) exceeding 26% and perovskite-silicon tandem values surpassing 34%, yet a substantial gap with the Shockley-Queisser (S-Q) limit persists-primarily due to grain-boundary (GB) defects that drive non-radiative recombination, ion migration, and degradation. Rational passivation demands a non-contact tool capable of identifying and quantifying specific defect species in device-relevant thin films, a capability absent from conventional probes. This short review demonstrates that terahertz time-domain spectroscopy (THz-TDS, 0.3-3.0 THz) fulfills this role. Across four OHP compositions fabricated by sequential vacuum evaporation (SVE)-MAPbI3, MAPbBr3, FAPbI3, and CsPbI3-the THz spectral window captures both intrinsic phonon modes and GB-localized molecular defect vibrations, enabling species-specific, quantitative characterization at room temperature. Notably, the oscillator strength of the SVE-specific 1.58 THz absorption in MAPbI3 scales linearly with XPS-quantified CH3NH2 defect concentration, establishing THz-TDS as a direct, non-destructive defect meter. Building on these findings, we propose a three-pillar framework for THz-guided defect engineering: (I) quantitative defect measurement via oscillator-strength analysis, (II) material-specific fingerprint identification from a systematically constructed THz library, and (III) fingerprint-guided defect elimination with real-time feedback-together defining a closed-loop quality-control cycle that connects spectroscopic diagnosis to passivation strategy and, ultimately, to enhanced solar-cell efficiency.

cond-mat.mtrl-sci

Detecting the Undetectable: Enhancing Unsupervised time series Anomaly Detection via Active Learning

Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time series data remains a critical yet unresolved challenge. In large-scale industrial applications, labeling time series data is often prohibitively expensive and time-consuming, making unsupervised learning a practical and widely adopted approach. However, existing unsupervised methods frequently struggle to distinguish near-normal anomalies from normal patterns and are vulnerable to noise contamination within normal samples. To address these limitations, we propose a novel framework that leverages active learning to iteratively enhance the performance of unsupervised models. Our framework's core contributions are (1) a masked time-series reconstruction feedback strategy that forces the model to learn robust temporal dependencies, and (2) a minimax learning strategy that promotes robustness by differentially treating normal and abnormal samples. This process encourages the model to better capture the dynamics of subtle and noisy patterns. The proposed framework is evaluated across 28 test cases involving four multivariate time-series datasets and seven unsupervised backbone models. Experimental results demonstrate a 12.39% improvement in AUC compared to the original models, confirming that our method can be readily integrated into existing unsupervised reconstruction-based anomaly detection systems to significantly enhance their performance.

cs.LG

Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers

Recent studies have explored large language models for time-series anomaly detection, yet existing approaches often rely on a single general-purpose model to directly infer anomaly indices or intervals, limiting controllability, interpretability, and reliability for complex anomaly patterns. We propose SAGE (Specialized Analyzer Group for Expert-like Detection), a multi-agent framework for structured anomaly diagnosis in univariate time series. It decomposes anomaly analysis into four specialized Analyzers for point, structural, seasonal, and pattern anomalies. Each Analyzer applies family-specific numerical tools and diagnostic visualizations to generate evidence, while an evidence-grounded Detector consolidates the evidence into confidence-scored anomaly records with intervals and candidate types. A Supervisor then converts these structured records into analyst-facing diagnostic reports. SAGE further constructs synthetic in-context examples from normal-reference training segments, without using real anomalous segments or anomaly-type labels as in-context examples. Across three benchmarks, SAGE achieves the best average performance among strong ML/DL and language-model-based baselines. Ablation studies and human evaluation further show that the proposed framework improves detection reliability and the practical usefulness of diagnostic outputs.

cs.AI

Spatial Disparities in Fire Shelter Accessibility: Capacity Challenges in the Palisades and Eaton Fires

The increasing frequency and severity of wildfire in California, exacerbated by prolonged drought and environmental changes, pose significant challenges to urban community resilience and equitable emergency response. The study investigates issues of accessibility to shelters during the Palisades and Eaton Fires which started in January 2025 in Southern California that led to over 180,000 displacements and the loss of 16,000 structures. Despite coordinated efforts of many organizations' emergency assistance, shelter shortages left many evacuees without safety or accessible refuge. This research aims to measure shelter accessibility during the fires' peak, evaluate whether existing shelter capacity met the demand, and identify spatial disparities in access. Findings reveal severe shelter shortages and pronounced inequities in access to shelters, particularly in geographically isolated regions and mountainous areas. To address these challenges, we implemented shelter placement strategies using both capacity-based and distance-based approaches, demonstrating potential improvements in accessibility and equity. The findings underscore the critical need for strategic shelter planning and infrastructure development to enhance disaster readiness and reduce vulnerability in regions that frequently experience wildfires.

cs.CY

Forecasting Anomaly Precursors via Uncertainty-Aware Time-Series Ensembles

Detecting anomalies in time-series data is critical in domains such as industrial operations, finance, and cybersecurity, where early identification of abnormal patterns is essential for ensuring system reliability and enabling preventive maintenance. However, most existing methods are reactive: they detect anomalies only after they occur and lack the capability to provide proactive early warning signals. In this paper, we propose FATE (Forecasting Anomalies with Time-series Ensembles), a novel unsupervised framework for detecting Precursors-of-Anomaly (PoA) by quantifying predictive uncertainty from a diverse ensemble of time-series forecasting models. Unlike prior approaches that rely on reconstruction errors or require ground-truth labels, FATE anticipates future values and leverages ensemble disagreement to signal early signs of potential anomalies without access to target values at inference time. To rigorously evaluate PoA detection, we introduce Precursor Time-series Aware Precision and Recall (PTaPR), a new metric that extends the traditional Time-series Aware Precision and Recall (TaPR) by jointly assessing segment-level accuracy, within-segment coverage, and temporal promptness of early predictions. This enables a more holistic assessment of early warning capabilities that existing metrics overlook. Experiments on five real-world benchmark datasets show that FATE achieves an average improvement of 19.9 percentage points in PTaPR AUC and 20.02 percentage points in early detection F1 score, outperforming baselines while requiring no anomaly labels. These results demonstrate the effectiveness and practicality of FATE for real-time unsupervised early warning in complex time-series environments.

cs.LG

Modeling and Optimizing the Provisioning of Exhaustible Capabilities for Simultaneous Task Allocation and Scheduling

Deploying heterogeneous robot teams to accomplish multiple tasks over extended time horizons presents significant computational challenges for task allocation and planning. In this paper, we present a comprehensive, time-extended, offline heterogeneous multi-robot task allocation framework, TRAITS, which we believe to be the first that can cope with the provisioning of exhaustible traits under battery and temporal constraints. Specifically, we introduce a nonlinear programming-based trait distribution module that can optimize the trait-provisioning rate of coalitions to yield feasible and time-efficient solutions. TRAITS provides a more accurate feasibility assessment and estimation of task execution times and makespan by leveraging trait-provisioning rates while optimizing battery consumption -- an advantage that state-of-the-art frameworks lack. We evaluate TRAITS against two state-of-the-art frameworks, with results demonstrating its advantage in satisfying complex trait and battery requirements while remaining computationally tractable.

cs.RO

COMET: Codebook-based Online-adaptive Multi-scale Embedding for Time-series Anomaly Detection

Time series anomaly detection is a critical task across various industrial domains. However, capturing temporal dependencies and multivariate correlations within patch-level representation learning remains underexplored, and reliance on single-scale patterns limits the detection of anomalies across different temporal ranges. Furthermore, focusing on normal data representations makes models vulnerable to distribution shifts at inference time. To address these limitations, we propose Codebook-based Online-adaptive Multi-scale Embedding for Time-series anomaly detection (COMET), which consists of three key components: (1) Multi-scale Patch Encoding captures temporal dependencies and inter-variable correlations across multiple patch scales. (2) Vector-Quantized Coreset learns representative normal patterns via codebook and detects anomalies with a dual-score combining quantization error and memory distance. (3) Online Codebook Adaptation generates pseudo-labels based on codebook entries and dynamically adapts the model at inference through contrastive learning. Experiments on five benchmark datasets demonstrate that COMET achieves the best performance in 36 out of 45 evaluation metrics, validating its effectiveness across diverse environments.

cs.LG

Learning and Optimizing the Efficacy of Spatio-Temporal Task Allocation under Temporal and Resource Constraints

Complex multi-robot missions often require heterogeneous teams to jointly optimize task allocation, scheduling, and path planning to improve team performance under strict constraints. We formalize these complexities into a new class of problems, dubbed Spatio-Temporal Efficacy-optimized Allocation for Multi-robot systems (STEAM). STEAM builds upon trait-based frameworks that model robots using their capabilities (e.g., payload and speed), but goes beyond the typical binary success-failure model by explicitly modeling the efficacy of allocations as trait-efficacy maps. These maps encode how the aggregated capabilities assigned to a task determine performance. Further, STEAM accommodates spatio-temporal constraints, including a user-specified time budget (i.e., maximum makespan). To solve STEAM problems, we contribute a novel algorithm named Efficacy-optimized Incremental Task Allocation Graph Search (E-ITAGS) that simultaneously optimizes task performance and respects time budgets by interleaving task allocation, scheduling, and path planning. Motivated by the fact that trait-efficacy maps are difficult, if not impossible, to specify, E-ITAGS efficiently learns them using a realizability-aware active learning module. Our approach is realizability-aware since it explicitly accounts for the fact that not all combinations of traits are realizable by the robots available during learning. Further, we derive experimentally-validated bounds on E-ITAGS' suboptimality with respect to efficacy. Detailed numerical simulations and experiments using an emergency response domain demonstrate that E-ITAGS generates allocations of higher efficacy compared to baselines, while respecting resource and spatio-temporal constraints. We also show that our active learning approach is sample efficient and establishes a principled tradeoff between data and computational efficiency.

cs.RO

SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMs

Large language models (LLMs) power many modern applications, but serving them at scale remains costly and resource-intensive. Current server-centric systems overlook consumer-grade GPUs at the edge. We introduce SpecEdge, an edge-assisted inference framework that splits LLM workloads between edge and server GPUs using a speculative decoding scheme, exchanging only token outputs over the network. SpecEdge employs proactive edge drafting to overlap edge token creation with server verification and pipeline-aware scheduling that interleaves multiple user requests to increase server-side throughput. Experiments show SpecEdge enhances overall cost efficiency by 1.91x through achieving 2.22x server throughput, and reduces inter token latency by 11.24% compared to a server-only baseline, introducing a scalable, cost-effective paradigm for LLM serving. The code is available at https://github.com/kaist-ina/specedge

cs.CL

Transfer Learning and Locally Linear Regression for Locally Stationary Time Series

This paper investigates locally linear regression for locally stationary time series and develops theoretical results for locally linear smoothing and transfer learning. Existing analyses have focused on local constant estimators and given samples, leaving the principles of transferring knowledge from auxiliary sources across heterogeneous time-varying domains insufficiently established. We derive uniform convergence for multivariate locally linear estimators under strong mixing. The resulting error expansion decomposes stochastic variation, smoothing bias, and a term induced by local stationarity. This additional term, originating from the locally stationary structure, has smaller order than in the Nadaraya-Watson benchmark, explaining the improved local linear performance. Building on these results, we propose bias-corrected transfer learned estimators that connect a sparsely observed series with densely observed related sources through a smoothly varying bias function defined over rescaled time and covariates. An additional refinement shows how local temporal adjustment of this bias enhances stability and enables efficient information borrowing across domains. Simulation studies and an empirical analysis of international fuel prices support the theoretical predictions and demonstrate the practical advantages of transfer learning.

math.ST

Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task Planning

Recent advances in large language models (LLMs) have enabled the automatic generation of executable code for task planning and control in embodied agents such as robots, demonstrating the potential of LLM-based embodied intelligence. However, these LLM-based code-as-policies approaches often suffer from limited environmental grounding, particularly in dynamic or partially observable settings, leading to suboptimal task success rates due to incorrect or incomplete code generation. In this work, we propose a neuro-symbolic embodied task planning framework that incorporates explicit symbolic verification and interactive validation processes during code generation. In the validation phase, the framework generates exploratory code that actively interacts with the environment to acquire missing observations while preserving task-relevant states. This integrated process enhances the grounding of generated code, resulting in improved task reliability and success rates in complex environments. We evaluate our framework on RLBench and in real-world settings across dynamic, partially observable scenarios. Experimental results demonstrate that our framework improves task success rates by 46.2% over Code-as-Policies baselines and attains over 86.8% executability of task-relevant actions, thereby enhancing the reliability of task planning in dynamic environments.

cs.AI

Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling

AI-driven materials discovery that couples automated experimentation with algorithmic decision-making requires process aware recipe to property predictors that are accurate, calibrated, and physically admissible. We approach this as a reasoning problem with large reasoning models (LRMs). To instill reasoning capability into language models, we curate reasoning traces from a teacher model to train a student model. However, most training pipelines select reasoning traces using binary correctness or learned preference signals that poorly reflect physical admissibility. We introduce Physics-aware Rejection Sampling (PaRS), a training-time trace selection scheme that favors traces consistent with fundamental physics and numerically close to targets, with lightweight halting to control compute. We instantiate our framework with a large student model fine-tuned on traces synthesized by a larger teacher model, and evaluate under matched token budgets against various rejection sampling baselines. Our method improves accuracy and calibration, reduces physics-violation rates, and lowers sampling cost relative to baselines. These results indicate that modest, domain-aware constraints combined with trace-level selection provide a practical path toward reliable, efficient LRMs for process-aware property prediction and closed-loop materials design.

cs.AI

Exciton Bohr radius of lead halide perovskites for photovoltaic and light-emitting applications

Exciton Bohr radius (a_B) and exciton binding energy (E_b) of metal halide perovskites are two prime quantities in their applications to both light-emitting diode displays and photovoltaic devices. We develop a reliable theoretical method of simultaneously finding a_B and ε_r^c (dielectric constant) based on the net exciton energy above the bulk band gap. It is estimated that a_B under the dielectric confinement is substantially smaller than a_B in the absence of dielectric confinement: 4.36 nm vs. 5.61 nm in the case of CH3NH3PbBr3. We attribute the enhanced a_B to variations of ε_r^c and the electron-hole correlation energy. We also develop a simple method of finding E_b based on the same net exciton energy. Using this, we attribute the well-known difference in E_b between organic bromide perovskites and iodide counterparts to ε_r^c and explain that iodide perovskites are more suited than bromide counterparts in photovoltaic applications, which require smaller E_b for efficient charge-carriers transport.

cond-mat.mtrl-sci

NeSyC: A Neuro-symbolic Continual Learner For Complex Embodied Tasks In Open Domains

We explore neuro-symbolic approaches to generalize actionable knowledge, enabling embodied agents to tackle complex tasks more effectively in open-domain environments. A key challenge for embodied agents is the generalization of knowledge across diverse environments and situations, as limited experiences often confine them to their prior knowledge. To address this issue, we introduce a novel framework, NeSyC, a neuro-symbolic continual learner that emulates the hypothetico-deductive model by continually formulating and validating knowledge from limited experiences through the combined use of Large Language Models (LLMs) and symbolic tools. Specifically, we devise a contrastive generality improvement scheme within NeSyC, which iteratively generates hypotheses using LLMs and conducts contrastive validation via symbolic tools. This scheme reinforces the justification for admissible actions while minimizing the inference of inadmissible ones. Additionally, we incorporate a memory-based monitoring scheme that efficiently detects action errors and triggers the knowledge refinement process across domains. Experiments conducted on diverse embodied task benchmarks-including ALFWorld, VirtualHome, Minecraft, RLBench, and a real-world robotic scenario-demonstrate that NeSyC is highly effective in solving complex embodied tasks across a range of open-domain environments.

cs.AI

Unveiling Key Aspects of Fine-Tuning in Sentence Embeddings: A Representation Rank Analysis

The latest advancements in unsupervised learning of sentence embeddings predominantly involve employing contrastive learning-based (CL-based) fine-tuning over pre-trained language models. In this study, we analyze the latest sentence embedding methods by adopting representation rank as the primary tool of analysis. We first define Phase 1 and Phase 2 of fine-tuning based on when representation rank peaks. Utilizing these phases, we conduct a thorough analysis and obtain essential findings across key aspects, including alignment and uniformity, linguistic abilities, and correlation between performance and rank. For instance, we find that the dynamics of the key aspects can undergo significant changes as fine-tuning transitions from Phase 1 to Phase 2. Based on these findings, we experiment with a rank reduction (RR) strategy that facilitates rapid and stable fine-tuning of the latest CL-based methods. Through empirical investigations, we showcase the efficacy of RR in enhancing the performance and stability of five state-of-the-art sentence embedding methods.

cs.CL