SearcharxivSearch

arXiv subjects

Yi Yang

Publications and source records attributed to Yi Yang.

At least 19 recordsLinked to original sources

Mapping the Emerging Social Science of Large Language Models

Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.

cs.CY

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation

Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5{\deg} of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2{\deg} on ClearGrasp and 3.1{\deg} on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.

cs.CV

SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness

Autonomous agent systems increasingly depend on reusable skill abstractions for consolidating experiential knowledge and domain expertise. These artifacts typically bundle free-form instructions with heterogeneous resources. However, ensuring their correctness remains challenging. Their failure modes transcend conventional code defects to subtle semantic inconsistencies such as intent conflicts, which manifest as silent failures masked by the underlying model. Moreover, skill correctness must be grounded in intended task boundaries and generalizability. We propose SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem. It transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts. For each node, SkillSpec derives an ExpectSpec from the surrounding declared intent, and infers FactSpecs from encoded behavior under partially disclosed intent. An intent mask regulates access to holistic, lineage, neighborhood, and local views to balance the bias introduced by excessive context against unsupported inference caused by insufficient context. SkillSpec jointly reasons over these views to flag candidate defects, and automatically validates them in an isolated sandbox. On 515 real-world skills from SkillsBench and widely downloaded repositories, SkillSpec identified 763 manually confirmed defects across 239 skills, achieving 61.2% precision. The node-level analysis across multiple model families shows that specification reasoning is consistently reliable for code nodes, whereas plain-text nodes remain a major bottleneck. Most defects arise at the boundaries between declared intent and implementation, demonstrating that explicit specifications provide a practical foundation for skill quality assurance in real-world agent ecosystems.

cs.SE

Visual Analysis of LLM-based Entity Resolution from Scientific Papers

This paper focuses on the visual analytics support for extracting domain-specific entity from extensive scientific literature, a task with inherent limitations using traditional named entity resolution methods. With the advent of large language models (LLMs) such as GPT-4, significant improvements over conventional machine learning approaches have been achieved due to LLM's capability on entity resolution integrate abilities such as understanding multiple types of text. This research introduces a new visual analysis pipeline that integrates these advanced LLMs with versatile visualization and interaction designs to support batch entity resolution. Specifically, we focus on a specific material science field of Metal-Organic Frameworks (MOFs) and a large data collection namely CSD-MOFs. Through collaboration with domain experts in material science, we obtain well-labeled synthesis paragraphs. We propose human-in-the-loop refinement over the entity resolution process using visual analytics techniques, which allows domain experts to interactively integrate insights into LLM intelligence, including error analysis and interpretation of the retrieval-augmented generation (RAG) algorithm. Our evaluation through the case study of example selection for RAG demonstrates that this human-machine collaborative approach improved single-document entity resolution accuracy by approximately 30%.

cs.IR

WireSeg-32K: A Physics-Grounded Synthetic Dataset for Wire Instance Segmentation

Deformable linear objects such as wires and cables are difficult to segment because they are thin, highly deformable, and frequently self-occluded, while large-scale instance-level annotations are expensive to obtain in real scenes. Existing resources either focus on cable tracing or semantic segmentation under constrained settings, or generate visually plausible images without physically grounded wire deformation. We present WireSeg-32k, a synthetic dataset for wire instance segmentation with 32,000 RGB images, instance masks, depth maps, and a complementary real-world test set with annotations. To generate this dataset, we develop DeformX, a co-simulation pipeline that couples Cosserat-rod dynamics with photorealistic Isaac Sim rendering, enabling physically plausible, contact-consistent wire shapes, CAD-based wire assets, and diverse visually grounded scenes. As a simple baseline, LoRA fine-tuning SAM3 on WireSeg-32k alone improves real-world mAP@75 by 10.2% over the off-the-shelf model, showing that physically grounded synthetic data can transfer to real wire perception.

cs.CV

Empirical variational principles for preimage entropies

Preimage entropy measures the complexity generated by the inverse-image structure of a non-invertible dynamical system. For a continuous map $f:X\to X$ on a compact metric space, Hurley's pointwise topological preimage entropies $h_m(f)$ and $h_p(f)$ are natural invariants measuring the complexity of preimage sets. The question of whether they admit unconditional variational principles in terms of suitable measure-theoretic counterparts remains open. In this paper we resolve it by using empirical metric preimage entropies $h^*_{m,\mu}(f)$ and $h^*_{p,\mu}(f)$, defined by restricting preimage fibers to orbit segments whose empirical measures are close to a prescribed invariant measure $\mu$. We prove the variational principles $$ h_m(f)=\sup_{\mu\in\mathcal M_f(X)}h^*_{m,\mu}(f), \qquad h_p(f)=\sup_{\mu\in\mathcal M_f(X)}h^*_{p,\mu}(f) $$ for every continuous map on a compact metric space. We also show that, in general, the set of all invariant measures in these formulas cannot be replaced by the set of ergodic invariant measures. We then compare the empirical entropies with the pointwise metric preimage entropy $h_{m,\mu}(f)$. For every ergodic invariant measure $\mu$, we prove $h^*_{p,\mu}(f)\ge h_{m,\mu}(f)$. Moreover, if $f$ has uniform separation of preimages, then for every ergodic invariant measure $\mu$, $$ h^*_{m,\mu}(f)=h^*_{p,\mu}(f)=h_{m,\mu}(f). $$ Examples show that $h_{m,\mu}(f)$ is not comparable with the empirical quantities in general. We further introduce a resolving-partition property, weaker than uniform separation of preimages, under which the variational principle for $h_m(f)$ and $h_{m,\mu}(f)$ holds. Finally, we establish corresponding variational principles for preimage pressure and give an example showing that uniform separation of preimages does not imply forward expansiveness.

math.DS

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

Multi-camera systems are increasingly practical for robotics, AR/VR, and autonomous driving because complementary views reduce depth ambiguity and preserve visibility under occlusion. Existing point-tracking benchmarks, however, focus on a single video or static multi-camera rigs. None test long-term 3D point tracking across several synchronized views under camera motion. We introduce TAPVid-MV (Tracking Any Point in Video across Multiple Views), the first benchmark for this setting. It contains a curated set of 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks across seven subsets spanning indoor and outdoor domains, from robotics and human activity to driving and synthetic procedural scenes. We obtain these trajectories using dataset-specific auxiliary modalities: sensor depth, LiDAR, SLAM and SfM points, human meshes, posed object meshes, and simulation. Every sequence and trajectory is visually verified by human annotators. Across more than 30 baselines, no method comes close to solving the task. Surprisingly, existing multi-view point trackers do not consistently outperform monocular point trackers. By evaluating reconstruction and point tracking on the same datasets, TAPVid-MV helps distinguish errors in recovered geometry from errors in point correspondence. Through this joint analysis, we identify geometry recovery as a major bottleneck for accurate 3D point tracking. Beyond multi-view 3D point tracking, our released annotations support monocular 2D and 3D point tracking, future-trajectory prediction, and 4D reconstruction.

cs.CV

Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still have certain shortcomings, hindering their further development as interactive design tools. Such schemes typically adhere to fixed input patterns, limiting flexibility in text input. Furthermore, their editing capabilities are constrained by a single or a few 2D visual models and require intricate pipeline design to integrate these models into 3D reconstruction processes. To address the aforementioned issues, we propose the Hash-Atlas network, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes. Building on this foundation, we introduce a dialogue-based 3D scene editing approach, termed CE3D++, which is centered on a large language model (LLM) that allows arbitrary textual input from users and interprets their intentions, subsequently facilitating the autonomous invocation of the corresponding visual models. Additionally, we extend CE3D++ to monocular 4D scenes by imposing motion constraints on moving objects and further fine-tuning the LLM by creating a trajectory dataset related to editing tasks, which enables the smaller LLM to schedule up to 30 different visual tools accurately. Experimental results demonstrate that CE3D++ effectively integrates multiple visual models to achieve diverse visual editing effects, possessing strong scene comprehension and multi-round dialog capabilities. The source codes and trained models are available at https://github.com/Fangkang515/CE3D.

cs.CV

Self-partitioned Interfacial Time Crystals

Nonequilibrium many-body systems can spontaneously break symmetry in time, as in time crystals, or in space, through self-organized domains and interfaces. Whether these two forms of symmetry breaking can intertwine so that an emergent interface alone hosts time-crystalline order remains unknown. In this work, by introducing the Rabi-Hatano-Nelson model, we unveil the existence and mechanism of a self-partitioned interfacial time crystal (SPITC), where a homogeneous system generates its own internal and tunable interface, at which the time-translation symmetry is also spontaneously broken. Such a SPITC phase is intrinsically induced by nonreciprocity and open boundary conditions, without external pumping or long-range interaction. The periodic and open boundary phase diagrams of the system are both mapped out; vacuum and Dicke-like superradiance with static or active orders are identified, with analytical phase boundaries in the weak coupling limit. The frequency of the SPITC is found to scale quadratically with the spin-photon coupling strength, as we derive analytically for the slow dynamics of the spins. The position of the SPITC boundary scales with a critical exponent of $-1$ as a function of the degree of nonreciprocity, in stark contrast to $-1/2$ for an otherwise stationary boundary. Our construction of SPITC establishes a route to spatiotemporal order in non-Hermitian many-body systems.

physics.optics

Scattering Equations as the lowest order K-identities in the calculation of Stringy Scaling of Hard String Scattering Amplitudes

We prove explicitly the n-point K-identities we proposed previously in the calculation of one tensor hard string scattering amplitudes (HSSA). These K-identities were the key to obtain the stringy scaling behavior in the saddle point calculation of the n-point HSSA and, on the other hand, to consistently match with the calculation of decoupling of zero norm states. Moreover, we introduce a G function to generate an infinite set of generalized K-identities (GKI). The lowest order set of these GKI is the scattering equations (SE) used in the calculation of field theory amplitudes in the CHY formalism. The next to leading order set of these GKI is the K-identities used previously in the calculation of one tensor HSSA. We conjecture that the higher order sets of these GKI can be used to calculate higher order HSSA and/or multi-tensor HSSA.

hep-th

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes, or attributes reappear. Existing memory mechanisms expose models to nonlocal history, but access alone does not ensure effective use. Our analysis reveals that video DiT layers exhibit distinct preferences for current, recent, and distant context, suggesting that long-range memory requires deciding both what to retrieve and where to use it. We introduce LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere. To reduce reliance on scarce high-quality long-horizon videos and explicit memory-allocation labels, we further propose Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space. Across 100 multi-shot evaluation prompts, LayerRecall achieves the best overall results on MemoBench and MovieBench while matching its backbone on VBench-Long, demonstrating stronger long-range recovery without sacrificing local continuity. Qualitative analyses further reveal memory-guided self-correction, whereby initially mismatched local attributes return to their historical appearance without resetting ongoing motion or scene structure. Additional analyses show cross-backbone portability and negligible inference overhead.

cs.CV

Encore: Infinite Audio-Video Generation with Adaptive Signal Routing

Existing audio-video generation methods produce well-synchronized clips but are limited to short durations, while long-video generation methods extend duration through chunk-based iterative synthesis yet lack audio entirely. Generating long audio-video jointly is fundamentally harder than either task alone: each chunk must simultaneously maintain video temporal coherence, audio temporal coherence, and cross-modal synchronization, whose conditioning signals enter the model through different pathways. In this work, we present Encore for long-form synchronized audio-video generation. Our key insight is to factor this challenge into: (1) local continuity handled by iterative generation with explicit cross-chunk context propagation, and (2) global consistency enforced via reference audio-video signals with shifted position embedding. Building on this design, we propose Adaptive Signal Routing (ASR), which introduces learnable attention biases within self-attention and learnable residual scales on cross-attention outputs, enabling the model to adaptively modulate the influence of each conditioning signal. Trained end-to-end for joint audio-video generation, Encore also supports infinite-length audio-to-video and video-to-audio synthesis at inference by conditioning on the ground-truth modality throughout the denoising process. Experiments on our extended VerseBench for long audio-video evaluation demonstrate that Encore significantly outperforms existing methods in both generation quality and temporal coherence. Code and data for this paper are at https://github.com/shaohua-pan/Encore.

cs.MM

Corrections induced by the GUP to the Lamb shift of an accelerated atom interacting with a quantum scalar field

We investigate the effect of the GUP on the Lamb shift of a two-level atom interacting with a real massless scalar quantum field, within the DDC formalism. For an atom undergoing inertial motion, uniform acceleration, and uniform circular motion, we analyze the separate contributions of vacuum fluctuations and radiation reaction. We first derive the statistical functions of the field along the atom's trajectories for the three types of motion, expressing them as frequency integrals, and then employ them to calculate the vacuum fluctuation and radiation reaction contributions to the radiative level shift. We show that the GUP-modified Lamb shift of the two-level atom arises entirely from vacuum fluctuations and acquires additional corrections proportional to $\beta$. We focus in particular on the acceleration-dependent GUP corrections. For a uniformly accelerated atom, the GUP corrections comprise thermal and nonthermal parts. At low accelerations, the thermal part exhibits nonmonotonic behavior, and is proportional to $a^4$ in the limit $a/\omega_0 \to 0$; the nonthermal part, by contrast, grows nonlinearly and monotonically, exceeding the thermal part by nearly two orders of magnitude at large accelerations. For an atom in uniform circular motion, the GUP corrections are purely nonthermal and also display a nonlinear, monotonic dependence on acceleration, increasing or decreasing steeply according to the sign of $\beta$. For the same $\beta$, the corrections are larger in uniform circular motion than in uniformly accelerated motion, since the former involves terms proportional to both $a^2$ and $a^3$, whereas the latter contains only $a^2$ terms.

quant-ph

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents

Large Language Model (LLM) agents increasingly solve long-horizon tasks through multi-turn interactions with users and external tools. In these settings, relevant task information often unfolds over time rather than being fully specified at the initial prompt. Service agents make this challenge especially concrete: users may clarify or revise their goals, while tool responses provide information needed for subsequent decisions. Thus, a final reward alone cannot indicate which actions contributed to resolving the task. Recent methods rely on comparative evidence from other trajectories or resampled continuations, or on separately constructed step-level learning signals, to refine credit. However, a completed rollout already records how information and errors flow between agent actions. We introduce Influence-Aware Policy Optimization (IAPO), which represents each rollout as a typed influence-dependency graph over trainable agent actions, with user and tool observations serving as evidence. IAPO converts support-use and failed-use structure into routing weights that redistribute the same trajectory-level advantage. Experiments with Qwen3-4B and Qwen3-8B demonstrate superior performance over multi-turn reinforcement learning (RL) baselines across three service-agent benchmarks: ${\tau^2}$-Bench, UserBench, and AgentChangeBench. BFCL-v4 Multi-Turn further shows that these gains do not compromise multi-turn function-calling performance. This work advances the understanding of credit assignment in multi-turn user interactions and provides a principled approach to training service agents from sparse outcome feedback.

cs.LG

Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models

LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show that this hidden convention is fragile: across 2 VLMs and 7 image and video benchmarks, the default layer is sub-optimal in 13 of 14 model-task pairs, and the best layer shifts with both task and visual backbone. Finding that layer by exhaustive layer-wise inference is prohibitively expensive, and no better fixed default exists. We therefore ask whether layer usefulness can instead be predicted from representation geometry. We study matrix-based entropy, introduced for unimodal layer analysis, which we compute over sample-level visual embeddings as Visual Dataset Entropy (VDE); and Gromov-Wasserstein (GW) distance, introduced for encoder-level VLM model selection, which we repurpose as a layer-wise visual--language alignment signal. Transferring these to LLaVA-based models is not obvious a priori: the vision tower is frozen while the multimodal projector is trained, so we profile both sides of the projector. We find that VDE transfers, and GW does not. Computed from 100 unlabeled task samples without downstream inference, pre-projector VDE tracks layer-wise accuracy and its top-ranked layers cover the oracle best layer on every task for the SigLIP-based LLaVA-Video, while giving region-level guidance for the CLIP-based Video-LLaVA. Post-projector profiles show that the projector reshapes visual geometry but does not erase the performance-relevant trend, leaving $\mathrm{VDE}_{\mathrm{pre}}$ the stronger signal. GW instead flattens after projection and is best read as an alignment diagnostic rather than a selector. VDE thus offers an interpretable, training-free policy that narrows the visual-layer search to a handful of candidates for limited downstream verification.

cs.CV

RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction

Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progress, they autonomously learn reaction patterns from extensive datasets with limited integration of established chemical knowledge as priors. To address this limitation, we introduce RetroMPA, a molecular property-aware, post-hoc enhancement module that injects chemical knowledge into the retrosynthesis pipeline. Rather than functioning as an independent SMILES sequence generator, RetroMPA is a broadly applicable, model-agnostic chemical filter designed to recalibrate and optimize the predictive pathways of existing algorithms. This plug-and-play framework integrates seamlessly with a range of data-driven retrosynthesis methods, enhancing outputs without modifying model architecture or requiring resource-intensive retraining. By leveraging a property-aware latent embedding space, RetroMPA consistently improves top-1 accuracy across eight representative retrosynthesis models by an average of 5.50% on USPTO-50K. Furthermore, we validate its scalability on the large-scale USPTO-Full dataset, achieving an average improvement of about 2.03% across both template-based and template-free architectures. Wet-lab experiments provide preliminary support for the practical utility of the framework. These syntheses confirmed viable, previously unreported substrate combinations for classic reaction paradigms---specifically, Suzuki-Miyaura coupling, Bucherer reaction, and Friedel-Crafts acylation---suggesting that RetroMPA can operate beyond mere data fitting. The code is open-sourced at https://github.com/MengzhouLu/RetroMPA.

cs.LG

OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation

Dynamic scene graphs (DSGs) capture spatio-temporal interactions across videos as $\langle$subject, predicate, object$\rangle$ triplets, and underpin downstream tasks such as video captioning, video question answering, and action analysis. However, end-to-end dynamic scene graph generation (DSGG) methods are closed-set: they recognize only objects and predicates from a fixed training vocabulary and struggle with the long-tailed distribution of rare concepts, severely limiting their real-world applicability. Existing open-vocabulary models typically inherit pretrained large language models, resulting in multi-stage training and inference with substantial cost. We introduce OvDSGG, the first end-to-end framework for open-vocabulary DSGG. OvDSGG builds on top of an open-vocabulary Spatial Backbone and a Temporal Backbone; we further propose a Triplet Feature Extraction Module that bridges them, and a Visual-Language Alignment Module that preserves open-vocabulary recognition by learning an adaptive decision boundary in the joint visual-language feature space, without expensive knowledge distillation in existing methods. We further introduce a rigorous open-vocabulary DSGG benchmark adapted from Action Genome, with disjoint Base/Novel splits for both objects and predicates. OvDSGG significantly outperforms open-vocabulary baselines across all metrics, with zero-shot Recall@$K$ scores 10.0--20.4 percentage point higher than the next-best baseline, while on closed-set DSGG remaining competitive with state-of-the-art models. Code and benchmark are publicly available at https://github.com/jhelsby/OvDSGG/.

cs.CV

JWST Spectroscopy of Type Ia Supernova 2025rbs from Maximum Light to the Nebular Phase

We present JWST observations of the Type Ia supernova (SN Ia) 2025rbs ($D=$14.5 Mpc) at +1, +23, and +84 days after B-band maximum, spanning peak light through a wavelength-dependent transition toward the nebular phase. Combined with ground-based optical and near-infrared (NIR) data, our panchromatic spectra (0.4-14 $\mu$m) include the first maximum-light mid-infrared (MIR) spectrum and the earliest MIR spectroscopic sequence of an SN Ia to date. At peak light, the MIR spectrum exhibits a continuum with permitted and forbidden features, including Si II, Ni II, and early-emerging [Ni III-IV] and [Ar II-III]. By +23 days the MIR is dominated by forbidden lines with a weak continuum, and by +84 days it is fully nebular, whereas the optical/NIR spectra remain transitional. The nebular spectrum reveals strongly stratified ejecta, with stable Ni concentrated at the lowest velocities, radioactive Co at intermediate velocities but absent within ~2000 km s$^{-1}$, and Ar occupying an outer shell. We detect small-scale substructure in [Ca IV] 3.21 $\mu$m with fractional amplitudes of a few percent and a characteristic velocity scale of ~800 km s$^{-1}$, which may reflect compositional structure, ionization variations, or both. Radiative-transfer calculations substantially underpredict these MIR Mg II features despite approximately reproducing the NIR Mg II 1.0927 $\mu$m line, suggesting that the relative strengths of these transitions are sensitive to the treatment of Mg ionization and excitation. These observations demonstrate that MIR spectroscopy beginning near maximum light simultaneously probes the emerging inner ejecta and rapidly fading outer burning products, providing new constraints for explosion and radiative-transfer models.

astro-ph.HE