SearcharxivSearch

arXiv subjects

Han Lin

Publications and source records attributed to Han Lin.

At least 37 records · Page 2Linked to original sources

VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations

Video temporal grounding (VTG) aims to locate precise segments in videos based on language queries, which is a fundamental challenge in video understanding. While recent Multimodal Large Language Models (MLLMs) have shown promise in tackling VTG through reinforcement learning (RL), they overlook the challenges arising from both the quality and difficulty of training samples. (1) Partially annotated samples. Many samples contain relevant segments beyond the annotated interval, introducing ambiguous supervision. (2) Hard-to-ground samples. Samples with poor zero-shot performance produce consistently low and indistinguishable rewards during RL training, exhibiting no clear preference among multiple outputs and thus hindering learning efficiency. To address these challenges, we propose VideoTG-R1, a novel curriculum RL framework with reflected boundary annotations, enabling data-efficient training. Specifically, we propose a Boundary Reflection Agent that utilizes MLLMs to predict query-relevant timestamps outside the annotated intervals, allowing us to identify and filter out partially annotated samples, thereby reducing ambiguity. Furthermore, we introduce a Difficulty Estimation Agent to assess the training difficulty of each sample and design a curriculum RL strategy that dynamically masks the videos of hard-to-ground samples according to the training steps, easing the training difficulty and providing clearer preference. Experiments on the VTG and grounded VideoQA tasks demonstrate the effectiveness of our method. Remarkably, with only 10% of the training samples and 21% of the computational budget, VideoTG-R1 outperforms full-data counterparts under both group relative policy optimization (GRPO) and supervised fine-tuning (SFT). The code is available at https://github.com/ldong1111/VideoTG-R1.

cs.CV

Fully automatic fabrication of fibre Bragg gratings using an AI-powered femtosecond laser inscription system

Fibre Bragg gratings (FBGs) are widely used in optical sensing and communication systems. Femtosecond laser inscription (FLI) enables hydrogen-free, thermally stable, high-resolution, and complex structures of FBG fabrication, but its practical application is limited by manual operation, low throughput, and sensitivity to laser alignment. In this study, we present an AI-powered FLI system that enables automated, stable, and efficient FBG fabrication. By integrating a Multi-Layer Perceptron (MLP) model for real-time fabrication position correction, the system maintains precise laser alignment (-0.6 to 0.2 microns of the fibre core plane) and ensures consistent processing. Strong and weak FBGs were fabricated in different types of fibres, and their spectral characteristics-including central wavelength, reflectivity, and FWHM-exhibited high stability and repeatability. The results demonstrate that the proposed AI-powered FLI system significantly reduces manual intervention while achieving reliable FBG performance. This approach holds great promise for scalable, high-throughput FBG production and can be extended to the fabrication of arbitrary FBG structures across various fibre types. With further training and model refinement, the AI-powered FLI provides a scalable and intelligent platform for next-generation automated FBG manufacturing.

physics.optics

Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents

There is growing interest in integrating high-fidelity visual synthesis capabilities into large language models (LLMs) without compromising their strong reasoning capabilities. Existing methods that directly train LLMs or bridge LLMs and diffusion models usually suffer from costly training since the backbone LLMs have not seen image representations during pretraining. We present Bifrost-1, a unified framework that bridges pretrained multimodal LLMs (MLLMs) and diffusion models using patch-level CLIP image embeddings as latent variables, which are natively aligned with the MLLM's CLIP visual encoder. These patch-level image embeddings are integrated into the diffusion model with a lightweight adaptation of its ControlNet. To retain the original multimodal reasoning capabilities of MLLMs, we equip the MLLM with a visual generation branch initialized from the original MLLM parameters when predicting the patch-level image embeddings. By seamlessly integrating pretrained MLLMs and diffusion models with patch-level CLIP latents, our framework enables high-fidelity controllable image generation with significant training efficiency. Our experiments demonstrate that Bifrost-1 achieves comparable or better performance than previous methods in terms of visual fidelity and multimodal understanding, with substantially lower compute during training. We also provide comprehensive ablation studies showing the effectiveness of our design choices.

cs.CV

EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance

Recent approaches for video generation with camera control often create anchor videos (i.e., rendered videos that approximate desired camera motions) to guide diffusion models as a structured prior, by rendering from estimated point clouds following camera trajectories. However, errors in point cloud and camera trajectory estimation often lead to inaccurate anchor videos with higher training cost and low efficiency, as the model is forced to compensate for rendering misalignments. To address these limitations, we introduce EPiC, an efficient and precise camera control learning framework that constructs well-aligned training anchor videos without the need for camera pose or point cloud estimation. Concretely, we create highly precise anchor videos by masking source videos based on first-frame visibility, which ensures strong alignment, eliminates the need for camera/point cloud estimation, and thus can be readily applied to any in-the-wild video. Furthermore, we introduce Anchor-ControlNet, a lightweight module that integrates anchor video guidance in visible regions to pretrained video diffusion models, with less than 1% of additional parameters. EPiC achieves efficient training with substantially fewer parameters, training steps, and less data, and generalizes robustly to anchor videos made with point clouds at test time, enabling precise 3D-informed camera control. EPiC achieves SoTA performance on RealEstate10K and MiraData for I2V camera control task. Notably, EPiC also exhibits strong zero-shot generalization to video-to-video (V2V) scenarios.

cs.CV

Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization

Recent advancements in text-to-video (T2V) diffusion models have significantly enhanced the visual quality of the generated videos. However, even recent T2V models find it challenging to follow text descriptions accurately, especially when the prompt requires accurate control of spatial layouts or object trajectories. A recent line of research uses layout guidance for T2V models that require fine-tuning or iterative manipulation of the attention map during inference time. This significantly increases the memory requirement, making it difficult to adopt a large T2V model as a backbone. To address this, we introduce Video-MSG, a training-free Guidance method for T2V generation based on Multimodal planning and Structured noise initialization. Video-MSG consists of three steps, where in the first two steps, Video-MSG creates Video Sketch, a fine-grained spatio-temporal plan for the final video, specifying background, foreground, and object trajectories, in the form of draft video frames. In the last step, Video-MSG guides a downstream T2V diffusion model with Video Sketch through noise inversion and denoising. Notably, Video-MSG does not need fine-tuning or attention manipulation with additional memory during inference time, making it easier to adopt large T2V models. Video-MSG demonstrates its effectiveness in enhancing text alignment with multiple T2V backbones (VideoCrafter2 and CogVideoX-5B) on popular T2V generation benchmarks (T2VCompBench and VBench). We provide comprehensive ablation studies about noise inversion ratio, different background generators, background object detection, and foreground object segmentation.

cs.CV

Supernovae at Distances < 40 Mpc: I.Catalogues and fractions of Supernovae in a Complete Sample

Context.This is the first paper of a series aiming to determine the fractions and birth rates of various types of supernovae (SNe) in the local Universe. Aims. In this paper, we aim to construct a complete sample of SNe in the nearby universe and provide more precise measurement of subtype fractions. Methods.We carefully selected our SN sample at a distance of < 40 Mpc mainly from wide-field surveys conducted over the years from 2016 to 2023. Results.The sample contains a total of 211 SNe, including 109 SNe II, 69 SNe Ia, and 33 SNe Ibc. With the aid of sufficient spectra, we can obtain relatively accurate subtype classifications for all SNe in this sample. After corrections for the Malmquist bias, this volume-limited sample gives fractions of SNe Ia, SNe Ibc, and SNe II as $30.4^{+3.7}_{-11.5}\%$, $16.3^{+3.7}_{-7.4}\%$, and $53.3^{+9.5}_{-18.7}\%$, respectively.In the SN Ia sample, the fraction of the 91T-like subtype becomes relatively low (~5.4\%), while that of the 02cx-like subtype shows a moderate increase (~6.8\%). In the SN Ibc sample, we find significant fractions of broadlined SNe Ic (~18.0\%) and SNe Ibn (~8.8\%). The fraction of 87A-like subtype is determined as ~2.3\% for the first time, indicating rare explosions from blue supergiant stars. We find that SNe Ia show a double peak number distribution in S0- and Sc-type host galaxies, which may serve as a straightforward evidence for the presence of "prompt" and "delayed" progenitor components giving rise to SN Ia explosions. Several subtypes of SNe such as 02cx-like SNe Ia, broadlined SNe Ic, SNe IIn (and perhaps SNe Ibn) are found to occur preferentially in less massive spiral galaxies, favoring their associations with young stellar progenitors. Moreover, the 02cx-like subtype shows a trend of exploding in the outer skirt of their hosts, suggestive of metal-poor progenitors.

astro-ph.HE

Supernovae at Distances < 40 Mpc: II. Supernova Rate in the Local Universe

Context.This is the second paper of a series aiming to determine the birth rates of supernovae in the local Universe. Aims. In this paper, we aim to estimate the SN rates in the local universe and fit the delay-time distribution of SNe Ia to put constraints on their progenitor scenarios. Methods.We performed a Monte-Carlo simulation to estimate the volumetric rates with the nearby SN sample introduced in Paper I of the series. The rate evolution of core-collapse SNe well traces the evolution of cosmic star formation history; while that of SNe Ia involves the convolution of cosmic star-formation history and a two-component delay-time distribution including a power law and a Gaussian component. Results.The volumetric rates of type Ia, Ibc and II SNe are derived as $0.325\pm0.040^{+0.016}_{-0.010}$, $0.160\pm0.028^{+0.044}_{-0.014}$, and $0.528\pm0.051^{+0.162}_{-0.013}$ (in unit of $10^{-4} yr^{-1} Mpc^{-3} h^3_{70}$), respectively. The rate of CCSNe is consistent with previous estimates. The newly derived local SN Ia rate is larger than existing results given at redshifts 0.01 < z < 0.1, favoring an increased rate from the universe at z ~ 0.1 to the local universe. A two-component model can well fit the rate variation, with the power law component accounting for the rate evolution at larger redshifts and the Gaussian component with a delay time of 12.63$\pm$0.38 Gyr accounting for the local rate evolution. This delayed component with such a longer delay time suggests that the progenitors of these SNe Ia were formed at around 1 Gyr after the birth of the universe, which could only be explained by a double-degenerate progenitor scenario. This is evidenced by the comparison with the PTF sample of SNe Ia at z = 0.073, which reveals that the increase in SN Ia rate at z < 0.01 is primarily due to the SNe Ia of massive E and S0 galaxies with old stellar populations.

astro-ph.HE

DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation

Storytelling video generation (SVG) aims to produce coherent and visually rich multi-scene videos that follow a structured narrative. Existing methods primarily employ LLM for high-level planning to decompose a story into scene-level descriptions, which are then independently generated and stitched together. However, these approaches struggle with generating high-quality videos aligned with the complex single-scene description, as visualizing such complex description involves coherent composition of multiple characters and events, complex motion synthesis and multi-character customization. To address these challenges, we propose DREAMRUNNER, a novel story-to-video generation method: First, we structure the input script using a large language model (LLM) to facilitate both coarse-grained scene planning as well as fine-grained object-level layout planning. Next, DREAMRUNNER presents retrieval-augmented test-time adaptation to capture target motion priors for objects in each scene, supporting diverse motion customization based on retrieved videos, thus facilitating the generation of new videos with complex, scripted motions. Lastly, we propose a novel spatial-temporal region-based 3D attention and prior injection module SR3AI for fine-grained object-motion binding and frame-by-frame spatial-temporal semantic control. We compare DREAMRUNNER with various SVG baselines, demonstrating state-of-the-art performance in character consistency, text alignment, and smooth transitions. Additionally, DREAMRUNNER exhibits strong fine-grained condition-following ability in compositional text-to-video generation, significantly outperforming baselines on T2V-ComBench. Finally, we validate DREAMRUNNER's robust ability to generate multi-object interactions with qualitative examples.

cs.CV

D-commuting SYK model: building quantum chaos from integrable blocks

We construct a new family of quantum chaotic models by combining multiple copies of integrable commuting SYK models. As each copy of the commuting SYK model does not commute with others, this construction breaks the integrability of each commuting SYK and the family of models demonstrates the emergence of quantum chaos. We study the spectrum of this model analytically in the double-scaled limit. As the number of copies tends to infinity, the spectrum becomes compact and equivalent to the regular SYK model. For finite $d$ copies, the spectrum is close to the regular SYK model in UV but has an exponential tail $e^{E/T_c}$ in the IR. We identify the reciprocal of the exponent in the tail as a critical temperature $T_c$, above which the model should be quantum chaotic. $T_c$ monotonically decreases as $d$ increases, which expands the chaotic regime over the non-chaotic regime. We propose the existence of a new phase around $T_c$, and the dynamics should be very different in two phases. We further carry out numeric analysis at finite $d$, which supports our proposal. Given any finite dimensional local Hamiltonian, by decomposing it into $d$ groups, in which all terms in one group commute with each other but terms from different groups may not, our analysis can give an estimate of the critical temperature for quantum chaos based on the decomposition. We also comment on the implication of the critical temperature to future quantum simulations of quantum chaos and quantum gravity.

hep-th

VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning

Procedural video representation learning is an active research area where the objective is to learn an agent which can anticipate and forecast the future given the present video input, typically in conjunction with textual annotations. Prior works often rely on large-scale pretraining of visual encoders and prediction models with language supervision. However, the necessity and effectiveness of extending compute intensive pretraining to learn video clip sequences with noisy text supervision have not yet been fully validated by previous works. In this work, we show that a strong off-the-shelf frozen pretrained visual encoder, along with a well designed prediction model, can achieve state-of-the-art (SoTA) performance in forecasting and procedural planning without the need for pretraining the prediction model, nor requiring additional supervision from language or ASR. Instead of learning representations from pixel space, our method utilizes the latent embedding space of publicly available vision encoders. By conditioning on frozen clip-level embeddings from observed steps to predict the actions of unseen steps, our prediction model is able to learn robust representations for forecasting through iterative denoising - leveraging the recent advances in diffusion transformers (Peebles & Xie, 2023). Empirical studies over a total of five procedural learning tasks across four datasets (NIV, CrossTask, COIN and Ego4D-v2) show that our model advances the strong baselines in long-horizon action anticipation (+2.6% in Verb ED@20, +3.1% in Noun ED@20), and significantly improves the SoTA in step forecasting (+5.0%), task classification (+3.8%), and procedure planning tasks (up to +2.28% in success rate, +3.39% in mAcc, and +0.90% in mIoU).

cs.CV

Probing the Shock Breakout Signal of SN 2024ggi from the Transformation of Early Flash Spectroscopy

We present early-time, hour-to-day cadence spectroscopy of the nearby type II supernova (SN II) 2024ggi, which was discovered at a phase when the SN shock just emerged from the red-supergiant (RSG) progenitor star. Over the first few days after the first light, SN 2024ggi exhibited prominent narrow emission lines formed through intense and persistent photoionization of the nearby circumstellar material (CSM). In the first 63 hours, spectral lines of He, C, N, and O revealed a rapid rise in ionization, as a result of the progressive sweeping-up of the CSM by the shock. The duration of the IIn-like spectra indicates a dense and relatively confined CSM distribution extending up to $\sim 4 \times 10^{14}$ cm. Spectral modeling reveals a CSM mass loss rate at this region exceeding $5 \times 10^{-3}{\rm M}_{\odot}$ yr$^{-1}$ is required to reproduce low-ionization emissions, which dramatically exceeds that of an RSG. Analyzing H$α$ emission shift implies the velocity of the unshocked outer CSM to be between 20 and 40 km s$^{-1}$, matching the typical wind velocity of an RSG. The differences between the inner and outer layers of the CSM and an RSG progenitor highlight a complex mass loss history before the explosion of SN 2024ggi.

astro-ph.HE

SN 2021dbg: A Luminous Type IIP-IIL Supernova Exploding from a Massive Star with a Layered Shell

We present extensive observations and analysis of supernova (SN) 2021dbg, utilizing optical photometry and spectroscopy. For approximately 385 days following the explosion, SN 2021dbg exhibited remarkable luminosity, surpassing most SNe II. This initial high luminosity is potentially attributed to the interaction between the ejected material and the surrounding circumstellar material (CSM), as evidenced by the pronounced interaction signatures observed in its spectra. The subsequent high luminosity is primarily due to the significant $^{56}$Ni ($0.17 \pm 0.05$ M$_{\odot}$) produced in the explosion. Based on the flux of flash emission lines detected in the initial spectra, we estimate that the CSM mass near the progenitor amounted to $\sim$(1.0--2.0) $\times 10^{-3}$ M$_{\odot}$, likely resulting from intense stellar wind activity 2--3 yr preceding the explosion. Considering the bolometric light curve, nebular spectrum modeling, and mass-loss rate, we suggest that the progenitor of SN 2021dbg was a red supergiant (RSG) with a mass of $\sim 20$ M$_{\odot}$ and a radius of 1200 R$_{\odot}$. This RSG featured a thick hydrogen shell, which may have contained a region with a sharp decrease in material density, electron density, and temperature, contributing to its layered structure. This object demonstrates mixed features of SNe IIP and SNe IIL, making it as a transitional event linking the above two subclasses of SNe II.

astro-ph.HE

DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning

Text-to-image (T2I) generation has seen significant growth over the past few years. Despite this, there has been little work on generating diagrams with T2I models. A diagram is a symbolic/schematic representation that explains information using structurally rich and spatially complex visualizations (e.g., a dense combination of related objects, text labels, directional arrows/lines, etc.). Existing state-of-the-art T2I models often fail at diagram generation because they lack fine-grained object layout control when many objects are densely connected via complex relations such as arrows/lines, and also often fail to render comprehensible text labels. To address this gap, we present DiagrammerGPT, a novel two-stage text-to-diagram generation framework leveraging the layout guidance capabilities of LLMs to generate more accurate diagrams. In the first stage, we use LLMs to generate and iteratively refine 'diagram plans' (in a planner-auditor feedback loop). In the second stage, we use a diagram generator, DiagramGLIGEN, and a text label rendering module to generate diagrams (with clear text labels) following the diagram plans. To benchmark the text-to-diagram generation task, we introduce AI2D-Caption, a densely annotated diagram dataset built on top of the AI2D dataset. We show that our DiagrammerGPT framework produces more accurate diagrams, outperforming existing T2I models. We also provide comprehensive analysis, including open-domain diagram generation, multi-platform vector graphic diagram generation, human-in-the-loop editing, and multimodal planner/auditor LLMs.

cs.CV

VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Recent text-to-video (T2V) generation methods have seen significant advancements. However, the majority of these works focus on producing short video clips of a single event (i.e., single-scene videos). Meanwhile, recent large language models (LLMs) have demonstrated their capability in generating layouts and programs to control downstream visual modules. This prompts an important question: can we leverage the knowledge embedded in these LLMs for temporally consistent long video generation? In this paper, we propose VideoDirectorGPT, a novel framework for consistent multi-scene video generation that uses the knowledge of LLMs for video content planning and grounded video generation. Specifically, given a single text prompt, we first ask our video planner LLM (GPT-4) to expand it into a 'video plan', which includes the scene descriptions, the entities with their respective layouts, the background for each scene, and consistency groupings of the entities. Next, guided by this video plan, our video generator, named Layout2Vid, has explicit control over spatial layouts and can maintain temporal consistency of entities across multiple scenes, while being trained only with image-level annotations. Our experiments demonstrate that our proposed VideoDirectorGPT framework substantially improves layout and movement control in both single- and multi-scene video generation and can generate multi-scene videos with consistency, while achieving competitive performance with SOTAs in open-domain single-scene T2V generation. Detailed ablation studies, including dynamic adjustment of layout control strength with an LLM and video generation with user-provided images, confirm the effectiveness of each component of our framework and its future potential.

cs.CV

EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents

Recent SOTA approaches for embodied learning via interaction directly employ large language models (LLMs) as agents to determine the next steps in an environment. Due to their world knowledge and reasoning capabilities, LLM agents achieve stronger performance than previous smaller agents based on reinforcement learning (RL); however, frequently calling LLMs is slow and expensive. Instead of directly employing LLMs as agents, can we use LLMs' reasoning capabilities to adaptively create training environments to help smaller RL agents learn useful skills that they are weak at? We propose EnvGen, a novel framework to address this question. We first prompt an LLM to generate training environments by giving it the task description and simulator objectives that the agents should learn and then asking it to generate a set of environment configurations (e.g., different terrains, items initially given to agents, etc.). Next, we train a small RL agent in a mixture of the original and LLM-generated environments. Then, we enable the LLM to continuously adapt the generated environments to progressively improve the skills that the agent is weak at, by providing feedback to the LLM in the form of the agent's performance. We demonstrate the usefulness of EnvGen with comprehensive experiments in Crafter and Heist environments. We find that a small RL agent trained with EnvGen can outperform SOTA methods, including a GPT-4 agent, and learns long-horizon tasks significantly faster. We also show that using an LLM to adapt environments dynamically outperforms curriculum learning approaches and how the environments are adapted to help improve RL agents' weaker skills over time. Additionally, EnvGen is substantially more efficient as it only uses a small number of LLM calls (e.g., 4 in total), whereas LLM agents require thousands of calls. Lastly, we present detailed ablation studies for EnvGen design choices.

cs.CL

Fast Tree-Field Integrators: From Low Displacement Rank to Topological Transformers

We present a new class of fast polylog-linear algorithms based on the theory of structured matrices (in particular low displacement rank) for integrating tensor fields defined on weighted trees. Several applications of the resulting fast tree-field integrators (FTFIs) are presented, including (a) approximation of graph metrics with tree metrics, (b) graph classification, (c) modeling on meshes, and finally (d) Topological Transformers (TTs) (Choromanski et al., 2022) for images. For Topological Transformers, we propose new relative position encoding (RPE) masking mechanisms with as few as three extra learnable parameters per Transformer layer, leading to 1.0-1.5%+ accuracy gains. Importantly, most of FTFIs are exact methods, thus numerically equivalent to their brute-force counterparts. When applied to graphs with thousands of nodes, those exact algorithms provide 5.7-13x speedups. We also provide an extensive theoretical analysis of our methods.

cs.LG

Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model

ControlNets are widely used for adding spatial control to text-to-image diffusion models with different conditions, such as depth maps, scribbles/sketches, and human poses. However, when it comes to controllable video generation, ControlNets cannot be directly integrated into new backbones due to feature space mismatches, and training ControlNets for new backbones can be a significant burden for many users. Furthermore, applying ControlNets independently to different frames cannot effectively maintain object temporal consistency. To address these challenges, we introduce Ctrl-Adapter, an efficient and versatile framework that adds diverse controls to any image/video diffusion model through the adaptation of pretrained ControlNets. Ctrl-Adapter offers strong and diverse capabilities, including image and video control, sparse-frame video control, fine-grained patch-level multi-condition control (via an MoE router), zero-shot adaptation to unseen conditions, and supports a variety of downstream tasks beyond spatial control, including video editing, video style transfer, and text-guided motion control. With six diverse U-Net/DiT-based image/video diffusion models (SDXL, PixArt-$α$, I2VGen-XL, SVD, Latte, Hotshot-XL), Ctrl-Adapter matches the performance of pretrained ControlNets on COCO and achieves the state-of-the-art on DAVIS 2017 with significantly lower computation (< 10 GPU hours).

cs.CV

The Red Supergiant Progenitor of Type II Supernova 2024ggi

We present a detailed analysis of the progenitor and its local environment for the recently discovered type II supernova (SN) 2024ggi at a distance of about 6.7~Mpc, by utilizing the pre-explosion images from the Hubble Space Telescope (HST) and \textit{Spitzer} Space Telescope. The progenitor is identified as a red, bright variable star, with absolute $F814W$-band magnitudes being $-$6.2 mag in 1995 to $-$7.2 mag in 2003, respectively, consistent with that of a normal red supergiant (RSG) star. Combining with the historical mid-infrared light curves, a pulsational period of about 379~days can be inferred for the progenitor star. Fitting its spectral energy distribution with stellar spectral models yields the stellar parameters of temperature, radius and bolometric luminosity as $T_*=3290_{-27}^{+19}$~K, $R_*=887_{-51}^{+60}$~R$_{\odot}$, and log($L$/L$_{\odot}$)$=4.92_{-0.04}^{+0.05}$, respectively. The above parameters indicate that the progenitor of SN 2024ggi is consistent with the stellar evolutionary track of a solar-metallicity massive star with an initial mass of $13_{-1}^{+1}$~M$_{\odot}$. Moreover, our analysis indicates a relatively low mass loss rate (i.e., $< 3\times10^{-6}$~M$_{\odot}$~yr$^{-1}$) for the progenitor compared to that inferred from the flashed spectra and X-ray detection (i.e., $10^{-2}$$-$$ 10$$^{-5}$~M$_{\odot}$~yr$^{-1}$), implying a significant enhancement in mass loss within a few years prior to the explosion.

astro-ph.HE