SearcharxivSearch

arXiv subjects

Tengfei Song

Publications and source records attributed to Tengfei Song.

At least 19 recordsLinked to original sources

SN 2022erq: A Superluminous Thermonuclear Supernova with Escalating Preexplosion Mass Loss

We present a photometric and spectroscopic study of the superluminous Type Ia supernova SN 2022erq. Its early spectra, dominated by iron-group elements with weak intermediate-mass features, might indicate highly efficient nuclear burning, broadly similar to that inferred for some overluminous SNe Ia. The rapid emergence and persistence of narrow Balmer emission lines superposed on this iron-rich spectrum provide clear evidence of long-lived interaction with a hydrogen-rich circumstellar medium (CSM), establishing SN 2022erq as a member of the rare Ia-CSM class. SN 2022erq reached a peak bolometric luminosity of about 8 x 10^43 erg/s and exhibited an exceptionally slow post-peak decline, indicating that its light curve is dominated by long-duration ejecta-CSM interaction. By combining H-alpha diagnostics with bolometric light-curve modeling, we reconstruct the pre-explosion mass-loss history of the progenitor. The mass-loss rate escalated by one order of magnitude over the final decades, rising from about 0.04 to about 0.6 solar masses per year. This surge produced a massive, extended CSM shell of about 3 solar masses out to about 3.5 x 10^16 cm. The young stellar environment (about 100 Myr) together with this substantial, extensive CSM points to a progenitor system consisting of a white dwarf and an intermediate-mass companion that underwent increasing mass loss prior to explosion.

astro-ph.HE

Research on the Flat Field Measurement Method of Coronagraph

The solar corona has an extremely low density, and its brightness is only about one millionth of that of the photosphere. High-dynamic-range imaging of its faint structure is therefore essential for studying coronal heating, coronal mass ejections, and space weather. Quantitative coronagraph imaging requires flat-field measurement and calibration, which underpin intensity calibration, small-scale feature detection, and long-term cyclic analysis. This paper analyzes the coronagraph imaging chain and the origins of flat-field errors, including optical aberrations, stray light, and pixel-response non-uniformity, and summarizes the resulting calibration requirements of next-generation coronagraphs. On this basis, ground-based and space-based flat-fielding methods are systematically reviewed: the ground-based methods include integrating-sphere uniform light sources, opal glass/diffuser plates, clear-sky and thin-cloud backgrounds, and solar-disk scanning, while the space-based methods include internal light sources and diffuser plates, attitude-roll and off-corona offset observations, and multi-phase statistical self-consistent flat-fielding. Their accuracy, resource cost, and applicability are compared. The review shows that no single method is simultaneously high-precision, easy to update, and engineer-friendly; a hierarchical, multi-method calibration framework is therefore recommended. Finally, a new method is proposed in which lithographically generated structured light fields, combined with Fourier-optics and machine-learning inversion, are used to estimate the pixel-response function. Preliminary experiments show that this method achieves a lower residual error than the integrating-sphere and opal-glass methods, providing a high-precision reference for future wide-band, high-resolution coronagraph calibration.

astro-ph.SR

Observational Technological Innovations and Future Development of the Lijiang Coronagraph

As a core ground-based coronal observation facility in China's low-latitude high-altitude regions, the Lijiang Coronagraph leverages the natural advantages of Lijiang Astronomical Observation Station, including its 3200 m altitude and low atmospheric turbulence. It has undergone a full development process, from introduction via Chinese-Japanese cooperation to independent innovation and iteration. This paper systematically summarizes its core technological innovations: upgrade of the automatic operating system, integration of the dual-band observation system, stray light suppression based on image differencing before and after cleaning, and high-precision image calibration and registration. These advances have significantly improved observation efficiency and data quality, laying a solid foundation for high-quality observations. Scientifically, the data reveal that 1.1 solar radii is a highly correlated region between coronal green line brightness and magnetic field intensity. The study also confirms a strong correlation between the coronal green line and the SDO/AIA 21.1 nm extreme ultraviolet band (correlation coefficient: 0.89-0.99), supporting early warning research on Coronal Mass Ejections (CMEs). These results provide key data for verifying coronal heating mechanisms and exploring the origin of the slow solar wind. The experience from the Lijiang Coronagraph not only lays a foundation for China's next-generation large-aperture coronagraphs, but also accelerates progress in low coronal observation capabilities, enabling the country to build internationally competitive capabilities in this field. The system is also an important part of the global coronal observation network.

astro-ph.SR

Forward Modeling of Dust-Induced Stray Light in Ground-Based Coronagraphs: A Dual-Path Monitoring Approach for High-Precision Inner Corona Observations

High-precision ground-based observations of the inner corona (1.05-2.0 R_sun) are fundamentally constrained by instrumental stray light, particularly the additive background from dynamic dust accumulation on the objective lens. To address this issue, we propose a correction method for the Spectral Imaging Coronagraph (SICG) based on dual-path real-time monitoring and forward physical modeling. By simultaneously imaging the objective lens surface, we obtain deterministic prior information on dust distribution. We construct a physical point-spread function using optical defocus parameters and reconstruct the nonuniform scattering background via convolution. Model parameters are retrieved through data-driven inversion constrained by polar coronal holes. The method demonstrates excellent robustness under varying contamination conditions. After correction, the rms noise in the polar background is reduced by approximately 67% on average, and the signal-to-background ratio improves by a factor of up to 3.7 under heavy contamination conditions. Comparisons with space-based Solar Dynamics Observatory/Atmospheric Imaging Assembly observations indicate that the corrected images recover the morphological structures of streamers with high fidelity. Further radial intensity analysis reveals that the correction process successfully restores the hydrostatic exponential decay characteristic of inner coronal radiation. The fitted decay coefficient corresponds to a plasma temperature of approximately 2.0 MK, consistent with the characteristic formation temperature of the Fe XIV 530.3 nm line. These results demonstrate that the method effectively eliminates the dominant systematic bias in ground-based observations, providing a reliable data foundation for high-precision coronal thermodynamic and dynamic research with the SICG.

astro-ph.SR

Atmospheric turbulence profiling with the Multistar Turbulence Monitor

Accurate characterization of atmospheric optical turbulence is essential for evaluating astronomical sites and optimizing adaptive optics systems. The Multistar Turbulence Monitor (MTM) infers the vertical distribution of the refractive-index structure constant Cn2(z) from differential image motion measured between multiple stellar pairs in short-exposure frames. We present a comprehensive investigation of the MTM method, combining theoretical analysis, instrument-performance assessment, numerical simulations, and on-sky observations obtained at the Daocheng Astronomical Site. Simulations based on a standard HV turbulence model demonstrate that the inversion pipeline robustly recovers both the integrated seeing and the vertical turbulence profile under realistic centroiding noise and varying pixel scales. The Markov Chain Monte Carlo (MCMC) inversion achieves stable results with thirteen discrete height nodes and provides reliable uncertainties. Three nights of MTM measurements at the Daocheng Astronomical Site show that MTM-derived seeing closely tracks simultaneous Differential Image Motion Monitor (DIMM) results, accurately reproducing both short-term fluctuations and nightly averages. These results confirm that MTM provides a simple, portable, and versatile solution for atmospheric turbulence profiling and routine seeing monitoring.

astro-ph.IM

SN 2023axu: A Type IIP Supernova Interacted with a Low-Density Stellar Wind

We present photometric and spectroscopic observations of Type IIP supernova SN 2023axu, spanning $\sim$400 d after the explosion. Its light curve is typical of normal SNe IIP, with a V-band peak of $-17.25 \pm 0.06$ mag and no early-time excess indicative of strong circumstellar interaction. The early spectra exhibit a distinctive broad "ledge" near 4600 \AA. Through spectral modeling and comparison, we attribute this feature to a blend of C, N, and He lines excited by weak interaction between the ejecta and a low-density stellar wind. The late-time photometric evolution shows no discernible contribution from interaction, arguing against strong late-time circumstellar material engagement and supporting the low-density wind scenario. From modeling, this SN synthesized $\sim 0.055\,M_\odot$ of $^{56}$Ni, and nebular spectrum analysis indicates a progenitor mass near $15\,M_\odot$. SN 2023axu thus exemplifies weak ejecta-wind interaction and highlights the diversity of mass-loss histories and circumstellar environments of SNe II progenitors.

astro-ph.HE

Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation

Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions.

cs.CV

The R2Pub Telescopes for Surveying: An Overview and Performance Evaluation of the System

The R2Pub telescope, built by the Beijing Planetarium, is a 60 cm equatorial binocular telescope located at the Daocheng site of Yunnan Observatories in China, at an altitude of about 4700 m. This paper presents an overview of the R2Pub telescope system, including its design, instrumentation, and survey capabilities, and reports an initial evaluation of its system performance. R2Pub is a prime-focus binocular system, with each optical tube covering a field of view of approximately 18 square degrees. It is designed to detect a wide range of transient and variable sources in the local universe, such as variable stars, eclipsing binaries, supernovae, gamma-ray burst afterglows, tidal disruption events, active galactic nuclei, and other unknown transients. The observatory infrastructure, including the dome, equatorial mount, optical tubes, and associated subsystems, has been fully constructed and installed, and the system has entered the commissioning phase. Benefiting from the high-altitude location, good seeing conditions, and dark sky background at the Daocheng site, performance tests during commissioning show that the R2Pub system can reach a 5-sigma limiting magnitude of about 18.7 mag in the Pan-STARRS r' band with a 60 s exposure. Ongoing observations with R2Pub are expected to contribute to studies of variable and transient phenomena and to enhance public outreach in astronomy. The binocular design enables simultaneous dual-band observations, providing instantaneous color information for transient sources and improving the classification and physical characterization of their properties and evolution.

astro-ph.IM

The First Scientific Flight and Observations of the 50-mm Balloon-Borne White-Light Coronagraph

A 50-mm balloon-borne white-light coronagraph (BBWLC) to observe whitelight solar corona over the altitude range from 1.08 to 1.50 solar radii has recently been indigenously developed by Yunnan Observatories in collaboration with Shangdong University (in Weihai) and Changchun Institute of Optics, Fine Mechanics and Physics, which will significantly improve the ability of China to detect and measure inner corona. On 2022 October 4, its first scientific flight took place at the Dachaidan area in Qinghai province of China. We describe briefly the BBWLC mission including its optical design, mechanical structure, pointing system, the first flight and results associated with the data processing approach. Preliminary analysis of the data shows that BBWLC imaged the Kcorona with three streamer structures on the west limb of the Sun. To further confirm the coronal signals obtained by BBWLC, comparisonswere made with observations of the Kcoronagraph of the High Altitude Observatory and the Atmospheric ImagingAssembly on board the Solar Dynamics Observatory. We conclude that BBWLC eventually observed the white-light corona in its first scientific flight.

astro-ph.SR

Mapping ground-based coronagraphic images to Helioprojective-Cartesian coordinate system by image registration

A few ground-based solar coronagraphs have been installed in western China for observing the low-layer corona in recent years. However, determining the Helioprojective Coordinates for the coronagraphic data with high precision is an important but challenging step for further research with other multi-wavelength data. In this paper, we propose an automatic coronal image registration method that combines local statistical correlation and feature point matching to achieve accurate registration between ground-based coronal green-line images and space-based 211 {\AA} images. Then, the accurate field of view information of the coronal green-line images can be derived, allowing the images to be mapped to the Helioprojective Cartesian Coordinates with an accuracy of no less than 0.1''. This method has been extensively validated using 100 days of coronal data spanning an 11-year period, demonstrating its broad applicability to ground-based coronagraphs equipped with green-line observations. It significantly enhances the scientific value of ground-based coronal data, enabling comprehensive studies of coronal transient activities and facilitating the joint analysis of data from multiple instruments. Additionally, it holds potential for future applications in improving the pointing accuracy of coronagraphs.

astro-ph.SR

PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts

Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecified factor values, often overlooking individual differences in emotion perception and expression. In this paper, we introduce PromptEVC that utilizes natural language prompts for precise and flexible emotion control. To bridge text descriptions with emotional speech, we propose emotion descriptor and prompt mapper to generate fine-grained emotion embeddings, trained jointly with reference embeddings. To enhance naturalness, we present a prosody modeling and control pipeline that adjusts the rhythm based on linguistic content and emotional cues. Additionally, a speaker encoder is incorporated to preserve identity. Experimental results demonstrate that PromptEVC outperforms state-of-the-art controllable EVC methods in emotion conversion, intensity control, mixed emotion synthesis, and prosody manipulation. Speech samples are available at https://jeremychee4.github.io/PromptEVC/.

eess.AS

Multimodal Machine Translation with Visual Scene Graph Pruning

Multimodal machine translation (MMT) seeks to address the challenges posed by linguistic polysemy and ambiguity in translation tasks by incorporating visual information. A key bottleneck in current MMT research is the effective utilization of visual data. Previous approaches have focused on extracting global or region-level image features and using attention or gating mechanisms for multimodal information fusion. However, these methods have not adequately tackled the issue of visual information redundancy in MMT, nor have they proposed effective solutions. In this paper, we introduce a novel approach--multimodal machine translation with visual Scene Graph Pruning (PSG), which leverages language scene graph information to guide the pruning of redundant nodes in visual scene graphs, thereby reducing noise in downstream translation tasks. Through extensive comparative experiments with state-of-the-art methods and ablation studies, we demonstrate the effectiveness of the PSG model. Our results also highlight the promising potential of visual information pruning in advancing the field of MMT.

cs.CV

Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation

Facial emotion perception in the vision large language model (VLLM) is crucial for achieving natural human-machine interaction. However, creating high-quality annotations for both coarse- and fine-grained facial emotion analysis demands costly expertise. The lack of such high-quality instruction data limits the performance of VLLMs in facial emotion perception. To address this, we propose a self-verification approach with emotion knowledge enhancement (SEKE), which generates high-quality instruction data for multi-grained emotion analysis cost-effectively using closed-source VLLM. This approach integrates prior human knowledge to VLLM inference, guided by the inherent correlations between three grained levels of emotion descriptions, i.e., discrete expression, valence-arousal, and action unit, to reliably generate comprehensive annotations. A self-verification strategy with Uncertainty-Aware Monte Carlo sampling (SV-UAMC) is further embedded to efficiently extract more accurate VLLM predictions, further improving annotation reliability. Consequently, we construct a facial emotion instruction dataset (FEID) containing three comprehensive descriptions, which provides coarse- and fine-grained emotional information for effective model training. Additionally, we introduce a facial emotion analysis benchmark (FEAB) to measure the VLLM's corresponding ability. Our method significantly outperforms state-of-the-art methods on three downstream facial emotion analysis tasks.

cs.LG

Memory Reviving, Continuing Learning and Beyond: Evaluation of Pre-trained Encoders and Decoders for Multimodal Machine Translation

Multimodal Machine Translation (MMT) aims to improve translation quality by leveraging auxiliary modalities such as images alongside textual input. While recent advances in large-scale pre-trained language and vision models have significantly benefited unimodal natural language processing tasks, their effectiveness and role in MMT remain underexplored. In this work, we conduct a systematic study on the impact of pre-trained encoders and decoders in multimodal translation models. Specifically, we analyze how different training strategies, from training from scratch to using pre-trained and partially frozen components, affect translation performance under a unified MMT framework. Experiments are carried out on the Multi30K and CoMMuTE dataset across English-German and English-French translation tasks. Our results reveal that pre-training plays a crucial yet asymmetrical role in multimodal settings: pre-trained decoders consistently yield more fluent and accurate outputs, while pre-trained encoders show varied effects depending on the quality of visual-text alignment. Furthermore, we provide insights into the interplay between modality fusion and pre-trained components, offering guidance for future architecture design in multimodal translation systems.

cs.CL

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model

This paper presents the technical solution proposed by Huawei Translation Service Center (HW-TSC) for the "End-to-End Document Image Machine Translation for Complex Layouts" competition at the 19th International Conference on Document Analysis and Recognition (DIMT25@ICDAR2025). Leveraging state-of-the-art open-source large vision-language model (LVLM), we introduce a training framework that combines multi-task learning with perceptual chain-of-thought to develop a comprehensive end-to-end document translation system. During the inference phase, we apply minimum Bayesian decoding and post-processing strategies to further enhance the system's translation capabilities. Our solution uniquely addresses both OCR-based and OCR-free document image translation tasks within a unified framework. This paper systematically details the training methods, inference strategies, LVLM base models, training data, experimental setups, and results, demonstrating an effective approach to document image machine translation.

cs.CV

Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models

The rapid advancement of large vision-language models (LVLMs) has significantly propelled applications in document understanding, particularly in optical character recognition (OCR) and multilingual translation. However, current evaluations of LVLMs, like the widely used OCRBench, mainly focus on verifying the correctness of their short-text responses and long-text responses with simple layout, while the evaluation of their ability to understand long texts with complex layout design is highly significant but largely overlooked. In this paper, we propose Menu OCR and Translation Benchmark (MOTBench), a specialized evaluation framework emphasizing the pivotal role of menu translation in cross-cultural communication. MOTBench requires LVLMs to accurately recognize and translate each dish, along with its price and unit items on a menu, providing a comprehensive assessment of their visual understanding and language processing capabilities. Our benchmark is comprised of a collection of Chinese and English menus, characterized by intricate layouts, a variety of fonts, and culturally specific elements across different languages, along with precise human annotations. Experiments show that our automatic evaluation results are highly consistent with professional human evaluation. We evaluate a range of publicly available state-of-the-art LVLMs, and through analyzing their output to identify the strengths and weaknesses in their performance, offering valuable insights to guide future advancements in LVLM development. MOTBench is available at https://github.com/gitwzl/MOTBench.

cs.LG

Characterization and Correction of the Scattering Background Produced by Dust on the Objective Lens of the Lijiang 10-cm Coronagraph

Scattered light from the objective lens, directly exposed to the intense sunlight, is a dominant source of stray light in internally occulted coronagraphs. The variable stray light, such as the scatter from dust on the objective lens, can produce varying scattering backgrounds in coronal images, significantly impacting image quality and data analysis. Using data acquired by the Lijiang 10-cm Coronagraph, the quantitative relationship between the distribution of dust on the objective lens and the resulting scattering backgrounds background is analyzed. Two empirical models for the scattering background are derived, and used to correct the raw coronal data. The second model, which depends on three parameters and performs better, shows that the scattering-background distribution varies with angle, weakens with increasing height, and enhances with increasing dust level on the objective lens. Moreover, we find that the dust on the center of the objective lens can contribute more significantly to the scattering background than on the edge. This study not only quantitatively confirms the significant impact of the stray light produced by dust on the objective lens of the coronagraph, but also corrects the coronal data with this stray light for the first time. Correcting for dust-scattered light is crucial for the high-precision calibration of ground-based coronagraph data, enabling a more accurate analysis of coronal structures. Furthermore, our model is envisioned to support the provision of reliable observational data for future routine coronal magnetic-field measurements using ground-based coronagraphs.

astro-ph.IM

A Novel Bi-hemispheric Discrepancy Model for EEG Emotion Recognition

The neuroscience study has revealed the discrepancy of emotion expression between left and right hemispheres of human brain. Inspired by this study, in this paper, we propose a novel bi-hemispheric discrepancy model (BiHDM) to learn the asymmetric differences between two hemispheres for electroencephalograph (EEG) emotion recognition. Concretely, we first employ four directed recurrent neural networks (RNNs) based on two spatial orientations to traverse electrode signals on two separate brain regions, which enables the model to obtain the deep representations of all the EEG electrodes' signals while keeping the intrinsic spatial dependence. Then we design a pairwise subnetwork to capture the discrepancy information between two hemispheres and extract higher-level features for final classification. Besides, in order to reduce the domain shift between training and testing data, we use a domain discriminator that adversarially induces the overall feature learning module to generate emotion-related but domain-invariant feature, which can further promote EEG emotion recognition. We conduct experiments on three public EEG emotional datasets, and the experiments show that the new state-of-the-art results can be achieved.

q-bio.NC