SearcharxivSearch

arXiv subjects

Han He

Publications and source records attributed to Han He.

At least 37 records · Page 2Linked to original sources

Near-Infrared Ca II Triplet As An Stellar Activity Indicator: Library and Comparative Study

We have established and released a new stellar index library of the Ca II Triplet, which serves as an indicator for characterizing the chromospheric activity of stars. The library is based on data from the Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST) Low-Resolution Spectroscopic Survey (LRS) Data Release 9 (DR9). To better reflect the chromospheric activity of stars, we have defined new indices $R$ and $R^{+}$. The library includes measurements of $R$ and $R^{+}$ for each Ca II infrared triplet (IRT) from 699,348 spectra of 562,863 F, G and K-type solar-like stars with Signal-to-Noise Ratio (SNR) higher than 100, as well as the stellar atmospheric parameters and basic information inherited from the LAMOST LRS Catalog. We compared the differences between the 3 individual index of the Ca II Triplet and also conducted a comparative analysis of $R^{+}_{\lambda8542}$ to the Ca II H&K $S$ and $R^+_{HK}$ index database. We find the fraction of low active stars decreases with $T_{eff}$ and the fraction of high active first decrease with decreasing temperature and turn to increase with decreasing temperature at 5800K. We also find a significant fraction of stars that show high activity index in both Ca II H&K and IRT are binaries with low activity, some of them could be discriminated in Ca II H&K $S$ index and $R^{+}_{\lambda8542}$ space. This newly stellar library serves as a valuable resource for studying chromospheric activity in stars and can be used to improve our comprehension of stellar magnetic activity and other astrophysical phenomena.

astro-ph.SR

Widely Interpretable Semantic Representation: Frameless Meaning Representation for Broader Applicability

This paper presents a novel semantic representation, WISeR, that overcomes challenges for Abstract Meaning Representation (AMR). Despite its strengths, AMR is not easily applied to languages or domains without predefined semantic frames, and its use of numbered arguments results in semantic role labels, which are not directly interpretable and are semantically overloaded for parsers. We examine the numbered arguments of predicates in AMR and convert them to thematic roles that do not require reference to semantic frames. We create a new corpus of 1K English dialogue sentences annotated in both WISeR and AMR. WISeR shows stronger inter-annotator agreement for beginner and experienced annotators, with beginners becoming proficient in WISeR annotation more quickly. Finally, we train a state-of-the-art parser on the AMR 3.0 corpus and a WISeR corpus converted from AMR 3.0. The parser is evaluated on these corpora and our dialogue corpus. The WISeR model exhibits higher accuracy than its AMR counterpart across the board, demonstrating that WISeR is easier for parsers to learn.

cs.CL

$\mathrm{H}α$ chromospheric activity of F-, G-, and K-type stars observed by the LAMOST Medium-Resolution Spectroscopic Survey

The distribution of stellar $\mathrm{H}α$ chromospheric activity with respect to stellar atmospheric parameters (effective temperature $T_\mathrm{eff}$, surface gravity $\log\,g$, and metallicity $\mathrm{[Fe/H]}$) and main-sequence/giant categories is investigated for the F-, G-, and K-type stars observed by the LAMOST Medium-Resolution Spectroscopic Survey (MRS). A total of 329,294 MRS spectra from LAMOST DR8 are utilized in the analysis. The $\mathrm{H}α$ activity index ($I_{\mathrm{H}α}$) and the $\mathrm{H}α$ $R$-index ($R_{\mathrm{H}α}$) are evaluated for the MRS spectra. The $\mathrm{H}α$ chromospheric activity distributions with individual stellar parameters as well as in the $T_\mathrm{eff}$ -- $\log\,g$ and $T_\mathrm{eff}$ -- $\mathrm{[Fe/H]}$ parameter spaces are analyzed based on the $R_{\mathrm{H}α}$ index data. It is found that: (1) for the main-sequence sample, the $R_{\mathrm{H}α}$ distribution with $T_\mathrm{eff}$ has a bowl-shaped lower envelope with a minimum at about 6200 K, a hill-shaped middle envelope with a maximum at about 5600 K, and an upper envelope continuing to increase from hotter to cooler stars; (2) for the giant sample, the middle and upper envelopes of the $R_{\mathrm{H}α}$ distribution first increase with a decrease of $T_\mathrm{eff}$ and then drop to a lower activity level at about 4300 K, revealing different activity characteristics at different stages of stellar evolution; (3) for both the main-sequence and giant samples, the upper envelope of the $R_{\mathrm{H}α}$ distribution with metallicity is higher for stars with $\mathrm{[Fe/H]}$ greater than about $-1.0$, and the lowest-metallicity stars hardly exhibit high $\mathrm{H}α$ indices. A dataset of $\mathrm{H}α$ activity indices for the LAMOST MRS spectra analyzed is provided with this paper.

astro-ph.SR

Unleashing the True Potential of Sequence-to-Sequence Models for Sequence Tagging and Structure Parsing

Sequence-to-Sequence (S2S) models have achieved remarkable success on various text generation tasks. However, learning complex structures with S2S models remains challenging as external neural modules and additional lexicons are often supplemented to predict non-textual outputs. We present a systematic study of S2S modeling using contained decoding on four core tasks: part-of-speech tagging, named entity recognition, constituency and dependency parsing, to develop efficient exploitation methods costing zero extra parameters. In particular, 3 lexically diverse linearization schemas and corresponding constrained decoding methods are designed and evaluated. Experiments show that although more lexicalized schemas yield longer output sequences that require heavier training, their sequences being closer to natural language makes them easier to learn. Moreover, S2S models using our constrained decoding outperform other S2S approaches using external resources. Our best models perform better than or comparably to the state-of-the-art for all 4 tasks, lighting a promise for S2S models to generate non-sequential structures.

cs.CL

DFEE: Interactive DataFlow Execution and Evaluation Kit

DataFlow has been emerging as a new paradigm for building task-oriented chatbots due to its expressive semantic representations of the dialogue tasks. Despite the availability of a large dataset SMCalFlow and a simplified syntax, the development and evaluation of DataFlow-based chatbots remain challenging due to the system complexity and the lack of downstream toolchains. In this demonstration, we present DFEE, an interactive DataFlow Execution and Evaluation toolkit that supports execution, visualization and benchmarking of semantic parsers given dialogue input and backend database. We demonstrate the system via a complex dialog task: event scheduling that involves temporal reasoning. It also supports diagnosing the parsing results via a friendly interface that allows developers to examine dynamic DataFlow and the corresponding execution results. To illustrate how to benchmark SoTA models, we propose a novel benchmark that covers more sophisticated event scheduling scenarios and a new metric on task success evaluation. The codes of DFEE have been released on https://github.com/amazonscience/dataflow-evaluation-toolkit.

cs.CL

Stellar Chromospheric Activity Database of Solar-like Stars Based on the LAMOST Low-Resolution Spectroscopic Survey

$\require{mediawiki-texvc}$A stellar chromospheric activity database of solar-like stars is constructed based on the Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST) Low-Resolution Spectroscopic Survey (LRS). The database contains spectral bandpass fluxes and indexes of Ca II H&K lines derived from 1,330,654 high-quality LRS spectra of solar-like stars. We measure the mean fluxes at line cores of the Ca II H&K lines using a 1 $Å$ rectangular bandpass as well as a 1.09 $Å$ full width at half maximum (FWHM) triangular bandpass, and the mean fluxes of two 20 $Å$ pseudo-continuum bands on the two sides of the lines. Three activity indexes, $S_{\rm rec}$ based on the 1 $Å$ rectangular bandpass, and $S_{\rm tri}$ and $S_L$ based on the 1.09 $Å$ FWHM triangular bandpass, are evaluated from the measured fluxes to quantitatively indicate the chromospheric activity level. The uncertainties of all the obtained parameters are estimated. We also produce spectrum diagrams of Ca II H&K lines for all the spectra in the database. The entity of the database is composed of a catalog of spectral sample and activity parameters, and a library of spectrum diagrams. Statistics reveal that the solar-like stars with high level of chromospheric activity ($S_{\rm rec}>0.6$) tend to appear in the parameter range of $T_{\rm eff}\text{ (effective temperature)}<5500\,{\rm K}$, $4.3<\log\,g\text{ (surface gravity)}<4.6$, and $-0.2<[{\rm Fe/H}]\text{ (metallicity)}<0.3$. This database with more than one million high-quality LAMOST LRS spectra of Ca II H&K lines and basal chromospheric activity parameters can be further used for investigating activity characteristics of solar-like stars and solar-stellar connection.

astro-ph.SR

Broadening and redward asymmetry of H$α$ line profiles observed by LAMOST during a stellar flare on an M-type star

Stellar flares are characterized by sudden enhancement of electromagnetic radiation in stellar atmospheres. So far much of our understanding of stellar flares comes from photometric observations, from which plasma motions in flare regions could not be detected. From the spectroscopic data of LAMOST DR7, we have found one stellar flare that is characterized by an impulsive increase followed by a gradual decrease in the H$α$ line intensity on an M4-type star, and the total energy radiated through H$α$ is estimated to be on the order of $10^{33}$ erg. The H$α$ line appears to have a Voigt profile during the flare, which is likely caused by Stark pressure broadening due to the dramatic increase of electron density and/or opacity broadening due to the occurrence of strong non-thermal heating. Obvious enhancement has been identified at the red wing of the H$α$ line profile after the impulsive increase of the H$α$ line intensity. The red wing enhancement corresponds to plasma moving away from the Earth at a velocity of 100$-$200 km s$^{-1}$. According to the current knowledge of solar flares, this red wing enhancement may originate from: (1) flare-driven coronal rain, (2) chromospheric condensation, or (3) a filament/prominence eruption that either with a non-radial backward propagation or with strong magnetic suppression. The total mass of the moving plasma is estimated to be on the order of $10^{15}$ kg.

astro-ph.SR

An Approach to Inference-Driven Dialogue Management within a Social Chatbot

We present a chatbot implementing a novel dialogue management approach based on logical inference. Instead of framing conversation a sequence of response generation tasks, we model conversation as a collaborative inference process in which speakers share information to synthesize new knowledge in real time. Our chatbot pipeline accomplishes this modelling in three broad stages. The first stage translates user utterances into a symbolic predicate representation. The second stage then uses this structured representation in conjunction with a larger knowledge base to synthesize new predicates using efficient graph matching. In the third and final stage, our bot selects a small subset of predicates and translates them into an English response. This approach lends itself to understanding latent semantics of user inputs, flexible initiative taking, and responses that are novel and coherent with the dialogue context.

cs.CL

The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer Encoders

Multi-task learning with transformer encoders (MTL) has emerged as a powerful technique to improve performance on closely-related tasks for both accuracy and efficiency while a question still remains whether or not it would perform as well on tasks that are distinct in nature. We first present MTL results on five NLP tasks, POS, NER, DEP, CON, and SRL, and depict its deficiency over single-task learning. We then conduct an extensive pruning analysis to show that a certain set of attention heads get claimed by most tasks during MTL, who interfere with one another to fine-tune those heads for their own objectives. Based on this finding, we propose the Stem Cell Hypothesis to reveal the existence of attention heads naturally talented for many tasks that cannot be jointly trained to create adequate embeddings for all of those tasks. Finally, we design novel parameter-free probes to justify our hypothesis and demonstrate how attention heads are transformed across the five tasks during MTL through label analysis.

cs.CL

ELIT: Emory Language and Information Toolkit

We introduce ELIT, the Emory Language and Information Toolkit, which is a comprehensive NLP framework providing transformer-based end-to-end models for core tasks with a special focus on memory efficiency while maintaining state-of-the-art accuracy and speed. Compared to existing toolkits, ELIT features an efficient Multi-Task Learning (MTL) model with many downstream tasks that include lemmatization, part-of-speech tagging, named entity recognition, dependency parsing, constituency parsing, semantic role labeling, and AMR parsing. The backbone of ELIT's MTL framework is a pre-trained transformer encoder that is shared across tasks to speed up their inference. ELIT provides pre-trained models developed on a remix of eight datasets. To scale up its service, ELIT also integrates a RESTful Client/Server combination. On the server side, ELIT extends its functionality to cover other tasks such as tokenization and coreference resolution, providing an end user with agile research experience. All resources including the source codes, documentation, and pre-trained models are publicly available at https://github.com/emorynlp/elit.

cs.CL

Levi Graph AMR Parser using Heterogeneous Attention

Coupled with biaffine decoders, transformers have been effectively adapted to text-to-graph transduction and achieved state-of-the-art performance on AMR parsing. Many prior works, however, rely on the biaffine decoder for either or both arc and label predictions although most features used by the decoder may be learned by the transformer already. This paper presents a novel approach to AMR parsing by combining heterogeneous data (tokens, concepts, labels) as one input to a transformer to learn attention, and use only attention matrices from the transformer to predict all elements in AMR graphs (concepts, arcs, labels). Although our models use significantly fewer parameters than the previous state-of-the-art graph parser, they show similar or better accuracy on AMR 2.0 and 3.0.

cs.CL

Flare Activity and Magnetic Feature Analysis of the Flare Stars II: Sub-Giant Branch

We present an investigation of the magnetic activity and flare characteristics of the sub-giant stars mostly from F and G spectral types and compare the results with the main-sequence (MS) stars. The light curve of 352 stars on the sub-giant branch (SGB) from the Kepler mission is analyzed in order to infer stability, relative coverage and contrast of the magnetic structures and also flare properties using three flare indexes. The results show that: (i) Relative coverage and contrast of the magnetic features along with rate, power and magnitude of flares increase on the SGB due to the deepening of the convective zone and more vigorous magnetic field production (ii) Magnetic activity of the F and G-type stars on the SGB does not show dependency to the rotation rate and does not obey the saturation regime. This is the opposite of what we saw for the main sequence, in which the G-, K- and M-type stars show clear dependency to the Rossby number; (iii) The positive relationship between the magnetic features stability and their relative coverage and contrast remains true on the SGB, though it has lower dependency coefficient in comparison with the MS; (iv) Magnetic proxies and flare indexes of the SGB stars increase with increasing the relative mass of the convective zone.

astro-ph.SR

Handwritten Character Recognition from Wearable Passive RFID

In this paper we study the recognition of handwritten characters from data captured by a novel wearable electro-textile sensor panel. The data is collected sequentially, such that we record both the stroke order and the resulting bitmap. We propose a preprocessing pipeline that fuses the sequence and bitmap representations together. The data is collected from ten subjects containing altogether 7500 characters. We also propose a convolutional neural network architecture, whose novel upsampling structure enables successful use of conventional ImageNet pretrained networks, despite the small input size of only 10x10 pixels. The proposed model reaches 72\% accuracy in experimental tests, which can be considered good accuracy for this challenging dataset. Both the data and the model are released to the public.

cs.CV

Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean

In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that are more faithful to Korean grammar. For compatibility to the rest of UD corpora, we follow the UDv2 guidelines, and extensively revise the part-of-speech tags and the dependency relations to reflect morphological features and flexible word-order aspects in Korean. The original and the revised versions of PKT-UD are experimented with transformer-based parsing models using biaffine attention. The parsing model trained on the revised corpus shows a significant improvement of 3.0% in labeled attachment score over the model trained on the previous corpus. Our error analysis demonstrates that this revision allows the parsing model to learn relations more robustly, reducing several critical errors that used to be made by the previous model.

cs.CL

Establishing Strong Baselines for the New Decade: Sequence Tagging, Syntactic and Semantic Parsing with BERT

This paper presents new state-of-the-art models for three tasks, part-of-speech tagging, syntactic parsing, and semantic parsing, using the cutting-edge contextualized embedding framework known as BERT. For each task, we first replicate and simplify the current state-of-the-art approach to enhance its model efficiency. We then evaluate our simplified approaches on those three tasks using token embeddings generated by BERT. 12 datasets in both English and Chinese are used for our experiments. The BERT models outperform the previously best-performing models by 2.5% on average (7.5% for the most significant case). Moreover, an in-depth analysis on the impact of BERT embeddings is provided using self-attention, which helps understanding in this rich yet representation. All models and source codes are available in public so that researchers can improve upon and utilize them to establish strong baselines for the next decade.

cs.CL

Chirality and magnetic configuration associated with two-ribbon solar flares: AR 10930 versus AR 11158

The structural property of the magnetic field in flare-bearing solar active regions (ARs) is one of the key aspects for understanding and forecasting solar flares. In this paper, we make a comparative analysis on the chirality and magnetic configurations associated with two X-class two-ribbon flares happening in AR 10930 and AR 11158. The photospheric magnetic fields of the two ARs were observed by space-based instruments, and the corresponding coronal magnetic fields were calculated based on the nonlinear force-free field model. The analysis shows that the electric current in the two ARs was distributed mostly around the main polarity inversion lines (PILs) where the flares happened, and the magnetic chirality (indicated by the signs of force-free factor $α$) along the main PILs is opposite for the two ARs, i.e., left-handed ($α<0$) for AR 10930 and right-handed ($α>0$) for AR 11158. It is found that, for both the flare events, a prominent magnetic connectivity (featured by co-localized strong $α$ and strong current density distributions) was formed along the main PIL before flare and was totally broken after flare eruption. The two branches of the broken magnetic connectivity, combined with the prominent magnetic connectivity before flare, compose the opposite magnetic configurations in the two ARs owing to their opposite chirality, i.e., Z-shaped configuration in AR 10930 with left-handed chirality and inverse Z-shaped configuration in AR 11158 with right-handed chirality. It is speculated that two-ribbon flares can be generally classified to these two magnetic configurations by chirality in the flare source regions of ARs.

astro-ph.SR

Magnetic Activity of F-, G-, and K-type Stars in the LAMOST-Kepler Field

Monitoring chromospheric and photospheric indexes of magnetic activity can provide valuable information, especially the interaction between different parts of the atmosphere and their response to magnetic fields. We extract chromospheric indexes, S and Rhk+, for 59,816 stars from LAMOST spectra in the LAMOST-Kepler program, and photospheric index, Reff, for 5575 stars from Kepler light curves. The log Reff shows positive correlation with log Rhk+. We estimate the power-law indexes between Reff and Rhk+ for F-, G-, and K-type stars, respectively. We also confirm the dependence of both chromospheric and photospheric activity on stellar rotation. Ca II H and K emissions and photospheric variations generally decrease with increasing rotation periods for stars with rotation periods exceeding a few days. The power-law indexes in exponential decay regimes show different characteristics in the two activity-rotation relations. The updated largest sample including the activity proxies and reported rotation periods provides more information to understand the magnetic activity for cool stars.

astro-ph.SR