SearcharxivSearch

arXiv subjects

Geza Kovacs

Publications and source records attributed to Geza Kovacs.

At least 19 recordsLinked to original sources

Pulsation-driven helium transport as a potential source of the Blazhko effect

We present a highly simplified nonlinear hydrodynamical model to emulate the main observed features of amplitude modulation (commonly known as Blazhko effect) in RR Lyrae stars. The model is based on the assumption that the periodic flow generated by the pulsation carries surplus helium in the ionization zones He I and II. Once this extra helium reaches a critical amount, a Rayleigh-Taylor-type instability leads to a back-flow of the surplus helium and the process starts over again, due to the continuing effect of pumping helium upward by the pulsation. This periodic variation of helium leads to various efficiency of radiation flux blocking in the helium ionization zone that shows up as a long-term periodic variation of the pulsation amplitude.

astro-ph.SR

TranslateGemma Technical Report

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the translation task, we employ a two-stage fine-tuning process. First, supervised fine-tuning is performed using a rich mixture of high-quality large-scale synthetic parallel data generated via state-of-the-art models and human-translated parallel data. This is followed by a reinforcement learning phase, where we optimize translation quality using an ensemble of reward models, including MetricX-QE and AutoMQM, targeting translation quality. We demonstrate the effectiveness of TranslateGemma with human evaluation on the WMT25 test set across 10 language pairs and with automatic evaluation on the WMT24++ benchmark across 55 language pairs. Automatic metrics show consistent and substantial gains over the baseline Gemma 3 models across all sizes. Notably, smaller TranslateGemma models often achieve performance comparable to larger baseline models, offering improved efficiency. We also show that TranslateGemma models retain strong multimodal capabilities, with enhanced performance on the Vistra image translation benchmark. The release of the open TranslateGemma models aims to provide the research community with powerful and adaptable tools for machine translation.

cs.CL

Secondary eclipses of two brown dwarfs in the K2 fields: detection by multiple dataset merging

By using various data sources for the stellar fluxes in overlapping campaign fields and employing full time series modeling, we report the detection of the secondary eclipses of two brown dwarfs (CWW 89Ab = EPIC 219388192b and HSHJ 430b = EPIC 211946007b). The detections yielded timings in agreement with the orbital elements derived from the earlier radial velocity measurements and eclipse depths of 70+/-12 ppm (CWW 89Ab) and 852+/-123 ppm (HSHJ 430b). While the high depth in the Kepler waveband for HSHJ 430b is in agreement with the assumption that the emitted flux comes mostly from the internal heat source and the absorbed stellar irradiation, the case of CWW 89Ab suggests very high albedo, because of the lack of sufficient thermal radiation in the Kepler waveband. Assuming completely reflective dayside hemisphere, without circulation, the maximum value of the eclipse depth due to the reflection of the stellar light is 56 ppm. By making the extreme assumption that the true eclipse depth is 3 sigma less than the observed depth, the minimum geometric albedo becomes ~0.6.

astro-ph.EP

MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task

In this paper, we present our submissions to the unified WMT25 Translation Evaluation Shared Task. For the Quality Score Prediction subtask, we create a new generation of MetricX with improvements in the input format and the training protocol, while for the Error Span Detection subtask we develop a new model, GemSpanEval, trained to predict error spans along with their severities and categories. Both systems are based on the state-of-the-art multilingual open-weights model Gemma 3, fine-tuned on publicly available WMT data. We demonstrate that MetricX-25, adapting Gemma 3 to an encoder-only architecture with a regression head on top, can be trained to effectively predict both MQM and ESA quality scores, and significantly outperforms its predecessor. Our decoder-only GemSpanEval model, on the other hand, we show to be competitive in error span detection with xCOMET, a strong encoder-only sequence-tagging baseline. With error span detection formulated as a generative task, we instruct the model to also output the context for each predicted error span, thus ensuring that error spans are identified unambiguously.

cs.CL

Digging Deeper for RR Lyrae Stars with Low Modulation Amplitudes

With the goal of searching for very low modulation amplitudes among fundamental mode RR Lyrae stars and assess their incidence rate, we performed a survey of 36 stars observed by the Kepler satellite during the entire four-year period of its mission. The search was conducted by a task-oriented code, designed to find low-amplitude signals in the presence of high-amplitude components and instrumental systematics. We found 7 new modulated stars and negate one earlier claimed star, whereby increasing the number of known Blazhko stars from 17 to 24 and yielding an observed occurrence rate of 67% for the Kepler field. Six of the new stars have the lowest modulation amplitudes found so far, with ~250 ppm Fourier side-lobe amplitudes near the fundamental mode frequency. Because of the small sample size in the Kepler field, we extended the survey to 12 campaign fields observed by K2, the ``two-wheeled'' mission of Kepler. From the 1061 stars we found 514 Blazhko stars. After correcting for the short duration of the time spent on each field, and for the noise dependence of the detections, we arrived at an underlying occurrence rate of ~75% - likely a lower limit for the true rate of Blazhko stars in the K2 fields.

astro-ph.SR

WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects

As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual performance, including on tasks like machine translation (MT). In this work, we extend the WMT24 dataset to cover 55 languages by collecting new human-written references and post-edits for 46 new languages and dialects in addition to post-edits of the references in 8 out of 9 languages in the original WMT24 dataset. The dataset covers four domains: literary, news, social, and speech. We benchmark a variety of MT providers and LLMs on the collected dataset using automatic metrics and find that LLMs are the best-performing MT systems in all 55 languages. These results should be confirmed using a human-based evaluation, which we leave for future work.

cs.CL

SMOL: Professionally translated parallel data for 115 under-represented languages

We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine translation for low-resource languages. SMOL has been translated into 124 (and growing) under-resourced languages (125 language pairs), including many for which there exist no previous public resources, for a total of 6.1M translated tokens. SMOL comprises two sub-datasets, each carefully chosen for maximum impact given its size: SMOLSENT, a set of sentences chosen for broad unique token coverage, and SMOLDOC, a document-level resource focusing on a broad topic coverage. They join the already released GATITOS for a trifecta of paragraph, sentence, and token-level content. We demonstrate that using SMOL to prompt or fine-tune Large Language Models yields robust chrF improvements. In addition to translation, we provide factuality ratings and rationales for all documents in SMOLDOC, yielding the first factuality datasets for most of these languages.

cs.CL

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

As LLMs continue to become more powerful and versatile, human evaluation has quickly become intractable at scale and reliance on automatic metrics has become the norm. Recently, it has been shown that LLMs are themselves state-of-the-art evaluators for many tasks. These Autoraters are typically designed so that they generalize to new systems and test sets. In practice, however, evaluation is performed on a small set of fixed, canonical test sets, which are carefully curated to measure certain capabilities of interest and are not changed frequently. In this work, we design a method which specializes a prompted Autorater to a given test set, by leveraging historical ratings on the test set to construct in-context learning (ICL) examples. We evaluate our Specialist method on the task of fine-grained machine translation evaluation, and show that it dramatically outperforms the state-of-the-art XCOMET metric by 54% and 119% on the WMT'23 and WMT'24 test sets, respectively. We perform extensive analyses to understand the representations learned by our Specialist metrics, and how variability in rater behavior affects their performance. We also verify the generalizability and robustness of our Specialist method for designing automatic metrics across different numbers of ICL examples, LLM backbones, systems to evaluate, and evaluation tasks.

cs.CL

Mitigating Metric Bias in Minimum Bayes Risk Decoding

While Minimum Bayes Risk (MBR) decoding using metrics such as COMET or MetricX has outperformed traditional decoding methods such as greedy or beam search, it introduces a challenge we refer to as metric bias. As MBR decoding aims to produce translations that score highly according to a specific utility metric, this very process makes it impossible to use the same metric for both decoding and evaluation, as improvements might simply be due to reward hacking rather than reflecting real quality improvements. In this work we find that compared to human ratings, neural metrics not only overestimate the quality of MBR decoding when the same metric is used as the utility metric, but they also overestimate the quality of MBR/QE decoding with other neural utility metrics as well. We also show that the metric bias issue can be mitigated by using an ensemble of utility metrics during MBR decoding: human evaluations show that MBR decoding using an ensemble of utility metrics outperforms a single utility metric.

cs.CL

Transforming Wearable Data into Personal Health Insights using Large Language Model Agents

Deriving personalized insights from popular wearable trackers requires complex numerical reasoning that challenges standard LLMs, necessitating tool-based approaches like code generation. Large language model (LLM) agents present a promising yet largely untapped solution for this analysis at scale. We introduce the Personal Health Insights Agent (PHIA), a system leveraging multistep reasoning with code generation and information retrieval to analyze and interpret behavioral health data. To test its capabilities, we create and share two benchmark datasets with over 4000 health insights questions. A 650-hour human expert evaluation shows that PHIA significantly outperforms a strong code generation baseline, achieving 84% accuracy on objective, numerical questions and, for open-ended ones, earning 83% favorable ratings while being twice as likely to achieve the highest quality rating. This work can advance behavioral health by empowering individuals to understand their data, enabling a new era of accessible, personalized, and data-driven wellness for the wider population.

cs.AI

Toward more accurate RR Lyrae metallicities

By using a large sample of published spectroscopic iron abundances, we point out the importance of gravity correction in deriving more accurate metal abundances for RR Lyrae stars. For the 197 stars with multiple spectra we find overall [Fe/H] standard deviations of 0.167 (as published), 0.145 (shifted by data source zero points) and 0.121 (both zero point-shifted and gravity-corrected). These improvements are significant at the ~2 sigma level at each correction step, leading to a clearly significant improvement after both corrections applied. The higher quality of the gravity-corrected metallicities is strongly supported also by the tighter correlation with the metallicities predicted from the period and Fourier phase phi_31. This work underlines the need for using some external estimates of the temporal gravity in the chemical abundance analysis rather than relying on a full-fetched spectrum fit that leads to large correlated errors in the estimated parameters.

astro-ph.SR

Large Language Models are Few-Shot Health Learners

Large language models (LLMs) can capture rich representations of concepts that are useful for real-world tasks. However, language alone is limited. While existing LLMs excel at text-based inferences, health applications require that models be grounded in numerical data (e.g., vital signs, laboratory values in clinical domains; steps, movement in the wellness domain) that is not easily or readily expressed as text in existing training corpus. We demonstrate that with only few-shot tuning, a large language model is capable of grounding various physiological and behavioral time-series data and making meaningful inferences on numerous health tasks for both clinical and wellness contexts. Using data from wearable and medical sensor recordings, we evaluate these capabilities on the tasks of cardiac signal analysis, physical activity recognition, metabolic calculation (e.g., calories burned), and estimation of stress reports and mental health screeners.

cs.CL

A Study of Stellar Spins in 15 Open Clusters

We analyze spectroscopic and photometric data to determine the projected inclinations of stars in 11 open clusters, placing constraints on the spin-axis distributions of six clusters. We combine these results with four additional clusters studied by Healy & McCullough (2020) and Healy et al. (2021) to perform an ensemble analysis of their spins. We find that eight out of ten constrained clusters (80%) have spin-axis orientations consistent with isotropy, and we establish a lower limit of four out of ten (40%) isotropic clusters at 75% confidence, assuming no correlation of spins between clusters. We also identify two clusters whose spin-axis distributions can be better described by a model consisting of an aligned fraction of stars combined with an isotropic distribution. However, the inclination values of these stars may be influenced by systematic error, and the small number of stars modeled as aligned in these two clusters precludes the interpretation that their stellar subsets are physically aligned. Overall, no cluster displays an unambiguous signature of spin alignment, and 97% of the stars in our sample are consistent with isotropic orientations in their respective clusters. Our results offer support for the dominance of turbulence over ordered rotation in clumps and do not suggest alignment of rotation axes and magnetic fields in protostars.

astro-ph.SR

Automatic Correction of Human Translations

We introduce translation error correction (TEC), the task of automatically correcting human-generated translations. Imperfections in machine translations (MT) have long motivated systems for improving translations post-hoc with automatic post-editing. In contrast, little attention has been devoted to the problem of automatically correcting human translations, despite the intuition that humans make distinct errors that machines would be well-suited to assist with, from typos to inconsistencies in translation conventions. To investigate this, we build and release the Aced corpus with three TEC datasets. We show that human errors in TEC exhibit a more diverse range of errors and far fewer translation fluency errors than the MT errors in automatic post-editing datasets, suggesting the need for dedicated TEC models that are specialized to correct human errors. We show that pre-training instead on synthetic errors based on human errors improves TEC F-score by as much as 5.1 points. We conducted a human-in-the-loop user study with nine professional translation editors and found that the assistance of our TEC system led them to produce significantly higher quality revised translations.

cs.CL

Bright single-mode RR Lyrae stars: matching Gaia EDR3 with pulsation and evolutionary models

We combine observed metallicity, optical and infrared magnitudes with evolutionary and pulsation models to derive average luminosities for 156 single-mode RR Lyrae stars. These luminosities are compared with those obtained from the Gaia EDR3 parallaxes, and found in excellent agreement with the high accuracy subsample (62 stars, with relative parallax errors less than 2%). With the temperature and metallicity scale used, no parallax shift seems to be necessary when alpha-enhanced evolutionary models are employed. Some 10% of the sample shows curious `distance keeping' between the evolutionary and pulsation models. The cause of this behavior is not clear at this moment but can be cured by an excessive increase of the reddening.

astro-ph.SR

Probing galactic double-mode RR Lyrae stars against Gaia EDR3

Classical double-mode pulsators (RR Lyrae stars and delta Cepheids) are important for their simultaneous pulsation in low-order radial modes. This enables us to put stringent constraints on their physical parameters. We use 30 bright galactic double-mode RR~Lyrae (RRd) stars to estimate their luminosities and compare them with those derived from the parallaxes of the recent data release (EDR3) of the Gaia survey. We employ pulsation and evolutionary models, together with observationally determined effective temperatures to derive the basic stellar parameters. Excluding 6 outlying stars (e.g., with blending issues) the RRd and Gaia luminosities correlate well. With the adopted temperature zero point from one of the works based on the infrared flux method, we find it necessary to increase the Gaia parallaxes by 0.02 mas to bring the RRd and Gaia luminosities into agreement. This value is consonant with those derived from studies on binary stars in the context of Gaia. We examine also the resulting period-luminosity-metallicity (PLZ) relation in the 2MASS K band as follows from the RRd parameters. This leads to the verification of two independently derived other PLZs. No significant zero point differences are found. Furthermore, the predicted K absolute magnitudes agree within sigma=0.005-0.01mag.

astro-ph.SR

Reconstructing Detailed Browsing Activities from Browser History

Users' detailed browsing activity - such as what sites they are spending time on and for how long, and what tabs they have open and which one is focused at any given time - is useful for a number of research and practical applications. Gathering such data, however, requires that users install and use a monitoring tool over long periods of time. In contrast, browser extensions can gain instantaneous access months of browser history data. However, the browser history is incomplete: it records only navigation events, missing important information such as time spent or tab focused. In this work, we aim to reconstruct time spent on sites with only users' browsing histories. We gathered three months of browsing history and two weeks of ground-truth detailed browsing activity from 185 participants. We developed a machine learning algorithm that predicts whether the browser window is focused and active at one second-level granularity with an F1-score of 0.84. During periods when the browser is active, the algorithm can predict which the domain the user was looking at with 76.2% accuracy. We can use these results to reconstruct the total time spent online for each user with an R^2 value of 0.96, and the total time each user spent on each domain with an R^2 value of 0.92.

cs.HC

QuizCram: A Quiz-Driven Lecture Viewing Interface

QuizCram is an interface for navigating lecture videos that uses quizzes to help users determine what they should view. We developed it in response to observing peaks in video seeking behaviors centered around Coursera's in-video quizzes. QuizCram shows users a question to answer, with an associated video segment. Users can use these questions to navigate through video segments, and find video segments they need to review. We also allow users to review using a timeline of previously answered questions and videos. To encourage users to review the material, QuizCram keeps track of their question-answering and video-watching history and schedules sections they likely have not mastered for review. QuizCram-format materials can be generated from existing lectures with in-video quizzes. Our user study comparing QuizCram to in-video quizzes found that users practice answering and reviewing questions more when using QuizCram, and are better able to remember answers to questions they encountered.

cs.HC