SearcharxivSearch

arXiv subjects

Ilay Kamai

Publications and source records attributed to Ilay Kamai.

8 recordsLinked to original sources

EncoTESS: Age-Sensitive Encodings from Raw TESS Light Curves

Main sequence stars of spectral types late F through M exhibit systematic variability in photometric light curves, particularly when they are young. Rotational modulation of starspots manifests as quasi-sinusoidal variability, which enables the measurement of rotation periods. Variability can also be stochastic, as in stellar flaring. However, since measurements of stochastic processes depend on the time of observation, they are typically noisier. Considering that different manifestations of variability have unique observational nuances, models that naturally unify these are incredibly useful for stellar characterization. Towards this goal, we have developed EncoTESS: a Time Series Foundation Model (TSFM) trained on a subset of TESS 2-min light curves. EncoTESS is specifically designed to handle the observational noise, heteroskedastic measurements, irregular sampling, and large data gaps common to TESS data. It is also ~1% of the size of a typical literature TSFM, so can be run easily on a modern laptop. EncoTESS encodes light curves into a fixed-size latent parameter space, which can be used to infer physical stellar properties and recovers light curve summary statistics well. EncoTESS outperforms rotation period and variability amplitude as age indicators for stars that have not converged onto the slow rotator sequence yet; broadly these include K and M stars less than ~100 Myr, and M stars less than ~1 Gyr. We focus on age inference as an application of EncoTESS in this work, but other downstream tasks such as stellar classification could also be explored. The architecture of EncoTESS enables its future extension to TESS light curves of all cadences, and additional surveys such as Kepler and the upcoming PLATO mission. The core EncoTESS framework and library of encodings produced for the stars used in this work are publicly available at https://github.com/philvanlane/encotess.

astro-ph.SR

The Maunder Model and Catalog: Stellar Rotation, Bimodal Activity, and Magnetic Braking in Kepler Main-Sequence Stars

We present The Maunder, a machine learning pipeline and resulting catalog of rotation periods for 148,746 main-sequence stars in the Kepler field. To overcome single-catalog systematics and the simulation-to-reality gap, our architecture employs a hybrid training objective: a joint-embedding self-supervised loss applied to all light curves, combined with a supervised loss trained strictly on cross-catalog consensus labels. By processing multi-scale time- and frequency-domain inputs over rolling windows, the model leverages conformalized quantile regression to output calibrated predictive intervals, providing statistically robust per-star rotation uncertainty metrics. This rolling-window inference reveals that 31,953 stars (21.5$\%$) exhibit bimodal rotational signals. By incorporating APOGEE $v \sin i$ measurements, we demonstrate that for distinct (non-harmonic) bimodals, the longer mode represents the true rotation, exposing a systematic failure mode wherein classical single-pass periodograms lock onto shorter aliases. Filtering by our calibrated confidence intervals yields a highly reliable subset of 119,428 stars. The catalog resolves various rotation-related phenomena: the metallicity dependence of rotation at fixed stellar mass, pointing on the role of metallicity in magnetic braking processes; tracing equatorial velocity and specific angular momentum directly across the Kraft break; recovery of empirical gyrochronology sequences and identification of hierarchical triple candidates among the synchronized-binary population. \emph{The Maunder} provides reliable rotation periods for the largest main-sequence population in \textit{Kepler}, allowing for population-level studies of rotation-based phenomena.

astro-ph.SR

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds and when each fails --- a gap that leaves practitioners, especially in scientific domains with heterogeneous instruments and multiple levels of measurement, unable to diagnose why standard methods underperform the best single modality. We study both objectives under a spiked signal-plus-noise model with structured cross-modal nuisance correlation, the ingredient that breaks the classical recovery guarantees, and derive separation ratios that expose complementary failure modes: alignment whitens each modality and fails when nuisance is strongly correlated across views; prediction encodes whatever is cross-predictable through a one-sided whitening, with recovery governed by source-modality quality. The resulting phase diagram partitions multimodal problems into four regimes --- Both, CA only, CP only, and Neither --- refined by a recovery count that separates partial recovery from complete failure. We present a data-driven procedure to locate real-world datasets in this diagram using a small labeled subsample, identifying the preferred objective and prediction direction before any cross-modal training, and identifying when no objective in the CA/CP family can improve on the stronger modality alone. Experiments on synthetic data, stereo-vision benchmarks, image--caption pairs, and two real scientific domains --- astronomy and single-cell multi-omics --- validate the predictions in the nonlinear regime, including both faces of the Neither regime. Code to reproduce the results is available at https://github.com/IlayMalinyak/mm_align_vs_pred.

cs.LG

Talking with the Latents -- how to convert your LLM into an astronomer

Recent advances in Large Language Models (LLMs) offer unique opportunities for scientific tasks, yet their ability to reason over complex numerical data remains largely unexplored. We propose a simple mechanism to introduce domain-specific physical knowledge into LLMs by fusing pre-trained latent physical features with a pre-trained language model. Our method employs a teacher-student knowledge distillation framework where a large LLM (teacher) generates synthetic question-answer supervision to transfer physical reasoning to a smaller LLM (student). The student is conditioned on latent physical features and trained via a lightweight adapter and Low-Rank Adaptation (LoRA). We demonstrate that this approach, applied to models with 1B, 8B, and 32B parameters, enables effective reasoning over real scientific data. Our models substantially outperform strong baselines, such as Gemini 3 Pro, across multiple downstream tasks without task-specific fine-tuning. We show that the model combines latent information with general physical understanding to predict complex properties and can be "steered" by identifying physically meaningful directions in the latent space. This allows for explicit physical manipulation and natural language interpretation of latent structures. While our experiments focus on astrophysics, the framework is domain-agnostic and applicable to various scientific fields. Our main contribution is a general framework for using LLMs as interpretable interfaces to scientific latent spaces, enabling a single model to perform diverse tasks through natural language guidance. This work marks a step toward developing scientifically capable and useful LLMs.

astro-ph.IM

Sub-Snowline Formation of Gas-Giant Planets in Binary Systems

Gas-giant planets are thought to require conditions beyond the water snow line to build solid cores efficiently. In close binary star systems, the companion's gravity additionally limits the region of stable orbits, potentially excluding the zone where giants should form.} We aim to identify binary systems in which gas giants exist despite the snow line lying in the dynamically unstable zone, and to develop a physically motivated formation channel that explains and predicts their observed locations. We analyse a catalogue of 811 circumstellar binary systems from \citet{Thebault2025}, identifying those hosting gas giants. ($M_p \geq 0.15\,M_\mathrm{Jup}$) with snow lines larger than $0.8\,a_c$ as defined by \citet{Quarles_2020}. We compare their metallicity and eccentricity distributions with the background population, model snow-line evolution with MESA, and fit a linear relation between observed planet semi-major axes and the tidal truncation radius from \citet{Pichardo2005}.} Among 393 gas-giant hosts, we identify 17 systems whose snow line lies in the dynamically unstable zone. Their metallicity and eccentricity distributions are consistent with the background population. We propose that a dust trap formed near the tidal truncation radius of the protoplanetary disc can explain sub-snowline giant formation. The observed planet positions follow $a_\mathrm{planet} = (0.569 \pm 0.05)\,r_t$ ($R^2 = 0.94$), enabling system-by-system predictive power. Evolved systems deviate from this relation, independently supporting a second-generation planet origin for those cases. The tidal truncation of a protoplanetary disc by the stellar companion provides a natural mechanism for sub-snowline gas-giant formation in binaries. The resulting empirical relation yields testable predictions for binary eccentricities in systems lacking direct orbital measurements.

astro-ph.EP

Machine-learning inference of stellar properties using integrated photometric and spectroscopic data

Stellar astrophysics relies on diverse observational modalities-primarily photometric light curves and spectroscopic data from which fundamental stellar properties are inferred. While machine learning (ML) has advanced analysis within individual modalities, the complementary information encoded across modalities remains largely underexploited. We present DESA (Dual Embedding model for Stellar Astrophysics), a novel multi-modal foundation model that integrates light curves and spectra to learn a unified, physically meaningful latent space for stars. DESA first trains separate modality-specific encoders using a hybrid supervised/self-supervised scheme, and then aligns them through DualFormer, a Transformer-based cross-modal integration module tailored for astrophysical data. DualFormer combines cross- and self-attention, a novel dual-projection alignment loss, and a projection-space eigendecomposition that yields physically structured embeddings. We demonstrate that DESA significantly outperforms leading unimodal and self-supervised baselines across a range of tasks. In zero- and few-shot settings, DESA's learned representations recover stellar color-magnitude and Hertzsprung-Russell diagrams with high fidelity ($R^2 = 0.92$ for photometric regressions). In full fine-tuning, DESA achieves state-of-the-art accuracy for binary star detection (AUC = $0.99$, AP = $1.00$) and stellar age prediction (RMSE = $0.94$ Gyr). As a compelling case, DESA naturally separates synchronized binaries from young stars, two populations with nearly identical light curves, purely from their embedded positions in UMAP space, without requiring external kinematic or luminosity information. DESA thus offers a powerful new framework for multimodal, data-driven stellar population analysis, enabling both accurate prediction and novel discovery.

astro-ph.SR

Too fast to be single: Tidal evolution and photometric identification of stellar and planetary companions

Many stars, including those in binary or multiple systems, exhibit modified rotational evolution due to tidal interactions. While magnetic braking slows rotation in single stars, close binaries experience synchronization from tidal forces, resulting in high spin rates. Thus, fast rotators often signify synchronized binaries or planetary systems. We analyze stellar rotation in the Kepler field to photometrically identify non-single systems. Establishing an initial rotation-temperature relationship for individual stars via young clusters, we confirm our findings through magnitude excess and prior binary star system studies. Stars rotating faster than this relationship display a bimodal distribution in peculiar velocity, indicative of non-single or young stars. Leveraging this, we separate non-single stars when peculiar velocity is measurable, or estimate likelihood for those without. Our method identifies 2229 potential non-single star systems with rotation periods exceeding 3 days. For ultra-fast rotators ($P_{rot} < 3 days$), we compile a catalog of 1518 ultra-short-period binary candidates, often part of hierarchical triples, reinforcing rapid spin's association with multiplicity. Applying our method to planet-host stars uncovers Kepler-1184 as a potential circumbinary system and identifies Kepler-493 and Kepler-957 potentially synchronized by close-in planets, with three others as potential false positives. Analysis of known non-single stars reveals clear tidal effects: period synchronization, orbit circularization, and a minimal pericenter constraint for binaries ($r_p \propto (P_{orb}/P_{rot})^{0.77}$). These findings offer insights into tidal evolution, provide a robust method for identifying stellar multiplicity, and have implications for stellar evolution, binary formation, and exoplanet dynamics.

astro-ph.SR

Accurate and Robust Stellar Rotation Periods catalog for 82771 Kepler stars using deep learning

We propose a new framework to predict stellar properties from light curves. We analyze the light-curve data from the Kepler space mission and develop a novel tool for deriving the stellar rotation periods for main-sequence stars. Using this tool, we provide rotation periods for more than 80K stars. Our model, LightPred, is a novel deep-learning model designed to extract stellar rotation periods from light curves. The model utilizes a dual-branch architecture combining Long Short-Term Memory (LSTM) and Transformer components to capture temporal and global data features. We train LightPred on self-supervised contrastive pre-training and simulated light curves generated using a realistic spot model. Our evaluation demonstrates that LightPred outperforms classical methods like the Autocorrelation Function (ACF) in terms of accuracy and average error. We apply LightPred to the Kepler dataset, generating the largest catalog to date. Using error analysis based on learned confidence and consistency metric, we were able to filter the predictions and remove stellar types with variability which is different than spot-induced variability. Our analysis shows strong correlations between error levels and stellar parameters. Additionally, we confirm tidal synchronization in eclipsing binaries with orbital periods shorter than 10 days. Our findings highlight the potential of deep learning in extracting fundamental stellar properties from light curves, opening new avenues for understanding stellar evolution and population demographics.

astro-ph.SR