SearcharxivSearch

arXiv subjects

Ishan Mishra

Publications and source records attributed to Ishan Mishra.

14 recordsLinked to original sources

Aerosols and hydrocarbons in the atmosphere of a white dwarf planet

Most stars, including our Sun, will one day evolve into red giants and, subsequently, white dwarfs. Several planet candidates have recently been identified orbiting white dwarfs, demonstrating that planets can survive the stellar post-main-sequence stage intact. Little is known about the atmospheric composition of post-main-sequence planets, with the most evolved transiting planets with atmospheric detections to date orbiting subgiants. Here we report an atmospheric detection for the white dwarf planet WD 1856 b, achieved through transmission spectroscopy with the JWST NIRSpec PRISM. Our 0.5-5.0 $μ$m spectrum reveals the presence of hydrocarbons (odds ratio of $167:1$ to $5377:1$, with $\mathrm{CH}_4$ preferred at $17:1$ to $30:1$), aerosols ($2 \times 10^5:1$ to $2 \times 10^6:1$), and thermal emission from the planetary nightside ($2 \times 10^{63}:1$ to $2 \times 10^{73}:1$). Our spectral analysis constrains WD 1856 b's mass to $4.3$ to $10.9 \mathrm{M}_J$, finds a carbon-enriched atmosphere (with a $\mathrm{CH}_4$ abundance of $\approx 7\%$), and an effective temperature exceeding the expected planetary equilibrium temperature ($390$ to $412 \, \mathrm{K}$ vs. $160 \, \mathrm{K}$). Based on cooling models, these results suggest that WD 1856 b underwent a migration-related reheating event $3.0$ to $5.5 \, \mathrm{Gyr}$ into the white dwarf phase, consistent with post-main-sequence tidal evolution to the present-day $0.02 \, \mathrm{au}$ circular orbit. Our results provide a window into the ultimate fate of giant planets orbiting stars with masses similar to our Sun.

astro-ph.EP

The Rocky Planet Picture Show: Implementation of Surface Reflection and Emission in $\texttt{POSEIDON}$ with Application to and Interpretation of JWST Data

The surface characterization of rocky exoplanets via emission spectroscopy represents a frontier of current (JWST) and future (HWO) observational efforts. Here, we implement new features in the open-source retrieval code $\texttt{POSEIDON (v1.4)}$ to fully account for an emitting and reflecting planetary surface and an overlying absorbing and scattering atmosphere. We show that realistic rocky surfaces (with wavelength-dependent albedos derived from laboratory measurements) affect emission spectra by imparting mid-infrared diagnostic absorption features, imprinting pseudo-features due to atmospheric transparency windows, and flipping absorption features to emission via surface-atmosphere interface pseudo-temperature inversions. We demonstrate that current JWST spectral data can distinguish between tenuous (low surface pressure, $\leq$ 1 bar) and thick (high surface pressures, $\geq$ 0.1 bar) atmospheres by performing atmosphere + surface retrievals on published JWST emission data of the rocky worlds TOI-1685b and 55 Cancri e. We then explore JWST MIRI LRS's capability to constrain surface geology of rocky worlds, finding that with sufficient SNR retrievals can distinguish between granite-like and basaltic surfaces for synthetic datasets. Finally, we provide an open-source database of lab-derived surface albedos (in the form of directional-hemispherical reflectances), organized by geologic classification and include supplemental tables developed to foster future collaboration between geology and exoplanet science. Our atmosphere + surface retrieval technique provides a pathway to probe geologic processes on rocky exoplanets, showing that upcoming JWST data for terrestrial worlds will enable a deeper exploration of rocky surfaces beyond our Solar System.

astro-ph.EP

Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs

Open-weight AI systems offer unique benefits, including enhanced transparency, open research, and decentralized access. However, they are vulnerable to tampering attacks which can efficiently elicit harmful behaviors by modifying weights or activations. Currently, there is not yet a robust science of open-weight model risk management. Existing safety fine-tuning methods and other post-training techniques have struggled to make LLMs resistant to more than a few dozen steps of adversarial fine-tuning. In this paper, we investigate whether filtering text about dual-use topics from training data can prevent unwanted capabilities and serve as a more tamper-resistant safeguard. We introduce a multi-stage pipeline for scalable data filtering and show that it offers a tractable and effective method for minimizing biothreat proxy knowledge in LLMs. We pretrain multiple 6.9B-parameter models from scratch and find that they exhibit substantial resistance to adversarial fine-tuning attacks on up to 10,000 steps and 300M tokens of biothreat-related text -- outperforming existing post-training baselines by over an order of magnitude -- with no observed degradation to unrelated capabilities. However, while filtered models lack internalized dangerous knowledge, we find that they can still leverage such information when it is provided in context (e.g., via search tool augmentation), demonstrating a need for a defense-in-depth approach. Overall, these findings help to establish pretraining data curation as a promising layer of defense for open-weight AI systems.

cs.LG

Selective Prior Synchronization via SYNC Loss

Prediction under uncertainty is a critical requirement for the deep neural network to succeed responsibly. This paper focuses on selective prediction, which allows DNNs to make informed decisions about when to predict or abstain based on the uncertainty level of their predictions. Current methods are either ad-hoc such as SelectiveNet, focusing on how to modify the network architecture or objective function, or post-hoc such as softmax response, achieving selective prediction through analyzing the model's probabilistic outputs. We observe that post-hoc methods implicitly generate uncertainty information, termed the selective prior, which has traditionally been used only during inference. We argue that the selective prior provided by the selection mechanism is equally vital during the training stage. Therefore, we propose the SYNC loss which introduces a novel integration of ad-hoc and post-hoc method. Specifically, our approach incorporates the softmax response into the training process of SelectiveNet, enhancing its selective prediction capabilities by examining the selective prior. Evaluated across various datasets, including CIFAR-100, ImageNet-100, and Stanford Cars, our method not only enhances the model's generalization capabilities but also surpasses previous works in selective prediction performance, and sets new benchmarks for state-of-the-art performance.

cs.CV

An Assessment of Organics Detection and Characterization on the Surface of Europa with Infrared Spectroscopy

Organics, if they do exist on Europa, may only be present in trace amounts on the surface. NASA's upcoming mission Europa Clipper is going to provide global, high quality data of the surface of Europa in the near-infrared (NIR), specifically the 3-5~$μ$m region, where organics are rich in spectroscopic features. In this work we investigate Europa Clipper's ability to constrain the abundance of selected trace species of interest that span different chemical bonds found in organics, such as C-H, C=C, C$\equiv$C, C=O and C$\equiv$N, via NIR spectroscopy in the 3-5~$μ$m wavelength region. We simulate reflectance spectra of these trace species mixed with water ice, at varying SNR and abundance fractions. The evidence for the trace species in a mixture is evaluated using two approaches: 1) calculating average strength of absorption feature(s), and 2) Bayesian model comparison (BMC) analysis. Our simulations show that sharp and strong spectroscopic features of trace ($\sim 5\%$ abundance by number) organic species should be detectable at $> 3σ$ significance in Europa Clipper quality data. A BMC analysis pushes the $3σ$ detection threshold of trace species even lower to $<1 \%$ abundance. We also consider an example with all trace species mixed together, with overlapping features, and BMC is able to retrieve strong evidence for all of them and also provide constraints on their abundance. These results are promising for Europa Clipper's capability to detect trace organic species, which would allow correlations to be drawn between the composition and geological regions with possibly endogenic material.

astro-ph.EP

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

We present RASO, a foundation model designed to Recognize Any Surgical Object, offering robust open-set recognition capabilities across a broad range of surgical procedures and object classes, in both surgical images and videos. RASO leverages a novel weakly-supervised learning framework that generates tag-image-text pairs automatically from large-scale unannotated surgical lecture videos, significantly reducing the need for manual annotations. Our scalable data generation pipeline gathers 2,200 surgical procedures and produces 3.6 million tag annotations across 2,066 unique surgical tags. Our experiments show that RASO achieves improvements of 2.9 mAP, 4.5 mAP, 10.6 mAP, and 7.2 mAP on four standard surgical benchmarks, respectively, in zero-shot settings, and surpasses state-of-the-art models in supervised surgical action recognition tasks. Code, model, and demo are available at https://ntlm1686.github.io/raso.

cs.CV

Improving Undergraduate Astronomy Students' Skills with Research Literature via Accessible Summaries: An Exploratory Case Study with Astrobites-based Reading Assignments

Undergraduate physics and astronomy students are expected to engage with scientific literature as they begin their research careers, but reading comprehension skills are rarely explicitly taught in major courses. We seek to determine the efficacy of a reading assignment designed to improve undergraduate astronomy (or related) majors' perceived ability to engage with research literature by using accessible summaries of current research written by experts in the field. During the 2022-2023 academic year, faculty members from six institutions incorporated reading assignments using accessible summaries from Astrobites into their undergraduate astronomy major courses, surveyed their students before and after the activities, and participated in follow-up interviews with our research team. Quantitative and qualitative survey data from 52 students show that students' perceptions of their abilities with jargon and identifying main takeaways of a paper significantly improved with use of the tested assignment template. Additionally, students report increased confidence of their abilities within astronomy after exposure to these assignments, and instructors valued a ready-to-use resource to incorporate reading comprehension in their pedagogy. This exploratory case study with Astrobites-based assignments suggests that incorporating current research in the undergraduate classroom through accessible literature summaries may increase students' confidence and ability to engage with research literature, assisting in their preparation for participation in research careers.

physics.ed-ph

Probabilistic Trust Intervals for Out of Distribution Detection

The ability of a deep learning network to distinguish between in-distribution (ID) and out-of-distribution (OOD) inputs is crucial for ensuring the reliability and trustworthiness of AI systems. Existing OOD detection methods often involve complex architectural innovations, such as ensemble models, which, while enhancing detection accuracy, significantly increase model complexity and training time. Other methods utilize surrogate samples to simulate OOD inputs, but these may not generalize well across different types of OOD data. In this paper, we propose a straightforward yet novel technique to enhance OOD detection in pre-trained networks without altering its original parameters. Our approach defines probabilistic trust intervals for each network weight, determined using in-distribution data. During inference, additional weight values are sampled, and the resulting disagreements among outputs are utilized for OOD detection. We propose a metric to quantify this disagreement and validate its effectiveness with empirical evidence. Our method significantly outperforms various baseline methods across multiple OOD datasets without requiring actual or surrogate OOD samples. We evaluate our approach on MNIST, Fashion-MNIST, CIFAR-10, CIFAR-100 and CIFAR-10-C (a corruption-augmented version of CIFAR-10), across various neural network architectures (e.g., VGG-16, ResNet-20, DenseNet-100). On the MNIST-FashionMNIST setup, our method achieves a False Positive Rate (FPR) of 12.46\% at 95\% True Positive Rate (TPR), compared to 27.09\% achieved by the best baseline. On adversarial and corrupted datasets such as CIFAR-10-C, our proposed method easily differentiate between clean and noisy inputs. These results demonstrate the robustness of our approach in identifying corrupted and adversarial inputs, all without requiring OOD samples during training.

cs.LG

Client Contribution Normalization for Enhanced Federated Learning

Mobile devices, including smartphones and laptops, generate decentralized and heterogeneous data, presenting significant challenges for traditional centralized machine learning models due to substantial communication costs and privacy risks. Federated Learning (FL) offers a promising alternative by enabling collaborative training of a global model across decentralized devices without data sharing. However, FL faces challenges due to statistical heterogeneity among clients, where non-independent and identically distributed (non-IID) data impedes model convergence and performance. This paper focuses on data-dependent heterogeneity in FL and proposes a novel approach leveraging mean latent representations extracted from locally trained models. The proposed method normalizes client contributions based on these representations, allowing the central server to estimate and adjust for heterogeneity during aggregation. This normalization enhances the global model's generalization and mitigates the limitations of conventional federated averaging methods. The main contributions include introducing a normalization scheme using mean latent representations to handle statistical heterogeneity in FL, demonstrating the seamless integration with existing FL algorithms to improve performance in non-IID settings, and validating the approach through extensive experiments on diverse datasets. Results show significant improvements in model accuracy and consistency across skewed distributions. Our experiments with six FL schemes: FedAvg, FedProx, FedBABU, FedNova, SCAFFOLD, and SGDM highlight the robustness of our approach. This research advances FL by providing a practical and computationally efficient solution for statistical heterogeneity, contributing to the development of more reliable and generalized machine learning models.

cs.LG

Distilling Calibrated Student from an Uncalibrated Teacher

Knowledge distillation is a common technique for improving the performance of a shallow student network by transferring information from a teacher network, which in general, is comparatively large and deep. These teacher networks are pre-trained and often uncalibrated, as no calibration technique is applied to the teacher model while training. Calibration of a network measures the probability of correctness for any of its predictions, which is critical in high-risk domains. In this paper, we study how to obtain a calibrated student from an uncalibrated teacher. Our approach relies on the fusion of the data-augmentation techniques, including but not limited to cutout, mixup, and CutMix, with knowledge distillation. We extend our approach beyond traditional knowledge distillation and find it suitable for Relational Knowledge Distillation and Contrastive Representation Distillation as well. The novelty of the work is that it provides a framework to distill a calibrated student from an uncalibrated teacher model without compromising the accuracy of the distilled student. We perform extensive experiments to validate our approach on various datasets, including CIFAR-10, CIFAR-100, CINIC-10 and TinyImageNet, and obtained calibrated student models. We also observe robust performance of our approach while evaluating it on corrupted CIFAR-100C data.

cs.CV

A comprehensive revisit of select Galileo/NIMS observations of Europa

The Galileo Near Infrared Mapping Spectrometer (NIMS) collected spectra of Europa in the 0.7-5.2 $μ$m wavelength region, which have been critical to improving our understanding of the surface composition of this moon. However, most of the work done to get constraints on abundances of species like water ice, hydrated sulfuric acid, hydrated salts and oxides have used proxy methods, such as absorption strength of spectral features or fitting a linear mixture of laboratory generated spectra. Such techniques neglect the effect of parameters degenerate with the abundances, such as the average grain-size of particles, or the porosity of the regolith. In this work we revisit three Galileo NIMS spectra, collected from observations of the trailing hemisphere of Europa, and use a Bayesian inference framework, with the Hapke reflectance model, to reassess Europa's surface composition. Our framework has several quantitative improvements relative to prior analyses: (1) simultaneous inclusion of amorphous and crystalline water ice, sulfuric-acid-octahydrate (SAO), CO$_2$, and SO$_2$; (2) physical parameters like regolith porosity and radiation-induced band-center shift; and (3) tools to quantify confidence in the presence of each species included in the model, constrain their parameters, and explore solution degeneracies. We find that SAO strongly dominates the composition in the spectra considered in this study, while both forms of water ice are detected at varying confidence levels. We find no evidence of either CO$_2$ or SO$_2$ in any of the spectra; we further show through a theoretical analysis that it is highly unlikely that these species are detectable in any 1-2.5 $μ$m Galileo NIMS data.

astro-ph.EP

Bayesian analysis of Juno/JIRAM's NIR observations of Europa

Juno spacecraft's spectrometer JIRAM recently observed the moon Europa in the 2-5 μm wavelength region. Here we present analysis of the average spectrum of a set of observations near 20°N and 40°W, focusing on the two forms of water-ice - amorphous and crystalline. We also take this as an opportunity to present a novel Bayesian spectral inversion framework for reflectance spectroscopy. We first validate this framework using simulated spectra of amorphous and crystalline ice mixtures and a laboratory spectrum of crystalline ice. We next analyze the JIRAM data and, through Bayesian model comparisons, find that a two-component intimately mixed model (TC-IM model) of amorphous and crystalline ice is strongly preferred (at 26σ confidence) over a two-component model of the same species but where their spectra are areally/linearly mixed. We also find that the TC-IM model is strongly preferred (at > 30σ confidence) over single-component models with only amorphous or crystalline ice, indicating the presence of both these phases of water ice in the data. For the highest SNR estimates of the JIRAM data, the TC-IM model solution corresponds to a mixture with a very large number density fraction (99.952 +/- 0.001 \%) of small (23.12 +/- 1.01 microns) amorphous ice grains, and a very small fraction (0.048 +/- 0.001 \%) of large (565.34 +/- 1.01 microns) crystalline ice grains. The overabundance of small amorphous ice grains we find is consistent with previous studies. The maximum-likelihood spectrum of the TC-IM model, however, is in tension with the data in the regions around 2.5 and 3.6 μm, and indicates the presence of non-ice components not currently included in our model, primarily due to the limited availability of cryogenic optical constants.

astro-ph.EP

Into the UV: The Atmosphere of the Hot Jupiter HAT-P-41b Revealed

For solar-system objects, ultraviolet spectroscopy has been critical in identifying sources for stratospheric heating and measuring the abundances of a variety of hydrocarbon and sulfur-bearing species, produced via photochemical mechanisms, as well as oxygen and ozone. To date, less than 20 exoplanets have been probed in this critical wavelength range (0.2-0.4 um). Here we use data from Hubble's newly implemented WFC3 UVIS G280 grism to probe the atmosphere of the hot Jupiter HAT-P-41b in the ultraviolet through optical in combination with observations at infrared wavelengths. We analyze and interpret HAT-P-41b's 0.2-5.0 um transmission spectrum using a broad range of methodologies including multiple treatments of data systematics as well as comparisons with atmospheric forward, cloud microphysical, and multiple atmospheric retrieval models. Although some analysis and interpretation methods favor the presence of clouds or potentially a combination of Na, VO, AlO, and CrH to explain the ultraviolet through optical portions of HAT-P-41b's transmission spectrum, we find that the presence of a significant H- opacity provides the most robust explanation. We obtain a constraint for the abundance of H-, log(H-) = -8.65 +/- 0.62 in HAT-P-41b's atmosphere, which is several orders of magnitude larger than predictions from equilibrium chemistry for a 1700 - 1950 K hot Jupiter. We show that a combination of photochemical and collisional processes on hot hydrogen-dominated exoplanets can readily supply the necessary amount of H- and suggest that such processes are at work in HAT-P-41b and many other hot Jupiter atmospheres.

astro-ph.EP

Disintegration of the Aged Open Cluster Berkeley 17

We present the analysis of the morphological shape of Berkeley 17, the oldest known open cluster (~10 Gyr), using a probabilistic star counting of Pan-STARRS point sources, and confirm its core-tail shape, plus an antitail, previously detected with the 2MASS data. The stellar population, as diagnosed by the color-magnitude diagram and theoretical isochrones, shows many massive members in the cluster core, whereas there is a paucity of such members in both tails. This manifests mass segregation in this aged star cluster with the low-mass members being stripped away from the system. It has been claimed that Berkeley 17 is associated with an excessive number of blue straggler candidates. Comparison of nearby reference fields indicates that about half of these may be field contamination.

astro-ph.SR