SearcharxivSearch

arXiv subjects

Michael Zhang

Publications and source records attributed to Michael Zhang.

At least 19 recordsLinked to original sources

JWST MIRI reveals a potential atmosphere on the ultra-hot rocky planet TOI-431b

An open question in exoplanet science is whether rocky exoplanets extremely close in to their host stars, `lava worlds', can retain significant atmospheres. Thermal emission observations via secondary eclipse can be used to determine whether a rocky exoplanet possesses an atmosphere. It is expected that an atmosphere could increase the planet's albedo and/or redistribute heat away from the dayside, reducing the secondary eclipse depth measured. Recent eclipse observations of lava planets, rocky planets hot enough to have liquid magma surfaces, have suggested the presence of atmospheres. Here, we present a single JWST MIRI/LRS partial secondary eclipse of the lava planet TOI-431b. We measure an eclipse depth of 81 $\pm$ 16 ppm, which corresponds to a brightness temperature of $1967^{+237}_{-253}$ K and a brightness temperature ratio R = $0.81\pm0.10$, being 1.2 $\sigma$ higher than the value reported by Spitzer. The observed brightness temperature ratio is 1.8$\sigma$ below that of a zero-albedo, zero heat redistribution bare rock (R=1). Given that magma pools are expected to have low albedos, we find that our results are best explained by the presence of an atmosphere. Future work to measure the exact composition of TOI-431b's atmosphere and improve models of lava planet atmospheres would better constrain the nature and evolution of lava planets.

astro-ph.EP

Helium escaping from the atmosphere of a nearby rocky exoplanet orbiting in a habitable zone

Observations of highly irradiated gas giant exoplanets have shown helium escaping from their atmospheres. There is limited evidence for atmospheres on rocky exoplanets, perhaps because they have already escaped. We report spectroscopic observations of LHS 1140b, a rocky exoplanet that orbits in the habitable zone of a nearby low-mass star. The near-infrared transit spectra show absorption by helium escaping from the planet's atmosphere. Helium absorption is detected in 2024 but not in 2025, indicating time-variable atmospheric escape. We interpret these results as indicating an upper atmosphere dominated by helium and depleted in hydrogen, with other volatile species trapped at lower altitudes, consistent with atmospheric fractionation models. No helium absorption is detected for LHS 1140c, a smaller and more heavily irradiated exoplanet in the same system.

astro-ph.EP

RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning

Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibility in open-ended and long-horizon tasks that require structured reasoning over evolving states. We introduce RoBoSR, an intermediate structural representation that formulates manipulation as step-wise state transitions over semantically grounded, object-centric scene graphs. By modeling object states and their spatial relations at the perception-action interface, RoBoSR disentangles high-level task reasoning from raw inputs and enables structured reasoning over preconditions, effects, and goal states. This representation endows the agent with causal reasoning capability, enforcing subtask dependencies and supporting coherent long-horizon task planning. To learn such structure-aware reasoning, we construct Manip-Cognition-1.6M, an open-world dataset that jointly supervises scene understanding, instruction interpretation, and subtask planning across diverse tasks. Across several benchmarks and real-world demonstrations, our method consistently outperforms prompting-based methods and classical TAMP baselines in zero-shot generalization and long-horizon tasks. The results underscore structured intermediate representations as a critical inductive bias for scalable embodied reasoning.

cs.RO

C, N, O, S, and photochemistry in a temperate giant planet orbiting a late M dwarf

We report the JWST NIRSpec/PRISM transit spectrum of TOI-6894b, an exceptional 420 K sub-Saturn that is the only known giant planet transiting a late M dwarf. Remarkably, both the light curve and the transit spectrum exhibit almost no stellar contamination. The spectrum is dominated by prominent absorption features from CH$_4$ and the photochemical product CS$_2$. For the first time in a transit spectrum, NH$_3$ is visually evident, while subtler features from H$_2$O, and CO$_2$ can also be seen. We significantly improve upon state-of-the-art photochemical reaction networks, and use our new network to run radiative-convective photochemical models at different metallicities. These models show that the spectrum--in particular the size of the NH$_3$ and CO$_2$ features relative to the CH$_4$ and H$_2$O features--is most consistent with a metallicity of 3--10$\times$ solar. Using a semi-free retrieval framework that perturbs the self-consistent model's abundance and temperature profiles to fit the data, we find that the planet's C/O, N/O, and S/O ratios are broadly consistent with solar values. A grid retrieval on 1D radiative-convective photochemical equilibrium (RCPE) models reveals a similar result: $[M/H]=0.46 \pm 0.08$ and C/O=$0.69 \pm 0.06$. The planet's atmospheric metallicity, abundance ratios, and bulk metal fraction are all strikingly similar to that of Jupiter, Saturn, and other gas giant exoplanets, despite orbiting a very low-mass star.

astro-ph.EP

Hydrogen airglow from an escaping ultrahot Jupiter atmosphere

Intense high-energy irradiation of close-in gaseous exoplanets drives the rapid escape of their atmospheres, fundamentally shaping planetary demographics. While atmospheric loss is routinely observed via transit absorption in atomic hydrogen, helium, and metal ions, the underlying physical properties, specifically the thermal structure, outflow dynamics, and mass-loss rate, remain poorly constrained due to inherent degeneracies in the transmission geometry. Here we report the first detection of atomic hydrogen emission from the escaping atmosphere of a gas giant. Using high-resolution spectroscopy of the ultrahot Jupiter KELT-9 b, we detect a hydrogen Balmer line (H{\alpha} 6564.6 {\AA}) emission signature originating from the planetary dayside. The emission line profile features a distinctive double-peaked shape with 0.1-0.15% peak amplitudes at +/-30 km/s and central self-absorption. This profile breaks transmission degeneracies, providing direct observational constraints on the vertical thermal structure, excited-state hydrogen populations, and wind dynamics in the upper atmosphere of KELT-9 b. Initial modeling reveals a vigorous outflow with a mass-loss rate above 10^{13} g/s, among the highest measured to date for gaseous exoplanets. Our results establish hydrogen airglow emission as a powerful diagnostic of atmospheric escape, opening a new observational window into the evolution of worlds in extreme radiation environments.

astro-ph.EP

Photochemical Production of CS2 in Temperate-to-Warm Gas Giant Exoplanet Atmospheres

Sulfur chemistry has emerged as an important probe of exoplanet atmospheres in the JWST era, although observational constraints have thus far been largely limited to SO2 and H2S in warm and hot exoplanets. Recent JWST observations have revealed CS2 in several cooler gas-giant exoplanets, yielding a new tracer of sulfur chemistry. However, the detailed chemical pathways responsible for the formation of CS2 remain poorly understood. Here, we use TOI-6894 b, a temperate gas giant with evidence for CS2, as a test case for one-dimensional photochemical kinetic-transport modeling and sensitivity analyses of CS2 chemistry. We show that CS2 is produced through coupled thermochemical and photochemical processes involving CH4 and H2S as the primary carbon and sulfur reservoirs, with S2 photolysis driving disequilibrium sulfur chemistry. Our models provide a physically consistent explanation for the observed CS2 feature in TOI-6894 b. Extending our analysis to gas giant exoplanets spanning a wide range of Teq, we find that CS2 abundance peaks in temperate to warm atmospheres (Teq ~ 500 - 700 K), and declines toward both lower and higher temperatures. This temperature dependence provides a unified framework for interpreting current CS2 observations, accounting for reported detections in temperate to warm planets and the lack of detections in colder and hotter giant exoplanets. Our results establish CS2 as a complementary probe of sulfur inventories and atmospheric metallicity in cool gas giant exoplanets

astro-ph.EP

Revisiting the Exo-Mercury Candidate GJ 367 b with ESPRESSO and a Self-Consistent Tidal Distortion Model

We report revised mass and radius measurements for GJ 367 b, an ultra-short-period (7.7 hr) sub-Earth in a multi-planet system orbiting a nearby (~9 pc) M dwarf host. Previous mass and radius measurements have suggested GJ 367 b has an anomalously high bulk density, close to that of solid iron. The existence of such an iron-rich planet is in tension with established planet formation scenarios. We utilized newly available TESS short-cadence photometry to constrain the radius of GJ 367 b to 0.736 +/- 0.035 R_Earth. We consider observational and modeling effects such as photometric dilution, stellar activity, and tidal distortion to account for possible inaccuracies in the star and planet radius measurements. From our radial velocity (RV) analysis using VLT/ESPRESSO data covering nearly the full orbit in a single night, we find a mass of 0.503 +/- 0.078 M_Earth, corresponding to a bulk density of 6.9 +1.6/-1.4 g cm-1. We present a new tidal distortion and interior composition modeling framework to assess the iron mass fraction of GJ 367 b. Considering several different interior composition assumptions and radial aspect ratios, we find an iron fraction of ~50-70%, which is broadly consistent with that of Mercury and not as iron rich as previously suggested.

astro-ph.EP

The role of the Hubble Space Telescope in advancing our understanding of atmospheric escape in exoplanets

An important evolutionary pathway for planetary atmospheres is escape to space, which has been studied on Earth and Mars for several decades and more recently in exoplanets. A particularly important regime is the hydrodynamic escape, wherein atmospheric mass escapes the planet at high rates in a collisional fluid outflow. This process is used to partly explain the early evolution of rocky planets in and out of the Solar System, as well as key aspects of exoplanet demographics. Hydrodynamic escape is not occurring in the Solar System planets, so our only option for such observations is through exoplanets. The ultraviolet (UV) capabilities of the Hubble Space Telescope (HST) are fundamental to detect hydrodynamic escape and measure the resulting mass-loss rates for a range of planetary systems and to identify targets for surveys with the Habitable Worlds Observatory. We discuss here what kinds of observations and instrument modes are necessary to continue studying atmospheric escape in exoplanets for the next decade, as well as how to advance our understanding of planetary evolution and habitability.

astro-ph.IM

Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

Steering large language models (LLMs) is usually done by either instruction prompting or activation steering. Prompting often gives strong control, but caches guidance tokens at every layer and can clutter long interactions; activation steering is compact but typically weaker and does not support large structured reminders. We introduce memory inception (MI), a training-free method that steers in latent attention space by inserting text-derived key-value (KV) banks only at selected layers. Rather than materializing reminder content throughout the prompt cache, MI treats steering as selective KV allocation, injecting latent slots only where the model routes to them. On matched personality-steering tasks, MI gives the best overall control--drift trade-off, remaining competitive with prompting while consistently outperforming CAA. On updateable guidance, MI supports mid-conversation behavior shifts without rewriting the visible transcript, achieving the highest post-shift alignment on Qwen3. On structured reasoning, MI outperforms visible prompting on HARDMath and PHYSICS (10/12 subject$\times$mode cells), serving as proxies for structured reasoning in verifiable domains, while cutting content-matched KV storage by up to 118$\times$. These results position MI as a powerful steering method when guidance is persistent, structured, or expensive to keep in the visible transcript.

cs.LG

Democratizing the medieval English legal tradition

The record of the beginning of the most widespread legal system in the world is contained in millions of pages of handwritten text. Most of the records of the first centuries of the Anglo-American legal system are hand-written in a highly abbreviated form of medieval Latin which only a few dozen scholars in the world are trained to read. In this interdisciplinary project, we construct a dataset of 4029 lines of text across 193 medieval criminal and civil cases. We then use the dataset to train an open-source end-to-end pipeline for transcribing these manuscripts. We first train standard neural network architectures for line segmentation and handwriting recognition (R-Blla and CNN+LSTM with CTC decoding, respectively) and show that they can already achieve 79% word accuracy, despite the relatively small training set and the challenge of expanding abbreviations. We then demonstrate that simple post-processing significantly boosts accuracy: adding an n-gram language model to the CTC decoder improves word accuracy to 82%, while asking Gemini Pro 3 to correct mistakes boosts accuracy to 88%. Finally, we compare the CNN+LSTM architecture with TrOCR, a transformer-based OCR architecture, demonstrating that TrOCR shows comparable word accuracy but worse character accuracy due to its over-willingness to guess, making it harder for humans to infer the correct reading. We incorporated our pipeline into a web portal (glyphmachina.com), opening up the English legal tradition to legal scholars, medievalists, and students.

cs.CV

Real-Time Cellist Postural Evaluation With On-Device Computer Vision

Posture is a critical factor for beginning instrumental learners. Most students receive instruction only once a week, and during the intervals between lessons they have little or no feedback on their physical posture. As a result, posture often deteriorates, increasing the risk of musculoskeletal injury and inefficient technique. Recent advances in computer vision and machine learning make it possible to evaluate posture without the constant presence of a human expert. However, current solutions have been extremely limited in availability and convenience due to their reliance on computationally expensive hardware or multi-sensor setups. We present Cello Evaluator, a real-time postural feedback system for practicing cellists. Through this optimization for on-device computer vision inference, we provide access to cellist postural evaluation to anyone with a current generation Android phone and thus reduces the postural feedback voids within individual practice. To validate our mobile application, we conduct a heuristic evaluation consisting of cellist and UX experts. Overall feedback from the evaluation found the app to be user friendly and helpful.

cs.HC

Evidence for an Atmosphere on the Ultra-Short Period super-Earth HD 3167 b

'Lava worlds'-Earth-sized planets hot enough (Teq >~ 1100 K) to melt their dayside silicate surfaces-have emerged as promising candidates for atmospheric detection and characterization. Thermal emission observations show an apparent dichotomy: the hottest lava worlds have colder daysides than the temperature of a maximally emitting bare rock, indicating the likely presence of thick and/or reflective atmospheres while the coldest ones do not. However, where in instellation flux this potential bifurcation occurs is uncertain. We present a JWST MIRI LRS eclipse of the ultra-short period (USP) lava world HD 3167 b (Teq = 1786 K, R = 1.6 Rearth, P = 0.96 d) that helps bridge this gap. We measure the white light eclipse depth to be 38 +/- 11 ppm, more than 5 sigma lower than the expected eclipse depth of a dark, maximally hot bare rock. We use this to derive a dayside brightness temperature that is best explained by the presence of an atmosphere that cools the dayside by reflecting incoming starlight and/or efficiently redistributing heat to the planet's nightside. An atmosphere is further compatible with the planet's slight under-density compared to an Earth-like composition. The corresponding dayside emission spectrum is not precise enough to constrain atmospheric composition, motivating follow-up spectroscopic observations with JWST NIRSpec. Lastly, we use our observation and existing data to refine key planetary parameters of the HD 3167 system. HD 3167 b is currently the least irradiated USP super-Earth with evidence for an atmosphere.

astro-ph.EP

OMNI-PoseX: A Fast Vision Model for 6D Object Pose Estimation in Embodied Tasks

Accurate 6D object pose estimation is a fundamental capability for embodied agents, yet remains highly challenging in open-world environments. Many existing methods often rely on closed-set assumptions or geometry-agnostic regression schemes, limiting their generalization, stability, and real-time applicability in robotic systems. We present OMNI-PoseX, a vision foundation model that introduces a novel network architecture unifying open-vocabulary perception with an SO(3)-aware reflected flow matching pose predictor. The architecture decouples object-level understanding from geometry-consistent rotation inference, and employs a lightweight multi-modal fusion strategy that conditions rotation-sensitive geometric features on compact semantic embeddings, enabling efficient and stable 6D pose estimation. To enhance robustness and generalization, the model is trained on large-scale 6D pose datasets, leveraging broad object diversity, viewpoint variation, and scene complexity to build a scalable open-world pose backbone. Comprehensive evaluations across benchmark pose estimation, ablation studies, zero-shot generalization, and system-level robotic grasping integration demonstrate the effectiveness of OMNI-PoseX. The OMNI-PoseX achieves SOTA pose accuracy and real-time efficiency, while delivering geometrically consistent predictions that enable reliable grasping of diverse, previously unseen objects.

cs.RO

PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs

Retrieval-augmented language models can retrieve relevant evidence yet still commit to answers before explicitly checking whether the retrieved context supports the conclusion. We present PAVE (Premise-Grounded Answer Validation and Editing), an inference-time validation layer for evidence-grounded question answering. PAVE decomposes retrieved context into question-conditioned atomic facts, drafts an answer, scores how well that draft is supported by the extracted premises, and revises low-support outputs before finalization. The resulting trace makes answer commitment auditable at the level of explicit premises, support scores, and revision decisions. In controlled ablations with a fixed retriever and backbone, PAVE outperforms simpler post-retrieval baselines in two evidence-grounded QA settings, with the largest gain reaching 32.7 accuracy points on a span-grounded benchmark. We view these findings as proof-of-concept evidence that explicit premise extraction plus support-gated revision can strengthen evidence-grounded consistency in retrieval-augmented LLM systems.

cs.CL

GSR: Learning Structured Reasoning for Embodied Manipulation

Despite rapid progress, embodied agents still struggle with long-horizon manipulation that requires maintaining spatial consistency, causal dependencies, and goal constraints. A key limitation of existing approaches is that task reasoning is implicitly embedded in high-dimensional latent representations, making it challenging to separate task structure from perceptual variability. We introduce Grounded Scene-graph Reasoning (GSR), a structured reasoning paradigm that explicitly models world-state evolution as transitions over semantically grounded scene graphs. By reasoning step-wise over object states and spatial relations, rather than directly mapping perception to actions, GSR enables explicit reasoning about action preconditions, consequences, and goal satisfaction in a physically grounded space. To support learning such reasoning, we construct Manip-Cognition-1.6M, a large-scale dataset that jointly supervises world understanding, action planning, and goal interpretation. Extensive evaluations across RLBench, LIBERO, GSR-benchmark, and real-world robotic tasks show that GSR significantly improves zero-shot generalization and long-horizon task completion over prompting-based baselines. These results highlight explicit world-state representations as a key inductive bias for scalable embodied reasoning.

cs.RO

Self-Interest and Systemic Benefits: Emergence of Collective Rationality in Mixed Autonomy Traffic Through Deep Reinforcement Learning

Autonomous vehicles (AVs) are expected to be commercially available in the near future, leading to mixed autonomy traffic consisting of both AVs and human-driven vehicles (HVs). Although numerous studies have shown that AVs can be deployed to benefit the overall traffic system performance by incorporating system-level goals into their decision making, it is not clear whether the benefits still exist when agents act out of self-interest -- a trait common to all driving agents, both human and autonomous. This study aims to understand whether self-interested AVs can bring benefits to all driving agents in mixed autonomy traffic systems. The research is centered on the concept of collective rationality (CR). This concept, originating from game theory and behavioral economics, means that driving agents may cooperate collectively even when pursuing individual interests. Our recent research has proven the existence of CR in an analytical game-theoretical model and empirically in mixed human-driven traffic. In this paper, we demonstrate that CR can be attained among driving agents trained using deep reinforcement learning (DRL) with a simple reward design. We examine the extent to which self-interested traffic agents can achieve CR without directly incorporating system-level objectives. Results show that CR consistently emerges in various scenarios, which indicates the robustness of this property. We also postulate a mechanism to explain the emergence of CR in the microscopic and dynamic environment and verify it based on simulation evidence. This research suggests the possibility of leveraging advanced learning methods (such as federated learning) to achieve collective cooperation among self-interested driving agents in mixed-autonomy systems.

cs.LG

Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow

As software systems grow in scale and complexity, understanding the distribution of programming language topics within source code becomes increasingly important for guiding technical decisions, improving onboarding, and informing tooling and education. This paper presents the design, implementation, and evaluation of a novel programming language topic classification workflow. Our approach combines a multi-label Support Vector Machine (SVM) with a sliding window and voting strategy to enable fine-grained localization of core language concepts such as operator overloading, virtual functions, inheritance, and templates. Trained on the IBM Project CodeNet dataset, our model achieves an average F1 score of 0.90 across topics and 0.75 in code-topic highlight. Our findings contribute empirical insights and a reusable pipeline for researchers and practitioners interested in code analysis and data-driven software engineering.

cs.SE

Reverse Engineering User Stories from Code using Large Language Models

User stories are essential in agile development, yet often missing or outdated in legacy and poorly documented systems. We investigate whether large language models (LLMs) can automatically recover user stories directly from source code and how prompt design impacts output quality. Using 1,750 annotated C++ snippets of varying complexity, we evaluate five state-of-the-art LLMs across six prompting strategies. Results show that all models achieve, on average, an F1 score of 0.8 for code up to 200 NLOC. Our findings show that a single illustrative example enables the smallest model (8B) to match the performance of a much larger 70B model. In contrast, structured reasoning via Chain-of-Thought offers only marginal gains, primarily for larger models.

cs.SE