SearcharxivSearch

arXiv subjects

Suman Saha

Publications and source records attributed to Suman Saha.

At least 19 recordsLinked to original sources

NGTS clusters survey - VI: Stellar rotation in seven young open clusters within the PLATO LOPS2 field

We present NGTS measurements of rotation period distributions for FGKM stars in seven young open clusters spanning ~40--700 Myr within PLATO's first long-stare LOPS2 field. We measure 1063 rotation periods, of which 479 are newly analysed as part of cluster specific rotation studies, whilst 63 are unique periods not reported in the recent TESS All-Sky Rotation Survey. Of the 1063 rotation periods, 285 are identified as likely binary or higher order multiple systems using colour-magnitude diagrams and Gaia astrometry. These are the first comprehensive rotation period distributions for Trumpler 10, NGC 2451B and Alessi 3, while extending existing distributions for NGC 2451A, NGC 2516, Collinder 135 and IC 2391, to create a fuller picture on the rotation state of young stars in PLATO's LOPS2 field. We find that main-sequence solar-mass stars in the ~40 Myr old NGC 2451B cluster, form a slow sequence that can be distinguished from their counterparts at ~70--80 Myr, thereby significantly reducing the age at which young stellar groups can be relatively aged via their rotation sequences. We also observe stalled spin down from the age of NGC 2451A to at least that of NGC 2516 (~70--150 Myr) at masses $\gtrsim1$ $M_{\odot}$, supporting previous predictions that angular momentum redistribution and removal should result in a wave of stalled spin down that propagates as a function of both mass and age. Finally, we provide a new age estimate for Alessi 3 of 687$\pm$106 Myr using differential gyrochronology age dating.

astro-ph.SR

Tool-Guided Retrieval-Augmented Repair for Securing LLM-Generated C Code

Large language models can generate C code from natural-language descriptions, but resulting programs often contain security vulnerabilities and compilation errors, posing risks for embedded and resource-constrained systems. This work investigates how feedback and retrieval improve reliability of LLM-generated C code. We present an analysis-and-repair workflow that combines compilation diagnostics, CodeQL static analysis, and KLEE symbolic execution with retrieval of prior repair patterns for iterative refinement. Evaluated on 5,000 C programming tasks exercising embedded relevant vulnerabilities, baseline models show substantial reliability gaps, with compilation failure rates up to 46% and security defect rates up to 49%. Our approach improves both metrics. For CodeLlama 7B, security defect rates decrease from 49% to 19% and total CodeQL errors drop from 15,088 to 2,463 (83.7%). For DeepSeek Coder 1.3B, compilation failures are reduced from 42% to 22% and security defects from 35% to 15%. These results show that integrating lightweight analysis tools can improve the safety of LLM-generated code for embedded development.

cs.SE

NGTS-39 b: A 58 d transiting warm Jupiter in an eccentric orbit

We report the discovery and characterisation of NGTS-39 b (TIC 453147896 b), a warm Jupiter transiting a Sun-like star on a 58.2 day, eccentric (e = 0.386 +/- 0.019) orbit. NGTS-39 b was first identified from a TESS single-transit event, and subsequently confirmed with NGTS photometry and radial-velocity measurements from CORALIE and HARPS. The host star is a bright (Tmag = 11.02) F9 dwarf with an effective temperature of Teff = 6053 +67/-30 K. NGTS-39 b is a Jupiter-sized gas giant with a radius of 1.088 +/- 0.012 RJ and a mass of 1.467 +/- 0.081 MJ. Its equilibrium temperature is 519 +6/-5 K, placing it between short-period hot Jupiters and cold, Jupiter-like giants. The high orbital eccentricity and intermediate equilibrium temperature of NGTS-39 b make it a valuable test case for formation and migration models, particularly in the poorly sampled regime of long-period gas giants. The RV data show a linear trend of gamma dot = -17.75 m s^-1 yr^-1, which indicates the presence of an outer companion. The discovery of NGTS-39 b contributes to the small but growing population of transiting warm Jupiters with P > 50 days orbiting bright stars.

astro-ph.EP

The Longest-period Young Transiting Exoplanets. A Duo of Puffy Giants inside a Debris Disk

We identify two large-radius planets around the F-type star HD 114082 as the longest-period young transiting exoplanets known. From the first transit, detected by NASA's Transiting Exoplanet Survey Satellite (TESS), and a second dip, spotted by the Next-Generation Transit Survey (NGTS), we predicted mid-transit times for HD 114082 b (planet b). We pinpoint its orbit (period Pb= 225.5504$\pm$0.0004 days) from a third transit captured with the ESA's CHaracterising ExOplanet Satellite and the upgraded Antarctic Search for Transiting ExoPlanets telescope (ASTEP+), alongside orbit-discriminating observations. Another dimming partly covered by ASTEP+ completes the four-transit series. We support with dynamical evidence the planetary nature of a deeper transit detected with TESS and NGTS, identifying planet c. Additionally, we reexamine the debris disk, fitting its excess emission with two dust components. Fundamental stellar parameters are inferred from stellar evolution models, while a joint modeling of photometric and radial-velocity time series yields the planetary parameters, with masses further constrained using an N-body code. For planet b, the semimajor axis a$_b$= 0.791$\pm$0.008 au, eccentricity eb$\approx$ 0, inclination ib= 89.791$\pm$0.014 degrees, radius Rb= 1.046$\pm$0.014 R$_J$, and 95 % confidence upper limit on its mass M$_{95\%,b}$= 1.6 M$_J$. For planet c, a$_c$= 0.99$^{+0.03}_{-0.04}$ au, ec$\approx$ 0, i$_c$= 89.701$\pm$0.011 degrees, R$_c$= 1.36$\pm$0.03 R$_J$, and M$_{95\%,c}$= 2.0 M$_J$ (0.24 M$_J$ if adding transit timing variation constrains). They seem to be moderate-to-low-mass giants in nearly resonant, coplanar, circular orbits that formed in situ, or beyond the snowline, and migrated inwards, shaping the disk.

astro-ph.EP

Glossy Silicate Clouds on the Scorched Dayside of LTT9779b

Discovered deep within the "Neptunian desert", LTT9779b remains the only known ultra-hot Neptune, prompting significant speculation regarding its unique formation and evolutionary history. Its exceptionally high geometric albedo has previously been attributed either to the presence of clouds or to an extremely metal-rich atmosphere. Here, we present a comprehensive panchromatic analysis of its dayside atmosphere using JWST NIRISS and NIRSpec/G395H observations to characterize its atmospheric structure and composition. Leveraging the exceptional signal-to-noise ratio (S/N) in the observed spectra, we report a 3-to-5$\sigma$ detection of dayside clouds, with strong evidence for Mg$_2$SiO$_4$(s) (silicate) condensation. This constitutes the first statistically significant detection of clouds on the dayside of a Neptunian-mass exoplanet. We demonstrate that a highly reflective cloud deck, rather than an extremely high-metallicity atmosphere, is the most likely explanation for the planet's anomalously high optical albedo. Furthermore, our atmospheric retrievals yield robust detections of both CO ($\sim$4.88$\sigma$) and CO$_2$ ($\sim$8.76$\sigma$), while providing tentative constraints on the H$_2$O abundance and upper limits on SiO, TiO, and VO. Finally, our analysis places a robust constraint on the C/O ratio of 0.984 $\pm$ 0.019. This aligns LTT9779b with other known ultra-hot Jupiters exhibiting super-solar C/O ratios, suggesting a broader trend driven by the sequestration of oxygen-bearing condensates in ultra-hot atmospheres.

astro-ph.EP

A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration

Future planetary exploration envisions autonomous robotic agents operating under severe communication constraints, without global positioning, and with minimal human intervention. In such environments, agents must not only perceive and act, but also reason over mission objectives, operational constraints, and evolving environmental conditions. While prior work has largely focused on perception and control, the translation of high-level mission knowledge into structured, machine-interpretable representations remains underexplored. We introduce a pilot benchmark for translating natural language (NL) into First-Order Logic (FOL) within the domain of planetary exploration. The dataset is constructed from real mission documentation sourced from NASA's Planetary Data System (PDS), spanning missions from 2003 to 2013. These documents describe mission phases such as launch, boost, coast, cruise, and orbital operations in rich natural language. We manually annotate these documents with corresponding FOL representations that capture temporal structure, agent roles, and operational dependencies. In addition, we provide structured predicate vocabularies and typed constants to enable controlled experimentation with varying levels of prior knowledge. This pilot benchmark provides a foundation for research at the intersection of language understanding and formal reasoning, grounded in real-world, safety-critical mission data. The dataset is provided at: https://github.com/HaydenMM/planetary-logic-benchmark/blob/main/pilot_benchmark.json

cs.CL

Evaluating LLM-Generated Obfuscated XSS Payloads for Machine Learning-Based Detection

Cross-site scripting (XSS) remains a persistent web security vulnerability, especially because obfuscation can change the surface form of a malicious payload while preserving its behavior. These transformations make it difficult for traditional and machine learning-based detection systems to reliably identify attacks. Existing approaches for generating obfuscated payloads often emphasize syntactic diversity, but they do not always ensure that the generated samples remain behaviorally valid. This paper presents a structured pipeline for generating and evaluating obfuscated XSS payloads using large language models (LLMs). The pipeline combines deterministic transformation techniques with LLM-based generation and uses a browser- based runtime evaluation procedure to compare payload behavior in a controlled execution environment. This allows generated samples to be assessed through observable runtime behavior rather than syntactic similarity alone. In the evaluation, an untuned baseline language model achieves a runtime behavior match rate of 0.15, while fine-tuning on behavior-preserving source-target obfuscation pairs improves the match rate to 0.22. Although this represents a measurable improvement, the results show that current LLMs still struggle to generate obfuscations that preserve observed runtime behavior. A downstream classifier evaluation further shows that adding generated payloads does not improve detection performance in this setting, although behavior- filtered generated samples can be incorporated without materially degrading performance. Overall, the study demonstrates both the promise and the limits of applying generative models to adversarial security data generation and emphasizes the importance of runtime behavior checks in improving the quality of generated data for downstream detection systems.

cs.CR

High-Precision Photometry with a scientific CMOS Camera: II On-Sky Testing of the Marana camera at the NGTS facility

Modern scientific CMOS cameras offer very fast readout speeds and low read noise. In this study, we evaluate the performance of the Andor Marana CMOS camera through on-sky testing carried out at the NGTS facility at the ESO Paranal Observatory in Chile. We mount the Marana camera to an NGTS telescope, and conduct photometric observations of bright stars. In particular, we target transit events around eight known bright exoplanet host stars. Simultaneous observations are carried out using an existing Andor iKon-L CCD camera on a neighbouring NGTS telescope. This allows for a direct comparison of the photometric precision between the CMOS and CCD cameras. We find that the Marana CMOS exhibits a similar level of photometric performance to the CCD camera, achieving 500\,ppm at a 30-minute timescale for a T $=10$\,mag star. Although the CCD has a slightly better quantum efficiency over the NGTS filter range (520-890\,nm), we find that the faster readout speed of the CMOS compared to the CCD means that the CMOS camera detects 20\,\% more photons per unit time for a solar-type star in our standard 10\,s exposure time operation mode. This results in the CMOS performing slightly better photometry in the photon-limited regime. We conclude that modern CMOS cameras, such as the Marana, are very well-suited for astronomical time-series photometry applications.

astro-ph.IM

Cold giant discoveries from a joint radial-velocity and astrometry framework

The population of long-period giant planets shapes planetary system architectures and formation pathways, but these cold Jupiters remain relatively unexplored. Radial velocity (RV) surveys lose sensitivity at multi-AU separations, while transit surveys have poor detection probability at long periods. Absolute astrometry from the Hipparcos and Gaia missions offer an additional source for stellar motion that can break the orbital inclination degeneracy and strengthen detection confidence. This is especially timely ahead Gaia DR4/DR5, expected to enable routine astrometric vetting and true-mass measurements for long-period RV planets. Extending the Chile-Hertfordshire ExoPlanet Survey (CHEPS) by combining RVs spanning up to 16 years with absolute astrometry, we search for and characterise cold giants around metal-rich FGK stars. We upgrade the EMPEROR framework, incorporating astrometric differencing to jointly fit RVs and astrometry for five CHEPS targets, performing Bayesian model comparison and quantify the astrometric contribution. Our analysis characterises orbital parameters for two known planets in HIP 21850 and detects five new: a warm Jupiter--HIP 10090c, orbital period $P=321.8 \pm 0.5$ d and mass $M=0.85 \pm 0.08$ $M_J$, and four Jupiter analogues--HIP 8923b, with $P=14.1 \pm 0.06$ yr and $M=9.98\pm 0.47 M_J$, HIP 10090b with $P=8.1\pm 0.3$ yr and $M=3.87\pm 0.63$ $M_J$, HIP 39330b with $P=12.7\pm 0.7$ yr and $M=1.68\pm 0.15$ $M_J$, and HIP 98599b with $P=7.3\pm 0.1$ yr and $M=6.85\pm 0.16$ $M_J$. Adding astrometry reduces period and mass uncertainties by factors between 3 and 10 and increases the Bayes factor by up to 60. The synergy of long-baseline RVs and absolute astrometry provides a robust pathway to discover and characterise cold giant planets. Our results demonstrate that astrometry meaningfully improves detection confidence and converts minimum masses into true masses.

astro-ph.EP

TIC-65910228 b / NGTS-38 b, a 180 day transiting warm super-Jupiter

We present the discovery of TIC-65910228 b / NGTS-38 b, a giant exoplanet with a radius of $1.081\pm0.047$ R$_\text{J}$ and a mass of $4.78_{-0.37}^{+0.39}$ M$_\text{J}$ on a long-period ($180.52797\pm0.00036$ day), moderately eccentric ($e=0.3086\pm0.010$) orbit transiting a bright (V=$10.230\pm0.020$ mag) metal rich ([Fe/H]=$0.33\pm0.09$, 'dex') F6V-F7V type host star. The planet was initially detected from a single transit in TESS Sector 33. A photometric monitoring campaign of 228 nights with NGTS detected a transit egress of the planet, which together with spectroscopic radial velocity monitoring with CORALIE and HARPS identified an orbital period of ~180.5,d. These radial velocity measurements also showed the mass of the companion to be planetary. Additional transit observations coordinated by the TESS follow-up observing program allowed further confirmation and refinement of this period. With its relatively cool equilibrium temperature of $457\pm11$ K, NGTS-38 b joins a small but growing population of well characterised transiting warm-Jupiters and has one of the longest periods of any discovered to date. The target is situated in the LOPS2 field of the upcoming PLATO mission which will allow for greater refinement of the system parameters and potential for the discovery of additional companions too small and/or too long-period to be seen by TESS or NGTS. NGTS-38 b's bright host star and wide orbital separation make it an attractive target for further study, including potential measurement of its spin-orbit alignment or targeted exomoon/ring searches.

astro-ph.EP

Mass estimates of the young TOI-451 transiting planets: Multidimensional Gaussian Process on stellar spectroscopic and photometric signals

The young TOI-451 planetary system, aged 125 Myr, provides a unique opportunity to test theories of planetary internal structures and atmospheric mass loss through examination of its three transiting planets. We present an exhaustive photometric and spectroscopic follow-up to determine the orbital and physical properties of the system. We perform multidimensional Gaussian Process regression with the code pyaneti on spectroscopic time-series and NGTS/LCO light curves to disentangle the stellar and planetary signal in ESPRESSO radial velocities. We show how contemporaneous photometry serves as an activity indicator to inform RV modelling within a multidimensional Gaussian Processes framework. We argue that this can be exploited when spectroscopic observations are adversely affected by low signal-to-noise and/or poor sampling. We estimate the Doppler semi-amplitudes of Kb = 2.6(+1.1,-1.2) m/s, Kc = 1.2(+1.0,-0.8) m/s and Kd = 2.7 +/- 1.2 m/s. This translates into 2-sigma mass estimates for TOI-451 b and d of Mb = 4.7(+2.1,-2.2) Earth masses and Md = 10.2(+4.6,-4.5) Earth masses, as well as a mass upper limit for TOI-451 c of Mc < 11.5 Earth masses. The derived planetary properties suggest that planets c and d contain significant hydrogen-rich envelopes. The inferred parameters of TOI-451 b are consistent with either a rocky world that still retains a small hydrogen envelope or a water world. These insights make the TOI-451 system an ideal laboratory for future follow-up studies aimed at measuring atmospheric compositions, detecting atmospheric mass-loss signatures, and further exploring planetary formation and evolution processes.

astro-ph.EP

SecureCodeRL: Security-Aware Reinforcement Learning for Code Generation with Partial-Credit Rewards

Large Language Models (LLMs) can generate plausible code, but in settings that require exact stdin/stdout behavior they frequently produce programs that compile yet fail tests, and in some cases they introduce security-sensitive patterns. This paper presents SecureCodeRL, a reinforcement learning (RL) pipeline for security-aware code generation that optimizes a combined reward R = {\alpha}Rfunc + \b{eta}Rsec. The key idea is a partial-credit functional reward that assigns intermediate scores for syntactic validity, successful execution, and producing output, reducing reward sparsity that otherwise stalls learning on competitive programming style tasks. I evaluate supervised fine-tuning (SFT) and PPO variants on a small held-out prompt set from APPS+ and observe that PPO with partial credit (using a continued-training variant) improves syntax validity from 45% (SFT) to 60% and achieves the only non-zero test success signal in this pilot evaluation (5% at-least-one-test-pass), while remaining 100% clean under Bandit static analysis. Although Bandit findings were absent in this small evaluation, the security term is integrated into training to discourage insecure shortcuts when they appear.

cs.CR

Rule-Based Approaches to Atomic Sentence Extraction

Natural language often combines multiple ideas into complex sentences. Atomic sentence extraction, the task of decomposing complex sentences into simpler sentences that each express a single idea, improves performance in information retrieval, question answering, and automated reasoning systems. Previous work has formalized the "split-and-rephrase" task and established evaluation metrics, and machine learning approaches using large language models have improved extraction accuracy. However, these methods lack interpretability and provide limited insight into which linguistic structures cause extraction failures. Although some studies have explored dependency-based extraction of subject-verb-object triples and clauses, no principled analysis has examined which specific clause structures and dependencies lead to extraction difficulties. This study addresses this gap by analyzing how complex sentence structures, including relative clauses, adverbial clauses, coordination patterns, and passive constructions, affect the performance of rule-based atomic sentence extraction. Using the WikiSplit dataset, we implemented dependency-based extraction rules in spaCy, generated 100 gold=standard atomic sentence sets, and evaluated performance using ROUGE and BERTScore. The system achieved ROUGE-1 F1 = 0.6714, ROUGE-2 F1 = 0.478, ROUGE-L F1 = 0.650, and BERTScore F1 = 0.5898, indicating moderate-to-high lexical, structural, and semantic alignment. Challenging structures included relative clauses, appositions, coordinated predicates, adverbial clauses, and passive constructions. Overall, rule-based extraction is reasonably accurate but sensitive to syntactic complexity.

cs.CL

Improving LLM-Assisted Secure Code Generation through Retrieval-Augmented-Generation and Multi-Tool Feedback

Large Language Models (LLMs) can generate code but often introduce security vulnerabilities, logical inconsistencies, and compilation errors. Prior work demonstrates that LLMs benefit substantially from structured feedback, static analysis, retrieval augmentation, and execution-based refinement. We propose a retrieval-augmented, multi-tool repair workflow in which a single code-generating LLM iteratively refines its outputs using compiler diagnostics, CodeQL security scanning, and KLEE symbolic execution. A lightweight embedding model is used for semantic retrieval of previously successful repairs, providing security-focused examples that guide generation. Evaluated on a combined dataset of 3,242 programs generated by DeepSeek-Coder-1.3B and CodeLlama-7B, the system demonstrates significant improvements in robustness. For DeepSeek, security vulnerabilities were reduced by 96%. For the larger CodeLlama model, the critical security defect rate was decreased from 58.55% to 22.19%, highlighting the efficacy of tool-assisted self-repair even on "stubborn" models.

cs.CR

Evolution of Buffer Management in Database Systems: From Classical Algorithms to Machine Learning and Disaggregated Memory

Buffer management remains a critical component of database and operating system performance, serving as the primary mechanism for bridging the persistent latency gap between CPU processing speeds and storage access times. This paper provides a comprehensive survey of buffer management evolution spanning four decades of research. We systematically analyze the progression from foundational algorithms like LRU-K, 2Q, LIRS, and ARC to contemporary machine learning-augmented policies and disaggregated memory architectures. Our survey examines the historical OS-DBMS architectural divergence, production system implementations in PostgreSQL, Oracle, and Linux, and emerging trends including eBPF-based kernel extensibility, NVM-aware tiering strategies, and RDMA-enabled memory disaggregation. Through analysis of over 50 seminal papers from leading conferences (SIGMOD, VLDB, OSDI, FAST), we identify key architectural patterns, performance trade-offs, and open research challenges. We conclude by outlining a research direction that integrates machine learning with kernel extensibility mechanisms to enable adaptive, cross-layer buffer management for heterogeneous memory hierarchies in modern database systems.

cs.DB

A 43 day transiting Neptune and two 25 day Saturns from TESS, NGTS and ASTEP

Beyond orbital periods of 10 days, there is a dearth of known transiting gas giants. On longer orbits, planets are less affected by their host star, and become ideal probes of planet formation, migration and evolution. We report the discovery of a long period Neptune and two Saturns, each initially identified as single transits in the TESS photometry, and solved through additional transits from ground-based follow-up photometric observations by NGTS and ASTEP. High-resolution radial velocity mass measurements using CORALIE and HARPS confirm their planetary nature. From joint modelling of the photometric and spectroscopic data, we determine an orbital period of $43.12655_{-0.00017}^{+0.00012}~$days, radius of $3.65\pm0.22~\mathrm{R_{\oplus}}$, and mass of $19.1_{-4.5}^{+4.9}~\mathrm{M_{\oplus}}$ for NGTS-34b, making it one of the longest period well-characterized transiting Neptunes. Orbiting a late F-type star, bright in the K-band (Kmag$~\simeq7.9$), it is amenable for cool atmosphere studies using JWST or Ariel. TOI-4940b is a small Saturn on a $25.867811_{-0.000056}^{+0.000058}~$day orbit with a radius of $6.61\pm0.37~\mathrm{R_{\oplus}}$ and an upper mass limit $<89~\mathrm{M_{\oplus}}$. NGTS-35b(=TOI-6669b) is a larger Saturn on a $25.241192\pm0.000022~$day, moderately eccentric orbit ($e = 0.192_{-0.033}^{+0.037}$), with a radius of $10.90\pm0.65~\mathrm{R_{\oplus}}$ and a mass of $152_{-19}^{+22}~\mathrm{M_{\oplus}}$. With an assumed albedo $A=0.3$, each of these planets has an equilibrium temperature below 700K, with NGTS-35b especially cold at $450~$K. These three giants add to the small but growing population of long period planets that can further our understanding of planet formation mechanisms.

astro-ph.EP

Early-Warning Signals of Political Risk in Stablecoin Markets: Human and Algorithmic Behavior Around the 2024 U.S. Election

We study how the 2024 U.S. presidential election, viewed as a major political risk event, affected cryptocurrency markets by distinguishing human-driven peer-to-peer stablecoin transactions from automated algorithmic activity. Using structural break analysis, we find that human-driven Ethereum Request for Comment 20 (ERC-20) transactions shifted on November 3, two days before the election, while exchange trading volumes reacted only on Election Day. Automated smart-contract activity adjusted much later, with structural breaks appearing in January 2025. We validate these shifts using surrogate-based robustness tests. Complementary energy-spectrum analysis of Bitcoin and Ethereum identifies pronounced post-election turbulence, and a structural vector autoregression confirms a regime shift in stablecoin dynamics. Overall, human-driven stablecoin flows act as early-warning indicators of political stress, preceding both exchange behavior and algorithmic responses.

q-fin.ST

Can LLMs Recover Program Semantics? A Systematic Evaluation with Symbolic Execution

Obfuscation poses a persistent challenge for software engineering tasks such as program comprehension, maintenance, testing, and vulnerability detection. While compiler optimizations and third-party code often introduce transformations that obscure program intent, existing analysis tools and large language models (LLMs) struggle to recover the original semantics. In this work, we investigate whether LLMs, when fine-tuned with symbolic execution artifacts, can effectively deobfuscate programs and restore analyzability. We construct a benchmark by applying four widely studied transformations-control-flow flattening, opaque predicates, arithmetic encoding, and branch encoding-across diverse C programs from TUM Obfuscation Benchmarks, the LLVM test suite, and algorithmic repositories. We then compare three state-of-the-art LLMs under two training configurations: baseline fine-tuning on obfuscated/original code pairs, and enhanced fine-tuning with additional KLEE artifacts such as SMT constraints, path statistics, and test cases. Our evaluation examines syntactic correctness (compilation success), semantic fidelity (behavioral equivalence under symbolic execution), and code quality (readability and structure). Results show that GPT-4.1-mini achieves the strongest deobfuscation overall, and that incorporating KLEE artifacts consistently improves semantic preservation and compilation success across models. These findings highlight deobfuscation as a broader software engineering concern, demonstrating that combining LLMs with symbolic execution can strengthen automated testing, static analysis, and program comprehension in the presence of obfuscation.

cs.SE