SearcharxivSearch

arXiv subjects

Lu Feng

Publications and source records attributed to Lu Feng.

At least 19 recordsLinked to original sources

Drive the Thoughts: Runtime Monitoring of VLA Reasoning-Trajectory Consistency

Autonomous vehicles (AVs) operate in complex environments where failures are consequential. Sophisticated machine learning models for perception and planning are key to overcoming at least part of that complexity, but their black-box nature complicates validation and verification (V&V). The recent integration of Vision-Language-Action (VLA) models into AVs introduces a unique opportunity: besides generating trajectories, these models produce an explicit Chain-of-Thought (CoT) explaining their underlying rationale. This CoT provides a rich specification to cross-check model outputs and detect inconsistencies that may expose unsafe or unintended behavior. This paper assesses whether CoTs from a recent open driving VLA can support such monitoring. We curate DriveAlignBench, a specialized dataset from NVIDIA's Alpamayo 1.5 VLA for AVs containing 150 CoT-trajectory pairs, which we manually annotate for reliability, trajectory consistency, and safety. Our analysis reveals that 33.3% of CoTs are unreliable. Among reliable CoTs, the generated trajectory is consistent with the CoT in 74% of cases. Leveraging this potential, we propose integrating a CoT-trajectory consistency check into a runtime monitor. The check is nontrivial: CoTs express open-vocabulary, scene-relative driving commitments, while trajectories are low-level ego-motion sequences whose semantics depend on road geometry and motion context. To bridge this gap, we develop a family of automated consistency monitors. Our best monitor, lane-relative F-LLM with GPT-5.5, achieves F1 = 0.75, improving over the strongest raw-waypoint LLM baseline by +0.13 absolute F1 and over a rule-based monitor by +0.38. We release DriveAlignBench, the monitor implementations, and annotation tools at https://github.com/776styjsu/drive-the-thoughts.

cs.SE

Radio Activity Across Accretion State Changes in Changing-look AGNs: Insights from FIRST and VLASS over Two Decades

Changing-look active galactic nuclei (CL-AGNs) provide a unique opportunity to probe the coupling between accretion flows and relativistic jets in supermassive black holes. We investigate the long-term radio behavior of CL-AGNs over approximately 20 years by combining FIRST and VLASS observations with quasi-simultaneous optical spectroscopy and photometry. From a parent sample of 1092 CL-AGNs, we identify 58 sources with radio detections. Radio-detected CL-AGNs exhibit systematically higher radio kinetic efficiency, quantified by \(P_{\rm j}/L_{\rm bol}\), than both typical radio-detected AGNs and radio transients, consistent with their preference for low Eddington ratios. At the population level, the expected anti-correlation between radio emission and accretion rate is weak. However, a clear source-by-source anti-correlation emerges in a small subset of CL-AGNs with continuous multi-epoch coverage. We further identify four radio transients, including both radio turn-on and turn-off events, and one source exhibiting a multiwavelength flare that may be indicative of tidal disruption event-like activity. These results suggest that radio activity in CL-AGNs is not governed by instantaneous accretion state changes but is instead regulated by long-term accretion history and jet evolution, with additional stochastic or transient channels contributing in rare cases.

astro-ph.GA

$J$ and $H$ band sky brightness measurements from polar day to polar night at Dome A, Antarctica

The near-infrared (NIR) sky brightness is a fundamental parameter for evaluating the performance of ground-based infrared observatories. Dome~A on the Antarctic plateau offers exceptional atmospheric conditions, yet its NIR sky background has not been continuously monitored. We present the first continuous $J/H$-band measurements of the sky background at Dome~A from polar day to polar night, and characterize their median levels and temporal variability. The Antarctic Infrared Binocular Telescope (AIRBT), operating in the $J$ and $H$ bands, obtained continuous fixed-pointing observations from February to May 2024, which were used to measure the NIR sky background. The median sky brightness is $5.2/2.9$ and $15.3/13.4~\mathrm{mag~arcsec^{-2}}$ in $J/H$ bands during daytime and nighttime, respectively. The twilight--nighttime boundaries occur at solar elevations of $-9.3^\circ$ in $J$ and $-7.4^\circ$ in $H$. At the same solar elevation, the NIR sky background during the polar night is darker by about $0.1$ and $0.4~\mathrm{mag~arcsec^{-2}}$ in the $J$ and $H$ bands compared with the period of regular day--night alternation. During the polar-night period, the nighttime sky brightness in the $H$ band shows a more evident association with the sunspot number, while the corresponding trend in the $J$ band is weaker. These results reveal systematic differences in sky background between polar and non-polar environments and between polar night and regular day--night cycles. The measured sky brightness may be elevated, as the observations were conducted near solar maximum, highlighting the importance of long-term monitoring across the solar cycle.

astro-ph.IM

Neutrino mass constraints in interacting dark energy models after DESI DR2

Recent DESI observations indicate a deviation from the $\Lambda$CDM model, showing a preference for dynamical dark energy and thereby relaxing the upper limit on the neutrino mass within this framework. This deviation can also be explained by the presence of an interaction between dark energy and dark matter. In this work, we investigate the cosmological upper bounds on the total neutrino mass ($\sum m_{\nu}$) across four different interacting dark energy (IDE) models. The present analysis employs the latest DESI baryon acoustic oscillation, cosmic microwave background, and type Ia supernova datasets. These results demonstrate that the upper bounds on $\sum m_{\nu}$ exhibit profound sensitivity to the specific phenomenological formulation of the interaction term. While the I$\Lambda$CDM2 model ($Q \propto H \rho_{\mathrm{c}}$) substantially relaxes the stringent upper limit ($\sum m_{\nu} < 0.129$ eV at 95% confidence level), notably the I$\Lambda$CDM3 model ($Q \propto H_0 \rho_{\mathrm{de}}$), severely compresses the allowed parameter space, yielding a highly restrictive bound of $\sum m_{\nu} < 0.051$ eV. Furthermore, rigorous goodness-of-fit evaluations utilizing the Deviance Information Criterion and $\Delta\chi^2_{\mathrm{MAP}}$ indicate that the current observational data statistically favor these mass-suppressing IDE models. This establishes an exacerbated statistical tension between the observationally preferred IDE scenarios and the normal hierarchy lower bound ($\sim 0.06$ eV) determined by terrestrial neutrino oscillation experiments.

astro-ph.CO

Report on the Designing Accountable Software Systems Workshop

The Workshop on Designing Accountable Software Systems (DASS) was convened in November 2024 with support from the U.S. National Science Foundation to engage a wide range of current and future stakeholders from government, academia, and industry on the cross-disciplinary topic of accountability in software systems. Over two days, attendees engaged in a series of panels, invited talks, and breakout sessions covering: (1) the dimensions of accountability, including legal compliance as well as business and societal aspects and drivers; (2) a conceptual model of the various structures needed to realize accountability; (3) the sources of legal requirements that affect software; (4) the operationalization of legal requirements in software; (5) the requirements to preserve evidence needed to conduct investigations; and (6) a range of challenges and contextual factors beyond software that affect why some accountability structures succeed, while others fail. The workshop was conducted as a collaborative systematization of knowledge that culminated in several research directions. The findings include the importance of clarifying definitions and responsibilities within accountable organizations, which can affect whether those researching accountability are making assumptions that limit the generalizability of findings. Further research was also identified as needed to study the ways to improve the translation of accountability structures into the software design process while improving engagement with stakeholders, such as legislators, regulators, business executives and system developers. Finally, a key finding was the high demands that DASS-like research projects place on interdisciplinary teams: both in terms of team formation and sustainment, as well as, the specific demands of cross-disciplinary learning that covers both research methods, research dissemination, and career development.

cs.SE

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cost under a safety budget. Under out-of-distribution (OOD) deployment shifts, pretraining-only safe ICRL can give poor reward-safety tradeoffs because the remaining budget affects behavior only through frozen policy conditioning, not an explicit action-level check against predicted future cost. We propose a latent Q-Barrier shield that learns a context representation, latent dynamics, and an ensemble cost critic before deployment. Without parameter updates, the shield infers context from history and filters or softly reweights candidate actions using the remaining budget and predicted future cost. We prove a conditional, error-decomposed barrier-margin result: a Q-Barrier-satisfying action leaves the next latent-budget state with an approximately budget-safe continuation under the learned critic, up to Bellman and latent-prediction errors. Across five safe ICRL benchmarks, the shield improves deployment-time reward-safety tradeoffs over a strong safe-ICRL baseline: after a short context window, it achieves higher return in four of five benchmarks while matching or lowering average episode cost in all five.

cs.LG

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple perspectives, experiments fail and inform the next attempt, and lessons accumulate across cycles. Existing autonomous research systems often model this process as a linear pipeline: they rely on single-agent reasoning, stop when execution fails, and do not carry experience across runs. We present AutoResearchClaw, a multi-agent autonomous research pipeline built on five mechanisms: structured multi-agent debate for hypothesis generation and result analysis, a self-healing executor with a \textsc{Pivot}/\textsc{Refine} decision loop that transforms failures into information, verifiable result reporting that prevents fabricated numbers and hallucinated citations, human-in-the-loop collaboration with seven intervention modes spanning full autonomy to step-by-step oversight, and cross-run evolution that converts past mistakes into future safeguards. On ARC-Bench, a 25-topic experiment-stage benchmark, AutoResearchClaw outperforms AI Scientist v2 by 54.7%. A human-in-the-loop ablation across seven intervention modes reveals that precise, targeted collaboration at high-leverage decision points consistently outperforms both full autonomy and exhaustive step-by-step oversight. We position AutoResearchClaw as a research amplifier that augments rather than replaces human scientific judgment. Code is available at https://github.com/aiming-lab/AutoResearchClaw.

cs.AI

SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation

Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee safe execution. Many safety failures are temporal: a robot may touch a clean surface after contamination or release an object before it is fully inside an enclosure. We introduce SafeManip, a property-driven benchmark to explicitly evaluate temporal safety properties in robotic manipulation, moving beyond prior evaluations that largely focus on task completion or per-state constraint violations. SafeManip defines reusable safety templates over finite executions using Linear Temporal Logic over finite traces (LTLf). It maps observed rollouts to symbolic predicate traces and evaluates them with LTLf-based monitors. Its property suite covers eight manipulation safety categories: collision and contact safety, grasp stability, release stability, cross-contamination, action onset, mechanism recovery, object containment, and enclosure access. Templates can be instantiated with task-specific objects, fixtures, regions, or skills, allowing the same safety specifications to generalize across tasks and environments. We evaluate SafeManip on six vision-language-action policies, including $\pi_0$, $\pi_{0.5}$, GR00T, and their training variants, across 50 RoboCasa365 household tasks. Results show that even strong models often behave unsafely. Task-success gains do not reliably translate into safer execution: many successful rollouts remain unsafe, while longer-horizon or more complex tasks expose more violations. SafeManip provides a reusable evaluation layer for diagnosing temporal safety failures and measuring safe success beyond task completion.

cs.RO

Measuring neutrino mass in light of ACT DR6 and DESI DR2

The recent release of high-precision cosmological data, particularly the small-scale cosmic microwave background (CMB) measurements from ACT and baryon acoustic oscillation (BAO) data from DESI, has opened a new landscape for probing the neutrino mass. In this work, we present updated constraints on the total neutrino mass, $\sum m_\nu$, and its hierarchy within the $\Lambda$CDM, $w$CDM, holographic dark energy (HDE), and $w_0w_a$CDM models, using the latest ACT DR6, DESI DR2, and DESY5 datasets. We find that the upper limits on $\sum m_\nu$ are critically governed by the evolutionary behavior of the dark energy equation of state. Specifically, models exhibiting early-time quintessence features (e.g., HDE) yield the most stringent constraints, whereas those allowing for early-time phantom behavior (e.g., $w_0w_a$CDM) result in significantly looser bounds. Despite these model-dependent variations, we observe a robust hierarchy dependence across all scenarios, where the inverted hierarchy consistently yields weaker constraints and the degenerate hierarchy consistently yields tightest constraints. Our analysis demonstrates that the improved small-scale CMB information from ACT, combined with high-precision BAO data, systematically tightens the limits on $\sum m_\nu$, providing a crucial benchmark for future neutrino mass measurement.

astro-ph.CO

Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks

As AI systems increasingly support human decision making, a central challenge is determining what information helps people recognize when to rely on AI predictions and when to question or override them. Across three controlled human-subject studies spanning abstract visual reasoning with RAVEN matrices and deductive logical reasoning with LSAT problems, we examine how different forms of AI support affect human--AI team performance. A multi-stage reveal study shows that AI predictions and explanations can affect objective accuracy and subjective confidence differently. In visual reasoning, LLM explanations do not improve accuracy beyond the predicted answer alone, and no additional support format significantly outperforms prediction-only support; predicted probabilities show the highest descriptive accuracy and error recovery, while a derived selective-automation policy provides a higher-performing reference benchmark. In language-based logical reasoning, by contrast, LLM explanations yield the highest accuracy and error recovery, outperforming expert-written explanations and probability-based support. These results show that no single support strategy is universally effective. Human--AI interfaces should instead be designed to support calibrated reliance and effective error recovery by matching the form of assistance to the task and the evidence available to users.

cs.HC

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safety guarantees transfer to deployment under distribution shift, using diabetes management as a safety-critical testbed. We benchmark safe RL algorithms on a unified clinical simulator and reveal a safety generalization gap: policies satisfying constraints during training frequently violate safety requirements on unseen patients. We demonstrate that test-time shielding, which filters unsafe actions using learned dynamics models, effectively restores safety across algorithms and patient populations. Across eight safe RL algorithms, three diabetes types, and three age groups, shielding achieves Time-in-Range gains of 13--14\% for strong baselines such as PPO-Lag and CPO while reducing clinical risk index and glucose variability. Our simulator and benchmark provide a platform for studying safety under distribution shift in safety-critical control domains. Code is available at https://github.com/safe-autonomy-lab/GlucoSim and https://github.com/safe-autonomy-lab/GlucoAlg.

cs.LG

Antarctic Infrared Binocular Telescope: Early Data Release of observations in the 1.4 {\mu}m water-vapor-absorption band

Ground-based observations around 1.4 $\mu$m are normally limited by strong absorption of telluric water-vapor. However, Dome A, Antarctica has exceptionally dry conditions that offer a unique opportunity for observations in this band. We designed a new filter covering 1.34--1.48 $\mu$m, namely $W'$, and installed it on the Antarctic Infrared Binocular Telescope (AIRBT) at Dome A in 2025. AIRBT comprises two identical 15 cm optical tube assemblies and two InGaAs cameras equipped with $J$ and $W'$ filters, respectively. With this Early Data Release (EDR), we aim to evaluate the performance of the $W'$ band at Dome A to observe objects with water-vapor features. This EDR covers $\thicksim 20 \ \mathrm{deg^2}$ in the Galactic plane using $\thicksim 20,000$ images in three nights. For 2 s exposures, the 5 $\sigma$ limiting magnitude histogram peaks at $J \thicksim 11.5$ mag (Vega) and $W' \thicksim 9.9$ mag, respectively. The $J-W'$ vs $J-H$ color-color diagram distinguishes ultracool candidates with water-vapor-absorption features from reddened early type stars. Furthermore, later-type stars tend to exhibit stronger water-vapor absorption. Some sources show larger $\Delta W'$ than $\Delta J$ across the three nights, which we attribute to variations of their water-vapor-absorption depth. We conclude that it will be efficient to search for ultracool stars and estimate their spectral subtypes using $W'$ band imaging at Dome A, where the atmospheric transmission is high and stable.

astro-ph.IM

Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning

Vision-language agents have achieved remarkable progress in a variety of multimodal reasoning tasks; however, their learning remains constrained by the limitations of human-annotated supervision. Recent self-rewarding approaches attempt to overcome this constraint by allowing models to act as their own critics or reward providers. Yet, purely text-based self-evaluation struggles to verify complex visual reasoning steps and often suffers from evaluation hallucinations. To address these challenges, inspired by recent advances in tool-integrated reasoning, we propose Agent0-VL, a self-evolving vision-language agent that achieves continual improvement with tool-integrated reasoning. Agent0-VL incorporates tool usage not only into reasoning but also into self-evaluation and self-repair, enabling the model to introspect, verify, and refine its reasoning through evidence-grounded analysis. It unifies two synergistic roles within a single LVLM: a Solver that performs multi-turn tool-integrated reasoning, and a Verifier that generates structured feedback and fine-grained self-rewards through tool-grounded critique. These roles interact through a Self-Evolving Reasoning Cycle, where tool-based verification and reinforcement learning jointly align the reasoning and evaluation distributions for stable self-improvement. Through this zero-external-reward evolution, Agent0-VL aligns its reasoning and verification behaviors without any human annotation or external reward models, achieving continual self-improvement. Experiments on geometric problem solving and visual scientific analysis show that Agent0-VL achieves an 12.5% improvement over the base model. Our code is available at https://github.com/aiming-lab/Agent0.

cs.CV

Explaining Decentralized Multi-Agent Reinforcement Learning Policies

Multi-Agent Reinforcement Learning (MARL) has gained significant interest in recent years, enabling sequential decision-making across multiple agents in various domains. However, most existing explanation methods focus on centralized MARL, failing to address the uncertainty and nondeterminism inherent in decentralized settings. We propose methods to generate policy summarizations that capture task ordering and agent cooperation in decentralized MARL policies, along with query-based explanations for When, Why Not, and What types of user queries about specific agent behaviors. We evaluate our approach across four MARL domains and two decentralized MARL algorithms, demonstrating its generalizability and computational efficiency. User studies show that our summarizations and explanations significantly improve user question-answering performance and enhance subjective ratings on metrics such as understanding and satisfaction.

cs.AI

Galaxy clusters from the DESI Legacy Imaging Surveys -- III. Star-forming fraction of brightest cluster galaxies

This study investigates the evolution of the star-forming fraction ($F_{\mathrm{sf}}$) of Brightest Cluster Galaxies (BCGs) at $z<0.8$, using the galaxy clusters identified from the Legacy Imaging Surveys from the Dark Energy Spectroscopic Instrument (DESI). Star-forming galaxies are identified using the $g-z$ color, and $F_{\mathrm{sf}}$ is measured as a function of redshift, cluster halo mass, and galaxy stellar mass. Field galaxies are used as a comparison sample to reduce selection effects. For BCGs, $F_{\mathrm{sf}}$ increases with redshift, showing a slow rise below $z \sim 0.4 - 0.5$ and a more rapid increase above this range. In contrast, $F_{\mathrm{sf}}$ decreases with increasing cluster halo mass and BCG stellar mass. At the low stellar mass end, BCGs exhibit higher star-forming fractions than field galaxies, suggesting enhanced star formation likely fueled by cold gas accretion from the intracluster medium. Also, star-forming BCGs tend to show larger projected offsets from the optical cluster density peak than quenching BCGs, indicating ongoing assembly. The analysis of the specific star formation rate (sSFR) further indicates a transition in the dominant mechanism driving star formation in BCGs: cooling flows are likely responsible at low redshift, while gas-rich mergers play a greater role at higher redshift. The shift in dominance occurs around $z \sim 0.5$, aligning with the steep rise in $F_{\mathrm{sf}}$ of BCG.

astro-ph.GA

Optimization-Based Robust Permissive Synthesis for Interval MDPs

We present an optimization-based framework for robust permissive synthesis for Interval Markov Decision Processes (IMDPs), motivated by robotic decision-making under transition uncertainty. In many robotic systems, model inaccuracies and sensing noise lead to interval-valued transition probabilities. While robust IMDP synthesis typically yields a single policy and permissive synthesis assumes exact models, we show that robust permissive synthesis under interval uncertainty can be cast as a global mixed-integer linear program (MILP) that directly encodes robust Bellman constraints. The formulation maximizes a quantitative permissiveness metric (the number of enabled state-action pairs), while guaranteeing that every compliant strategy satisfies probabilistic reachability or expected reward specifications under all admissible transition realizations. To address the exponential complexity of vertex-based uncertainty representations, we derive a dualization-based encoding that eliminates explicit vertex enumeration and scales linearly with the number of successors. Experimental evaluation on four representative robotic benchmark domains demonstrates scalability to IMDPs with hundreds of thousands of states. The proposed framework provides a practical and general foundation for uncertainty-aware, flexibility-preserving controller synthesis in robotic systems.

cs.RO

Safe In-Context Reinforcement Learning

In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history. While ICRL has shown impressive generalization, safety during this adaptation process remains unexplored, limiting its applicability in real-world deployments where test-time behavior is expected to be safe. In this work, we propose SCARED: Safe Contextual Adaptive Reinforcement via Exact-penalty Dual, the first method that promotes safe adaptation of ICRL under the constrained Markov decision process framework. During the parameter-update-free adaptation process, our agent not only maximizes the reward but also keeps the accumulated cost within a user-specified safety budget. We also demonstrate that the agent actively reacts to the safety budget; with a higher safety budget, the agent behaves more aggressively, and with a lower safety budget the agent behaves more conservatively. Across challenging benchmarks, SCARED consistently enables safe and robust in-context adaptation, outperforming existing ICRL and safe meta-RL baselines.

cs.LG

Temperature and wind characteristics of Lenghu site for ventilation and structural design of large telescope enclosure

In recent years, a significant number of observatories and universities have been planning to construct optical and infrared telescopes at the Lenghu site in Qinghai Province due to the site's excellent seeing and clear night sky fraction. Although astronomical performances of the Lenghu site have been reported in detail by numerous papers, there were few reports showing statistics of temperature and wind characteristics in the traditional way required for the design of steel structures of large astronomical telescopes and enclosures, as well as the ventilation and air conditioning systems of these enclosures. This paper aims to present such new statistical data on temperature and wind conditions at the site, which could be helpful to inform and aid in such design decisions at the Lenghu site

astro-ph.IM