SearcharxivSearch

arXiv subjects

Jianxiang Liu

Publications and source records attributed to Jianxiang Liu.

11 recordsLinked to original sources

Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

Ensuring reliability in uncertain environments remains difficult for long-horizon robotic manipulation. End-to-end VLA models are data-heavy and opaque, making diagnosis and verification difficult. Hierarchical pipelines are more interpretable, but their plans are often weakly grounded in observations, weakly aligned with low-level actions, and computed without online feedback, leading to open-loop behavior and hallucinations. To address these issues, we introduce the Triplet-to-Track System (TTS), a closed-loop long-horizon imitation learning system that uses human videos to reduce reliance on robot-collected data. TTS represents high-level subgoals as instance-grounded triplets, translates them into continuous track priors for execution, and monitors task progress from observations for online replanning. Across diverse real-world long-horizon tasks, TTS achieves a 74.8\% average success rate and supports object-level and compositional generalization.

cs.RO

Lyman-$α$ forest constraints on pure and mixed fuzzy dark matter

Fuzzy dark matter (FDM), often realized as an ultralight scalar field, can suppress the growth of small-scale structures and could be strictly tested with Lyman-$α$ forest measurements. In this work, we constrain both pure and mixed FDM models (PFDM and MFDM) using measurements of the one-dimensional (1D) Lyman-$α$ forest flux power spectrum at $z=5.0$, 4.6, and 4.2. We perform cosmological hydrodynamical simulations with modified initial conditions and construct a two-stage neural network emulator for accurate analysis. The first stage predicts the cold dark matter (CDM) 1D flux power spectrum, while the second stage predicts the MFDM effect relative to the CDM baseline. This construction improves the sensitivity to weak FDM effects, enforces the correct CDM limit, and enables robust interpolation across a broad range of FDM masses and fractions. After marginalizing over the intergalactic medium parameters, we obtain the FDM mass $m_{\mathrm{FDM}}>1.9\times10^{-21}~\mathrm{eV}$ at 95\% credible level for the PFDM model. For the MFDM model, we find the FDM fraction of dark matter $f_{\mathrm{FDM}}<0.07$, $0.12$, and $0.65$ at 95\% credible level for $\log_{10}(m_{\mathrm{FDM}}/\mathrm{eV})=-23.0$, $-22.0$, and $-21.0$, respectively. When $\log_{10}(m_{\mathrm{FDM}}/\mathrm{eV})\gtrsim -20$, the current data do not provide an effective upper limit on $f_{\mathrm{FDM}}$.

astro-ph.CO

Probing Dark Matter Substructure with Image Number Anomaly in Strong Lensing Systems

Gravitational lensing observables, including anomalies in image positions, flux ratios, and time delays, serve as usual probes of dark matter (DM) substructure. When dark matter substructure possesses sufficient perturbations, it may lead to the formation of extra images in otherwise canonical doubly or quadruply imaged systems. With the advent of increasingly precise observational instruments, previously undetectable images may become measurable and image number anomalies therefore could be an increasingly viable method. In this paper, we utilize the gravitational lensing phenomenon of image number anomaly to derive constraints on dark matter substructure. We present the extra images induced by distinct forms of DM substructure, specifically primordial black holes (PBHs) and fuzzy dark matter (FDM) and show that higher angular resolution observations increase the probability of detecting additional lensed images. Based on a null detection of image number anomalies in a sample of 3500 lens systems generated from the \textit{Strong Lensing Halo model-based mock catalogs} (SL-Hammocks), we derive upper limits on the abundance of PBHs. At the 95\% confidence level, the PBH abundance is constrained to $\lesssim 0.125\%$, $0.08\%$, and $0.04\%$ for PBH masses in the range $\sim 10^{7}$--$10^{9}~M_{\odot}$, corresponding to angular resolutions of $0.1''$, $0.05''$, and $0.01''$, respectively. Similarly, we exclude particle masses below $0.4$, $0.6$, and $3.5 \times 10^{-22} \ \mathrm{eV}$ for FDM at the same confidence level for the respective resolutions. Furthermore, the abundance of PBHs $\lesssim 0.9\%$ could be constrained at an angular resolution of $0.5''$ for the Legacy Survey of Space and Time (LSST) Observations. Finally, we discuss methodologies for identifying image number anomalies in special cases and demonstrate feasibility using a fitting procedure.

astro-ph.CO

BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances

Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations to executable robot behavior remains challenging. Prior work either relies on robot data for downstream adaptation or learns affordance representations that remain at the perception level and do not directly support real-world execution. We present BridgeACT, an affordance-driven framework that learns robotic manipulation directly from human videos without requiring any robot demonstration data. Our key idea is to model affordance as an embodiment-agnostic intermediate representation that bridges human demonstrations and robot actions. BridgeACT decomposes manipulation into two complementary problems: where to grasp and how to move. To this end, BridgeACT first grounds task-relevant affordance regions in the current scene, and then predicts task-conditioned 3D motion affordances from human demonstrations. The resulting affordances are mapped to robot actions through a grasping module and a lightweight closed-loop motion controller, enabling direct deployment on real robots. In addition, we represent complex manipulation tasks as compositions of affordance operations, which allows a unified treatment of diverse tasks and object-to-object interactions. Experiments on real-world manipulation tasks show that BridgeACT outperforms prior baselines and generalizes to unseen objects, scenes, and viewpoints.

cs.RO

Joint Constraints on Fuzzy and Warm Dark Matter from Satellite Populations of the Milky Way and Andromeda

We perform a joint analysis of the Milky Way (MW) and Andromeda (M31) satellite populations to constrain the properties of fuzzy dark matter (FDM) and thermal-relic warm dark matter (WDM). We combine MW satellite observations from the Dark Energy Survey (DES) and Pan-STARRS1 (PS1) with M31 satellite data from the Pan-Andromeda Archaeological Survey (PAndAS), and model the corresponding observable satellite populations using the empirical galaxy--halo connection model described in Nadler et al. (2020) together with the appropriate selection functions. Uncertainties in the virial masses of the MW and M31 are incorporated through host-mass priors that linearly scale the relevant model parameters, allowing us to infer the full posterior distributions of all parameters. For the FDM case, we obtain $m_{\mathrm{FDM}} > 1.74 \times 10^{-20}~\mathrm{eV}$ (95\% CL) and $m_{\mathrm{FDM}} > 1.42 \times 10^{-20}~\mathrm{eV}$ (20:1 posterior ratio). For thermal-relic WDM, we find $m_{\mathrm{WDM}} > 6.20~\mathrm{keV}$ (95\% CL) and $m_{\mathrm{WDM}} > 5.75~\mathrm{keV}$ (20:1 posterior ratio). These results represent a moderate improvement over MW-only constraints, and provide the strongest constraints to date on the FDM and WDM derived from satellite galaxy populations in the Local Group.

astro-ph.CO

Preference-Guided Reinforcement Learning for Efficient Exploration

In this paper, we investigate preference-based reinforcement learning (PbRL), which enables reinforcement learning (RL) agents to learn from human feedback. This is particularly valuable when defining a fine-grain reward function is not feasible. However, this approach is inefficient and impractical for promoting deep exploration in hard-exploration tasks with long horizons and sparse rewards. To tackle this issue, we introduce LOPE: \textbf{L}earning \textbf{O}nline with trajectory \textbf{P}reference guidanc\textbf{E}, an end-to-end preference-guided RL framework that enhances exploration efficiency in hard-exploration tasks. Our intuition is that LOPE directly adjusts the focus of online exploration by considering human feedback as guidance, thereby avoiding the need to learn a separate reward model from preferences. Specifically, LOPE includes a two-step sequential policy optimization technique consisting of trust-region-based policy improvement and preference guidance steps. We reformulate preference guidance as a trajectory-wise state marginal matching problem that minimizes the maximum mean discrepancy distance between the preferred trajectories and the learned policy. Furthermore, we provide a theoretical analysis to characterize the performance improvement bound and evaluate the effectiveness of the LOPE. When assessed in various challenging hard-exploration environments, LOPE outperforms several state-of-the-art methods in terms of convergence rate and overall performance.The code used in this study is available at https://github.com/buaawgj/LOPE.

cs.LG

Detecting dark matter substructure with lensed quasars in optical bands

Flux ratios of multiple images in strong gravitational lensing systems provide a powerful probe of dark matter substructure. Optical flux ratios of lensed quasars are typically affected by stellar microlensing, and thus studies of dark matter substructure often rely on emission regions that are sufficiently extended to avoid microlensing effects. To expand the accessible wavelength range for studying dark matter substructure through flux ratios and to reduce reliance on specific instruments, we confront the challenges posed by microlensing and propose a method to detect dark matter substructure using optical flux ratios of lensed quasars. We select 100 strong lensing systems consisting of 90 doubles and 10 quads to represent the overall population and adopt the Kolmogorov--Smirnov (KS) test as our statistical method. By introducing different types of dark matter substructure into these strong lensing systems, we demonstrate that using quads alone provides the strongest constraints on dark matter and that several tens to a few hundred independent flux ratio measurements from quads can be used to study the properties of dark matter substructure and place constraints on dark matter parameters. Furthermore, we suggest that the use of multi-band flux ratios can substantially reduce the required number of quads. Such sample sizes will be readily available from ongoing and upcoming wide-field surveys.

astro-ph.CO

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the $Q$-value overestimation for out-of-distribution (OOD) actions. Existing methods address this issue by imposing constraints; however, they often become overly conservative when evaluating OOD regions, which constrains the $Q$-function generalization. This over-constraint issue results in poor $Q$-value estimation and hinders policy improvement. In this paper, we introduce a novel approach to achieve better $Q$-value estimation by enhancing $Q$-function generalization in OOD regions within Convex Hull and its Neighborhood (CHN). Under the safety generalization guarantees of the CHN, we propose the Smooth Bellman Operator (SBO), which updates OOD $Q$-values by smoothing them with neighboring in-sample $Q$-values. We theoretically show that SBO approximates true $Q$-values for both in-sample and OOD actions within the CHN. Our practical algorithm, Smooth Q-function OOD Generalization (SQOG), empirically alleviates the over-constraint issue, achieving near-accurate $Q$-value estimation. On the D4RL benchmarks, SQOG outperforms existing state-of-the-art methods in both performance and computational efficiency.

cs.LG

Time Delay Anomalies of Fuzzy Gravitational Lenses

Fuzzy dark matter is a promising alternative to the standard cold dark matter. It has quite recently been noticed, that they can not only successfully explain the large-scale structure in the Universe, but can also solve problems of position and flux anomalies in galaxy strong lensing systems. In this paper we focus on the perturbation of time delays in strong lensing systems caused by fuzzy dark matter, thus making an important extension of previous works. We select a specific system HS 0810+2554 for the study of time delay anomalies. Then, based on the nature of the fuzzy dark matter fluctuations, we obtain theoretical relationship between the magnitude of the perturbation caused by fuzzy dark matter, its content in the galaxy, and its de Broglie wavelength $λ_{\mathrm{dB}}$. It turns out that, the perturbation of strong lensing time delays due to fuzzy dark matter quantified as standard deviation is $σ_{Δt} \propto λ_{\mathrm{dB}}^{1.5}$. We also verify our results through simulations. Relatively strong fuzzy dark matter fluctuations in a lensing galaxy make it possible to to destroy the topological structure of the lens system and change the arrival order between the saddle point and the minimum point in the time delays surface. Finally, we stress the unique opportunity for studying properties of fuzzy dark matter created by possible precise time delay measurements from strongly lensed transients like fast radio bursts, supernovae or gravitational wave signals.

astro-ph.GA

Learning Diverse Policies with Soft Self-Generated Guidance

Reinforcement learning (RL) with sparse and deceptive rewards is challenging because non-zero rewards are rarely obtained. Hence, the gradient calculated by the agent can be stochastic and without valid information. Recent studies that utilize memory buffers of previous experiences can lead to a more efficient learning process. However, existing methods often require these experiences to be successful and may overly exploit them, which can cause the agent to adopt suboptimal behaviors. This paper develops an approach that uses diverse past trajectories for faster and more efficient online RL, even if these trajectories are suboptimal or not highly rewarded. The proposed algorithm combines a policy improvement step with an additional exploration step using offline demonstration data. The main contribution of this paper is that by regarding diverse past trajectories as guidance, instead of imitating them, our method directs its policy to follow and expand past trajectories while still being able to learn without rewards and approach optimality. Furthermore, a novel diversity measurement is introduced to maintain the team's diversity and regulate exploration. The proposed algorithm is evaluated on discrete and continuous control tasks with sparse and deceptive rewards. Compared with the existing RL methods, the experimental results indicate that our proposed algorithm is significantly better than the baseline methods regarding diverse exploration and avoiding local optima.

cs.LG