Searcharxiv⌕ Search

arXiv subjects

Haozhi Wang

Publications and source records attributed to Haozhi Wang.

At least 19 recordsLinked to original sources

In-Context Learning for Robots: Methods and Applications

General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clarifies their transfer assumptions and the roles of training, correspondence, and memory in making context useful. Across manipulation and navigation, we examine how these mechanisms preserve taught requirements as objects, environments, and execution conditions change. This analysis links method design to evaluation practices that distinguish responsiveness to teaching, physical transfer, and benefits from retained experience. The resulting agenda connects compositional task acquisition and faithful transfer with physical recursive self-improvement, in which experience improves the ability to learn subsequent tasks.

cs.RO↗

Asteroseismic Detection of a Massive Unseen Companion to the $δ$ Scuti Star TIC 160582982

We present TIC 160582982, an eccentric $δ$ Scuti binary with a dynamically massive unseen companion. Six coherent high-frequency pressure modes used as pulsation clocks yield a phase-modulation orbit with $P_{\rm orb}=89.29\pm0.19$ d, $e=0.539^{+0.040}_{-0.037}$, and $a_1\sin i/c=145.6^{+3.7}_{-3.4}$ s, corresponding to $K_1=42.2^{+2.2}_{-1.8}$ km s$^{-1}$ and $f(M)=0.416^{+0.032}_{-0.028}\,M_\odot$. An independent frequency-modulation analysis of the two strongest modes gives a consistent orbital solution. Rotating MIST isochrones constrained by the atmospheric parameters and spectral energy distribution (SED) give $M_1=1.93\pm0.19\,M_\odot$ and $R_1=2.65\pm0.23\,R_\odot$. The mass function implies a formal edge-on minimum companion mass of approximately $1.80\,M_\odot$. Assuming a luminous main-sequence companion and including the non-eclipsing geometry gives $M_2=2.19^{+1.24}_{-0.34}\,M_\odot$ and $i=61.5^{+16.6}_{-19.6}$ deg. However, no clear evidence for the presence of a luminous counterpart is found, either photometrically or spectroscopically. This suggests the presence of a massive compact object, such as a neutron star or even a stellar-mass black hole, as the massive unseen binary companion. Future phase-resolved high-resolution spectroscopy will be crucial for obtaining an independent orbital constraint and clarifying the nature of the companion.

astro-ph.SR↗

Discovery and Characterization of Three New High-amplitude delta Scuti Stars from TESS Observations

We report the discovery and detailed analysis of three new high-amplitude δ Scuti (HADS) stars, TIC 408074920, TIC 189714989, and TIC 34137913, identified from Transiting Exoplanet Survey Satellite short-cadence observations. Fourier analysis reveals dominant radial modes accompanied by rich harmonic structures in all three stars, confirming their HADS classification. In addition, TIC 189714989 exhibits clear symmetric side peaks around the radial harmonics, indicative of amplitude and/or phase modulation. Broadband spectral energy distribution (SED) fitting constrained by Gaia DR3 parallaxes provides independent estimates of the stellar effective temperatures and radii. These are compared with results from stellar evolutionary and seismic modeling based on MESA and GYRE. For two targets, the SED-derived and seismic radii are mutually consistent, while TIC 408074920 displays a significant discrepancy, plausibly attributable to the different sensitivities of the two methods and additional systematic effects. These findings highlight the diversity of amplitude variability among classical HADS and emphasize the importance of the long-term monitoring that future missions such as PLATO will allow in order to explore the underlying physical mechanisms of this variability.

astro-ph.SR↗

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large-scale human-annotated dataset with complete editing examples and a standardized evaluator for comparing editing systems. Existing resources are limited by small scale, missing edited outputs, or the absence of human quality labels, while current evaluation often relies on expensive manual inspection or generic vision-language model judges that are not specialized for editing quality. We introduce VEFX-Dataset, a human-annotated dataset containing 5,049 video editing examples across 9 major editing categories and 32 subcategories, each labeled along three decoupled dimensions: Instruction Following, Rendering Quality, and Edit Exclusivity. Building on VEFX-Dataset, we propose VEFX-Reward, a reward model designed specifically for video editing quality assessment. VEFX-Reward jointly processes the source video, the editing instruction, and the edited video, and predicts per-dimension quality scores via ordinal regression. We further release VEFX-Bench, a benchmark of 300 curated video-prompt pairs for standardized comparison of editing systems. Experiments show that VEFX-Reward aligns more strongly with human judgments than generic VLM judges and prior reward models on both standard IQA/VQA metrics and group-wise preference evaluation. Using VEFX-Reward as an evaluator, we benchmark representative commercial and open-source video editing systems, revealing a persistent gap between visual plausibility, instruction following, and edit locality in current models. Our project page is https://xiangbogaobarry.github.io/VEFX-Bench/.

cs.CV↗

Generative Models in Decision Making: A Survey

Generative models have fundamentally reshaped the landscape of decision-making, reframing the problem from pure scalar reward maximization to high-fidelity trajectory generation and distribution matching. This paradigm shift addresses intrinsic limitations in classical Reinforcement Learning (RL), particularly the limited expressivity of standard unimodal policy distributions in capturing complex, multi-modal behaviors embedded in diverse datasets. However, current literature often treats these models as isolated algorithmic improvements, rarely synthesizing them into a single comprehensive framework. This survey proposes a principled taxonomy grounding generative decision-making within the probabilistic framework of Control as Inference. By performing a variational factorization of the trajectory posterior, we conceptualize four distinct functional roles: Controllers for amortized policy inference, Modelers for dynamics priors, Optimizers for iterative trajectory refinement, and Evaluators for trajectory guidance and value assessment. Unlike existing architecture-centric reviews, this function-centric framework allows us to critically analyze representative generative families across distinct dimensions. Furthermore, we examine deployment in high-stakes domains, specifically Embodied AI, Autonomous Driving, and AI for Science, highlighting systemic risks such as dynamics hallucination in world models and proxy exploitation. Finally, we chart the path toward Generalist Physical Intelligence, identifying pivotal challenges in inference efficiency, trustworthiness, and the emergence of Physical Foundation Models.

cs.LG↗

ASCNet: Research on all-sky camera images classification at the Muztagh-ata site

Cloud coverage is one of the crucial elements of site testing in astronomy. All-sky camera (ASC) images are beneficial for our research on cloud coverage. In this paper, we propose ASCNet, an innovative model specifically designed for classifying nighttime ASC images collected at the Muztagh-ata site from 2022 March to 2024 June. ASCNet integrates ResNet34 with an ASCModule, which employs Depthwise Dilated Convolution and embeds lightweight Squeeze-and-Excitation attention within its branches to extract fine-grained texture information from the luminance channel. The data set is partitioned by category, with 70% of images assigned to the training set and 30% to the test set. The model's performance is assessed by comparing its predictions on the test set with manually annotated labels, yielding a consistency rate of 92.7%. All evaluation metrics of ASCNet are as follows: Accuracy 92.66%, Precision 83.26%, Recall 84.25%, and F1-Score 83.67%, and both ablation and comparative experiments demonstrate significant superiority over other models. A confusion matrix is utilized to analyze the differences between manual classification and model classification. The statistical results demonstrate the model's excellent classification performance and its robust generalization ability, illustrating that ASCNet has potential for application in future astronomical image classifications.

astro-ph.IM↗

A Theory of Multi-Agent Generative Flow Networks

Generative flow networks utilize a flow-matching loss to learn a stochastic policy for generating objects from a sequence of actions, such that the probability of generating a pattern can be proportional to the corresponding given reward. However, a theoretical framework for multi-agent generative flow networks (MA-GFlowNets) has not yet been proposed. In this paper, we propose the theory framework of MA-GFlowNets, which can be applied to multiple agents to generate objects collaboratively through a series of joint actions. We further propose four algorithms: a centralized flow network for centralized training of MA-GFlowNets, an independent flow network for decentralized execution, a joint flow network for achieving centralized training with decentralized execution, and its updated conditional version. Joint Flow training is based on a local-global principle allowing to train a collection of (local) GFN as a unique (global) GFN. This principle provides a loss of reasonable complexity and allows to leverage usual results on GFN to provide theoretical guarantees that the independent policies generate samples with probability proportional to the reward function. Experimental results demonstrate the superiority of the proposed framework compared to reinforcement learning and MCMC-based methods.

cs.LG↗

Charge Parity Rates in Transmon Qubits with Different Shunting Capacitors

The presence of non-equilibrium quasiparticles in superconducting resonators and qubits operating at millikelvin temperature has been known for decades. One metric for the number of quasiparticles affecting qubits is the rate of single-electron change in charge on the qubit island ($\textit i.e.$ the charge parity rate). Here, we have utilized a Ramsey-like pulse sequence to monitor changes in the parity states of five transmon qubits. The five qubits have shunting capacitors with two different geometries and fabricated from both Al and Ta. The charge parity rate differs by a factor of two for the two transmon designs studied here but does not depend on the material of the shunting capacitor. The underlying mechanism of the source of parity switching is further investigated in one of the qubit devices by increasing the quasiparticle trapping rate using induced vortices in the electrodes of the device. The charge parity rate exhibited a weak dependence on the quasiparticle trapping rate, indicating that the main source of charge parity events is from the production of quasiparticles across the Josephson junction. To estimate this source of quasiparticle production, we simulate and estimate pair-breaking photon absorption rates for our two qubit geometries and find a similar factor of two in the absorption rate for a background blackbody radiation temperature of $T^*\sim$ 350 mK.

quant-ph↗

Fabrication of Metal Air Bridges for Superconducting Circuits using Two-photon Lithography

Extraneous high frequency chip modes parasitic to superconducting quantum circuits can result in decoherence when these modes are excited. To suppress these modes, superconducting air bridges (AB) are commonly used to electrically connect ground planes together when interrupted by transmission lines. Here, we demonstrate the use of two-photon photolithography to build a supporting 3D resist structure in conjunction with a lift-off process to create AB. The resulting aluminum AB, have a superconducting transition temperature $T_{c} = 1.08$ K and exhibit good mechanical strength up to lengths of 100 $μ$m. A measurable amount of microwave loss is observed when 35 AB were placed over a high-$Q$ Ta quarter-wave coplanar waveguide resonator.

quant-ph↗

Unicorn: Unified Neural Image Compression with One Number Reconstruction

Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit image compression (IIC) based on implicit neural representations (INR). The former is encountering impasses of either leveling off bitrate reduction at a cost of tremendous complexity while the latter suffers from excessive smoothing quality as well as lengthy decoder models. In this paper, we propose an innovative paradigm, which we dub \textbf{Unicorn} (\textbf{U}nified \textbf{N}eural \textbf{I}mage \textbf{C}ompression with \textbf{O}ne \textbf{N}number \textbf{R}econstruction). By conceptualizing the images as index-image pairs and learning the inherent distribution of pairs in a subtle neural network model, Unicorn can reconstruct a visually pleasing image from a randomly generated noise with only one index number. The neural model serves as the unified decoder of images while the noises and indexes corresponds to explicit representations. As a proof of concept, we propose an effective and efficient prototype of Unicorn based on latent diffusion models with tailored model designs. Quantitive and qualitative experimental results demonstrate that our prototype achieves significant bitrates reduction compared with EIC and IIC algorithms. More impressively, benefitting from the unified decoder, our compression ratio escalates as the quantity of images increases. We envision that more advanced model designs will endow Unicorn with greater potential in image compression. We will release our codes in \url{https://github.com/uniqzheng/Unicorn-Laduree}.

cs.CV↗

Optical turbulence in the atmospheric surface layer at the Pamir Plateau Muztagh-ata site

In this paper, we conducted a detailed analysis of optical turbulence in the Atmospheric Surface Layer (ASL) at Muztagh-ata site during on-site testing. We utilized ultrasonic anemometers positioned on a 30-meter tower to collect and process data at five height levels, obtaining data from October 1, 2021 to the present. We investigated the behavior of optical turbulence parameters (\(C_n^2\) and seeing \(\varepsilon\)) in the ASL. Nighttime \(C_n^2\) primarily fluctuated in the range of \(10^{-16}\) to \(10^{-14}\), exhibiting an exponential decrease with height. During the day, it showed a \(h^{-0.82}\) dependency, while at night, it displayed a \(h^{-0.48}\) dependency. Additionally, we presented the distribution of seeing across different layers within the ASL, showing a gradual decrease with increasing height, with a median seeing of 0.24 arcseconds at nighttime and 0.48 arcseconds at daytime between 6-30m. We investigated the relationship between surface temperature inversion, seeing in the ASL, and wind speed at the site. Our results show that under temperature inversion conditions, seeing significantly improves and is often accompanied by low to moderate wind speeds, while high wind speeds are usually associated with poorer seeing. Preliminary calculations and observational results, combined with the high altitude and unique geographical location, suggest that Muztagh-ata site has the potential to be an outstanding optical astronomical observatory in the western plateau of china.

physics.ao-ph↗

KIC 10855535: An elegant Delta Scuti pulsator with Amplitude and Phase Modulation

We investigated the pulsating behavior of KIC 10855535 using Kepler 4-year long cadence data. Two independent frequencies were detected: a pulsation frequency F0 = 17.733260(5)d-1 and a low frequency f8=0.412643(8)d-1 We identify F0 as the fundamental frequency, at which a equidistant quintuplet is centered, suggesting that the star orbits in a binary system. The fitted orbital parameters align well with those reported in previous literature. Long-term phase modulation caused by binarity has been confirmed by considering TESS light curve. Through adjusting light time via removing the light time effect, we measured a linear change in period of order $\dot{P}/P \simeq 1.44\times 10^{-7}yr^{-1}$, a value that could be indicative of stellar evolution. The star also exhibits a gradual and stable amplitude growth, thereby raising the possibility of structural changes during its evolution. We attributed f8 and its two harmonics to rotation and surface spots, with further analysis suggesting evolving characteristics over time. Based on the hypothesis, KIC 10855535 may rotate slowly for its type, with a speed of 37(2)km/s. Overall, KIC 10855535 presents an exceptionally clean spectrum and a relatively slow rotation as a δ Sct pulsator, exhibiting a single pulsation mode that undergoes both amplitude and phase modulation.

astro-ph.SR↗

Generalized geometric speed limits for quantum observables

Leveraging quantum information geometry, we derive generalized quantum speed limits on the rate of change of the expectation values of observables. These bounds subsume and, for Hilbert space dimension $\geq 3$, tighten existing bounds -- in some cases by an arbitrarily large multiplicative constant. The generalized bounds can be used to design "fast" Hamiltonians that enable the rapid driving of the expectation values of observables with potential applications e.g.~to quantum annealing, optimal control, variational quantum algorithms, and quantum sensing. Our theoretical results are supported by illustrative examples and an experimental demonstration using a superconducting qutrit. Possibly of independent interest, along the way to one of our bounds we derive a novel upper bound on the generalized quantum Fisher information with respect to time (including the standard symmetric logarithmic derivative quantum Fisher information) for unitary dynamics in terms of the variance of the associated Hamiltonian and the condition number of the density matrix.

quant-ph↗

Molecular tuning of DNA framework-programmed silicification by cationic silica cluster attachment

The organizational complexity of biominerals has long fascinated scientists seeking to understand biological programming and implement new developments in biomimetic materials chemistry. Nonclassical crystallization pathways have been observed and analyzed in typical crystalline biominerals, involving the controlled attachment and reconfiguration of nanoparticles and clusters on organic templates. However, the understanding of templated amorphous silica mineralization remains limited, hindering the rational design of complex silica-based materials. Here, we present a systematic study on the stabilization of self-capping cationic silica cluster (CSC) and their assembly dynamics using DNA nanostructures as programmable attachment templates. By tuning the composition and structure of CSC, we demonstrate high-fidelity silicification at single-cluster resolution, revealing a process of adaptive templating involving cooperative adjustments of both the DNA framework and cluster morphology. Our results provide a unified model of silicification by cluster attachment and pave the way towards the molecular tuning of pre- and post-nucleation stages of sol-gel reactions. Overall, our findings provide new insights for the design of silica-based materials with controlled organization and functionality, bridging the gap between biomineralization principles and the rational design of biomimetic material.

physics.chem-ph↗

Twisted DNA origami-based chiral monolayers for spin filtering

DNA monolayers with inherent chirality play a pivotal role across various domains, including biosensors, DNA chips, and bioelectronics. Nonetheless, conventional DNA chiral monolayers, typically constructed from single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA), often lack structural orderliness and design flexibility at the interface. Structural DNA nanotechnology emerges as a promising solution to tackle these challenges. In this study, we present a strategy for crafting highly adaptable twisted DNA origami-based chiral monolayers. These structures exhibit distinct interfacial assembly characteristics and effectively mitigate the structural disorder of dsDNA monolayers, which is constrained by a limited persistence length of ~50 nm of dsDNA. We highlight the spin-filtering capabilities of four representative DNA origami-based chiral monolayers, demonstrating a maximal one-order-of-magnitude increase in spin-filtering efficiency per unit area compared to conventional dsDNA chiral monolayers. Intriguingly, our findings reveal that the higher-order, tertiary, chiral structure of twisted DNA origami further enhances the spin-filtering efficiency. This work paves the way for the rational design of DNA chiral monolayers.

physics.chem-ph↗

Cluster aggregates surrounding Pismis 5 in the Vela Molecular Ridge

Context. In the Gaia era, the precision of astrometric data is unprecedented. High-quality data make it easier to find more cluster aggregates and support further confirmation of these open clusters. Aims. We use Gaia DR3 to redetermine the open clusters surrounding Pismis 5 in the Vela Molecular Ridge. We also investigate the basic properties of these clusters. Methods. We apply two clustering algorithms (StarGO and pyUPMASK) to identify the open cluster members in a five-dimensional space with Gaia DR3. Results. We identify eight open clusters surrounding Pismis 5 in the Vela Molecular Ridge. The open cluster QZ 1 is newly discovered. Through investigating the comprehensive properties of the clusters, one open binary cluster candidate (Alessi 43 and Collinder 197) and one triple open cluster candidate (Pismis 5, Pismis 5A, and Pismis 5B) are discussed. Conclusions. Binary and triple open cluster candidates have been identified as potential primordial aggregates based on their similar age, position, and motion. According to kinematic speculations, the two aggregate candidates will gradually separate, and their interiors will slowly disintegrate.

astro-ph.GA↗

Identification and Mitigation of Conducting Package Losses for Quantum Superconducting Devices

Low-loss superconducting rf devices are required when used for quantum computation. Here, we present a series of measurements and simulations showing that conducting losses in the packaging of our superconducting resonator devices affect the maximum achievable internal quality factors (Qi) for a series of thin-film Al quarter-wave resonators with fundamental resonant frequencies varying between 4.9 and 5.8 GHz. By utilizing resonators with different widths and gaps, different volumes of the stored electromagnetic energy were sampled thus affecting Qi. When the backside of the sapphire substrate of the resonator device is adhered to a Cu package with a conducting silver glue, a monotonic decrease in the maximum achievable Qi is found as the electromagnetic sampling volume is increased. This is a result of induced currents in large surface resistance regions and dissipation underneath the substrate. By placing a hole underneath the substrate and using superconducting material for the package, we decrease the ohmic losses and increase the maximum Qi for the larger size resonators.

quant-ph↗

Meta Generative Flow Networks with Personalization for Task-Specific Adaptation

Multi-task reinforcement learning and meta-reinforcement learning have been developed to quickly adapt to new tasks, but they tend to focus on tasks with higher rewards and more frequent occurrences, leading to poor performance on tasks with sparse rewards. To address this issue, GFlowNets can be integrated into meta-learning algorithms (GFlowMeta) by leveraging the advantages of GFlowNets on tasks with sparse rewards. However, GFlowMeta suffers from performance degradation when encountering heterogeneous transitions from distinct tasks. To overcome this challenge, this paper proposes a personalized approach named pGFlowMeta, which combines task-specific personalized policies with a meta policy. Each personalized policy balances the loss on its personalized task and the difference from the meta policy, while the meta policy aims to minimize the average loss of all tasks. The theoretical analysis shows that the algorithm converges at a sublinear rate. Extensive experiments demonstrate that the proposed algorithm outperforms state-of-the-art reinforcement learning algorithms in discrete environments.

cs.LG↗