SearcharxivSearch

arXiv subjects

Zhen Yuan

Publications and source records attributed to Zhen Yuan.

At least 19 recordsLinked to original sources

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-{\alpha}, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-{\alpha} delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

cs.RO

Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions

Manipulation involving rigid-deformable interactions, such as hanging clothes or dressing humans, is common in daily life, making it essential for household robots. Compared to single-object manipulation or interactions between rigid bodies, these tasks are particularly challenging due to the rich multi-point contacts and the complex dynamics of the deformable bodies during interaction. Therefore, object-centric representations such as 6D poses or structural points without task-specific information become insufficient for these interactions. In this work, we propose a hybrid correspondence-based representation tailored for rigid-deformable interactions. First, to capture intricate interaction information, we introduce structure-, task-, and interaction-aware sparse keypoints. The keypoints are generated based on the global structures of both rigid and deformable objects, and filtered by their local interaction contacts. However, tracking these sparse keypoints through the interaction remains difficult due to the high-dimensional dynamics of deformable objects. Therefore, we further construct dense correspondences on the deformable objects for accurate keypoint tracking throughout the manipulation. This hybrid design combines the advantages of both representations: sparse keypoints encode rich, task-specific information for fine-grained manipulation, while dense correspondences ensure efficient tracking and generalization to novel deformations, shapes, and scenarios. Together, they enable one-shot transfer to new tasks with minimal demonstrations. Extensive experiments demonstrate the effectiveness and broad applicability of our method.

cs.RO

Characterizing the disruption of B\"ootes III: a missing link in the Galactic halo?

The B\"ootes III (Boo3) dwarf galaxy has long been suspected of being the progenitor of Styx, a ~50{\deg}-long stellar stream that was simultaneously discovered in the same region of sky. Boo3's diffuse morphology, large velocity dispersion, small pericenter, and excess of candidate stars at large radii suggest it is undergoing active tidal disruption. A link to Styx is therefore logical; however, a clear connection between these structures has not yet been clearly demonstrated. Here, we re-examine the Boo3-Styx association by searching for Boo3's tidal debris using a combination of Gaia-selected members, new CaHK narrow-band imaging with CFHT/MegaCam, and stellar tracer catalogues of blue horizontal branch and red giant branch stars. We also conduct a broad search for a putative stream using matched filter techniques applied to SDSS DR17 and DELVE DR2. Despite our extensive search, we find no observational evidence directly linking Boo3 to Styx. Furthermore, our results suggest that either Boo3's extended substructure is too diffuse to be detected with current data, or that its particular orbit may have erased a coherent tidal signature. Boo3 thus remains an enigmatic system, and exemplifies the need for spectroscopic follow-up to properly disentangle the nature between this faint Milky Way satellite and nearby stream.

astro-ph.GA

Exoplanets in ancient stellar populations: occurrence constraints and hot-Jupiter candidates in the Galactic halo

The Galactic halo preserves a record of the Milky Way's earliest assembly and contains both in-situ stars and stars accreted from dwarf galaxies. Possible planets around these stars, therefore, probe formation in ancient, metal-poor environments, including systems of extragalactic origin. We present a search for short-period transiting planets around kinematically selected halo dwarfs using Gaia DR3 and TESS, focusing on planets with periods of $1 < P < 10$ days. We identify two hot-Jupiter (HJ) candidates, one in the in-situ and one in the accreted halo, although the latter is highly grazing and excluded from the occurrence analysis. The accreted candidate, if confirmed, would orbit the most metal-poor HJ host known ([Fe/H] $\approx -1$). Using injection--recovery tests and automated vetting, we constrain occurrence in the full halo, in-situ, and accreted samples. In the HJ regime ($8\,R_\oplus < R_{\rm p} < 22\,R_\oplus$, $1\,\text{day}\ < P < 10$ days), the non-grazing candidate implies an overall halo occurrence rate of $0.13^{+0.12}_{-0.07}\%$ if planetary, while the absence of confirmed detections gives a corresponding $1\sigma$ upper limit of $<0.14\%$. For the in-situ halo, we infer $0.17^{+0.17}_{-0.10}\%$ (or $<0.19\%$ assuming no detections), while for the accreted halo we derive an upper limit of $<0.56\%$. These rates lie well below the corresponding short-period giant-planet occurrence measured in the Galactic disc. A forward model assuming Kepler-like occurrence also predicts $10 \pm 3$ detections compared with at most one observed. We find no significant occurrence difference between the in-situ and accreted halo populations, strengthening the evidence that close-in giant planets are rare across the old, metal-poor halo.

astro-ph.EP

The formation of the C-19 progenitor: a primordial cluster heated by gas expulsion

The extremely metal-poor nature of the C-19 stream indicates that its progenitor was a primordial stellar system born in the very early Universe. Current observations show that it has a small metallicity dispersion (0.18 at the 95% confidence level), which is the signature of a globular cluster origin, while at the same time displaying an unusually large velocity dispersion ($\sim10$ km/s) typical of dwarf galaxies. To reconcile this conflicting observational evidence, previous simulations have focused on potential interactions with dark matter subhalos, which can efficiently make a cluster stream dynamically hot. In this work, we explore internal dynamical processes in star cluster formation, focusing on initial conditions shaped by gas expulsion and a top-heavy initial mass function. We find that the large observed velocity dispersion and broad stream morphology can be reproduced by a cluster that underwent severe gas expulsion and expansion during its birth phase, which is potentially a typical formation scenario of extremely metal-poor star clusters. A top-heavy IMF and binaries can also increase the velocity dispersion. The formation of C-19 may involve a combination of these effects.

astro-ph.GA

M100: An Orchestrated Dataflow Architecture Powering General AI Computing

As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility for diverse AI workloads, they often fall short in efficiency and cost-effectiveness. Various Domain-Specific Architectures (DSAs) excel at particular AI tasks but struggle to extend across broader applications or adapt to the rapidly evolving AI landscape. M100 is Li Auto's response: a performant, cost-effective architecture for AI inference in Autonomous Driving (AD), Large Language Models (LLMs), and intelligent human interactions, domains crucial to today's most competitive automobile platforms. M100 employs a dataflow parallel architecture, where compiler-architecture co-design orchestrates not only computation but, more critically, data movement across time and space. Leveraging dataflow computing efficiency, our hardware-software co-design improves system performance while reducing hardware complexity and cost. M100 largely eliminates caching: tensor computations are driven by compiler- and runtime-managed data streams flowing between computing elements and on/off-chip memories, yielding greater efficiency and scalability than cache-based systems. Another key principle was selecting the right operational granularity for scheduling, issuing, and execution across compiler, firmware, and hardware. Recognizing commonalities in AI workloads, we chose the tensor as the fundamental data element. M100 demonstrates general AI computing capability across diverse inference applications, including UniAD (for AD) and LLaMA (for LLMs). Benchmarks show M100 outperforms GPGPU architectures in AD applications with higher utilization, representing a promising direction for future general AI computing.

cs.LG

MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models

Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet their positional encoding mechanisms remain suboptimal. Existing approaches uniformly assign positional indices to all tokens, overlooking variations in information density within and across modalities, which leads to inefficient attention allocation where redundant visual regions dominate while informative content is underrepresented. We identify positional granularity as an implicit resource and propose MODIX (Multimodal Information-Driven Positional IndeX Scaling), a training-free framework that dynamically adapts positional strides based on modality-specific contributions. MODIX jointly models intra-modal density via covariance-based entropy and inter-modal interaction via cross-modal alignment to derive unified scores, which rescale positional indices to allocate finer granularity to informative modalities while compressing redundant ones, without requiring any modification to model parameters or architecture. Experiments across diverse architectures and benchmarks demonstrate that MODIX consistently improves multimodal reasoning and adaptively reallocates attention according to task-dependent information distributions, suggesting that positional encoding should be treated as an adaptive resource in Transformers for multimodal sequence modeling.

cs.CV

A Tale of Two Origins: In-Situ versus Accreted Nitrogen-Rich Field Stars in the MW

Spectroscopic surveys have identified significant numbers of metal-poor nitrogen-rich (N-rich) field stars. These stars are strong candidates for escapees from globular clusters (GCs), as their distinctive nitrogen enhancement mirrors the chemical patterns observed in some of the members of GCs. As part of the effort to characterize their chemodynamical properties, we derived abundances for up to 25 elements in a sample of 33 N-rich field giant stars (18 of them are studied for the first time) using high-resolution optical spectroscopy. We confirm their elevated abundances of N, Na, and Al, strongly supporting a GC origin. Given that Galactic GCs themselves formed within diverse progenitor galaxies, we sought to identify the ancestral systems of these N-rich field stars. By analyzing their dynamical parameters, we separated the sample into high-energy (HE) and low-energy (LE) groups. The HE group exhibits lower [{\alpha}/Fe] and enhanced r-process abundances compared to the LE group. This indicates that the HE stars likely escaped from GCs accreted from massive dwarf galaxies (e.g., Gaia-Sausage-Enceladus), while the LE stars probably originated from in-situ GCs. We also find that the chemical pattern of these N-rich stars with [Fe/H] {\lessapprox} -1.0 are similar to the high-redshift ''N-emitters''. Furthermore, orbital integrations revealed a close encounter between one N-rich field star and the globular cluster NGC 6235. Our work demonstrates the potential of using chemodynamical analyses to trace Galactic assembly through chemical peculiar stars, while highlighting that larger samples and more precise data in the future are crucial to establish definitive origins.

astro-ph.GA

The primordial nature of the C-19 stellar stream

Stellar streams, remnants of compact star systems stretched out by the tidal forces of the Milky Way, offer a unique way to study stellar populations that formed billions of years ago. A particularly unique stream is C-19, the most metal-poor stellar stream known at less than a thousandth of the Sun's metallicity. The nature of C-19 is not yet clear, with properties that resemble both star clusters and ultra faint dwarf galaxies, yet in either case its extremely low metallicity indicates very early star formation, <1 Gyr after the Big Bang. Here, we present the first detailed study on the nature of C-19 based on the chemical abundances of 14 member stars from high-resolution spectroscopy. These reveal that C-19 formed stars in an early, rapid, and prolific star formation event, with mild inhomogeneous mixing of elements produced in massive stars. There is otherwise no evidence for subsequent star formation, multiple stellar populations, nor chemical evolution. Although C-19 is currently disrupted in the Milky Way halo, it offers a rare and complementary window into the details of star formation and chemical evolution in the early universe, ideal for comparisons with current studies of primordial star formation in the high-redshift universe.

astro-ph.GA

High-fidelity stellar extinction with Gaia and APOGEE -- I. The method and a new extinction curve

The scarcity of high-fidelity extinction measurements remains a bottleneck in deriving accurate stellar properties from Gaia parallaxes. In this work, we aim to derive precision extinction estimates for APOGEE DR19 stars, establishing a new benchmark for Galactic stellar population studies. We first determine reddening by comparing observed colorsr, etrieved from photometric surveys or standardized synthetic magnitudes from Gaia BP/RP spectra, to intrinsic colors predicted via an XGBoost model. The model is trained on minimally reddened stars to infer intrinsic colors and their associated uncertainties, using APOGEE stellar parameters (Teff, logg, [Fe/H], and [alpha/Fe]). The derived reddening values are then converted into extinctions using an anchor ratio of A_BP / A_RP = 1.694 +/- 0.004, derived from red-clump-like stars. Here, we provide extinction measurements in 39 filters across 10 photometric systems and introduce a new empirical extinction curve optimized for broadband passbands. Our extinction estimates (Av) outperform existing results (Bayestar19, StarHorse, SEDEX), achieving a typical precision of 0.03 mag in Av. Notably, we identify systematic deviations of up to 30% between monochromatic and passband-integrated extinction ratios at wavelengths greater than 700 nm. This result highlights the necessity of adopting passband-specific coefficients when correcting extinction to derive stellar parameters. The derived extinction and reddening data are available to the community for download through Zenodo.

astro-ph.SR

HR-GO II: chemical abundances of low-$E$ retrograde dynamically-tagged-groups: Revealing Thamnos as a very metal-poor substructure

Milky Way halo substructures identified in dynamical space are known to suffer from contamination from the Milky Way in-situ stars, which makes their accreted origins uncertain. We present detailed chemical abundances of 35 stars belonging to two sets of dynamically tagged groups, Rg8 and Rg9, to investigate their accreted nature. Both groups are composed of stars with low orbital energy and very retrograde orbits. We find that Rg8 and Rg9 are chemically indistinguishable across all elements, from C to Eu, strongly indicating that they belong to the same structure. The iron-abundance distribution of this low-$E$ retrograde group has a prominent peak at [Fe/H] $\approx-2.1$, revealing that its main population is very metal-poor, and a secondary peak at [Fe/H] $\approx-1.5$, very likely due to contamination from Milky Way in-situ stars. These groups also heavily overlap with the Thamnos substructure in dynamical space, and we thus use them to investigate the chemical properties of Thamnos. The dominant, low-metallicity population provides strong evidence for the ex-situ origin of Thamnos, as well as its very metal-poor nature. We do not see any evidence of an $\alpha$ knee in our sample, which is consistent with previous studies. Comparison with the Cetus-Palca stream in the chemical space shows similar abundance distributions, and thus it suggests that the Thamnos progenitor dwarf galaxy had a truncated star formation history due to its early merger with the Milky Way.

astro-ph.GA

OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models

Vision-Language Models (VLMs) have demonstrated strong performance across various multimodal tasks, where position encoding plays a vital role in modeling both the sequential structure of textual information and the spatial structure of visual information. However, current VLMs commonly adopt modality-unified 1D or 2D positional indexing strategies, which treat textual and visual tokens uniformly without accounting for their distinct structural properties and sequential continuity for text and spatial coherence for vision. To address this limitation, we propose OMEGA, a novel position encoding framework that employs Modality-Specific Position Encoding (MSPE) to assign positional indices while preserving the inherent structures of each modality across separate coordinate dimensions. Additionally, to align the information density of multimodal data in the positional index space, OMEGA introduces Global Adaptive Encoding Step Scaling (GAESS), which adaptively adjusts the position encoding step size of visual tokens based on the embedding entropy of both modalities. Experimental results demonstrate that OMEGA consistently enhances VLM performance across diverse architectures and VQA benchmarks. On visual-intensive tasks, OMEGA achieves up to 3.43% improvement over baseline position encoding strategies on Qwen2.5-VL-3B, with consistent gains observed across larger models including Qwen2.5-VL-7B and LLaVA-v1.5-7B.

cs.CV

EEG Dynamic Microstate Patterns Induced by Pulsed Wave Transcranial Photobiomodulation Therapy

Transcranial photobiomodulation (tPBM) therapy is an emerging, non-invasive neuromodulation technique that has demonstrated considerable potential in the field of neuropsychiatric disorders. Several studies have found that pulsed wave (PW) tPBM therapy yields superior biomodulatory effects. However, its neural mechanisms are still unknown which poses a significant barrier to the development of an optimized protocol. A randomized, single-blind study including 29 participants was conducted using a crossover design, with sham and continuous wave (CW) groups as controls. The EEG microstate analysis was utilized to explore the relative variations in temporal parameters and brain functional connectivity. To further elucidate the dynamic activity patterns of microstates, a 10-repeat 10-fold cross-validation with nine machine learning algorithms and kernel Shapley additive explanations analysis was employed. Results indicated that the pulsed wave mode enhanced the global efficiency, local efficiency, and betweenness centrality of microstate C in brain functional networks as well as the mean durations parameter achieving a middle to large effect size, with superior effects compared to the sham and continuous wave groups. Furthermore, the support vector machine based on the radial basis function method with kernel Shapley additive explanations analysis demonstrated the best performance with an area under the curve (AUC) reaching 0.956, and found that the 8 of top-10 microstate features related to microstate C contributed most significantly to the PW mode. In conclusion, the EEG microstate analysis found that PW tPBM therapy modulates the microstate C-specific patterns in the human brain, suggesting that microstate dynamics may serve as a state-dependent biomarker for the optimization of tPBM protocol.

q-bio.NC

Deep Learning-Based Fetal Lung Segmentation from Diffusion-weighted MRI Images and Lung Maturity Evaluation for Fetal Growth Restriction

Fetal lung maturity is a critical indicator for predicting neonatal outcomes and the need for post-natal intervention, especially for pregnancies affected by fetal growth restriction. Intra-voxel incoherent motion analysis has shown promising results for non-invasive assessment of fetal lung development, but its reliance on manual segmentation is time-consuming, thus limiting its clinical applicability. In this work, we present an automated lung maturity evaluation pipeline for diffusion-weighted magnetic resonance images that consists of a deep learning-based fetal lung segmentation model and a model-fitting lung maturity assessment. A 3D nnU-Net model was trained on manually segmented images selected from the baseline frames of 4D diffusion-weighted MRI scans. The segmentation model demonstrated robust performance, yielding a mean Dice coefficient of 82.14%. Next, voxel-wise model fitting was performed based on both the nnU-Net-predicted and manual lung segmentations to quantify IVIM parameters reflecting tissue microstructure and perfusion. The results suggested no differences between the two. Our work shows that a fully automated pipeline is possible for supporting fetal lung maturity assessment and clinical decision-making.

cs.CV

scMamba: A Scalable Foundation Model for Single-Cell Multi-Omics Integration Beyond Highly Variable Feature Selection

The advent of single-cell multi-omics technologies has enabled the simultaneous profiling of diverse omics layers within individual cells. Integrating such multimodal data provides unprecedented insights into cellular identity, regulatory processes, and disease mechanisms. However, it remains challenging, as current methods often rely on selecting highly variable genes or peaks during preprocessing, which may inadvertently discard crucial biological information. Here, we present scMamba, a foundation model designed to integrate single-cell multi-omics data without the need for prior feature selection while preserving genomic positional information. scMamba introduces a patch-based cell tokenization strategy that treats genomics regions as words (tokens) and cells as sentences. Building upon the concept of state space duality, scMamba distills rich biological insights from high-dimensional, sparse single-cell multi-omics data. Additionally, our novel contrastive learning approach, enhanced with cosine similarity regularization, enables superior alignment across omics layers compared to traditional methods. Systematic benchmarking across multiple datasets demonstrates that scMamba significantly outperforms state-of-the-art methods in preserving biological variation, aligning omics layers, and enhancing key downstream tasks such as clustering, cell type annotation, and trajectory inference. Our findings position scMamba as a powerful tool for large-scale single-cell multi-omics integration, capable of handling large-scale atlases and advancing biological discovery.

q-bio.CB

ViT-based Local Volume dwarf galaxy Identificationin (VIDA) in the CSST survey

Identifying dwarf galaxies within the Local Volume is crucial for constraining the luminosity function of satellite galaxies in the nearby universe. We report the detection capabilities of dwarf galaxies within the Local Volume using the Chinese Space Station Telescope (CSST). Based on the simulated imaging data of CSST, we develop a detection and classification pipeline that combines traditional image-based search techniques with advanced machine learning classification models. The simulated Local Volume dwarf galaxies can be identified using a pre-processing method for "extended source detection", followed by classification with a pretrained ViT-Base model. This pipeline achieves a true positive rate (TPR) exceeding 85% with a false positive rate (FPR) of only 0.1%. We quantify the detection completeness of Local Volume dwarf galaxies across a three-dimensional parameter space defined by absolute magnitude ($M_V$), half-light radius ($R_h$), and heliocentric distance, based on simulated single-exposure CSST wide-field imaging survey data. For unresolved or semi-resolved dwarf galaxies, our method achieves a significantly deeper absolute magnitude detection limit compared to catalog-based approaches, reaching $M_V = -7$ within 10 \Mpc. By combining this image-based approach with traditional stellar catalog-based "matched filter" techniques, our automated framework established in this work can identify dwarf galaxies within 20 \Mpc for the CSST mission.

astro-ph.GA

DeepSPV: A Deep Learning Pipeline for 3D Spleen Volume Estimation from 2D Ultrasound Images

Splenomegaly, the enlargement of the spleen, is an important clinical indicator for various associated medical conditions, such as sickle cell disease (SCD). Spleen length measured from 2D ultrasound is the most widely used metric for characterising spleen size. However, it is still considered a surrogate measure, and spleen volume remains the gold standard for assessing spleen size. Accurate spleen volume measurement typically requires 3D imaging modalities, such as computed tomography or magnetic resonance imaging, but these are not widely available, especially in the Global South which has a high prevalence of SCD. In this work, we introduce a deep learning pipeline, DeepSPV, for precise spleen volume estimation from single or dual 2D ultrasound images. The pipeline involves a segmentation network and a variational autoencoder for learning low-dimensional representations from the estimated segmentations. We investigate three approaches for spleen volume estimation and our best model achieves 86.62%/92.5% mean relative volume accuracy (MRVA) under single-view/dual-view settings, surpassing the performance of human experts. In addition, the pipeline can provide confidence intervals for the volume estimates as well as offering benefits in terms of interpretability, which further support clinicians in decision-making when identifying splenomegaly. We evaluate the full pipeline using a highly realistic synthetic dataset generated by a diffusion model, achieving an overall MRVA of 83.0% from a single 2D ultrasound image. Our proposed DeepSPV is the first work to use deep learning to estimate 3D spleen volume from 2D ultrasound images and can be seamlessly integrated into the current clinical workflow for spleen assessment.

eess.IV

Dc-EEMF: Pushing depth-of-field limit of photoacoustic microscopy via decision-level constrained learning

Photoacoustic microscopy holds the potential to measure biomarkers' structural and functional status without labels, which significantly aids in comprehending pathophysiological conditions in biomedical research. However, conventional optical-resolution photoacoustic microscopy (OR-PAM) is hindered by a limited depth-of-field (DoF) due to the narrow depth range focused on a Gaussian beam. Consequently, it fails to resolve sufficient details in the depth direction. Herein, we propose a decision-level constrained end-to-end multi-focus image fusion (Dc-EEMF) to push DoF limit of PAM. The DC-EEMF method is a lightweight siamese network that incorporates an artifact-resistant channel-wise spatial frequency as its feature fusion rule. The meticulously crafted U-Net-based perceptual loss function for decision-level focus properties in end-to-end fusion seamlessly integrates the complementary advantages of spatial domain and transform domain methods within Dc-EEMF. This approach can be trained end-to-end without necessitating post-processing procedures. Experimental results and numerical analyses collectively demonstrate our method's robust performance, achieving an impressive fusion result for PAM images without a substantial sacrifice in lateral resolution. The utilization of Dc-EEMF-powered PAM has the potential to serve as a practical tool in preclinical and clinical studies requiring extended DoF for various applications.

eess.IV