SearcharxivSearch

arXiv subjects

Jiaxuan Li

Publications and source records attributed to Jiaxuan Li.

At least 19 recordsLinked to original sources

Rubin LSST DP2 unveils almost-dark galaxies in the Virgo Cluster

Galaxies with the faintest surface brightness are currently known only in the Local Group. Similar objects should exist beyond our vicinity and are crucial for understanding galaxy evolution, structure, and dark matter content, yet surveys have not reached the depth required to detect them systematically. We present a population of seven almost-dark galaxies identified in Data Preview 2 of the Rubin Legacy Survey of Space and Time. They surround M49 in the vicinity of the Virgo Cluster, and exhibit central surface brightnesses in the range of $26.9 - 28.5 \, \mathrm{mag \, arcsec^{-2}}$ in the $g$-band, with half-light radii of $0.6 - 4.6 \, \mathrm{kpc}$ at the distance of Virgo, and stellar masses of $10^6 - 10^7 \, \mathrm{M_\odot}$. Their characteristics are analogous to those of the faintest and low-mass galaxies identified among satellite galaxies And XXI, And XXIII, and And XXV in the Local Group. This discovery demonstrates the power of the forthcoming Rubin LSST 10-year survey to uncover extremely faint galaxies at scale, promising the large statistical samples needed to constrain the faint-end luminosity function and the nature of dark matter.

astro-ph.GA

Wearable Multimodal Human-Machine Interface for Integrated Hand Intentions Decoding in Dynamic Teleoperation

Under ubiquitous teleoperation environments with optically challenging conditions, an interface for tele-operated grasping that combines wearability with precise decoding of hand intentions (hand pose, gestures, and grasping force) is essential. Yet, existing interfaces often fall short in meeting these demands, compromising either the diversity of multiple intentions decoding or wearability. To address this, we developed a novel Multiple Intentions Decoding Human-Machine Interface (MI-DHMI) that integrates high-throughput surface electromyography (sEMG) sensors with hand-mounted and forearm-mounted inertial measurement units (IMUs). The developed interface is supported by a unified framework for simultaneous multiple intentions decoding. By employing multimodal deep learning and hardware design with a low noise floor, the decoding framework selectively focuses on the sEMG components that are genuinely associated with finger movements. This effectively reduces decoding errors caused by sEMG variability during unconstrained upper-limb motions, thereby significantly enhancing robustness. Even under unconstrained wrist and forearm motion, the interface achieves a gesture recognition accuracy exceeding 97%, grasping force estimation with $R^2 = 0.95$, and hand pose decoding consistent with the actual hand pose, outperforming baseline devices and algorithms. Ablation studies further validate the effectiveness of the proposed decoding framework. Finally, two online experiments were conducted to validate the device, demonstrating its superior performance in high-stability tasks, including a pouring task and object grasping. The developed interface provides a new solution of a fully wearable, multiple intentions decoding system, offering effective support for ubiquitous teleoperation and contributing to the advancement of human-machine interaction research.

cs.RO

Data-Centric Neuromotor Interfaces for Portable Human-Machine Interaction

Dexterous human-machine interaction requires intuitive and expressive interfaces that can be efficiently deployed on constrained edge devices. Flexible material-based neuromotor interfaces hold considerable promise, as they decode human movement intention into natural control. Although emerging flexible electronic skins enable wearable high-fidelity data acquisition, practical deployment inevitably involves trade-offs between computational resources and portability. We present a data-centric paradigm where physiological features yield fundamental separability, providing sufficient discriminative cues for recognition. A wireless, high-bandwidth system developed for collecting various electrophysiological signals, when integrated with muscle-specific electrodes, forms a surface electromyography-based interface. Exploiting highly separable data, a 2,210-parameter model achieves 94.36% accuracy across 34 gestures and can be rapidly deployed on edge devices, establishing a new thousand-parameter benchmark for dexterous decoding. The underlying data-algorithm interactions in the data-centric paradigm are further clarified, demonstrating its feasibility in real-world scenarios. This study provides a principled and validated pathway for practical deployment of reliable neuromotor interfaces.

cs.RO

Do VLMs Share Safety Neurons Across Modalities?

Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts for self-repair, and introduce two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals. Our analysis reveals: (i) Text safety in VLMs is localizable: $\sim$88 neurons ($<$0.01%) whose targeted ablation substantially reduces refusal. (ii) Text safety neurons constitute the dominant refusal pathway: ablating them is the only intervention that consistently and substantially reduces refusal across all models. (iii) Visual safety is high-dimensional and diffuse at the single-neuron level: text safety concentrates in $\sim$5 subspace directions while visual safety requires $\geq$50. This gap holds across architectures, explaining why current alignment has not closed the visual safety gap. Project page is at: https://jiaxuan-li.github.io/vlm-safety-neuron/ Warning: this paper may include examples of harmful content.

cs.LG

ELVES-Dwarf. II. A Systematic Search for Satellite Systems of Dwarf Galaxies in the Local Volume

We present the Exploration of Local VolumE Satellites of Dwarf Galaxies (ELVES-Dwarf) survey, a systematic census of satellite systems around dwarf hosts in the Local Volume. Our final sample comprises 39 predominantly isolated hosts with stellar masses $10^{7} 10^5\,M_\odot$ around the 39 hosts. Above our fiducial completeness threshold of $M_\star\gtrsim10^{5.7}\,M_\odot$ and within the projected virial radius, 21 hosts have no confirmed satellites, 10 have one, six have two, and two have four, revealing substantial host-to-host scatter in satellite abundance. Overall, the observed satellite abundances and stellar mass functions are broadly consistent with predictions from the cosmological simulation TNG50 and galaxy formation models calibrated using Milky Way satellites. The projected radial distribution of the satellites is also consistent with theoretical expectations and with satellite populations around Milky Way-mass hosts. In contrast, the quenched fraction of satellites around dwarf hosts is substantially lower than around Milky Way-mass hosts, suggesting that environmental quenching is less efficient in dwarf halos. ELVES-Dwarf provides the first large, homogeneous, distance-confirmed sample of satellites around dwarf hosts and establishes a foundation for understanding galaxy formation and evolution in less-dense environments.

astro-ph.GA

Multimodal Adaptive Control for Safe Robotic Craniotomy Under Partial Observability

Autonomous robotic craniotomy requires continuous regulation of tool-tissue interactions to mitigate mechanical overload and thermal damage while maintaining surgical efficiency. However, this process is inherently partially observable due to unknown, time-varying tissue properties and the inability to directly measure cutting temperatures under physical occlusion. To address these challenges, we propose RL-MACRO, a cybernetic closed-loop intelligence framework that couples multimodal perception, adaptive decision-making, and robotic execution. This framework empowers the surgical robot to autonomously perceive inaccessible states from partial sensory feedback and dynamically optimize its behaviors under uncertain environment. A CNN-LSTM observer first fuses force and sound feedback to reconstruct the hidden temperature state (R^2=0.939, MAE = 1.717 deg C). This reconstructed temperature, alongside multi-sensor features, forms the belief state for an offline Implicit Q-Learning (IQL) policy. A novel dual-head Actor dynamically coordinates the feed rate, spindle speed, and cutting depth to optimize efficiency within strict safety bounds. These decisions are seamlessly translated into spatial motions via online trajectory re-planning and velocity servoing. Experiments on bovine ribs and six ex vivo goat skulls validate the system's robust perception, adaptive recovery from force/temperature excursions, and smooth execution on irregular surfaces, establishing a data-driven cybernetic paradigm for safe and efficient autonomous bone cutting.

cs.RO

DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions. Existing vision-language segmentation methods rely on deterministic cross-modal matching, which overlooks aleatoric uncertainty from ambiguous boundaries and epistemic uncertainty from limited training data, leading to fragile performance under domain shift. To address this issue, we propose DistMedVL, a probabilistic vision-language framework that introduces a lightweight Probabilistic Cross-Modal Adapter (PCM-Adapter) upon frozen encoders to explicitly model representational uncertainty. Specifically, the PCM-Adapter comprises two sequential modules for progressive probabilistic alignment. We first devise a Mahalanobis Alignment Module (MAM) that models textual tokens as Gaussian distributions and computes patch-text compatibility via Mahalanobis distance, yielding variance-conditioned matching that downweights unreliable feature dimensions. Moreover, we devise a Distribution Flow Module (DFM) that estimates modality-wise confidence parameters and performs vision-guided refinement of textual distributions, accommodating distributional variation across imaging modalities. Extensive experiments across eight medical segmentation benchmarks demonstrate that DistMedVL outperforms state-of-the-art methods with only 6.3M trainable parameters, exhibiting superior data efficiency, perturbation robustness and cross-dataset generalization.

cs.CV

Agentic Economic Modeling

We introduce Agentic Economic Modeling (AEM), a framework that aligns synthetic LLM choices with small-sample human evidence for econometric inference. AEM first generates task-conditioned synthetic choices via LLMs, then learns a bias-correction mapping from task features and raw LLM choices to human-aligned choices, upon which standard econometric estimators perform inference to recover demand elasticities and treatment effects. We validate AEM in two experiments. In a large scale conjoint study, using only 10% of the original data to fit the correction model lowers the error of the demand-parameter estimates, while uncorrected LLM choices increase the errors. In a regional field experiment, a mixture model calibrated on 10% of geographic regions estimates a treatment effect of -65$\pm$10 bps on the hold-out regions, closely matching the full human experiment (-60$\pm$8 bps). These results demonstrate AEM's potential to improve RCT efficiency and represent a step toward LLM-based counterfactual generation.

econ.EM

Zangetsu: A Candidate of Isolated, Quiescent, and Backsplash Ultra-Diffuse Galaxy in the COSMOS Field

Deep imaging surveys have changed our view of the low surface brightness (LSB) Universe. The ``renaissance'' of the LSB galaxy population, as a prime example of this recent development, continues to challenge our understanding of galaxy formation. Here, we report the serendipitous discovery of Zangetsu, an isolated, quiescent, and distorted ultra-diffuse galaxy (UDG) candidate in the COSMOS field, using images from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). Zangetsu exhibits an extremely low central surface brightness ($\mathrm{μ_{0,g}}=26.60\pm0.01$ mag arcsec$^{-2}$), a very shallow inner surface brightness profile ($\mathrm{n}_{\rm Sersic}=0.40\pm0.01$), and a large angular size ($\mathrm{R_e}\approx 10.44$ arcsec). Surprisingly, Zangetsu also has a quiescent stellar population ($\mathrm{g-i}=0.96$), an unusually elongated shape ($\mathrm{b/a}\sim 0.25$), and mild morphological asymmetry, making it a rare case among known UDGs. Surface brightness fluctuation analysis of HSC and Hubble Space Telescope (HST) images only provides a distance lower limit of $D>25.4$ Mpc (thus $\mathrm{R_e}>1.38$ kpc). However, Zangetsu remains an extreme outlier in the luminosity-size relation of known LSB galaxies, suggesting that it could be an exceptionally large and/or diffuse system. Classic internal or external UDG formation mechanisms alone struggle to explain such a system. A backsplash origin may account for its isolation and quiescent nature. This finding also raises the possibility that current works may overlook similarly extreme, elongated systems that could further our understanding of the LSB Universe.

astro-ph.GA

Thermal Conductivity and Temperature-Induced Band Gap Renormalization in Crystalline and Amorphous Ga$_2$O$_3$

The lattice thermal conductivity (LTC) and electron-phonon interactions in crystalline and amorphous gallium oxide are herein determined by coupling a machine-learned interatomic potential, namely the moment tensor potential (MTP) model, to first-principles calculations. Crystalline $β$-Ga$_2$O$_3$ exhibits a substantial band gap renormalization (BGR) of $\sim$0.45 eV at 700 K, with $\sim$0.2 eV caused by zero-point BGR. The computed temperature dependence of BGR induced by classical nuclear motion in $β$-Ga$_2$O$_3$ is stronger than that in amorphous Ga$_2$O$_3$, with the difference in BGR reaching $\sim$0.18 eV at 900 K. Thermal transport calculations reveal that the LTC of amorphous Ga$_2$O$_3$ remains near $0.9$ W$\cdot$ m$^{-1}$$\cdot$K$^{-1}$ for temperatures between 300 K and 700 K, which is approximately an order of magnitude lower than that of crystalline $β$-Ga$_2$O$_3$. Overall, the presented framework provides a computationally tractable and reliable route for predicting properties of semiconductors (both crystalline and amorphous) under operating conditions relevant to microelectronics and optoelectronics.

cond-mat.mtrl-sci

An Integrated Hardware-Software Design for Low-Data Spatial Defect Detection in Robotic Visual Inspection with Hybrid Optoelectronic Neural Networks

To address data overload and inefficient shape-level annotation in robotic visual inspection, this paper proposes a hardware-software integrated optoelectronic architecture. A non-imaging, low-data paradigm is established to minimize annotation dependency. First, a sensor-in-the-loop strategy reconfigures a Digital Micromirror Device (DMD) as a physical optical convolutional layer, enabling photonic-domain feature extraction that unifies sensing hardware and processing software. To suppress data volume at the source, a block-based compressed sensing strategy encodes spatial information into low-dimensional temporal signals, drastically reducing redundancy. Subsequently, to bypass laborious manual defect shape annotation, natural language descriptions guide the network to align with highly generalizable features from Contrastive Language-Image Pre-training (CLIP), steering the attention maps of the optoelectronic neural network toward defect shapes. Furthermore, a Localization Accuracy for Attention (LAA) metric is proposed to quantify shape-level defect localization performance. Experiments on transparent material defect detection validate the system's effectiveness. Parametric analysis reveals how measurement matrices, compression ratios, and block sizes affect accuracy. Results show that, compared to traditional imaging, the proposed architecture maintains equivalent accuracy while reducing data volume by 90% for Vision Transformers and computational workload by 60% for Convolutional Neural Networks. This low-data paradigm offers an efficient solution for industrial automation scenarios involving massive data streams, high acquisition costs, or constrained edge resources.

cs.RO

Structure leads and dominates comprehension in naturalistic reading

The hierarchical account and statistical or sequential account have long been framed as rival theories in explaining online comprehension. A lot of evidence has shown that both hierarchical and non-hierarchical factors can shape comprehension and the open question is no longer whether hierarchy contributes, but when and how strongly it does. We addressed the question with co-registered EEG and eye-tracking, treating syntactic depth as the variable for operationalizing hierarchical structure. On timing, hierarchical structure influenced reading before the eyes fixated a word: its neural effect emerged as early as 108 ms before fixation onset, over right-central regions, and the scanpath showed an anticipatory bias toward structurally central words. Both the transitional-probability analysis and the regression on fixation-related potentials supported this pre-fixational timing. In the transitional-probability analysis, readers preferentially moved between syntactically central words rather than following serial word order, showing that scanpaths are organized by syntactic depth rather than by linear adjacency. On strength, Bayesian network modeling showed that syntactic depth was the strongest predictor of departures from linear, word-by-word reading, outweighing lexical familiarity and surprisal. Taken together, the results indicate that hierarchical structure anticipatorily guides online comprehension at both the behavioral and neural levels, and dominates the reading path relative to statistical features.

cs.CL

Regret Pre-training: Bridging Prior and Posterior Views for Enhanced Knowledge Grounding

Causal language models factorize sequence probabilities using only preceding context, leaving future information unexploited during training despite its availability in the training data. This paper introduces Regret Pre-training, a self-supervised framework grounded in the Learning Using Privileged Information (LUPI) paradigm. The framework employs a dual-view architecture in which a single model generates both a causal Student distribution and a future-conditioned Teacher distribution. The training objective augments standard language modeling with a regret loss that minimizes the KL divergence from teacher to student, transferring future-aware signals to the causal representations. We investigate two teacher configurations on the OLMoE-1B-7B architecture:LocalRegret, which extends attention by one future token, andGlobalRegret, which conditions on bidirectional context with the target position masked. Experiments on nine downstream tasks following 4 billion tokens of training demonstrate that both configurations consistently outperform the baseline. On average,GlobalRegret andLocalRegret achieve 33.9% and 32.2% accuracy respectively, surpassing the baseline's 30.2%. Most notably,GlobalRegret improves BoolQ performance by 18.1 percentage points (61.0% vs 42.9%). The framework introduces no additional parameters and requires only one extra inference-mode forward pass per training step.

cs.CL

(LRDs)$^2$: The Low-ReDshift Little Red Dots Survey. II. DESI DR1 Sample

JWST has revealed a substantial population of "Little Red Dots" (LRDs) at $z>4$, challenging conventional AGN frameworks. However, the low-redshift regime remains largely unexplored. In the second paper of the (LRDs)$^2$ series, we present a systematic selection from DESI DR1 and identify 27 LRDs at $z=0.2-0.9$, yielding a number density lower limit of $7.5 \times 10^{-10}$ cMpc$^{-3}$. We conducted near-IR spectroscopic follow-up observations for 18 of them, revealing their full SED shapes and emission lines. These low-$z$ LRDs share the hallmark properties of their high-$z$ counterparts: compact morphology, V-shaped UV-optical continua, broad Balmer emission with extreme decrements (median H$α$/H$β\sim 16$), frequent Balmer absorption (67%), and blackbody-like optical-to-near-IR continua. All have low metallicity, occupy the same regions in the BPT diagram as high-$z$ LRDs, and have softer ionizing spectra than typical AGNs. The consistency between low-$z$ and high-$z$ LRD properties indicates the same physical processes at work. The correlation between broad-line Balmer luminosity and $L_{5100}$ deviates from that of local type-1 AGNs, limiting the direct application of local BH mass calibrations. Ionized [O III] outflows are ubiquitous (78%). One LRD at $z=0.196$, J1717+3807, shows robust long-term variability in $i$ and WISE bands. The optical-to-NIR continua of LRDs reveal a wide range of temperatures $\sim 2000-4700$ K (peak $0.6-1.5$ $μ$m), with a subset showing cooler and larger envelopes than those at high $z$. Low-$z$ LRDs serve not only as proximate laboratories for probing the nature of LRDs, but also trace the cosmic evolution of this population from the cosmic dawn to the present day.

astro-ph.GA

Supply Chain Coordination Mechanism Design: Consensus Planning Protocol Meets Vickrey-Clarke-Groves Mechanism

This paper introduces the theoretical framework for combining Vickrey-Clarke-Groves (VCG) mechanisms with the Consensus Planning Protocol (CPP) to enable truthful and efficient collaboration between a retailer and vendors to lower joint Cost to Serve. We demonstrate how this integration preserves both dominant-strategy incentive compatibility and efficiency in high-dimensional environments. We further introduce an activity fee design to improve its revenue property for the retailer while maintaining the mechanism's desirable properties. This CPP-VCG framework serves as the theoretical foundation for designing collaborative mechanisms to coordinates distributed, agent-based optimization between retailers and suppliers.

econ.TH

Carbon-Aware Compute--Power Scheduling for AI Data Centers with Microgrid Prosumer Operations

AI data centers are increasingly becoming tightly coupled compute--energy systems, where workload placement, cooling demand, electricity procurement, storage operation, and carbon emissions interact over time. This paper studies carbon-aware compute--power scheduling for geographically distributed AI data centers with microgrid prosumer capabilities. We propose a mixed-integer linear programming (MILP) framework that jointly schedules rigid training jobs, routes elastic inference workloads, dispatches local generation and battery storage, and manages bidirectional grid interaction under latency, continuity, power-balance, and carbon-budget constraints. The model captures two key features of emerging AI infrastructure: heterogeneous workload flexibility and site-level energy prosumer operation. Experiments on synthetic yet practically motivated instances show that the proposed joint MILP substantially improves total operational benefit over compute-only and energy-only baselines while reducing emissions. The results further indicate that inference-routing flexibility is a major source of value, battery storage provides useful temporal flexibility, and local-generation-rich settings are particularly favorable. The framework provides a tractable optimization abstraction for sustainable and grid-interactive AI data centers.

cs.CE

Joint Communication and Trajectory Design for Movable Antenna Systems

Movable antennas (MAs) have attracted significant attention in wireless communications due to their ability to reconfigure channel conditions by flexibly adjusting the antenna positions within a confined region. However, MA movement generally incurs a non-negligible delay, which may significantly limit the data transmission time at optimized positions. To tackle this challenge, this paper investigates a new joint communication and trajectory optimization problem, where each MA transmits while moving along an optimized trajectory to prolong the effective data transmission time. Focusing on a single-MA system, our goal is to maximize the average data rate by optimizing the MA's positions over time, subject to its maximum velocity constraints. However, this continuous-time antenna position optimization problem is highly non-convex and challenging to solve. To tackle this challenge, we first consider a special case with two channel paths and derive the optimal MA trajectory in closed form. For other general cases, we ingeniously reformulate the average rate maximization problem into a fixed-hop shortest path problem in graph theory by sampling the antenna movement region into a multitude of discrete points, and solve it optimally. Simulation results demonstrate that our proposed algorithm can significantly improve the data rate compared to other baseline schemes.

eess.SP

Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID

Any-Time Person Re-identification (AT-ReID) necessitates the robust retrieval of target individuals under arbitrary conditions, encompassing both modality shifts (daytime and nighttime) and extensive clothing-change scenarios, ranging from short-term to long-term intervals. However, existing methods are highly relying on pure visual features, which are prone to change due to environmental and time factors, resulting in significantly performance deterioration under scenarios involving illumination caused modality shifts or cloth-change. In this paper, we propose Semantic-driven Token Filtering and Expert Routing (STFER), a novel framework that leverages the ability of Large Vision-Language Models (LVLMs) to generate identity consistency text, which provides identity-discriminative features that are robust to both clothing variations and cross-modality shifts between RGB and IR. Specifically, we employ instructions to guide the LVLM in generating identity-intrinsic semantic text that captures biometric constants for the semantic model driven. The text token is further used for Semantic-driven Visual Token Filtering (SVTF), which enhances informative visual regions and suppresses redundant background noise. Meanwhile, the text token is also used for Semantic-driven Expert Routing (SER), which integrates the semantic text into expert routing, resulting in more robust multi-scenario gating. Extensive experiments on the Any-Time ReID dataset (AT-USTC) demonstrate that our model achieves state-of-the-art results. Moreover, the model trained on AT-USTC was evaluated across 5 widely-used ReID benchmarks demonstrating superior generalization capabilities with highly competitive results. Our code will be available soon.

cs.CV