SearcharxivSearch

arXiv subjects

Zheng Ma

Publications and source records attributed to Zheng Ma.

At least 19 recordsLinked to original sources

China's Shrinking Home Bias and Rising Disruptive Impact: Evidence from a Global Citation Network Analysis

China has become the world's largest producer of scientific publications, yet concerns persist that this growth is inflated by excessive domestic citation practices. In this study, we analyze a citation network of over 45 million publications from Web of Science (1980-2025) to investigate China's home citation bias and research impact. Using a network reshuffling null model to control for the structural effect of publication volume, we find that China's home citation bias is less pronounced than commonly assumed and has been steadily declining over the past two decades. Chinese researchers do not exhibit a significantly stronger home citation preference than other major countries, indicating increasing internationalization rather than insularity. Furthermore, using the persistent disruption framework, we show that Chinese papers are converging toward American papers in their capacity to produce paradigm-shifting work. These findings challenge prevailing narratives about Chinese scientific home bias and suggest that China's advances in research impact.

physics.soc-ph

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.

cs.AI

Low Ly$\alpha$ Visibility in Galaxy Overdensities: Reionization Topology and Neutral-Fraction Ceilings from DIVER over $4.8<z<11$

Ly-alpha emission is widely used to trace cosmic reionization, but its interpretation depends on how Ly-alpha visibility varies with galaxy environment. We use deep JWST/NIRSpec observations from Deep Insights into UV Spectroscopy at the Epoch of Reionization (DIVER) in GOODS-N to measure Ly-alpha visibility for 250 galaxies at 4.8 25 A. We combine these measurements with H-alpha and [O III] emitters from JWST/NIRCam wide-field slitless spectroscopy to map the density field around each DIVER galaxy. Galaxies with high Ly-alpha equivalent widths (W_Lyalpha>25 A) or high effective Ly-alpha escape fractions (f_esc,Lyalpha^eff>0.05) tend to lie farther from nearby H-alpha and [O III] emitters than galaxies with lower Ly-alpha visibility. The clearest signal occurs near the prominent GOODS-N overdensity at z~5.2, where fewer than 15% of galaxies show strong Ly-alpha emission. This trend is opposite to the simplest inside-out reionization expectation that overdensities produce larger ionized regions and enhance Ly-alpha visibility. Possible explanations include circumgalactic and local intergalactic opacity, dense absorbers, and gas kinematics. We also derive an empirical upper envelope for f_esc,Lyalpha^eff and calibrate it with reionization simulations. Interpreting this envelope as a limiting IGM-attenuation signal gives neutral-fraction ceilings of _max=0.36, 0.76, 0.74, 0.84, and 1.0 at z~5.2, 5.8, 6.7, 7.7, and 9.8, respectively. The z~8 ceiling disfavors an almost completely neutral IGM at this epoch. These results support patchy reionization already underway by z~8 and show that galaxy Ly-alpha visibility encodes both large-scale ionization topology and near-source gas structure.

astro-ph.GA

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.

cs.CL

Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing

Accurate traffic forecasting is essential for proactive resource management in edge computing, where service demand evolves dynamically across both space and time. In practical cellular edge systems, traffic exhibits strong spatial correlations among neighboring service regions and long-range temporal dependencies driven by user mobility and application behavior. Existing recurrent forecasting approaches can capture short-term dynamics but often struggle to model long-horizon traffic evolution under non-stationary conditions. To address this challenge, we propose a spatiotemporal graph Transformer framework that jointly models spatial interactions and temporal dependencies for traffic forecasting in edge computing. The framework employs graph neural networks to capture spatial correlations among service regions and leverages Transformer-based self-attention to learn long-range temporal patterns from historical traffic observations. By decoupling spatial representation learning from temporal reasoning, the proposed approach provides an effective mechanism for large-scale spatiotemporal traffic modeling. Extensive experiments on a real-world cellular network dataset demonstrate that the proposed graph Transformer consistently outperforms recurrent graph-based baselines, including GCN-RNN, GCN-LSTM, and GCN-GRU models, across multiple forecasting horizons. The resulting forecasts enable more effective proactive resource provisioning and reduce overload risk compared with reactive management strategies. These results highlight the potential of graph-enhanced attention mechanisms for building intelligent and adaptive edge computing systems.

cs.LG

A quasar hatching from a buried red phase at z = 3.7

We present JADES-GS 209777, previously cataloged as CANDELS J033238.02-274626.2, hereafter "the Hatchling," a red quasar at $z=3.711$. While the source has been reported in earlier deep-field catalogs, our multiwavelength analysis reveals a visible active nucleus still embedded in a dense gas- and dust-rich environment. Red quasar continua are often attributed to dust attenuation, including non-standard extinction curves, but the highly comprehensive multiwavelength data for this source provide direct constraints on the material being cleared. Using JWST/NIRSpec, NIRCam, MIRI, HST, MUSE, Chandra, ALMA, and VLA data, we detect broad emission lines and strong X-ray emission, showing that the active nucleus is at least partially exposed. We also detect H$\alpha$ and He I absorption, indicating dense gas close to the nucleus. Kinematically disturbed O I, Mg II, Na D, and [O III] features, together with extended Ly$\alpha$ emission over $\gtrsim 20$ kpc, further show that multiphase gas is being accelerated from the nuclear region into the host-galaxy environment. The ALMA detection reveals strong dust emission, with the inferred infrared luminosity placing the system in the ULIRG regime. The continuum is red and sharply declining toward the rest-frame UV, resembling compact red AGNs, and may reflect extreme dust attenuation, gas reprocessing, possible BAL-like absorption, or a combination of these effects. Regardless of which mechanism dominates the continuum shape, the line diagnostics show that the visible nucleus remains partially obscured by nearby material. The Hatchling therefore represents a unique opportunity to explore a poorly known transition phase in which feedback is likely clearing an enshrouded quasar and allowing it to emerge toward a more unobscured active nucleus.

astro-ph.GA

PhotoIFU: NIRCam as a Photometric Integral Field Unit for Mapping Feedback in Galaxies

We present PhotoIFU, a workflow that uses deep multi-band imaging as a low-resolution photometric integral field unit. Applied to PSF-matched JWST/NIRCam imaging, PhotoIFU treats each spatial pixel as a coarse SED element and fits the pixel SEDs with Prospector to map resolved stellar-population and ISM-related properties. We apply this approach to three galaxies at $z=1.3$--3.7 in JADES: two systems with extended ionized line emission and one post-starburst galaxy with an exceptionally strong neutral outflow. Pixel-by-pixel SED fitting gives maps of stellar-mass surface density, specific star formation rate, dust attenuation, gas-phase metallicity, and recent star-formation history. We find that regions selected from the extended-emission or outflow geometry occupy distinct parts of the resolved SED-property distribution compared with the full host. In the systems with extended ionized emission, these regions are generally less dusty, consistent with ionized emission being observed along dust-poor, low-column-density pathways through the host. In the neutral-outflow system, the selected regions show enhanced recent star formation, suggesting that compact rejuvenation may mark the aftermath of an earlier energetic phase. These results show that galactic outflows and extended emission-line structures can be spatially associated with measurable differences in resolved host-galaxy stellar populations and ISM-related properties. PhotoIFU provides an imaging-based method for resolved SED mapping of feedback-related structures in larger galaxy samples where full spectroscopic integral-field mapping is unavailable.

astro-ph.GA

Analytic Distribution of Classifier-Free Guidance for Schedule Design

Classifier-free guidance (CFG) is the default mechanism for conditional generation in diffusion models, but the distribution sampled by its deterministic guided dynamics is not captured by the usual product-distribution heuristic $p_0^\omega q_0^{1-\omega}$. We analyze CFG through the probability flow ODE and derive exact analytic path-integral representations of the induced distributions for both constant and time-dependent guidance. The resulting formulas show that CFG modifies $p_{t_0}$ by an exponential path-integral correction, and that a time-dependent schedule enters this correction through the weight $\omega(t)-1$. This characterization explains how score discrepancies accumulate along sampling trajectories and motivates Distribution-Guided CFG (DG-CFG), a schedule that balances timestep contributions while accounting for signal strength and low-noise score-error amplification. A toy model with analytic scores closely verifies the predicted distributions. Across Stable Diffusion~1.5, Stable Diffusion~2.1, and Stable Diffusion~XL, DG-CFG yields a stronger diversity--fidelity trade-off and robustly mitigates the saturation and quality degradation caused by strong constant or heuristic guidance. Complete NFE experiments on Stable Diffusion~1.5 and Stable Diffusion~2.1 confirm that these gains persist across sampling budgets, while fixed-quality experiments on both backbones show that DG-CFG reaches target metrics with fewer sampling steps.

cs.LG

An (in)complete NIRSpec census of Balmer absorption in Type 1 AGN -- radiation-driven outflows in little red dots, quasars and variable stars

A notable achievement of the first generation of JWST surveys was the discovery of an abundance of faint high-redshift ($z > 4$) broad-line active galactic nuclei (AGN) in the $L_{\rm bol} < 10^{45}$~erg~s$^{-1}$ regime, previously only accessible at $z < 1$. The high prevalence of absorption features in their broad hydrogen lines is one of the peculiarities of a significant fraction of this new population. In this paper, we conduct a broad census of Balmer absorption in a sample of 47 Type-1 AGN spanning $2 < z < 7$. Accounting for incompleteness of JWST spectroscopic data, we estimate that $44_{-6}^{+21}$~\% of Little Red Dots (LRDs) have absorption in \Has while the incidence rate is $< 25$~\% (at 2$\sigma$) in Little Blue Dots (LBDs). Additionally, R1000 JWST/NIRSpec data is strongly resolution limited, implying an inability to disentangle Doppler broadening in the absorption from optical depth effects. This is alleviated with R2700 observations, although some degeneracies remain. We find a significant correlation between the velocity of the Balmer absorption and the [OIII]5007 narrow line luminosity, suggesting a common driving mechanism. We do not find any other significant correlations (and do not confirm previously claimed correlations) between Balmer absorption and other spectral properties of LRDs (except for the expected correlation between \Has and \Hbs absorption velocities). Comparing LRDs, quasars and stellar \Has absorbers, we establish that absorption in both LRDs and broad absorption line quasars is consistent with radiatively driven outflows, echoing the physics of variable stars yet occurring at vastly different scales.

astro-ph.GA

Multimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES Embeddings

Molecular property prediction often relies on isolated data modalities, where continuous 3D graph neural networks (GNNs) struggle to efficiently capture long-range topological dependencies and exact macroscopic heuristics. In this work, we introduce a parameter-efficient Tri-Branch Modular Fusion Neural Network that synthesizes three orthogonal modalities: 3D spatial geometry (SchNet), discrete topological grammar (SMILES via ChemBERTa), and explicit macroscopic physicochemical descriptors (Deep & Cross Network). By bypassing standard scalar readouts and employing a shared late-fusion architecture, the framework establishes a mathematically rigorous multimodal latent space that effectively resolves the arithmetic and oversmoothing limitations of local message passing. We evaluate the proposed architecture on the QM9 benchmark, targeting the extensive thermodynamic property of atomization energy at 0 K ($U_0^{\mathrm{atom}}$). Through systematic combinatorial ablation and latent bottleneck optimization ($d_e=64$), the tri-modal framework achieves a validation Mean Absolute Error (MAE) of 0.0207 eV. Operating with fewer than one million parameters, this architecture decisively surpasses the sub-chemical accuracy threshold and yields a substantial 20.6% error reduction over a strictly controlled geometric baseline. Ultimately, our findings demonstrate that integrating orthogonal macroscopic and topological data streams provides a synergistic, $\mathcal{O}(1)$ physical shortcut. This multimodal alignment offers a highly efficient alternative to brute-force parameter scaling, establishing a robust surrogate model for high-throughput virtual screening (HTVS) pipelines.

cs.LG

Decoupled Latent Optimization of Diffusion Models for Full Waveform Inversion

Full waveform inversion (FWI) recovers subsurface velocity from seismic recordings by solving a severely ill-posed, nonconvex PDE-constrained optimization. Classical regularizers stabilize the inversion but fail to reproduce realistic geological structures; recent diffusion-prior methods improve realism at the cost of a fragile trade-off between data fidelity and prior consistency. We propose Decoupled Latent Optimization (DLO), which relaxes the standard latent-optimization formulation into a quadratic-penalty objective over an auxiliary physical variable and a latent variable. The data-fidelity gradient acts in physical space, the diffusion sampler contributes only through a decoded prior sample, and the standard smoothed-velocity initialization of classical FWI is preserved. On the OpenFWI benchmark, DLO outperforms classical regularizers and existing diffusion-based methods under clean, noisy, and missing-trace acquisitions. The prior, trained on 70*70 OpenFWI models, transfers directly to the Marmousi and Overthrust benchmarks, where DLO recovers intricate fault structures and remains robust to initialization smoothing and measurement noise.

cs.LG

GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning

Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address this, we propose the GeoSym Engine, an automated and scalable neuro-symbolic framework. By leveraging a type-conditional grammar and an analytic SymGT Solver, it derives exact symbolic ground truths and seamlessly integrates with a robust rendering pipeline to produce high-precision geometric diagrams. Using this engine, we construct GeoSym127K, a difficulty-stratified dataset featuring 51K high-resolution images, 127K questions with symbolic ground truths, and 55K answer-verified CoT QA pairs. We also introduce GeoSym-Bench, an expert-curated suite of 511 complex samples for rigorous evaluation. Through extensive supervised fine-tuning (SFT), we demonstrate that GeoSym drives concentrated improvements specifically on diagram-dependent and multi-step geometry tasks. Our Qwen3-VL-8B model gains an absolute +22.21% on the MathVerse Vision-Only subset and reaches 61.52% (+6.19% improvement) on WeMath, mitigating long-horizon logic fragmentation and outperforming advanced closed-source models like Doubao-1.8. Furthermore, applying Reinforcement Learning with Verifiable Rewards (RLVR) via GRPO reveals that initializing from structural SFT checkpoints substantially elevates the performance ceiling over zero-shot RL. Driven by deterministic exact-match signals, this showcases the robust scaling potential of our verifiable reasoning synthesis. Datasets and code are available at https://huggingface.co/datasets/Tomie0506/GeoSym127K and https://github.com/Tomie56/GeoSym127K.

cs.CV

GraphPI: Efficient Protein Inference with Graph Neural Networks

The integration of deep learning approaches in biomedical research has been transformative, enabling breakthroughs in various applications. Despite these strides, its application in protein inference is impeded by the scarcity of extensively labeled datasets, a challenge compounded by the high costs and complexities of accurate protein annotation. In this study, we introduce GraphPI, a novel framework that treats protein inference as a node classification problem. We treat proteins as interconnected nodes within a protein-peptide-PSM graph, utilizing a Graph Neural Network-based architecture to elucidate their interrelations. To address label scarcity, we train the model on a set of unlabeled public protein datasets with pseudo-labels derived from an existing protein inference algorithm, enhanced by self-training to iteratively refine labels based on confidence scores. Contrary to prevalent methodologies necessitating dataset-specific training, our research illustrates that GraphPI, due to the well normalized nature of Percolator features, exhibits universal applicability without dataset-specific fine-tuning, a feature that not only mitigates the risk of overfitting but also enhances computational efficiency. Our empirical experiments reveal notable performance on various test datasets and deliver significantly reduced computation times compared to common protein inference algorithms.

cs.LG

Early metal-enriched baryon cycling before the midpoint of cosmic reionization

Models predict that chemical enrichment and gas redistribution should begin rapidly once star formation starts, but direct constraints at the earliest epochs have been scarce. Here we show that metal-enriched gas in multiple ionic phases was already present around galaxies before the midpoint of cosmic reionization. Using JWST/NIRSpec rest-frame ultraviolet spectroscopy from SPURS, we detect blueshifted metal absorption in three galaxies at $7.2<z<9.3$. The detected transitions span neutral, low-ionization, and high-ionization species, including O I, Si II, C II, Si IV, and C IV, with velocity offsets of $|\Delta v|\sim 50$--$250\,\mathrm{km\,s^{-1}}$ relative to nebular systemic redshifts. The ionic coexistence, overlapping velocity structure, and equivalent-width ratios are consistent with outflowing or otherwise kinematically disturbed galaxy-associated gas, implying rapid metal enrichment. These results show that key conditions for baryon cycling were established in at least a subset of luminous galaxies within the first several hundred million years of cosmic time, well before the completion of reionization.

astro-ph.GA

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achieving a marked performance leap, e.g., nearly 70% success on AndroidWorld. However, these systems keep their training data closed and remain opaque about their task and trajectory synthesis recipes. We present OpenMobile, an open-source framework that synthesizes high-quality task instructions and agent trajectories, with two key components: (1) The first is a scalable task synthesis pipeline that constructs a global environment memory from exploration, then leverages it to generate diverse and grounded instructions. and (2) a policy-switching strategy for trajectory rollout. By alternating between learner and expert models, it captures essential error-recovery data often missing in standard imitation learning. Agents trained on our data achieve competitive results across three dynamic mobile agent benchmarks: notably, our fine-tuned Qwen2.5-VL and Qwen3-VL reach 51.7% and 64.7% on AndroidWorld, far surpassing existing open-data approaches. Furthermore, we conduct transparent analyses on the overlap between our synthetic instructions and benchmark test sets, and verify that performance gains stem from broad functionality coverage rather than benchmark overfitting. We release data and code at https://njucckevin.github.io/openmobile/ to bridge the data gap and facilitate broader mobile agent research.

cs.AI

The Way We Tally Becomes the Tale: the Impact of Selection Strategies on the Inferred Evolution of Little Red Dots Across Cosmic Time

Little Red Dots (LRDs) have emerged as a key population linked to early black hole growth, yet photometric selections have predominantly targeted only the most extreme red systems, thereby shaping our current understanding of this new population of objects. In this work, we deliberately explore a broad range of optical redness while enforcing stringent compactness and visual inspection to ensure robustness and minimize contamination. Leveraging the depth and multiwavelength coverage of the JWST Advanced Deep Extragalactic Survey (JADES) data in the GOODS-North and GOODS-South fields, we construct the largest photometric census of LRDs to date in these fields, comprising 412 sources over $z\approx2\text{--}11$ across $\approx349.6$ arcmin$^2$. We show that classic extreme color cuts isolate only a minor fraction of this population ($\lesssim25\%$), while the majority of LRDs span a broader, largely unexplored parameter space. We quantify how selection strategies impact UV and optical luminosity functions and number density evolution, finding that current demographic trends of LRDs are strongly driven by selection biases and further limited by incomplete identification at both high and low redshift. Spectroscopically confirmed LRDs reveal a continuous range of spectral shapes, consistent with varying Active Galactic Nucleus (AGN) and host contributions in agreement with recent findings. Our results demonstrate that commonly adopted, purity-driven selections bias current demographic constraints toward the most extreme systems, potentially misrepresenting the diversity and evolution of the LRD population. Accounting for these selection effects is essential for interpreting LRDs and their role in early black hole growth.

astro-ph.GA

Light-modulated exchange bias in multiferroic heterostructures

Magnetic straintronics, the strain-mediated control of magnetic anisotropy, has emerged as a key direction for next-generation energy-efficient technologies. In multiferroic heterostructures, magnetoelectric coupling is typically achieved by applying an electric field on a ferroelectric phase, inducing strain through the converse piezoelectric effect, which is then transferred to the adjacent ferromagnetic phase. As an alternative, strain can be remotely modulated through the photostrictive effect induced by light. While light-driven control of magnetic anisotropy has been explored, optical modulation of more complex phenomena such as exchange bias remains largely unaddressed. Here, we demonstrate significant light-induced modulation of exchange bias and magnetization switching at room temperature in a Pb(Mg1/3Nb2/3)O3-Pb(Zr,Ti)O3 (PMN-PZT)/Fe80Ga20(FeGa)/Ir20Mn80(IrMn) multiferroic heterostructure, driven by visible-light-photostriction. The magnetization state correlates with the light intensity, enabling multi-level states with light power densities as low as 0.1 W cm-2. These findings suggest a promising route toward low-power, multistate, and wireless opto-magnetic memory applications.

cond-mat.mtrl-sci

FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair

Multimodal Automated Program Repair (MAPR) extends traditional program repair by requiring models to jointly reason over source code, textual issue descriptions, and visual artifacts such as GUI screenshots. While recent LLM-based repair systems have shown promising results, existing approaches face several limitations: rigid workflow pipelines restrict exploration during debugging, visual reasoning is often performed over full-page screenshots without localized grounding, and failed repair attempts are rarely transformed into reusable knowledge. To address these challenges, we propose FailureMem, a multimodal repair framework that integrates three key mechanisms: a hybrid workflow-agent architecture that balances structured localization with flexible reasoning, active perception tools that enable region-level visual grounding, and a Failure Memory Bank that converts past repair attempts into reusable guidance. Experiments on SWE-bench Multimodal demonstrate FailureMem improves the resolved rate over GUIRepair by 3.7%.

cs.SE