SearcharxivSearch

arXiv subjects

Yan Ma

Publications and source records attributed to Yan Ma.

At least 19 recordsLinked to original sources

VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.

cs.CV

SleepWalking: Privileged Representation Shaping for End-to-End Blind Locomotion in Legged Robots

Partially observable locomotion requires a policy to act when task-relevant properties of the robot--environment state are not fully specified by instantaneous observations. Existing approaches often address this challenge by explicitly estimating missing physical variables or processing extended observation histories through structured architectures. We take a different view: partial observability is fundamentally an information-retention problem. The decisive question is not how task-relevant information enters the network, but whether the policy's internal state retains it. Guided by this perspective, we propose SleepWalking for Robot Locomotion (SWAQ), a one-stage end-to-end framework that uses next-step privileged physical reconstruction to shape what a recurrent history representation retains during policy learning, while the deployed actor uses only a direct history-to-action pathway. Under aligned training settings, SWAQ achieves a 15.0\% higher peak mean terrain level than DWAQ, the strongest non-exteroceptive baseline, while using 44.4\% fewer inference MACs per control step. Layerwise probes further show that information associated with the reconstructed physical variables remains linearly decodable through the policy head up to the layer preceding the action output. Complementary theoretical analysis relates privileged-variable recoverability to the achievable-return gap between history-based and privileged-information policy classes. These results suggest that semantic objectives can structure learning without requiring a corresponding architectural decomposition of the deployed controller.

cs.RO

AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty

Multimodal large language models (MLLMs) are widely used for automated annotation, yet their per-class accuracy varies widely (e.g., 12%-98% across the 13 classes of three classroom sub-datasets) and is expensive to measure: evaluating one 27B MLLM on 5,416 validation images takes roughly 14 hours, whereas a frozen-CLIP pass over the same images completes in about 3 minutes. A low-cost signal for ranking classes by expected MLLM annotation difficulty a priori remains underexplored. Building on the AnchorProxy construct (per-class zero-shot CLIP accuracy) introduced in the companion study, this paper systematically evaluates its full-frame formulation, termed AnchorScore here, as an a priori diagnostic that flags the classes MLLMs are least likely to annotate reliably. On classroom behavior data (SCB5, 13 classes, 6 MLLMs), AnchorScore correlates with per-class MLLM accuracy (Spearman rho = 0.769, p = 0.002, n = 13). None of the alternative difficulty predictors (DINOv2, ResNet-50, SigLIP, or MLLM self-verbalized uncertainty) showed a significant class-level correlation at n = 13. A cross-model consensus control suggests AnchorScore primarily captures a shared class-difficulty factor rather than a CLIP-specific signal. An independent replication on Stanford40 Actions yields a nearly identical effect (rho = 0.817, p < 0.001); the association is strongest on activity-recognition data and attenuates on medical and satellite imagery. Three practical applications follow: a deployable hybrid CLIP/MLLM routing strategy (predicted-class routing: up to +23 pp over CLIP-only at roughly 44% fewer MLLM calls), prompt disambiguation on hard classes (exploratory), and review-priority prediction for human verification. AnchorScore does not estimate exact MLLM accuracy; it provides a low-cost ranking signal that directs expensive MLLM evaluation to the classes where it is most informative.

cs.CV

Highly integrated quantum key distribution transmitter enabled by silicon photonics

Quantum key distribution (QKD) provides information-theoretic security independent of computational assumptions, yet the bulk and cost of current systems hinder large-scale deployment. Although integrated photonic technologies have enabled highly integrated QKD chips, practical QKD transmitters are still predominantly implemented as rack-mounted systems. Here, we demonstrate a highly integrated standalone QKD transmitter that integrates all essential functionalities required for practical QKD operation within a compact platform built around a silicon photonic encoding chip. The transmitter occupies only $167 \times 56 \times 21~\mathrm{mm}^3$ ($\sim 0.2~\mathrm{dm}^3$), comparable in size to a half-height, half-width PCIe card and more than 30 times smaller in volume than a conventional 1U rack-mounted system. Paired with a conventional discrete-component receiver, it achieves a secure key rate of 219.1 kbps over 51.3 km of standard single-mode fiber, with performance comparable to that of conventional discrete-component implementations. This work bridges the gap between photonic chip integration and deployable QKD hardware, marking an important step toward transitioning QKD transmitters from conventional rack-mounted equipment to compact board-level platforms.

quant-ph

How is Water released in Hydrogen-Based Metal Oxide Reduction? Unraveling the Kinetic Bottleneck in Sustainable Metal Production

Hydrogen-based direct reduction of metal oxides is a ubiquitous solid-gas redox process central to geophysics, sustainable metallurgy, redox energy cycles and catalysis. During this process, hydrogen removes lattice oxygen to form water, yet product water has long been regarded as a passive exhaust, and its nanoscale formation, trapping and removal remain poorly understood. Here, we directly observe redox-product water release from iron oxide during hydrogen-based direct reduction. Because water removal emerges from coupled structural, chemical and crystallographic evolution across multiple length-scales under realistic non-equilibrium reaction-conditions, we establish a correlative multiscale in-situ approach that links pore evolution, molecular water signatures, phase transformation and chemical-state evolution during hematite reduction. We uncover a mechanism in which oxygen removal induces closed nanopores spatially delocalized from reaction surfaces, causing transient trapping of water vapor. Water is released only when these pores coalesce into a percolating network connected to the surface, coinciding with and accelerating the onset of the hematite-to-magnetite transformation. These findings show that dynamically evolving pore topology governs mass transport and redox kinetics in solid-gas reactions, closing a critical mechanistic gap in product-water removal and providing nanoscale guidance for hydrogen-based metal extraction, reactor design, and sustainable redox energy technologies under practical conditions.

physics.chem-ph

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the level of individual visual claims. A good dense caption should be both faithful and informative, avoiding hallucination without omitting salient details. Yet pairwise preferences, reference-based metrics, and holistic scalar rewards compress these local errors into a single sequence-level signal, obscuring the tradeoff between factuality and coverage. We introduce ClaimDiff-RL, a framework that uses reference-conditioned atomic claim differences as the reward unit for caption RL. Given an image, an actor caption, and a reference caption, a multimodal judge enumerates visually grounded differences, verifies each difference against the image, assigns open-vocabulary error types and severity levels, and produces per-difference statistics for reward composition. This makes hallucinated claims and omitted salient facts separately measurable and tunable. Experiments show that holistic scalar rewards can reduce hallucination by increasing missing facts, while ClaimDiff-RL exposes this faithfulness and coverage tradeoff and enables more balanced operating points. On a 160-image human-labeled diagnostic benchmark, public captioning benchmarks, and VQA benchmarks, ClaimDiff-RL improves the hallucination--missing-fact balance, preserves general capability, and even surpasses Gemini-3-Pro-Preview on several fine-grained Capability dimensions such as object counting, spatial relations, and scene recognition. These results suggest that typed, verifiable claim differences are an effective reward unit for fine-grained and diagnosable caption RL.

cs.LG

Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model

We present daVinci-MagiHuman, an open-source audio-video generative foundation model for human-centric generation. daVinci-MagiHuman jointly generates synchronized video and audio using a single-stream Transformer that processes text, video, and audio within a unified token sequence via self-attention only. This single-stream design avoids the complexity of multi-stream or cross-attention architectures while remaining easy to optimize with standard training and inference infrastructure. The model is particularly strong in human-centric scenarios, producing expressive facial performance, natural speech-expression coordination, realistic body motion, and precise audio-video synchronization. It supports multilingual spoken generation across Chinese (Mandarin and Cantonese), English, Japanese, Korean, German, and French. For efficient inference, we combine the single-stream backbone with model distillation, latent-space super-resolution, and a Turbo VAE decoder, enabling generation of a 5-second 256p video in 2 seconds on a single H100 GPU. In automatic evaluation, daVinci-MagiHuman achieves the highest visual quality and text alignment among leading open models, along with the lowest word error rate (14.60%) for speech intelligibility. In pairwise human evaluation, it achieves win rates of 80.0% against Ovi 1.1 and 60.9% against LTX 2.3 over 2000 comparisons. We open-source the complete model stack, including the base model, the distilled model, the super-resolution model, and the inference codebase.

cs.CV

KeySense: LLM-Powered Hands-Down, Ten-Finger Typing on Commodity Touchscreens

Existing touchscreen software keyboards prevent users from resting their hands, forcing slow and fatiguing index-finger tapping ("chicken typing") instead of familiar hands-down ten-finger typing. We present KeySense, a purely software solution that preserves physical keyboard motor skills. KeySense isolates intentional taps from resting-finger noise using cognitive-motor timing patterns, and then uses a fine-tuned LLM decoder to convert the resulting noisy letter sequence into the intended word. In controlled component tests, the decoder substantially outperforms two statistical baselines (top-1 accuracy 84.8% vs 75.7% and 79.3%). A 12-participant study shows clear ergonomic and performance benefits: compared with the conventional hover-style keyboard, users rated KeySense as markedly less physically demanding (NASA-TLX median 1.5 vs 4.0), and after brief practice typed significantly faster (WPM 28.3 vs 26.2, p < 0.01). These results indicate that KeySense enables accurate, efficient, and comfortable ten-finger text entry on commodity touchscreens without any extra hardware.

cs.HC

What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

Vision tool-use reinforcement learning (RL) can equip vision language models with visual operators such as crop-and-zoom and achieves strong performance gains, yet it remains unclear whether these gains are driven by improvements in tool use or evolving intrinsic capabilities. We introduce MED (Measure--Explain--Diagnose), a coarse-to-fine framework that disentangles intrinsic capability changes from tool-induced effects, decomposes the tool-induced performance difference into gain and harm terms, and probes the mechanisms driving their evolution. Across checkpoint-level analyses in the crop-and-zoom setting on two VLMs with different tool priors and six benchmarks, we find that improvements are dominated by intrinsic learning, while tool-use RL mainly reduces tool-induced harm (e.g., fewer call-induced errors and weaker tool schema interference) and yields limited progress in tool-based correction of intrinsic failures. Overall, in the crop-and-zoom setting studied here, current vision tool-use RL learns to coexist safely with tools rather than master them.

cs.CV

Adaptively trained Physics-informed Radial Basis Function Neural Networks for Solving Multi-asset Option Pricing Problems

The present study investigates the numerical solution of Black-Scholes partial differential equation (PDE) for option valuation with multiple underlying assets. We develop a physics-informed (PI) machine learning algorithm based on a radial basis function neural network (RBFNN) that concurrently optimizes the network architecture and predicts the target option price. The physics-informed radial basis function neural network (PIRBFNN) combines the strengths of the traditional radial basis function collocation method and the physics-informed neural network machine learning approach to effectively solve PDE problems in the financial context. By employing a PDE residual-based technique to adaptively refine the distribution of hidden neurons during the training process, the PIRBFNN facilitates accurate and efficient handling of multidimensional option pricing models featuring non-smooth payoff conditions. The validity of the proposed method is demonstrated through a set of experiments encompassing a single-asset European put option, a double-asset exchange option, and a four-asset basket call option.

cs.LG

VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images

Interactive image segmentation(IIS) plays a critical role in generating precise annotations for remote sensing imagery, where objects often exhibit scale variations, irregular boundaries and complex backgrounds. However, existing IIS methods, primarily designed for natural images, struggle to generalize to remote sensing domains due to limited annotated data and computational overhead. To address these challenges, we proposed RS-ISRefiner, a novel click-based IIS framework tailored for remote sensing images. The framework employs an adapter-based tuning strategy that preserves the general representations of Vision Foundation Models while enabling efficient learning of remote sensing-specific spatial and boundary characteristics. A hybrid attention mechanism integrating convolutional local modeling with Transformer-based global reasoning enhances robustness against scale diversity and scene complexity. Furthermore, an improved probability map modulation scheme effectively incorporates historical user interactions, yielding more stable iterative refinement and higher boundary accuracy. Comprehensive experiments on six remote sensing datasets, including iSAID, ISPRS Potsdam, SandBar, NWPU, LoveDA Urban and WHUBuilding, demonstrate that RS-ISRefiner consistently outperforms state-of-the-art IIS methods in terms of segmentation accuracy, efficiency and interaction cost. These results confirm the effectiveness and generalizability of our framework, making it highly suitable for high-quality instance segmentation in practical remote sensing scenarios. The codes are available at https://github.com/wondelyan/VFM-ISRefiner .

cs.CV

HuBMAP Data Portal: a resource for multimodal spatial and single-cell data of healthy human tissues

The NIH Human BioMolecular Atlas Program (HuBMAP) Data Portal (https://portal.hubmapconsortium.org/) serves as a comprehensive repository for multimodal, multi-scale spatial and single-cell data from healthy human tissues. As of August 2026, the portal hosts 9,316 public datasets from 26 data types spanning 29 organ classes across 501 donors. Portal infrastructure and user interfaces support data search and discovery, visualization, and analysis directly in web browsers. These capabilities include metadata- and data-driven search, collaborative Workspaces with access to high-performance compute, and interactive Vitessce visualizations across non-spatial, 2D, and 3D spatial datasets. Data-type-specific uniform processing pipelines and rigorous quality control processes ensure comparability of results across laboratories, organs, and donors, while externally processed community-contributed datasets provide complementary perspectives. Here we describe portal functionality, infrastructure, and design, and highlight its role as a platform for large-scale spatial single-cell research across diverse data types, organs, and scales.

q-bio.QM

Mechanistic insights into hydrogen reduction of multicomponent oxides via in-situ high-energy X-ray diffraction

Co-reduction of multicomponent oxides with hydrogen provides a carbon-neutral approach toward sustainable alloy design. Herein, we investigate the hydrogen-based direct reduction, using in-situ high-energy X-ray diffraction of two precursor variants: mechanically mixed powders and pre-sintered oxide mixtures, targeting an equiatomic CoFeMnNi alloy. We find distinct reduction pathways and microstructure evolution depending on initial precursors. Mixed powders at 700 {\deg}C are reduced to body-centered-cubic, face-centered-cubic, and MnO phases via halite, spinel, and Mn3O4 intermediates, whereas the pre-sintered material directly transforms into a mixture of metallic and oxide phases. The post-reduction microstructures are also different: mixed oxides show loosely packed morphology, whereas pre-sintered material reveals metallic nanoparticles supported on nanoporous MnO. The formation of nanoporous metallic networks is strongly governed by the precursor state, highlighting the role of initial precursors on the final microstructure. This precursor design strategy offers a single-step route to nanoporous alloys with potential applications in catalysis and energy technologies.

cond-mat.mtrl-sci

Visual Programmability: A Guide for Code-as-Thought in Chart Understanding

Chart understanding presents a critical test to the reasoning capabilities of Vision-Language Models (VLMs). Prior approaches face critical limitations: some rely on external tools, making them brittle and constrained by a predefined toolkit, while others fine-tune specialist models that often adopt a single reasoning strategy, such as text-based chain-of-thought (CoT). The intermediate steps of text-based reasoning are difficult to verify, which complicates the use of reinforcement-learning signals that reward factual accuracy. To address this, we propose a Code-as-Thought (CaT) approach to represent the visual information of a chart in a verifiable, symbolic format. Our key insight is that this strategy must be adaptive: a fixed, code-only implementation consistently fails on complex charts where symbolic representation is unsuitable. This finding leads us to introduce Visual Programmability: a learnable property that determines if a chart-question pair is better solved with code or direct visual analysis. We implement this concept in an adaptive framework where a VLM learns to choose between the CaT pathway and a direct visual reasoning pathway. The selection policy of the model is trained with reinforcement learning using a novel dual-reward system. This system combines a data-accuracy reward to ground the model in facts and prevent numerical hallucination, with a decision reward that teaches the model when to use each strategy, preventing it from defaulting to a single reasoning mode. Experiments demonstrate strong and robust performance across diverse chart-understanding benchmarks. Our work shows that VLMs can be taught not only to reason but also how to reason, dynamically selecting the optimal reasoning pathway for each task.

cs.CV

Robust End-to-End FSO Transmission with Joint Coding Modulation and BiLSTM-Based Channel Modeling under Atmospheric Turbulence

Free space optical (FSO) communication is considered a promising solution in next_generation communication networks. However, its performance is significantly influenced by atmospheric turbulence. To enhance system robustness to turbulence, we propose a turbulence_robust end_to_end FSO communication system (TRFSO) that integrates a data-driven channel model with joint source_channel coding modulation (JSCCM). Specifically, a bidirectional long short-term memory (BiLSTM)_based channel model is developed and trained on data collected over a physical FSO link under varying turbulence conditions. This model accurately captures real_world channel distortions, achieving a minimum Kullback_Leibler (KL) divergence of 0.0019 in amplitude distribution matching. Experimental results show that the TRFSO system trained with the BiLSTM_based channel model outperforms the same architecture trained under the additive white Gaussian noise (AWGN) channel, achieving an average 3.5 dB improvement in multi-scale structural similarity (MS_SSIM) under strong atmospheric turbulence. These results demonstrate the effectiveness of the proposed TRFSO in achieving robust and reliable transmission under dynamic atmospheric turbulence.

physics.optics

Interaction as Intelligence: Deep Research With Human-AI Partnership

This paper introduces "Interaction as Intelligence" research series, presenting a reconceptualization of human-AI relationships in deep research tasks. Traditional approaches treat interaction merely as an interface for accessing AI capabilities-a conduit between human intent and machine output. We propose that interaction itself constitutes a fundamental dimension of intelligence. As AI systems engage in extended thinking processes for research tasks, meaningful interaction transitions from an optional enhancement to an essential component of effective intelligence. Current deep research systems adopt an "input-wait-output" paradigm where users initiate queries and receive results after black-box processing. This approach leads to error cascade effects, inflexible research boundaries that prevent question refinement during investigation, and missed opportunities for expertise integration. To address these limitations, we introduce Deep Cognition, a system that transforms the human role from giving instructions to cognitive oversight-a mode of engagement where humans guide AI thinking processes through strategic intervention at critical junctures. Deep cognition implements three key innovations: (1)Transparent, controllable, and interruptible interaction that reveals AI reasoning and enables intervention at any point; (2)Fine-grained bidirectional dialogue; and (3)Shared cognitive context where the system observes and adapts to user behaviors without explicit instruction. User evaluation demonstrates that this cognitive oversight paradigm outperforms the strongest baseline across six key metrics: Transparency(+20.0%), Fine-Grained Interaction(+29.2%), Real-Time Intervention(+18.5%), Ease of Collaboration(+27.7%), Results-Worth-Effort(+8.8%), and Interruptibility(+20.7%). Evaluations on challenging research problems show 31.8% to 50.0% points of improvements over deep research systems.

cs.CL

Sustainable Pre-reduction of Ferromanganese Oxides with Hydrogen: Heating Rate-Dependent Reduction Pathways and Microstructure Evolution

The reduction of ferromanganese ores into metallic feedstock is an energy-intensive process with substantial carbon emissions, necessitating sustainable alternatives. Hydrogen-based pre-reduction of manganese-rich ores offers a low-emission pathway to augment subsequent thermic Fe-Mn alloy production. However, reduction dynamics and microstructure evolution under varying thermal conditions remain poorly understood. This study investigates the influence of heating rate on the hydrogen-based direct reduction of natural Nchwaning ferromanganese ore and a synthetic analog. Non-isothermal thermogravimetric analysis revealed a complex multistep reduction process with overlapping kinetic regimes. Isoconversional kinetic analysis showed increased activation energy with reduction degree, indicating a transition from surface-reaction to diffusion-controlled reduction mechanisms. Interrupted X-ray diffraction experiments suggested that slow heating enables complete conversion to MnO and metallic Fe, while rapid heating promotes Fe- and Mn-oxides intermixing. Thermodynamic calculations for the Fe-Mn-O system predicted the equilibrium phase evolution, indicating Mn stabilized Fe-containing spinel and halite phases. Microstructural analysis revealed that slow heating rate yields fine and dispersed Fe particles in a porous MnO matrix, while fast heating leads to sporadic Fe-rich agglomerates. These findings suggest heating rate as a critical parameter governing reduction pathway, phase distribution, and microstructure evolution, thus offering key insights for optimizing hydrogen-based pre-reduction strategies towards more efficient and sustainable ferromanganese production.

cond-mat.mtrl-sci

Hydrogen-based direct reduction of multicomponent oxides: Insights from powder and pre-sintered precursors toward sustainable alloy design

The co-reduction of metal oxide mixtures using hydrogen as a reductant in conjunction with compaction and sintering of the evolving metallic blends offers a promising alternative toward sustainable alloy production through a single, integrated, and synergistic process. Herein, we provide fundamental insights into hydrogen-based direct reduction (HyDR) of distinct oxide precursors that differ by phase composition and morphology. Specifically, we investigate the co-reduction of multicomponent metal oxides targeting a 25Co-25Fe-25Mn-25Ni (at.%) alloy, by using either a compacted powder (mechanically mixed oxides) comprising Co3O4-Fe2O3-Mn2O3-NiO or a pre-sintered compound (chemically mixed oxides) comprising a Co,Ni-rich halite and a Fe,Mn-rich spinel. Thermogravimetric analysis (TGA) at a heating rate of 10 {\deg}C/min reveals that the reduction onset temperature for the compacted powder was ~175 {\deg}C, whereas it was significantly delayed to ~525 {\deg}C for the pre-sintered sample. Nevertheless, both sample types attained a similar reduction degree (~80%) after isothermal holding for 1 h at 700 {\deg}C. Phase analysis and microstructural characterization of reduced samples confirmed the presence of metallic Co, Fe, and Ni alongside MnO. A minor fraction of Fe remains unreduced, stabilized in the (Fe,Mn)O halite phase, in accord with thermodynamic calculations. Furthermore, ~1 wt.% of BCC phase was found only in the reduced pre-sintered sample, owing to the different reduction pathways. The kinetics and thermodynamics effects were decoupled by performing HyDR experiments on pulverized pre-sintered samples. These findings demonstrate that initial precursor states influence both the reduction behavior and the microstructural evolution, providing critical insights for the sustainable production of multicomponent alloys.

cond-mat.mtrl-sci