SearcharxivSearch

arXiv subjects

Yucheng Li

Publications and source records attributed to Yucheng Li.

At least 19 recordsLinked to original sources

Intensity-based scattering correction enables in vivo two-photon imaging beyond 1 mm

Optical imaging of the deep brain with subcellular resolution is essential for neuroscience, but noninvasive imaging beyond the cortex, through scattering white matter and into the hippocampus, has generally required three-photon microscopy at longer excitation wavelengths. Here, we introduce deep-learning-enhanced Fourier-domain intensity coupling for scattering correction (DeepFOCUS), an intensity-based two-photon approach that uses deep learning to compute intensity-modulation masks for real-time modulation of excitation light during image acquisition. Unlike deep-learning-based image restoration, this method directly improves image formation by computing intensity-modulation masks that shape the excitation light in real time, with each mask experimentally validated by the acquired fluorescence signal to avoid hallucination artifacts. Using 1035 nm excitation, we achieved in vivo two-photon imaging beyond 1 mm depth in the intact mouse brain, resolving YFP-labeled neurons and FITC-labeled blood vessels through the entire cortex and white matter down to the CA1 region of the hippocampus. DeepFOCUS extends two-photon imaging to depths previously accessible mainly with three-photon microscopy and could enable broader adoption of hippocampal imaging by upgrading existing two-photon systems.

physics.optics

CyberNeuro: A Privacy-Preserving Agentic Workbench for Cohort-Scale Neuroimage and Clinical Data Analysis

Despite tremendous success in neuroimaging methodology, making large-scale, high-dimensional datasets ready for AI/ML applications remains a critical operational bottleneck. Conventional workflows require extensive manual effort across metadata curation, pipeline execution, post-processing quality control, and data management, a burden that disproportionately excludes laboratories with limited manpower and computational infrastructure. To address this real-world barrier, there is an urgent need for scalable, cost-effective computational platforms that democratize advanced neuroimaging analytics and accelerate discoveries in mental health and clinical translation. Capitalizing on multi-agent LLM breakthroughs, we introduce CyberNeuro, an agentic workbench with a tailored local LLM-model ('WandaMind') for automated neuroimaging and health-data analysis. Driven by four dedicated agents (Planner, Validator, Dispatcher, and Reporter) communicating via a secure MCP bridge and a pinned execution layer, CyberNeuro enables researchers to execute complex workflows using natural language while maintaining clinical-grade data privacy. On the public NeuroBench suite, CyberNeuro increases held-out domain accuracy from 40% to 69% over the baseline model. Beyond automated metrics, the platform integrates a human-in-the-loop verification panel to ensure rigorous biomedical quality control. Across the same end-to-end 10-batch cohort workflow suite, the local WandaMind configuration completed all tasks with an estimated aggregate token count of about 10.6% using WandaMind and 61.7% using cloud providers of token usage, compared to Neuroclaw, respectively. The platform and its production-ready modules are available at https://wanda-cyberbench.com.

cs.MA

NeuroPilot: An Agent-Driven Smart Pipeline for Processing, Quality Control, and Managing Neuroimages

Transforming raw neuroimage archives into analysis-ready derivatives relies on three brittle stages: data standardization, modality-specific preprocessing, and quality control (QC). While individual neuroimaging tools are well developed, their orchestration requires project-specific scripts, environment-adaptive tuning, and labor-intensive manual QC. To address this, we introduce NeuroPilot, a multi-agent system that digitalizes the expertise of neuroimage processing, QC, and data management into three LLM-invocable skills: dcm2bids-skill, neuroimage-pre-skill, and qc-agent-skill. The LLM-driven agent autonomously orchestrates workflows, generalizing various infrastructure settings into a single configuration to achieve the highest scalability. Demonstrating the system's generalizability, we deployed NeuroPilot across 17 cohorts (>123,000 subjects) spanning infant to aging populations and multiple MRI modalities (structural, diffusion, functional). In practice, after standardizing data via the dcm2bids-skill, the agent dynamically routes datasets to the optimal neuroimage-pre-skill based on available modalities and cohort traits (e.g., dispatching T1w and fMRI data to fMRIPrep, or selecting specialized pipelines for infant cohorts). The qc-agent-skill then drives an evidence-based, semi-automated QC via a 3-D browser dashboard, utilizing a multi-tiered verification system to optimize failed cases and escalate complex issues for supervisor inspection. Quantitatively, our QC agent screened 558 production subjects, validating its automated flags against FreeSurfer's topology-defect metrics. The infant processing pipeline achieved a 100% (201/201) completion rate on QC-validated inputs. Importantly, NeuroPilot compresses the traditional 2--3 month timeline for training staff and processing complete datasets into a single week. NeuroPilot is deployed in https://wanda-cyberbench.com/.

cs.CV

Qwen-AgentWorld: Language World Models for General Agents

A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling based on language models can further push the boundaries of general agents. (i) We first focus on building foundation models for agentic environment simulation. We introduce Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, the first language world models capable of simulating agentic environments covering 7 domains via long chain-of-thought reasoning. Leveraging more than 10M environment interaction trajectories of 7 domains in real-world environments, we develop Qwen-AgentWorld through a three-stage training pipeline: CPT injects general-purpose world modeling capabilities from the state transition dynamics and augmented professional corpora, SFT activates next-state-prediction reasoning, and RL sharpens simulation fidelity through a tailored framework with hybrid rubric-and-rule rewards. To evaluate language world models, we present AgentWorldBench, a comprehensive benchmark constructed from real-world interactions of 5 frontier models on 9 established benchmarks. Empirical results demonstrate that Qwen-AgentWorld significantly outperforms existing frontier models. (ii) Beyond foundation models, we further investigate two complementary paradigms through which world modeling enhances general agents. First, as a decoupled environment simulator, Qwen-AgentWorld supports scalable and controllable simulation of thousands of real-world environments for agentic RL, yielding gains that surpass real-environment training alone. Second, as a unified agent foundation model, world-model training acts as a highly effective warm-up that improves downstream performance across 7 agentic benchmarks. Code: https://github.com/QwenLM/Qwen-AgentWorld

cs.CL

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Reinforcement learning (RL) has become a key component in modern large language models, yet the rollout stage remains the key bottleneck in RL training pipelines. Although Multi-Token Prediction (MTP) offers a natural solution to accelerate rollouts through speculative decoding, many studies have observed that MTP acceptance rates degrade significantly during RL training, leading to limited speedup performance. To address this bottleneck, we present Bebop, a systematic study of MTP in LLM post-training, and offer practical recipes to integrate MTP into large-scale RL pipelines. First, we reveal that the MTP acceptance rate is fundamentally bounded by the fluctuation of model entropy, which demonstrates a clear negative linear relationship with the rise of entropy in the RL stage. Second, we show that probabilistic rejection sampling largely alleviates the disturbance introduced by entropy in RL compared to greedy draft sampling. We further identify that the conventional MTP training objectives (cross-entropy or KL) are suboptimal in such settings, and therefore we propose a novel end-to-end TV loss that directly optimizes multi-step rejection sampling acceptance rate, yielding ~10% acceptance rate improvements, achieving up to 95% acceptance rates and up to 25% extra inference throughput gains across mathematical reasoning, code generation, and agentic tasks. Third, we test various online MTP training strategies during RL and show that pre-RL MTP training with e2e TV loss and rejection sampling achieves a consistent acceptance rate and speedup throughout the entire RL, eliminating the need for costly online MTP updating. We provide extensive experiments and analysis that validate our findings. Experimental results show our method achieves up to 1.8x end-to-end acceleration in async RL training of Qwen3.5, Qwen3.6, and Qwen3.7 models.

cs.LG

Superconductivity in Al-based high-entropy alloys TiHfNbTaAl and TaNbHfZrAl

Since the first report of a high entropy alloy (HEA) superconductor in 2014, HEAs have continued to captivate the interest of superconducting researchers. Owing to the significant degree of disorder inherent in these systems, they serve as exemplary models for examining the properties of materials that exist in states intermediate between crystalline and amorphous structures. Here we present the superconductivity properties and crystal structure of TaNbHfZrAl and TiHfNbTaAl HEAs, which both have the VEC of 4.2 and body centered cubic (BCC) structure. Through resistivity, magnetic, and specific heat measurements, we prove that both samples are the bulk type-II superconductors with a critical temperature Tc of 5.5 K for TaNbHfZrAl and Tc of 3.2 K for TiHfNbTaAl. The Tc of HEA superconductors is influenced by the VEC and the element composition. And the incorporation of Al in high disorder HEA superconductors causes a more crystallinelike Tc dependence.

cond-mat.supr-con

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL

Modern reasoning agents are increasingly evaluated on their ability to generate multiple valid solution paths, plans, or tool-use traces for a given input. Standard reward-maximizing RL tends to collapse onto the most easily reinforced high-reward mode, whereas distribution-matching RL aims to allocate probability mass across the entire reward-shaped solution set. Achieving this objective requires computing a prompt-dependent partition function over the trajectory space. Because existing distribution-matching methods learn this partition function online alongside the policy, calibration errors in the partition function directly distort policy updates and remain impossible to diagnose independently. We introduce DISA, short for Decoupled Importance-Sampled Anchoring, which moves this calibration problem outside the RL loop. DISA draws proposal trajectories offline, estimates the partition function via importance sampling, and freezes the resulting partition-function estimate before policy optimization begins. This decoupling preserves the distribution-matching objective while strictly separating partition-function estimation from policy learning in data, gradients, loss, and diagnostics. Empirically, on two open-weight backbones across six math and three code benchmarks, DISA matches or exceeds the online-coupled distribution-matching baseline FlowRL, outperforms rewardmaximization baselines GRPO and GSPO on math averages, and exceeds LoRASFT distillation by up to 13.8 Mean@8 points on the same offline trajectories. An LLM-as-judge evaluation further shows that DISA retains substantially more strategy-level diversity than reward-maximization baselines, and sensitivity studies on the proposal strength and inverse temperature follow the bias-variance pattern predicted by the analysis.

cs.LG

Superconductivity in the A15-type V3(Os1-2xSixGex) medium-entropy alloys

Cubic A15-type superconducting alloys continue to fascinate the academic and industrial fields because they mainly support the largest market for low-temperature superconducting applications and show exotic physical properties. Medium-/high-entropy alloys (MEAs-HEAs) can be employed stably under extreme conditions due to their high mechanical hardness and excellent irradiation tolerance. Combining with the features of the A15-type superconductor and MEAs-HEAs, we design a series of previously unreported A15-type V3(Os1-2xSixGex) (x = 0.333, 0.375, 0.425) MEA superconductors, which can be obtained by an arc melting method. Resistivity, magnetic susceptibility, and specific heat measurements indicate that all of them are type-II bulk superconductors. The superconducting transition temperature (Tc) exhibits an upward trend with the systematic reduction of Os concentration. Additionally, the upper critical field of the V3(Os0.333Si0.333Ge0.333) sample is larger than the Pauli limit, suggesting it may be robust against magnetic fields due to spin-orbit coupling induced by the heavy Os atoms. These findings not only advance our understanding of emergent phenomena in entropy-stabilized A15-type alloys but also expand the members of new superconductors.

cond-mat.supr-con

Oxygen permeability and stability in the entropy-stabilized Co-based Perovskite oxygen permeable membranes

Oxygen transport membranes (OTMs), enabling catalytic reaction and gas separation, support crucial chemical engineering processes and decarbonization technologies, but their applications are hindered by limited oxygen permeation fluxes and inadequate long-term stability during operation. Here, a series of high-entropy perovskite OTMs based on La0.5Sr0.5CoO3 were designed and synthesized by the simple sol-gel method. The impact of varying doping ratios on the structure, surface morphology, oxygen permeability, and stability of these high-entropy OTMs was thoroughly examined. At 950 {\deg}C, the optimal composition, La0.25Sr0.25Gd0.2Nd0.2Pr0.1CoO3, achieved oxygen permeation fluxes of 1.62 mL min-1 cm-2 under air/He gradient and 1.46 mL min-1 cm-2 under air/CO2, respectively. Remarkably, all high-entropy OTMs demonstrated stable operation for over 100 h in a pure CO2 environment without a significant decline in performance. This finding paves a new way to enhance the structural and oxygen permeation stability of OTMs, and further promotes the application of OTMs in oxy-fuel combustion technologies aimed at improving CO2 capture and storage efficiency.

cond-mat.mtrl-sci

BrainSegNet: A Novel Framework for Whole-Brain MRI Parcellation Enhanced by Large Models

Whole-brain parcellation from MRI is a critical yet challenging task due to the complexity of subdividing the brain into numerous small, irregular shaped regions. Traditionally, template-registration methods were used, but recent advances have shifted to deep learning for faster workflows. While large models like the Segment Anything Model (SAM) offer transferable feature representations, they are not tailored for the high precision required in brain parcellation. To address this, we propose BrainSegNet, a novel framework that adapts SAM for accurate whole-brain parcellation into 95 regions. We enhance SAM by integrating U-Net skip connections and specialized modules into its encoder and decoder, enabling fine-grained anatomical precision. Key components include a hybrid encoder combining U-Net skip connections with SAM's transformer blocks, a multi-scale attention decoder with pyramid pooling for varying-sized structures, and a boundary refinement module to sharpen edges. Experimental results on the Human Connectome Project (HCP) dataset demonstrate that BrainSegNet outperforms several state-of-the-art methods, achieving higher accuracy and robustness in complex, multi-label parcellation.

cs.CV

RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation

Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support. We present RIGOURATE, a two-stage multimodal framework that retrieves supporting evidence from a paper's body and assigns each claim an overstatement score. The framework consists of a dataset of over 10K claim-evidence sets from ICLR and NeurIPS papers, annotated using eight LLMs, with overstatement scores calibrated using peer-review comments and validated through human evaluation. It employes a fine-tuned reranker for evidence retrieval and a fine-tuned model to predict overstatement scores with justification. Compared to strong baselines, RIGOURATE enables improved evidence retrieval and overstatement detection. Overall, our work operationalises evidential proportionality and supports clearer, more transparent scientific communication.

cs.CL

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their capacity for complex reasoning and broaden their applicability across diverse scenarios. Dynamic sparse attention is a promising approach for reducing the computational cost of long-context. However, efficiently training LLMs with dynamic sparse attention on ultra-long contexts-especially in distributed settings-remains a significant challenge, due in large part to worker- and step-level imbalance. This paper introduces MTraining, a novel distributed methodology leveraging dynamic sparse attention to enable efficient training for LLMs with ultra-long contexts. Specifically, MTraining integrates three key components: a dynamic sparse training pattern, balanced sparse ring attention, and hierarchical sparse ring attention. These components are designed to synergistically address the computational imbalance and communication overheads inherent in dynamic sparse attention mechanisms during the training of models with extensive context lengths. We demonstrate the efficacy of MTraining by training Qwen2.5-3B, successfully expanding its context window from 32K to 512K tokens on a cluster of 32 A100 GPUs. Our evaluations on a comprehensive suite of downstream tasks, including RULER, PG-19, InfiniteBench, and Needle In A Haystack, reveal that MTraining achieves up to a 6x higher training throughput while preserving model accuracy. Our code is available at https://github.com/microsoft/MInference/tree/main/MTraining.

cs.CL

LogiNumSynth: Synthesizing Joint Logical-Numerical Reasoning Problems for Language Models

Joint logical-numerical reasoning remains a major challenge for language models, yet existing datasets rely on fixed rule sets and offer limited control over task complexity, constraining their generalizability for evaluation and training. We present LogiNumSynth, a flexible natural language problem synthesizer that synthesizes tasks requiring proficiency in joint logical reasoning (e.g., rule-based reasoning) and numerical reasoning (e.g., arithmetic computation). LogiNumSynth supports fine-grained control over reasoning world richness, logical reasoning depth, and the complexity of numerical computations, enabling flexible data synthesis across difficulty levels. We demonstrate three key contributions: (1) Synthesizer -- synthesizing fully controllable joint reasoning tasks over natural language; (2) Evaluation & Process Analysis -- evaluating both process accuracy and answer accuracy; (3) Targeted Training -- using synthesized data to enhance LLMs' reasoning performance. Experiments with multiple LLMs highlight persistent weaknesses in logical-numerical reasoning, showing that LogiNumSynth can serve as both a diagnostic tool and a source of targeted supervision for advancing integrated reasoning skills.

cs.CL

Strongly correlated electronic superconductivity in the noncentrosymmetric Re-Os-based high/medium-entropy alloys

The class of unconventional superconductors, particularly noncentrosymmetric superconductors, has been highly considered as potential materials for understanding the complex properties of quantum materials. Here, five previously unreported Re3.5Os3.5Ta0.5Hf0.5Nb3, Re3Os3Ta0.5Hf0.5Nb3, Re3.5Os3.5Mo0.5Hf0.5Nb3, Re3.5Os3.5Mo0.5W0.5Nb3, and Re3Os3Mo0.5Hf0.5Nb3 Re-Os-based high/medium-entropy alloys (MEAs-HEAs) with valence electron count ranging from 6.45 to 6.81 were synthesized and investigated using x-ray diffraction, transport, magnetization, and specific heat measurements. Our analyses confirm that all five compounds crystallize in a noncentrosymmetric {\alpha}-Mn-type structure and exhibit type-II superconductivity with Tc values from 4.20 K to 5.11 K, respectively. Unexpectedly, despite being immersed in an acidic environment for one month, the structures and superconducting properties of HEAs remain stable. Our findings indicate that the Tc increases with an increasing valence electron count in MEAs-HEAs. Furthermore, these noncentrosymmetric {\alpha}-Mn-type HEA superconductors have large Kadowaki-Woods ratios (KWR), implying the presence of strong electronic correlations.

cond-mat.supr-con

Efficient Byzantine Consensus MechanismBased on Reputation in IoT Blockchain

Blockchain technology has advanced rapidly in recent years and is now widely used in a variety of fields. Blockchain appears to be one of the best solutions for managing massive heterogeneous devices while achieving advanced data security and data reputation, particularly in the field of large-scale IoT (Internet of Things) networks. Despite the numerous advantages, there are still challenges while deploying IoT applications on blockchain systems due to the limited storage, power, and computing capability of IoT devices, and some of these problems are caused by the consensus algorithm, which plays a significant role in blockchain systems by ensuring overall system reliability and robustness. Nonetheless, most existing consensus algorithms are prone to poor node reliability, low transaction per second (TPS) rates, and scalability issues. Aiming at some critical problems in the existing consensus algorithms, this paper proposes the Efficient Byzantine Reputation-based Consensus (EBRC) mechanism to resolve the issues raised above. In comparison to traditional algorithms, we reinvented ways to evaluate node reliability and robustness and manage active nodes. Our experiments show that the EBRC algorithm has lower consensus delay, higher throughput, improved security, and lower verification costs. It offers new reference ideas for solving the Internet of Things+blockchain+Internet court construction problem.

cs.DC

Strong correlation behavior and Strong coupling superconductivity in (Ti1/4Hf1/4Nb1/4Ta1/4)1-xNix with the rich magnetic element Ni

Searching for new superconductors, especially unconventional superconductors, has been studied extensively for decades but remains one of the major outstanding challenges in condensed matter physics. Medium/high-entropy alloys (MEAs-HEAs) are new fertile soils of unconventional superconductors and generate widespread interest and questions on the existence of superconductivity in highly disordered materials. Here, we report on the effect of Ni-doped on the crystal structure and superconductivity properties of strongly coupled TiHfNbTa MEA. XRD results indicate that the maximum solid solution of (Ti1/4Hf1/4Nb1/4Ta1/4)1-xNix is about 7.7%. Resistivity, magnetic susceptibility, and specific heat measurements demonstrated that (Ti1/4Hf1/4Nb1/4Ta1/4)1-xNix HEAs are all bulk type-II superconductors and follow the trend of the increase of Tc with the increase of Ni-doped contents. The specific heat jump of all (Ti1/4Hf1/4Nb1/4Ta1/4)1-xNix are much larger than the BCS value of 1.43, suggesting all these HEAs are strongly coupled superconductors. Additionally, large Kadawaki-Woods ratio values suggest that there is a strong electron correlation effect in this system. The (Ti1/4Hf1/4Nb1/4Ta1/4)1-xNix HEA system is a new ideal material platform for the study of strong correlation behavior and strongly coupled superconductivity, which provides an insight into the physics of high-temperature superconductors or other unconventional superconductors.

cond-mat.supr-con

SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression

Large language models (LLMs) have achieved widespread adoption across numerous applications. However, many LLMs are vulnerable to malicious attacks even after safety alignment. These attacks typically bypass LLMs' safety guardrails by wrapping the original malicious instructions inside adversarial jailbreaks prompts. Previous research has proposed methods such as adversarial training and prompt rephrasing to mitigate these safety vulnerabilities, but these methods often reduce the utility of LLMs or lead to significant computational overhead and online latency. In this paper, we propose SecurityLingua, an effective and efficient approach to defend LLMs against jailbreak attacks via security-oriented prompt compression. Specifically, we train a prompt compressor designed to discern the "true intention" of the input prompt, with a particular focus on detecting the malicious intentions of adversarial prompts. Then, in addition to the original prompt, the intention is passed via the system prompt to the target LLM to help it identify the true intention of the request. SecurityLingua ensures a consistent user experience by leaving the original input prompt intact while revealing the user's potentially malicious intention and stimulating the built-in safety guardrails of the LLM. Moreover, thanks to prompt compression, SecurityLingua incurs only a negligible overhead and extra token cost compared to all existing defense methods, making it an especially practical solution for LLM defense. Experimental results demonstrate that SecurityLingua can effectively defend against malicious attacks and maintain utility of the LLM with negligible compute and latency overhead. Our code is available at https://aka.ms/SecurityLingua.

cs.CR

R-KV: Redundancy-aware KV Cache Compression for Reasoning Models

Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 16% of the KV cache. This KV-cache reduction also leads to a 90% memory saving and a 6.6X throughput over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets.

cs.CL