Searcharxiv⌕ Search

arXiv subjects

Lu Lu

Publications and source records attributed to Lu Lu.

At least 55 records · Page 3Linked to original sources

Spectral bounds for vertex-weighted Laplacians of simplicial complexes

The vertex-weighted Laplacian naturally extends the combinatorial Laplacian for simplicial complexes. Inspired by Lew's foundational techniques for vertex-weighted Laplacians, we present a comprehensive spectral analysis of this operator. First, we determine how basic operations, including joins, complements, and Alexander duals, affect its spectrum. This yields a sharp upper bound on the spectral radius in terms of vertex weights, along with a lower bound on the multiplicity at which this bound is attained. Second, we establish a sharp lower bound for the spectral gap and characterize when the equality holds. Third, explicit lower bounds for the remaining eigenvalues are derived, linking the vertex-weighted Laplacian spectrum to that of a related weighted graph. Finally, we reveal new spectral relations between a simplicial complex and its subcomplexes. These results not only generalize numerous known theorems on combinatorial Laplacians but also provide deeper spectral insights into simplicial structures, ultimately unifying and extending a broad range of earlier work in this field.

math.CO↗

FunDiff: Diffusion Models over Function Spaces for Physics-Informed Generative Modeling

Recent advances in generative modeling -- particularly diffusion models and flow matching -- have achieved remarkable success in synthesizing discrete data such as images and videos. However, adapting these models to physical applications remains challenging, as the quantities of interest are continuous functions governed by complex physical laws. Here, we introduce $\textbf{FunDiff}$, a novel framework for generative modeling in function spaces. FunDiff combines a latent diffusion process with a function autoencoder architecture to handle input functions with varying discretizations, generate continuous functions evaluable at arbitrary locations, and seamlessly incorporate physical priors. These priors are enforced through architectural constraints or physics-informed loss functions, ensuring that generated samples satisfy fundamental physical laws. We theoretically establish minimax optimality guarantees for density estimation in function spaces, showing that diffusion-based estimators achieve optimal convergence rates under suitable regularity conditions. We demonstrate the practical effectiveness of FunDiff across diverse applications in fluid dynamics and solid mechanics. Empirical results show that our method generates physically consistent samples with high fidelity to the target distribution and exhibits robustness to noisy and low-resolution data. Code and datasets are publicly available at https://github.com/sifanexisted/fundiff.

cs.LG↗

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud-based benchmarking infrastructure comprising over 700,000 expert-curated tasks spanning 24 primary and 91 secondary specialties, with dedicated tracks for LLMs, multimodal models, and agents. Items undergo multi-stage refinement and multi-round review by clinicians from more than 500 institutions, and open-ended responses are scored by an LLM-as-a-judge calibrated to human ratings. We evaluate 15 frontier models. Base LLMs reach a mean overall score of 54.1/100 (best: Claude Sonnet 4.5, 62.5/100), but safety and ethics remain low (18.4/100). Multimodal models perform worse overall (mean 47.5/100; best: GPT-5, 54.9/100), with solid perception yet weaker cross-modal reasoning. Agents built on the same backbones substantially improve end-to-end performance (mean 79.8/100), with Claude Sonnet 4.5-based agents achieving up to 85.3/100 overall and 88.9/100 on safety tasks. MedBench v4 thus reveals persisting gaps in multimodal reasoning and safety for base models, while showing that governance-aware agentic orchestration can markedly enhance benchmarked clinical readiness without sacrificing capability. By aligning tasks with Chinese clinical guidelines and regulatory priorities, the platform offers a practical reference for hospitals, developers, and policymakers auditing medical AI.

cs.CL↗

TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine

Large language models (LLMs) have demonstrated exceptional capabilities in general domains, yet their application in highly specialized and culturally-rich fields like Traditional Chinese Medicine (TCM) requires rigorous and nuanced evaluation. Building upon prior foundational work such as TCM-3CEval, which highlighted systemic knowledge gaps and the importance of cultural-contextual alignment, we introduce TCM-5CEval, a more granular and comprehensive benchmark. TCM-5CEval is designed to assess LLMs across five critical dimensions: (1) Core Knowledge (TCM-Exam), (2) Classical Literacy (TCM-LitQA), (3) Clinical Decision-making (TCM-MRCD), (4) Chinese Materia Medica (TCM-CMM), and (5) Clinical Non-pharmacological Therapy (TCM-ClinNPT). We conducted a thorough evaluation of fifteen prominent LLMs, revealing significant performance disparities and identifying top-performing models like deepseek\_r1 and gemini\_2\_5\_pro. Our findings show that while models exhibit proficiency in recalling foundational knowledge, they struggle with the interpretative complexities of classical texts. Critically, permutation-based consistency testing reveals widespread fragilities in model inference. All evaluated models, including the highest-scoring ones, displayed a substantial performance degradation when faced with varied question option ordering, indicating a pervasive sensitivity to positional bias and a lack of robust understanding. TCM-5CEval not only provides a more detailed diagnostic tool for LLM capabilities in TCM but aldso exposes fundamental weaknesses in their reasoning stability. To promote further research and standardized comparison, TCM-5CEval has been uploaded to the Medbench platform, joining its predecessor in the "In-depth Challenge for Comprehensive TCM Abilities" special track.

cs.CL↗

Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking

In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities of foundation models, towards more comprehensive understanding, more natural generation and more human-like interaction. Audio, as a modality rich in semantic, emotional, and contextual cues, plays a vital role in achieving naturalistic and embodied machine intelligence. This survey provides a comprehensive review of recent progress in integrating audio into LLMs, with a focus on four key areas: audio comprehension, audio generation, speech-based interaction, and audio-visual understanding. We analyze how LLMs are reshaping audio perception and reasoning, enabling systems to understand sound at a deeper semantic level, generate expressive audio outputs, and engage in human-like spoken interaction. Furthermore, we explore how the fusion of audio and visual modalities enhances situational awareness and cross-modal reasoning, pushing the boundaries of multimodal intelligence. This survey not only synthesizes existing research but also identifies critical challenges and future directions for building audio-native AGI systems capable of perceiving, understanding, and interacting through sound as naturally as humans do.

eess.AS↗

2D-to-3D transformation of ring origami via snap-folding instabilities

Ring origami, consisting of closed-loop rods, is a class of shape-morphing structures that undergo shape transformation through folding enabled by snap-buckling instabilities, referred to as snap-folding instabilities. Previous studies have shown that 2D ring origami composed of rod segments with in-plane natural curvature (i.e., the stress-free curved state lies in the plane of the planar ring) can achieve diverse and intriguing 2D-to-2D shape transformations. Here, we propose a 2D-to-3D shape transformation strategy for ring origami by introducing out-of-plane natural curvature (i.e., the stress-free curved state lies in a plane perpendicular to the planar ring) into the rod segments. Due to natural curvature-induced out-of-plane bending moments, a 2D elastic ring spontaneously snaps out-of-plane and reaches equilibrium in a 3D configuration. These snapping-induced out-of-plane shape transitions not only enable self-guided, spontaneous shape morphing, but also allow the construction of complex structures from simple geometries, making them promising for the design of functional deployable and foldable structures. By combining a multi-segment Kirchhoff rod model with finite element simulations and experiments, we systematically investigate the 3D equilibrium states and transition behavior of these systems. Using square and hexagonal rings as representative examples, we demonstrate that by rationally designing the out-of-plane natural curvature of rod segments, 2D rings can exhibit a range of functional behaviors, including spontaneous 2D-to-3D shape transformation (e.g., planar square to sphere) via snap-folding, multistability with various 3D configurations, and monostability with a compact zero-energy 3D configuration.

physics.app-ph↗

Identifying Trustworthiness Challenges in Deep Learning Models for Continental-Scale Water Quality Prediction

Water quality is foundational to environmental sustainability, ecosystem resilience, and public health. Deep learning offers transformative potential for large-scale water quality prediction and scientific insights generation. However, their widespread adoption in high-stakes operational decision-making, such as pollution mitigation and equitable resource allocation, is prevented by unresolved trustworthiness challenges, including performance disparity, robustness, uncertainty, interpretability, generalizability, and reproducibility. In this work, we present a multi-dimensional, quantitative evaluation of trustworthiness benchmarking three state-of-the-art deep learning architectures: recurrent (LSTM), operator-learning (DeepONet), and transformer-based (Informer), trained on 37 years of data from 482 U.S. basins to predict 20 water quality variables. Our investigation reveals systematic performance disparities tied to process complexity, data availability, and basin heterogeneity. Management-critical variables remain the least predictable and most uncertain. Robustness tests reveal pronounced sensitivity to outliers and corrupted targets; notably, the architecture with the strongest baseline performance (LSTM) proves most vulnerable under data corruption. Attribution analyses align for simple variables but diverge for nutrients, underscoring the need for multi-method interpretability. Spatial generalization to ungauged basins remains poor across all models. This work serves as a timely call to action for advancing trustworthy data-driven methods for water resources management and provides a pathway to offering critical insights for researchers, decision-makers, and practitioners seeking to leverage artificial intelligence (AI) responsibly in environmental management.

cs.LG↗

TANTE: Time-Adaptive Operator Learning via Neural Taylor Expansion

Operator learning for time-dependent partial differential equations (PDEs) has seen rapid progress in recent years, enabling efficient approximation of complex spatiotemporal dynamics. However, most existing methods rely on fixed time step sizes during rollout, which limits their ability to adapt to varying temporal complexity and often leads to error accumulation. Here, we propose the Time-Adaptive Transformer with Neural Taylor Expansion (TANTE), a novel operator-learning framework that produces continuous-time predictions with adaptive step sizes. TANTE predicts future states by performing a Taylor expansion at the current state, where neural networks learn both the higher-order temporal derivatives and the local radius of convergence. This allows the model to dynamically adjust its rollout based on the local behavior of the solution, thereby reducing cumulative error and improving computational efficiency. We demonstrate the effectiveness of TANTE across a wide range of PDE benchmarks, achieving superior accuracy and adaptability compared to fixed-step baselines, delivering accuracy gains of 60-80 % and speed-ups of 30-40 % at inference time. The code is publicly available at https://github.com/zwu88/TANTE for transparency and reproducibility.

cs.LG↗

Learning the detector in optical tomography

We propose a method to reconstruct the optical absorption of a highly-scattering medium probed by diffuse light. The method consists of learning the optical detection system and then using this result to reconstruct the absorption. Our results are illustrated by numerical simulations.

physics.optics↗

Neural Networks as Surrogate Solvers for Time-Dependent Accretion Disk Dynamics

Accretion disks are ubiquitous in astrophysics, appearing in diverse environments from planet-forming systems to X-ray binaries and active galactic nuclei. Traditionally, modeling their dynamics requires computationally intensive (magneto)hydrodynamic simulations. Recently, Physics-Informed Neural Networks (PINNs) have emerged as a promising alternative. This approach trains neural networks directly on physical laws without requiring data. We for the first time demonstrate PINNs for solving the two-dimensional, time-dependent hydrodynamics of non-self-gravitating accretion disks. Our models provide solutions at arbitrary times and locations within the training domain, and successfully reproduce key physical phenomena, including the excitation and propagation of spiral density waves and gap formation from disk-companion interactions. Notably, the boundary-free approach enabled by PINNs naturally eliminates the spurious wave reflections at disk edges, which are challenging to suppress in numerical simulations. These results highlight how advanced machine learning techniques can enable physics-driven, data-free modeling of complex astrophysical systems, potentially offering an alternative to traditional numerical simulations in the future.

astro-ph.EP↗

Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents

Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To train agents for this dynamic process, particularly in multi-modal contexts, we introduce a sandbox environment for reinforcement learning (RL) that supports interleaved speech-text rollouts. Our core strategy, Turn-level Adjudicated Reinforcement Learning (TARL), addresses the challenge of credit assignment in long-horizon tasks by employing a Large Language Model (LLM) as a judge to provide turn-level evaluation. To enhance exploration, we integrate a mixed-task training curriculum with mathematical reasoning problems. This unified approach boosts the task pass rate on the text-based $τ$-bench by over 6% compared to strong RL baselines. Crucially, we demonstrate our framework's suitability for fine-tuning a multi-modal foundation model for agentic tasks. By training a base multi-modal LLM on interleaved speech-text rollouts, we equip it with tool-use abilities, paving the way for more natural, voice-driven interactive agents.

cs.CL↗

Extreme High-Energy Neutrinos: IceCube vs. KM3NeT

We review the state of the art in the detection of extreme high-energy neutrinos, focusing on the IceCube and KM3NeT neutrino telescopes. IceCube, operating deep in Antarctic ice, and KM3NeT, a new array in the Mediterranean Sea, employ distinct designs to capture Cherenkov light from neutrino interactions. We examine their detector architectures, readout and reconstruction performance for PeV-scale and higher-energy neutrinos. Recent candidate events above 5 PeV are highlighted. These include a ~120 PeV muon track observed by KM3NeT in 2023, and IceCube's highest-energy detections, which comprise several-PeV showers and tracks. We outline current approaches to neutrino energy reconstruction and explore scenarios that might explain the apparent differences in observed event characteristics. Finally, we summarize future prospects for extreme-energy neutrino observations and their implications for astrophysical source populations and cosmogenic neutrinos.

astro-ph.HE↗

Measuring the Astrophysical Galactic Plane Neutrino Flux and Searching for Galactic PeVatrons using the IceCube Multi-Flavor Astrophysical Neutrino Sample

The IceCube Neutrino Observatory has provided new insights into the high-energy universe, in particular, unveiling neutrinos from the galactic plane. However, galactic neutrino sources are still unresolved. The recent detection of multi-PeV photons by LHAASO from the Cygnus region highlights its potential as a galactic neutrino source. Additionally, LHAASO, HAWC, and HESS have reported over forty galactic gamma-ray sources with energies above 100 TeV. Detecting neutrinos correlated with high-energy gamma-ray sources would provide compelling evidence of cosmic-ray acceleration in these galactic sources. In this work, we compile a 12.3-year, full-sky, all-flavor dataset, the IceCube Multi-Flavor Astrophysics Neutrino sample (ICEMAN). ICEMAN is the combination of three largely independent neutrino samples of different event morphologies and builds upon the previous work of the DNN-based cascade sample, Enhanced Starting Track Event Selection, and the Northern Track sample. Recent improvements in ice modeling and detector calibration are also incorporated into the cascade reconstruction. In addition to revisiting the galactic plane, we adopt two different analysis methods to search for galactic PeVatrons. First, we use a template-based approach to probe the Cygnus Cocoon region. Second, we use a point source hypothesis to find correlations between IceCube neutrinos and gamma-ray sources detected at energies greater than 100 TeV.

astro-ph.HE↗

RAMS: Residual-based adversarial-gradient moving sample method for scientific machine learning in solving partial differential equations

Physics-informed neural networks (PINNs) and neural operators, two leading scientific machine learning (SciML) paradigms, have emerged as powerful tools for solving partial differential equations (PDEs). Although increasing the training sample size generally enhances network performance, it also increases computational costs for physics-informed or data-driven training. To address this trade-off, different sampling strategies have been developed to sample more points in regions with high PDE residuals. However, existing sampling methods are computationally demanding for high-dimensional problems, such as high-dimensional PDEs or operator learning tasks. Here, we propose a residual-based adversarial-gradient moving sample (RAMS) method, which moves samples according to the adversarial gradient direction to maximize the PDE residual via gradient-based optimization. RAMS can be easily integrated into existing sampling methods. Extensive experiments, ranging from PINN applied to high-dimensional PDEs to physics-informed and data-driven operator learning problems, have been conducted to demonstrate the effectiveness of RAMS. Notably, RAMS represents the first efficient adaptive sampling approach for operator learning, marking a significant advancement in the SciML field.

cs.CE↗

Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters

Multilingual translation stands as a challenging task for large language models (LLMs) to handle intricate language patterns and stilted translations that arise in automated translations. In this paper, we introduce Seed-X, a family of open-source LLMs comprising instruct and reasoning models, pushing the limits of translation capability with 7B parameter size. The base model is pre-trained on a diverse, high-quality dataset encompassing both monolingual and bilingual content across 28 languages, harnessing the full potential of multilingual data. The instruct model is then finetuned to translate by Chain-of-Thought (CoT) reasoning and further enhanced through reinforcement learning (RL) to achieve better generalization across diverse language pairs. Seed-X achieves performance comparable to leading closed-source models, including Gemini-2.5 and GPT-4o, across 28 languages, and significantly outperforms larger open-source models in both automatic metrics and human evaluations. We share the best practices through our optimization process, and make the parameter public available for advancing translation research and applications.

cs.CL↗

DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

We present DuPO, a dual learning-based preference optimization framework that generates annotation-free feedback via a generalized duality. DuPO addresses two key limitations: Reinforcement Learning with Verifiable Rewards (RLVR)'s reliance on costly labels and applicability restricted to verifiable tasks, and traditional dual learning's restriction to strictly dual task pairs (e.g., translation and back-translation). Specifically, DuPO decomposes a primal task's input into known and unknown components, then constructs its dual task to reconstruct the unknown part using the primal output and known information (e.g., reversing math solutions to recover hidden variables), broadening applicability to non-invertible tasks. The quality of this reconstruction serves as a self-supervised reward to optimize the primal task, synergizing with LLMs' ability to instantiate both tasks via a single model. Empirically, DuPO achieves substantial gains across diverse tasks: it enhances the average translation quality by 2.13 COMET over 756 directions, boosts the mathematical reasoning accuracy by an average of 6.4 points on three challenge benchmarks, and enhances performance by 9.3 points as an inference-time reranker (trading computation for accuracy). These results position DuPO as a scalable, general, and annotation-free paradigm for LLM optimization.

cs.LG↗

Combinatorial Laplacians and Relative Homology of Complex Pairs

As a discretization of the Hodge Laplacian, the combinatorial Laplacian of simplicial complexes has garnered significant attention. In this paper, we study combinatorial Laplacians for complex pairs $(X, A)$, where $A$ is a subcomplex of a simplicial complex $X$. We establish a relative version of the matrix-tree theorem for complex pairs, which generalizes both the matrix-tree theorem for simplicial complexes proved by Duval, Klivans, and Martin (2009) and the result for Dirichlet eigenvalues of graph pairs by Chung (1996). Furthermore, we derive several lower bounds for the spectral gaps of complex pairs and characterize the equality case for one sharp lower bound. As by-products, we obtain sufficient conditions for the vanishing of relative homology. Our results demonstrate that the combinatorial Laplacians for complex pairs are closely related to relative homology.

math.CO↗

Measurement of All Flavor PeV Neutrino Flux using Combined Datasets from IceCube

Recently, the IceCube Neutrino Observatory has reported a deviation from the single power law in the extragalactic diffuse neutrino flux. A neural network-based event selection of contained and uncontained cascade events from IceCube, in which uncontained events have interaction vertices at the edge or outside of the detector instrumentation volume, has a factor ~3 gain in effective area over the cascade events used in the novel combined tracks and cascades selection which reported the deviation. Systematic improvements and rigorously updated modeling of the atmospheric neutrino background is incorporated into this high statistics contained and uncontained cascade event selection to clarify features of the astrophysical neutrino spectrum across energies from 1 TeV up to 100 PeV.

astro-ph.HE↗