Searcharxiv⌕ Search

arXiv subjects

Lu Lu

Publications and source records attributed to Lu Lu.

At least 37 records · Page 2Linked to original sources

Data-driven discovery of governing differential equations across physical systems

Differential equations play a critical role in scientific discovery because they provide a mathematical framework to describe the behaviour of physical phenomena. As a promising alternative to traditional first principles, data-driven differential equation discovery has attracted increasing attention for its ability to infer governing laws directly from experimental or simulated data, especially when the underlying physics is unclear. However, the field has expanded rapidly along diverse methodological directions, particularly with the emergence of AI-based approaches, and still lacks a clear organizing perspective. In this Review, we propose a problem-oriented perspective on data-driven differential equation discovery. We first introduce a two-dimensional phase diagram of equation discoverability, where discovery problems are organized according to structural complexity and coefficient complexity. This phase diagram shows how the field has moved from the discovery of sparse equations with simple coefficients toward more complex governing laws with richer structures and more flexible parameterizations. It also clarifies why different methodological families succeed or fail in different problem settings. We then present the representation-evaluation-optimization (REO) framework as a fundamental abstraction of the discovery process. By identifying the core problems of equation discovery that persist across algorithmic variations, REO shifts the discussion from individual algorithms to the fundamental principles that determine discoverability. We connect these perspectives to applications across physics and adjacent sciences, and argue that the next challenge is not merely recovering equations, but using them to revise existing theories, distil mechanisms and form new scientific concepts.

cs.LG↗

Foundation Models for Wireless Communications: From PHY Intelligence to Network Autonomy

6G networks will introduce unprecedented complexity, which calls for a paradigm shift in network optimization and management. Artificial intelligence (AI)-based solutions, especially those enabled by the recently developed foundation models, have been recognized as promising candidates. Foundation models are large-scale AI models with general-purpose feature extraction capabilities, and once trained on massive amounts of data, they can be adapted to solve a wide range of downstream tasks, either in a zero-shot manner or with few-shot fine-tuning. This article provides a comprehensive overview of how foundation models are reshaping physical-layer processing and wireless resource management across three progressive paradigms. First, we examine the adaptation of off-the-shelf pre-trained foundation models to various wireless tasks. Second, we explore wireless-native foundation models, built from scratch on wireless data to bridge cross-domain modality gaps and capture universal wireless-domain physical characteristics. Third, we highlight agentic foundation models, which elevate static data processing into autonomous, reasoning-driven network orchestration. Furthermore, we discuss the impact of applying foundation models to emerging 6G frontiers, including integrated sensing and communications (ISAC), new multiple-input multiple-output (MIMO) architectures, semantic communications, and system-level network autonomy. Finally, we identify critical open challenges and opportunities, charting a promising path toward fully intelligent and adaptive wireless networks.

eess.SP↗

EarlyTom: Early Token Compression Completes Fast Video Understanding

Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amounts of visual tokens. Although recent approaches achieve extremely low token retention ratios while maintaining accuracy comparable to full-token baselines, most of them perform compression only at the late stage of prefilling, leaving the efficiency of the vision encoder unoptimized. In this paper, we first show that vision encoding contributes a large portion to the time-to-first-token (TTFT). Therefore, instead of compressing visual tokens only after the vision encoder, performing compression inside the encoder still leaves substantial room for exploration. Based on this insight, we propose EarlyTom, a training-free token compression framework that performs early-stage visual token compression inside the vision encoder, enabling significantly better TTFT reduction and higher throughput. In addition, we introduce a decoupled spatial token selection strategy that improves the overall compression effectiveness. EarlyTom reduces TTFT by up to 2.65x and FLOPs by up to 61% on a single NVIDIA A100 GPU for the LLaVA-OneVision-7B model, while maintaining accuracy comparable to the full-token baseline. These improvements substantially enhance the practicality of deploying Video-LLMs in real-world production scenarios.

cs.CV↗

Helmholzian Spectra of Graphs: Novel Properties

Let $\grad$, $\curl$, and $\dv$ be the graph-theoretic analogues of the gradient, curl, and divergence operators from multivariate calculus. The graph Laplacian $-\dv \grad$ gives rise to the celebrated Laplacian matrix, while the matrix representation of the graph Helmholtzian $\grad \grad^* + \curl^* \curl$ is called the Helmholtzian matrix. In this paper, we present a new graph-theoretic proof that the Helmholtzian matrix indeed represents the graph Helmholtzian. We then investigate the spectral properties of this matrix. Our main results are as follows: (i) a classification of graphs having exactly two distinct Helmholtzian eigenvalues; (ii) the nullity of the Helmholtzian matrix; and (iii) a combinatorial interpretation of the coefficients of the Helmholtzian polynomial. Furthermore, we determine the Helmholtzian spectrum for certain graph products and characterize Helmholtzian integral graphs, as well as derive bounds for the smallest Helmholtzian eigenvalue. Meanwhile, we pose some open problems for future research.

math.CO↗

Helmholzian spectra of graphs: basic properties

The Helmholtzian matrix of a graph $G=(V(G),E(G))$ is a graph-theoretic analogue of the vector Laplacian (or Helmholtz operator) [S. Li, L. Lu, J.F. Wang, A graph discretization of vector Laplacian, 379 (2026) 446--460]. Motivated by the applications of graph Helmholtzian in simplicial networks, we will investiagte its basic spectral properties. As the first graph matrix indexed by edge set, we find that Helmholtzian matrix is positive semi-definite and its non-negativity correlates with the odd cycles in $G$ and the orientation on $E(G)$, while its irreducibility relates to the signed graphs with loops. We show that the eigenvalues of Helmholtzian matrix are independent of the orientation and further investigate the eigenvalue interlacing under edge additions. One of striking findings is that the non-zero eigenvalues of the Laplacian matrix are those of Helmholtzian matrix of every graph. All these discoveries reveal that the Helmholtzian spectrum of $G$ balances and bridges the oriented graphs, weighted graphs and signed graphs as well as their adjacency or Laplacian spectra.

math.CO↗

End-to-end Listen, Look, Speak and Act

Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and interruptions. Realizing these capabilities is essential for building models simulating humans. We present ELLSA (End-to-end Listen, Look, Speak and Act), which, to our knowledge, is the first full-duplex, end-to-end model that simultaneously perceives and generates across vision, text, speech, and action within a single architecture, enabling interaction patterns previously out of reach, yielding more natural, human-like behaviors. At its core is a novel SA-MoE architecture (Self-Attention Mixture-of-Experts) that routes each modality to specialized experts and fuses them through a unified attention backbone. This provides a generalizable solution for joint multimodal perception and concurrent generation, leveraging strong pre-trained components while enabling efficient modality integration and mitigating modality interference. On speech-interaction and robot-manipulation benchmarks, ELLSA matches modality-specific baselines, while uniquely supporting advanced multimodal and full-duplex behaviors such as dialogue and action turn-taking, defective instruction rejection, speaking-while-acting, context-grounded visual question answering, and action barge-ins. We contend that ELLSA represents a step toward more natural and general interactive intelligence, contributing to the broader pursuit of artificial general intelligence. All data, code and model checkpoints will be released at https://github.com/bytedance/SALMONN/tree/ELLSA.

cs.AI↗

Nested Fourier-enhanced neural operator for efficient modeling of radiation transfer in fires

Computational fluid dynamics (CFD) has become an essential tool for predicting fire behavior, yet maintaining both efficiency and accuracy remains challenging. A major source of computational cost in fire simulations is the modeling of radiation transfer, which is usually the dominant heat transfer mechanism in fires. Solving the high-dimensional radiative transfer equation (RTE) with traditional numerical methods can be a performance bottleneck. Here, we present a machine learning framework based on Fourier-enhanced multiple-input neural operators (Fourier-MIONet) as an efficient alternative to direct numerical integration of the RTE. We first investigate the performance of neural operator architectures for a small-scale 2D pool fire and find that Fourier-MIONet provides the most accurate radiative solution predictions. The approach is then extended to 3D CFD fire simulations, where the computational mesh is locally refined across multiple levels. In these high-resolution settings, monolithic surrogate models for direct field-to-field mapping become difficult to train and computationally inefficient. To address this issue, a nested Fourier-MIONet is proposed to predict radiation solutions across multiple mesh-refinement levels. We validate the approach on 3D McCaffrey pool fires simulated with FireFOAM, including fixed fire sizes and a unified model trained over a continuous range of heat release rates (HRRs). The proposed method achieves global relative errors of 2-4% for 3D varying-HRR scenarios while providing faster inference than the estimated cost of one finite-volume radiation solve in FireFOAM for the 16-solid-angle case. With fast and accurate inference, the surrogate makes higher-fidelity radiation treatments practical and enables the incorporation of more spectrally resolved radiation models into CFD fire simulations for engineering applications.

physics.flu-dyn↗

Optimization of Magnetic Milli-Spinner for Robotic Endovascular Intervention

Vascular diseases such as atherosclerosis, thrombosis, and aneurysms can lead to life-threatening medical events. Conventional catheter- or guidewire-based interventional devices often struggle to navigate through highly tortuous vasculature. The recently developed multifunctional magnetic milli-spinner offers a promising wireless solution by integrating a central through-hole and side slits into a cylindrical body with helical fins, enabling rapid and stable navigation for clot debulking, targeted drug delivery, and aneurysm treatment. Here, we combine computational fluid dynamics simulations with experimental validation to optimize the milli-spinner's structural design for high-velocity propulsion and high-efficiency clot debulking in tubular flow environments. By systematically investigating the effects of through-hole radius, fin number, fin helical angle, and slit dimension on propulsion performance, the optimized milli-spinner achieves swimming velocities of 55 cm/s (175 body lengths per second) in saline water and 44 cm/s (140 body lengths per second) in a fluid with viscosity (3.5 mPa.s) comparable to that of arterial blood at high shear rates, far exceeding existing untethered magnetic robots in tubular environments (less than 80 body lengths per second). This exceptional velocity enables stable upstream operation against strong physiological flows representative of major arteries and veins, establishing the milli-spinner as a robust untethered navigation platform for operation in high-flow, tortuous vasculature.

physics.med-ph↗

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues - such as emotion, tone, and speaker attributes - and to respond appropriately in both content and style remains under-explored. Progress is further hindered by the scarcity of high-quality and expressive demonstrations. To address this, we introduce a new reinforcement learning (RL) framework for paralinguistic-aware S2S, ParaS2S, which evaluates and optimizes both response content and speaking style directly at the waveform level. We first construct ParaS2SBench, a benchmark that evaluates the naturalness of input-output pairs in terms of content and speaking style using expressive and challenging queries. For the automatic judge, we propose a PolyTone training strategy and a multi-stage framework, preventing the style hallucination of end-to-end audio LLM judging. Our judge correlates well with human preferences and is scalable, enabling the model to interact and learn from unlabeled speech via RL. Experiments show that existing S2S models fail to respond appropriately to paralinguistic attributes, performing no better than pipeline-based baselines. Our RL approach (ParaS2SAlign) achieves a 10% relative improvement in the appropriateness of response content and speaking style on ParaS2SBench over supervised fine-tuning (SFT), surpassing all prior models while requiring substantially fewer paired demonstrations than pure SFT. Our findings highlight the need for a scalable and accurate automatic evaluator for speech-to-speech interaction.

eess.AS↗

Subspace variations of the weighted skew Bollobás theorem

Let $V$ be a finite-dimensional real vector space. A collection $\mathcal{P} = \{(A_i,B_i)\}_{i=1}^m$ of pairs of subspaces of $V$ is called a skew Bollobás system if $\dim(A_i\cap B_i)=0$ for each $i\in [m]$ and $\dim(A_i\cap B_j)>0$ for all $1\leq i<j \leq m$. Assume that $V = V^{(1)}\oplus \cdots \oplus V^{(r)}$ and $\mathcal{P}= \{(A_i,B_i)\}_{i=1}^m$ is a skew Bollobás system of subspaces of $V$ satisfying $ A_i = \bigoplus_{k=1}^r (A_i \cap V^{(k)})$ and $ B_i = \bigoplus_{k=1}^r (B_i \cap V^{(k)})$ for each $i\in [m]$. Denote $a_{i,k} = \dim(A_i \cap V^{(k)})$ and $b_{i,k} = \dim(B_i \cap V^{(k)})$. Suppose that $a_{1,k} \le \cdots \le a_{m,k}$ and $b_{1,k} \ge \cdots \ge b_{m,k}$ for each $k\in [r]$. Using the exterior algebraic method developed by Lovász and Scott--Wilmer, we prove that $$ \sum_{i=1}^{m} \frac{1}{\prod_{k=1}^{r} \binom{a_{i,k}+b_{i,k}}{a_{i,k}}} \le 1 . $$ This generalizes the results of Alon (JCTA, 1985) and Scott--Wilmer (JLMS, 2021) to multipart weighted setting. Secondly, we solve a conjecture of Hegedüs (AJC, 2015) concerning projective subspaces, showing that any skew Bollobás system of projective subspaces in an $n$-dimensional projective space contains at most $2^{n+1} - 2$ pairs. Thirdly, we prove that if $\mathcal{P}= \{(A_i,B_i)\}_{i=1}^m$ is a skew Bollobás system of subspaces of $V$ with $a_i=\dim (A_i)$ and $b_i=\dim (B_i)$, then $$ \sum_{i=1}^m \frac{1}{(a_i+ b_i+1)\binom{a_i+b_i}{a_i}} \le 1. $$ This gives an extension to the subspace setting of the results of Hegedüs--Frankl (EUJC, 2024) and Yue (DM, 2026). Finally, we extend the above inequality to systems of $d$-tuples of subspaces, giving a unified bound that implies the corresponding results for $d$-tuples of subsets.

math.CO↗

RED-DiffEq: Regularization by denoising diffusion models for solving inverse PDE problems with application to full waveform inversion

Partial differential equation (PDE)-governed inverse problems are fundamental across various scientific and engineering applications; yet they face significant challenges due to nonlinearity, ill-posedness, and sensitivity to noise. Here, we introduce a new computational framework, RED-DiffEq, by integrating physics-driven inversion and data-driven learning. RED-DiffEq leverages pretrained diffusion models as a regularization mechanism for PDE-governed inverse problems. We apply RED-DiffEq to solve the full waveform inversion problem in geophysics, a challenging seismic imaging technique that seeks to reconstruct high-resolution subsurface velocity models from seismic measurement data. Our method shows enhanced accuracy and robustness compared to conventional methods. Additionally, it exhibits strong generalization ability to more complex velocity models that the diffusion model is not trained on. Our framework can also be directly applied to diverse PDE-governed inverse problems.

cs.LG↗

Active operator learning with predictive uncertainty quantification for partial differential equations

With the increased prevalence of neural operators being used to provide rapid solutions to partial differential equations (PDEs), understanding the accuracy of model predictions and the associated error levels is necessary for deploying reliable surrogate models in scientific applications. Existing uncertainty quantification (UQ) frameworks employ ensembles or Bayesian methods, which can incur substantial computational costs during both training and inference. We propose a lightweight predictive UQ method tailored for Deep operator networks (DeepONets) that also generalizes to other operator networks. Numerical experiments on linear and nonlinear PDEs demonstrate that the framework's uncertainty estimates are unbiased and provide accurate out-of-distribution uncertainty predictions with a sufficiently large training dataset. Our framework provides fast inference and uncertainty estimates that can efficiently drive outer-loop analyses that would be prohibitively expensive with conventional solvers. We demonstrate how predictive uncertainties can be used in the context of Bayesian optimization and active learning problems to yield improvements in accuracy and data-efficiency for outer-loop optimization procedures. In the active learning setup, we extend the framework to Fourier Neural Operators (FNO) and describe a generalized method for other operator networks. To enable real-time deployment, we introduce an inference strategy based on precomputed trunk outputs and a sparse placement matrix, reducing evaluation time by more than a factor of five. Our method provides a practical route to uncertainty-aware operator learning in time-sensitive settings.

cs.LG↗

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a reinforcement-learning framework for audio-driven portrait animation built on a multimodal backbone for autoregressive audio-to-video generation. FlowPortrait introduces a human-aligned evaluation system based on Multimodal Large Language Models (MLLMs) to assess lip-sync accuracy, expressiveness, and motion quality. These signals are combined with perceptual and temporal consistency regularizers to form a stable composite reward, which is used to post-train the generator via Group Relative Policy Optimization (GRPO). Extensive experiments, including both automatic evaluations and human preference studies, demonstrate that FlowPortrait consistently produces higher-quality talking-head videos, highlighting the effectiveness of reinforcement learning for portrait animation.

cs.CV↗

Probabilistic modeling of Cherenkov emission from particle showers

Subatomic particles can interact with target nuclei in matter or decay in flight, and an individual high-energy particle can induce a particle shower composed of numerous, lower-energy secondaries. These particle showers broadly exhibit universality across diverse media, including air, water, ice, and other materials, with their development governed by the Standard Model. Full Monte Carlo simulation of particle showers, where each secondary is individually tracked and propagated, can be a computational challenge to perform at scale. Experiments thus resort to parametrized approximations when efficient simulation becomes necessary. Here, we construct distributions of parameters capable of describing the Cherenkov light yield from particle showers in ice or water. Sampling from the distributions allows for a much improved description of event-to-event fluctuations, in amplitude and shape, along the shower axis. Including these effects is essential for a more accurate simulation of signal and background events in current and next-generation neutrino telescopes.

astro-ph.HE↗

What Quality Engineers Need to Know about Degradation Models

Degradation models play a critical role in quality engineering by enabling the assessment and prediction of system reliability based on data. The objective of this paper is to provide an accessible introduction to degradation models. We explore commonly used degradation data types, including repeated measures degradation data and accelerated destructive degradation test data, and review modeling approaches such as general path models and stochastic process models. Key inference problems, including reliability estimation and prediction, are addressed. Applications across diverse fields, including material science, renewable energy, civil engineering, aerospace, and pharmaceuticals, illustrate the broad impact of degradation models in industry. We also discuss best practices for quality engineers, software implementations, and challenges in applying these models. This paper aims to provide quality engineers with a foundational understanding of degradation models, equipping them with the knowledge necessary to apply these techniques effectively in real-world scenarios.

stat.AP↗

Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual tokens engenders a substantial computational load and key-value (KV) cache footprint bottleneck. Existing approaches focus on token-wise optimization, leveraging diverse intricate token pruning techniques to eliminate non-crucial visual tokens. Nevertheless, these methods often unavoidably undermine the integrity of the KV cache, resulting in failures in long-text generation tasks. To this end, we conduct an in-depth investigation towards the attention mechanism of the model from a new perspective, and discern that attention within more than half of all decode layers are semantic similar. Upon this finding, we contend that the attention in certain layers can be streamlined by inheriting the attention from their preceding layers. Consequently, we propose Lazy Attention, an efficient attention mechanism that enables cross-layer sharing of similar attention patterns. It ingeniously reduces layer-wise redundant computation in attention. In Lazy Attention, we develop a novel layer-shared cache, Q Cache, tailored for MLLMs, which facilitates the reuse of queries across adjacent layers. In particular, Q Cache is lightweight and fully compatible with existing inference frameworks, including Flash Attention and KV cache. Additionally, our method is highly flexible as it is orthogonal to existing token-wise techniques and can be deployed independently or combined with token pruning approaches. Empirical evaluations on multiple benchmarks demonstrate that our method can reduce KV cache usage by over 35% and achieve 1.5x throughput improvement, while sacrificing only approximately 1% of performance on various MLLMs. Compared with SOTA token-wise methods, our technique achieves superior accuracy preservation.

cs.CV↗

PI-MFM: Physics-informed multimodal foundation model for solving partial differential equations

Partial differential equations (PDEs) govern a wide range of physical systems, and recent multimodal foundation models have shown promise for learning PDE solution operators across diverse equation families. However, existing multi-operator learning approaches are data-hungry and neglect physics during training. Here, we propose a physics-informed multimodal foundation model (PI-MFM) framework that directly enforces governing equations during pretraining and adaptation. PI-MFM takes symbolic representations of PDEs as the input, and automatically assembles PDE residual losses from the input expression via a vectorized derivative computation. These designs enable any PDE-encoding multimodal foundation model to be trained or adapted with unified physics-informed objectives across equation families. On a benchmark of 13 parametric one-dimensional time-dependent PDE families, PI-MFM consistently outperforms purely data-driven counterparts, especially with sparse labeled spatiotemporal points, partially observed time domains, or few labeled function pairs. Physics losses further improve robustness against noise, and simple strategies such as resampling collocation points substantially improve accuracy. We also analyze the accuracy, precision, and computational cost of automatic differentiation and finite differences for derivative computation within PI-MFM. Finally, we demonstrate zero-shot physics-informed fine-tuning to unseen PDE families: starting from a physics-informed pretrained model, adapting using only PDE residuals and initial/boundary conditions, without any labeled solution data, rapidly reduces test errors to around 1% and clearly outperforms physics-only training from scratch. These results show that PI-MFM provides a practical and scalable path toward data-efficient, transferable PDE solvers.

cs.LG↗

Spectral bounds for vertex-weighted Laplacians of simplicial complexes

The vertex-weighted Laplacian naturally extends the combinatorial Laplacian for simplicial complexes. Inspired by Lew's foundational techniques for vertex-weighted Laplacians, we present a comprehensive spectral analysis of this operator. First, we determine how basic operations, including joins, complements, and Alexander duals, affect its spectrum. This yields a sharp upper bound on the spectral radius in terms of vertex weights, along with a lower bound on the multiplicity at which this bound is attained. Second, we establish a sharp lower bound for the spectral gap and characterize when the equality holds. Third, explicit lower bounds for the remaining eigenvalues are derived, linking the vertex-weighted Laplacian spectrum to that of a related weighted graph. Finally, we reveal new spectral relations between a simplicial complex and its subcomplexes. These results not only generalize numerous known theorems on combinatorial Laplacians but also provide deeper spectral insights into simplicial structures, ultimately unifying and extending a broad range of earlier work in this field.

math.CO↗