Searcharxiv⌕ Search

arXiv subjects

Lu Lu

Publications and source records attributed to Lu Lu.

At least 73 records · Page 4Linked to original sources

Spectral extremal problem of the $p$th power of cycles

For a cycle $C_k$ on $k$ vertices, its $p$-th power, denoted $C_k^p$, is the graph obtained by adding edges between all pairs of vertices at distance at most $p$ in $C_k$. Let $\ex(n, F)$ and $\spex(n, F)$ denote the maximum possible number of edges and the maximum possible spectral radius, respectively, among all $n$-vertex $F$-free graphs. In this paper, we determine precisely the unique extremal graph achieving $\ex(n, C_k^p)$ and $\spex(n, C_k^p)$ for sufficiently large $n$.

math.CO↗

From Link Diversity to Cross-Band Feedback Collaboration: A New Perspective on Hybrid Optical-RF Systems

We suggest a re-examination of the conventional view that hybrid optical-radio frequency (O-RF) systems are primarily diversity-driven networks that switch between RF and optical links for robustness. Instead, we uncover a new architectural opportunity: repurposing the optical downlink to enable real-time feedback channel coding over the RF uplink, where structured decoder feedback is delivered from the access point to guide the transmitter's coding strategy. This insight marks a conceptual paradigm shift from passive link diversity to active cross-band collaboration, where the wideband, interference-free optical wireless communication (OWC) is no longer merely a downlink backup but a functional enabler of uplink reliability. To realize this vision, we propose a novel architecture, O-RF with Cross-Band Feedback (O-RF-CBF), that exploits the optical downlink feedback to facilitate adaptive RF uplink coding. Numerical results reveal that O-RF-CBF achieves significant uplink throughput gains over traditional O-RF systems. Our findings highlight that inter-band synergy, not redundancy, is the key to unlocking the full potential of hybrid wireless networks.

cs.IT↗

Characterizing the Astrophysical Neutrino Flux Using Contained and Uncontained Cascade Events

Recently, the IceCube Neutrino Observatory has reported a deviation from the single power law in the extragalactic diffuse neutrino flux. A neural network-based event selection of contained and uncontained cascade events from IceCube, in which uncontained events have interaction vertices at the edge or outside of the detector instrumentation volume, has a factor ~3 gain in effective area over the cascade events used in the novel combined tracks and cascades selection which reported the deviation. Systematic improvements and rigorously updated modeling of the atmospheric neutrino background is incorporated into this high statistics contained and uncontained cascade event selection to clarify features of the astrophysical neutrino spectrum across energies from 1 TeV up to 100 PeV.

astro-ph.HE↗

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.

cs.CL↗

Enhancements to the IceCube Extremely High Energy Neutrino Selection using Graph & Transformer Based Neural Networks

KM3NeT has recently reported the detection of a very high-energy neutrino event, while IceCube has previously set upper limits on the differential neutrino flux above 100 PeV but has yet to observe a neutrino event with an energy comparable to that of the KM3NeT detection. To improve diffuse measurements above 10 PeV, we apply machine learning techniques to enhance atmospheric muon background rejection and directional reconstruction. We utilize a Graph Neural Network (GNN) to perform a classification task that distinguishes neutrinos from high-energy atmospheric muons. The method allows for the rejection of early hits from laterally spread, lower-energy muons in cosmic ray showers without relying on directional reconstruction as a prior. Additionally, a Transformer-based Neural Network is implemented for directional reconstruction. Unlike previous likelihood-based rapid reconstruction algorithms that assume a single muon track, this method makes no prior assumptions about event topology of the particle inside the detector. We demonstrate improved background rejection and reconstruction performance using machine learning techniques. Applications to the development of future Extremely High Energy (EHE) selections are also discussed.

astro-ph.HE↗

Spectral extremal problem for the odd prism

The spectral Turán number $\spex(n, F)$ denotes the maximum spectral radius $ρ(G)$ of an $F$-free graph $G$ of order $n$. This paper determines $\spex\left(n, C_{2k+1}^{\square}\right)$ for all sufficiently large $n$, establishing the unique extremal graph. Here, $C_{2k+1}^{\square}$ is the odd prism -- the Cartesian product $C_{2k+1} \square K_2$ -- where the Cartesian product $G \square F$ has vertex set $V(G) \times V(F)$, and edges between $(u_1,v_1)$ and $(u_2,v_2)$ if either $u_1 = u_2$ and $v_1v_2 \in E(F)$, or ($v_1 = v_2$ and $u_1u_2 \in E(G)$).

math.CO↗

Enhancing searches for astrophysical neutrino sources in IceCube with machine learning and improved spatial modeling

Searches for astrophysical neutrino sources in IceCube rely on an unbinned likelihood that consists of an energy and spatial component. Accurate modeling of the detector, ice, and spatial distributions leads to improved directional and energy reconstructions, resulting in increased sensitivity. In this work, we utilize our best knowledge of the detector ice properties and detector calibrations to reconstruct in-ice particle showers. The spatial component of the likelihood is parameterized either by a 2D Gaussian or a von Mises Fisher (vMF) distribution at small and large angular uncertainties, respectively. Here, we use a gradient-boosted decision tree with a vMF spatial likelihood loss function, reparameterized through two coordinate transformations, to predict per-event point spread functions (PSF). Additionally, we discuss the search for PeV cosmic ray sources using the IceCube Multi-Flavor Astrophysical Neutrino (ICEMAN) sample. Our search contains both an analysis of individual neutrino sources coincident with greater than 100 TeV gamma-ray sources and also a stacking analysis. We outline the prospects for extended neutrino emission originating from the Cygnus Cocoon region.

astro-ph.HE↗

Stochastic Operator Network: A Stochastic Maximum Principle Based Approach to Operator Learning

We develop a novel framework for uncertainty quantification in operator learning, the Stochastic Operator Network (SON). SON combines the stochastic optimal control concepts of the Stochastic Neural Network (SNN) with the DeepONet. By formulating the branch net as an SDE and backpropagating through the adjoint BSDE, we replace the gradient of the loss function with the gradient of the Hamiltonian from Stohastic Maximum Principle in the SGD update. This allows SON to learn the uncertainty present in operators through its diffusion parameters. We then demonstrate the effectiveness of SON when replicating several noisy operators in 2D and 3D.

cs.LG↗

Neural-operator element method: Efficient and scalable finite element method enabled by reusable neural operators

The finite element method (FEM) is a well-established numerical method for solving partial differential equations (PDEs). However, its mesh-based nature gives rise to substantial computational costs, especially for complex multiscale simulations. Emerging machine learning-based methods (e.g., neural operators) provide data-driven solutions to PDEs, yet they present challenges, including high training cost and low model reusability. Here, we propose the neural-operator element method (NOEM) by synergistically combining FEM with operator learning to address these challenges. NOEM leverages neural operators (NOs) to simulate subdomains where a large number of finite elements would be required if FEM was used. In each subdomain, an NO is used to build a single element, namely a neural-operator element (NOE). NOEs are then integrated with standard finite elements to represent the entire solution through the variational framework. Thereby, NOEM does not necessitate dense meshing and offers efficient simulations. We demonstrate the accuracy, efficiency, and scalability of NOEM by performing extensive and systematic numerical experiments, including nonlinear PDEs, multiscale problems, PDEs on complex geometries, and discontinuous coefficient fields.

cs.CE↗

QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset will be released at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.

eess.AS↗

Quantum DeepONet: Neural operators accelerated by quantum computing

In the realm of computational science and engineering, constructing models that reflect real-world phenomena requires solving partial differential equations (PDEs) with different conditions. Recent advancements in neural operators, such as deep operator network (DeepONet), which learn mappings between infinite-dimensional function spaces, promise efficient computation of PDE solutions for a new condition in a single forward pass. However, classical DeepONet entails quadratic complexity concerning input dimensions during evaluation. Given the progress in quantum algorithms and hardware, here we propose to utilize quantum computing to accelerate DeepONet evaluations, yielding complexity that is linear in input dimensions. Our proposed quantum DeepONet integrates unary encoding and orthogonal quantum layers. We benchmark our quantum DeepONet using a variety of PDEs, including the antiderivative operator, advection equation, and Burgers' equation. We demonstrate the method's efficacy in both ideal and noisy conditions. Furthermore, we show that our quantum DeepONet can also be informed by physics, minimizing its reliance on extensive data collection. Quantum DeepONet will be particularly advantageous in applications in outer loop problems which require exploring parameter space and solving the corresponding PDEs, such as uncertainty quantification and optimal experimental design.

quant-ph↗

Capacity-Optimized Pre-Equalizer Design for Visible Light Communication Systems

Since commercial LEDs are primarily designed for illumination rather than data transmission, their modulation bandwidth is inherently limited to a few MHz. This becomes a major bottleneck in the implementation of visible light communication (VLC) systems necessiating the design of pre-equalizers. While state-of-the-art equalizer designs primarily focus on the data rate increasing through bandwidth expansion, they often overlook the accompanying degradation in signal-to-noise ratio (SNR). Achieving effective bandwidth extension without introducing excessive SNR penalties remains a significant challenge, since the channel capacity is a non-linear function of both parameters. In this paper, we present a fundamental analysis of how the parameters of the LED and pre-equalization circuits influence the channel capacity in intensity modulation and direct detection (IMDD)-based VLC systems. We derive a closed-form expression for channel capacity model that is an explicitly function of analog pre-equalizer circuit parameters. Building upon the derived capacity expression, we propose a systematic design methodology for analog pre-equalizers that effectively balances bandwidth and SNR, thereby maximizing the overall channel capacity across a wide range of channel attenuations. We present extensive numerical results to validate the effectiveness of the proposed design and demonstrate the improvements over conventional bandwidth-optimized pre-equalizer designs.

cs.IT↗

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

In order to enable fluid and natural human-machine speech interaction, existing full-duplex conversational systems often adopt modular architectures with auxiliary components such as voice activity detectors, interrupters, conversation state predictors, or multiple LLMs. These systems, however, suffer from error accumulation across modules and struggle with key challenges such as context-dependent barge-in and echo cancellation. Recent approaches, most notably Moshi, simplify the pipeline by injecting audio codecs into the token space of a single LLM. However, such methods still incur significant performance degradation when operating on the speech rather than text modality. In this paper, we introduce SALMONN-omni, the first single, standalone full-duplex speech LLM that operates without audio codecs in its token space. It features a novel dynamic thinking mechanism within the LLM backbone, enabling the model to learn when to transition between speaking and listening states. Experiments on widely used benchmarks for spoken question answering and open-domain dialogue show that SALMONN-omni achieves at least 30\% relative performance improvement over existing open-source full-duplex models and performs highly competitively to half-duplex and turn-based systems, despite using substantially less training data. Moreover, SALMONN-omni demonstrates strong performance in complex conversational scenarios, including turn-taking, backchanneling, echo cancellation and context-dependent barge-in, with further improvements achieved through reinforcement learning. Some demo conversations between user and SALMONN-omni are provided in the following repository https://github.com/bytedance/SALMONN.

cs.CL↗

Genetic Algorithm-Accelerated Computational Discovery of Liquid Crystal Polymers with Enhanced Optical Properties

Liquid crystal polymers with exceptional optical properties are highly promising for next-generation virtual, augmented, and mixed reality (VR/AR/MR) technologies, serving as high-performance, compact, lightweight, and cost-effective optical components. However, the growing demands for optical transparency and high refractive index in advanced optical devices present a challenge for material discovery. In this study, we develop a novel approach that integrates first-principles calculations with genetic algorithms to accelerate the discovery of liquid crystal polymers with low visible absorption and high refractive index. By iterating within a predefined space of molecular building blocks, our approach rapidly identifies reactive mesogens that meet target specifications. Additionally, it provides valuable insights into the relationships between molecular structure and properties. This strategy not only accelerates material screening but also uncovers key molecular design principles, offering a systematic and scalable alternative to traditional trial-and-error methods.

cond-mat.soft↗

CreoPep: A Universal Deep Learning Framework for Target-Specific Peptide Design and Optimization

Target-specific peptides, such as conotoxins, exhibit exceptional binding affinity and selectivity toward ion channels and receptors. However, their therapeutic potential remains underutilized due to the limited diversity of natural variants and the labor-intensive nature of traditional optimization strategies. Here, we present CreoPep, a deep learning-based conditional generative framework that integrates masked language modeling with a progressive masking scheme to design high-affinity peptide mutants while uncovering novel structural motifs. CreoPep employs an integrative augmentation pipeline, combining FoldX-based energy screening with temperature-controlled multinomial sampling, to generate structurally and functionally diverse peptides that retain key pharmacological properties. We validate this approach by designing conotoxin inhibitors targeting the $α$7 nicotinic acetylcholine receptor, achieving submicromolar potency in electrophysiological assays. Structural analysis reveals that CreoPep-generated variants engage in both conserved and novel binding modes, including disulfide-deficient forms, thus expanding beyond conventional design paradigms. Overall, CreoPep offers a robust and generalizable platform that bridges computational peptide design with experimental validation, accelerating the discovery of next-generation peptide therapeutics.

q-bio.BM↗

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) \etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A/B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A/B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints can be found at https://github.com/bytedance/SALMONN.

eess.AS↗

Buckling and post-buckling of cylindrical shells under combined torsional and axial loads

The buckling behavior of cylindrical shells has gained significant interest over the past century due to its rich nonlinear behavior and broad engineering applications. While the buckling of cylindrical shells under a single load (e.g., compression or torsion) has been extensively studied, the buckling behavior under combined torsional and axial loads remains largely unexplored. In this paper, based on a combination of experiments, theoretical modeling, and finite element simulations, we systematically investigate the buckling and post-buckling behavior of cylindrical shells under combined torsional and axial loads. Three different types of combined loads are considered: compression with pre-torsion, torsion with pre-tension, and torsion with pre-compression. The theoretical model is established within the framework of the Donnell shell theory and solved using the Galerkin method, through which the critical buckling load, critical circumferential wavenumber, buckling pattern, and post-buckling equilibrium path of clamped-clamped thin cylindrical shells under various types of loads can be determined. The theoretical predictions agree well with finite element simulations and qualitatively capture the various buckling phenomena observed in the experiments. It is found that cylindrical shells exhibit quite different post-buckling behavior under combined loads compared to under a single compressive or torsional load. For instance, when a clamped-clamped thin cylindrical shell is subjected to pure torsion or torsion with a relatively small pre-compression, it consistently shows a diagonal-shaped pattern during deformation. However, with a relatively large pre-compression, the shell transitions from a diagonal-shaped pattern to a twisted diamond-shaped pattern.

nlin.PS↗

Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context

Large Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio. However, most existing models primarily focus on analyzing input signals using text instructions, overlooking scenarios in which speech instructions and audio are mixed and serve as inputs to the model. To address these challenges, we introduce Solla, a novel framework designed to understand speech-based questions and hear the acoustic context concurrently. Solla incorporates an audio tagging module to effectively identify and represent audio events, as well as an ASR-assisted prediction method to improve comprehension of spoken content. To rigorously evaluate Solla and other publicly available models, we propose a new benchmark dataset called SA-Eval, which includes three tasks: audio event classification, audio captioning, and audio question answering. SA-Eval has diverse speech instruction with various speaking styles, encompassing two difficulty levels, easy and hard, to capture the range of real-world acoustic conditions. Experimental results show that Solla performs on par with or outperforms baseline models on both the easy and hard test sets, underscoring its effectiveness in jointly understanding speech and audio.

eess.AS↗