SearcharxivSearch

arXiv subjects

Fan Bu

Publications and source records attributed to Fan Bu.

At least 19 recordsLinked to original sources

Real-Variable Characterizations and Their Applications of Anisotropic Besov Spaces with Matrix $\mathcal A_\infty$ Weights

Let $\alpha\in\mathbb{R}$, $p\in(0,\infty)$, and $q\in(0,\infty]$. In this article, we develop a theory of matrix-weighted anisotropic Besov spaces associated with an expansive matrix $A$ and an $\mathcal A_{p,\infty}$-matrix weight $W$. We first introduce the homogeneous spaces $\dot B_{p,q}^{\alpha}(A,W)$ and establish their $\varphi$-transform characterization. Then we construct counterexamples to show that the assumption $W\in\mathcal A_{p,\infty}$ in this characterization cannot be relaxed to $W\in\bigcup_{r\in(0,\infty)}\mathcal A_r$. The same counterexamples also show that this weaker condition $W\in\bigcup_{r\in(0,\infty)}\mathcal A_r$ is insufficient to ensure the well-definedness of $\dot B_{p,q}^{\alpha}(A,W)$. Next we characterize $\mathcal A_{p,\infty}$-matrix weights via the rescaled maximal operator, which leads naturally to a new concept of the critical rescaling index that quantitatively captures the self-improving behavior of matrix weights. In terms of this index, we obtain optimal boundedness for almost diagonal operators on the associated sequence spaces $\dot b_{p,q}^{\alpha}(A,W)$. Based on these, we further establish the molecular characterization of $\dot B_{p,q}^{\alpha}(A,W)$ and some sharp boundedness results for pseudo-differential operators on these spaces.

math.CA

Multi-Signal Safety Surveillance with Bayesian Latent Factor Modeling and Bias Correction

Safety surveillance increasingly involves repeated monitoring of many exposure-outcome signals in observational healthcare data, where sparse information, dependence across related signals, and systematic error can complicate inference. Existing frameworks typically focus on either correcting residual bias using negative controls or borrowing information across exposure-outcome pairs, but not both. We propose a multi-signal Bayesian sequential surveillance framework that integrates empirical bias correction with low-rank latent factor modeling. At each analysis time, a hierarchical Bayesian model learns exposure-specific bias distributions from negative control outcomes assumed to have null latent effects. Conditional on these distributions, low-rank latent factors are estimated across exposures and outcomes of interest to share information across correlated signals. As new data accrue, posterior inference is updated sequentially, yielding bias-corrected posterior summaries of effect sizes across multiple monitored signals. We illustrate the method in a postmarket vaccine safety surveillance study using a large US insurance claims database.

stat.ME

Bayesian Joint Modeling of Longitudinal Symptomatology Scale Responses and Fall Outcomes via Heterogeneous Latent Transition Analysis

The Study of Women's Health Across the Nation (SWAN) has followed women for over 30 years, from midlife premenopause until later life. The study has 16 surveys at approximately 2 years intervals that cover a wide range of physical and psychological symptoms. These multivariate categorical survey responses potentially contain rich health-related information. Temporal trajectories of the survey responses can be characterized by both the responses profiles and the evolving dynamics of the responses over time. To capture those two features and investigate how they inform subsequent health outcomes, we propose a joint multi-layer latent transition model. We combine a latent transition model that classifies individuals based on their response profiles over time with an additional layer of clustering of these latent class transition sequences, with the goal of connecting these cluster profiles with health outcomes: in this application, self-reported falls. In addition, we evaluate the operating characteristics of the method through simulation studies.

stat.AP

GeoID-PINN: Identifiability-Aware Regional Epidemic Inference with Geographic Coupling

Regional surveillance data reflect local transmission, reporting, seeding, and external infection pressure, which are difficult to identify separately. We introduce GeoID-PINN, a physics-informed neural network (PINN) for susceptible-infectious-recovered-deceased (SIRD) dynamics. The model represents spatial dependence with a row-stochastic source-composition matrix whose rows assign nonnegative source weights that sum to one. We regularize this matrix toward a spatial prior constructed from distance, adjacency, commuting, or lead-lag information. In a four-region simulation with known truth, a compatible distance prior gives source-composition error 0.099. The error rises to 0.159 without regularization and 0.577 under a strongly misspecified prior, while trajectory fit and transmission-scale estimates remain similar. Accurate trajectories therefore do not guarantee recovery of the regional dependence structure. We also evaluate GeoID-PINN retrospectively using COVID-19 data from 64 Louisiana counties. Relative to an autoregressive negative-binomial baseline, Forecast-Trained Geo-PINN reduces mean squared error (MSE) from 32,957 to 11,468 and mean absolute error (MAE) from 70.60 to 57.73. The baseline has lower negative log likelihood (NLL), 5.158 versus 5.346, indicating better distributional fit but worse point accuracy. In a controlled 15-county comparison, county adjacency reduces MSE by 6.85 percent and MAE by 3.1 percent. Similar performance across plausible priors supports structured regularization but not unique edge recovery. These results require prior-sensitivity and observation-model checks before interpretation.

cs.LG

MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck

Melody skeleton extraction aims to derive a shorter melody that preserves structural notes while removing ornaments. Prior methods rely on hand-crafted reduction rules or note-wise salience classifiers trained with heuristically or procedurally generated pseudo-labels. Such supervision can inherit generator bias and does not explicitly optimize a coherent reduced melody. We introduce MeloBottleneck, a self-supervised framework that represents a skeleton as a length-controlled, order-preserving latent subsequence. A hard-bottleneck extractor selects note events, a rhythmic-closure operator produces a self-consistent skeleton, and a re-ornamentation decoder reconstructs the input melody. Training combines reconstruction, a frozen autoregressive melody prior, ornament-invariant consistency across procedurally ornamented views, and ornament exclusion. We evaluate three regimes: synthetic out-of-distribution ornament-to-skeleton, TAVERN variation-to-theme, and Jiugong ornamented-to-gongche. A matched pseudo-label classifier excels on the synthetic benchmark, while MeloBottleneck transfers better, achieving competitive selection quality on TAVERN and Jiugong. Skeletonized melodies also improve BM25-based fragment retrieval, boosting Recall@K and MRR while reducing query time. Overall, the results suggest that learning skeletons as latent subsequences yields more robust transfer than pseudo-label imitation.

cs.SD

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.

cs.CL

HiMed: Incentivizing Hindi Reasoning in Medical LLMs

Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, their performance degrades sharply in Hindi, particularly on Indian systems of medicine. We argue that robust cross-lingual medical transfer requires Hindi reasoning. To this end, we introduce HiMed, a Hindi reasoning medical corpus and benchmark suite covering both Western and Indian medicine. We further propose HiMed-8B, a Hindi-form medical reasoning LLM, through the design of decaying scaffolding reward. Extensive experiments demonstrate improvement in Hindi medical reasoning performance and reduction in the English--Hindi accuracy gap. Ablation studies validate the contribution of each training stage and reward component. All data and code are available on GitHub: https://github.com/FreedomIntelligence/HiMed.

cs.CL

HYVE: Hybrid Views for LLM Context Engineering over Machine Data

Machine data is central to observability and diagnosis in modern computing systems, appearing in logs, metrics, telemetry traces, and configuration snapshots. When provided to large language models (LLMs), this data typically arrives as a mixture of natural language and structured payloads such as JSON or Python/AST literals. Yet LLMs remain brittle on such inputs, particularly when they are long, deeply nested, and dominated by repetitive structure. We present HYVE (HYbrid ViEw), a framework for LLM context engineering for inputs containing large machine-data payloads, inspired by database management principles. HYVE surrounds model invocation with coordinated preprocessing and postprocessing, centered on a request-scoped datastore augmented with schema information. During preprocessing, HYVE detects repetitive structure in raw inputs, materializes it in the datastore, transforms it into hybrid columnar and row-oriented views, and selectively exposes only the most relevant representation to the LLM. During postprocessing, HYVE either returns the model output directly, queries the datastore to recover omitted information, or performs a bounded additional LLM call for SQL-augmented semantic synthesis. We evaluate HYVE on diverse real-world workloads spanning knowledge QA, chart generation, anomaly detection, and multi-step network troubleshooting. Across these benchmarks, HYVE reduces token usage by 50-90% while maintaining or improving output quality. On structured generation tasks, it improves chart-generation accuracy by up to 132% and reduces latency by up to 83%. Overall, HYVE offers a practical approximation to an effectively unbounded context window for prompts dominated by large machine-data payloads.

cs.AI

Real-variable theory of matrix-weighted multi-parameter Besov--Triebel--Lizorkin-type spaces

We develop a comprehensive theory for a general class of multi-parameter function spaces of Besov-Triebel-Lizorkin type, with a matrix weight. We prove the equivalence of different quasi-norms, the identification of function and sequence spaces via the $\varphi$-transform, the boundedness of almost diagonal operators and multi-parameter singular integrals under minimal assumptions, molecular and wavelet characterisations, and Sobolev-type embedding theorems. We identify matrix-weighted $L^p$ spaces, Sobolev spaces, and multi-parameter BMO spaces as examples of our general scale of spaces. Thus, our result on the boundedness of multi-parameter singular integrals on these spaces is seen as an extension, with a different method, of a recent theorem of Domelevo et al. [J. Math. Anal. Appl. 2024] on matrix-weighted $L^p$ spaces. For this theory, we develop several tools of independent interest. Many previous results were restricted to integrability exponents $p\in(1,\infty)$, while Besov-Triebel-Lizorkin spaces naturally involve the full range $p\in(0,\infty)$. We extend the definition of multi-parameter $A_p$ matrix weights to $p\in(0,1]$ and establish their basic properties, culminating in the $L^p$-boundedness of a matrix-weighted strong maximal operator (suitably rescaled when $p\in(0,1]$) for all $p\in(0,\infty)$. For $p\in(1,\infty)$, this is due to Vuorinen [Adv. Math. 2024] by convex-set-valued techniques of Bownik and Cruz-Uribe [arXiv 2022; Math. Ann. (to appear)]; the lack of convexity requires us to develop a new approach that works for all $p\in(0,\infty)$. We also need and prove a multi-parameter extension of Carleson-type embeddings from Frazier and Roudenko [Math. Ann. 2021] but attributed by them to F. Nazarov. We prove the necessity of the conditions of the new embedding using a nontrivial elaboration of Carleson's classical counterexample [Mittag-Leffler Rep. 1974].

math.FA

The Renaissance of Expert Systems: Optical Recognition of Printed Chinese Jianpu Musical Scores with Lyrics

Large-scale optical music recognition (OMR) research has focused mainly on Western staff notation, leaving Chinese Jianpu (numbered notation) and its rich lyric resources underexplored. We present a modular expert-system pipeline that converts printed Jianpu scores with lyrics into machine-readable MusicXML and MIDI, without requiring massive annotated training data. Our approach adopts a top-down expert-system design, leveraging traditional computer-vision techniques (e.g., phrase correlation, skeleton analysis) to capitalize on prior knowledge, while integrating unsupervised deep-learning modules for image feature embeddings. This hybrid strategy strikes a balance between interpretability and accuracy. Evaluated on The Anthology of Chinese Folk Songs, our system massively digitizes (i) a melody-only collection of more than 5,000 songs (> 300,000 notes) and (ii) a curated subset with lyrics comprising over 1,400 songs (> 100,000 notes). The system achieves high-precision recognition on both melody (note-wise F1 = 0.951) and aligned lyrics (character-wise F1 = 0.931).

cs.CV

Matrix-Weighted Besov-Triebel-Lizorkin Spaces of Optimal Scale: Real-Variable Characterizations, Invariance on Integrable Index, and Sobolev-Type Embedding

In this article, using growth functions we introduce generalized matrix-weighted Besov-Triebel-Lizorkin-type spaces with matrix $\mathcal{A}_{\infty}$ weights. We first characterize these spaces, respectively, in terms of the $\varphi$-transform, the Peetre-type maximal function, and the Littlewood-Paley functions. Furthermore, after establishing the boundedness of almost diagonal operators on the corresponding sequence spaces, we obtain the molecular and the wavelet characterizations of these spaces. As applications, we find the sufficient and necessary conditions for the invariance of those Triebel-Lizorkin-type spaces on the integrable index and also for the Sobolev-type embedding of all these spaces. The main novelty exists in that these results are of wide generality, the growth condition of growth functions is not only sufficient but also necessary for the boundedness of almost diagonal operators and hence this new framework of Besov-Triebel-Lizorkin-type is optimal, some results either are new or improve the known ones even for known matrix-weighted Besov-Triebel-Lizorkin spaces, and, furthermore, even in the scalar-valued setting, all the results are also new.

math.FA

S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

Recent advances in large language models (LLMs) have fundamentally reshaped speech-to-speech (S2S) systems, enabling increasingly natural spoken interaction. However, existing benchmarks still rely heavily on text-based evaluation and largely ignore paralinguistic cues such as prosody, emotion, and speaker traits, which are central to expressive and human-like communication. We introduce S2S-Arena, a speech-native benchmark for evaluating instruction-following S2S models with explicit assessment of both semantic understanding and paralinguistic expression. S2S-Arena features a four-level interaction protocol that systematically probes models under increasing paralinguistic complexity, a two-stage data construction pipeline that produces 1,243 speech samples spanning 100+ real-world tasks, and an arena-style evaluation framework that enables reference-free, pairwise comparison directly in the speech modality. Benchmarking 10 state-of-the-art S2S systems over 1,000+ comparisons reveals substantial performance gaps (especially under complex paralinguistic demands) between current academic and industrial systems. Our analysis further identifies key design factors governing expressive instruction following, providing actionable insights for building more natural, robust, and human-aligned speech agents.

cs.CL

Soundwave: Less is More for Speech-Text Alignment in LLMs

Existing end-to-end speech large language models (LLMs) usually rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth. We focus on two fundamental problems between speech and text: the representation space gap and sequence length inconsistency. We propose Soundwave, which utilizes an efficient training strategy and a novel architecture to address these issues. Results show that Soundwave outperforms the advanced Qwen2-Audio in speech translation and AIR-Bench speech tasks, using only one-fiftieth of the training data. Further analysis shows that Soundwave still retains its intelligence during conversation. The project is available at https://github.com/FreedomIntelligence/Soundwave.

cs.CL

Maximal Function and Atomic Characterizations of Matrix-Weighted Hardy Spaces with Their Applications to Boundedness of Calder\'on--Zygmund Operators

Let $p\in(0,1]$ and $W$ be an $A_p$-matrix weight, which in scalar case is exactly a Muckenhoupt $A_1$ weight. In this article, we introduce matrix-weighted Hardy spaces $H^p_W$ via the matrix-weighted grand non-tangential maximal function and characterize them, respectively, in terms of various other maximal functions and atoms, both of which are closely related to matrix weights under consideration and their corresponding reducing operators. As applications, we first establish the finite atomic characterization of $H^p_W$, then using it we give a criterion on the boundedness of sublinear operators from $H^p_W$ to any $\gamma$-quasi-Banach space, and finally applying this criterion we further obtain the boundedness of Calder\'on--Zygmund operators on $H^p_W$. The main novelty of these results lies in that the aforementioned maximal functions related to reducing operators are new even in the scalar weight case and we characterize these matrix-weighted Hardy spaces by a fresh and natural variant of classical weighted atoms via first establishing a Calder\'on--Zygmund decomposition which is also new even in the scalar weight case.

math.FA

Besov--Triebel--Lizorkin-Type Spaces with Matrix $A_\infty$ Weights

Introduced by A. Volberg, matrix $A_{p,\infty}$ weights provide a suitable generalization of Muckenhoupt $A_\infty$ weights from the classical theory. In our previous work, we established new characterizations of these weights. Here, we use these results to study inhomogeneous Besov-type and Triebel--Lizorkin-type spaces with such weights. In particular, we characterize these spaces, in terms of the $\varphi$-transform, molecules, and wavelets, and obtain the boundedness of almost diagonal operators, pseudo-differential operators, trace operators, pointwise multipliers, and Calder\'on--Zygmund operators on these spaces. This is the first systematic study of inhomogeneous Besov--Triebel--Lizorkin-type spaces with $A_{p,\infty}$-matrix weights, but some of the results are new even when specialized to the scalar unweighted case.

math.FA

An Investigation into Value Misalignment in LLM-Generated Texts for Cultural Heritage

As Large Language Models (LLMs) become increasingly prevalent in tasks related to cultural heritage, such as generating descriptions of historical monuments, translating ancient texts, preserving oral traditions, and creating educational content, their ability to produce accurate and culturally aligned texts is being increasingly relied upon by users and researchers. However, cultural value misalignments may exist in generated texts, such as the misrepresentation of historical facts, the erosion of cultural identity, and the oversimplification of complex cultural narratives, which may lead to severe consequences. Therefore, investigating value misalignment in the context of LLM for cultural heritage is crucial for mitigating these risks, yet there has been a significant lack of systematic and comprehensive study and investigation in this area. To fill this gap, we systematically assess the reliability of LLMs in generating culturally aligned texts for cultural heritage-related tasks. We conduct a comprehensive evaluation by compiling an extensive set of 1066 query tasks covering 5 widely recognized categories with 17 aspects within the knowledge framework of cultural heritage across 5 open-source LLMs, and examine both the type and rate of cultural value misalignments in the generated texts. Using both automated and manual approaches, we effectively detect and analyze the cultural value misalignments in LLM-generated texts. Our findings are concerning: over 65% of the generated texts exhibit notable cultural misalignments, with certain tasks demonstrating almost complete misalignment with key cultural values. Beyond these findings, this paper introduces a benchmark dataset and a comprehensive evaluation workflow that can serve as a valuable resource for future research aimed at enhancing the cultural sensitivity and reliability of LLMs.

cs.CL

Neural Posterior Estimation for Stochastic Epidemic Modeling

Stochastic infectious disease models capture uncertainty in public health outcomes and have become increasingly popular in epidemiological practice. However, calibrating these models to observed data is challenging with existing methods for parameter estimation. Stochastic epidemic models are nonlinear dynamical systems with potentially large latent state spaces, resulting in computationally intractable likelihood densities. We develop an approach to calibrating complex epidemiological models to high-dimensional data using Neural Posterior Estimation, a novel technique for simulation-based inference. In NPE, a neural conditional density estimator trained on simulated data learns to "invert" a stochastic simulator, returning a parametric approximation to the posterior distribution. We introduce a stochastic, discrete-time Susceptible Infected (SI) model with heterogeneous transmission for healthcare-associated infections (HAIs). HAIs are a major burden on healthcare systems. They exhibit high rates of asymptotic carriage, making it difficult to estimate infection rates. Through extensive simulation experiments, we show that NPE produces accurate posterior estimates of infection rates with greater sample efficiency compared to Approximate Bayesian Computation (ABC). We then use NPE to fit our SI model to an outbreak of carbapenem-resistant Klebsiella pneumoniae in a long-term acute care facility, finding evidence of location-based heterogeneity in patient-to-patient transmission risk. We argue that our methodology can be fruitfully applied to a wide range of mechanistic transmission models and problems in the epidemiology of infectious disease.

stat.ME

Roadmap towards Superhuman Speech Understanding using Large Language Models

The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o, highlight the potential for end-to-end speech LLMs, which preserves non-semantic information and world knowledge for deeper speech understanding. To guide the development of speech LLMs, we propose a five-level roadmap, ranging from basic automatic speech recognition (ASR) to advanced superhuman models capable of integrating non-semantic information with abstract acoustic knowledge for complex tasks. Moreover, we design a benchmark, SAGI Bechmark, that standardizes critical aspects across various tasks in these five levels, uncovering challenges in using abstract acoustic knowledge and completeness of capability. Our findings reveal gaps in handling paralinguistic cues and abstract acoustic knowledge, and we offer future directions. This paper outlines a roadmap for advancing speech LLMs, introduces a benchmark for evaluation, and provides key insights into their current limitations and potential.

cs.CL