SearcharxivSearch

arXiv subjects

Yifei Wu

Publications and source records attributed to Yifei Wu.

At least 19 recordsLinked to original sources

Temporal Error Growth of Strang Splitting Method for the Periodic Cubic NLS

We study the temporal error growth of the Strang splitting method for the periodic cubic nonlinear Schrödinger (NLS) equation. The leading error is governed by a forced linearized equation, whose growth depends sharply on dimension and the sign of the nonlinearity. In the 1D defocusing case, we prove a uniform quadratic upper bound $C(1+T^2)τ^2$ using the global Birkhoff transformation and the resulting degenerate structure of the linearized flow. In the higher-dimensional defocusing case, the exponential growth is constructed using arbitrarily small unstable standing waves. Moreover, to prove the higher-dimensional defocusing upper bound, we use the periodic Strichartz estimates of Killip and Vişan to show that $\int_0^T\|u(t)\|_{L^\infty}^2\,dt\le C(u_0)(1+T)$. This yields an exponential rate independent of the final time, despite possible growth of higher Sobolev norms. The exponential lower bound in the focusing case is obtained from the unstable linearized dynamics around a plane wave, with the Akhmediev breather providing the underlying mechanism. In addition, for 1D focusing case with small initial data, the appliction of the local Birkhoff transformation ensures that the error grows at most quadratically in time. Various numerical experiments confirm the sharp linear, quadratic, and exponential error growth rates.

math.NA

Modified wave operators for nonlinear Schrödinger equations in the full subcritical long range regime

We construct modified wave operators for the nonlinear Schrödinger equation $i\partial_tu+\frac12Δu=|u|^pu$ in the full subcritical long-range case $0<p<2/d$, with small, nonvanishing, analytic final data $W(x)$ with bounded logarithmic gradients. Previous results established large-time asymptotics for selected classes of Cauchy data. Moreover, the exact asymptotic expansion for $p<1/d$ remained unknown. When $1/d<p<2/d$, our result gives the approximation $$\frac{1}{(it)^{\frac{d}{2}}} e^{\frac{i|x|^2}{2t}} W\left(\frac{x}{t}\right) \exp\left[ -i\frac{t^{1-\frac{dp}{2}}-1}{1-\frac{dp}{2}} \left|W\left(\frac{x}{t}\right)\right|^p \right]$$ The wave operator is constructed by an iteration in the analytic spaces with decreasing radius. When $p\le 1/d$, we construct the profile from a finite truncation of a Fuchsian equation coupled with a transport equation. This profile still leaves a long-range triangular coupling whose terminal integral does not preserve the required fast decay class. The construction yields quantitative $L^q$ asymptotics for $2\le q\le\infty$ and uniqueness in the prescribed analytic asymptotic classes. The central new ingredients are a nonlinear final-state normal form that removes this long-range coupling and a mixed iteration in particular analytic spaces.

math.AP

Praxist: From Experimental Artifacts to Solution Lineages

Autonomous R\&D agents now write, run, and improve executable artifacts under automated evaluation---but largely as laboratory instruments: shown on curated benchmarks, with gains that are hard to trace to a cause and costs well above what sustained engineering practice absorbs. The limitation is structural. Most systems treat each attempt as nearly self-contained, so logs, memories, and search trees record what happened without establishing which design element produced an improvement, whether its evidence survived validation, or how it recombines with others. Long campaigns therefore keep re-learning the same lessons. We introduce Praxist, a lineage-centered generational system that converts reproducible artifacts and evaluator outcomes into a typed evidence graph of findings, lane-structured frontiers, and agendas. Separating local artifact construction from cohort-level evidence synthesis lets later attempts inherit validated mechanisms, unresolved claims, and useful constraints, and leaves results attached to an inspectable lineage. On the standardized 75-task MLE-bench suite, the finalized official-grader results give Praxist 60 medals (80.0\%), 49 of them gold, against 55 medals (73.3\%) and 34 gold for a Claude Code baseline on Claude Opus 4.8---at a recorded model spend of US\$3,054 versus US\$38,370, roughly a twelfth of the cost. Four case studies---quantitative trading, LiDAR-inertial-visual SLAM, tokamak magnetic control, and rocket landing---carry the same process into open-ended engineering problems, improving on each task-native baseline in headline accuracy, survival, or resource cost, with the discovery path on record. Stronger artifacts at an order of magnitude less spend, each backed by an auditable lineage, are, to our knowledge, first brought together here: the operating profile production research requires, not the one a benchmark demonstration establishes.

cs.MA

Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their corresponding intensity distributions highly entangled and ambiguous. We observe that once the spatial organization of anatomical regions is established, estimating their CT intensities becomes substantially more tractable. Motivated by this, we propose LiftXR, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction. Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays, providing spatial guidance for an intensity renderer to reconstruct a CT volume. An anatomical parser then performs volumetric perception on the reconstruction, exploiting its spatially resolved boundary and intensity cues to recover a refined anatomical layout. This transition from projection-conditioned layout generation to reconstruction-conditioned anatomical perception allows the parsed layout to provide feedback for region-specific intensity calibration. Extensive experiments on two public datasets demonstrate that LiftXR consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art. Moreover, the reconstructed CT achieves superior performance in external downstream segmentation, indicating improved anatomical fidelity. Code will be released.

cs.CV

The subsonic limit of the 3D Zakharov system

We obtain the optimal convergence rates in the subsonic limit of the three-dimensional Zakharov system for initial data belonging to the low-regularity Sobolev space $\HH^s=H^s\times H^{s-1}\times H^{s-1}$. For the Schrödinger component, we prove first-order convergence in $L^2$ for initial data in $\HH^3$, and second-order convergence under the compatibility condition for data in $\HH^4$. For the wave component, we obtain first-order convergence in $L^2$ for data in $\HH^3$ and second-order convergence for data in $\HH^4$. The obtained rates are optimal and coincide with those predicted by the formal asymptotic expansion. No localization assumptions, smallness or high-order regularity hypotheses are required. This improves all previous results on the subsonic limit of the Zakharov system and resolves the optimality issue at the Sobolev regularity level. The proof relies on a uniform local well-posedness theory that remains valid in the subsonic limit. A key ingredient is a refined normal form analysis combined with bilinear Strichartz estimates in atomic function spaces, which allows us to fully exploit the dispersive structure of the Zakharov system at low regularity and to overcome the derivative losses arising from the singular coupling.

math.AP

Ultra-Low-Cost Hybrid Beamforming: A New Static-Connection Architecture with Sparse Phase-Shifter Sharing

Hybrid beamforming is a promising solution for high-frequency multi-antenna wireless systems, but its implementation is constrained by the cost and complexity of analog phase-shifter (PS) networks. Although sub-connected architectures simplify the analog network, their conventional realization still requires a dedicated PS for each antenna, causing considerable layout area, wiring, calibration, and control overheads. To address this issue, this paper proposes a novel static-connection architecture with sparse PSs for ultra-low-cost sub-connected hybrid beamforming, where antennas within each sub-array share a PS through an optimized fixed PS-to-antenna connection matrix. The proposed architecture preserves static connections while enabling dynamic beam control via adaptive PS phase-shift adjustments and digital precoding. For the single-radio-frequency (RF)-chain scenario, the sparse-PS connection design is transformed into an antenna-grouping problem, with analytically characterized structural properties and an efficient algorithm. For the multi-RF-chain scenario, we develop a quality-of-service (QoS)-majorization-minimization (MM) algorithm to handle the mixed discrete-continuous optimization problem. Numerical results demonstrate that the proposed architecture reduces the PS count while preserving most beamforming capability of the traditional full-PS sub-connected architecture. In particular, the proposed design achieves PS-count reductions of 37.5% and 62.5% in single-RF-chain and multi-RF-chain systems, respectively, while avoiding deep-null and grating-lobe degradations associated with deterministic connection schemes. These results provide engineering insights into static sparse-PS sharing: the key to hardware-efficient hybrid beamforming is not merely reducing the PS count, but also preserving essential analog-domain degrees of freedom through optimized PS connection topologies.

cs.IT

World Models in Pieces: Structural Certification for General Agents

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrelevant failures. We first formalize this limitation by proving that general agents are not universal, rendering standard worst-case analysis uninformative. To overcome this, we introduce structural certification, a transition-local framework that maps bounded goal-conditioned performance to entry-wise guarantees on the agent's internal world model. Our main contribution is constructive. We provide algorithms that filter specific transitions using deep compositional goals and prove that a general agent on these goals has a structural world model with a $\mathcal{O}(1/n) + \mathcal{O}(δ)$ error bound. Conversely, this bound is tight in the small-$δ$ regime, whose existence is explicitly guaranteed by our certification. These results enable the certifiable deployment of general agents by localizing the specific transitions where long-horizon planning is reliable.

cs.AI

Movable-Antenna-Enhanced ISAC: Optimal Antenna Trajectory and Beamforming Design

Integrated sensing and communication (ISAC) is a key enabling technology for next-generation wireless networks. However, most existing ISAC systems rely on fixed-position antennas, which restrict performance when balancing sensing and communication objectives. Movable antenna (MA) technology introduces additional spatial degrees of freedom through antenna mobility, yet existing studies on MA-enabled ISAC schemes mainly consider static antenna repositioning and fail to fully exploit this capability. By leveraging spatio-temporal sampling enabled by antenna motion, optimized MA trajectories can synthesize large virtual aperture arrays, thereby improving angular resolution and reducing sensing ambiguity. To this end, this paper investigates a dynamic MA-enabled ISAC system and studies the joint design of MA trajectories and transmit beamforming. We formulate a joint trajectory and beamforming optimization problem to minimize sensing beampattern mismatch under communication quality-of-service constraints. A branch-and-bound-based algorithm is developed to obtain the globally optimal solution. Numerical results show that the proposed framework significantly outperforms baseline schemes with only one or two antenna repositioning steps, demonstrating its practical feasibility.

eess.SP

HRM-Text: Efficient Pretraining Beyond Scaling

The current pretraining paradigm for large language models relies on massive compute and internet-scale raw text, creating a significant barrier to foundational research. In contrast, biological systems demonstrate highly sample-efficient learning through multi-timescale processing, such as the functional organization of the frontoparietal loop. Taking this as inspiration, we introduce HRM-Text, which replaces standard Transformers with a Hierarchical Recurrent Model (HRM) that decouples computation into slow-evolving strategic and fast-evolving execution layers. To stabilize this deep recurrence for language modeling, we introduce MagicNorm and warmup deep credit assignment. Furthermore, instead of standard raw-text pretraining, we train exclusively on instruction-response pairs using a task-completion objective and PrefixLM masking. Serving as an empirical existence proof of efficient pretraining, a 1B-parameter HRM-Text model trained from scratch on only 40 billion unique tokens and $1,500 budget achieves 60.7% on MMLU, 81.9% on ARC-C, 82.2% on DROP, 84.5% on GSM8K, and 56.2% on MATH. Despite utilizing roughly 100-900x fewer training tokens and 96-432x less estimated compute than standard baselines, HRM-Text performs competitively with 2-7B parameter open models. These results demonstrate that co-designing architectures and objectives can radically reduce the compute-to-performance ratio, making pretraining from scratch accessible to the broader research community.

cs.CL

Optimal error bounds on the exponential wave integrator for nonlinear Schrödinger equations with highly singular potential

We establish error estimates of the first-order exponential wave integrator (EWI) for the nonlinear Schrödinger equation (NLSE) with a highly singular potential in $\mathbb{R}^d$ with $1\leq d \leq 3$. Our results deal with singular potentials in $L^p_\text{loc}(\mathbb{R}^d)$ with $p>\frac{d}{2}$ and $p\geq 1$, which is (almost) the weakest regularity of the potential required by the well-posedness of the NLSE. First, for $L^p_\text{loc}$-potentials with $p>2$, we establish an optimal first-order $L^2$-norm convergence for the EWI, with the convergence order slightly reduced to $1^-$ when $p=2$. To the best of our knowledge, the optimal first-order convergence for the three-dimensional $L^2$-potential is for the first time in the literature. The optimality of such an error bound is two-fold: (i) the first-order $L^2$-norm convergence is optimal for the EWI (and its higher-order versions) under the given $L^2$-regularity assumption on the potential, and (ii) to achieve the first-order $L^2$-norm convergence for the EWI, such an assumption is optimally weak. For more singular potentials in $L^p_\text{loc}(\mathbb{R}^d)$ with $\frac{d}{2} < p < 2$ and $p\geq 1$, we prove that the $L^2$-norm convergence is (almost) of $(1-α)$-order when $d=1,2$, and of $(1-\frac{3}{2}α)$-order when $d=3$, where $α:=d(1/p - 1/2)$ when $d =1,2,3$, $p>1$ and $α:=\frac{1}{2}^+$ when $d=1$, $p=1$. Notably, this result pushes the error estimate to the threshold regularity of the potential that matches the threshold regularity for the well-posedness of the NLSE, which is also for the first time. Two main ingredients are adopted in the proof: (i) the use of discrete space-time Lebesgue spaces together with discrete Strichartz estimates to establish the stability of the numerical scheme, and (ii) the use of normal form transformation and frequency decompositions to obtain optimal error bounds.

math.NA

Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation

We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence.

cs.CV

From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition

Managing extensive context remains a critical bottleneck for Large Language Models (LLMs), particularly in applications like long-document question answering and autonomous agents where lengthy inputs incur high computational costs and introduce noise. Existing compression techniques often disrupt local coherence through discrete token removal or rely on implicit latent encoding that suffers from positional bias and incompatibility with closed-source APIs. To address these limitations, we introduce the EDU-based Context Compressor, a novel explicit compression framework designed to preserve both global structure and fine-grained details. Our approach reformulates context compression as a structure-then-select process. First, our LingoEDU transforms linear text into a structural relation tree of Elementary Discourse Units (EDUs) which are anchored strictly to source indices to eliminate hallucination. Second, a lightweight ranking module selects query-relevant sub-trees for linearization. To rigorously evaluate structural understanding, we release StructBench, a manually annotated dataset of 248 diverse documents. Empirical results demonstrate that our method achieves state-of-the-art structural prediction accuracy and significantly outperforms frontier LLMs while reducing costs. Furthermore, our structure-aware compression substantially enhances performance across downstream tasks ranging from long-context tasks to complex Deep Search scenarios.

cs.CL

RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context

Large language models are evolving from single-turn responders into tool-using agents capable of sustained reasoning and decision-making for deep research. Prevailing systems adopt a linear pipeline of plan to search to write to a report, which suffers from error accumulation and context rot due to the lack of explicit control over both model behavior and context. We introduce RhinoInsight, a deep research framework that adds two control mechanisms to enhance robustness, traceability, and overall quality without parameter updates. First, a Verifiable Checklist module transforms user requirements into traceable and verifiable sub-goals, incorporates human or LLM critics for refinement, and compiles a hierarchical outline to anchor subsequent actions and prevent non-executable planning. Second, an Evidence Audit module structures search content, iteratively updates the outline, and prunes noisy context, while a critic ranks and binds high-quality evidence to drafted content to ensure verifiability and reduce hallucinations. Our experiments demonstrate that RhinoInsight achieves state-of-the-art performance on deep research tasks while remaining competitive on deep search tasks.

cs.CL

Regularization for the Schrödinger equation with rough potential: one-dimensional case

In this work, we investigate the following Schrödinger equation with a spatial potential \begin{align*} i\partial_t u+\partial_x^2 u+ηu=0, \end{align*} where $η$ is a given spatial potential (including the delta potential and $|x|^{-γ}$-potential). Our goal is to provide the regularization mechanism of this model when the potential $η\in L_x^r+L_x^\infty$ is rough. In this paper, we mainly focus on one-dimensional case and establish the following results: 1) When the potential $η\in L_x^1+L_x^\infty(\mathbb{R})$, then the solution is in $H_x^{\frac 32-}(\mathbb{R})$; however, there exists some $η\in L_x^1+L_x^\infty(\mathbb{R})$ such that the solution is not in $H_x^{\frac 32}(\mathbb{R})$; 2) When the potential $η\in L_x^r+L_x^\infty(\mathbb{R})$ for $1 2$, then the solution is in $H_x^{2}(\mathbb{R})$; however, there exists some $η\in L_x^r+L_x^\infty(\mathbb{R})$ such that the solution is not in $H_x^{2+}(\mathbb{R})$. Hence, we provide a complete classification of the regularity mechanism. Our proof is mainly based on the application of the commutator, local smoothing effect and normal form method. Additionally, we also discuss, without proof, the influence of the existence of nonlinearity on the regularity of solution.

math.AP

Regularization for the Schrödinger equation with rough potential: high-dimensional case

In this work, we investigate the regularization mechanisms of the Schrödinger equation with a spatial potential $$ i\partial_t u+Δu+ηu =0, $$ where $η$ denotes a given spatial potential. The regularity of solutions constitutes one of the central problems in the theory of dispersive equations. Recent works \cite{Bai-Lian-Wu-2024, M-Wu-Z24} have established the sharp regularization mechanisms for this model in the whole space $\mathbb{R}$ and on the torus $\mathbb{T}$, with $η$ being a rough potential. The present paper extends the line of research to the high-dimensional setting with rough potentials $η\in L_x^r+L_x^{\infty}$. More precisely, we first show that when $1\leq r <\frac d2$, there exists some $η\in L_x^r+L_x^{\infty}$ such that the equation is ill-posed in $H_x^γ$ for any $γ\in \mathbb{R}$. Conversely, when $\frac d2 \leq r \leq \infty$, the expected optimal regularity is given by $$H_x^{γ_*}, \quad γ_*=\mbox{min}\{2+\frac d2-\frac dr, 2\}.$$ We establish a comprehensive characterization of the regularity, with the exception of two dimensional endpoint case $d=2, r=1$. Our novel theoretical framework combines several fundamental ingredients: the construction of counterexamples, the proposal of splitting normal form method, and the iterative Duhamel construction. Furthermore, we briefly discuss the effect of the interaction between rough potentials and nonlinear terms on the regularity of solutions.

math.AP

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-based free-form editing. To solve this challenge, we introduce a novel framework that unifies speech understanding, generation, and editing. The core of our unified model is a unified continuous speech tokenizer MingTok-Audio, the first continuous tokenizer to effectively integrate semantic and acoustic features, which makes it suitable for both understanding and generation tasks. Based on this unified continuous audio tokenizer, we developed the speech language model Ming-UniAudio, which achieved a balance between generation and understanding capabilities. Ming-UniAudio sets new state-of-the-art (SOTA) records on 8 out of 12 metrics on the ContextASR benchmark. Notably, for Chinese voice cloning, it achieves a highly competitive Seed-TTS-WER of 0.95. Leveraging this foundational model, we further trained a dedicated speech editing model Ming-UniAudio-Edit, the first speech language model that enables universal, free-form speech editing guided solely by natural language instructions, handling both semantic and acoustic modifications without timestamp condition. To rigorously assess the editing capability and establish a foundation for future research, we introduce Ming-Freeform-Audio-Edit, the first comprehensive benchmark tailored for instruction-based free-form speech editing, featuring diverse scenarios and evaluation dimensions spanning semantic correctness, acoustic quality, and instruction alignment. We open-sourced the continuous audio tokenizer, the unified foundational model, and the free-form instruction-based editing model to facilitate the development of unified audio understanding, generation, and manipulation.

cs.CL

ScenGAN: Attention-Intensive Generative Model for Uncertainty-Aware Renewable Scenario Forecasting

To address the intermittency of renewable energy source (RES) generation, scenario forecasting offers a series of stochastic realizations for predictive objects with superior flexibility and direct views. Based on a long time-series perspective, this paper explores uncertainties in the realms of renewable power and deep learning. Then, an uncertainty-aware model is meticulously designed for renewable scenario forecasting, which leverages an attention mechanism and generative adversarial networks (GANs) to precisely capture complex spatial-temporal dynamics. To improve the interpretability of uncertain behavior in RES generation, Bayesian deep learning and adaptive instance normalization (AdaIN) are incorporated to simulate typical patterns and variations. Additionally, the integration of meteorological information, forecasts, and historical trajectories in the processing layer improves the synergistic forecasting capability for multiscale periodic regularities. Numerical experiments and case analyses demonstrate that the proposed approach provides an appropriate interpretation for renewable uncertainty representation, including both aleatoric and epistemic uncertainties, and shows superior performance over state-of-the-art methods.

cs.LG

System-Level Performance and Communication Tradeoff in Networked Control with Predictions

Distributed control of large-scale systems is challenging due to the need for scalable and localized communication and computation. In this work, we introduce a Predictive System-Level Synthesis PredSLS framework that designs controllers by jointly integrating communication constraints and local disturbance predictions into an affine feedback structure. Rather than focusing on the worst-case uncertainty, PredSLS leverages both current state feedback and future system disturbance predictions to achieve distributed control of networked systems. In particular, PredSLS enables a unified system synthesis of the optimal $κ$-localized controller, therefore outperforms approaches with post hoc communication truncation, as was commonly seen in the literature. The PredSLS framework can be naturally decomposed into spatial and temporal components for efficient and parallelizable computation across the network, yielding a regret upper bound that explicitly depends on the prediction error and communication range. Our regret analysis not only reveals a non-monotonic trade-off between control performance and communication range when prediction errors are present, but also guides the identification of an optimal size for local communication neighborhoods, thereby enabling the co-design of controller and its underlying communication topology.

eess.SY