SearcharxivSearch

arXiv subjects

Jingyu Liu

Publications and source records attributed to Jingyu Liu.

At least 19 recordsLinked to original sources

Approximate Inversion of Discrete Fourier Integral Operators via Hierarchically Semiseparable Matrices

This paper introduces a novel method for approximating the inverse of discrete Fourier integral operators (FIOs). Given an $N \times N$ matrix representation $K$ of an FIO, the proposed algorithm consists of two stages. In the offline stage, we first construct a butterfly factorization (BF) $\tilde{K}$ of $K$, which enables fast forward matrix-vector multiplication. We then construct a hierarchically semiseparable (HSS) approximation $\tilde{G} \approx G$, where $G = K^{*} K$, using fast applications of $\tilde{K}$ and $\tilde{K}^{*}$ to random matrices. Finally, we apply the ULV factorization to the HSS matrix $\tilde{G}$ to obtain an approximation $\tilde{F} \approx G^{-1}$. Combining these approximations yields an approximate inverse $K^{-1} \approx \tilde{F} \tilde{K}^{*}$. The offline stage has complexity $O(N \log^{2} N)$ for 1D problems and $O(N^{1.5} \log N)$ for 2D problems. In the online stage, the proposed method approximates $K^{-1} u$ for a given input vector $u$ with complexity $O(N \log N)$ for both 1D and 2D problems. The proposed method can be used either as a direct solver or as a preconditioner for iterative methods. Numerical results for 1D and 2D FIOs demonstrate the effectiveness of the proposed method.

math.NA

A Gradient-based yet Spike-Timing-Dependent Solution to the Feedback Learning Problem in Neural Microcircuits

The brain uses discrete spikes for dynamic computation, yet, how neural microcircuits (NMCs) solve temporal credit assignment using local spike timing remains a fundamental open question. Dominant spiking neural network (SNN) approaches circumvent this by approximating backpropagation through surrogate gradients, decoupling learning from biological spike timing. Here, we reformulate temporal credit assignment as a state separation problem: extracting task-required components induced by historical perturbations directly from the current neural state. This enables an online feedback learning framework for NMCs through a gradient tunneling (GT) algorithm and the lead-lag expansion technique that derives credit assignment from local synaptic spike timing, while remaining compatible with ANN-SNN hybrid architectures. Experimentally, GT-trained NMCs excel at long-timescale evidence integration and noise-robust memory retention, and perform comparably to leading SNN online learning methods on real-world benchmarks with far fewer parameters. The proposed framework addresses the two-decade-old NMC feedback learning problem and suggests a computationally plausible explanation for the brain's learning mechanisms.

cs.NE

MSR-IVA: Masked Structural Residual Independent Vector Analysis for State-Aware Fusion of Structural MRI and Dynamic Functional Network Connectivity

Multimodal fusion of structural MRI (sMRI) and dynamic functional network connectivity (dFNC) can reveal how brain structure relates to changing functional states. When the same structural latent representation is coupled with multiple states, applying independent vector analysis (IVA) separately to each state can produce unrelated structural decompositions, while forcing identical decompositions may suppress state-specific relationships. In addition, not every subject expresses every dynamic state. We propose masked structural residual IVA (MSR-IVA), a state-aware framework that combines a shared structural representation with state-specific residual adaptations and masks for incomplete state expression. On an Alzheimer's Disease Neuroimaging Initiative cohort, MSR-IVA improved matched source coupling by 6.5% and reduced unmatched dependence by 15.7% relative to the independent pairwise IVA baseline. Among subjects expressing both states, mean absolute cross-state structural source correlation was 0.9177 for MSR-IVA versus 0.2978 for no sharing, demonstrating controlled structural sharing that preserves source correspondence while allowing state-specific adaptation.

cs.LG

AdaptICA: Data-Adaptive Transformation Learning for Independent Component Analysis

Independent component analysis (ICA) is widely used to recover latent structure from signal and imaging data, but standard ICA assumes that the observed measurement scale preserves a linear mixing structure. This assumption may fail for features produced through nonlinear preprocessing, such as band-specific power in motor-imagery EEG. We propose AdaptICA, an adaptive transformation-based framework that jointly learns grouped componentwise transformations and the demixing structure using a profiled mutual-information criterion. Because the transformation and demixing parameters may compensate for one another, their joint estimation introduces new identifiability and asymptotic challenges. We establish identifiability, consistency, and asymptotic normality of the transformation estimator, together with joint strong consistency of the transformation and demixing estimators. AdaptICA selects the transformation structure data-adaptively and includes the identity transformation as a candidate, thereby reducing to standard ICA when no scale adjustment is needed. Extensive simulations support the theoretical results. Applications demonstrate that AdaptICA can recover more independent and interpretable sources when transformation is beneficial while retaining standard ICA when the original measurement scale is adequate.

stat.ME

From PBS to ePBS: the Microstructure of Block Building

Ethereum's Glamsterdam upgrade introduces enshrined proposer-builder separation (ePBS), replacing relay-centric PBS with direct builder bids to proposers. We study how this shift changes the block-building microstructure through a general imperfect-information two-stage auction with verifiable messages, where an early bid serves as both a price offer and a signal. PBS and ePBS are modeled as restrictions of the same block-building game: PBS fixes stopping and disclosure exogenously, while ePBS lets the proposer choose stopping and disclosure ex post. Latency heterogeneity is captured by asymmetric information updates: fast builders observe disclosed early information before rebidding, while slow builders do not. We combine exact perfect Bayesian equilibrium characterizations in tractable cases with calibrated no-regret learning in finite games. For PBS, we show that separating equilibria preserve the standard first-price-auction payoff benchmark and provide conditions for their existence. For ePBS, we demonstrate a ratchet effect: because the proposer can defer block proposal and use early bid information in the second stage, builders anticipate ex-post extraction and shade or pool early bids, generating allocation inefficiency and revenue-efficiency valleys. We interpret this ratchet distortion as a commitment failure. Under full commitment, the optimal policy collapses to the static Myerson auction and removes the ratchet channel. To realize part of this commitment advantage in a feasible mechanism, we propose a Trusted Execution Environment (TEE) sidecar that enforces limited commitment. We formulate the revenue-maximizing TEE mechanism as a bilinear optimization problem. In conservative finite benchmarks, the TEE design increases the proposer revenue relative to the first-price benchmark by approximately \(25\%\).

cs.GT

Sketch-and-Restart: Randomized Sketching in Quadrature-Based Restarting for Matrix Functions

We develop a sketch-and-restart framework for computing the action of a matrix function on a vector, $f(A) b$, where $A$ is large, sparse, and non-Hermitian. The framework combines quadrature-based restarting with Arnoldi-like decompositions generated by sketched or truncated Arnoldi processes. Within this framework, we develop two classes of restarted algorithms. The first uses a fixed Krylov subspace dimension and is based either on the sketched Arnoldi process or on a new sketched harmonic Arnoldi process proposed in this work. The second class chooses the Krylov subspace dimension adaptively by running the truncated Arnoldi process until the condition number of the generated basis, estimated from its sketch, exceeds a prescribed threshold. We also establish the convergence of the restarted sketched harmonic Arnoldi method for Stieltjes functions under the assumption that $A$ is positive real. Numerical experiments demonstrate the effectiveness of the proposed framework, including the computational savings achieved through sketching, the storage reduction enabled by adaptive truncation, and the acceleration obtained from thick restarting.

math.NA

Long-range coupling enabled multiband group-velocity control of topological edge states from slow light to light stopping

Topological edge states provide robust optical transport immune to disorder, yet their propagation velocity is usually constrained by the intrinsic band dispersion, limiting dynamic control of topological light transport. We introduce long-range next-nearest-neighbor (NNN) couplings into a Harper--Hofstadter photonic lattice and establish a versatile platform for group-velocity engineering. We demonstrate that the NNN couplings play two distinct roles: the vertical coupling opens a previously closed topological band gap by lifting the degeneracy of bulk bands, while the horizontal coupling reshapes the edge-state dispersion through momentum-dependent corrections, enabling controllable topological slow-light transport. Furthermore, the band-gap Chern numbers associated with different gaps exhibit opposite signs, giving rise to topological edge states with opposite chiralities. Propagation simulations reveal robust unidirectional transport of these counter-chiral edge states with reduced group velocities. By continuously tuning the NNN coupling strength, the group velocity of topological edge modes can be reduced toward zero at specific momenta, resulting in topological light-stopping effects. These results demonstrate that long-range NNN couplings provide an effective mechanism for engineering momentum-dependent topological group velocities and offer new possibilities for robust slow-light devices, optical delay lines, and multiband integrated photonic systems.

physics.optics

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.

cs.CL

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in \textit{GigaWorld-1}, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.

cs.RO

A Superfast Direct Solver for 2D Type-II Inverse Nonuniform Discrete Fourier Transform Based on Hierarchically Semiseparable Matrix

This paper proposes a direct inversion method for the 2D type-II nonuniform discrete Fourier transform~(NUDFT). The NUDFT matrix $A$ is factored as $A = G F$, where $G$ can be expressed as a kernel matrix and $F$ is the 2D DFT matrix. We show that $G$ can be approximated by a hierarchically semiseparable~(HSS) matrix and give an estimate of the HSS rank. Then, using the least-squares solver for HSS matrix and the two-dimensional inverse fast Fourier transform, the inverse NUDFT problem can be solved efficiently. Our algorithm has an offline complexity of $O\bigl(M+ N^{3 / 2} \log^{3} N\bigr)$ where $M$ and $N$ are the size of rows and columns of the NUDFT matrix, respectively. Once the direct solver is built, it can be applied to a vector with an online complexity of $O\bigl(M+ N \log^{3} N\bigr)$. The proposed method can be used as a preconditioner for iterative methods, especially when the sample points are distributed on a grid such that $A$ is ill-conditioned. Numerical results are provided to show the scaling performance of the inversion method and demonstrate the efficiency and robustness of it as a preconditioner.

math.NA

Diagnosing Task Insensitivity in Language Agents

Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identify a key source of this failure as task insensitivity: when faced with similar but distinct tasks, models might apply patterns learned during training and fail to solve the task at hand. We show that models often continue with actions aligned with the original task even when the instruction is semantically corrupted and cannot be directly answered. We further find that, when we replace the task description in a trained prompt with another similar but distinct task, the model may still output the same action. This behavior is accompanied by a consistent training-time attention drift away from task tokens and toward local observations, suggesting an optimization bias toward shortcuts. To mitigate this problem, we propose Task-Perturbed NLL Optimization, a lightweight contrastive regularizer that explicitly encourages action dependence on the task instruction. Extensive evaluations show that our intervention improves task sensitivity and OOD generalization while preserving more stable attention to task tokens.

cs.AI

Where Do CoT Training Gains Land in LLM based Agents?

Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may instead reflect post-hoc reasoning, which means the model already knows the answer before reasoning. We therefore ask what CoT training is actually improving: is the model getting better at changing its action through generated reasoning, or is it getting better at predicting the action directly from the prompt? We study this question by comparing \emph{prompt actions} (predicting action without CoT) with CoT actions (predicting action with CoT). Across checkpoints, prompt-action quality improves substantially. While interacting with the environment, the relative advantage of CoT actions over prompt actions remains similar, showing that CoT training does not widen the advantage of CoT reasoning, and it helps to improve the quality of prompt actions. We further find that later checkpoints are less likely to revise the action in response to CoT, suggesting greater reliance on the prompt. Motivated by these patterns, we selectively mask action-token supervision on a fraction of training examples. This intervention improves out-of-domain generalization.

cs.AI

M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking

Building a general-purpose whole-body controller is essential for enabling diverse motion capabilities in humanoid robots across a wide range of downstream tasks, including locomotion and loco-manipulation. Different tasks rely on distinct motion reference modalities: locomotion primarily depends on coordinated robot joint trajectories, whereas manipulation requires precise end-effector trajectory tracking. Existing methods often overlook the representational mismatch between dense robot joint angles and sparse end-effector poses. To address this, we propose Multi-Modal Mimic (M3imic), a versatile multi-modal whole-body control framework that unifies heterogeneous motion reference modalities, including robot joint angles, human pose trajectories, and end-effector poses, using modality-specific encoders to map them into a shared latent space. Leveraging large-scale reinforcement learning in the simulator, we train a single policy that achieves sim-to-real transfer across multiple motion reference modalities without modality-specific retraining. Extensive simulation and real-world experiments on the Unitree G1 robot are conducted to evaluate the proposed framework. In simulation, the policy achieves a peak success rate of 98.42\% on an unseen test dataset, demonstrating its exceptional generalization capability. The code is available at https://github.com/Renforce-Dynamics/MultiModalWBC

cs.RO

Isolating Nonlinear Independent Sources in fMRI with $β$-TCVAE Models

Learning meaningful latent representations from nonlinear fMRI data remains a fundamental challenge in neuroimaging analysis. Traditional independent component analysis, widely used due to its ability to estimate interpretable functional brain networks, relies on a linear mixing assumption for latent sources, limiting its ability to capture the inherently nonlinear and complex organization of brain dynamics. More recently, deep representation learning methods have emerged as promising alternatives for modeling nonlinear latent structure. However, many of these approaches have been evaluated primarily on simulated datasets or natural image benchmarks, with comparatively limited validation on real-world neuroimaging data such as fMRI. In this work, we are motivated by the $β$-TCVAE (Total Correlation Variational Autoencoder), a refinement of the $β$-VAE framework for learning latent representations without introducing additional hyperparameters during training. We adapt and modify this model to fMRI data for nonlinear source disentanglement, aiming to separate mixed spatial and temporal brain signals into interpretable components. We show that the $β$-TCVAE framework can recover meaningful nonlinear spatial components with biological relevance, including well-established intrinsic connectivity networks such as the default mode network. Furthermore, we evaluate the learned representations using functional network connectivity, showing that the latent structure captures coherent and interpretable brain organization patterns. This study provides a pilot investigation that bridges nonlinear representation learning and fMRI analysis.

cs.LG

NeuroGAN-3D: Enhancing Intrinsic Functional Brain Networks via High-Fidelity 3D Generative Super-Resolution

Recent advances in neuroimaging have deepened our understanding of the brain's complex functional and structural organization. Among these, functional Magnetic Resonance Imaging (fMRI) - particularly resting-state fMRI (rs-fMRI) - has emerged as a tool for identifying biomarkers of intrinsic brain connectivity and delineating large-scale neural networks. These networks are typically represented as volumetric spatial maps that capture functionally coherent brain regions and reflect individual differences in brain activity and structure. The spatial resolution of these maps plays an important role, as it determines the ability to localize functional units with precision, perform reliable brain parcellation, and detect subtle, spatially specific neurobiological alterations associated with development, aging, or disease. Therefore, improving the effective resolution of neuroimaging-derived maps holds significant promise for enabling more detailed insights into brain architecture and its relationship to behavior and pathology. To address this need, we propose NeuroGAN-3D, a novel 3D generative super-resolution model tailored to the computational demands of volumetric neuroimaging. Our model leverages a generative adversarial network architecture to enhance the spatial resolution of rs-fMRI spatial maps, significantly outperforming a conventional baseline.

cs.CV

Probabilistic Assessment of Rare Transient Instability Events via Kriging-based Active Learning Framework

The increasing uncertainty in modern power systems, driven by the integration of intermittent energy sources and variable loads, underscores the need for probabilistic transient stability assessment. However, existing assessment methods primarily focus on average system stability behavior and may struggle or incur high computational cost when identifying rare transient instability events, which in turn are critical for ensuring system resilience. To address this, the paper proposes a Kriging-based active learning framework to accurately characterize rare instability regions within the input uncertainty space and estimate the associated small instability probability, while requiring only a limited number of expensive time-domain simulations. The proposed active learning (AL) framework is tested on a modified IEEE 59-bus system with simulated load and wind uncertainties, and a WECC 240-bus system incorporating real-world wind and solar generation data. Comparative studies with the existing random forest-based active learning method and three non-AL methods demonstrate that the proposed AL framework achieves superior accuracy and computational efficiency.

eess.SY

Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving

Prefill-Decode (PD) disaggregation has become the standard architecture for modern LLM inference engines, which alleviates the interference of two distinctive workloads. With the growing demand for multi-turn interactions in chatbots and agentic systems, we re-examined PD in this case and found two fundamental inefficiencies: (1) every turn requires prefilling the new prompt and response from the last turn, and (2) repeated KV transfers between prefill and decode nodes saturate the bandwidth, leading to high latency and even service degradation. Our key insight is that not all prefill operations are equally disruptive: append-prefill, which processes only the new input tokens while reusing cached KV states, incurs an order-of-magnitude smaller decoding slowdown than full prefill. This motivates routing append-prefill to decode nodes locally. However, through comprehensive analysis, we show that no single fixed routing strategy satisfies all Service Level Objectives (SLOs) simultaneously. Based on this insight, we propose Prefill Prefill-capable Decode (PPD) disaggregation, a dynamic routing system that decides when to process Turn 2+ requests locally on decode nodes using cached KV states. PPD adapts to varying SLOs via configurable weights and seamlessly integrates with traditional PD deployments. With extensive evaluations, we show that PPD reduces Turn 2+ time-to-first-token (TTFT) by $\sim$68\% while maintaining competitive time-per-output-token (TPOT), effectively alleviating KV transfer congestion under high load. PPD provides a flexible and efficient paradigm for multi-turn LLM serving.

cs.NI

Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices

Large Multimodal Models (LMMs) are inherently modular, comprising vision and audio encoders, a projector, and a language backbone. Yet existing systems execute them monolithically, underutilizing the heterogeneous accelerators (NPUs, GPUs, DSPs) on modern SoCs and inflating end-to-end latency. We present Nanomind, a hardware-software co-design inference framework that decomposes each LMM into modular "bricks"--vision, projector, language, and audio--and maps each brick to its best-suited compute units. A Token-Aware Buffer Manager (TABM) enables zero-copy embedding transfer across accelerators on unified-memory SoCs, bypassing CPU bottlenecks. Combined with customized hardware, a battery-aware scheduler, and fused low-bit GEMM kernels, Nanomind runs entirely on a compact, battery-powered prototype that operates fully offline. Nanomind reduces end-to-end energy by 42.3% against mainstream edge frameworks and devkits; in its on-demand low-power mode, the prototype runs LLaVA-OneVision-Qwen2-0.5B with a camera for nearly 18.8 hours on a single 2,000 mAh battery.

cs.DC