SearcharxivSearch

arXiv subjects

Jie Ni

Publications and source records attributed to Jie Ni.

6 recordsLinked to original sources

A Multi-View Coupled Tensor Decomposition for Lightweight Online Adaptive Traffic Prediction

Accurate online traffic prediction is essential for intelligent transportation systems, where forecasting must be performed continuously under imperfect sensing conditions. Missing observations and anomalous disturbances make this task challenging, particularly when prediction relies on a single traffic view. This paper proposes a Multi-View Coupled Tensor Decomposition (MVCTD) model for online traffic prediction from imperfect multi-view observations, such as speed, flow, and occupancy. The proposed model uses coupled tensor decomposition to build a structured latent forecasting space, in which shared spatial structures across traffic views and view-specific temporal dynamics are jointly modeled. A group sparse regularization is further introduced to capture correlated abnormal responses induced by real traffic anomalies and thus reduce their influence on forecasts. For streaming deployment, MVCTD performs iterative refinement only on the current latent tensor, while the remaining model variables are updated by lightweight closed-form steps based on summarized historical information, thereby avoiding repeated optimization over the full historical sequence. Experiments on real-world traffic datasets demonstrate that MVCTD achieves accurate forecasts with favorable runtime under severe missingness, confirming its suitability for online traffic prediction.

math.OC

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. As a result, training continues to allocate gradient budget to already-solved objectives instead of focusing on those with greater remaining headroom. We introduce \textbf{Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization} (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This dynamically reallocates optimization effort toward under-optimized objectives while empirically maintaining performance on those that are already well satisfied. We further show that saturation-aware reweighting can reverse the sign of an update, rather than merely rescale its magnitude. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5\%$ on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by $3.8\%$ on average and up to $9.2 \%$ on AMC23, and on coding benchmarks it improves pass rate by up to $2.3\%$, while in all settings maintaining the easier objectives near their already satisfied levels.

cs.LG

SafeAgent: A Runtime Protection Architecture for Agentic Systems

Large language model (LLM) agents are vulnerable to prompt-injection attacks that propagate through multi-step workflows, tool interactions, and persistent context, making input-output filtering alone insufficient for reliable protection. This paper presents SafeAgent, a runtime security architecture that treats agent safety as a stateful decision problem over evolving interaction trajectories. The proposed design separates execution governance from semantic risk reasoning through two coordinated components: a runtime controller that mediates actions around the agent loop and a context-aware decision core that operates over persistent session state. The core is formalized as a context-aware advanced machine intelligence and instantiated through operators for risk encoding, utility-cost evaluation, consequence modeling, policy arbitration, and state synchronization. Experiments on Agent Security Bench (ASB) and InjecAgent show that SafeAgent consistently improves robustness over baseline and text-level guardrail methods while maintaining competitive benign-task performance. Ablation studies further show that recovery confidence and policy weighting determine distinct safety-utility operating points.

cs.AI

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike previous agents that only scale up model size or context length, MiroThinker explores interaction scaling at the model level, systematically training the model to handle deeper and more frequent agent-environment interactions as a third dimension of performance improvement. Unlike LLM test-time scaling, which operates in isolation and risks degradation with longer reasoning chains, interactive scaling leverages environment feedback and external information acquisition to correct errors and refine trajectories. Through reinforcement learning, the model achieves efficient interaction scaling: with a 256K context window, it can perform up to 600 tool calls per task, enabling sustained multi-turn reasoning and complex real-world research workflows. Across four representative benchmarks-GAIA, HLE, BrowseComp, and BrowseComp-ZH-the 72B variant achieves up to 81.9%, 37.7%, 47.1%, and 55.6% accuracy respectively, surpassing previous open-source agents and approaching commercial counterparts such as GPT-5-high. Our analysis reveals that MiroThinker benefits from interactive scaling consistently: research performance improves predictably as the model engages in deeper and more frequent agent-environment interactions, demonstrating that interaction depth exhibits scaling behaviors analogous to model size and context length. These findings establish interaction scaling as a third critical dimension for building next-generation open research agents, complementing model capacity and context windows.

cs.CL

Omnidirectionally manipulated skyrmions in an orientationally chiral system

Skyrmions, originally from condensed matter physics, have been widely explored in various physical systems, including soft matter. A crucial challenge in manipulating topological solitary waves like skyrmions is controlling their flow on demand. Here, we control the arbitrary moving direction of skyrmions in a chiral liquid crystal system by adjusting the bias of the applied alternate current electric field. Specifically, the velocity, including both moving direction and speed can be continuously changed. The motion control of skyrmions originates from the symmetry breaking of the topological structure induced by flexoelectric-polarization effect. The omnidirectional control of topological solitons opens new avenues in light-steering and racetrack memories.

cond-mat.soft

The thermal power generation and economic growth in the central and western China: A heterogeneous mixed panel Granger-Causality approach

The problem of the new energy economy has become a global hot issue. This study examines the causal relationship between the ratio of thermal power in total power generation (RTPG) and economic growth (GDP) in the western and central China by using the heterogeneous mixed panel Granger causality approach that accounts for both slope heterogeneity and cross-sectional dependence. For the overall panel, the empirical findings support the presence of unidirectional causality running from GDP to RTPG (in northwest China), and from RTPG to GDP (in central). At the provincial level, there is causality from GDP to RTPG in NeiMongol and Ningxia, and causality from RTPG to GDP in Shanxi, Anhui, and Jiangxi. As for the cross regions relationships, we find that GDP (in western) Granger-cause RTPG (in central), and RTPG (in southwest) Granger-cause GDP (in central and northwest). Moreover, panel regressions show the negative impact from GDP to RTPG in the northwest, and RTPG to GDP in the central. However, RTPG has a positive influence on GDP in the northwest. Therefore, to improve economic development without compromising the regions' competitiveness in central and western China, we can adjust the power generation structure, and increase investments in the renewable energy supply and energy efficiency.

stat.AP