SearcharxivSearch

arXiv subjects

Jian Sun

Publications and source records attributed to Jian Sun.

At least 19 recordsLinked to original sources

Cartesian tensor equivariant machine-learning force field for spin-dependent atomistic simulations

Magnetic materials exhibit an intricate coupling between atomic structure and spin degrees of freedom, posing a fundamental challenge for atomistic simulations across experimentally relevant length and time scales. Here we introduce HotPP-Spin, a spin-dependent extension of HotPP for magnetic machine learning interatomic potentials, built on Cartesian tensor equivariant message passing. Atomic magnetic moments are treated as explicit axial-vector degrees of freedom, while spatial-inversion and time-reversal parities are propagated through the tensor couplings. This construction provides a unified representation of exchange-dominated and spin-orbit-induced interactions without imposing predefined analytical interaction forms. A scalar spin-dependent potential energy surface yields energy-conserving atomic forces and magnetic effective fields through differentiation. Benchmarks spanning collinear magnetism, noncollinear magnetism, and spin-orbit-coupling-induced magnetic anisotropy show that HotPP-Spin accurately describes magnetic energy landscapes, magnetic forces, and magnetic-order-dependent energy-volume relations within the same general framework. For H-phase monolayer VSe\(_2\), stochastic spin-dynamics simulations using the learned magnetic effective fields locate the finite-size magnetic ordering crossover at 415--435~K, in close numerical agreement with the reported experimental value of \(418.5\pm7.8\)~K. These results establish Cartesian tensor message passing as a general route for connecting first-principles magnetic energetics with large-scale atomistic simulations of coupled structural and spin phenomena.

physics.comp-ph

HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory

Long-term memory is crucial for personalized responses and long-horizon agent interactions. Existing methods often rely on LLMs to compress or rewrite dialogue histories and use the transformed memories as retrieval evidence. Despite the progress in organizing fragmented contexts, two major drawbacks persist: (1) information loss from compression, which discards fine-grained but later useful details, and (2) semantic drift from rewriting, which erodes the original tone and situated context. In this work, we propose a novel Human-profile Enhanced Retrieval Optimization framework for long-term agent memory (HERO). Specifically, HERO converts the dialogue history into a traceable heterogeneous memory graph that preserves raw dialogue text as evidence for reasoning, thereby mitigating information loss. For retrieval, HERO extracts initial anchors from the current query and incorporates human profiles via an iterative graph traversal; these anchors and profiles provide guidance signals that adaptively activate the most informative regions of the graph. Experiments on two benchmark datasets show that HERO outperforms strong baselines on both factual and personalized reasoning, while providing more faithful access to raw dialogue evidence.

cs.AI

Coupled Optimal Transport with Landmark Constraints

Existing optimal transport (OT) models primarily seek an OT map or plan between distributions by minimizing a prescribed transport cost or distortion. However, minimizing transport cost or distortion alone may fail to identify a geometrically meaningful transformation between the two distributions. To address this limitation, this paper proposes a novel coupled OT framework that leverages a small number of annotated landmarks to guide the recovery of an underlying deformation governing the distribution transformation. The coupled OT framework integrates the optimization of the transport plan and the deformation field into a unified model, where the landmark-guided deformation field and the cost-driven transport plan are coupled through a mutual-consistency constraint. As a result, the deformation is jointly determined by the annotated landmarks and cost-driven distribution matching. The proposed framework provides a principled connection between landmark-based registration and transport-based distribution matching, enabling the recovery of transport maps from sparse geometric supervision. We establish the well-definedness of the proposed model in a general variational setting and develop a finite-element-based numerical algorithm for computation whose convergence properties are systematically analyzed. The practical effectiveness of the proposed approach is verified in shape matching.

cs.CV

Remarkable Enhancement of High Harmonic Generation from Superhard Material under High Pressure

High harmonic generation (HHG) in solids offers a pathway to develop compact extreme ultraviolet (EUV) sources crucial for attosecond science and advanced spectroscopy. Here, we demonstrate theoretically that high pressure dramatically enhances HHG in superhard hexagonal tungsten nitride (h-WN6). Compared with solid-state systems at ambient pressure, the reshaped electronic environment under high pressure leads to a unique band-gap widening in h-WN6, which raises the material's damage threshold, allowing the use of stronger laser fields and enabling access to higher-energy bands. This pressure-induced band-gap widening offers a promising strategy to overcome the cutoff limitation of solid-state EUV light sources.

physics.optics

DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing

End-to-end (E2E) autonomous driving algorithms require rigorous closed-loop validation in simulation environments offering high visual fidelity, strong interactivity, and real-time performance. Existing approaches, from game engines to static neural rendering, inherently trade off these requirements and struggle with the dynamic scene composition essential for E2E testing. To bridge this gap, we propose a novel decoupled 3D Gaussian Splatting (3DGS) framework tailored for large-scale E2E evaluation. We fundamentally decompose scenes into a high-fidelity static background and manipulable dynamic agents using an object-centric canonical representation. To resolve resulting representational conflicts, we introduce three targeted modules: (1) asset compression via perceptual pruning and vector quantization for real-time traffic rendering; (2) map-guided geometric registration leveraging semantic topology to strictly align trajectories; and (3) proxy-based relighting transferring ambient illumination for seamless photometric integration. Extensive experiments demonstrate that DecoupleGS achieves a balanced fidelity-efficiency trade-off, improves metric and photometric consistency, and provides a practical closed-loop sensor simulation platform for E2E autonomous driving evaluation.

cs.CV

A Self-Triggered Agentic Push Recommendation System

Push notification is a critical recommendation scenario on large-scale platforms, allowing the system to proactively reach users outside the application to improve long-term re-engagement. However, designing an optimal push system requires handling a complex action space for the "whether and when" delivery problem under strict system resource constraints. Existing solutions typically fall into two passive paradigms: pre-planned frequency methods that allocate delivery times via offline modeling, limiting real-time adaptability; and fixed-interval triggering methods that periodically poll the system, creating a strict dilemma between excessive computational overhead and diminished optimal timing capture. Furthermore, such multi-stage frameworks severely suffer from local optima. To overcome these limitations, in this paper, we propose STEPS, a proactive, Self-Triggered End-to-end Agentic Push Recommendation System, which is already fully deployed at Douyin with over 1 billion users. STEPS reformulates push recommendation as a self-triggered agentic process in which the system decides not only whether to send a push, but also when to invoke itself again, thereby forming a closed loop that balances real-time effectiveness and efficiency. Specifically, STEPS consists of two decision transformer-based agents: a planning agent that schedules the next system invocation using a gated ordinal regression method, and an execution agent that decides whether to send a push based on trajectory rewards. Furthermore, we introduce a lightweight filtering agent to both control computational overhead and act as a crucial safeguard against unreasonable planning behaviors. Online A/B testing demonstrates that STEPS significantly increases user active days by 0.2843% and reduces the push permission disablement rate by 1.9089%, while the filtering agent reduces computational overhead by 79.42%.

cs.IR

Superconducting ternary compounds Li-X-B (X=Mo, W) within the mild pressure range: First-principles predictions

Among the superconducting hydrides under high pressure, a number of studies concentrate on the ternary compounds to explore unique superconductors, which are capable of reducing the stable pressure and maintain superconductivity. In this work, to verify our proposed strategy of ternary composition lines (TCLs) to explore ternary compounds, we combined the first-principles calculations and crystal structure predictions to study the ternary compounds Li-X-B (X=Mo, W) under high pressure. After calculations along five and four TCLs in Li-W-B and Li-Mo-B, respectively, five Li-W-B compounds and four Li-Mo-B compounds were predicted. The compositions of LiWB4, Li4MoB2 and LiMo2B2 could be thermodynamically stable under high pressure, and Li2WB6 is around 0.02 eV/atom above the convex hull at 0 GPa, which has potential for synthesizing. Both of the predicted Li2WB6 P6/mmm and Li2WB4 R-3m are superconducting and their Tc are around 11 K, which are similar to the Tc of WB2 P6/mmm around 100 GPa. An anomalous increase of Tc was found in Li4MoB2 C2/m upon compression. We carried out full ternary search (FTS) to evaluate the validity of the TCLs strategy in Li-W-B system at 0 GPa. Our results are helpful for understanding the phase diagram of Li-X-B (X=Mo, W) under high pressure and the introducing of Li atoms provide candidate structures to reduce the measured stable pressure from ~100 GPa in WB2 P6/mmm to 0 GPa. Meanwhile, we preliminary validate the strategy of TCLs in structure predictions and we expect to improve this strategy in the future, shedding light on the studies of ternary compounds.

cond-mat.supr-con

AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight

Language-conditioned quadrotor flight requires a policy to ground semantic goals, anticipate the visual consequences of ego-motion, and output control references that remain smooth and dynamically executable under rapidly changing first-person views. Existing aerial vision-language navigation and vision-language-action methods commonly use discrete actions, high-level waypoints, or instantaneous velocity commands, which provide limited supervision about how flight actions change future observations. We present AeroAct, an action-centered world-action model (WAM) for quadrotor navigation. To the best of our knowledge, AeroAct is the first WAM instantiated and demonstrated for real-world aerial flight. The model adapts a pretrained video diffusion Transformer to predict local trajectory-action chunks from egocentric visual history, proprioception, and language. Future first-person frames are used during training as dense consequence supervision, while deployment directly decodes actions without generating future video. To obtain aligned visual, state, language, and dynamically feasible action data, we build a DiffAero-based pipeline with complementary Isaac Lab and 3D Gaussian splatting renderers. We further introduce a low-cost handheld collection device that couples camera observations with motion estimates to recreate flight-like egocentric trajectories, and a self-guidance procedure that improves temporal consistency across overlapping trajectory chunks. Closed-loop simulation and real-world experiments show that temporal visual context improves target tracking and object-search performance, and that WAM-based policies can be executed on a physical quadrotor.

cs.RO

Cross-Field Channel Parameter Estimation and Channel Characterization at THz Bands in Indoor Scenarios

The terahertz (THz) frequency band offers the potential for ultra-high data rate transmission in future wireless communication systems. To extend the transmission distance and enhance spectral efficiency, the deployment of large-scale antenna arrays emerges as a promising solution in the THz band. This paper targets the critical challenge of cross-field (hybrid near-field/far-field) channel parameter estimation and channel characterization in such configurations. We first establish a 260-380 GHz virtual uniform linear array (ULA) measurement framework in an indoor scenario, capturing high-resolution channel transfer functions (CTFs) that reveal spatial non-stationarity and cross-field wavefront characteristics. Building upon these empirical observations, we propose a cross-field space-alternating generalized expectation-maximization (SAGE) algorithm that discriminatively estimates near-field and far-field multipath components (MPCs) via Bayesian phase-curvature classification, while explicitly tracking spatial birth-death phenomena through visibility region estimation. Analysis of the measurement data validates the algorithm's effectiveness in resolving cross-field MPCs and quantifies that near-field MPCs account for over 90% of total MPCs at 2 m transmission distance (380 GHz). We observe that spatial non-stationarity intensifies as the carrier frequency increases and the transmission distance decreases. These findings offer quantitative guidelines for channel modeling and system design in wireless THz communication systems.

eess.SP

World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning

Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods often rely on external scenario generators, heuristic perturbations, or simulator-heavy rollouts, which makes them difficult to integrate with modern autoregressive planners. Here, we cast adversarially robust planner learning as a constrained min-max game and propose Adversarial World Modeling (AWM), a theoretically grounded multi-agent self-play fine-tuning framework. Since solving the exact game is intractable, AWM introduces a principled decoupled solver. In the inner minimization, the planner's predictive world model is converted into a role-conditioned adversary that learns sparse, scene-adaptive attack coalitions via counterfactual credit assignment. In the outer maximization, the ego planner optimizes a regret-aware robust best response against the frozen AWM, utilizing tail-risk weighting and reference-anchored trust regions to improve hard-case recovery while preserving nominal driving behavior. Experiments on the nuPlan and InterPlan benchmarks demonstrate that our method generates transferable adversarial interactions and yields a robust planner that achieves competitive closed-loop performance in both nominal and highly interactive long-tail scenarios. Theoretical analysis justifies the decoupled solver and the main optimization components.

cs.RO

Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring

Wound monitoring is a critical yet underserved clinical challenge, where timely identification of severe adverse events (SAEs) such as infection, tissue deterioration, and delayed healing can significantly impact patient outcomes. While vision-language models (VLMs) show strong multimodal reasoning, they often lack domain-specific grounding to integrate wound imagery with heterogeneous clinical information, and provide limited mechanisms for detecting cases that diverge from the training distribution. We present a multimodal framework for automated wound monitoring and SAE detection. Our approach leverages paired clinical notes and wound descriptions capturing visual characteristics such as appearance, surrounding skin condition, color changes, and signs of inflammation or healing progression, encoded through a dual-stream Low-Rank Adaptation (LoRA) framework built on a frozen BiomedCLIP backbone. We introduce a cross-contextual LoRA fusion mechanism enabling information exchange between clinical semantics and visual wound descriptors, producing context-aware multimodal representations without full model fine-tuning. To identify personalized SAEs, we propose a wound-specific out-of-distribution (OOD) detection framework combining semantic matching, visual typicality, caption-text alignment, and caption-visual alignment into a unified SAE (OOD) score. To capture healing dynamics, we incorporate covariate consistency and temporal drift penalties that leverage changes in wound characteristics across visits. Experiments on a longitudinal wound dataset collected through clinical visits show promising performance on both wound healing assessment and SAE detection, highlighting the potential of semantically enriched, temporally aware vision-language systems for clinical wound monitoring and early risk identification.

cs.CV

High-order tensor neural network for iteration-free structure relaxation

Structure relaxation is important for the discovery of new materials, yet conventional ab initio optimization remains a major bottleneck in high-throughput screening workflows. Machine learning potentials have accelerated relaxation by orders of magnitude, but they still rely on iterative optimization and high-quality DFT force labels. Here, we present HotRelax, a high-order tensor message-passing neural network for one-shot, end-to-end prediction of relaxed structures. Trained directly on paired unrelaxed and relaxed structures, HotRelax requires no DFT force labels and predicts relaxed structures in a single forward pass, without iterative inference or post-processing. Across five diverse datasets spanning 3D bulk crystals, 2D layered materials and catalysts, HotRelax shows strong performance relative to state-of-the-art end-to-end relaxation models, achieving lower prediction errors on several benchmarks while maintaining a compact model size and efficient inference. Extensive DFT calculations further show that the predicted structures are close in energy to their DFT-relaxed counterparts. When integrated into catalytic workflows, HotRelax also improves the accuracy and generalization of relaxed-state energy prediction models. Together, these results support HotRelax as an efficient and widely applicable framework for end-to-end structure relaxation, with strong potential to accelerate high-throughput materials discovery.

physics.comp-ph

Data-Driven Robust MPC for Unknown Nonlinear Systems via Set-Membership Learning

Data-driven model predictive control (MPC) has become an attractive approach for controlling unknown systems, especially when data are corrupted by noise. However, most existing data-driven MPC methods focus on linear systems, and little attention has been given to nonlinear dynamics under disturbances. To fill this gap, we propose a robust data-driven min-max MPC scheme for unknown nonlinear systems with process disturbances. We represent the unknown nonlinear dynamics using vector fields built from a dictionary of basis functions, yielding an equivalent linear form with unknown matrices. These unknown matrices are characterized by a set-membership representation derived from noisy input-state data. Using this uncertainty description, we formulate a min-max MPC problem. Two online scenarios are studied: i) when state measurements are noise-free, and, ii) when they are corrupted by process disturbance. For each case, we derive a Lyapunov-based semidefinite program (SDP) to compute a stabilizing state-feedback controller. The resulting schemes are shown to guarantee recursive feasibility and either exponential or robust stability of the closed-loop system depending on whether there is process disturbance. Simulation studies on benchmark examples illustrate the effectiveness and competitive performance of the proposed approach compared to existing data-driven and model-based controllers.

eess.SY

Endogenous Randomness from Adversarial Market Learning

We propose a deterministic adversarial market model in which apparent randomness emerges endogenously from the interaction between a market mechanism and a population of predictive traders. Unlike a classical generative adversarial network, the model does not attempt to imitate an external empirical data distribution and does not inject random noise into a generator. The market is represented by a deterministic binary return path, while traders learn predictive strategies from observed in-sample history and trade on an out-of-sample continuation. The market then adapts against the traders by reducing their predictive and trading edge. The central experiment begins with a smooth, highly predictable market path. Traders with multiple lookback windows and multiple holding periods learn to predict future cumulative returns. Initially, these traders earn large out-of-sample profits. After adversarial market adaptation, their out-of-sample profitability collapses toward zero. Importantly, in the final clean specification, no explicit sign-balance, transition-rate, or autocorrelation penalties are imposed. Nevertheless, the out-of-sample return sequence becomes balanced, has transition rate close to one half, has low autocorrelation, and passes block-based distributional diagnostics. In a medium-size experiment with $T_{\mathrm{IS}}=2000$ and $T_{\mathrm{OOS}}=10000$, the out-of-sample positive-return fraction is $0.5010$, the transition rate is $0.4896$, and the maximum absolute autocorrelation is $0.0275$. Binary return blocks transformed into dyadic variables are close to uniform on $[0,1]$, and normalized block sums are broadly consistent with a standard normal law. These results support the hypothesis that market randomness can arise as the endogenous residue of arbitrage pressure rather than from exogenous stochastic shocks.

q-fin.TR

Monotonicity of Normalized Implied-Volatility Coordinates under No-Arbitrage

For a fixed maturity, an arbitrage-free option smile induces natural normalized strike coordinates. This paper makes three contributions. First, it gives an elementary discrete no-arbitrage proof of monotonicity for the central Black--Scholes normalized coordinate \(k/v(k)\), using only finite-strike comparisons, convexity, monotonicity, and put--call parity. Thus the argument applies directly to finitely quoted option chains and does not require a continuously quoted smile, differentiability of option prices, differentiability of implied volatility, digital prices, or density extraction. Second, it extends the same monotonicity principle to the normal, or Bachelier, implied volatility formula, proving that the normalized coordinate \((F-K)/\sigma_N(K)\) is decreasing in strike under static no-arbitrage. Third, it proves a model-free normal-variance identity: remaining normal variance can be represented as a normal-density weighted integral of squared Bachelier implied volatility in the normalized coordinate. This third result is the normal/Bachelier analogue of Fukasawa's lognormal variance identity, which expresses variance-type quantities through Black implied variance in normalized coordinates. The paper therefore complements Fukasawa's continuous-strike normalizing transformation theory with a finite-quote no-arbitrage proof and a new normal-variance counterpart, while connecting the results to the volatility-derivatives literature surveyed by Carr and Lee.

q-fin.MF

Binary Decompilation LLM with Feedback-Driven Multi-Turn Refinement

Binary decompilation is fundamental to security tasks such as vulnerability discovery, malware inspection, and executable-only program understanding. Recent LLM-based decompilation methods have shown promising results, but most still follow a single-turn generation paradigm: given assembly code or decompiler-produced pseudo-code, the model generates one output and stops. Consequently, the generated code may appear readable or even compile successfully, yet still deviate from the behavior of the original binary and mislead downstream analysis. This paper presents AutoDecompiler, a decompilation-specialized LLM trained with reinforcement learning for feedback-driven multi-turn binary decompilation. Instead of treating decompilation as one-shot code generation, AutoDecompiler formulates it as an iterative refinement process, where the model revises generated code based on compilation, execution, and input/output testing feedback. To enable this process, we design decompilation-specific rewards that capture code validity, recompilability, execution consistency, and semantic fidelity. We further construct stage-aware diagnostic feedback from compiler errors, execution failures, and failed test cases, and introduce progress-aware trajectory rewarding and turn-aware advantage reweighting to encourage beneficial revisions while suppressing regressions. We train the AutoDecompiler family and evaluate it across different input settings, model scales, and benchmarks. Experimental results show that AutoDecompiler consistently outperforms its single-turn counterparts under the same model size and input setting, achieving clear improvements in behavioral re-executability. These results demonstrate that learning to exploit program feedback with reinforcement learning is an effective direction for improving the functional correctness of LLM-based binary decompilation.

cs.SE

From Attacks to Curricula: Learnability-Guided Adversarial Training for Safe Autonomous Driving

Closed-loop adversarial training improves autonomous driving safety by exposing policies to rare safety-critical scenarios. Standard pipelines first generate adversarial scenarios and then sample them for policy optimization. However, most existing frameworks remain attack-oriented: collision-driven generators often synthesize unsolvable extreme situations, which can degrade learning, while heuristic samplers ignore the evolving capability of the driving policy, causing sample inefficiency and delayed convergence. We propose AlignADV, a learnability-guided closed-loop adversarial training framework that converts adversarial scenarios into resolvable and capability-aligned curricula. First, we reformulate adversarial scenario generation as a preference alignment problem and employ direct preference optimization to guide the generator toward critical yet resolvable scenarios. Second, we introduce behavioral fingerprints to capture the intrinsic characteristics of the evolving policy and construct a multi-modal capability prediction model that estimates policy performance without expensive closed-loop simulations. By combining resolvability-aligned scenarios with capability predictions, AlignADV develops a dynamic curriculum sampling mechanism that prioritizes scenarios targeting the current policy's vulnerabilities. Experiments on the Waymo Open Motion Dataset demonstrate that AlignADV improves convergence efficiency and final performance, reducing training steps by up to 40.6 percent compared with baseline methods while lowering collision rate and improving route completion under both normal and adversarial traffic conditions. These results highlight a shift from attack-oriented scenario generation to learnability-guided policy improvement, offering a principled direction for safer and more efficient autonomous driving training. Project page: https://meiyuewen.github.io/AlignADV/.

cs.RO

MAD: Mapping-Aware World Models for Agile Quadrotor Flight

Agile quadrotor flight in cluttered scenes requires more than a reactive mapping from a depth image to a control command: the vehicle must remember which regions have been observed, infer nearby occupied space, and act under partial visibility and tight latency. In this paper, we present Mapping-Aware Dreamer (MAD), a geometry-aware world model for vision-based quadrotor flight. Instead of using raw-image reconstruction as the main self-supervised objective, MAD learns recurrent latent dynamics that reconstruct robocentric occupancy and visibility grid maps together with proprioceptive states. This design forces the latent state to encode local geometry, visibility history, and ego-motion in a form that is directly relevant to collision avoidance. MAD is trained in DiffAero using a GPU-parallel map-construction module that provides high-throughput supervision for occupancy and visibility. The learned representation is used in three policy-learning modes: imagination-based MAD-Dreamer and feature-extractor variants based on PPO and SHAC. Across visual navigation and racing tasks, MAD-based agents achieve higher success rates, faster flight, and better cross-task transfer than corresponding vision-only baselines. The model also produces interpretable map predictions and accurate ego-motion estimates from depth observations. We further deploy the learned policy on a physical quadrotor with an Intel RealSense D435i and demonstrate safe indoor and outdoor flight under limited sensing, reaching 9.66 m/s in simulation and 5.05 m/s in real-world forest experiments. These results show that mapping-aware world models provide a practical middle ground between modular aerial navigation and end-to-end learning.

cs.RO