SearcharxivSearch

arXiv subjects

Feng Zhang

Publications and source records attributed to Feng Zhang.

At least 19 recordsLinked to original sources

Semi-Tensor Product-Based Multi-Term Randomized T-SVD and Its Visual Applications

Tensor singular value decomposition (T-SVD), which is built upon the tensor-tensor product (t-product), has emerged as a powerful tool for processing high-dimensional visual data such as color images and videos. However, the standard t-product imposes strict dimensional compatibility constraints. Although extensions based on the semi-tensor product (STP) relax this restriction, their single-term formulations still suffer from limited approximation accuracy. Moreover, these deterministic methods incur high computational costs when processing large-scale tensor data. To address these issues, this paper introduces a novel semi-tensor product for third-order tensors under the t-product framework induced by arbitrary invertible linear transforms. The resulting tensor semi-tensor product breaks the rigid dimension matching requirement of the standard t-product, while retaining the closed-form property of T-SVD. Based on this construction, we develop a multi-term semi-tensor product singular value decomposition (MSTP-SVD), which integrates multiple orthogonal decomposition terms to significantly improve low-rank approximation accuracy compared with single-term schemes. To reduce the computational cost of multi-term modeling, we incorporate randomized projection and power iteration techniques into the MSTP-SVD framework, yielding an accelerated multi-term randomized semi-tensor product SVD (MRSTP-SVD) algorithm that achieves a balance between reconstruction accuracy and computational efficiency. Experiments on image and video compression and completion tasks demonstrate the effectiveness of the proposed method.

cs.LG

Two-parameter variational estimates for averages over tori

One-parameter variational inequalities are well developed, whereas their multi-parameter counterparts remain much less understood. We explore two-parameter variational inequalities for averages over tori in $\mathbb{R}^3$. To capture the underlying two-parameter structure, we introduce a local two-parameter $r$-variation norm that combines rectangular increments with variations along the boundary. The resulting variation operator pointwise dominates the corresponding two-parameter local maximal function and, unlike the maximal function, also captures oscillation across the two parameters. We establish sharp $L^p$--$L^q$ bounds for this variation operator up to endpoints. For comparison, we also obtain sharp $L^p$ bounds up to endpoints for the corresponding local one-parameter variation operator, revealing a genuine difference between the one- and two-parameter boundedness regions. The proof combines square function estimates for two-parameter propagators with local smoothing estimates through mixed-norm interpolation.

math.CA

Minute-Scale High-Fidelity Gyrokinetic Simulations with Portability from Laptop to Supercomputer

Global gyrokinetic particle simulations remain computationally expensive, as they demand both adequate marker statistics and three-dimensional field solvers. In this work, we present a hybrid spectral method within the particle-in-Fourier (PIF) framework and implement it in the electrostatic model of GTC. Charge scatter and field gather are performed between particles and fields on a two-dimensional poloidal mesh, while the corresponding Poisson solver is discretized using radial finite differences and poloidal $m$-harmonics. Truncated spectral transforms are employed to connect multiple representations for fields, avoiding costly particle-grid operations for each individual $m$-harmonic within the particle loop. Benchmarks against conventional particle-in-cell (PIC) simulations successfully reproduce single-$n$ ion temperature gradient (ITG) mode structures and dispersion relations, as well as multi-$n$ nonlinear ITG transport and its regulation by zonal flows. Compared to conventional PIC, the proposed method reduces the effective problem size by more than a factor of 48 and achieves a speedup of over two orders of magnitude for single-$n$ cases. A 2000-step single-$n$ simulation with approximately 2 million markers completes in 78.2 seconds on a laptop GPU, while multi-$n$ turbulence simulation also completes within minutes. Furthermore, the elimination of toroidal particle-shift communication yields promising preliminary scaling performance on multiple NVIDIA A100 GPUs. The numerical scheme is broadly applicable for accelerating particle simulations on platforms ranging from laptops to supercomputers.

physics.plasm-ph

Environment Evolution for Terminal Agents

Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits generalization and the continuous provision of learning signals as the model becomes stronger. In this paper, we propose environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals. We derive three evolution directions that influence environment difficulty from the multi-turn learning objective and then implement evolution along these directions through a loop-engineered multi-agent harness. Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments. We validate its effectiveness on Qwen3.6-27B and Qwen3.6-35B-A3B through simple long-horizon RL training, improving their performance by 14.4 and 18.0 percentage points on Terminal-Bench 2.1, respectively.

cs.AI

Emergent Noncollinearity and Near-Degenerate Magnetic Superlattices in AT6X6 Kagome Metals

Ferromagnetic AT6X6 Kagome compounds are a popular class of systems in which quantum magnetism with topological features has been observed. These systems allow easy chemical substitution, creating an opportunity to fine-tune their properties. In this paper, we present electronic-structure and magnetic ground-state studies of several AT6X6 compounds with relatively low magnetic-ground-state stability. We find unusual magnetic orderings, including complex spin-spiral states and the formation of magnetic long-range superstructures. While LiFe6Ga6 and TiMn6Ge6 retain collinear AFM ground states with low-energy FM/AFM layer sequences, competing spin-spiral and long-period antiferromagnetic structures in MgFe6Ga6 and a double-spin-spiral ground state in TiFe6Ga6 were determined. Magnetism in all these systems appears local, with adiabatic energy profiles suggesting non-Heisenberg long-range interactions, including a strong biquadratic term. In TiMn6Ge6, we found the conditions for magnetic tunneling. Our results show that, in addition to traditional magnetic topological features in such FM Kagome systems, near-degenerate magnetic superstructures suitable for spintronic switching applications can form naturally. Overall, these systems represent a potentially rich playground for neutron diffraction and spintronics experimental studies.

cond-mat.mtrl-sci

Quantifying geographic domain shift to decouple the geospatial transferability of human mobility flow generation models

Human mobility serves as an essential proxy for understanding social, economic, and environmental dynamics in urban systems. Geospatial transferability, which measures a model's capability in a new location or unseen region, is a critical dimension for comparing different human mobility generation models. However, few studies have studied the intrinsic characteristics of geospatial transferability. To this end, this study systematically investigates the geospatial transferability of four representative human mobility generation models using a large-scale benchmark dataset of census tract level commuting flows across 2265 counties in the United States. Inspired by the domain adaptation theory in machine learning, we introduce geographic domain shift to describe the intrinsic differences in geographic feature distributions and spatial structures between source and target regions, which may jointly affect model transferability. Moreover, we propose two metrics, mutual information and spatial shift, to quantify the geographic domain shift. To examine their associations with model transferability, we employ linear mixed-effects regression to analyze the associations between geographic domain shifts and transferability. Our results reveal substantial spatial heterogeneity and asymmetry in transfer performance across regions. Both information shift and spatial shift exhibit statistically significant and complementary explanatory power. This indicates that geospatial transferability depends not only on model design but also on intrinsic geographic differences. These findings provide a novel methodological framework for evaluating and improving the geospatial transferability of human mobility generation models and support more robust and fair human mobility data synthesis across diverse regions. It also offers insights on spatial transferability for GeoAI model development.

cs.AI

NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting

With the rapid development and adoption of 3D Gaussian Splatting (3DGS), the need for effective copyright protection has become increasingly critical. Existing watermarking techniques for 3DGS mainly focus on protecting rendered images via pre-trained decoders, leaving the underlying 3D Gaussian primitives vulnerable to misuse. In particular, they are ineffective against Partial Infringement, where an adversary extracts and reuses only a subset of Gaussians. In this paper, we propose NGS-Marker, a novel native watermarking framework for 3DGS. It integrates a jointly trained watermark injector and message decoder, and employs a gradientbased progressive injection strategy to ensure full-scene coverage. This enables robust ownership decoding from any local region. We further extend NGS-Marker with hybrid protection (combining native and indirect watermarks) and support for multimodal watermarking. Extensive experiments demonstrate that NGS-Marker effectively defends against partial infringement while offering practical flexibility for real-world deployment.

cs.CV

DiffImaginE: Imagine to Verify Entity Types with Diffusion

Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.

cs.AI

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging. At corpus scale, suffix-array-based methods are commonly applied independently within shards, leaving cross-shard duplicates undetected and making the resulting retention behavior sensitive to the sharding configuration. Hash-based methods enable global exact duplicate counting, but often rely on fixed copy-retention policies that cannot accommodate heterogeneous repetition patterns. We propose a scalable subdocument deduplication framework that decouples duplicate detection from copy retention. It identifies duplicate groups through natural-boundary segmentation, normalized exact hashing, and distributed aggregation, and then applies an explicit frequency- and length-aware retention policy that allocates an adaptive copy budget to each group, retaining more copies of low-frequency or short repetitions while more aggressively deleting high-frequency or long ones. Experiments on FineWeb-Edu and a code-containing web corpus show that models trained on data processed by our method achieve the best overall performance among the evaluated settings. These results underscore the importance of explicit copy-retention control.

cs.CL

Deep Research Pretraining via Predictive Navigation

Deep research agents are often trained on expensive, environment-grounded tool-use trajectories that require repeated retrieval, document inspection, and report evaluation. We introduce Deep Research Pretraining (DRP), an offline framework that derives predictive navigation supervision from naturally occurring evidence structures. Given a citation-bearing or hyperlinked passage, DRP constructs a proxy research objective, recovers linked evidence and graph-related alternatives, and converts them into search-open-write trajectories. This teaches models what to search for, which documents to inspect, and how to synthesize evidence, without a live retrieval environment or executed policy rollout. We instantiate DRP on scholarly citation graphs (DRP-Paper) and Wikipedia hyperlinks (DRP-Web), continually pretrain separate Qwen3-14B-Base models on 1B tokens, and fine-tune them on controlled fractions of 13K agent trajectories. Across five independently sampled subsets at each low-data budget, both variants consistently outperform matched no-DRP models on DeepResearch Bench. With one quarter of the SFT data, DRP-Web even surpasses a fixed no-DRP full-data checkpoint, with gains transferring to ResearchQA, WebWalkerQA, and SimpleQA. Starting from matched low-data SFT checkpoints, the DRP-Web advantage also persists through subsequent agentic RL. Source-matched and evidence-mismatch controls indicate that these improvements arise from evidence-conditioned navigation rather than domain exposure or agent-format imitation. DRP thus provides a promising complementary approach to trajectory-based agent training.

cs.CL

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards

Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it remains hard to optimize in the high-accuracy regime because such rewards become weakly discriminative: near-correct outputs receive very similar scores, providing limited learning signal for hard cases. We propose Step-Aware Annealing (SAA), a plug-and-play reward sharpening mechanism that progressively increases reward curvature during training, amplifying subtle quality differences among high-scoring samples while preserving stability in early learning. Built on SAA, we introduce DocPO, a document policy optimization framework with element-specific, reference-based rewards anchored by edit-distance signals: normalized string edit distance (NED) for text, tree edit distance similarity (TEDS) for tables, and a hybrid Rubric+edit reward for formulas. Experiments on OmniDocBench and DocElemHard show that SAA consistently improves GRPO-style RL across document elements over non-annealed rewards, without requiring additional human supervision for reward construction.

cs.CV

Alchemical thermodynamic integration for ab initio free-energy calculations in solutions

We develop an alchemical thermodynamic integration scheme that couples ab initio force calculators on the fly during Monte Carlo and molecular dynamics simulations. The implementation is validated against existing hybrid-Hamiltonian approaches. The scheme yields ab initio free energies of high-pressure Fe-Ni and ambient Li-Na liquid solutions that agree with previous calculations and reproduce the experimentally observed Li-Na miscibility gap. The code has an efficiency comparable to standard ab initio molecular dynamics. These results establish this scheme as a practical alchemical-integration framework for first-principles free-energy calculations in solutions.

cond-mat.mtrl-sci

exa-PD: A scalable high-performance workflow for multi-element phase diagram construction

Exa-PD is a highly parallelizable workflow designed for the construction of multi-element phase diagrams (PDs). It uses standard sampling techniques, molecular dynamics (MD) and Monte Carlo (MC) as implemented in the LAMMPS package, to simultaneously sample multiple phases over a fine temperature-composition mesh for free-energy calculations. Parsl serves as the global workflow engine, coordinating large ensembles of MD and MC tasks to achieve massive parallelization with strong scalability. The resulting free energies of liquid and solid phases are then fed to CALPHAD modeling via the PyCalphad package to construct multi-element PDs.

cond-mat.mtrl-sci

VSC: A Zero-Dimensional Fusion Design Platform for Multiple Magnetic Configurations

The VeloAlpha System Code (VSC) is a computational framework for zero-dimensional fusion power-balance studies across five magnetic-confinement configurations: tokamaks, magnetic mirrors, field-reversed configurations (FRCs), dipoles, and stellarators. A common power-balance formulation connects fusion production, charged-particle deposition, radiation, transport loss, external heating, and fusion gain, while each configuration retains its own geometry, profile weights, confinement model, and operating constraints. The same solver interface supports both single-point calculations and two-dimensional plasma operating contour (POPCON) scans, producing fusion and heating powers, gain, radiation and transport losses, geometry quantities, and configuration-specific validity indicators. VSC therefore makes it possible to study how assumptions about density, temperature, magnetic field, confinement, and geometry shape the accessible operating space of different fusion concepts within one traceable framework. By combining reduced-order physics models with a unified computational platform, VSC enables rapid assessment and comparative analysis of candidate fusion reactor concepts during the early design stage.

physics.plasm-ph

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms to incorporate constraint information during the generation process. While existing approaches attempt to mitigate this by relying on external tools or task decomposition, they fail to enhance the model's intrinsic constraint awareness. To address this, we propose Constraint-Aware Reinforcement Learning (CARL), a novel RL framework designed to strengthen LLMs' intrinsic focus on constraints. CARL introduces a constraint-aware reward by comparing the model's output distributions under constrained and unconstrained inputs, encouraging constraint focus and penalizing neglect. Compatible with various RL frameworks and requiring no external solvers or top models, CARL enables scalable, end-to-end constraint-aware planning. Extensive experiments on BlocksWorld, TravelPlanner, and T-Eval demonstrate that CARL significantly outperforms standard Reinforcement Fine-Tuning (RFT) baselines and state-of-the-art reasoning models, exhibiting a markedly increased focus on constraints.

cs.AI

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to trajectory neglect, in which agents lose focus on the task goal and interaction history at intermediate steps. Prior work has explored step-level supervision using Shannon-entropy-based uncertainty signals, which conflate inherent state complexity with agent confidence and therefore provide unreliable estimates of decision reliability. To address this issue, we propose normalized entropy, which measures confidence deviations relative to an agent's average behavior under a given state, thereby strengthening the association between low-quality actions and trajectory neglect. Building on this insight, we introduce Selective Trajectory-Aware Policy Optimization (STAPO), a hierarchical group-based RL framework. STAPO leverages normalized entropy to locate outlier steps associated with trajectory neglect and optimizes them via a joint mechanism of trajectory-aware reward and trajectory-independent penalty, enhancing trajectory awareness while preserving training stability. Extensive experiments on ALFWorld, WebShop, and Search-Augmented QA demonstrate that STAPO achieves state-of-the-art performance while substantially alleviating trajectory neglect, validating its effectiveness and robustness for agentic tasks.

cs.AI

MORE: A Multilingual Document Parsing Benchmark and Evaluation

Multilingual documents encapsulate rich regional cultures, scientific discoveries, and historical records. Parsing this content into structured, machine-readable formats is critical for unlocking global knowledge. However, existing benchmarks predominantly focus on high-resource languages like English and Chinese, creating an evaluation blind spot concerning model performance on other languages. While recent Vision-Language Models (VLMs) claim support for hundreds of languages, the lack of ground truth makes it impossible to empirically verify these capabilities. To bridge this gap, we introduce MORE, a large-scale benchmark designed for multilingual document parsing evaluation. MORE distinguishes itself through three key dimensions: (1) Unprecedented Scale: It covers 149 languages, making it the most linguistically diverse benchmark to date; (2) Structural Complexity: Unlike previous works, it extends evaluation beyond plain text to include structural elements such as code blocks, tables, and catalogs; and (3) Data Authenticity: All samples are curated from real-world documents via a model-assisted, human-refined annotation pipeline. We evaluate state-of-the-art models using MORE, establishing new performance baselines for long-tail languages and validating the benchmark's effectiveness in diagnosing model capabilities in realistic, diverse scenarios. The MORE dataset will be available at https://github.com/zimoqingfeng/MORE.

cs.CV

Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design

We present Home3D 1.0, a modular image-to-3D generation system that produces high-quality 3D assets from a single reference image, targeting interior design and e-commerce applications. Given a photograph of a furniture or decor item, the system outputs a mesh with physically-based rendering (PBR) materials, and the mesh can be decomposed into material-specific components. The pipeline is organized into four tightly coupled modules: Geometry reconstructs a watertight mesh through latent SDF modelling with a geometry VAE and a coarse-to-fine flow-matching DiT; Texture predicts multiview albedo observations, reprojects them onto the mesh, and completes unseen surface regions with a 3D texture field; Material uses MatWeaver to obtain component masks through video-based segmentation and UV-space voting, then retrieves and bakes PBR maps from a curated material library through hierarchical multi-modal matching; and Parts generates material-editable semantic part meshes with a PartVAE and PartDiT, decoding multi-head part-specific SDF fields in one pass. Each module is evaluated independently with dedicated metrics, highlighting both the current system capability and the remaining gaps toward broader deployment.

cs.CV