SearcharxivSearch

arXiv subjects

Jing Zhang

Publications and source records attributed to Jing Zhang.

At least 19 recordsLinked to original sources

A non-trivially non-trivial automorphism at an inaccessible cardinal

We show that, relative to the existence of an inaccessible cardinal, it is consistent that for an inaccessible cardinal $\kappa$, there is a non-trivial automorphism of $\mathcal{P}(\kappa)/\mathrm{fin}$ whose restriction to $\mathcal{P}(\beta)/\mathrm{fin}$ is induced by a bijection on $\beta$ for every $\beta<\kappa$. This answers questions posed by Larson--McKenney, Shelah--Stepr\=ans, and Veli\v{c}kovi\'{c}.

math.LO

A Graph Foundation Model for Large-Scale MIMO Detection

Large-scale multiple-input multiple-output (MIMO) detection is fundamental to modern wireless networks but constrained by performance-complexity trade-offs. Existing detectors, whether classical or learning-based, often fall short in either scalability or generalizability across heterogeneous scenarios. To overcome these limitations, we introduce a wireless-native graph foundation model (GFM) tailored for large-scale MIMO detection. The proposed GFM employs a physics-informed hybrid architecture, integrating the local correlation extraction of message passing neural networks with the global attention of graph Transformers, encoding the physical interference patterns from the expectation propagation algorithm. Via extensive pre-training, this synergy enables the learning of a general-purpose detection mapping scalable across antenna dimensions and channel conditions. For rapid downstream deployment, parameter-efficient fine-tuning is leveraged to adapt the GFM to specific non-ideal system regimes with minimal overhead. To enhance inference efficiency, a mixture-of-experts mechanism is embedded at downstream deployment to dynamically activate only the necessary sub-modules. Evaluations show that the proposed GFM consistently outperforms classical detectors and advanced data-driven baselines in accuracy, configuration generality, and cross-scenario transferability across various challenging zero-shot and few-shot conditions.

cs.IT

Measurement-Based Feedback of Open Quantum Systems: A Control-Theoretic Review and Tutorial

This review develops a control-theoretic perspective on measurement-based feedback for continuously monitored open quantum systems, with the main analysis focused on finite-dimensional systems governed by diffusive stochastic master equations. We introduce the relevant state-space, invariant-subspace, and quantum non-demolition structures, interpret quantum filtering as nonlinear observer dynamics, and review open-loop asymptotics, filter stability, state-feedback stabilization, robustness, and reduced-order observer-based control. Particular emphasis is placed on a recurrence--contraction framework, in which Hamiltonian feedback removes non-target invariant obstructions while measurement-induced dynamics provide local exponential contraction. Although the detailed analysis is developed for finite-dimensional diffusive models, the underlying measurement--estimation--feedback architecture is relevant across a broad range of quantum platforms. We further discuss implementation challenges and open problems involving scalable estimation, sampling and delay, adaptation, hybrid and non-Markovian dynamics, practical stability, optimal control, and learning-based design. By organizing these developments around invariance, estimation, recurrence, and contraction, the review provides a tutorial bridge between measurement-based quantum feedback and nonlinear stochastic control.

quant-ph

ISAC with Co-Prime Arrays: Virtual-Aperture Sensing and uplink downlink communications

Integrated sensing and communication (ISAC) enables simultaneous communication and environmental sensing in unmanned aerial vehicle (UAV) networks, but its performance is constrained by the physical antenna aperture and residual self-interference (SI) in full-duplex (FD) sensing. To address these issues, we propose a shared-aperture ISAC architecture in which a sparse co-prime array (CPA) is embedded in a uniform linear array (ULA) grid for FD sensing, while the remaining antenna positions support time-division duplexing (TDD) communication. We characterize the sensing performance through an order-wise Cramer-Rao bound (CRB) analysis, showing that the CPA achieves a stronger asymptotic sensing gain than the partitioned ULA benchmark in both single-target and nondegenerate multi-target scenarios. We further reveal a space-time sampling tradeoff under the same physical aperture. Based on the proposed architecture, we formulate a non-convex joint resource allocation problem that maximizes the weighted downlink-uplink sum rate by jointly designing the sensing transmit covariance, downlink precoder, and uplink receive beamformers under sensing accuracy, BS transmit-power, communication QoS, and residual SI constraints. An alternating-optimization-based algorithm is developed. Simulations demonstrate consistent performance gains over the considered baselines and confirm the complementary benefits of the CPA virtual aperture and sensing covariance optimization.

cs.IT

Magnetic-Field-Calibration-Free Determination of the Hyperfine Constant $A$ in Ultracold Fermi gases of $^{40}$K

Hyperfine constant $A$ is a key parameter of the hyperfine structure and underpins precision spectroscopy and metrology. In this Letter, we develop a magnetic-field-calibration-free method for determining the ground-state hyperfine constant $A$ in an ultracold $^{40}$K Fermi gas by utilizing a pair of magnetically insensitive ("clock") transitions. This overcomes the stringent magnetic-field calibration requirements of conventional methods. We measure the transition frequency between these two magnetically insensitive transitions with Hz-level resolution over a range of magnetic fields, and obtain the ground-state hyperfine constant $A = -h\times 285.730536(2)\,\mathrm{MHz}$, corresponding to an absolute uncertainty of about $2\,\mathrm{Hz}$. Our value reduces the uncertainty by nearly three orders of magnitude compared with previous determinations, providing a substantially improved reference for high-precision spectroscopy and metrology with $^{40}$K.

cond-mat.quant-gas

From Test Performance to Risk-Based Effect Sizes: A Unified Wald-Type Framework to Design Clinical Validation Studies for Binary and Survival Outcomes

Clinical validation studies of predictive tests are usually designed to focus on sensitivity ($Se$) and specificity ($Sp$), while statistical power is often calculated on regression-effect scales (e.g., risk ratio, hazard ratio). However, these quantities are statistically connected. Here, we provide closed-form links from sensitivity, specificity, and disease prevalence ($\pi$) to predictive risks, risk contrasts, and Wald-type variance, power, and sample-size formulas for binary and fixed-horizon survival outcomes. Analyses of statistical efficiency via C- and D-optimal principles demonstrate how prevalence and threshold choices affect study efficiency, supporting rapid decisions in preliminary studies and informing the design of subsequent, larger studies. Simulations show good calibration across most realistic scenarios; when events are rare and test effects are simultaneously very large, continuity and minimum-event corrections are needed to stabilize the approximation. We illustrate the framework with a case study describing use of the coronary artery calcium score for predicting incident cardiovascular disease in patients with type 2 diabetes mellitus. The formulas let investigators check power and required enrollment directly from $(Se,Sp,\pi)$, without running a separate simulation for each design candidate.

stat.ME

AI Control Scientist: LLM-driven Agentic System for Automated Control Design

Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual parameter tuning, resulting in limited efficiency and scalability. To this end, this paper proposes AI Control Scientist (AICS), the first large language model (LLM)-driven agent capable of automatically generating optimized controller from language design requirements. Specifically, a Task Modeling Agent interprets user requirements to engineering constraints; a Controller Design Agent generate candidate controller structures and executable code; and a Parameter Tuning Agent refine controller parameters under closed-loop performance criteria. Experiments demonstrate that the proposed agentic system can automatically generate multiple representative control systems, outperforms existing automated baselines in both design success rate and optimization efficiency. This work has the potential to transform control system design from human-driven to agent-driven, paving the way for model predictive control and other advanced control systems design.

cs.AI

A Geometry-Driven, Framework-Agnostic Optimization for Object Pose Estimation

Current object pose estimation research remains predominantly model-centric, focusing on architectural innovations and post-processing refinements. This paper introduces a data-centric optimization by proposing a novel, physically grounded rotation representation through principal axes alignment. Our method aligns the object's coordinate system with its inherent geometric axes, derived from inertial properties, yielding three key advantages: Inherent Stability-leveraging the energy-minimizing property of principal axes provides a robust representation that is less sensitive to noise and occlusions; Symmetry-Aware Canonicalization-explicitly resolving rotational ambiguities for symmetric objects at the data level, which fundamentally eliminates label confusion during network training; and Framework Agnosticism-the optimization is applied purely at the dataset level, ensuring plug-and-play compatibility with existing networks without any architectural modification. We validate the framework across diverse category-level and instance-level models. Extensive experiments demonstrate consistent and significant accuracy improvements, while preserving the integrity of the baseline network. This work establishes a new, geometry-driven direction for enhancing pose estimation, circumventing the need for complex network redesign.

cs.CV

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

Ancient Chinese artifact text recognition is fundamental to heritage digitization, and benchmarks for ancient texts are essential for evaluating current model capabilities. However, existing benchmarks suffer from ''fragmentation'', manifested in limited temporal coverage, limited medium diversity, and incomplete script types. Therefore, we present Ancient-Bench, a comprehensive benchmark of 2,700 images for ancient Chinese artifact text recognition, featuring three dimensions: Multi-millennial (spanning 3,000 years of character evolution), Multi-medium (covering nine artifact categories), and Multi-script (encompassing seven historical script forms). To enable consistent and fair evaluation across heterogeneous media, we further define three annotation standards tailored to the medium-specific characteristics of ancient texts: symbol standardization, character standardization, and parsing standardization. Extensive experiments on Ancient-Bench covering general Vision-Language Models (VLMs) and OCR-specialist models reveal that ancient Chinese artifact text recognition remains fundamentally unsolved, with persistent challenges in variant characters, specialized symbols, and hallucination. The dataset is available at https://github.com/SCUT-DLVCLab/Ancient_Bench.

cs.CV

DeepONet-LSTM Neural Operator for Output Feedback Control of Reaction Diffusion PDEs

This paper presents a neural operator-based approach for the output feedback boundary stabilization of reaction diffusion PDEs. The classical output feedback backstepping design requires solving control and observer kernel equations for each reaction coefficient. To avoid computing these kernel functions, the output feedback control law is reformulated as a causal boundary operator that maps the reaction coefficient and the boundary measurement to the boundary control input. A hybrid DeepONet-LSTM neural operator is proposed to approximate this causal operator, where DeepONet encodes the spatial coefficient and LSTM captures the temporal dependence of the measurement history. We analyze the Lipschitz continuity of the boundary operator and prove the closed-loop practical stability with the learned controller. A modified loss is also introduced to improve the temporal regularity of the learned boundary input. Numerical results illustrate that the proposed neural operator controller effectively stabilizes the system.

eess.SY

What Matters for Latent Actions in Robot Learning

Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite rapid progress, research on LAM remains highly fragmented, with existing methods evaluating different design choices in isolation under inconsistent experimental settings, making it difficult to identify the factors that truly determine downstream robotic manipulation performance. In this work, we present the first comprehensive empirical study of latent action learning for robotic manipulation. We unify representative LAM methods within a common autoencoding framework and systematically investigate 41 LAM design choices across three dimensions, including latent action modeling paradigms, learning objectives and regularization methods, and latent action integration strategies. We further examine four proxy metrics for evaluating latent action quality and assess their ability to reliably predict downstream robotic manipulation performance. Extensive experiments on three widely used benchmarks provide strong empirical evidence that fine-tuning vision-language model (VLM) backbones with latent actions provides a stronger initialization for downstream policy learning, with further validation on real-world robot manipulation tasks.

cs.RO

ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models

Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines. Existing SQL correction approaches either rely on large-scale, high-quality training data with substantial overhead, or adopt single-path agentic workflows that are brittle to early mistakes and prone to error propagation. To develop a practical SQL correctness system for industrial scenarios, we present a training-free framework that formulates SQL correction as a plan-guided, tree-structured debugging process. By maintaining multiple correction strategies and enabling backtracking, the framework mitigates error accumulation during iterative refinement. We further integrate execution-based verification and clause-level diagnostic tools to support strategy pruning and precise error localization. We evaluate the system on the BIRD-Critic benchmark and observe consistent accuracy gains over strong LLM backbones and representative agent-based baselines, achieving a 9.42% improvement over the previous state-of-the-art method. The framework is also deployed in the Torch Log Service (TLS) of Volcano Engine to support an online Text-to-TLS API. In production, it improves execution accuracy from 36.77% to 53.61% on real user queries with a representative strong LLM backbone (GPT-5). These results demonstrate the effectiveness and stability of our approach in real-world deployments.

cs.AI

MVFM-3DAD: Multi-view Flow Matching for 3D Anomaly Detection via Density Proxy Estimation

In 3D anomaly detection (3DAD), most existing methods rely on Memory bank retrieval or reconstruction. However, memory-based methods are constrained by the coverage of stored normal features, while reconstruction-based methods may learn identity shortcuts that also reconstruct anomalous inputs well. These limitations motivate a density-oriented approach that evaluates whether a test sample follows the learned normal distribution. To this end, we propose MVFM-3DAD, a flow-based framework that reframes 3DAD as density proxy estimation over the normal data distribution. MVFM-3DAD introduces a Bidirectional Geometric Projector (BGP), whose forward process converts irregular point clouds into structured multi-view representations. The Flow-guided Density Proxy Estimator (FDPE) estimates a reference density for each view feature, after which the backward process of BGP maps these multi-view density estimates to their corresponding 3D points. Building on it, anomalous features can be identified by their terminal normality. Unlike conventional flow-based likelihood estimation, our formulation requires neither input reconstruction nor explicit Jacobian evaluation, yielding a simple and efficient anomaly-scoring mechanism. Extensive experiments show that MVFM-3DAD outperforms the strongest competing methods on Real3D-AD and MVTec3D-AD. Code is available at https://github.com/lil-wayne-0319/MV3D-AD

cs.GR

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained $\pi_{0.5}$ instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.

cs.RO

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators. However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of systematic frameworks for quantifying alignment with human reviewers. To address this gap, we propose SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation. We collect and annotate 675 survey papers with 1,630 review reports. We structure authentic peer-review reports by converting free-form comments into four-dimensional scores (Readability, Criticalness, Comprehensiveness, Structure) paired with supporting rationales. We further release standardized train/test splits and an evaluation protocol to measure alignment between automatic evaluators and human reviewers. To validate the benchmark, we develop SurveyAlign, a strong baseline evaluator by fine-tuning Qwen3-32B with LoRA on our annotated data, augmented with external knowledge for knowledge-intensive dimensions. On the test set, SurveyAlign substantially improves reviewer alignment over prompt-based judging with GPT-5.2, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 across all four dimensions. Our contributions are twofold: (1) we establish the first multi-dimensional, reviewer-aligned dataset with a reproducible evaluation framework for survey reviewing; (2) we develop a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research. Our code and data are available at https://surveyreview.github.io

cs.CL

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models are predominantly trained on natural images, while existing medical localization resources remain fragmented across imaging modalities, datasets, and task formulations. To address this gap, we construct LocAnyMed-200K, a multimodal medical visual grounding dataset containing approximately 200K image-query-answer examples across computed tomography, optical medical imaging, ultrasound, and X-ray. We harmonize heterogeneous detection and localization resources into a unified free-form instruction format that supports one or multiple bounding boxes, point coordinates, and no-target outputs for negative queries. Full-parameter fine-tuning of LocateAnything-3B on LocAnyMed-200K improves F1@IoU 0.50 from 10.64 to 85.59 on a held-out evaluation split, demonstrating that large-scale domain-specific supervision can equip a general grounding model with effective medical localization capabilities. Beyond spatial coordinates, a clinically interpretable grounding system should also communicate the evidence supporting its prediction. We therefore derive LocAnyMed-CoT-20K, a rationale-augmented subset that connects anatomical context, visual observations, and spatial conclusions through structured reasoning and further improves cross-source generalization through fine-tuning. Together, these resources provide a unified foundation for studying both localization accuracy and rationale quality across heterogeneous medical imaging modalities. The code is publicly available at https://github.com/MiliLab/LocAnyMed.

cs.CV

IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks

Semantic segmentation models are vulnerable to transferable adversarial perturbations, yet evaluating transfer attacks on dense prediction models can be computationally expensive. Existing ensemble attacks often rely on multiple surrogate models, increasing the computation cost, even harder for segmentation. This paper studies an efficient single-source alternative for transferable attacks on semantic segmentation. We formulate transferable attack composition as a chained computation over differentiable attack components, allowing the expensive source-model gradient computation to be shared. To reduce the update instability introduced by chained composition, we further use an integrated-gradient-style path-averaged direction as an empirical stabilization heuristic. Experiments on Pascal VOC and Cityscapes evaluate the resulting transferability efficiency trade-off across CNN- and transformer-based segmentation models. IGME achieves competitive transferability compared with single-source baselines and favorable runtime compared with model-ensemble attacks, while requiring access to only one source model.

cs.CV

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts. A natural solution is to equip MLLMs with zoom-in tools for active local inspection. However, through a pilot study on XLRS-Bench, we find that zoom-in is only partially effective: it resolves easy and medium-level tasks with locally recoverable evidence, but saturates on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning. Motivated by this finding, we move beyond single-tool zoom-in and introduce GeoMTVR, a large-scale Geospatial Multi-Tool Visual Reasoning dataset built from wide-area satellite imagery. GeoMTVR contains 13K UHR VQA samples with interleaved reasoning trajectories, diverse visual tool calls, and returned visual observations, enabling models to learn question decomposition, tool selection, regional inspection, object-level grounding, auxiliary visual reasoning, and cross-tool evidence integration. Beyond supervised fine-tuning, we propose a tool-attention-focused reinforcement learning algorithm that concentrates optimization on critical tool-use decisions, including when to invoke tools, which tool to select, where to apply it, and how to interpret tool outputs. By combining SFT on GeoMTVR with our RL algorithm, we develop GeoLens, a multi-tool visual reasoning MLLM for UHR RS. Experiments show that GeoLens consistently outperforms direct reasoning and single-tool zoom-in baselines, achieving stronger accuracy, better evidence grounding, and more efficient tool-use trajectories.

cs.CV