Searcharxiv⌕ Search

arXiv · 2609.34424

Construction-Reuse Trade-offs for Exact Certificates in Fixed-Rank Threshold Screening

Abstract

Repeated threshold queries may reuse selected identities without reusing stale reports, but cheaper certificates need not shorten the complete response. We study selected-lower, atomic-upper (SLA) certificates for fixed-rank conjunctive screening with explicit missing-information semantics. An endpoint characterization and a counterexample separate same-source containment from policy-dependent online behavior. The original 320-session experiment reduces summed construction medians by 31.18% against an exclusion-cover certificate, yet its Cover/SLA full-API geometric time ratio is 0.9682 (95% conditional blocked interval 0.9593-0.9769), and SLA takes 12.94% more summed time than uncached Bitmap. Three separately launched complete repeats preserve this adverse ordering, with Cover/SLA ratios of 0.9665-0.9718. An additional 720-configuration exploration varies catalogue size, construction period, requested count and query locality on empirically resampled tables. SLA is faster in 295 configurations against Cover and 271 against Bitmap, descriptive counts that do not establish universal superiority. Separately instrumented additive costs distinguish construction savings from retrieval and report costs. A plane-stress component case adds an independent analytical displacement check and 4,608 boundary-challenging queries: five implementations agree exactly, while medium- and fine-mesh selections differ at 459 positions. A state-stratified public bolt-record exercise preserves 185 incomplete positions among 1,479 requests. Raw timings, complete configuration results and a tested clean-environment package support reproducibility. The SLA construction was explored and refined through the self-evolving AI system ZiYor; the named authors specified, implemented and evaluated it. This is a bounded mechanics-to-query study, not physical joint qualification or universal speedup.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zisu Li, Jingyang Du. 2026-09-28. Construction-Reuse Trade-offs for Exact Certificates in Fixed-Rank Threshold Screening. https://arxiv.org/abs/2609.34424

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

ScentGen: Hierarchical Multimodal Olfactory Semantic Modeling for Molecular Odor Description Generation

In this paper, we introduce a molecular odor description generation task, which aims to generate natural language odor descriptions from molecular structures. Unlike conventional methods that describe molecular odor using discrete labels, this task generates expressive and human-interpretable sensory descriptions. To address this task, we propose a hierarchical multimodal olfactory semantic modeling framework, named ScentGen. ScentGen consists of three key components: an odor semantic planner, a semantic adapter, and a description generator. The odor semantic planner integrates complementary molecular information from 1D SMILES sequences, 2D molecular graphs, and 3D molecular conformations to learn discriminative and structured olfactory semantics. The semantic adapter further maps the learned olfactory representation into the hidden space of a large language model, transforming molecular odor semantics into language-compatible continuous prompts. Conditioned on these prompts, the description generator produces coherent odor descriptions that reflect plausible sensory characteristics of the input molecule. Considering the lack of molecular datasets with natural language odor descriptions, we further construct a molecular odor description dataset containing paired multimodal molecular representations and human-interpretable odor descriptions. Extensive experiments demonstrate that ScentGen generates coherent and expressive odor descriptions, providing a more flexible solution for molecular odor understanding beyond discrete odor label prediction.

cs.CE↗

SymbolicLM: Training Language Models as Symbolic Regressors

Large Language Models (LLMs) have shown promising capabilities in scientific reasoning, yet scientific discovery ultimately requires deriving precise laws directly from observational data, known as Symbolic Regression (SR). This poses a challenge for LLMs due to the gap between probabilistic text generation and the exact structural requirements of SR. Existing approaches rely on complex external scaffolds, which are computationally expensive and separate symbolic reasoning from the model itself. To address this limitation, we propose to directly equip LLMs with symbolic regression capabilities through dedicated numerical-symbolic and physical supervision. We introduce PhysSymbArena, a large-scale benchmark containing over 160,000 equations and 1.8B tokens of numerical-symbolic data with physical descriptions, enabling systematic training and evaluation. Based on PhysSymbArena, we develop SymbolicLM, which enhances the symbolic regression ability of LLMs through mathematical and physical supervision. During inference, we further introduce SymbolicSGA, a refinement framework that leverages quantitative feedback to iteratively improve generated equations. Experiments on multiple symbolic regression benchmarks show that SymbolicLM substantially improves structural recovery while maintaining competitive numerical fitting performance. These results demonstrate that symbolic regression can be explicitly learned as an intrinsic capability of LLMs.

cs.CE↗

Decompose Dynamics Before Learning Dependencies in Spatiotemporal Systems

Relations in networked spatiotemporal systems are often learned from observations that entangle dynamics governed by different mechanisms, obscuring what evolves locally and how it propagates across nodes. We introduce Component-Aware Network Dynamics with Ordered Relations (CANDOR), which decomposes local dynamics before learning their dependencies. CANDOR represents each trajectory through a persistent background, gradual accumulation and release, and sparse shocks. Conditioned on these components, a delay-aware physical branch models edge and sample-dependent propagation over directed topology, while a topology-unconstrained functional branch discovers latent dependencies from background dynamics. Context-adaptive fusion combines functional, forward-propagation, and reverse-support forecasts, with training objectives encouraging specialized and semantically consistent representations. Experiments on two traffic benchmarks and three long-horizon water-quality datasets span two distinct spatiotemporal systems: human-driven urban traffic and naturally evolving river water quality. CANDOR consistently outperforms the strongest baselines, reducing MAE by up to 4.31% for traffic and MSE by up to 5.47% for water-quality forecasting. These results establish decomposition before dependency learning as an effective principle for spatiotemporal representation learning.

cs.CE↗