Searcharxiv⌕ Search

arXiv · 2610.07749

Cost-Effective Numerical QA over Semi-Structured Table: Structuring, Resolution, Planning

Abstract

Semi-structured tables encode rich semantic information through diverse layout elements, such as hierarchical row and column headers. Answering numerical questions over such data is challenging because it requires a joint effort of accurate structural understanding of tables, question uncertainty resolution, and complex query intents interpretation. Existing methods are rarely effective in handling the above multifaceted challenges in a holistic manner; moreover, they rely on LLMs without considering financial implications. We propose SemiBonsai, a cost-effective, LLM-powered framework featuring: (1) a Multiway Layered Table Structurer that converts a table into a multiway layered tree organized by structural elements; (2) a Context-Aware Uncertainty Resolver that grounds underspecified phrases to context-coherent structural elements; (3) a Plan-Guided Reasoner that decomposes complex query intents into a query plan for reasoning; (4) a Budget-Aware Bandit Router that learns per-instance LLM utilities and optimizes LLM selection under fixed monetary constraints. Experiments on three benchmarks show that SemiBonsai improves both QA effectiveness and cost-effectiveness. Notably, the Table Structurer improves answer accuracy by 46% over alternative tree construction methods, and Router achieves a 5% average relative gain in answer accuracy compared with existing routing methods across budgets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Feng Luo, Hui Luo, Zhifeng Bao, J. Shane Culpepper, Xiaoli Wang, Shazia Sadiq. 2026-10-06. Cost-Effective Numerical QA over Semi-Structured Table: Structuring, Resolution, Planning. https://arxiv.org/abs/2610.07749

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Crane: An Accurate and Scalable Neural Sketch for Graph Stream Summarization

Graph streams are rapidly evolving sequences of edges that convey continuously changing relationships among entities, playing a crucial role in domains such as networking, finance, and cybersecurity. Their massive scale and high dynamism make obtaining accurate statistics challenging with limited memory constraints. Traditional methods summarize graph streams through hand-crafted sketches, while recent studies have begun to replace these sketches with neural counterparts to improve adaptability and accuracy. However, this shift faces a major challenge: under limited memory, dominant frequent items tend to overshadow rare ones, hindering the neural network's ability to recover accurate statistics. To address this, we propose Crane, a hierarchical neural sketch architecture for graph stream summarization. Crane uses a hierarchical carry mechanism that automatically elevates frequent items to higher memory layers, reducing interference between frequent and infrequent items within the same layer. To better accommodate real-world deployment, Crane further adopts an adaptive memory expansion strategy that dynamically adds new layers once the occupancy of the top layer exceeds a threshold, enabling scalability across diverse data magnitudes. Experimental results show that Crane significantly reduces estimation error compared to state-of-the-art (SOTA) methods, while also offering competitive throughput.

cs.DB↗

Hierarchical Decomposition of Separable Workflow-Nets

The Partially Ordered Workflow Language (POWL) has recently emerged as a process modeling notation, offering strong quality guarantees and high expressiveness. While early versions of POWL relied on strict block-structured operators for choices and loops, the language has recently evolved into POWL 2.0, introducing choice graphs to enable the modeling of non-block-structured decisions and cycles. To bridge the gap between the theoretical advantages of POWL and the practical need for compatibility with established notations, robust model transformations are required. This paper presents a novel algorithm for transforming safe and sound workflow nets (WF-nets) into equivalent POWL 2.0 models. The algorithm recursively identifies structural patterns within the WF-net and translates them into their POWL representation. Unlike the previous approach that required separate detection strategies for exclusive choices and loops, our new algorithm utilizes choice graphs to capture generalized decision and cyclic patterns. We formally prove the correctness of our approach, showing that the generated POWL model preserves the language of the input WF-net. Furthermore, we prove the completeness of our algorithm on the class of separable WF-nets, which corresponds to nets constructed via the hierarchical nesting of state machines and marked graphs. We evaluate our algorithm on large-scale process models to demonstrate its high scalability. Furthermore, to test its practical expressiveness, we applied it to a benchmark of 1,493 industrial and synthetic process models. Our algorithm successfully transformed all models in this benchmark, suggesting that POWL 2.0's expressive power is generally sufficient to capture the complex logic found in real-world business processes. This work paves the way for broader adoption of POWL in practical process analysis and improvement applications.

cs.DB↗

Disk-Resident Graph ANN Search: An Experimental Evaluation

As data volumes grow while memory capacity remains limited, disk-resident graph-based approximate nearest neighbor (ANN) methods have become a practical alternative to memory-resident designs, shifting the bottleneck from computation to disk I/O. However, since their technical designs diverge widely across storage, layout, and execution paradigms, a systematic understanding of their fundamental performance trade-offs remains elusive. This paper presents a comprehensive experimental study of disk-resident graph-based ANN methods. First, we decompose such systems into five key technical components, i.e., storage strategy, disk layout, cache management, query execution, and update mechanism, and build a unified taxonomy of existing designs across these components. Second, we conduct fine-grained evaluations of representative strategies for each technical component to analyze the trade-offs in throughput, recall, and resource utilization. Third, we perform comprehensive end-to-end experiments and parameter-sensitivity analyses to evaluate overall system performance under diverse configurations. Fourth, our study reveals several non-obvious findings: (1) vector dimensionality fundamentally reshapes component effectiveness, necessitating dimension-aware design; (2) existing layout strategies exhibit surprisingly low I/O utilization (less than or equal to 15%); (3) page size critically affects feasibility and efficiency, with smaller pages preferred when layouts are carefully optimized; and (4) update strategies present clear workload-dependent trade-offs between in-place and out-of-place designs. Based on these findings, we derive practical guidelines for system design and configuration, and outline promising directions for future research.

cs.DB↗