SearcharxivSearch

arXiv subjects

Renhao Xue

Publications and source records attributed to Renhao Xue.

7 recordsLinked to original sources

Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction

Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Existing document benchmarks are dominated by clean, high-quality samples, leaving low accuracy regions too sparse for calibration assessment. We introduce ConfBench, the first calibration-specific benchmark for key information extraction (KIE), built by applying 20 controlled degradation pipelines to a diverse document set, yielding 1,346 variants and 70K+ entity-level evaluations spanning the full accuracy spectrum. We evaluate four proprietary and three open-weight VLMs under verbalized and log-probability confidence estimation methods across three input modalities, and find: (i) OCR+Image modality results in more accurate confidence estimates; (ii) model capability is the dominant factor: within the Claude family confidence quality scales monotonically with capability, while across families parameter count is a poor predictor; (iii) calibration quality varies widely across models, from near-perfect to severely overconfident, and per-model post-hoc correction rescales these absolute confidence values for threshold-based routing without altering ranking-based operational metrics; and (iv) log-probability with first-token aggregation consistently outperforms mean-token and margin aggregations. We also introduce ECARB, a review-budget metric translating discriminative gains into operational savings. We release ConfBench publicly to enable systematic study of confidence estimators and calibration methods for trustworthy IDP application deployment.

cs.AI

From Errors to Rules: Iterative Prompt Optimization for Text Classification

Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-driven diagnosis, each with known but incompletely characterized strengths and limitations. We conduct a comprehensive empirical study across diverse classification benchmarks (2 to 150 classes) comparing these paradigms through both quantitative evaluation and qualitative analysis of optimization traces, revealing that each paradigm excels on structurally different task types and that no single method dominates. Guided by these insights, we propose Error-Guided Optimization (ERGO), an error-driven method that iterates over the full training set in non-overlapping batches, diagnoses classification failures, and generates targeted decision rules through a diagnose-prescribe-rewrite feedback loop. ERGO achieves the best accuracy on tasks where errors concentrate in specific confused label pairs (which we term boundary-learnable tasks): TREC: 90.0%, CLINC150: 94.4%, converges in 3-5 iterations, and produces interpretable decision rules. While ERGO does not achieve the highest overall average, it fills a complementary role: demonstration-based ICL wins on coverage-dependent tasks, exploration-based search wins on many-class intent, and ERGO wins where decision boundaries are learnable from error patterns. We provide a complementarity framework linking task characteristics to optimal paradigm selection, offering practical guidance for practitioners.

cs.AI

PRISM: Position-encoded Regressive Inverse Spectral Model for Multilayer Thin-Film Design

The inverse problem of multilayer thin-film optical coatings design represents a complex combinatorial-continuous optimization challenge. We present PRISM (Position-encoded Regressive Inverse Spectral Model), a unified decoder-only autoregressive transformer that streamlines this process by jointly predicting discrete material selection and continuous thickness regression within a single backbone. PRISM introduces two primary architectural innovations: (1) spectrum prefix conditioning, which utilizes standard prefix tokens for in-context target injection, and (2) cumulative-depth Rotary Position Embeddings, which encode continuous thickness directly into the positional representation to preserve the physical spatial relationships of the stack. Our benchmarks demonstrate that a PRISM-13M model reduces MAE by over 50\% compared to other transformer baselines while utilizing only one-fifth of the parameters. Furthermore, a 44M-parameter variant achieves state-of-the-art performance (MAE = 0.010) on our in-distribution validation benchmark and operates significantly faster than simulated annealing, offering a highly efficient alternative to classical optimization methods.

cs.LG

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a shared dense computation path despite substantial heterogeneity in temporal structure. Mixture-of-Experts (MoE) offers a natural alternative by enabling conditional computation, but standard MoE routing leaves expert specialization weakly identified and often unstable during downstream adaptation. We propose AME-TS, a structure-guided sparse time series foundation model that aligns expert routing with interpretable temporal structure. AME-TS first uses a lightweight regime predictor to estimate series-level descriptors, including forecastability, seasonality, trend, and sparsity, and maps them to a soft structural prior over experts. This series-level prior guides token-level routing during training, encouraging structure-aligned specialization. On the GIFT-Eval benchmark, AME-TS delivers a strong accuracy-efficiency tradeoff across model scales: it substantially outperforms existing time series foundation models at small model scales and remains competitive with the strongest models at larger scales, while activating substantially fewer parameters through sparse routing. We further show that AME-TS learns more interpretable routing geometry and substantially more stable expert specialization than standard MoE during fine-tuning on the M5 dataset. These results suggest that structure-aware routing is an effective and reliable way to realize the benefits of sparse expert models for time series forecasting.

cs.LG

Fully Atomic-Layer-Deposited Vertical Complementary FeRAM with Ultra-High 2Pr > 100 uC/cm2 and High Endurance > 1E10 cycles

A limited remanent polarization (Pr) in HfO2-based FeRAM remains a key obstacle to density scaling and reliability, while material and process optimizations offer only incremental improvements. This limitation fundamentally originates from the thickness-constrained switchable polarization and the intrinsic polarization ceiling of HfO2-based ferroelectrics. Here, we propose an all-ALD-grown vertical complementary FeRAM (VCF) architecture, in which the top and bottom stacked FeRAM cells maintain complementary polarization. This complementary dipole configuration converts the readout from a single-layer polarization response into a differential polarization summation, thereby amplifying the effective charge window without increasing the switching field of each individual layer or incurring area overhead. Viewed from top to bottom, an "up-down" polarization pair stores logic '1', whereas a "down-up" pair stores logic '0'. Using a complementary polarization write-read scheme, the VCF achieves an effective differential polarization above 100 uC/cm^2 and retains above 90 uC/cm^2 after 1e10 switching cycles without electrical breakdown. Robust retention (longer than 1e4 s at 85 degC) and strong disturb immunity are demonstrated, with an effective differential polarization above 80 uC/cm^2 under a V/3 scheme after 1e6 disturb pulses. Array-level operation is validated in a 5 x 5 selector-free crosspoint array. The performance enhancement of the VCF arises from the co-optimization of the all-ALD-grown process, device architecture, and operation scheme, enabling high density, a wide memory window, and strong reliability for scalable FeRAM integration.

cond-mat.mtrl-sci

Exploration and Application of AI in 6G Field

The recent upsurge of diversified mobile applications, especially those supported by AI, is spurring heated discussions on the future evolution of wireless communications. While 5G is being deployed around the world, efforts from industry and academia have started to look beyond 5G and conceptualize 6G. We envision 6G to experience an unprecedented transformation that will make it completely different from the previous generations of wireless systems. In particular, 6G will go beyond mobile Internet and will be required to support AI services. Meanwhile, AI will play a critical role in designing and optimizing 6G architectures, protocols and operations. In this article, we discuss the features of 6G, and the difficulties of carrying out 6G, and AI-enabled methods for 6G network design and optimization.

cs.NI

Symbol Rate and Carries Estimation in OFDM Framework: A high Accuracy Technique under Low SNR

Under a low Signal-to-Noise Ratio (SNR), the Orthogonal Frequency-Division Multiplexing (OFDM) signal symbol rate is limited. Existing carrier number estimation algorithms lack adequate methods to deal with low SNR. This paper proposes an algorithm with a low error rate under low SNR by correlating the signal and applying a Fast Fourier Transform (FFT) operation. By improving existing algorithms, we improve the performance of the OFDM carrier count algorithm. The performance of the OFDM's useful symbol time estimation algorithm is improved by estimating the number of carriers and symbol rate.

cs.IT