SearcharxivSearch

arXiv subjects

Bowen Zhao

Publications and source records attributed to Bowen Zhao.

At least 19 recordsLinked to original sources

Redshift-Dependent Intrinsic Dispersion in the Quasar UV/X-ray Luminosity Relation

Accurate modeling of the intrinsic dispersion in the quasar UV/X-ray luminosity relation is essential for reliable cosmological inference. We investigate its redshift dependence using luminosity distances reconstructed from cosmic chronometer and baryon acoustic oscillation measurements through Gaussian-process (GP) regression. Bayesian model comparison and posterior constraints show that the intrinsic dispersion is not well described by a single redshift-independent constant over $0.7<z<2.6$. It remains approximately constant at $0.7<z<1.6$, but shows an overall decreasing trend in the higher-redshift interval $1.6<z<2.6$, where the redshift-dependent intrinsic-dispersion model is decisively favored. This conclusion remains qualitatively robust against changes in the scaling-relation parameterization, GP kernel, and redshift binning scheme. We further examine its impact on cosmological inference in the flat $\Lambda$CDM model and find that, under the adopted calibration setup, the redshift-dependent intrinsic-dispersion model shifts the posterior median of $\Omega_{\rm m0}$ by $\Delta\Omega_{\rm m0}\simeq 0.025$. This indicates that intrinsic-dispersion modeling is a non-negligible component of the systematic-error budget for quasar cosmology and should be accounted for in future precision analyses.

astro-ph.CO

VEPHand: View-Efficient Photometric Hand Performance Capture at Scale

Robust, high-fidelity 3D hand capture, while fundamental to digital human creation, remains challenging with practical multi-view systems that balance rich photometry with the geometric ambiguities of reconstruction arising from limited viewpoint density. This paper presents an end-to-end pipeline for dynamic hand performance capture and registration, specifically designed for view-efficient setups ($\sim$20 views). We address key challenges with two primary innovations. First, to overcome reconstruction difficulties like limited view overlap and background clutter, our mask-free neural method robustly extracts detailed hand geometry and appearance from unmasked images using scene parameterization and scenario-specific density regularization. Second, addressing registration challenges such as accurately capturing non-linear skin deformations and ensuring plausible results during severe self-contact, we propose a physics-inspired framework. It aligns reconstructions to a personalized hand model by optimizing intrinsic volumetric offsets within its canonical tetrahedral mesh, alongside pose parameters. This approach, supported by robust losses and optimization, captures fine surface deformations, ensures plausible results under severe articulation and self-contact, and demonstrates strong tolerance to input noise. We demonstrate the scalability and robustness of our automated pipeline on an extensive dataset of over 12,000 sequences, from which we also derive a large-scale, high-quality synthetic 2D/3D hand dataset for training downstream tasks. This showcases its effectiveness for single hands, intricate two-hand interactions, and natural hand-object manipulations. Our method achieves state-of-the-art reconstruction fidelity in view-efficient, unmasked scenarios and highly accurate registration. Our project page are available at https://vephand.github.io/.

cs.CV

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent five-year evaluation subset with 12,484 scored questions, including 9,905 text-only questions and 2,579 questions with images. We evaluate 21 proprietary, open-source, and medical-specific models, reporting text-only and with-image performance separately. Because these subsets contain different questions, we further introduce a paired image-removal audit that evaluates questions with images before and after removing visual content to explore four answer-transition states. The audit shows that proprietary and open source models gain substantially from images, whereas medical-specific systems show limited observable use of visual evidence, with many correct answers persisting after image removal. Even among proprietary models, the net image-removal effect varies sevenfold across professions, from +5.7 points on Physician questions to +39.8 points on Public Health Nurse questions. We release JMed48k to support reproducible, profession-stratified evaluation of vision-language models in medical licensing settings.

cs.CV

Memento-Skills: Let Agents Design Agents

We introduce \emph{Memento-Skills}, a generalist, continually-learnable LLM agent system that functions as an \emph{agent-designing agent}: it autonomously constructs, adapts, and improves task-specific agents through experience. The system is built on a memory-based reinforcement learning framework with \emph{stateful prompts}, where reusable skills (stored as structured markdown files) serve as persistent, evolving memory. These skills encode both behaviour and context, enabling the agent to carry forward knowledge across interactions. Starting from simple elementary skills (like Web search and terminal operations), the agent continually improves via the \emph{Read--Write Reflective Learning} mechanism introduced in \emph{Memento~2}~\cite{wang2025memento2}. In the \emph{read} phase, a behaviour-trainable skill router selects the most relevant skill conditioned on the current stateful prompt; in the \emph{write} phase, the agent updates and expands its skill library based on new experience. This closed-loop design enables \emph{continual learning without updating LLM parameters}, as all adaptation is realised through the evolution of externalised skills and prompts. Unlike prior approaches that rely on human-designed agents, Memento-Skills enables a generalist agent to \emph{design agents end-to-end} for new tasks. Through iterative skill generation and refinement, the system progressively improves its own capabilities. Experiments on the \emph{General AI Assistants} benchmark and \emph{Humanity's Last Exam} demonstrate sustained gains, achieving 26.2\% and 116.2\% relative improvements in overall accuracy, respectively. Code is available at https://github.com/Memento-Teams/Memento-Skills.

cs.AI

Graph models for covariant holographic entropy I

We construct a graph model for holographic entropies in general time-dependent spacetimes. In static settings, such models arise from Ryu-Takayanagi surfaces on a common Cauchy slice and imply that the holographic entropy cone is polyhedral. Extending this construction to the covariant Hubeny-Rangamani-Takayanagi (HRT) setting is obstructed by the absence of a preferred time slice, raising the possibility of unphysical "short-cuts" built from partial HRT surfaces. We identify a geometric condition--the existence of exposed regions for each pair of HRT surfaces--under which this obstruction is removed. Under this condition, we construct weight functions by projecting along null generators of entanglement horizons and prove a Conditional No-Short-Cut Theorem: any graph cut is dominated by a surface composed of complete HRT surfaces. Consequently, the graph model reproduces HRT entropies, establishing the equivalence between the covariant and static holographic entropy cones in this regime. We further show that configurations in which exposed regions are absent due to nesting of interaction regions can be partially resolved by grouping HRT surfaces into timelike clusters. This provides evidence that the graph model extends beyond the exposed-region regime and suggests a path toward a complete covariant construction.

hep-th

New conditions for multipartite entanglement wedge connectivity in $n$-to-$n$ holographic scattering

We investigate the geometry of entanglement wedges for asymptotic $n$-to-$n$ scattering configurations in asymptotically AdS$_3$ spacetimes. Extending the $2$-to-$2$ Connected Wedge Theorem, we establish a strictly weaker sufficient condition for the input entanglement wedge $\mathcal{E}(V_1\cup\cdots\cup V_n)$ to be connected: the existence of a single pair of input regions satisfying a $2$-to-all causal intersection condition already forces full multipartite wedge connectivity. We also derive novel necessary conditions, showing that when the input wedge is connected, certain output ridges must enter the input wedge, and we organize these consequences into a layered reduction on the boundary lattice. Furthermore, we analyze the generalized bulk scattering region $\mathcal{S}_E = \mathcal{E}(V_1\cup\cdots\cup V_n)\cap \mathcal{E}(W_1\cup\cdots\cup W_n)$ and obtain necessary conditions for it to be nonempty; for $n>2$ these conditions are stronger than mere wedge connectedness. Our results provide new geometric restrictions on multipartite entanglement in holography and clarify the holographic dictionary for multi-partite scattering processes, while also highlighting intrinsic limitations for $n>2$.

hep-th

Single-hole spectral functions in one-dimensional quantum magnets with different ground states

Recent advances in numerical analytic continuation with physics-motivated constraints allow sharp spectral features to be extracted from imaginary-time quantum Monte Carlo (QMC) data. We apply these methods to one-dimensional $S=1/2$ spin systems with a single ejected fermion, computing the momentum- and energy-dependent single-hole spectral function $A(k,\omega)$. The real-space Green's function $G(r,\tau)$ is evaluated using Angelucci's canonical transformation [Phys. Rev. B 51, 11580 (1995)] implemented within stochastic series expansion QMC, and $A(k,\omega)$ is obtained by constrained stochastic analytic continuation. We contrast systems exhibiting spin-charge separation with those forming a spin polaron through effective spin-charge attraction. For the conventional $t$-$J$ chain, we recover the established signatures of spin-charge separation. Adding a multispin interaction $Q$ drives the system into a spontaneously dimerized valence-bond-solid (VBS) state; spin-charge-separation features persist up to the transition. Although the spectra generally agree with the conventional analytical ansatz, we find a gap between two holon bands that the ansatz predicts to be degenerate at $k=0$ and $k=\pi$. Deep in the VBS phase, the spectra provide evidence for spinon-holon binding at large $Q/J$. In a statically dimerized $t$-$J$ chain, we observe equally spaced spin-polaron bands associated with increasingly large bound states and two internal modes, even and odd under parton permutation. These results demonstrate the power of constrained analytic continuation combined with large-scale QMC for resolving sharp spectral features and distinguishing fractionalized from bound excitations.

cond-mat.str-el

PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting

The attention mechanism has demonstrated remarkable potential in sequence modeling, exemplified by its successful application in natural language processing with models such as Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformer (GPT). Despite these advancements, its utilization in time series forecasting (TSF) has yet to meet expectations. Exploring a better network structure for attention in TSF holds immense significance across various domains. In this paper, we present PeriodNet with a brand new structure to forecast univariate and multivariate time series. PeriodNet incorporates period attention and sparse period attention mechanism for analyzing adjacent periods. It enhances the mining of local characteristics, periodic patterns, and global dependencies. For efficient cross-variable modeling, we introduce an iterative grouping mechanism which can directly reduce the cross-variable redundancy. To fully leverage the extracted features on the encoder side, we redesign the entire architecture of the vanilla Transformer and propose a period diffuser for precise multi-period prediction. Through comprehensive experiments conducted on eight datasets, we demonstrate that PeriodNet outperforms six state-of-the-art models in both univariate and multivariate TSF scenarios in terms of mean square error and mean absolute error. In particular, PeriodNet achieves a relative improvement of 22% when forecasting time series with a length of 720, in comparison to other models based on the conventional encoder-decoder Transformer architecture.

cs.LG

Learning-Based Blockage-Resilient Beam Training in Near-Field Terahertz Communications

Terahertz (THz) band is considered a promising candidate to meet the high-throughput requirement for future sixth-generation (6G) wireless communications due to its ultrawide bandwidth. However, due to the high penetration loss at high-frequencies, blockage becomes a serious problem in THz communications, especially in near-field indoor communications with numerous obstacles. To address this issue, this paper investigates blockage-resilient near-field beam training based on self-accelerating Airy beam, which can propagate along a curved trajectory to circumvent obstacles. Specifically, we first analyze the trajectory of the Airy beam and the beam pattern at the receiver using a discrete Fourier transform (DFT) codebook in the presence of obstacles. Interestingly, we reveal that the beam pattern not only captures the receiver's location information but also implicitly encodes the spatial relationship between the receiver and obstacle, which facilitates identifying the optimal Airy beam configuration. Based on this insight, we formulate the blockage-resilient beam training task as a multitask learning problem and propose a lightweight attention-based multi-parameter beam training network (AMPBT-Net) to jointly predict the angle, distance, and curvature parameters of the optimal Airy beam based on the beam pattern. Finally, simulation results demonstrate that the Airy beam effectively mitigates blockage effects and the proposed scheme achieves comparable performance to exhaustive beam sweeping while significantly reducing training overhead.

eess.SP

A proof of Generalized Connected Wedge Theorem

In the context of asymptotic $2$-to-$2$ scattering process in AdS/CFT, the Connected Wedge Theorem identifies the existence of $O(1/G_N)$ mutual information between suitable boundary subregions, referred to as decision regions, as a necessary but not sufficient condition for bulk-only scattering processes, i.e., nonempty bulk scattering region $S_0$. Recently, Liu and Leutheusser proposed an enlarged bulk scattering region $S_E$ and conjectured that the non-emptiness of $S_E$ fully characterizes the existence of $O(1/G_N)$ mutual information between decision regions. Here, we provide a geometrical or general relativity proof for a slightly modified version of their conjecture.

hep-th

NNQS-AFQMC: Neural network quantum states enhanced fermionic quantum Monte Carlo

We introduce an efficient approach to implement neural network quantum states (NNQS) as trial wavefunctions in auxiliary-field quantum Monte Carlo (AFQMC). NNQS are a recently developed class of variational ans\"atze capable of flexibly representing many-body wavefunctions, though they often incur a high computational cost during optimization. AFQMC, on the other hand, is a powerful stochastic projector approach for ground-state calculations, but it normally requires an approximate constraint via a trial wavefunction or trial density matrix, whose quality affects the accuracy. Recently it has been shown (Xiao et al, arXiv2505.18519) that a broad class of highly correlated wave-functions can be integrated into AFQMC through stochastic sampling techniques. In this work, we apply this approach and present a direct integration of NNQS with AFQMC, allowing NNQS to serve as high-quality trial wavefunctions for AFQMC with manageable computational cost. We test the NNQS-AFQMC method on the challenging nitrogen molecule (N$_2$) at stretched geometries. Our results demonstrate that AFQMC with an NNQS trial wavefunction can attain near-exact total energies, highlighting the potential of AFQMC with NNQS to overcome longstanding challenges in strongly correlated electronic structure calculations. We also outline future research directions for improving this promising methodology.

physics.chem-ph

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

Despite tremendous recent advances in large model reasoning ability, vision-language models (VLMs) still struggle with detailed visual reasoning, especially when compute resources are limited. To address this challenge, we draw inspiration from methods like Deepseek-r1 for VLMs and train smaller-scale models with Group Relative Policy Optimization (GRPO) to use external tools such as zoom. The greatest benefit is obtained with a combination of GRPO learning, a simple reward structure, a simplified tool-calling interface, allocating additional tokens to the result of the tool call, and a training data mix that over-represents visually difficult examples. Compared to similarly-sized baseline models, our method achieves better performance on some visual question-answering (VQA) tasks, thanks to the detailed visual information gathered from the external tool.

cs.LG

Pura: An Efficient Privacy-Preserving Solution for Face Recognition

Face recognition is an effective technology for identifying a target person by facial images. However, sensitive facial images raises privacy concerns. Although privacy-preserving face recognition is one of potential solutions, this solution neither fully addresses the privacy concerns nor is efficient enough. To this end, we propose an efficient privacy-preserving solution for face recognition, named Pura, which sufficiently protects facial privacy and supports face recognition over encrypted data efficiently. Specifically, we propose a privacy-preserving and non-interactive architecture for face recognition through the threshold Paillier cryptosystem. Additionally, we carefully design a suite of underlying secure computing protocols to enable efficient operations of face recognition over encrypted data directly. Furthermore, we introduce a parallel computing mechanism to enhance the performance of the proposed secure computing protocols. Privacy analysis demonstrates that Pura fully safeguards personal facial privacy. Experimental evaluations demonstrate that Pura achieves recognition speeds up to 16 times faster than the state-of-the-art.

cs.CV

Unknown Word Detection for English as a Second Language (ESL) Learners Using Gaze and Pre-trained Language Models

English as a Second Language (ESL) learners often encounter unknown words that hinder their text comprehension. Automatically detecting these words as users read can enable computing systems to provide just-in-time definitions, synonyms, or contextual explanations, thereby helping users learn vocabulary in a natural and seamless manner. This paper presents EyeLingo, a transformer-based machine learning method that predicts the probability of unknown words based on text content and eye gaze trajectory in real time with high accuracy. A 20-participant user study revealed that our method can achieve an accuracy of 97.6%, and an F1-score of 71.1%. We implemented a real-time reading assistance prototype to show the effectiveness of EyeLingo. The user study shows improvement in willingness to use and usefulness compared to baseline methods.

cs.HC

TRUST: A Toolkit for TEE-Assisted Secure Outsourced Computation over Integers

Secure outsourced computation (SOC) provides secure computing services by taking advantage of the computation power of cloud computing and the technology of privacy computing (e.g., homomorphic encryption). Expanding computational operations on encrypted data (e.g., enabling complex calculations directly over ciphertexts) and broadening the applicability of SOC across diverse use cases remain critical yet challenging research topics in the field. Nevertheless, previous SOC solutions frequently lack the computational efficiency and adaptability required to fully meet evolving demands. To this end, in this paper, we propose a toolkit for TEE-assisted (Trusted Execution Environment) SOC over integers, named TRUST. In terms of system architecture, TRUST falls in a single TEE-equipped cloud server only through seamlessly integrating the computation of REE (Rich Execution Environment) and TEE. In consideration of TEE being difficult to permanently store data and being vulnerable to attacks, we introduce a (2, 2)-threshold homomorphic cryptosystem to fit the hybrid computation between REE and TEE. Additionally, we carefully design a suite of SOC protocols supporting unary, binary and ternary operations. To achieve applications, we present \texttt{SEAT}, secure data trading based on TRUST. Security analysis demonstrates that TRUST enables SOC, avoids collusion attacks among multiple cloud servers, and mitigates potential secret leakage risks within TEE (e.g., from side-channel attacks). Experimental evaluations indicate that TRUST outperforms the state-of-the-art and requires no alignment of data as well as any network communications. Furthermore, \texttt{SEAT} is as effective as the \texttt{Baseline} without any data protection.

cs.CR

CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart

Multimodal Question Answering (MMQA) is crucial as it enables comprehensive understanding and accurate responses by integrating insights from diverse data representations such as tables, charts, and text. Most existing researches in MMQA only focus on two modalities such as image-text QA, table-text QA and chart-text QA, and there remains a notable scarcity in studies that investigate the joint analysis of text, tables, and charts. In this paper, we present C$\text{T}^2$C-QA, a pioneering Chinese reasoning-based QA dataset that includes an extensive collection of text, tables, and charts, meticulously compiled from 200 selectively sourced webpages. Our dataset simulates real webpages and serves as a great test for the capability of the model to analyze and reason with multimodal data, because the answer to a question could appear in various modalities, or even potentially not exist at all. Additionally, we present AED (\textbf{A}llocating, \textbf{E}xpert and \textbf{D}esicion), a multi-agent system implemented through collaborative deployment, information interaction, and collective decision-making among different agents. Specifically, the Assignment Agent is in charge of selecting and activating expert agents, including those proficient in text, tables, and charts. The Decision Agent bears the responsibility of delivering the final verdict, drawing upon the analytical insights provided by these expert agents. We execute a comprehensive analysis, comparing AED with various state-of-the-art models in MMQA, including GPT-4. The experimental outcomes demonstrate that current methodologies, including GPT-4, are yet to meet the benchmarks set by our dataset.

cs.CL

MORSE: An Efficient Homomorphic Secret Sharing Scheme Enabling Non-Linear Operation

Homomorphic secret sharing (HSS) enables two servers to locally perform functions on encrypted data directly and obtain the results in the form of shares. A Paillier-based HSS solution seamlessly achieves multiplicative homomorphism and consumes less communication costs. Unfortunately, existing Paillier-based HSS schemes suffer from a large private key size, potential calculation error, expensive computation and storage overhead, and only valid on linear operations (e.g., addition and multiplication). To this end, inspired by the Paillier cryptosystem with fast encryption and decryption, we propose MORSE, an efficient homomorphic secret sharing scheme enabling non-linear operation, which enjoys a small key size, no calculation error and low overhead. In terms of functions, MORSE supports addition, subtraction, multiplication, scalar-multiplication, and comparison. Particularly, we carefully design two conversion protocols achieving the mutual conversion between one Paillier ciphertext and two secret shares, which allows MORSE to continuously perform the above operations. Rigorous analyses demonstrate that MORSE securely outputs correct results. Experimental results show that MORSE makes a runtime improvement of up to 9.3 times in terms of secure multiplication, and a communication costs reduction of up to 16.6% in secure comparison, compared to the state-of-the-art.

cs.CR

Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?

Large vision-language models (VLMs) have become state-of-the-art for many computer vision tasks, with in-context learning (ICL) as a popular adaptation strategy for new ones. But can VLMs learn novel concepts purely from visual demonstrations, or are they limited to adapting to the output format of ICL examples? We propose a new benchmark we call Spatial Visual Ambiguity Tasks (SVAT) that challenges state-of-the-art VLMs to learn new visuospatial tasks in-context. We find that VLMs fail to do this zero-shot, and sometimes continue to fail after finetuning. However, adding simpler data to the training by curriculum learning leads to improved ICL performance.

cs.CV