SearcharxivSearch

arXiv subjects

Zhou

Publications and source records attributed to Zhou.

12 recordsLinked to original sources

Infinite Drive: Optimal Urban Location of Dynamic Wireless Charging at Signalized Intersections

Dynamic Wireless Power Transfer (DWPT) could eliminate plug-in charging in cities, but optimal urban deployment is complex. This paper develops a mixed-integer programming model that optimizes DWPT location under signalized intersection dynamics -- acceleration, deceleration, and queue-position-dependent dwell time -- through probabilistic signal patterns and saturation headway-based modeling. A case study of Kawagoe City, Japan, shows that electrifying 1.5% of the road network is sufficient to sustain continuous urban EV operation without plug-in charging for the baseline scenario, and at most 2.9% suffices across all tested assumptions. Monte Carlo simulations of continuous trip chains averaging approximately 600 km and reaching up to approximately 800 km confirm that optimized 12 kWh-battery deployments sustain operation in all simulated runs, revealing an infrastructure-battery tradeoff corresponding to roughly 1.7-3.0 tonnes CO2e of avoided battery manufacturing emissions per vehicle relative to a conventional 40 kWh urban EV. These findings position DWPT deployment as an environmentally efficient pathway for sustainable urban mobility when deployed optimally.

physics.soc-ph

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike prior code generation benchmarks that focus on algorithmic tasks, VISTA targets realistic UI-centric development, where agents must produce functional, visually coherent applications from underspecified inputs. We define five prompt-information conditions that vary along two axes, visual/structural fidelity and stack constraint: (1) text only with free stack choice, (2) text with reference screenshots under three specified stacks, (3) text with reference screenshots under free stack choice, (4) text with screenshots and pruned Figma structure under a single specified stack, and (5) text with screenshots and pruned Figma structure under free stack choice. To enable robust evaluation, each page in the benchmark is manually annotated with interactive UI components and around three visual anchor points, addressing the well-known limitations of script-based testing tools such as Playwright in open-ended code generation settings. Evaluation combines DOM-grounded reference matching, behavior-specific browser tests, and CLIP-based visual similarity, jointly measuring structural alignment, behavioral completeness, and overall visual fidelity. We use VISTA to assess four agent systems drawn from two model families and two harnesses, finding that visual fidelity and functional correctness are partially decoupled across both input conditions and agents, and that agent editing style varies sharply but is largely orthogonal to task quality. VISTA establishes a rigorous and reproducible foundation for advancing agent-based software engineering research. Code is available at https://github.com/kaboider/VISTA_Bench.

cs.SE

Age-Specific Logistic Regression with Complex Event Time Data

In attempt to advance the current practice for assessing and predicting the primary ovarian insufficiency (POI) risk in female childhood cancer survivors, we propose two estimating function based approaches for age-specific logistic regression. Both approaches adapt the inverse probability of censoring weighting (IPCW) strategy and yield consistent estimators with asymptotic normality. The first approach modifies the IPCW weights used by Im et al. (2023) to account for doubly censoring. The second approach extends the outcome weighted IPCW approach to use the information of the subjects censored before the analysis time. We consider variance estimation for the estimators and explore by simulation the two approaches implemented in the situations where the conditional right-censoring time distribution required in the IPCW weighs is unknown and approximated using the survival random forest approaches, stratified empirical distribution functions, or the estimator under the Cox proportional hazards model. The numerical studies indicate that the second approach is more efficient when right-censoring is relatively heavy, whereas the first approach is preferable when the right-censoring is light. We also observe that the performance of the two approaches heavily relies on the estimation of censoring distribution in our simulation settings. The POI data from a childhood cancer survivor study are employed throughout the paper for motivation and illustration. Our data analysis provides new insight into understanding the POI risk among cancer survivors.

stat.ME

COMPAS: A Distributed Multi-Party SWAP Test for Parallel Quantum Algorithms

The limited number of qubits per chip remains a critical bottleneck in quantum computing, motivating the use of distributed architectures that interconnect multiple quantum processing units (QPUs). However, executing quantum algorithms across distributed systems requires careful co-design of algorithmic primitives and hardware architectures to manage circuit depth and entanglement overhead. We identify multivariate trace estimation as a key subroutine that is naturally suited for distribution, and broadly useful in tasks such as estimating R\'enyi entropies, virtual cooling and distillation, and certain applications of quantum signal processing. In this work, we introduce COMPAS, an architecture that realizes multivariate trace estimation across a multi-party network of interconnected modular and distributed QPUs by leveraging pre-shared entangled Bell pairs as resources. COMPAS adds only a constant depth overhead and consumes Bell pairs at a rate linear in circuit width, making it suitable for near-term hardware. Unlike other schemes, which must choose between asymptotic optimality in circuit depth or GHZ width, COMPAS achieves both at once. Additionally, we analyze network-level errors and simulate the effects of circuit-level noise on the architecture.

quant-ph

On Vanishing Variance in Transformer Length Generalization

It is a widely known issue that Transformers, when trained on shorter sequences, fail to generalize robustly to longer ones at test time. This raises the question of whether Transformer models are real reasoning engines, despite their impressive abilities in mathematical problem solving and code synthesis. In this paper, we offer a vanishing variance perspective on this issue. To the best of our knowledge, we are the first to demonstrate that even for today's frontier models, a longer sequence length results in a decrease in variance in the output of the multi-head attention modules. On the argmax retrieval and dictionary lookup tasks, our experiments show that applying layer normalization after the attention outputs leads to significantly better length generalization. Our analyses attribute this improvement to a reduction-though not a complete elimination-of the distribution shift caused by vanishing variance.

cs.LG

Formal justification of a continuum relaxation model for one-dimensional moir\'e materials

Mechanical relaxation in moir\'e materials is often modeled by a continuum model where linear elasticity is coupled to a stacking penalty known as the Generalized Stacking Fault Energy (GSFE). We review and compute minimizers of a one-dimensional version of this model, and then show how it can be formally derived from a natural atomistic model. Specifically, we show that the continuum model emerges in the limit $\epsilon \downarrow 0$ and $\delta \downarrow 0$ while holding the ratio $\eta := \frac{\epsilon^2}{\delta}$ fixed, where $\epsilon$ is the ratio of the monolayer lattice constant to the moir\'e lattice constant and $\delta$ is the ratio of the typical stacking energy to the monolayer stiffness.

math-ph

Implementing Bayesian inference on a stochastic CO2-based grey-box model for assessing indoor air quality in Canadian primary schools

The COVID-19 pandemic brought global attention to indoor air quality (IAQ), which is intrinsically linked to clean air change rates. Estimating the air change rate in indoor environments, however, remains challenging. It is primarily due to the uncertainties associated with the air change rate estimation, such as pollutant generation rates, dynamics including weather and occupancies, and the limitations of deterministic approaches to accommodate these factors. In this study, Bayesian inference was implemented on a stochastic CO2-based grey-box model to infer modeled parameters and quantify uncertainties. The accuracy and robustness of the ventilation rate and CO2 emission rate estimated by the model were confirmed with CO2 tracer gas experiments conducted in an airtight chamber. Both prior and posterior predictive checks (PPC) were performed to demonstrate the advantage of this approach. In addition, uncertainties in real-life contexts were quantified with an incremental variance {\sigma} for the Wiener process. This approach was later applied to evaluate the ventilation conditions within two primary school classrooms in Montreal. The Equivalent Clean Airflow Rate (ECAi) was calculated following ASHRAE 241, and an insufficient clean air supply within both classrooms was identified. A supplement of 800 cfm clear air delivery rate (CADR) from air-cleaning devices is recommended for a sufficient ECAi. Finally, steady-state CO2 thresholds (Climit, Ctarget, and Cideal) were carried out to indicate when ECAi requirements could be achieved under various mitigation strategies, such as portable air cleaners and in-room ultraviolet light, with CADR values ranging from 200 to 1000 cfm.

stat.AP

Innovation Diffusion in EV Charging Location Decisions: Integrating Demand & Supply through Market Dynamics

This paper offers a strategic approach to Electric Vehicles (EVs) charging network planning, emphasizing the integration of demand and supply dynamics via continuous-time fluid queue models and discrete flow refueling location modeling, all in the context of innovation diffusion principles. We employ a continuous-time approximation based on Ordinary Differential Equations (ODEs) to design multi-year supply curves, a method that stands in contrast to conventional practices which often overlook inter-year transitions and ongoing processes. For medium-term charging station location planning (CSLP), we apply a flow refueling location model (FRLM) within grid-based multi-level networks, considering both multiple-path networks and capacity constraints. The grid-based network planning strategy uses a three-tier (Macro-Meso-Micro) approach for thorough EV charging station placement, with the macro-level covering entire cities, the meso-level assessing detailed EV routes and bridging the macro to micro levels, and the micro-level focusing on precise station placement for accessibility and efficiency. Our investigation into overutilization and underutilization scenarios delivers valuable insights for policymaking and cost-benefit analyses. Illustrating our approach with the example of the Chicago sketch network, we introduce an integrated demand-supply model suitable for a single region and extendable to multiple regions, thereby addressing a gap in the existing literature. Our proposed methodology focuses on EV station placement, taking into account future needs, geographical capacities, and the importance of scenario analysis, which empowers strategic resource planning for EV charging networks over extended timeframes, thus aiding the transition towards a more sustainable and efficient transportation system.

math.OC

Modeling Household Online Shopping Demand in the U.S.: A Machine Learning Approach and Comparative Investigation between 2009 and 2017

Despite the rapid growth of online shopping and research interest in the relationship between online and in-store shopping, national-level modeling and investigation of the demand for online shopping with a prediction focus remain limited in the literature. This paper differs from prior work and leverages two recent releases of the U.S. National Household Travel Survey (NHTS) data for 2009 and 2017 to develop machine learning (ML) models, specifically gradient boosting machine (GBM), for predicting household-level online shopping purchases. The NHTS data allow for not only conducting nationwide investigation but also at the level of households, which is more appropriate than at the individual level given the connected consumption and shopping needs of members in a household. We follow a systematic procedure for model development including employing Recursive Feature Elimination algorithm to select input variables (features) in order to reduce the risk of model overfitting and increase model explainability. Extensive post-modeling investigation is conducted in a comparative manner between 2009 and 2017, including quantifying the importance of each input variable in predicting online shopping demand, and characterizing value-dependent relationships between demand and the input variables. In doing so, two latest advances in machine learning techniques, namely Shapley value-based feature importance and Accumulated Local Effects plots, are adopted to overcome inherent drawbacks of the popular techniques in current ML modeling. The modeling and investigation are performed both at the national level and for three of the largest cities (New York, Los Angeles, and Houston). The models developed and insights gained can be used for online shopping-related freight demand generation and may also be considered for evaluating the potential impact of relevant policies on online shopping demand.

cs.LG

Bio-inspired Structure Identification in Language Embeddings

Word embeddings are a popular way to improve downstream performances in contemporary language modeling. However, the underlying geometric structure of the embedding space is not well understood. We present a series of explorations using bio-inspired methodology to traverse and visualize word embeddings, demonstrating evidence of discernible structure. Moreover, our model also produces word similarity rankings that are plausible yet very different from common similarity metrics, mainly cosine similarity and Euclidean distance. We show that our bio-inspired model can be used to investigate how different word embedding techniques result in different semantic outputs, which can emphasize or obscure particular interpretations in textual data.

cs.CL

Nonlinear Function Estimation with Empirical Bayes and Approximate Message Passing

Nonlinear function estimation is core to modern machine learning applications. In this paper, to perform nonlinear function estimation, we reduce a nonlinear inverse problem to a linear one using a polynomial kernel expansion. These kernels increase the feature set, and may result in poorly conditioned matrices. Nonetheless, we show several examples where the matrix in our linear inverse problem contains only mild linear correlations among columns. The coefficients vector is modeled within a Bayesian setting for which approximate message passing (AMP), an algorithmic framework for signal reconstruction, offers Bayes-optimal signal reconstruction quality. While the Bayesian setting limits the scope of our work, it is a first step toward estimation of real world nonlinear functions. The coefficients vector is estimated using two AMP-based approaches, a Bayesian one and empirical Bayes. Numerical results confirm that our AMP-based approaches learn the function better than LASSO, offering markedly lower error in predicting test data.

cs.IT

An Approximate Message Passing Framework for Side Information

Approximate message passing (AMP) methods have gained recent traction in sparse signal recovery. Additional information about the signal, or \emph{side information} (SI), is commonly available and can aid in efficient signal recovery. This work presents an AMP-based framework that exploits SI and can be readily implemented in various settings for which the SI results in separable distributions. To illustrate the simplicity and applicability of our approach, this framework is applied to a Bernoulli-Gaussian (BG) model and a time-varying birth-death-drift (BDD) signal model, motivated by applications in channel estimation. We develop a suite of algorithms, called AMP-SI, and derive denoisers for the BDD and BG models. Numerical evidence demonstrating the advantages of our approach are presented alongside empirical evidence of the accuracy of a proposed state evolution.

cs.IT