SearcharxivSearch

arXiv subjects

Ziyang Xiong

Publications and source records attributed to Ziyang Xiong.

13 recordsLinked to original sources

OPENPATH: A Supervisor--Specialist Agent System for Personalized, Accessible, and Multi-stop Urban Trip Planning

Urban trip-planning systems are commonly optimized for travel time and cost, but they offer limited support for the heterogeneous needs that real travelers bring, such as personalized preferences, multi-stop itinerary construction, and end-to-end wheelchair accessibility. We present openpaths, a supervisor-specialist multi-agent system that handles all of these tasks within a single architecture. openpaths adopts a deliberate division of labor: LLM agents parse natural-language input, classify request intent, and orchestrate execution, while classical algorithms perform route optimization over curated mobility and accessibility data. This design ensures that the resulting trip honors heterogeneous user preferences and enforces strict accessibility requirements when requested. Beyond per-user planning, openpaths doubles as a measurement instrument for city-scale accessibility analysis: applied to NYC, the system reveals substantial ADA infrastructure gaps and quantifies their effect on job accessibility for wheelchair users. Overall, this study shows how a supervisor-specialist LLM agentic framework can support heterogeneous trip planning and transparent, equitable transportation analysis in real urban environments.

eess.SY

Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping

Continual incorporation of new knowledge is essential for the long-term evolution of large language models (LLMs). Existing approaches typically rely on parameter-update algorithms to mitigate catastrophic forgetting, yet they suffer from fundamental limitations: 1) forgetting is unavoidable as the amount of newly injected knowledge grows; and 2) model updates are often irreversible. As modern LLMs become increasingly expressive, it is natural to question whether large-scale weight updates are necessary for acquiring a small amount of new knowledge. In this work, we propose a principled framework that models autoregressive language generation as a Markov process over tokens, where model memory is represented by a Markov transition matrix. Under this formulation, incorporating new knowledge/tokens corresponds to extending the state space, and preserving existing transitions guarantees retention of previously learned knowledge. We then prove a sample complexity bound for incorporating new tokens via a token-to-dictionary mapping strategy. In particular, for learning the transition behavior of each new token, the required number of samples scales linearly with the number of existing tokens it is mapped to. To realize this mapping, we propose an embedding-tuning algorithm that requires minimal parameter updates and induces zero forgetting. Experimental results further demonstrate the effectiveness of our method and validate our theoretical findings.

cs.LG

Holistic Multi-Scale Inference of the Leverage Effect: Efficiency under Dependent Microstructure Noise

This paper addresses the long-standing challenge of estimating the leverage effect from high-frequency data contaminated by dependent, non-Gaussian microstructure noise. We depart from the conventional reliance on pre-averaging or volatility "plug-in" methods by introducing a holistic multi-scale framework that operates directly on the leverage effect. We propose two novel estimators: the Subsampling-and-Averaging Leverage Effect (SALE) and the Multi-Scale Leverage Effect (MSLE). Central to our approach is a shifted window technique that constructs a noise-unbiased base estimator, significantly simplifying the multi-scale architecture. We provide a rigorous theoretical foundation for these estimators, establishing central limit theorems and stable convergence results that remain valid under both noise-free and dependent-noise settings. The primary contribution to estimation efficiency is a specifically designed weighting strategy for the MSLE estimator. By optimizing the weights based on the asymptotic covariance structure across scales and incorporating finite-sample variance corrections, we achieve substantial efficiency gains over existing benchmarks. Extensive simulation studies and an empirical analysis of 30 U.S. assets demonstrate that our framework consistently yields smaller estimation errors and superior performance in realistic, noisy market environments.

stat.ME

FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases

Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert curators to search relevant papers, reconcile evidence across documents, and produce ontology-grounded annotations - a workflow that existing benchmarks, focused on isolated subtasks like named entity recognition or relation extraction, do not capture. We present FlyBench to evaluate AI agents on end-to-end agentic ontology curation from scientific literature. Given only a gene symbol, agents must search and read from a corpus of 16,898 full-text papers to produce structured annotations: Gene Ontology terms describing function, expression patterns, and historical synonyms linking decades of nomenclature. The benchmark includes 7,397 expert-curated annotations across 100 genes drawn from FlyBase, the Drosophila (fruit fly) knowledge base. We evaluate four baseline agent architectures: memorization, fixed pipeline, single-agent, and multi-agent. We find that architectural choices significantly impact performance, with multi-agent designs outperforming simpler alternatives, yet scaling backbone models yields diminishing returns. All baselines leave substantial room for improvement. Our analysis surfaces several findings to guide future development; for example, agents primarily use retrieval to confirm parametric knowledge rather than discover new information. We hope FlyBench will drive progress on retrieval-augmented scientific reasoning, a capability with broad applications across scientific domains.

cs.AI

Ultracompact high-Q whispering gallery mode microresonator in a non-closed waveguide path

Integrated photonic circuits are foundational for versatile applications, where high-performance traveling-wave optical resonators are critical. Conventional whispering-gallery mode microresonators (WGMRs) confine light in closed-loop waveguide paths, thus inevitably occupy large footprints. Here, we report an ultracompact high loaded Q silicon photonic WGMR in an open curved path instead. By leveraging spatial mode multiplexing, low-loss mode converter-based photonic routers enable reentrant photon recycling in a single non-closed waveguide. The fabricated device achieves a measured loaded Q-factor of 1.78*10^5 at 1554.3 nm with a 1.05 nm free spectral range in a ultracompact footprint of 0.00137 mm^2-6*smaller than standard WGMRs while delivering 100*higher Q-factor than photonic crystal counterparts. This work pioneers dense integration of high-performance WGMR arrays through open-path mode recirculation.

physics.optics

Integrated Silicon Photonic Multichannel Optical Hybrid for Broadband Parallel Coherent Reception

We design and demonstrate a monolithically integrated silicon photonic multichannel optical hybrid for versatile broadband coherent reception, addressing the critical limitations of current wavelength multiplexed systems in scalability and power efficiency. The device combines a phase-compensated 90-degree optical hybrid with four robust three-stage Mach-Zehnder interferometer lattice filters, enabling 34-port functionality (two inputs and 32 outputs) for simultaneous analog and digital signal processing. Leveraging multimode interferometer designs,the chip achieves a broadband response with sub-dB passband uniformity across eight 200 GHz-spaced wavelength channels, while maintaining phase errors below 4 degrees over a 13.5 nm (1539-1552.5 nm) bandwidth with only 2.5 mW thermal tuning power.Experimentally, we validate its parallel-processing capability through RF channelizer reception (showing an average spurious-free dynamic range of 80.8 dB*Hz2/3 and image rejection ratio of 33.26 dB) and coherent optical communication (achieving 1.024 Tb/s data rate for 32-QAM signals with bit error rates far below the 20% SD-FEC threshold). The scheme enhances system performance with fully passive wavelength multiplexing integration, supporting high-fidelity uniformity and projecting scalability to 1.468 Tb/s. This work promises advancements in high-performance optoelectronic devices for next-generation AI-driven data centers and 5G-XG networks.

physics.optics

Miniaturized Computational Dispersion-Engineered Silicon Photonic Vernier Caliper Spectrometer

The development of miniaturized spectrometers for cost-effective mobile applications remains challenging, as small footprints fundamentally degrade bandwidth and resolution. Typically, achieving high resolution necessitates extended and sophisticated optical paths for spectral decorrelation. These restrict bandwidth both physically (through resonant wavelength periodicity constraints) and mathematically (due to resulting ill-conditioned large matrix factorizations). Here, we report a spectrometer using a computational dispersion-engineered silicon photonic Vernier caliper. This deterministic design enables periodicity-suppressed orthogonal measurements by nature, thus overcoming the bandwidth-resolution-footprint limit of current chip-scale spectrometers. Leveraging the dispersion-engineered Vernier subwavelength grating microrings and factorization-free matrix computation, a spectral resolution of 1.4 pm is achieved throughout a bandwidth of >160 nm with a footprint of <55*35 μm2 in a single detection channel,establishing the highest bandwidth-to-resolution-to-footprint ratio (>57 μm-2) demonstrated to date. Furthermore, broadband densely overlapped molecular absorption spectra of hydrogen cyanide are precisely measured, resolving 49 R- and P-branch lines with linewidths ranging from 15 to 86 pm which is fundamentally challenging for compressive sensing approaches. Our chip-scale spectrometer provides a new path toward precise and real-time multi-species spectral analysis and facilitates their commercialization.

physics.optics

Versatile and reconfigurable integrated silicon nitride photonic microresonator

Unlocking the full potential of integrated photonics requires versatile, multi-functional devices that can adapt to diverse application demands. However, confronting this challenge with conventional single-function resonators often results in tedious and complex systems. We present an elegant solution: a versatile and reconfigurable dual-polarization Si3N4 microresonator that represents a paradigm shift in on-chip photonic designs. Our device, based on a binary-star orbital architecture, can be dynamically reconfigured into three distinct topologies: a Möbius-like microcavity, a Fabry-Pérot resonator, and a microring resonator. This unprecedented functionality is enabled by a tunable balanced Mach-Zehnder interferometer that facilitates controllable mutual mode coupling of counterpropagating lights using a single control knob. We experimentally demonstrate that the device not only supports polarization-diverse operation on a compact footprint but also gives rise to a rich variety of physical phenomena, including a standing wave cavity, a traveling wave cavity, free spectral range multiplication, and the photonic pinning effect. These behaviors are accurately modeled using the Transfer Matrix Method and intuitively explained by Temporal Coupled Mode Theory. Our results underscore the profound potential for a chip-scale platform to realize reconfigurable reconstructive spectrometers and on-chip synthetic dimensions for topological physics.

physics.optics

Million-Q Dual-Polarization Micro-Fabry-Perot Resonators in Silicon Nitride Photonic Integrated Circuits

Miniaturized Fabry-Perot standing-wave resonators and whispering-gallery travelling wave resonators constitute foundational building blocks for photonic integrated circuits. While both architectures offer transformative potential through high quality factors and dual-polarization operation, integrated Fabry-Perot resonators face significant challenges in simultaneously achieving ultra-high Q-factors and broadband thermal tunability for fundamental transverse magnetic (TM0) and transverse electric (TE0) modes within a compact footprint-primarily due to polarization-dependent losses in conventional chip-scale reflectors. Here, we overcome this limitation by demonstrating an integrated silicon nitride dual-polarization micro-Fabry-Perot resonator with polarization-insensitive Sagnac loop reflectors and multimode waveguides to effectively suppress losses and enable high-performances for both fundamental transverse magnetic (TM0) and transverse electric (TE0) modes. The device achieves record loaded quality factors of 2.38*106 (TM0) and 3.48*105 (TE0) respectively and intrinsic quality factors will be even higher. Moreover, both two modes are tuned over the whole free spectral range of around 0.111 nm (TM0) and 0.112 nm (TE0) with the thermal tuning efficiencies of approximately 1.04 pm/mW (TM0) and 1.24 pm/mW (TE0). These advances establish a new benchmark for compact, high-performance dual-polarization resonators in optical sensors, nonlinear and integrated quantum photonics.

physics.optics

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Recent research enhances language model reasoning by scaling test-time compute via longer chain-of-thought traces. This often improves accuracy but also introduces redundancy and high computational cost, especially for small language models distilled with supervised fine-tuning (SFT). In this work, we propose new algorithms to improve token-efficient reasoning with small-scale models by effectively trading off accuracy and computation. We first show that the post-SFT model fails to determine the optimal stopping point of the reasoning process, resulting in verbose and repetitive outputs. Verbosity also significantly varies across wrong vs correct responses. To address these issues, we propose two solutions: (1) Temperature scaling (TS) to control the stopping point for the thinking phase and thereby trace length, and (2) TLDR: a length-regularized reinforcement learning method based on GRPO that facilitates multi-level trace length control (e.g. short, medium, long reasoning). Experiments on four reasoning benchmarks, MATH500, AMC, AIME24 and OlympiadBench, demonstrate that TS is highly effective compared to s1's budget forcing approach and TLDR significantly improves token efficiency by about 50% with minimal to no accuracy loss over the SFT baseline. Moreover, TLDR also facilitates flexible control over the response length, offering a practical and effective solution for token-efficient reasoning in small models. Ultimately, our work reveals the importance of stopping time control, highlights shortcomings of pure SFT, and provides effective algorithmic recipes.

cs.LG

MapExplorer: New Content Generation from Low-Dimensional Visualizations

Low-dimensional visualizations, or "projection maps," are widely used in scientific and creative domains to interpret large-scale and complex datasets. These visualizations not only aid in understanding existing knowledge spaces but also implicitly guide exploration into unknown areas. Although techniques such as t-SNE and UMAP can generate these maps, there exists no systematic method for leveraging them to generate new content. To address this, we introduce MapExplorer, a novel knowledge discovery task that translates coordinates within any projection map into coherent, contextually aligned textual content. This allows users to interactively explore and uncover insights embedded in the maps. To evaluate the performance of MapExplorer methods, we propose Atometric, a fine-grained metric inspired by ROUGE that quantifies logical coherence and alignment between generated and reference text. Experiments on diverse datasets demonstrate the versatility of MapExplorer in generating scientific hypotheses, crafting synthetic personas, and devising strategies for attacking large language models-even with simple baseline methods. By bridging visualization and generation, our work highlights the potential of MapExplorer to enable intuitive human-AI collaboration in large-scale data exploration.

cs.AI

LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks

Safety is a paramount concern for large language models (LLMs) in open deployment, motivating the development of safeguard methods that enforce ethical and responsible use through safety alignment or guardrail mechanisms. Jailbreak attacks that exploit the \emph{false negatives} of safeguard methods have emerged as a prominent research focus in the field of LLM security. However, we found that the malicious attackers could also exploit false positives of safeguards, i.e., fooling the safeguard model to block safe content mistakenly, leading to a denial-of-service (DoS) affecting LLM users. To bridge the knowledge gap of this overlooked threat, we explore multiple attack methods that include inserting a short adversarial prompt into user prompt templates and corrupting the LLM on the server by poisoned fine-tuning. In both ways, the attack triggers safeguard rejections of user requests from the client. Our evaluation demonstrates the severity of this threat across multiple scenarios. For instance, in the scenario of white-box adversarial prompt injection, the attacker can use our optimization process to automatically generate seemingly safe adversarial prompts, approximately only 30 characters long, that universally block over 97% of user requests on Llama Guard 3. These findings reveal a new dimension in LLM safeguard evaluation -- adversarial robustness to false positives.

cs.CR

MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows

Scientific innovation relies on detailed workflows, which include critical steps such as analyzing literature, generating ideas, validating these ideas, interpreting results, and inspiring follow-up research. However, scientific publications that document these workflows are extensive and unstructured. This makes it difficult for both human researchers and AI systems to effectively navigate and explore the space of scientific innovation. To address this issue, we introduce MASSW, a comprehensive text dataset on Multi-Aspect Summarization of Scientific Workflows. MASSW includes more than 152,000 peer-reviewed publications from 17 leading computer science conferences spanning the past 50 years. Using Large Language Models (LLMs), we automatically extract five core aspects from these publications -- context, key idea, method, outcome, and projected impact -- which correspond to five key steps in the research workflow. These structured summaries facilitate a variety of downstream tasks and analyses. The quality of the LLM-extracted summaries is validated by comparing them with human annotations. We demonstrate the utility of MASSW through multiple novel machine-learning tasks that can be benchmarked using this new dataset, which make various types of predictions and recommendations along the scientific workflow. MASSW holds significant potential for researchers to create and benchmark new AI methods for optimizing scientific workflows and fostering scientific innovation in the field. Our dataset is openly available at \url{https://github.com/xingjian-zhang/massw}.

cs.CL