SearcharxivSearch

arXiv subjects

Yichen Shen

Publications and source records attributed to Yichen Shen.

At least 19 recordsLinked to original sources

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent. Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen-Verse/PAST-Bench

cs.CL

Conditional Predictive Inference for General Structured Data with Group Symmetries

We study distribution-free predictive inference for data with group symmetries, aiming to establish near-conditional coverage guarantees beyond exchangeability for structured data. While many predictive inference methods achieve a target coverage level, most provide marginal coverage. In practice, conditional predictive inference is often preferred, as it quantifies uncertainty for black-box predictions given observed attributes, thereby accommodating heterogeneity. Although many efforts have pursued efficient conditional coverage, existing methods rely on the i.i.d. or exchangeable assumption, often violated in structured settings such as networks, clusters, and imaging data. Recently, SymmPI introduced a unified approach to predictive inference under group symmetries beyond exchangeability; nevertheless, its guarantees remain marginal and do not account for population heterogeneity. To bridge this gap, we introduce C-SymmPI, a framework that achieves near-conditional coverage under general data structures with group symmetries, extending beyond exchangeability to cover networks, cluster-level data, and related structures. Inspired by relaxed multi-accuracy, our approach reformulates conditional coverage as miscoverage error over a user-specified function class. We establish theoretical guarantees under distributional invariance and distribution shift, and derive convergence rates for linear and RKHS function classes, recovering state-of-the-art results in the exchangeable setting as special cases. For computational efficiency, we develop two variants: a projection-based algorithm for high-dimensional observations, and a sampling-based algorithm for large or infinite groups. We demonstrate effectiveness on hierarchical and network data. Empirical results show that C-SymmPI delivers more informative and stable conditional coverage with improved accuracy compared to existing methods.

stat.ME

Quantum Dispersive Waves and Multimode Squeezing in Pure-Kerr Parametrically Driven Cavity Solitons

Parametrically driven cavity solitons (PDCS), unlike single-pumped cavity solitons, are localized optical pulses arising from parametric processes. These cavity solitons, recently discovered in pure-Kerr media, offer great promise for nonlinear dynamics studies and metrology. Here, we present the first multimode quantum description of pure-Kerr PDCS. In the below threshold regime, we verify single- and two-mode squeezing, while above threshold we uncover novel "quantum" dispersive waves - the quantum analog of soliton Cherenkov radiation. Besides revealing these unexplored quantum properties, we show that PDCS generates up to 20 dB of squeezing, only limited by overcoupling and intrinsic losses for experimentally routine parameters. We therefore provide a pathway to observe strong multimode quantum noise reduction in these systems.

quant-ph

Context-Adaptive Synthesis and Compression for Enhanced Retrieval-Augmented Generation in Complex Domains

Large Language Models (LLMs) excel in language tasks but are prone to hallucinations and outdated knowledge. Retrieval-Augmented Generation (RAG) mitigates these by grounding LLMs in external knowledge. However, in complex domains involving multiple, lengthy, or conflicting documents, traditional RAG suffers from information overload and inefficient synthesis, leading to inaccurate and untrustworthy answers. To address this, we propose CASC (Context-Adaptive Synthesis and Compression), a novel framework that intelligently processes retrieved contexts. CASC introduces a Context Analyzer & Synthesizer (CAS) module, powered by a fine-tuned smaller LLM, which performs key information extraction, cross-document consistency checking and conflict resolution, and question-oriented structured synthesis. This process transforms raw, scattered information into a highly condensed, structured, and semantically rich context, significantly reducing the token count and cognitive load for the final Reader LLM. We evaluate CASC on SciDocs-QA, a new challenging multi-document question answering dataset designed for complex scientific domains with inherent redundancies and conflicts. Our extensive experiments demonstrate that CASC consistently outperforms strong baselines.

cs.CL

Highly squeezed nanophotonic quantum microcombs with broadband frequency tunability

Squeezed light offers genuine quantum advantage in enhanced sensing and quantum computation; yet the level of squeezing or quantum noise reduction generated from nanophotonic chips has been limited. In addition to strong quantum noise reduction, key desiderata for such a nanophotonic squeezer include frequency agility or tunability over a broad frequency range, and simultaneous operation in many distinct, well-defined quantum modes (qumodes). Here we present a strongly overcoupled silicon nitride squeezer based on a below-threshold optical parametric amplifier (OPA) that produces directly detected squeezing of 5.6 dB $\pm$ 0.2 dB, surpassing previous demonstrations in both continuous-wave and pulsed regimes. We introduce a seed-assisted detection technique into such nanophotonic squeezers that reveals a quantum frequency comb (QFC) of 16 qumodes, with a separation of 11~THz between the furthest qumode pair, while maintaining a strong squeezing. Additionally, we report spectral tuning of a qumode comb pair over one free-spectral range of the OPA, thus bridging the spacing between the discrete modes of the QFC. Our results significantly advance both the generation and detection of nanophotonic squeezed light in a broadband and multimode platform, establishing a scalable, chip-integrated path for compact quantum sensors and continuous-variable quantum information processing systems.

physics.optics

CoProSketch: Controllable and Progressive Sketch Generation with Diffusion Model

Sketches serve as fundamental blueprints in artistic creation because sketch editing is easier and more intuitive than pixel-level RGB image editing for painting artists, yet sketch generation remains unexplored despite advancements in generative models. We propose a novel framework CoProSketch, providing prominent controllability and details for sketch generation with diffusion models. A straightforward method is fine-tuning a pretrained image generation diffusion model with binarized sketch images. However, we find that the diffusion models fail to generate clear binary images, which makes the produced sketches chaotic. We thus propose to represent the sketches by unsigned distance field (UDF), which is continuous and can be easily decoded to sketches through a lightweight network. With CoProSketch, users generate a rough sketch from a bounding box and a text prompt. The rough sketch can be manually edited and fed back into the model for iterative refinement and will be decoded to a detailed sketch as the final result. Additionally, we curate the first large-scale text-sketch paired dataset as the training data. Experiments demonstrate superior semantic consistency and controllability over baselines, offering a practical solution for integrating user feedback into generative workflows.

cs.CV

InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers

Scaling Large Language Model (LLM) training relies on multi-dimensional parallelism, where High-Bandwidth Domains (HBDs) are critical for communication-intensive parallelism like Tensor Parallelism. However, existing HBD architectures face fundamental limitations in scalability, cost, and fault resiliency: switch-centric HBDs (e.g., NVL-72) incur prohibitive scaling costs, while GPU-centric HBDs (e.g., TPUv3/Dojo) suffer from severe fault propagation. Switch-GPU hybrid HBDs (e.g., TPUv4) take a middle-ground approach, but the fault explosion radius remains large. We propose InfiniteHBD, a transceiver-centric HBD architecture that integrates connectivity and dynamic switching at the transceiver level by embedding Optical Circuit Switching (OCS) within each transceiver. It enables reconfigurable point-to-multipoint communication and scalable variable-size ring topologies. InfiniteHBD achieves datacenter-scale scalability without cost explosion, fault isolation at the node level, and full bandwidth utilization for healthy GPUs. Key innovations include a Silicon Photonic-based OCS transceiver (OCSTrx), a reconfigurable k-hop ring topology, and an HBD-DCN orchestration algorithm. The evaluation demonstrates that InfiniteHBD reduces cost to 31% of NVL-72, achieves a near-zero GPU waste ratio (over 10x lower than NVL-72 and TPUv4), maintains near-zero cross-ToR traffic under 7% node fault ratio, and improves Model FLOPs Utilization by 3.37x compared to NVIDIA DGX (8 GPUs/node).

cs.NI

Strong nanophotonic quantum squeezing exceeding 3.5 dB in a foundry-compatible Kerr microresonator

Squeezed light, with its quantum noise reduction capabilities, has emerged as a powerful resource in quantum information processing and precision metrology. To reach noise reduction levels such that a quantum advantage is achieved, off-chip squeezers are typically used. The development of on-chip squeezed light sources, particularly in nanophotonic platforms, has been challenging. We report 3.7 $\pm$ 0.2 dB of directly detected nanophotonic quantum squeezing using foundry-fabricated silicon nitride (Si$_3$N$_4$) microrings with an inferred squeezing level of 10.7 dB on-chip. The squeezing level is robust across multiple devices and pump detunings, and is consistent with the overcoupling degree without noticeable degradation from excess classical noise. We also offer insights to mitigate thermally-induced excess noise, that typically degrades squeezing, by using small-radius rings with a larger free spectral range (450 GHz) and consequently lower parametric oscillation thresholds. Our results demonstrate that Si$_3$N$_4$ is a viable platform for strong quantum noise reduction in a CMOS-compatible, scalable architecture.

physics.optics

BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation using RGB Frames and Events

Recent advances in event-based vision suggest that these systems complement traditional cameras by providing continuous observation without frame rate limitations and a high dynamic range, making them well-suited for correspondence tasks such as optical flow and point tracking. However, there is still a lack of comprehensive benchmarks for correspondence tasks that include both event data and images. To address this gap, we propose BlinkVision, a large-scale and diverse benchmark with multiple modalities and dense correspondence annotations. BlinkVision offers several valuable features: 1) Rich modalities: It includes both event data and RGB images. 2) Extensive annotations: It provides dense per-pixel annotations covering optical flow, scene flow, and point tracking. 3) Large vocabulary: It contains 410 everyday categories, sharing common classes with popular 2D and 3D datasets like LVIS and ShapeNet. 4) Naturalistic: It delivers photorealistic data and covers various naturalistic factors, such as camera shake and deformation. BlinkVision enables extensive benchmarks on three types of correspondence tasks (optical flow, point tracking, and scene flow estimation) for both image-based and event-based methods, offering new observations, practices, and insights for future research. The benchmark website is https://www.blinkvision.net/.

cs.CV

BlinkTrack: Feature Tracking over 80 FPS via Events and Images

Event cameras, known for their high temporal resolution and ability to capture asynchronous changes, have gained significant attention for their potential in feature tracking, especially in challenging conditions. However, event cameras lack the fine-grained texture information that conventional cameras provide, leading to error accumulation in tracking. To address this, we propose a novel framework, BlinkTrack, which integrates event data with grayscale images for high-frequency feature tracking. Our method extends the traditional Kalman filter into a learning-based framework, utilizing differentiable Kalman filters in both event and image branches. This approach improves single-modality tracking and effectively solves the data association and fusion from asynchronous event and image data. We also introduce new synthetic and augmented datasets to better evaluate our model. Experimental results indicate that BlinkTrack significantly outperforms existing methods, exceeding 80 FPS with multi-modality data and 100 FPS with preprocessed event data. Codes and dataset are available at https://github.com/ColieShen/BlinkTrack.

cs.CV

Photonic Integrated Neuro-Synaptic Core for Convolutional Spiking Neural Network

Neuromorphic photonic computing has emerged as a competitive computing paradigm to overcome the bottlenecks of the von-Neumann architecture. Linear weighting and nonlinear spiking activation are two fundamental functions of a photonic spiking neural network (PSNN). However, they are separately implemented with different photonic materials and devices, hindering the large-scale integration of PSNN. Here, we propose, fabricate and experimentally demonstrate a photonic neuro-synaptic chip enabling the simultaneous implementation of linear weighting and nonlinear spiking activation based on a distributed feedback (DFB) laser with a saturable absorber (DFB-SA). A prototypical system is experimentally constructed to demonstrate the parallel weighted function and nonlinear spike activation. Furthermore, a four-channel DFB-SA array is fabricated for realizing matrix convolution of a spiking convolutional neural network, achieving a recognition accuracy of 87% for the MNIST dataset. The fabricated neuro-synaptic chip offers a fundamental building block to construct the large-scale integrated PSNN chip.

physics.optics

A Brewster route to Cherenkov detectors

The Cherenkov effect enables a valuable tool, known as the Cherenkov detector, to identify high-energy particles via the measurement of the Cherenkov cone. However, the sensitivity and momentum coverage of such detectors are intrinsically limited by the refractive index of the host material. Especially, identifying particles with energy above multiple gigaelectronvolts requires host materials with a near-unity refractive index, which are often limited to large and bulky gas chambers. Overcoming this fundamental material limit is important for future particle detectors yet remains a long-standing scientific challenge. Here, we propose a different paradigm for Cherenkov detectors that utilizes the broadband angular filter made from stacks of variable one-dimensional photonic crystals. Owing to the Brewster effect, the angular filter is transparent only to Cherenkov photons from a precise incident angle, and particle identification is achieved by mapping each Cherenkov angle to the peak-intensity position of transmitted photons in the detection plane. This unique property of the angular filter is exceptionally beneficial to Cherenkov detection as it enables the realization of a non-dispersive pseudo refractive index over the entire visible spectrum. Moreover, such a pseudo refractive index can be flexibly tuned to arbitrary values, including those close to unity. Our angular-selective Brewster paradigm offers a feasible solution to implement compact and highly sensitive Cherenkov detectors especially in beam lines and it can cover a wide momentum range using readily available dielectric materials.

physics.optics

Vector-Vector-Matrix Architecture: A Novel Hardware-Aware Framework for Low-Latency Inference in NLP Applications

Deep neural networks have become the standard approach to building reliable Natural Language Processing (NLP) applications, ranging from Neural Machine Translation (NMT) to dialogue systems. However, improving accuracy by increasing the model size requires a large number of hardware computations, which can slow down NLP applications significantly at inference time. To address this issue, we propose a novel vector-vector-matrix architecture (VVMA), which greatly reduces the latency at inference time for NMT. This architecture takes advantage of specialized hardware that has low-latency vector-vector operations and higher-latency vector-matrix operations. It also reduces the number of parameters and FLOPs for virtually all models that rely on efficient matrix multipliers without significantly impacting accuracy. We present empirical results suggesting that our framework can reduce the latency of sequence-to-sequence and Transformer models used for NMT by a factor of four. Finally, we show evidence suggesting that our VVMA extends to other domains, and we discuss novel hardware for its efficient use.

cs.CL

Real-Time Uncertainty Estimation in Computer Vision via Uncertainty-Aware Distribution Distillation

Calibrated estimates of uncertainty are critical for many real-world computer vision applications of deep learning. While there are several widely-used uncertainty estimation methods, dropout inference stands out for its simplicity and efficacy. This technique, however, requires multiple forward passes through the network during inference and therefore can be too resource-intensive to be deployed in real-time applications. We propose a simple, easy-to-optimize distillation method for learning the conditional predictive distribution of a pre-trained dropout model for fast, sample-free uncertainty estimation in computer vision tasks. We empirically test the effectiveness of the proposed method on both semantic segmentation and depth estimation tasks and demonstrate our method can significantly reduce the inference time, enabling real-time uncertainty quantification, while achieving improved quality of both the uncertainty estimates and predictive performance over the regular dropout model.

cs.CV

A Recurrent Ising Machine in a Photonic Integrated Circuit

Conventional computing architectures have no known efficient algorithms for combinatorial optimization tasks, which are encountered in fundamental areas and real-world practical problems including logistics, social networks, and cryptography. Physical machines have recently been proposed and implemented as an alternative to conventional exact and heuristic solvers for the Ising problem, one such optimization task that requires finding the ground state spin configuration of an arbitrary Ising graph. However, these physical approaches usually suffer from decreased ground state convergence probability or universality for high edge-density graphs or arbitrary graph weights, respectively. We experimentally demonstrate a proof-of-principle integrated nanophotonic recurrent Ising sampler (INPRIS) capable of converging to the ground state of various 4-spin graphs with high probability. The INPRIS exploits experimental physical noise as a resource to speed up the ground state search. By injecting additional extrinsic noise during the algorithm iterations, the INPRIS explores larger regions of the phase space, thus allowing one to probe noise-dependent physical observables. Since the recurrent photonic transformation that our machine imparts is a fixed function of the graph problem, and could thus be implemented with optoelectronic architectures that enable GHz clock rates (such as passive or non-volatile photonic circuits that do not require reprogramming at each iteration), our work paves a way for orders-of-magnitude speedups in exploring the solution space of combinatorially hard problems.

physics.optics

Ambulatory Atrial Fibrillation Monitoring Using Wearable Photoplethysmography with Deep Learning

We develop an algorithm that accurately detects Atrial Fibrillation (AF) episodes from photoplethysmograms (PPG) recorded in ambulatory free-living conditions. We collect and annotate a dataset containing more than 4000 hours of PPG recorded from a wrist-worn device. Using a 50-layer convolutional neural network, we achieve a test AUC of 95% and show robustness to motion artifacts inherent to PPG signals. Continuous and accurate detection of AF from PPG has the potential to transform consumer wearable devices into clinically useful medical monitoring tools.

physics.med-ph

Heuristic Recurrent Algorithms for Photonic Ising Machines

The inability of conventional electronic architectures to efficiently solve large combinatorial problems motivates the development of novel computational hardware. There has been much effort recently toward developing novel, application-specific hardware, across many different fields of engineering, such as integrated circuits, memristors, and photonics. However, unleashing the true potential of such novel architectures requires the development of featured algorithms which optimally exploit their fundamental properties. We here present the Photonic Recurrent Ising Sampler (PRIS), a heuristic method tailored for parallel architectures that allows for fast and efficient sampling from distributions of combinatorially hard Ising problems. Since the PRIS relies essentially on vector-to-fixed matrix multiplications, we suggest the implementation of the PRIS in photonic parallel networks, which realize these operations at an unprecedented speed. The PRIS provides sample solutions to the ground state of arbitrary Ising models, by converging in probability to their associated Gibbs distribution. By running the PRIS at various noise levels, we probe the critical behavior of universality classes and their critical exponents. In addition to the attractive features of photonic networks, the PRIS relies on intrinsic dynamic noise and eigenvalue dropout to find ground states more efficiently. Our work suggests speedups in heuristic methods via photonic implementations of the PRIS. We also hint at a broader class of (meta)heuristic algorithms derived from the PRIS, such as combined simulated annealing on the noise and eigenvalue dropout levels. Our algorithm can also be implemented in a competitive manner on fast parallel electronic hardware, such as FPGAs and ASICs.

physics.app-ph

Migrating Knowledge between Physical Scenarios based on Artificial Neural Networks

Deep learning is known to be data-hungry, which hinders its application in many areas of science when datasets are small. Here, we propose to use transfer learning methods to migrate knowledge between different physical scenarios and significantly improve the prediction accuracy of artificial neural networks trained on a small dataset. This method can help reduce the demand for expensive data by making use of additional inexpensive data. First, we demonstrate that in predicting the transmission from multilayer photonic film, the relative error rate is reduced by 46.8% (26.5%) when the source data comes from 10-layer (8-layer) films and the target data comes from 8-layer (10-layer) films. Second, we show that the relative error rate is decreased by 22% when knowledge is transferred between two very different physical scenarios: transmission from multilayer films and scattering from multilayer nanoparticles. Finally, we propose a multi-task learning method to improve the performance of different physical scenarios simultaneously in which each task only has a small dataset.

cs.CV