SearcharxivSearch

arXiv subjects

Haodong Wu

Publications and source records attributed to Haodong Wu.

6 recordsLinked to original sources

From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning

Supervised fine-tuning (SFT) data selection is commonly formulated as instance ranking: score each example and retain a top-$k$ subset. However, effective SFT training subsets are often produced through ordered curation recipes, where filtering, mixing, and deduplication operators jointly shape the final data distribution. We formulate this problem as fixed-pool data recipe search: given a raw instruction pool and a library of grounded operators, the goal is to discover an executable recipe that constructs a high-quality selected subset under a limited budget of full SFT evaluations, without generating, rewriting, or augmenting training samples. We introduce AutoSelection, a two-layer solver that decouples fixed-pool materialization based on cached task-, data-, and model-side signals from expensive full evaluation, using warmup probes, realized subset states, local recipe edits, Gaussian-process-assisted ranking, and stagnation-triggered reseeding. Experiments on a 90K instruction pool show that AutoSelection achieves the strongest in-distribution reasoning average across three base models, outperforming full-data training, random recipe search, random top-$k$, and single-operator selectors. Additional Out-of-distribution graph-reasoning results, search-stability analyses, structural ablations, and 1.5B-to-7B transfer checks further show that recipe structure matters beyond individual selection operators. Code is available at https://github.com/w253/AutoSelection.

cs.LG

Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning ability of Large Language Models (LLMs), but sparse outcome rewards make token-level credit assignment difficult. We study token-level credit as a reward-conditioned shift from the behavior policy to a hindsight posterior. In autoregressive RLVR, this shift can be expressed through Conditional Mutual Information (CMI), which shows that token entropy upper-bounds possible hindsight credit. Entropy, however, indicates capacity rather than update direction, so we introduce the Four Quadrant Decomposition to separate updates by reward polarity and token entropy. Controlled interventions show that these two factors jointly shape token updates. Sustained reasoning gains concentrate in signed high-entropy quadrants, whereas low-entropy updates saturate quickly. Based on this analysis, we propose Hindsight-Aware Policy Optimization (HAPO), a sign-preserving modification to GRPO that performs capacity-guided advantage reallocation. Experiments on mathematical reasoning benchmarks in two model settings show that HAPO achieves competitive performance among entropy-aware baselines.

cs.LG

Reversible optical isolators and quasi-circulators using a magneto-optical Fabry-Pérot cavity

Nonreciprocal optical devices are essential for laser protection, modern optical communication and quantum information processing by enforcing one-way light propagation. The conventional Faraday magneto-optical nonreciprocal devices rely on a strong magnetic field, which is provided by a permanent magnet. As a result, the isolation direction of such devices is fixed and severely restricts their applications in quantum networks.In this work, we experimentally demonstrate the simultaneous one-way transmission and unidirectional reflection by using a magneto-optical Fabry-Pérot cavity and a magnetic field strength of $50~\milli\tesla$. An optical isolator and a three-port quasi-circulator are realized based on this nonreciprocal cavity system. The isolator achieves an isolation ratio of up to $22~\deci\bel$ and an averaged insertion loss down to $0.97~\deci\bel$. The quasi-circulator is realized with a fidelity exceeding $99\%$ and an overall survival probability of $89.9\%$, corresponding to an insertion loss of $\sim 0.46~\deci\bel$. The magnetic field is provided by an electromagnetic coil, thereby allowing for reversing the light circulating path. The reversible quasi-circulator paves the way for building reconfigurable quantum networks.

physics.optics

Generation of True Quantum Random Numbers with On-Demand Probability Distributions via Single-Photon Quantum Walks

Random numbers are at the heart of diverse fields, ranging from simulations of stochastic processes to classical and quantum cryptography. The requirement for true randomness in these applications has motivated various proposals for generating random numbers based on the inherent randomness of quantum systems. The generation of true random numbers with arbitrarily defined probability distributions is highly desirable for applications, but it is very challenging. Here we show that single-photon quantum walks can generate multi-bit random numbers with on-demand probability distributions, when the required ``coin'' parameters are found with the gradient descent (GD) algorithm. Our theoretical and experimental results exhibit high fidelity for various selected distributions. This GD-enhanced single-photon system provides a convenient way for building flexible and reliable quantum random number generators. Multi-bit random numbers are a necessary resource for high-dimensional quantum key distribution.

quant-ph

A passive bias-free ultrabroadband optical isolator based on unidirectional self-induced transparency

Achieving a broadband nonreciprocal device without gain and any external bias is very challenging and highly desirable for modern photonic technologies and quantum networks. Here, we theoretically propose a passive and bias-free all-optical isolator for a femtosecond laser pulse by exploiting a new mechanism of unidirectional self-induced transparency, obtained with a nonlinear medium followed by a normal absorbing medium at one side. The transmission contrast between the forward and backward directions can reach ~14.3 dB for a 2π5 fs laser pulse, implying isolation of a signal with an ultrabroad bandwidth of 200 THz. The 20 dB bandwidth is about 57 nm, already comparable with a magneto-optical isolator. This cavity-free optical isolator may pave the way to integrated nonmagnetic isolation of ultrashort laser pulses.

physics.optics

Towards On-Demand Heralded Single-Photon Sources via Photon Blockade

Spontaneous parametric down-conversion (SPDC) in a laser pumped optical nonlinear medium can produce heralded single photons with a high purity but a very low yield. Improving the yield by increasing the pump power in SPDC inevitably reduces the purity due to excitation of multi-photon events. We propose a scheme to overcome this purity-yield trade-off by suppressing multi-photon events in a cavity-enhanced SPDC via the photon blockade effect. By introducing a strong photon-photon interaction into the intracavity medium and increasing the pump power, we can improve the available single-photon yield to larger than $90\%$, while maintaining a high purity of $99\%$, towards on-demand generation of single photons through the SPDC process. Our quasi-on-demand SPDC sources may boost single-photon-based quantum information technology.

quant-ph