Searcharxiv⌕ Search

arXiv subjects

Xiaolong Zhu

Publications and source records attributed to Xiaolong Zhu.

At least 19 recordsLinked to original sources

SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching

Modeling long-term user behavior is central to sequential recommendation and billion-scale industrial recommender systems, yet production ranking models operate under strict latency, memory, communication, and training-throughput constraints. At the 100K scale, the challenge extends beyond attention complexity: raw sequence features must be stored, transferred, and repeatedly processed during training and online serving. Existing approaches based on history truncation, multi-stage behavior retrieval, compressed lifelong histories, or train-short/infer-long extrapolation either weaken end-to-end optimization or retain substantial length-dependent cost. We present SequenceO1, an end-to-end framework for ultra-long user behavior sequence modeling, deployed at full traffic on Douyin with histories of up to 100K interactions. SequenceO1 follows a compress-then-reason design. Its Sketch Attention (SA) uses learnable prototypes and prototype-wise normalization to compress the raw history into a fixed-size, target-agnostic user representation. Target-conditioned Stacked Target-to-History Cross Attention (STCA) then models complementary time scales: a recent 10K suffix for short-term interests and the compact sketch for long-term preferences. To make training and inference practical, SequenceO1 combines low-rank user representation caching, multi-request user-level batching, pipeline lift, and a fused FlashSA kernel to amortize feature storage, communication, and computation across targets, training instances, and consecutive requests. Production experiments show consistent offline and online gains, while the compact cached sketch retains most of the benefit of directly scaling end-to-end sequence ranking to 100K. These results provide a practical model-system approach to efficient attention, sequence compression, and scalable long-sequence and long-context recommendation systems.

cs.IR↗

Remote epitaxy beyond polarity

Remote epitaxy through a monolayer two-dimensional material-covered substrate establishes a crystallographic registry across the van der Waals (vdW) surface that enables the epitaxial growth, lift-off and transfer of single-crystalline films. A central belief in remote epitaxy is that the substrate facilitating the phenomenon must be a material with strong ionicity, as the interatomic electrostatic potential fluctuation in covalent and metallic materials is substantially attenuated by two-dimensional materials. Here, we show remote epitaxy is possible when the substrate is a metallic or covalently bonded material and experimentally demonstrate non-polar remote homo- and heteroepitaxy across a wide range of material systems, including both metals and semiconductors. The achieved non-polar remote interactions are designed and engineered by harnessing substrate conductivity and vicinal surface step-edge density. These findings indicate that remote epitaxy is universal and applicable to ionic, metallic, and covalent materials, expanding its capabilities and stimulating a plethora of new fundamental scientific questions about the mechanism of remote epitaxy.

cond-mat.mtrl-sci↗

State-resolved electron capture in low-energy Ar2+-Ar/N2 collisions

As a fundamental process in atomic physics, charge exchange relies on quantum state-resolved data that is crucial for various fields such as astrophysics and plasma physics. However, there remains a g in the research on multi-electron target systems. This study aims to investigate the dynamic mechanisms of single/double electron capture in collisions between Ar2+ ions and Ar atoms or N2 molecules at an energy of 40 keV, thereby supplementing high-precision experimental data in this field. The experiment is conducted on the electron beam ion source (EBIS) platform at the Institute of Modern Physics, Chinese Academy of Sciences, using the cold target recoil ion momentum spectroscopy (COLTRIMS) technique. An ion beam containing ground-state Ar2+ (3s^2 3p^(4 3) P) and metastable Ar2+ (3s^2 3p^(4 1) D,(_^1)S) is used as the projectile, colliding with a supersonic Ar/ N2 mixed gas target. Three-dimensional momentum of recoil ions is reconstructed through coincidence measurements of recoil ions and scattered ions, and the Q-value and scattering angle distribution are calculated. Theoretical comparisons are performed using the molecular Coulombic over barrier model (MCBM).

physics.atom-ph↗

Results of the NeurIPS 2023 Neural MMO Competition on Multi-task Reinforcement Learning

We present the results of the NeurIPS 2023 Neural MMO Competition, which attracted over 200 participants and submissions. Participants trained goal-conditional policies that generalize to tasks, maps, and opponents never seen during training. The top solution achieved a score 4x higher than our baseline within 8 hours of training on a single 4090 GPU. We open-source everything relating to Neural MMO and the competition under the MIT license, including the policy weights and training code for our baseline and for the top submissions.

cs.LG↗

Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model

Using reinforcement learning with human feedback (RLHF) has shown significant promise in fine-tuning diffusion models. Previous methods start by training a reward model that aligns with human preferences, then leverage RL techniques to fine-tune the underlying models. However, crafting an efficient reward model demands extensive datasets, optimal architecture, and manual hyperparameter tuning, making the process both time and cost-intensive. The direct preference optimization (DPO) method, effective in fine-tuning large language models, eliminates the necessity for a reward model. However, the extensive GPU memory requirement of the diffusion model's denoising process hinders the direct application of the DPO method. To address this issue, we introduce the Direct Preference for Denoising Diffusion Policy Optimization (D3PO) method to directly fine-tune diffusion models. The theoretical analysis demonstrates that although D3PO omits training a reward model, it effectively functions as the optimal reward model trained using human feedback data to guide the learning process. This approach requires no training of a reward model, proving to be more direct, cost-effective, and minimizing computational overhead. In experiments, our method uses the relative scale of objectives as a proxy for human preference, delivering comparable results to methods using ground-truth rewards. Moreover, D3PO demonstrates the ability to reduce image distortion rates and generate safer images, overcoming challenges lacking robust reward models. Our code is publicly available at https://github.com/yk7333/D3PO.

cs.LG↗

The NeurIPS 2022 Neural MMO Challenge: A Massively Multiagent Competition with Specialization and Trade

In this paper, we present the results of the NeurIPS-2022 Neural MMO Challenge, which attracted 500 participants and received over 1,600 submissions. Like the previous IJCAI-2022 Neural MMO Challenge, it involved agents from 16 populations surviving in procedurally generated worlds by collecting resources and defeating opponents. This year's competition runs on the latest v1.6 Neural MMO, which introduces new equipment, combat, trading, and a better scoring system. These elements combine to pose additional robustness and generalization challenges not present in previous competitions. This paper summarizes the design and results of the challenge, explores the potential of this environment as a benchmark for learning methods, and presents some practical reinforcement learning training approaches for complex tasks with sparse rewards. Additionally, we have open-sourced our baselines, including environment wrappers, benchmarks, and visualization tools for future research.

cs.AI↗

Neural MMO 2.0: A Massively Multi-task Addition to Massively Multi-agent Learning

Neural MMO 2.0 is a massively multi-agent environment for reinforcement learning research. The key feature of this new version is a flexible task system that allows users to define a broad range of objectives and reward signals. We challenge researchers to train agents capable of generalizing to tasks, maps, and opponents never seen during training. Neural MMO features procedurally generated maps with 128 agents in the standard setting and support for up to. Version 2.0 is a complete rewrite of its predecessor with three-fold improved performance and compatibility with CleanRL. We release the platform as free and open-source software with comprehensive documentation available at neuralmmo.github.io and an active community Discord. To spark initial research on this new platform, we are concurrently running a competition at NeurIPS 2023.

cs.AI↗

Benchmarking Robustness and Generalization in Multi-Agent Systems: A Case Study on Neural MMO

We present the results of the second Neural MMO challenge, hosted at IJCAI 2022, which received 1600+ submissions. This competition targets robustness and generalization in multi-agent systems: participants train teams of agents to complete a multi-task objective against opponents not seen during training. The competition combines relatively complex environment design with large numbers of agents in the environment. The top submissions demonstrate strong success on this task using mostly standard reinforcement learning (RL) methods combined with domain-specific engineering. We summarize the competition design and results and suggest that, as an academic community, competitions may be a powerful approach to solving hard problems and establishing a solid benchmark for algorithms. We will open-source our benchmark including the environment wrapper, baselines, a visualization tool, and selected policies for further research.

cs.AI↗

Emergent collective intelligence from massive-agent cooperation and competition

Inspired by organisms evolving through cooperation and competition between different populations on Earth, we study the emergence of artificial collective intelligence through massive-agent reinforcement learning. To this end, We propose a new massive-agent reinforcement learning environment, Lux, where dynamic and massive agents in two teams scramble for limited resources and fight off the darkness. In Lux, we build our agents through the standard reinforcement learning algorithm in curriculum learning phases and leverage centralized control via a pixel-to-pixel policy network. As agents co-evolve through self-play, we observe several stages of intelligence, from the acquisition of atomic skills to the development of group strategies. Since these learned group strategies arise from individual decisions without an explicit coordination mechanism, we claim that artificial collective intelligence emerges from massive-agent cooperation and competition. We further analyze the emergence of various learned strategies through metrics and ablation studies, aiming to provide insights for reinforcement learning implementations in massive-agent environments.

cs.AI↗

Multi-Agent Path Finding via Tree LSTM

In recent years, Multi-Agent Path Finding (MAPF) has attracted attention from the fields of both Operations Research (OR) and Reinforcement Learning (RL). However, in the 2021 Flatland3 Challenge, a competition on MAPF, the best RL method scored only 27.9, far less than the best OR method. This paper proposes a new RL solution to Flatland3 Challenge, which scores 125.3, several times higher than the best RL solution before. We creatively apply a novel network architecture, TreeLSTM, to MAPF in our solution. Together with several other RL techniques, including reward shaping, multiple-phase training, and centralized control, our solution is comparable to the top 2-3 OR methods.

cs.AI↗

A Multi-UAV System for Exploration and Target Finding in Cluttered and GPS-Denied Environments

The use of multi-rotor Unmanned Aerial Vehicles (UAVs) for search and rescue as well as remote sensing is rapidly increasing. Multi-rotor UAVs, however, have limited endurance. The range of UAV applications can be widened if teams of multiple UAVs are used. We propose a framework for a team of UAVs to cooperatively explore and find a target in complex GPS-denied environments with obstacles. The team of UAVs autonomously navigates, explores, detects, and finds the target in a cluttered environment with a known map. Examples of such environments include indoor scenarios, urban or natural canyons, caves, and tunnels, where the GPS signal is limited or blocked. The framework is based on a probabilistic decentralised Partially Observable Markov Decision Process which accounts for the uncertainties in sensing and the environment. The team can cooperate efficiently, with each UAV sharing only limited processed observations and their locations during the mission. The system is simulated using the Robotic Operating System and Gazebo. Performance of the system with an increasing number of UAVs in several indoor scenarios with obstacles is tested. Results indicate that the proposed multi-UAV system has improvements in terms of time-cost, the proportion of search area surveyed, as well as successful rates for search and rescue missions.

cs.RO↗

Ultra-compact graphene plasmonic photodetector with the bandwidth over 110GHz

Graphene-based photodetectors, taking advantage of high carrier mobility and broadband absorption in graphene, have recently experienced rapid development. However, their performances with respect to the responsivity and bandwidth are still limited by either weak light-graphene interaction or large resistance-capacitance product. Here, we demonstrate a waveguide coupled integrated graphene plasmonic photodetector on the silicon-on-insulator platform. Benefiting from plasmonic enhanced graphene-light interactions and subwavelength confinement of the optical energy, we present a small-footprint graphene-plasmonic photodetector with bandwidth beyond 110GHz and intrinsic responsivity of 360mA/W. Attributed to the unique electronic bandstructure of graphene and its ultra-broadband absorption, the operational wavelength range extending beyond mid-infrared, and possibly further, can be anticipated. Our results show that the combination of graphene with plasmonic devices has great potential to realize ultra-compact and high-speed optoelectronic devices for graphene-based optical interconnects.

physics.app-ph↗

Holographic Resonant Laser Printing of metasurfaces using plasmonic template

Laser printing with a spatial light modulator (SLM) has several advantages over conventional raster-writing and dot-matrix display (DMD) writing: multiple pixel exposure, high power endurance and existing software for computer generated holograms (CGH). We present a technique for the design and manufacturing of plasmonic metasurfaces based on ultrafast laser printing with an SLM. As a proof of principle, we have used this technique to laser print a plasmonic metalens as well as high resolution plasmonic color decorations. The high throughput holographic resonant laser printing (HRLP) approach enables on-demand mass-production of customized metasurfaces.

cond-mat.mes-hall↗

Effective electro-optic modulation in low-loss graphene-plasmonic slot waveguides

Surface plasmon polaritons enable light concentration within subwavelength regions, opening thereby new avenues for miniaturizing the device and strengthening light-matter interactions. Here we realize effective electro-optic modulation in low-loss plasmonic waveguides with the aid of graphene, and the devices are fully integrated in the silicon-on-insulator platform. By advantageously exploiting low-loss plasmonic slot-waveguide modes, which weakly leak into a substrate while feature strong fields within the two-layer-graphene covered slots in metal, we successfully achieve a tunability of 0.13 dB/um for our fabricated graphene-plasmonic waveguide devices with extremely low insertion loss, which outperforms previously reported graphene-plasmonic devices. Our results highlight the potential of graphene plasmonic leaky-mode hybrid waveguides to realized active ultra-compact devices for optoelectronic applications.

physics.optics↗

Slow-light-enhanced energy efficiency for the graphene microheater on silicon photonic crystal waveguides

Slow light has been widely utilized to obtain enhanced nonlinearities, enhanced spontaneous emissions, and increased phase shifts owing to its ability to promote light-matter interactions. By incorporating a graphene microheater on a slow-light silicon photonic crystal waveguide, we experimentally demonstrated an energy-efficient graphene microheater with a tuning efficiency of 1.07 nm/mW and power consumption per free spectral range of 3.99 mW. The rise and decay times (10% to 90%) were only 750 ns and 525 ns, which, to the best of our knowledge, are the fastest reported response times for microheaters in silicon photonics. The corresponding record-low figure of merit of the device was 2.543 nW.s, which is one order of magnitude lower than results reported in previous studies. The influences of the graphene-photonic crystal waveguide interaction length and the shape of the graphene heater were also investigated, providing valuable guidelines for enhancing the graphene microheater tuning efficiency.

physics.optics↗

Graphene-plasmon polaritons: From fundamental properties to potential applications

With the unique possibilities for controlling light in nanoscale devices, graphene plasmonics has opened new perspectives to the nanophotonics community with potential applications in metamaterials, modulators, photodetectors, and sensors. This paper briefly reviews the recent exciting progress in graphene plasmonics. We begin with a general description for optical properties of graphene, particularly focusing on the dispersion of graphene-plasmon polaritons. The dispersion relation of graphene-plasmon polaritons of spatially extended graphene is expressed in terms of the local response limit with intraband contribution. With this theoretical foundation of graphene-plasmon polaritons, we then discuss recent exciting progress, paying specific attention to the following topics: excitation of graphene plasmon polaritons, electron-phonon interactions in graphene on polar substrates, and tunable graphene plasmonics with applications in modulators and sensors. Finally, we seek to address some of the apparent challenges and promising perspectives of graphene plasmonics.

physics.optics↗

Effective electro-optical modulation with high extinction ratio by a graphene-silicon microring resonator

Graphene opens up for novel optoelectronic applications thanks to its high carrier mobility, ultra-large absorption bandwidth, and extremely fast material response. In particular, the opportunity to control optoelectronic properties through tuning of Fermi level enables electro-optical modulation, optical-optical switching, and other optoelectronics applications. However, achieving a high modulation depth remains a challenge because of the modest graphene-light interaction in the graphene-silicon devices, typically, utilizing only a monolayer or few layers of graphene. Here, we comprehensively study the interaction between graphene and a microring resonator, and its influence on the optical modulation depth. We demonstrate graphene-silicon microring devices showing a high modulation depth of 12.5 dB with a relatively low bias voltage of 8.8 V. On-off electro-optical switching with an extinction ratio of 3.8 dB is successfully demonstrated by applying a square-waveform with a 4 V peak-to-peak voltage.

physics.optics↗

Plasmon-phonon coupling in large-area graphene dot and antidot arrays

Nanostructured graphene on SiO2 substrates pave the way for enhanced light-matter interactions and explorations of strong plasmon-phonon hybridization in the mid-infrared regime. Unprecedented large-area graphene nanodot and antidot optical arrays are fabricated by nanosphere lithography, with structural control down to the sub-100 nanometer regime. The interaction between graphene plasmon modes and the substrate phonons is experimentally demonstrated and structural control is used to map out the hybridization of plasmons and phonons, showing coupling energies of the order 20 meV. Our findings are further supported by theoretical calculations and numerical simulations.

cond-mat.mes-hall↗