SearcharxivSearch

arXiv subjects

Chen Lai

Publications and source records attributed to Chen Lai.

6 recordsLinked to original sources

ExecuTorch -- A Unified PyTorch Solution to Run AI Models On-Device

Local execution of AI on edge devices is important for low latency and offline operation. However, deploying models on diverse hardware remains fragmented, often requiring model conversion or complete reimplementation outside the PyTorch ecosystem where the model was originally authored. We introduce ExecuTorch, a unified PyTorch-native deployment framework for edge AI. ExecuTorch enables seamless deployment of machine learning models across heterogeneous compute environments. It scales from embedded microcontrollers to complex system-on-chips (SoCs) with dedicated accelerators, powering devices ranging from wearables and smartphones to large compute clusters. ExecuTorch preserves PyTorch semantics while allowing customization, support for optimizations like quantization, and pluggable execution "backends". These features together enable fast experimentation, allowing researchers to validate deployment behavior entirely within PyTorch, bridging the gap between research and production.

cs.LG

MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes

The paradigm shift in large language models (LLMs) from instinctive responses to chain-of-thought (CoT) reasoning has fueled two prevailing assumptions: (1) reasoning capabilities only emerge in sufficiently large models, and (2) such capabilities require training on massive datasets. While the first assumption has already been challenged by recent sub-billion-parameter reasoning models such as Qwen3-0.6B and DeepSeek distilled variants, the second remains largely unquestioned. In this work, we revisit the necessity of scaling to extremely large corpora (>10T tokens) for reasoning emergence. By carefully curating and resampling open-source datasets that we identify as beneficial under our designed metrics, we demonstrate that strong reasoning abilities can emerge with far less data. Specifically, we show that only ~2T tokens of high-quality data are sufficient, and pre-training with 4.2T tokens on the dataset resampled from these ~2T tokens, followed by a established post-training procedure, enables the development of MobileLLM-R1, a series of sub-billion-parameter reasoning models that substantially outperform prior models trained on fully open-sourced data. For example, MobileLLM-R1-950M achieves an AIME score of 15.5, compared to just 0.6 for OLMo-2-1.48B and 0.3 for SmolLM-2-1.7B. Remarkably, despite being trained on only 11.7% of the tokens compared to Qwen3's proprietary 36T-token corpus for pretraining, MobileLLM-R1-950M matches or surpasses Qwen3-0.6B across multiple reasoning benchmarks. To facilitate further research in this direction, we have made the models (https://huggingface.co/collections/facebook/mobilellm-r1) and code (https://github.com/facebookresearch/MobileLLM-R1) publicly available, along with the complete training recipe, data sources, and data mixing ratios.

cs.CL

MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

This paper addresses the growing need for efficient large language models (LLMs) on mobile devices, driven by increasing cloud costs and latency concerns. We focus on designing top-quality LLMs with fewer than a billion parameters, a practical choice for mobile deployment. Contrary to prevailing belief emphasizing the pivotal role of data and parameter quantity in determining model quality, our investigation underscores the significance of model architecture for sub-billion scale LLMs. Leveraging deep and thin architectures, coupled with embedding sharing and grouped-query attention mechanisms, we establish a strong baseline network denoted as MobileLLM, which attains a remarkable 2.7%/4.3% accuracy boost over preceding 125M/350M state-of-the-art models. Additionally, we propose an immediate block-wise weight-sharing approach with no increase in model size and only marginal latency overhead. The resultant models, denoted as MobileLLM-LS, demonstrate a further accuracy enhancement of 0.7%/0.8% than MobileLLM 125M/350M. Moreover, MobileLLM model family shows significant improvements compared to previous sub-billion models on chat benchmarks, and demonstrates close correctness to LLaMA-v2 7B in API calling tasks, highlighting the capability of small models for common on-device use cases.

cs.LG

Sub-40nm Nanogratings Self-Organized in PVP-based Polymer Composite Film by Photoexcitation and Two Sequent Splitting under Femtosecond Laser Irradiation

Laser-induced periodic surface structures (LIPSSs) on various materials have been extensively investigated because of their wide applications. The combination of different materials allows for greater freedom in tailoring their functions and achieving responses not possible in a homogeneous material. By utilizing a femtosecond (fs) laser to irradiate the Fe-doped Polyvinyl Pyrrolidone (PVP) composite film, highly regular ultrafine nanogratings (U-nanogratings) with a period as small as 35.0 ($\pm$ 2.0) nm can be self-organized on the surface with extremely high efficiency. The period of the U-nanogratings can be controlled by varying the scanning speed of the laser beam (deposited energy) and the thickness of the composite film. Based on the experimental, theoretical, and simulation results, we propose a two-step formation mechanism: composite film excitation and two sequent grating-splitting. The high photosensitivity and low glass transition temperature of the composite film facilitate the fabrication of the ultrafine nanostructures. The proposed design method for the composite material and fabrication process could not only provide a strategy for obtaining highly regular U-nanogratings, but also offer a platform to explore the interaction physics between ultra-short pulses and matter under extreme conditions.

physics.optics

Recoil-sensitive lithium interferometer without a subrecoil sample

We report simultaneous conjugate Ramsey-Bordé interferometers with a sample of low-mass (lithium-7) atoms at 50 times the recoil temperature. We optically pump the atoms to a magnetically insensitive state using the $2S_{1/2} - 2P_{1/2}$ line. Fast stimulated Raman beam splitters address a broad velocity class and unavoidably drive two conjugate interferometers that overlap spatially. We show that detecting the summed interference signals of both interferometers, using state labeling, allows recoil measurements and suppression of phase noise from vibrations. The use of "warm" atoms allows for simple, efficient, and high-flux atom sources and broadens the applicability of recoil-sensitive interferometry to particles that remain difficult to trap and cool.

physics.atom-ph

High spatial frequency periodic structures induced on ferric ion-doped Polyvinyl Pyrrolidone film by femtosecond laser pulses

Utilizing continues-wave or pulsed laser to induce nano-structures on various material surfaces is one significant method in nano-fabrication technology. In this report, we investigate the formation of high spatial frequency periodic structures on Polyvinyl Pyrrolidone (PVP) film by a linearly polarized femtosecond laser. Ferric (Fe) ions are introduced into the film to improve the photosensitivity. Regular nano-gratings with spatial periods at the range of 60-100nm, which are about one tenth of the irradiating wavelength, can be induced. The period direction of the nano-gratings is perpendicular to the polarization of the femtosecond laser. By tuning the laser energy and scanning speed, we find that the nano-gratings can be formed in a wide range of experimental parameters. As high laser energy can excite not only metals, but also semiconductors and polymers, we believe the formation of the nano-gratings is due to the interaction between the incident femtosecond laser and surface plasmons. The laser processable PVP-based materials and the induced nano-gratings will have potential applications in biophotonics and nanophotonics.

physics.optics