SearcharxivSearch

arXiv subjects

Jiaqiang Liu

Publications and source records attributed to Jiaqiang Liu.

6 recordsLinked to original sources

OneLA: Scaling Linear-Attention Decoding to Large Beams in Generative Recommendation

Generative recommendation (GR) relies on large-beam decoding to generate hundreds of candidate items, creating a new scaling challenge for recurrent linear attention. Existing linear attention serving systems either materialize a full recurrent state for every beam or repeatedly replay shared history, incurring substantial memory and traffic overhead. To address this, we present OneLA, a linear-attention decoding framework that exploits the shared prompt and short divergent suffixes of GR workloads. Specifically, OneLA represents all beam states using a single shared prompt-derived state and compact, append-only records of their divergent transitions. Using this representation, OneLA computes only the state information required at each decoding step, without reconstructing a full recurrent state for every beam. Furthermore, OneLA uses a lightweight ancestry index to track the transition records that make up each beam's history, allowing beams to be updated without moving or copying existing records. A fused GPU kernel further reuses the shared state across beams. Our analysis shows that OneLA achieves 1.54-2.46x end-to-end decode speedups while substantially reducing recurrent-state memory use and data movement.

cs.AI

Quantized Inference for OneRec-V2

Quantized inference has demonstrated substantial system-level benefits in large language models while preserving model quality. In contrast, reliably applying low-precision quantization to recommender systems remains challenging in industrial settings. This difficulty arises from differences in training paradigms, architectural patterns, and computational characteristics, which lead to distinct numerical behaviors in weights and activations. Traditional recommender models often exhibit high-magnitude and high-variance weights and activations, making them more sensitive to quantization-induced perturbations. In addition, recommendation workloads frequently suffer from limited hardware utilization, limiting the practical gains of low-precision computation. In this work, we revisit low-precision inference in the context of generative recommendation. Through empirical distribution analysis, we show that the weight and activation statistics of OneRec-V2 are significantly more controlled and closer to those of large language models than traditional recommendation models. Moreover, OneRec-V2 exhibits a more compute-intensive inference pattern with substantially higher hardware utilization, enabling more end-to-end throughput gains with low-precision computation. Leveraging this property, we develop a FP8 post training quantization framework and integrate it into an optimized inference infrastructure. The proposed joint optimization achieves a 49\% reduction in end-to-end inference latency and a 92\% increase in throughput. Extensive online A/B testing further confirms that FP8 inference introduces no degradation in core metrics. These results suggest that as recommender systems evolve toward the paradigms of large language models, algorithm-level and system-level optimization techniques established in the LLM domain can be effectively adapted to large-scale recommendation workloads.

cs.IR

OneRec-V2 Technical Report

Recent breakthroughs in generative AI have transformed recommender systems through end-to-end generation. OneRec reformulates recommendation as an autoregressive generation task, achieving high Model FLOPs Utilization. While OneRec-V1 has shown significant empirical success in real-world deployment, two critical challenges hinder its scalability and performance: (1) inefficient computational allocation where 97.66% of resources are consumed by sequence encoding rather than generation, and (2) limitations in reinforcement learning relying solely on reward models. To address these challenges, we propose OneRec-V2, featuring: (1) Lazy Decoder-Only Architecture: Eliminates encoder bottlenecks, reducing total computation by 94% and training resources by 90%, enabling successful scaling to 8B parameters. (2) Preference Alignment with Real-World User Interactions: Incorporates Duration-Aware Reward Shaping and Adaptive Ratio Clipping to better align with user preferences using real-world feedback. Extensive A/B tests on Kuaishou demonstrate OneRec-V2's effectiveness, improving App Stay Time by 0.467%/0.741% while balancing multi-objective recommendations. This work advances generative recommendation scalability and alignment with real-world feedback, representing a step forward in the development of end-to-end recommender systems.

cs.IR

OneRec Technical Report

Recommender systems have been widely used in various large-scale user-oriented platforms for many years. However, compared to the rapid developments in the AI community, recommendation systems have not achieved a breakthrough in recent years. For instance, they still rely on a multi-stage cascaded architecture rather than an end-to-end approach, leading to computational fragmentation and optimization inconsistencies, and hindering the effective application of key breakthrough technologies from the AI community in recommendation scenarios. To address these issues, we propose OneRec, which reshapes the recommendation system through an end-to-end generative approach and achieves promising results. Firstly, we have enhanced the computational FLOPs of the current recommendation model by 10 $\times$ and have identified the scaling laws for recommendations within certain boundaries. Secondly, reinforcement learning techniques, previously difficult to apply for optimizing recommendations, show significant potential in this framework. Lastly, through infrastructure optimizations, we have achieved 23.7% and 28.8% Model FLOPs Utilization (MFU) on flagship GPUs during training and inference, respectively, aligning closely with the LLM community. This architecture significantly reduces communication and storage overhead, resulting in operating expense that is only 10.6% of traditional recommendation pipelines. Deployed in Kuaishou/Kuaishou Lite APP, it handles 25% of total queries per second, enhancing overall App Stay Time by 0.54% and 1.24%, respectively. Additionally, we have observed significant increases in metrics such as 7-day Lifetime, which is a crucial indicator of recommendation experience. We also provide practical lessons and insights derived from developing, optimizing, and maintaining a production-scale recommendation system with significant real-world impact.

cs.IR

Radiation Effects on Scientific CMOS Detectors for X-ray Astronomy: II. Total Ionizing Dose Irradiation

Complementary metal-oxide-semiconductor (CMOS) detectors are a competitive choice for current and upcoming astronomical missions. To understand the performance variations of CMOS detectors in space environment, we investigate the total ionizing dose effects on custom-made large-format X-ray CMOS detectors. Three CMOS detector samples were irradiated with a Co-60 source with a total dose of 70 krad and 105 krad. We test and compare the performance of these detectors before and after irradiation. After irradiation, the dark current increases by roughly 20 to 100 times, and the readout noise increases from 3 e- to 6 e-. The bias level at 50 ms integration time decreases by 13 to 18 Digital Number (DN) at -30 degree. The energy resolution increases from about 150 eV to about 170 eV at 4.5 keV at -30 degree. The conversion gain of the detectors varies for less than 2% after the irradiation. Furthermore, there are about 50 pixels whose bias at 50 ms has changed by more than 20 DN after the exposure to the radiation and about 30 to 140 pixels whose readout noise has increased by over 20 e- at -30 degree at 50 ms integration time. These results demonstrate that the performances of large-format CMOS detectors do not suffer significant degeneration in space environment.

astro-ph.IM

Radiation effects on scientific CMOS sensors for X-ray astronomy: I. proton irradiation

Complementary metal-oxide-semiconductor (CMOS) sensors are a competitive choice for future X-ray astronomy missions. Typically, CMOS sensors on space astronomical telescopes are exposed to a high dose of irradiation. We investigate the impact of irradiation on the performance of two scientific CMOS (sCMOS) sensors between -30 to 20 degree at high gain mode (7.5 times), including the bias map, readout noise, dark current, conversion gain, and energy resolution. The two sensors are irradiated with 50 MeV protons with a total dose of 5.3*10^10 p/cm^2. After the exposure, the bias map, readout noise and conversion gain at various temperatures are not significantly degraded, nor is the energy resolution at -30 degree. However, after the exposure the dark current has increased by hundreds of times, and for every 20 degree increase in temperature, the dark current also increases by an order of magnitude. Therefore, at room temperature, the fluctuations of the dark currents dominate the noise and lead to a serious degradation of the energy resolution. Moreover, among the 4k * 4k pixels, there are about 100 pixels whose bias at 50 ms has changed by more than 10 DN (~18 e-), and about 10 pixels whose readout noise has increased by over 15 e- at -30 degree. Fortunately, the influence of the dark current can be reduced by decreasing the integration time, and the degraded pixels can be masked by regular analysis of the dark images. Some future X-ray missions will likely operate at -30 degree, under which the dark current is too small to significantly affect the X-ray performance. Our investigations show the high tolerance of the sCMOS sensors for proton radiation and prove their suitability for X-ray astronomy applications.

astro-ph.IM