SearcharxivSearch

arXiv subjects

Boyang Ding

Publications and source records attributed to Boyang Ding.

11 recordsLinked to original sources

OneReason Technical Report

Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.

cs.IR

Kwai Summary Attention Technical Report

Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agentic intelligence and recommendation system. However, the standard softmax attention exhibits quadratic time complexity with respect to sequence length. As the sequence length increases, this incurs substantial overhead in long-context settings, leading the training and inference costs of extremely long sequences deteriorate rapidly. Existing solutions mitigate this issue through two technique routings: i) Reducing the KV cache per layer, such as from the head-level compression GQA, and the embedding dimension-level compression MLA, but the KV cache remains linearly dependent on the sequence length at a 1:1 ratio. ii) Interleaving with KV Cache friendly architecture, such as local attention SWA, linear kernel GDN, but often involve trade-offs among KV Cache and long-context modeling effectiveness. Besides the two technique routings, we argue that there exists an intermediate path not well explored: {Maintaining a linear relationship between the KV cache and sequence length, but performing semantic-level compression through a specific ratio $k$}. This $O(n/k)$ path does not pursue a ``minimum KV cache'', but rather trades acceptable memory costs for complete, referential, and interpretable retention of long distant dependency. Motivated by this, we propose Kwai Summary Attention (KSA), a novel attention mechanism that reduces sequence modeling cost by compressing historical contexts into learnable summary tokens.

cs.CL

Kelix Technical Report

Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which unifies comprehension and generation under self-supervision. Extending this paradigm to multimodal data requires a shared, discrete representation across modalities. However, most vision-language models (VLMs) still rely on a hybrid interface: discrete text tokens paired with continuous Vision Transformer (ViT) features. Because supervision is largely text-driven, these models are often biased toward understanding and cannot fully leverage large-scale self-supervised learning on non-text data. Recent work has explored discrete visual tokenization to enable fully autoregressive multimodal modeling, showing promising progress toward unified understanding and generation. Yet existing discrete vision tokens frequently lose information due to limited code capacity, resulting in noticeably weaker understanding than continuous-feature VLMs. We present Kelix, a fully discrete autoregressive unified model that closes the understanding gap between discrete and continuous visual representations.

cs.CV

Kwai Keye-VL 1.5 Technical Report

In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Models (MLLMs). However, video understanding remains a challenging area due to the dynamic and information-dense nature of videos. Existing models struggle with the trade-off between spatial resolution and temporal coverage when processing video content. We present Keye-VL-1.5, which addresses fundamental challenges in video comprehension through three key innovations. First, we introduce a novel Slow-Fast video encoding strategy that dynamically allocates computational resources based on inter-frame similarity, processing key frames with significant visual changes at higher resolution (Slow pathway) while handling relatively static frames with increased temporal coverage at lower resolution (Fast pathway). Second, we implement a progressive four-stage pre-training methodology that systematically extends the model's context length from 8K to 128K tokens, enabling processing of longer videos and more complex visual content. Third, we develop a comprehensive post-training pipeline focusing on reasoning enhancement and human preference alignment, incorporating a 5-step chain-of-thought data construction process, iterative GSPO-based reinforcement learning with progressive prompt hinting for difficult cases, and alignment training. Through extensive evaluation on public benchmarks and rigorous internal human assessment, Keye-VL-1.5 demonstrates significant improvements over existing models, particularly excelling in video understanding tasks while maintaining competitive performance on general multimodal benchmarks.

cs.CV

Characterizing the Observational Properties of the Sun's High-latitude m=1 Inertial Mode

Low-m inertial modes have been recently discovered in the Sun's high-latitude regions. In this study, we characterize the observational properties of the m = 1 mode by analyzing time-distance subsurface flow maps. Synoptic flow maps, constructed from daily subsurface flow maps using a tracking rate corresponding to the rotation at latitude 65 degrees, are filtered in both the spherical harmonic and Fourier domains to retain only the m = 1 mode and its dominant frequencies. Our analysis reveals a power distribution that is significantly stronger in the northern polar region. The mode's power exhibits an anti-correlation with solar activity, remaining strong and persistent during the solar activity minimum and becoming weaker and more fragmented during the solar maximum. Magnetic flux transported from low to high latitudes influences both the mode's power and lifetime, enhancing its power and shortening its lifetime upon arrival. The phases of the m = 1 mode in the northern and southern polar regions are near-antisymmetric for most of the time with short deviations. We also compute zonal and meridional phase velocities of the mode and find that it exhibits significantly less differential rotation than its surrounding plasma. The meridional phase velocity, comprising both the local plasma's meridional flow and the mode's intrinsic phase motion, is directed poleward below latitude 70 degrees and equatorward above this latitude. These observational findings underscore the need for a deeper understanding of the internal dynamics of the low-m modes, which may offer valuable insights into the structure and dynamics of the solar interior.

astro-ph.SR

Bandgap Control in Two-Dimensional Semiconductors via Coherent Doping of Plasmonic Hot Electrons

Bandgap control is of central importance for semiconductor technologies. The traditional means of control is to dope the lattice chemically, electrically or optically with charge carriers. Here, we demonstrate for the first time a widely tunable bandgap (renormalisation up to 650 meV at room-temperature) in two-dimensional (2D) semiconductors by coherently doping the lattice with plasmonic hot electrons. In particular, we integrate tungsten-disulfide (WS$_2$) monolayers into a self-assembled plasmonic crystal, which enables coherent coupling between semiconductor excitons and plasmon resonances. Accompanying this process, the plasmon-induced hot electrons can repeatedly fill the WS$_2$ conduction band, leading to population inversion and a significant reconstruction in band structures and exciton relaxations. Our findings provide an innovative and effective measure to engineer optical responses of 2D semiconductors, allowing a great flexiblity in design and optimisation of photonic and optoelectronic devices.

physics.optics

Revealing Strong Plasmon-Exciton Coupling Between Nano-gap Resonators and Two-Dimensional Semiconductors at Ambient Conditions

Strong coupling of two-dimensional semiconductor excitons with plasmonic resonators enables control of light-matter interaction at the subwavelength scale. Here we develop strong coupling in plasmonic nano-gap resonators that allow modification of exciton number contributing to the coupling. Using this system, we not only demonstrate a large vacuum Rabi splitting up to 163 meV and splitting features in photoluminescence spectra, but also reveal that the exciton number can be reduced down to single-digit level (N<10), which is an order lower than that of traditional systems, close to single-exciton based strong coupling. In addition, we prove that the strong coupling process is not affected by the large exciton coherence size that was previously believed to be detrimental to the formation of plasmon-exciton interaction. Our work provides a deeper understanding of storng coupling in two-dimensional semiconductors, paving the way for room temperature quantum optics applications.

physics.optics

Improving Knowledge Graph Embedding Using Simple Constraints

Embedding knowledge graphs (KGs) into continuous vector spaces is a focus of current research. Early works performed this task via simple models developed over KG triples. Recent attempts focused on either designing more complicated triple scoring models, or incorporating extra information beyond triples. This paper, by contrast, investigates the potential of using very simple constraints to improve KG embedding. We examine non-negativity constraints on entity representations and approximate entailment constraints on relation representations. The former help to learn compact and interpretable representations for entities. The latter further encode regularities of logical entailment between relations into their distributed representations. These constraints impose prior beliefs upon the structure of the embedding space, without negative impacts on efficiency or scalability. Evaluation on WordNet, Freebase, and DBpedia shows that our approach is simple yet surprisingly effective, significantly and consistently outperforming competitive baselines. The constraints imposed indeed improve model interpretability, leading to a substantially increased structuring of the embedding space. Code and data are available at https://github.com/iieir-km/ComplEx-NNE_AER.

cs.AI

Leveraging Text and Knowledge Bases for Triple Scoring: An Ensemble Approach - The BOKCHOY Triple Scorer at WSDM Cup 2017

We present our winning solution for the WSDM Cup 2017 triple scoring task. We devise an ensemble of four base scorers, so as to leverage the power of both text and knowledge bases for that task. Then we further refine the outputs of the ensemble by trigger word detection, achieving even better predictive accuracy. The code is available at https://github.com/wsdm-cup-2017/bokchoy.

cs.IR

Plasmonic Gas Sensing based on Cavity-Coupled Metallic Nanoparticles

Here we demonstrate the gas sensing ability of cavity-coupled metallic nanoparticle systems, comprising gold nanoparticles separated from a gold mirror with a polymer spacer. An increase in relative humidity (RH) causes the spacer to expand, which induces a significant reduction of nanoparticle scattering intensity, as the scattering is highly dependent on the cavity-nanoparticle coupling that closely relates to the nanoparticle-mirror distance. This lithography-free structure enables a remarkable averaging sensitivity at 0.12 dB/% RH and 0.25 dB/% RH over RH range (45-75%), possessing an estimated resolution better than 0.5% RH with full reversibility and almost zero-hysteresis, exhibiting notable gas sensing potentials.

physics.optics

Mode Modification of Plasmonic Gap Resonances induced by Strong Coupling with Molecular Excitons

Plasmonic cavities can be used to control the atom-photon coupling process at the nanoscale, since they provide ultrahigh density of optical states in an exceptionally small mode volume. Here we demonstrate strong coupling between molecular excitons and plasmonic resonances (so-called plexcitonic coupling) in a film-coupled nanocube cavity, which can induce profound and significant spectral and spatial modifications to the plasmonic gap modes. Within the spectral span of a single gap mode in the nanotube-film cavity with a 3-nm wide gap, the introduction of narrow-band J-aggregate dye molecules not only enables an anti-crossing behavior in the spectral response, but also splits the single spatial mode into two distinct modes that are easily identified by their far-field scattering profiles. Simulation results confirm the experimental findings and the sensitivity of the plexcitonic coupling is explored using digital control of the gap spacing. Our work opens up a new perspective to study the strong coupling process, greatly extending the functionality of nanophotonic systems, with the potential to be applied in cavity quantum electrodynamic systems.

physics.optics