SearcharxivSearch

arXiv subjects

Junmin Chen

Publications and source records attributed to Junmin Chen.

11 recordsLinked to original sources

Kwai Keye-VL-2.0 Technical Report

We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding and agentic intelligence. To address the challenges of ultra-long contexts, information redundancy, and prohibitive computational costs inherent in hour-level videos, Keye-VL-2.0 is the first to adapt DeepSeek Sparse Attention (DSA) to GQA-based multimodal architectures, enabling lossless 256K context processing while capturing critical frames and long-range temporal dependencies. This architecture is underpinned by a highly optimized training and inference infrastructure, including scalable video I/O, heterogeneous ViT-LM parallelism, and custom DSA kernels that significantly maximize throughput and minimize computational overhead. Furthermore, to overcome the algorithmic dilemma of catastrophic forgetting during multi-task alignment, we introduce Cross-Modal Multi-Teacher On-Policy Distillation (MOPD) paired with Context-RL and Video-RL. By distilling dense token-level teacher feedback from on-policy rollouts back into the MoE backbone, which activates only 3B parameters, Keye-VL-2.0 natively empowers advanced agent collaboration across Code, Tool, and Search scenarios with multimodal self-correction. Extensive evaluations across video understanding, temporal grounding, reasoning, STEM, and agent benchmarks demonstrate that Keye-VL-2.0-30B-A3B achieves state-of-the-art performance among models of similar scale, particularly excelling in fine-grained temporal localization on TimeLens and long-video comprehension on Video-MME-v2 and LongVideoBench. We release our model checkpoints to accelerate community progress toward scalable and robust multimodal agentic applications.

cs.CV

OneReason Technical Report

Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.

cs.IR

GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment

We present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR). Existing long-context RL methods often treat data construction as a matter of designing increasingly complex retrieval paths, leading to homogeneous task coverage and reward formulations that inadequately reflect practical long-context requirements. Our work offers two contributions. (1) Capability-oriented data construction with full open release. We openly release a dataset of 23K RLVR samples, the complete construction pipeline, and all training code. Guided by a taxonomy of long-context capabilities, the dataset spans 9 task types, each paired with its natural evaluation metric. It comprises curated open-source samples from established corpora and synthetic samples whose QA pairs are generated from real source documents such as books, academic papers, and multi-turn dialogues. Under the same vanilla GRPO setup, our dataset alone outperforms the closed-source QwenLong-L1.5 dataset. Moreover, our Qwen3-30B-A3B model trained on this data delivers long-context performance comparable to DeepSeek-R1-0528 and Qwen3-235B-A22B-Thinking-2507, suggesting that broader coverage and greater reward diversity substantially benefit long-context capability improvement. (2) TMN-Reweight for heterogeneous multitask optimization. To address optimization challenges from heterogeneous rewards, we propose TMN-Reweight, which combines task-level mean normalization for cross-task reward scale alignment with difficulty-adaptive weighting for more reliable advantage estimation. TMN-Reweight further improves average performance over vanilla GRPO, with general capabilities preserved or improved across reported evaluations.

cs.CL

Kwai Summary Attention Technical Report

Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agentic intelligence and recommendation system. However, the standard softmax attention exhibits quadratic time complexity with respect to sequence length. As the sequence length increases, this incurs substantial overhead in long-context settings, leading the training and inference costs of extremely long sequences deteriorate rapidly. Existing solutions mitigate this issue through two technique routings: i) Reducing the KV cache per layer, such as from the head-level compression GQA, and the embedding dimension-level compression MLA, but the KV cache remains linearly dependent on the sequence length at a 1:1 ratio. ii) Interleaving with KV Cache friendly architecture, such as local attention SWA, linear kernel GDN, but often involve trade-offs among KV Cache and long-context modeling effectiveness. Besides the two technique routings, we argue that there exists an intermediate path not well explored: {Maintaining a linear relationship between the KV cache and sequence length, but performing semantic-level compression through a specific ratio $k$}. This $O(n/k)$ path does not pursue a ``minimum KV cache'', but rather trades acceptable memory costs for complete, referential, and interpretable retention of long distant dependency. Motivated by this, we propose Kwai Summary Attention (KSA), a novel attention mechanism that reduces sequence modeling cost by compressing historical contexts into learnable summary tokens.

cs.CL

Differentiable hybrid force fields support scalable autonomous electrolyte discovery

Autonomous electrolyte discovery demands a computational engine that satisfies a critical trilemma: it must be fast enough for high-throughput screening, accurate enough for quantitative property prediction, and calibratable enough for online refinement. Classical empirical force fields (FFs) are fast but rely on error cancellation, while standard machine learning interatomic potentials (MLIPs) are computationally expensive. In this Perspective, we highlight that differentiable hybrid FFs resolve this trilemma by fusing physically motivated functional forms with neural-network short-range corrections. Grounded in Energy Decomposition Analysis (EDA), state-of-the-art models such as PhyNEO-Electrolyte and ByteFF-Pol achieve zero-shot generalization to bulk phases, delivering throughputs on the order of tens of ns/day (up to $\sim$50 ns/day, depending on model complexity) for 10,000-atom systems. Crucially, their physical skeletons provide a well-conditioned parameter space for differentiable molecular dynamics (dMD). This enables a dual-calibration paradigm: bottom-up \textit{ab initio} parameterization combined with top-down fine-tuning from macroscopic experimental observables. We propose that this architecture meets the requirements of a ``ChemRobot-ready'' digital twin by integrating physics-grounded simulation with experimentally calibratable refinement, thereby enabling closed-loop autonomous electrolyte discovery.

cond-mat.mtrl-sci

Refinement and Performance Benchmark for Range-Separated Water Force Field

In our previous work, we developed a CCSD(T)-level range-separated water force field that combines the power of physics-driven and machine learning models. However, it was found that expensive CCSD(T)/CBS calculations lead to limited number of QM data as well as the missing of force labels, both of which lead to training instability issues. Bulk properties show large variations that cannot be resolved by simply reducing the fitting error in small cluster QM dataset. Such instability in bulk phase simulation is a universal problem in the training of machine learning potentials (MLPs), and is particularly severe at CCSD(T) level of theory.In this work, using our range-separated water model as an example, we aim to overcome these limitations by developing a new training workflow. It is composed by several techniques including: 1. an active learning protocol that ensures more thorough sampling in different temperatures and densities; 2. an intermediate force label technique employing machine learning density functional; and 3. an ensemble knowledge distillation (EKD) method. These techniques significantly stabilize the resulting water model, consistently achieving sub-chemical accuracies in both cluster energies and experimental properties. Benchmarks are carried out for various properties including densities, radial distribution functions (RDFs), dielectric constants, diffusivity, and infrared spectra, all showing state-of-the-art (SOTA) performances and proving the effectiveness of the training protocol.

physics.chem-ph

From Tags to Trees: Structuring Fine-Grained Knowledge for Controllable Data Selection in LLM Instruction Tuning

Effective and controllable data selection is critical for LLM instruction tuning, especially with massive open-source datasets. Existing approaches primarily rely on instance-level quality scores, or diversity metrics based on embedding clusters or semantic tags. However, constrained by the flatness of embedding spaces or the coarseness of tags, these approaches overlook fine-grained knowledge and its intrinsic hierarchical dependencies, consequently hindering precise data valuation and knowledge-aligned sampling. To address this challenge, we propose Tree-aware Aligned Global Sampling (TAGS), a unified framework that leverages a knowledge tree built from fine-grained tags, thereby enabling joint control of global quality, diversity, and target alignment. Using an LLM-based tagger, we extract atomic knowledge concepts, which are organized into a global tree through bottom-up hierarchical clustering. By grounding data instances onto this tree, a tree-aware metric then quantifies data quality and diversity, facilitating effective sampling. Our controllable sampling strategy maximizes tree-level information gain and enforces leaf-level alignment via KL-divergence for specific domains. Extensive experiments demonstrate that TAGS significantly outperforms state-of-the-art baselines. Notably, it surpasses the full-dataset model by \textbf{+5.84\%} using only \textbf{5\%} of the data, while our aligned sampling strategy further boosts average performance by \textbf{+4.24\%}.

cs.CL

A Hybrid Physics-Driven Neural Network Force Field for Liquid Electrolytes

Electrolyte design plays an important role in the development of lithium-ion batteries and sodium-ion batteries. Battery electrolytes feature a large design space composed of different solvents, additives, and salts, which is difficult to explore experimentally. High-fidelity molecular simulation can accurately predict the bulk properties of electrolytes by employing accurate potential energy surfaces, thus guiding the molecule and formula engineering. At present, the overly simplified classic force fields rely heavily on experimental data for fine-tuning, thus its predictive power on microscopic level is under question. In contrast, the newly emerged machine learning interatomic potential (MLIP) can accurately reproduce the ab initio data, demonstrating excellent fitting ability. However, it is still haunted by problems such as low transferrability, insufficient stability in the prediction of bulk properties, and poor training cost scaling. Therefore, it cannot yet be used as a robust and universal tool for the exploration of electrolyte design space. In this work, we introduce a highly scalable and fully bottom-up force field construction strategy called PhyNEO-Electrolyte. It adopts a hybrid physics-driven and data-driven method that relies only on monomer and dimer EDA (energy deomposition analysis) data. With a careful separation of long/short-range and non-bonding/bonding interactions, we rigorously restore the long-range asymptotic behavior, which is critical in the description of electrolyte systems. Through this approach, we significantly improve the data efficiency of MLIP training, allowing us to achieve much larger chemical space coverage using much less data while retaining reliable quantitative prediction power in bulk phase calculations. PhyNEO-electrolyte thus serves as an important tool for future electrolyte optimization.

physics.chem-ph

Ion-modulated structure, proton transfer, and capacitance in the Pt(111)/water electric double layer

The electric double layer (EDL) governs electrocatalysis, energy conversion, and storage, yet its atomic structure, capacitance, and reactivity remain elusive. Here we introduce a machine learning interatomic potential framework that incorporates long-range electrostatics, enabling nanosecond simulations of metal-electrolyte interfaces under applied electric bias with near-quantum-mechanical accuracy. At the benchmark Pt(111)/water and Pt(111)/aqueous KF electrolyte interfaces, we resolve the molecular structure of the EDL, reveal proton-transfer mechanisms underlying anodic water dissociation and the diffusion of ionic water species, and compute differential capacitance. We find that the nominally inert K+ and F- ions, while leaving interfacial water structure largely unchanged, screen bulk fields, slow proton transfer, and generate a prominent capacitance peak near the potential of zero charge. These results show that ion-specific interactions, which are ignored in mean-field models, are central to capacitance and reactivity, providing a molecular basis for interpreting experiments and designing electrolytes.

physics.chem-ph

Foundation Models for Atomistic Simulation of Chemistry and Materials

Given the power of large language and large vision models, it is of profound and fundamental interest to ask if a foundational model based on data and parameter scaling laws and pre-training strategies is possible for learned simulations of chemistry and materials. The scaling of large and diverse datasets and highly expressive architectures for chemical and materials sciences should result in a foundation model that is more efficient and broadly transferable, robust to out-of-distribution challenges, and easily fine-tuned to a variety of downstream observables, when compared to specific training from scratch on targeted applications in atomistic simulation. In this Perspective we aim to cover the rapidly advancing field of machine learned interatomic potentials (MLIP), and to illustrate a path to create chemistry and materials MLIP foundation models at larger scale.

physics.chem-ph

Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models

Recently, Large language models (LLMs) have revolutionized Natural Language Processing (NLP). Pretrained LLMs, due to limited training context size, struggle with handling long token sequences, limiting their performance on various downstream tasks. Current solutions toward long context modeling often employ multi-stage continual pertaining, which progressively increases the effective context length through several continual pretraining stages. However, those approaches require extensive manual tuning and human expertise. In this paper, we introduce a novel single-stage continual pretraining method, Head-Adaptive Rotary Position Encoding (HARPE), to equip LLMs with long context modeling capabilities while simplifying the training process. Our HARPE leverages different Rotary Position Encoding (RoPE) base frequency values across different attention heads and directly trains LLMs on the target context length. Extensive experiments on 4 language modeling benchmarks, including the latest RULER benchmark, demonstrate that HARPE excels in understanding and integrating long-context tasks with single-stage training, matching and even outperforming existing multi-stage methods. Our results highlight that HARPE successfully breaks the stage barrier for training LLMs with long context modeling capabilities.

cs.CL