SearcharxivSearch

arXiv subjects

Tong Gao

Publications and source records attributed to Tong Gao.

At least 19 recordsLinked to original sources

Kimi K3: Open Frontier Intelligence

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

cs.CL

OpenCompass: A Universal Evaluation Platform for Large Language Models

In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.

cs.CL

Smooth-Rigid-Body Contact as a ReLCP: A Recursively Generated Linear Complementarity Problem

This paper reformulates complementarity-based time-stepping for frictionless nonsmooth contact between smooth rigid bodies as a recursively generated linear complementarity problem (ReLCP), involving a sequence of LCPs of increasing dimension. Starting from a classical single-constraint shared-normal signed-distance (SNSD) LCP, the method adds unilateral constraints only when the discrete-time update predicted by the current contact set would violate nonpenetration of the underlying smooth surfaces. The resulting procedure acts directly on smooth geometry, enforces nonpenetration to a prescribed tolerance, and avoids the oversampling inherent to proxy-surface contact models such as tessellations or multi-sphere decompositions, for which improved geometric fidelity can drive rapid growth in constraint count and cost. For strictly convex bodies, we prove that an initially overlap free configuration with sufficiently small timestep sizes, imply finite termination of the adaptive augmentation, and yield a unique discrete-time velocity update. In the small timestep limit and for any fixed overlap-free discrete state with a fixed geometric overlap tolerance, we prove that the recursion terminates after the initial solve, reducing the method to the classical single-constraint SNSD LCP and retaining the usual consistency of complementarity time-stepping with the underlying differential variational inequality. Numerical tests on colliding ellipsoids, compacting ellipsoid suspensions, growing bacterial colonies, and taut chainmail networks demonstrate stable large-timestep behavior, bounded interpenetration without discretization-induced surface roughness, and substantial reductions in both active constraint counts and runtime relative to representative discrete-surface complementarity formulations.

cs.RO

Kimi K2: Open Agentic Intelligence

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.

cs.LG

Kimi K2.5: Visual Agentic Intelligence

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to $4.5\times$ over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.

cs.CL

Computational modeling of Pulsed Field Ablation for pulmonary vein isolation

Pulsed field ablation (PFA) has emerged as a non-thermal alternative to traditional thermal ablation techniques for the treatment of atrial fibrillation (AF). This study presents a patient-specific 3D computational framework to model the effects of PFA on pulmonary vein isolation (PVI). The modeling framework is rigorously validated against published numerical and experimental data, demonstrating strong agreement across a range of scenarios. Using realistic left atrial (LA) anatomy, commercially available circular, flower, and basket catheter configurations are simulated to evaluate lesion formation across different applied voltages. The performance of each catheter type is quantitatively assessed using multiple metrics, including lesion volume, energy delivery efficiency and transmurality. Simulation results show that circular catheters provide the highest energy delivery efficiency and target coverage at lower voltages, while basket catheters produce the largest lesion volumes. This framework offers a useful basis for exploring catheter design and treatment planning in PFA applications.

physics.med-ph

A Novel RFID Authentication Protocol Based on A Block-Order-Modulus Variable Matrix Encryption Algorithm

In this paper, authentication for mobile radio frequency identification (RFID) systems with low-cost RFID sensor tags is studied. Firstly, an adaptive modulus (AM) encryption algorithm is proposed. Subsequently, in order to enhance the security without additional storage of new key matrices, a self-updating encryption order (SUEO) algorithm is designed. Furthermore, a diagonal block local transpose key matrix (DBLTKM) encryption algorithm is presented, which effectively expands the feasible domain of the key space. Based on the above three algorithms, a novel joint AM-SUEO-DBLTKM encryption algorithm is constructed. Making full use of the advantages of the proposed joint algorithm, a two-way RFID authentication protocol, named AM-SUEO-DBLTKM-RFID, is proposed for mobile RFID systems. In addition, the Burrows-Abadi-Needham (BAN) logic and security analysis indicate that the proposed AM-SUEO-DBLTKM-RFID protocol can effectively combat various typical attacks. Numerical results demonstrate that the proposed AM-SUEO-DBLTKM algorithm can save 99.59% of tag storage over traditional algorithms. Finally, the low computational complexity as well as the low storage cost of the proposed AM-SUEO-DBLTKM-RFID protocol facilitates deployment within low-cost RFID sensor tags.

cs.CR

Scaling behavior and giant field-enhancement of the thermal conductivity in the honeycomb antiferromagnet BaCo2(AsO4)2

The layered honeycomb material BaCo$_2$(AsO$_4$)$_2$ (BCAO) is of topical interest because its magnetic state is related to that of the Kitaev magnet $α$-RuCl$_3$. Using thermal transport to probe how magnetic excitations interact with phonons in the magnetically disordered regime, we have uncovered an unusually large enhancement of the thermal conductivity $κ_{xx}$ in an in-plane magnetic field ${\bf H}$. Just above the Néel temperature $T_{\rm N}$, a field of 13 T increases $κ_{xx}$ by a factor $\sim 211$, much larger than reported previously in any magnetic insulator. Interestingly, $κ_{xx}(H,T)$ exhibits a scaling behavior in the entire magnetically disordered region that surrounds the ordered zigzag state. The ratio $Δκ_{xx}(H,T)/κ_{xx}(13,T)$, measured throughout the disordered region, collapses to a one-parameter scaling function ${\rm exp}(-1/gx)$ (where $x = μ_{\rm B}B/k_{\rm B}T$ and $g$ is a constant).

cond-mat.str-el

Towards Automated Error Analysis: Learning to Characterize Errors

Characterizing the patterns of errors that a system makes helps researchers focus future development on increasing its accuracy and robustness. We propose a novel form of "meta learning" that automatically learns interpretable rules that characterize the types of errors that a system makes, and demonstrate these rules' ability to help understand and improve two NLP systems. Our approach works by collecting error cases on validation data, extracting meta-features describing these samples, and finally learning rules that characterize errors using these features. We apply our approach to VilBERT, for Visual Question Answering, and RoBERTa, for Common Sense Question Answering. Our system learns interpretable rules that provide insights into systemic errors these systems make on the given tasks. Using these insights, we are also able to "close the loop" and modestly improve performance of these systems.

cs.CL

Hydrodynamic instabilities and collective dynamics in activity-balanced pusher-puller mixtures

Microorganisms living in microfluidic environments often form multi-species swarms, where they can leverage collective motions to achieve enhanced transport and spreading. Nevertheless, there is a general lack of physical understandings of the origins of the multiscale unstable dynamics observed within these systems. Here, we build a computational model to study binary suspensions of rear- and front-actuated microswimmers, or respectively the so-called "pusher" and "puller" particles, that have different populations and swimming speeds. We perform direct particle simulations to reveal that collective system dynamics are possible even in the scenario of an "activity-balanced" mixture, which produces near zero mean extra stress. We first construct a continuum kinetic model to describe the initial transient period when the system is near uniform isotropy and then perform linear stability analysis to reveal the system's finite-wavelength hydrodynamic instabilities, in contrast with the long-wavelength instabilities of pure pusher/puller suspensions. Then, we carry out slender-body discrete particle simulations to resolve both the short time instabilities and the the longtime dynamics, which feature non-trivial density fluctuations and spatially-correlated motions, distinct from those of single-species.

physics.flu-dyn

The planar thermal Hall conductivity in the Kitaev magnet α-RuCl3

We report detailed measurements of the Onsager-like planar thermal Hall conductivity $κ_{xy}$ in $α$-RuCl$_3$, a spin-liquid candidate of topical interest. With the thermal current ${\bf J}_{\rm Q}$ and magnetic field $\bf B\parallel a$ (zigzag axis), the observed $κ_{xy}/T$ varies strongly with temperature $T$ (1-10 K). The results are well-described by bosonic edge excitations which evolve to topological magnons at large $B$. Fits to $κ_{xy}/T$ yield a Chern number $\sim 1$ and a band energy $ω_1\sim$1 meV, in agreement with sharp modes seen in electron spin-resonance experiments. The bosonic character is incompatible with half-quantization of $κ_{xy}/T$.

cond-mat.str-el

MMOCR: A Comprehensive Toolbox for Text Detection, Recognition and Understanding

We present MMOCR-an open-source toolbox which provides a comprehensive pipeline for text detection and recognition, as well as their downstream tasks such as named entity recognition and key information extraction. MMOCR implements 14 state-of-the-art algorithms, which is significantly more than all the existing open-source OCR projects we are aware of to date. To facilitate future research and industrial applications of text recognition-related problems, we also provide a large number of trained models and detailed benchmarks to give insights into the performance of text detection, recognition and understanding. MMOCR is publicly released at https://github.com/open-mmlab/mmocr.

cs.CV

Oscillations of the thermal conductivity observed in the spin-liquid state of $α$-RuCl$_3$

In the class of materials called spin liquids, a magnetically ordered state cannot be attained even at milliKelvin temperatures because of conflicting constraints on each spin (for e.g. from geometric or exchange frustration). The resulting quantum spin-liquid (QSL) state is currently of intense interest because it exhibits novel excitations as well as wave-function entanglement. The layered insulator $α$-RuCl$_3$ orders as a zigzag antiferromagnet below $\sim$7 K in zero magnetic field. The zigzag order is destroyed when a magnetic field $\bf H$ is applied parallel to the zigzag axis a. Within the field interval (7.3, 11) Tesla, there is growing evidence that a QSL state exists. Here we report the observation of oscillations in its thermal conductivity below 4 K. The oscillation amplitude is very large within the interval (7.3, 11) T and strongly suppressed on either side. Paradoxically, the oscillations are periodic in 1/\emph{H}, analogous to quantum oscillations in metals, even though $α$-RuCl$_3$ is an excellent insulator with a gap of 1.9 eV. By tilting $\bf H$ out of the plane, we find that the oscillation period is determined by the in-plane component $H_a$. As the temperature is raised above 0.5 K, the oscillation amplitude decreases exponentially. The decrease anticorrelates with the emergence above $\sim$2 K of an anomalous planar thermal Hall conductivity measured with $\bf H\parallel a$. To exclude extrinsic artifacts, we carried out several tests. The implications of the oscillations are discussed.

cond-mat.str-el

Systematic Generalization on gSCAN with Language Conditioned Embedding

Systematic Generalization refers to a learning algorithm's ability to extrapolate learned behavior to unseen situations that are distinct but semantically similar to its training data. As shown in recent work, state-of-the-art deep learning models fail dramatically even on tasks for which they are designed when the test set is systematically different from the training data. We hypothesize that explicitly modeling the relations between objects in their contexts while learning their representations will help achieve systematic generalization. Therefore, we propose a novel method that learns objects' contextualized embeddings with dynamic message passing conditioned on the input natural language and end-to-end trainable with other downstream deep learning modules. To our knowledge, this model is the first one that significantly outperforms the provided baseline and reaches state-of-the-art performance on grounded-SCAN (gSCAN), a grounded natural language navigation dataset designed to require systematic generalization in its test splits.

cs.AI

Weak-field induced nonmagnetic state in a Co-based honeycomb

Layered honeycomb magnets are of interest as potential realizations of the Kitaev quantum spin liquid (KQSL), a quantum state with long-range spin entanglement and an exactly solvable Hamiltonian. Conventional magnetically ordered states are present for all currently known candidate materials, however, because non-Kitaev terms in the Hamiltonians obscure the Kitaev physics. Current experimental studies of the KQSL are focused on 4d- or 5d-transition-metal-based honeycombs, in which strong spin-orbit coupling can be expected, yielding Kitaev interaction that dominate in an applied magnetic field. In contrast, for 3d-based layered honeycomb magnets, spin orbit coupling is weak and thus Kitaev-physics should be substantially less accessible. Here we report our studies on BaCo2(AsO4)2, for which we find that the magnetic order associated with the non-Kitaev interactions can be fully suppressed by a relatively low magnetic field, yielding a non-magnetic material and implying the presence of strong magnetic frustration and weak non-Kitaev interactions.

cond-mat.str-el

High mobility in a van der Waals layered antiferromagnetic metal

Magnetic van der Waals (vdW) materials have been heavily pursued for fundamental physics as well as for device design. Despite the rapid advances, so far magnetic vdW materials are mainly insulating or semiconducting, and none of them possesses a high electronic mobility - a property that is rare in layered vdW materials in general. The realization of a magnetic high-mobility vdW material would open the possibility for novel magnetic twistronic or spintronic devices. Here we report very high carrier mobility in the layered vdW antiferromagnet GdTe3. The electron mobility is beyond 60,000 cm2 V-1 s-1, which is the highest among all known layered magnetic materials, to the best of our knowledge. Among all known vdW materials, the mobility of bulk GdTe3 is comparable to that of black phosphorus, and is only surpassed by graphite. By mechanical exfoliation, we further demonstrate that GdTe3 can be exfoliated to ultrathin flakes of three monolayers, and that the magnetic order and relatively high mobility is retained in approximately 20-nm-thin flakes.

cond-mat.mtrl-sci

Large-Scale Answerer in Questioner's Mind for Visual Dialog Question Generation

Answerer in Questioner's Mind (AQM) is an information-theoretic framework that has been recently proposed for task-oriented dialog systems. AQM benefits from asking a question that would maximize the information gain when it is asked. However, due to its intrinsic nature of explicitly calculating the information gain, AQM has a limitation when the solution space is very large. To address this, we propose AQM+ that can deal with a large-scale problem and ask a question that is more coherent to the current context of the dialog. We evaluate our method on GuessWhich, a challenging task-oriented visual dialog problem, where the number of candidate classes is near 10K. Our experimental results and ablation studies show that AQM+ outperforms the state-of-the-art models by a remarkable margin with a reasonable approximation. In particular, the proposed AQM+ reduces more than 60% of error as the dialog proceeds, while the comparative algorithms diminish the error by less than 6%. Based on our results, we argue that AQM+ is a general task-oriented dialog algorithm that can be applied for non-yes-or-no responses.

cs.CL

A gap-protected zero-Hall effect state in the quantum limit of the nonsymmorphic metal KHgSb

A recurring theme in topological matter is the protection of unusual electronic states by symmetry, for example, protection of the surface states in Z2 topological insulators by time reversal symmetry [1-3]. Recently interest has turned to unusual surface states in the large class of nonsymmorphic materials [4-11]. In particular KHgSb is predicted to exhibit double quantum spin Hall (QSH) states [10]. Here we report observation of a novel feature of the Hall conductivity in KHgSb in strong magnetic field B. In the quantum limit, the Hall conductivity is observed to fall exponentially to zero, but the diagonal conductivity is finite. A large gap protects this unusual zero-Hall state. We propose that, in this limit, the chemical potential drops into the bulk gap, intersecting equal numbers of right and left-moving QSH surface modes to produce the zero-Hall state.

cond-mat.str-el