SearcharxivSearch

arXiv subjects

Yuzhi Wang

Publications and source records attributed to Yuzhi Wang.

At least 19 recordsLinked to original sources

Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture Slides

Generating mind maps from lecture slides can help learners efficiently assimilate fragmented knowledge, promising substantial benefits for intelligent education. However, dedicated automatic generation and evaluation frameworks remain underexplored and challenging, requiring a global-local knowledge focus balance and handling large-scale, heterogeneous slides. We formulate the Slides2MindMap task, which aims to reconstruct cognitively efficient knowledge hierarchies from a course's slide deck collection. For systematic evaluation, we introduce S2M-Bench, a benchmark comprising 12,774 slide pages with expert-annotated mind maps spanning 24 university courses. S2M-Bench includes a cognitive-science-grounded evaluation framework that integrates ground-truth-based comparison, structure conformity analysis, and VLM-as-a-Judge. To address this task, we propose AutoMindMap, an agentic framework inspired by the Structure Building Framework. AutoMindMap comprises Skeleton Laying for global scaffold anchoring, Iterative Knowledge Integration augmented by context-aware summarization, and Dual-Stage Refinement with a local-global decoupling mechanism. The framework reconciles local knowledge faithfulness with global coherence, and adapts to slide-specific features. Experiments on S2M-Bench demonstrate that AutoMindMap outperforms baselines and achieves superior robustness across different models and scenarios, underscoring its pedagogical application value.

cs.AI

LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop Participatory Urban Planning

Participatory Urban Planning (PUP) is increasingly supported by LLM-based agents, yet existing methods largely rely on static preference elicitation and one-shot stakeholder discussions, overlooking the cyclical nature of real-world planning, where residential life, experience collection, and plan adjustment continually interact. We propose Living-in-the-loop Participatory Urban Planning (LiPUP), a closed-loop paradigm that alternates between simulated residential living and experience-driven plan revision, while posing two key challenges: grounding scattered living experience in concrete urban contexts and translating subjective feedback into spatially coherent planning actions. To instantiate LiPUP, we introduce LiPUP-MA, an LLM-based multi-agent framework that constructs a Plan-centric Graph-based Experience Bank to organize urban-grounded residential feedback from living simulation and equips a Spatially-constrained Skill-augmented Planner agent to revise plans by harmonizing experiential, visual, and geospatial evidence. Experiments show that LiPUP-MA consistently outperforms baselines on both conventional static planning metrics and living-based metrics, while iterative LiPUP cycles further improve plan quality.

cs.AI

Attention Residuals

Residual connections with PreNorm are standard in modern LLMs, yet they accumulate all layer outputs with fixed unit weights. This uniform aggregation causes uncontrolled hidden-state growth with depth, progressively diluting each layer's contribution. We propose Attention Residuals (AttnRes), which replaces this fixed accumulation with softmax attention over preceding layer outputs, allowing each layer to selectively aggregate earlier representations with learned, input-dependent weights. To address the memory and communication overhead of attending over all preceding layer outputs for large-scale model training, we introduce Block AttnRes, which partitions layers into blocks and attends over block-level representations, reducing the memory footprint while preserving most of the gains of full AttnRes. Combined with cache-based pipeline communication and a two-phase computation strategy, Block AttnRes becomes a practical drop-in replacement for standard residual connections with minimal overhead. Scaling law experiments confirm that the improvement is consistent across model sizes, and ablations validate the benefit of content-dependent depth-wise selection. We further integrate AttnRes into the Kimi Linear architecture (48B total / 3B activated parameters) and pre-train on 1.4T tokens, where AttnRes mitigates PreNorm dilution, yielding more uniform output magnitudes and gradient distribution across depth, and improves downstream performance across all evaluated tasks.

cs.CL

Learning Physics-Informed Noise Models from Dark Frames for Low-Light Raw Image Denoising

Recently, the mainstream practice for training low-light raw image denoising methods has shifted towards employing synthetic data. Noise modeling, which focuses on characterizing the noise distribution of real-world sensors, profoundly influences the effectiveness and practicality of synthetic data. Currently, physics-based noise modeling struggles to characterize the entire real noise distribution, while learning-based noise modeling impractically depends on paired real data. In this paper, we propose a novel strategy: learning the noise model from dark frames instead of paired real data, to break down the data dependency. Based on this strategy, we introduce an efficient physics-informed noise neural proxy (PNNP) to approximate the real-world sensor noise model. Specifically, we integrate physical priors into neural proxies and introduce three efficient techniques: physics-guided noise decoupling (PND), physics-aware proxy model (PPM), and differentiable distribution loss (DDL). PND decouples the dark frame into different components and handles different levels of noise flexibly, which reduces the complexity of noise modeling. PPM incorporates physical priors to constrain the synthetic noise, which promotes the accuracy of noise modeling. DDL provides explicit and reliable supervision for noise distribution, which promotes the precision of noise modeling. PNNP exhibits powerful potential in characterizing the real noise distribution. Extensive experiments on public datasets demonstrate superior performance in practical low-light raw image denoising. The source code will be publicly available at the project homepage.

eess.IV

Kimi Linear: An Expressive, Efficient Attention Architecture

We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

cs.CL

General First-Principles Approach to Crystals in Finite Magnetic Fields

We introduce a general first-principles methodology for computing electronic structure in a finite uniform magnetic field which allows for an arbitrary rational magnetic flux and nonlocal pseudopotentials, at a comparable time complexity of conventional plane-wave pseudopotential approaches in zero-field conditions. The versatility of this method is demonstrated through comprehensive applications to both molecular and crystalline systems, including calculations of magnetizabilities, magnetically induced currents, and magnetic energy bands. Furthermore, we provide rigorous proofs of two properties for crystals in uniform magnetic fields: the "strong translational symmetry" and "magnetic bands shift" phenomena.

cond-mat.mtrl-sci

Topological nontrivial berry phase in altermagnet CrSb

The study of topological properties in magnetic materials has long been one of the forefront research areas in condensed matter physics. CrSb, as a prototypical candidate material for altermagnetism, has attracted significant attention due to its unique magnetic properties. This system provides a novel platform for exploring the intrinsic relationship between altermagnetic order and exotic topological states. In this study, we combine systematic electrical transport experiments with first-principles calculations to investigate the possible realization mechanisms of topological semimetal states in CrSb and their manifestations in quantum transport phenomena. Our high field magneto-transport measurements reveal that the magnetoresistance of CrSb exhibits no sign of saturation up to 35 T, following a distinct power-law dependence with an exponent of 1.48. The nonlinear Hall resistivity further indicates a multiband charge transport mechanism. Under high magnetic fields, we observe pronounced Shubnikov-de Haas (SdH) quantum oscillations and discernible Zeeman-effect-induced band splitting at 1.6 K. Systematic Fermi surface and band calculations combined with Berry phase analysis confirm the nontrivial topological character of this material (with a Berry phase approaching π). These findings not only provide crucial experimental evidence for understanding the electronic structure of CrSb, but also establish an important foundation for investigating topological quantum states in altermagnets.

cond-mat.mtrl-sci

Kimi-VL Technical Report

We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong agent capabilities - all while activating only 2.8B parameters in its language decoder (Kimi-VL-A3B). Kimi-VL demonstrates strong performance across challenging domains: as a general-purpose VLM, Kimi-VL excels in multi-turn agent tasks (e.g., OSWorld), matching flagship models. Furthermore, it exhibits remarkable capabilities across diverse challenging vision language tasks, including college-level image and video comprehension, OCR, mathematical reasoning, and multi-image understanding. In comparative evaluations, it effectively competes with cutting-edge efficient VLMs such as GPT-4o-mini, Qwen2.5-VL-7B, and Gemma-3-12B-IT, while surpassing GPT-4o in several key domains. Kimi-VL also advances in processing long contexts and perceiving clearly. With a 128K extended context window, Kimi-VL can process diverse long inputs, achieving impressive scores of 64.5 on LongVideoBench and 35.1 on MMLongBench-Doc. Its native-resolution vision encoder, MoonViT, further allows it to see and understand ultra-high-resolution visual inputs, achieving 83.2 on InfoVQA and 34.5 on ScreenSpot-Pro, while maintaining lower computational cost for common tasks. Building upon Kimi-VL, we introduce an advanced long-thinking variant: Kimi-VL-Thinking-2506. Developed through long chain-of-thought (CoT) supervised fine-tuning (SFT) and reinforcement learning (RL), the latest model exhibits strong long-horizon reasoning capabilities (64.0 on MMMU, 46.3 on MMMU-Pro, 56.9 on MathVision, 80.1 on MathVista, 65.2 on VideoMMMU) while obtaining robust general abilities. Code and models are publicly accessible at https://github.com/MoonshotAI/Kimi-VL.

cs.CV

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learning (RL) unlocks a new axis for the continued improvement of artificial intelligence, with the promise that large language models (LLMs) can scale their training data by learning to explore with rewards. However, prior published work has not produced competitive results. In light of this, we report on the training practice of Kimi k1.5, our latest multi-modal LLM trained with RL, including its RL training techniques, multi-modal data recipes, and infrastructure optimization. Long context scaling and improved policy optimization methods are key ingredients of our approach, which establishes a simplistic, effective RL framework without relying on more complex techniques such as Monte Carlo tree search, value functions, and process reward models. Notably, our system achieves state-of-the-art reasoning performance across multiple benchmarks and modalities -- e.g., 77.5 on AIME, 96.2 on MATH 500, 94-th percentile on Codeforces, 74.9 on MathVista -- matching OpenAI's o1. Moreover, we present effective long2short methods that use long-CoT techniques to improve short-CoT models, yielding state-of-the-art short-CoT reasoning results -- e.g., 60.8 on AIME, 94.6 on MATH500, 47.3 on LiveCodeBench -- outperforming existing short-CoT models such as GPT-4o and Claude Sonnet 3.5 by a large margin (up to +550%).

cs.AI

Materials discovery acceleration by using condition generative methodology

With the rapid advancement of AI technologies, generative models have been increasingly employed in the exploration of novel materials. By integrating traditional computational approaches such as density functional theory (DFT) and molecular dynamics (MD), existing generative models, including diffusion models and autoregressive models, have demonstrated remarkable potential in the discovery of novel materials. However, their efficiency in goal-directed materials design remains suboptimal. In this work we developed a highly transferable, efficient and robust conditional generation framework, PODGen, by integrating a general generative model with multiple property prediction models. Based on PODGen, we designed a workflow for the high-throughput crystals conditional generation which is used to search new topological insulators (TIs). Our results show that the success rate of generating TIs using our framework is 5.3 times higher than that of the unconstrained approach. More importantly, while general methods rarely produce gapped TIs, our framework succeeds consistently, highlighting an effectively $\infty$ improvement. This demonstrates that conditional generation significantly enhances the efficiency of targeted material discovery. Using this method, we generated tens of thousands of new topological materials and conducted further first-principles calculations on those with promising application potential. Furthermore, we identified promising, synthesizable topological (crystalline) insulators such as CsHgSb, NaLaB$_{12}$, Bi$_4$Sb$_2$Se$_3$, Be$_3$Ta$_2$Si and Be$_2$W.

cond-mat.mtrl-sci

Kimi-Audio Technical Report

We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio.

eess.AS

Muon is Scalable for LLM Training

Recently, the Muon optimizer based on matrix orthogonalization has demonstrated strong results in training small-scale language models, but the scalability to larger models has not been proven. We identify two crucial techniques for scaling up Muon: (1) adding weight decay and (2) carefully adjusting the per-parameter update scale. These techniques allow Muon to work out-of-the-box on large-scale training without the need of hyper-parameter tuning. Scaling law experiments indicate that Muon achieves $\sim\!2\times$ computational efficiency compared to AdamW with compute optimal training. Based on these improvements, we introduce Moonlight, a 3B/16B-parameter Mixture-of-Expert (MoE) model trained with 5.7T tokens using Muon. Our model improves the current Pareto frontier, achieving better performance with much fewer training FLOPs compared to prior models. We open-source our distributed Muon implementation that is memory optimal and communication efficient. We also release the pretrained, instruction-tuned, and intermediate checkpoints to support future research.

cs.LG

MoBA: Mixture of Block Attention for Long-Context LLMs

Scaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI). However, the quadratic increase in computational complexity inherent in traditional attention mechanisms presents a prohibitive overhead. Existing approaches either impose strongly biased structures, such as sink or window attention which are task-specific, or radically modify the attention mechanism into linear approximations, whose performance in complex reasoning tasks remains inadequately explored. In this work, we propose a solution that adheres to the ``less structure'' principle, allowing the model to determine where to attend autonomously, rather than introducing predefined biases. We introduce Mixture of Block Attention (MoBA), an innovative approach that applies the principles of Mixture of Experts (MoE) to the attention mechanism. This novel architecture demonstrates superior performance on long-context tasks while offering a key advantage: the ability to seamlessly transition between full and sparse attention, enhancing efficiency without the risk of compromising performance. MoBA has already been deployed to support Kimi's long-context requests and demonstrates significant advancements in efficient attention computation for LLMs. Our code is available at https://github.com/MoonshotAI/MoBA.

cs.LG

Simultaneous achievement of record-breaking colossal magnetoresistance and angular magnetoresistance in an antiferromagnetic semiconductor EuSe2

Magnetoresistance effect lays the foundation for spintronics, magnetic sensors and hard drives. The pursuit of magnetic materials with colossal magnetoresistance (CMR) and/or angular magnetoresistance (AMR) has attracted enduring research interest and extensive investigations over past decades. Here we report on the discovery of field-induced record-breaking CMR of ~ -10^14 % and AMR ~ 10^14% achieved simultaneously in an antiferromagnetic rare-earth dichalcogenide EuSe2. Such intriguing observations are attributed to strong magnetic anisotropy and magnetic-field induced antiferromagnetic to ferromagnetic transition of the localized Eu2+ spins, which in turn closes the bandgap by lifting the degeneracy of Se-5p bands near Fermi level. Our DFT calculations perfectly replicate the experimental findings based on the Brillouin function and carries transport model. The present work provides a potential simple antiferromagnetic material for achieving angle-sensitive spintronic devices.

cond-mat.str-el

Scaling Behavior of Magnetoresistance and Hall Resistivity in Altermagnet CrSb

The discovery of altermagnet (AM) marks a significant advancement in magnetic materials, combining characteristics of both ferromagnetism and antiferromagnetism. In this Letter, we focus on CrSb, which has been verified to be an AM and to exhibit substantial spin splitting near the Fermi level. After successfully growing high-quality CrSb single crystals, we performed comprehensive magnetization, magnetoresistance (MR), and Hall resistivity measurements, along with the electronic structure, and Fermi surface (FS) calculations, as well as the magneto-transport property numerical simulations. An antiferromagnetic transition occurring at $T_{N}$ = 712 K was reconfirmed. It was found that both experimental MR and Hall resistivity are consistent with the numerical simulation results, and exhibit obvious scaling behavior. The nonlinear Hall resistivity is due to its multi-band structure, rather than an anomalous Hall effect (AHE). Especially, the scaling behavior in Hall resistivity is first observed within an AM material. These findings demonstrate that the magneto-transport properties in CrSb originate from the intrinsic electronic structure and are dominated by the Lorentz force.

cond-mat.mtrl-sci

Observation of surface Fermi arcs in altermagnetic Weyl semimetal CrSb

As a special type of collinear antiferromagnetism (AFM), altermagnetism has garnered significant research interest recently. Altermagnets exhibit broken parity-time symmetry and zero net magnetization in real space, leading to substantial band splitting in momentum space even in the absence of spin-orbit coupling. Meanwhile, parity-time symmetry breaking always induce nontrivial band topology such as Weyl nodes. While Weyl semimetal states and nodal lines have been theoretically proposed in altermagnets, rare reports of experimental observation have been made up to this point. Using ARPES and first-principles calculations, we systematically studied the electronic structure of the room-temperature altermagnet candidate CrSb. At generic locations in momentum space, we clearly observed band spin splitting. Furthermore, we identified discrete surface Fermi arcs on the (100) cleaved side surface close to the Fermi level originating from bulk band topology. Our results imply that CrSb contains interesting nontrivial topological Weyl physics, in addition to being an excellent room temperature altermagnet.

cond-mat.mtrl-sci

Moiré Fractional Chern Insulators II: First-principles Calculations and Continuum Models of Rhombohedral Graphene Superlattices

The experimental discovery of fractional Chern insulators (FCIs) in rhombohedral pentalayer graphene twisted on hexagonal boron nitride (hBN) has preceded theoretical prediction. Supported by large-scale first principles relaxation calculations at the experimental twist angle of $0.77^\circ$, we obtain an accurate continuum model of $n=3,4,5,6,7$ layer rhombohedral graphene-hBN moiré systems. Focusing on the pentalayer case, we analytically explain the robust $|C|=0,5$ Chern numbers seen in the low-energy single-particle bands and their flattening with displacement field, making use of a minimal two-flavor continuum Hamiltonian derived from the full model. We then predict nonzero valley Chern numbers at the $ν= -4,0$ insulators observed in experiment. Our analysis makes clear the importance of displacement field and the moiré potential in producing localized "heavy fermion" charge density in the top valence band, in addition to the nearly free conduction band. Lastly, we study doubly aligned devices as additional platforms for moiré FCIs with higher Chern number bands.

cond-mat.mes-hall

Learnability Enhancement for Low-light Raw Denoising: Where Paired Real Data Meets Noise Modeling

Low-light raw denoising is an important and valuable task in computational photography where learning-based methods trained with paired real data are mainstream. However, the limited data volume and complicated noise distribution have constituted a learnability bottleneck for paired real data, which limits the denoising performance of learning-based methods. To address this issue, we present a learnability enhancement strategy to reform paired real data according to noise modeling. Our strategy consists of two efficient techniques: shot noise augmentation (SNA) and dark shading correction (DSC). Through noise model decoupling, SNA improves the precision of data mapping by increasing the data volume and DSC reduces the complexity of data mapping by reducing the noise complexity. Extensive results on the public datasets and real imaging scenarios collectively demonstrate the state-of-the-art performance of our method. Our code is available at: https://github.com/megvii-research/PMN.

cs.CV