SearcharxivSearch

arXiv subjects

Tongyu Wang

Publications and source records attributed to Tongyu Wang.

9 recordsLinked to original sources

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present daVinci-kernel, a reinforcement learning framework that couples skill discovery with skill exploitation through a dynamically evolving skill library. daVinci-kernel jointly trains three agents sharing one LLM backbone: a Skill Selection Agent that retrieves relevant techniques via BM25 and LLM reranking, a Policy Agent that generates multi-turn CUDA/Triton kernels conditioned on selected skills, and a Skill Summary Agent that distills successful rollouts into reusable skills. Candidate skills are added only after execution-based verification confirms reproducible speedups. All three agents share a single LLM backbone, are initialized via a structured SFT cold start on diversity-filtered data, and are then jointly optimized end-to-end with multi-turn REINFORCE and per-agent advantage estimation. On KernelBench, daVinci-kernel-14B achieves 37.2%, 70.6%, and 32.2% on Level 1, Level 2, and Level 3 under the Fast$_1$ threshold, outperforming the strongest prior RL-trained model, Dr\. Kernel-14B.

cs.LG

DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training

Effectively scaling Reinforcement Learning (RL) is crucial for enhancing the reasoning and alignment of Large Language Models. The massive data and complex execution flows inherent in these tasks require a distributed architecture capable of efficient scaling. However, to simplify programming and dependency management, mainstream frameworks often rely on a centralized architecture where a single node dispatches both control and data. This inherent coupling creates significant communication bottlenecks, severely limiting system scalability and efficiency. We present DISTFLOW, a novel, fully distributed RL framework that adopts a multi-controller paradigm. By decoupling data transmission from control dispatch, DISTFLOW establishes a parallelism-aware, decentralized Data Coordinator that leverages local caching, load balancing, and asynchronous double buffer to minimize communication overhead and mitigate straggler effects. For control logic, it introduces a task scheduler built upon Directed Acyclic Graph (DAG) that facilitates fine-grained, independent execution. Experimental results demonstrate that DISTFLOW achieves near-linear scalability up to 512 GPUs and delivers up to a 2.63x throughput improvement over state-of-the-art (SOTA) frameworks. The source code is available at: https://github.com/sii-research/siiRL.

cs.DC

Emergence of Homophily under Contextual Mechanisms

This paper introduces a tractable model to study incentive-compatible homophily under both external environments--such as exogenous shocks or policy constraints--and internal micromotives based on interactive attributes. We propose a set of invariants that capture main features of homophily and the well-defined partition dynamics leading to perfect global homophily. The criteria for homophily formation are characterized via isomorphism. Within this framework, we demonstrate the emergence of macro-complementarity coupled with micro-substitution, where local individuals' utility function is nonlinear and submodular. We discuss two types of financial networks and their differences: hierarchical structure emerges from short-term liquidity transactions, whereas core-periphery structure is based on a stock-based perspective.

econ.TH

LIMI: Less is More for Agency

We define Agency as the emergent capacity of AI systems to function as autonomous agents actively discovering problems, formulating hypotheses, and executing solutions through self-directed engagement with environments and tools. This fundamental capability marks the dawn of the Age of AI Agency, driven by a critical industry shift: the urgent need for AI systems that don't just think, but work. While current AI excels at reasoning and generating responses, industries demand autonomous agents that can execute tasks, operate tools, and drive real-world outcomes. As agentic intelligence becomes the defining characteristic separating cognitive systems from productive workers, efficiently cultivating machine autonomy becomes paramount. Current approaches assume that more data yields better agency, following traditional scaling laws from language modeling. We fundamentally challenge this paradigm. LIMI (Less Is More for Intelligent Agency) demonstrates that agency follows radically different development principles. Through strategic focus on collaborative software development and scientific research workflows, we show that sophisticated agentic intelligence can emerge from minimal but strategically curated demonstrations of autonomous behavior. Using only 78 carefully designed training samples, LIMI achieves 73.5% on comprehensive agency benchmarks, dramatically outperforming state-of-the-art models: Kimi-K2-Instruct (24.1%), DeepSeek-V3.1 (11.9%), Qwen3-235B-A22B-Instruct (27.5%), and GLM-4.5 (45.1%). Most strikingly, LIMI demonstrates 53.7% improvement over models trained on 10,000 samples-achieving superior agentic intelligence with 128 times fewer samples. Our findings establish the Agency Efficiency Principle: machine autonomy emerges not from data abundance but from strategic curation of high-quality agentic demonstrations.

cs.AI

Distributed Projection-free Algorithm for Constrained Aggregative Optimization

In this paper, we focus on solving a distributed convex aggregative optimization problem in a network, where each agent has its own cost function which depends not only on its own decision variables but also on the aggregated function of all agents' decision variables. The decision variable is constrained within a feasible set. In order to minimize the sum of the cost functions when each agent only knows its local cost function, we propose a distributed Frank-Wolfe algorithm based on gradient tracking for the aggregative optimization problem where each node maintains two estimates, namely an estimate of the sum of agents' decision variable and an estimate of the gradient of global function. The algorithm is projection-free, but only involves solving a linear optimization to get a search direction at each step. We show the convergence of the proposed algorithm for convex and smooth objective functions over a time-varying network. Finally, we demonstrate the convergence and computational efficiency of the proposed algorithm via numerical simulations.

math.OC

An Optimal Distributed Algorithm with Operator Extrapolation for Stochastic Aggregative Games

This work studies Nash equilibrium seeking for a class of stochastic aggregative games, where each player has an expectation-valued objective function depending on its local strategy and the aggregate of all players' strategies. We propose a distributed algorithm with operator extrapolation, in which each player maintains an estimate of this aggregate by exchanging this information with its neighbors over a time-varying network, and updates its decision through the mirror descent method. An operator extrapolation at the search direction is applied such that the two step historical gradient samples are utilized to accelerate the convergence. Under the strongly monotone assumption on the pseudo-gradient mapping, we prove that the proposed algorithm can achieve the optimal convergence rate of $\mathcal{O}(1/k)$ for Nash equilibrium seeking of stochastic games. Finally, the algorithm performance is demonstrated via numerical simulations.

math.OC

C-Arm Non-Circular Orbits: Geometric Calibration, Image Quality, and Avoidance of Metal Artifacts

Metal artifacts present a frequent challenge to cone-beam CT (CBCT) in image-guided surgery, obscuring visualization of metal instruments and adjacent anatomy. Recent advances in mobile C-arm systems have enabled 3D imaging capacity with non-circular orbits. We extend a previously proposed metal artifacts avoidance (MAA) method to reduce the influence of metal artifacts by prospectively defining a non-circular orbit that avoids metal-induced biases in projection domain. Accurate geometric calibration is an important challenge to accurate 3D image reconstruction for such orbits. We investigate the performance of interpolation-based calibration from a library of circular orbits for any non-circular orbit. We apply the method to non-circular scans acquired for MAA, which involves: (i) coarse 3D localization of metal objects via only two scout views using an end-to-end trained neural network; (ii) calculation of the metal-induced x-ray spectral shift for all possible views; and (iii) identification of the non-circular orbit that minimizes the variations in spectral shift. Non-circular orbits with interpolation-based geometric calibration yielded reasonably accurate 3D image reconstruction. The end-to-end neural network accurately localized metal implants with just two scout views even in complex anatomical scenes, improving Dice coefficient by ~42% compared to a more conventional cascade of separately trained U-nets. In a spine phantom with pedicle screw instrumentation, non-circular orbits identified by the MAA method reduced the magnitude of metal "blomming" artifacts (apparent width of the screw shaft) in CBCT reconstructions by ~70%. The proposed imaging and calibration methods present a practical means to improve image quality in mobile C-arm CBCT by identifying non-circular scan protocols that improve sampling and reduce metal-induced biases in the projection data.

physics.med-ph

Structure Sensitivity in Oxide Catalysis: First-Principles Kinetic Monte Carlo Simulations for CO Oxidation at RuO$_2$(111)

We present a density-functional theory based kinetic Monte Carlo study of CO oxidation at the (111) facet of RuO$_2$. We compare the detailed insight into elementary processes, steady-state surface coverages and catalytic activity to equivalent published simulation data for the frequently studied RuO$_2$(110) facet. Qualitative differences are identified in virtually every aspect ranging from binding energetics over lateral interactions to the interplay of elementary processes at the different active sites. Nevertheless, particularly at technologically relevant elevated temperatures, near-ambient pressures and near-stoichiometric feeds both facets exhibit almost identical catalytic activity. These findings challenge the traditional definition of structure sensitivity based on macroscopically observable turnover frequencies and allow to scrutinize the applicability of structure sensitivity classifications developed for metals to oxide catalysis.

cond-mat.mtrl-sci

Exploring Morphology-Activity Relationships: Ab Initio Wulff Construction for RuO2 Nanoparticles under Oxidizing Conditions

We present a density-functional theory based Wulff construction of the equilibrium shape of RuO2 particles in an oxygen environment. The obtained intricate variations of the crystal habit with the oxygen chemical potential allow for a detailed discussion of the dependence on the oxidizing pretreatment observed in recent powder catalyst studies. The analysis specifically indicates an incomplete particle shape equilibration in previously employed low temperature calcination. Equilibrated particles could be active CO oxidation catalysts with long-term stability in oxidizing feed and then represent an interesting alternative to the previously suggested core-shell concept.

cond-mat.mtrl-sci