SearcharxivSearch

arXiv subjects

Jinghui Wang

Publications and source records attributed to Jinghui Wang.

At least 19 recordsLinked to original sources

Addressing respiratory gating latency for accurate pulse delivery in preclinical electron FLASH irradiation on a clinical linear accelerator

Background: Clinical linear accelerators are an accessible platform for preclinical research on the biological effects of ultra rapid electron irradiation (FLASH). However, they are not inherently designed for the accurate pulse control required for experiments using a small number of relatively high-dose pulses, and available methods for beam control such as respiratory gating can be error prone owing to system latency. Here we experimentally characterize the temporal latency of the respiratory gating system for controlling beam-on and beam-off at the individual linac pulse level. Methods and Materials: We used programmable controller boards and a relay circuit to monitor and control delivery of specific numbers of pulses through the built-in monitor chamber and respiratory gating system of a Varian Trilogy linac. We implemented two methods an adaptive method using only the delivered pulse signal, and a synchronization method additionally using the linac internal pulse-timing signal and characterized their performance for standard and customized pulse sequences. Results: Characterizing the latency parameters permitted choosing optimal timing parameters that maximized the rate of successfully delivering the desired number of pulses using both adaptive and synchronization methods. Conclusions: We demonstrated that accounting for latency and/or using the ability to read the prior information on expected pulse timing can provide high accuracy in delivering specified numbers of pulses. This reliability is critical for accurate dose delivery in preclinical FLASH research of single fraction and especially fractionated dosing regimens. The ability to generate custom pulse sequences enables more detailed exploration of the temporal dependence of biological FLASH effects.

physics.med-ph

KAT-Coder-V2 Technical Report

We present KAT-Coder-V2, an agentic coding model developed by the KwaiKAT team at Kuaishou. KAT-Coder-V2 adopts a "Specialize-then-Unify" paradigm that decomposes agentic coding into five expert domains - SWE, WebCoding, Terminal, WebSearch, and General - each undergoing independent supervised fine-tuning and reinforcement learning, before being consolidated into a single model via on-policy distillation. We develop KwaiEnv, a modular infrastructure sustaining tens of thousands of concurrent sandbox instances, and scale RL training along task complexity, intent alignment, and scaffold generalization. We further propose MCLA for stabilizing MoE RL training and Tree Training for eliminating redundant computation over tree-structured trajectories with up to 6.2x speedup. KAT-Coder-V2 achieves 79.6% on SWE-bench Verified (vs. Claude Opus 4.6 at 80.8%), 88.7 on PinchBench (surpassing GLM-5 and MiniMax M2.7), ranks first across all three frontend aesthetics scenarios, and maintains strong generalist scores on Terminal-Bench Hard (46.8) and tau^2-Bench (93.9). Our model is publicly available at https://streamlake.com/product/kat-coder.

cs.CL

Synchronization in Traffic Dynamics: Mechanisms of Hysteresis

Starting from a second-order linear differential equation, we analyze the dynamical mechanisms of no behavior pattern (pure response), reaction and anticipation behaviors in traffic. As an emergence of the underlying dynamical evolution, the periodic evolution trajectories (3D hysteresis) in phase space ($v_i, v_j, d_{ji}$) exhibit fascinating characters. We investigate the emerging Time-Delay ($TD$) phenomena and the resulting analytical hysteresis, an equal frequency sets of Lissajous figures. By quantifying energy dissipation through individual and system perspectives, we demonstrate that $TD$ and Time-To-Collision ($TTC$) are direct metrics of zero-dissipation under equilibrium and synchronization states. Finally, a phase diagram based on $TD$ and $TTC$ is developed to bridge the dynamical behaviors in traffic across $\mathbb{R}^1$ and $\mathbb{R}^2$ spaces. Our results provide a theoretical foundation upon which many obscure mechanisms become self-evident, such as the $TD$-induced flip of the hysteresis (clockwise to counterclockwise in FD) and the crossed hysteresis, etc.

physics.soc-ph

Critical Thresholds in Non-Pharmaceutical Interventions for Epidemic Control

Non-pharmaceutical interventions, such as contact tracing and social distancing, are critical for controlling epidemic outbreaks, yet their dynamic interactions remain underexplored. We introduce a probabilistic framework to analyze the synergy between contact tracing speed, quantified by the contact tracing period $\tau$, and the average number of close contacts, $\bar{k}_+$, reflecting social distancing measures. We identify critical thresholds ($R=1$) that separate pandemic and contained phases in the $\bar{k}_{+}-\tau$ plane, validated using high-resolution data from Shenzhen's 2022 Omicron outbreak (1,187 cases, 86,451 contacts). Our findings show that contact tracing alone can contain diseases with $R_0 < 2.12$ (95% CI 2.07-2.16), covering 43.33% of major infectious diseases, while combining with social distancing extends control to $R_0 < 7.82$ (95% CI 7.70-7.93), encompassing 86.67% of pathogens. These results, supported by empirical data, highlight the efficacy of rapid tracing and targeted social distancing as alternatives to mass PCR testing. Our framework offers actionable insights for optimizing NPI strategies, though challenges in scaling to regions with higher tracing miss rates or weaker infrastructure underscore the need for adaptive, data-driven policies.

physics.soc-ph

Accelerated Discovery of Crystalline Materials with Record Ultralow Lattice Thermal Conductivity via a Universal Descriptor

Ultralow glass-like lattice thermal conductivity in crystalline materials is crucial for enhancing energy conversion efficiency in thermoelectrics and thermal insulators. We introduce a universal descriptor for thermal conductivity that relies only on the atomic number in the primitive cell and the sound velocity, enabling fast and scalable materials screening. Coupled with high-throughput workflows and universal machine learning potentials, we identify the candidate materials with ultralow thermal conductivity from over 25, 000 materials. We further validate this approach by experimentally confirming record-low thermal conductivity values of 0.15-0.16 W/m/K from 170 to 400 K in the halide metal CsAg2I3. Combining inelastic neutron scattering with first-principles calculations, we attribute the ultralow thermal conductivity to the intrinsically small sound velocity, strong anharmonicity, and structural complexity. Our work illustrates how a universal descriptor, combined with high-throughput screening, machine-learning potential and experiment, enables the efficient discovery of materials with ultralow thermal conductivity.

cond-mat.mtrl-sci

SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models

Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage, language bias, and insufficient alignment with real-world developer workflows. Existing benchmarks often focus on algorithmic problems or Python-centric bug fixing, leaving critical dimensions of software engineering underexplored. To address these gaps, we introduce SWE-Compass1, a comprehensive benchmark that unifies heterogeneous code-related evaluations into a structured and production-aligned framework. SWE-Compass spans 8 task types, 8 programming scenarios, and 10 programming languages, with 2000 high-quality instances curated from authentic GitHub pull requests and refined through systematic filtering and validation. We benchmark ten state-of-the-art LLMs under two agentic frameworks, SWE-Agent and Claude Code, revealing a clear hierarchy of difficulty across task types, languages, and scenarios. Moreover, by aligning evaluation with real-world developer practices, SWE-Compass provides a rigorous and reproducible foundation for diagnosing and advancing agentic coding capabilities in large language models.

cs.SE

Field-Tunable Anisotropic Fulde-Ferrell Phase in NbSe$_2$/CrSiTe$_3$ Heterostructures

The emergence of superconductivity in two-dimensional transition metal dichalcogenides with strong spin orbit coupling (SOC) has opened new avenues for exploring exotic superconducting states. Here, we report experimental observation of an anisotropic Fulde-Ferrell (FF) phase in few-layer NbSe$_2$/CrSiTe$_3$ heterostructures under in-plane magnetic fields. Through combined magnetoresistance and nonreciprocal transport measurements, we find that due to the couplings from the ferromagnetic CrSiTe$_3$, a half-dome-shaped region emerges in the magnetic field-temperature ($B$-$T$) diagram. Importantly, the half-dome-shaped region exhibits finite second harmonic resistance with in-plane anisotropy, indicating that the superconducting state is an anisotropic FF phase. Through a symmetry analysis combined with mean field calculations, we attribute the emergent anisotropic FF phase to the CrSiTe$_3$ layer induced Rashba SOC and three-fold rotational symmetry breaking. These results demonstrate that heterostructure stacking is a powerful tool for symmetry engineering in superconductors, which can advance the design of quantum devices in atomically thin superconducting materials.

cond-mat.supr-con

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse

Agentic large language model (LLM) training often involves multi-turn interaction trajectories that branch into multiple execution paths due to concurrent tool use, think-mode, sub-agent, context management and other runtime designs. As a result, the tokens produced by a single task naturally form a tree-structured token trajectory with shared prefixes, rather than a linear sequence. Existing training pipelines linearize such trajectories and treat each branch independently, leading to substantial redundant computation in both forward and backward passes. We derive that averaging the loss over all branches independently is algebraically identical to a per-token weighted loss, where each token's weight equals the fraction of branches passing through it. The problem therefore reduces to computing the log-probability of every token in the prefix tree exactly once, with no repeated computation across shared prefixes: we propose DFS serialization of the tree, which visits every token exactly once, and adapt full-attention and SSM layers to ensure the resulting log-probabilities match independent per-branch calculation exactly. In practice, a single trajectory tree can be too large to fit in GPU memory; we therefore propose Redundancy-Free Tree Partitioning, which handles memory-constrained settings with zero redundant computation and peak memory bounded by a single root-to-leaf path. Together, these contributions form Tree Training, an efficient framework for training LLMs on tree-structured trajectories, achieving up to 6.2x end-to-end training speedup on dense and MoE models for both supervised fine-tuning and reinforcement learning.

cs.LG

KAT-Coder Technical Report

Recent advances in large language models (LLMs) have enabled progress in agentic coding, where models autonomously reason, plan, and act within interactive software development workflows. However, bridging the gap between static text-based training and dynamic real-world agentic execution remains a core challenge. In this technical report, we present KAT-Coder, a large-scale agentic code model trained through a multi-stage curriculum encompassing Mid-Term Training, Supervised Fine-Tuning (SFT), Reinforcement Fine-Tuning (RFT), and Reinforcement-to-Deployment Adaptation. The Mid-Term stage enhances reasoning, planning, and reflection capabilities through a corpus of real software engineering data and synthetic agentic interactions. The SFT stage constructs a million-sample dataset balancing twenty programming languages, ten development contexts, and ten task archetypes. The RFT stage introduces a novel multi-ground-truth reward formulation for stable and sample-efficient policy optimization. Finally, the Reinforcement-to-Deployment phase adapts the model to production-grade IDE environments using Error-Masked SFT and Tree-Structured Trajectory Training. In summary, these stages enable KAT-Coder to achieve robust tool-use reliability, instruction alignment, and long-context reasoning, forming a deployable foundation for real-world intelligent coding agents. Our KAT series 32B model, KAT-Dev, has been open-sourced on https://huggingface.co/Kwaipilot/KAT-Dev.

cs.CL

SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior

In this paper, we study the content-aware layout generation problem, which aims to automatically generate layouts that are harmonious with a given background image. Existing methods usually deal with this task with a single-step reasoning framework. The lack of a feedback-based self-correction mechanism leads to their failure rates significantly increasing when faced with complex element layout planning. To address this challenge, we introduce SEGA, a novel Stepwise Evolution Paradigm for Content-Aware Layout Generation. Inspired by the systematic mode of human thinking, SEGA employs a hierarchical reasoning framework with a coarse-to-fine strategy: first, a coarse-level module roughly estimates the layout planning results; then, another refining module performs fine-level reasoning regarding the coarse planning results. Furthermore, we incorporate layout design principles as prior knowledge into the model to enhance its layout planning ability. Besides, we present GenPoster-100K that is a new large-scale poster dataset with rich meta-information annotation. The experiments demonstrate the effectiveness of our approach by achieving the state-of-the-art results on multiple benchmark datasets. Our project page is at: https://brucew91.github.io/SEGA.github.io/

cs.CV

Field-free Superconducting Diode Effect in FeTe$_{0.55}$Se$_{0.45}$

The superconducting diode effect (SDE) - the asymmetry of critical currents with respect to current direction - is a pivotal advancement in non-reciprocal superconductivity. While SDE has been realized in diverse systems, a fundamental challenge remains achieving field-free operation in iron-based superconductors with simple device geometries. Here, we report a non-volatile, field-free SDE in thin crystalline FeTe$_{0.55}$Se$_{0.45}$(FTS), showing asymmetric critical currents with a rectification coefficient of 1.9% and operating temperatures up to 9 K. Intriguingly, a pronounced non-zero second harmonic resistance emerges at the superconducting transition, exhibiting a sign reversal under varying current and temperature. The SDE persists at zero magnetic field and the rectification coefficient($\eta$) exhibits an even symmetric dependence on the magnetic field, distinguishing it from magnetic chirality anisotropy mechanisms. In addition to this, we systematically ruled out influences from dynamic superconducting domains, thermal gradients, and sample geometry, while establishing that localized stress amplifies the rectification coefficient, likely constituting one of the principal contributing factors. These results establish FTS as a robust platform for realizing field-free superconducting diodes in a structurally simple platform.

cond-mat.supr-con

Exploring Braess' paradox in pedestrian evacuation traffic: experiment and modeling

The emergence of the Braess' paradox in road traffic systems demonstrates the positive effect of transportation planning in improving efficiency. By contrast, the phenomenon has rarely been examined in pedestrian evacuation traffic. Yet the possibility that Braess' paradox could lengthen evacuation times under hazardous conditions has received little systematic attention. In this paper, we investigate Braess' paradox in pedestrian evacuation traffic through a series of supervised experiments and corresponding traffic assignment models to examine its potential occurrence. Our empirical and modeling results indicate that Braess' paradox is unlikely to be a prevalent phenomenon in pedestrian traffic systems. Specifically, under autonomous evacuation and the assumption of complete network knowledge, the paradox does not arise in high-demand evacuation contexts in our case studies. Under a more realistic assumption of limited network knowledge, however, the paradox can occur. These findings highlight the importance of information conditions for evacuation performance and provide guidance for the design and management of large public venues.

physics.soc-ph

SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

We introduce SeamlessFlow, a server based reinforcement learning (RL) framework that addresses two core challenges in industrial scale RL: (1) decoupling RL training from the complex execution flow of agents; (2) maximizing GPU utilization with minimal idle time while preserving the stability and scalability required for large-scale deployments. First, SeamlessFlow introduces a data plane that decouples the RL trainer from diverse, complex agent implementations while sustaining high throughput. A central trajectory manager maintains complete interaction histories and supports partial rollout, allowing rollout to pause for weight updates and resume seamlessly, keeping agents unaware of service interruptions. Second, we propose a tag driven scheduling paradigm that abstracts hardware into capability tagged resources, unifying colocated and disaggregated architectures. Based on this, SeamlessFlow introduces a spatiotemporal multiplexing pipeline that dynamically reassigns idle training nodes to rollout in a train rollout separated setup, eliminating pipeline bubbles and fully exploiting heterogeneous cluster resources. By combining these innovations, SeamlessFlow delivers both stability and high performance, making it well suited for multi agent, long horizon, and other complex RL tasks.

cs.LG

KAT-V1: Kwai-AutoThink Technical Report

We present Kwaipilot-AutoThink (KAT), an open-source 40B large language model developed to address the overthinking problem in reasoning-intensive tasks, where an automatic thinking training paradigm is proposed to dynamically switch between reasoning and non-reasoning modes based on task complexity. Specifically, first, we construct the dual-regime dataset based on a novel tagging pipeline and a multi-agent synthesis strategy, and then we apply Multi-Token Prediction (MTP)-enhanced knowledge distillation, enabling efficient and fine-grained reasoning transfer with minimal pretraining cost. Besides, we implement a cold-start initialization strategy that introduces mode-selection priors using majority-vote signals and intent-aware prompting. Finally, we propose Step-SRPO, a reinforcement learning algorithm that incorporates intermediate supervision into the GRPO framework, offering structured guidance over both reasoning-mode selection and response accuracy. Extensive experiments across multiple benchmarks demonstrate that KAT consistently matches or even outperforms current state-of-the-art models, including DeepSeek-R1-0528 and Qwen3-235B-A22B, across a wide range of reasoning-intensive tasks while reducing token usage. Notably, KAT outperforms all open-source models and even surpasses o3-mini on the leakage-controlled LiveCodeBench Pro. Beyond academic evaluation, KAT has been successfully deployed in Kwaipilot (i.e., Kuaishou's internal coding assistant), where it improves real-world development workflows with high accuracy, efficiency, and controllable reasoning behaviors. Moreover, we are actively training a 200B Mixture-of-Experts (MoE) model with 40B active parameters, and early results already show significant gains, further demonstrating the scalability of the AutoThink paradigm.

cs.CL

A Kinematic Constraint on Pedestrian Walking: Power-law Scaling between Critical Angular Velocity and Speed

This paper presents a statistical analysis of speed and angular velocity obtained from pedestrian experiments across nine distinct datasets. Experimental scenarios included crossing motion, unidirectional/bidirectional flows, bidirectional/four-directional crossing flows, pedestrian-vehicle interactions, unidirectional flow in a circular corridor, and circle antipode configurations. We applied filtering methods to reduce noise and analyzed the data at different sampling frequencies. The results reveal a universal power-law scaling between critical angular velocity and speed, with a scaling exponent of approximately -0.8. This relationship defines a bounded region in the speed-angular velocity phase space, suggesting a kinematic constraint on pedestrian motion.

physics.soc-ph

Dependence of the Radical Dynamics on the Beam Temporal Profile in FLASH Radiotherapy

Purpose: This study aims to investigate the impact of the beam temporal profile on the radical dynamics and inter-track interactions of FLASH radiotherapy, supporting parameter optimization for the equipment development and clinical implementation. Methods: MonteCarlo simulations based on the IRT method were performed to analyze the dynamics after irradiation, including single-pulse or multi-pulses irradiation, pulse repetition rate, width and dose. The physicochemical experiments were performed to measure the eaq-lifetimes for validation. The generation and recombination of OH and eaq-radicals were recorded under 6 MeV electron irradiation with varying beam temporal profiles. The radial distributions of the radicals were statistically analyzed, and the corresponding LETd and LETt were calculated. The inter-track interactions were assessed through a mathematical model. Results: The spatial distribution and temporal evolution of radicals were significantly affected by the beam time profiles. Compared with multi-pulses irradiation, single-pulse mode with a width less than 1/10 of the radical lifetime, a repetition interval longer than the radical lifetime, and a dose exceeding 1 Gy/pulse can lead to radicals rapid consumption, reducing the residual content. Instantaneous high dose rates induced radical tracks overlaps. When the single-pulse dose exceeded 1 Gy, the overlap probability approached 100%, aligning with the threshold for radical instantaneous combination. Conclusion: Under a low-duty cycle and high instantaneous dose-rate time profile, the radicals were rapidly consumed through track overlap hence reduced damage to normal tissues, inducing FLASH effect. The optimized time profile can be used to guide the development of equipment and parameter settings in clinical practice to maximize the FLASH effect, such as the laser accelerators and superconducting photocathode guns.

physics.med-ph

SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Recent advances of reasoning models, exemplified by OpenAI's o1 and DeepSeek's R1, highlight the significant potential of Reinforcement Learning (RL) to enhance the reasoning capabilities of Large Language Models (LLMs). However, replicating these advancements across diverse domains remains challenging due to limited methodological transparency. In this work, we present two-Staged history-Resampling Policy Optimization (SRPO), which surpasses the performance of DeepSeek-R1-Zero-32B on the AIME24 and LiveCodeBench benchmarks. SRPO achieves this using the same base model as DeepSeek (i.e. Qwen2.5-32B), using only about 1/10 of the training steps required by DeepSeek-R1-Zero-32B, demonstrating superior efficiency. Building upon Group Relative Policy Optimization (GRPO), we introduce two key methodological innovations: (1) a two-stage cross-domain training paradigm designed to balance the development of mathematical reasoning and coding proficiency, and (2) History Resampling (HR), a technique to address ineffective samples. Our comprehensive experiments validate the effectiveness of our approach, offering valuable insights into scaling LLM reasoning capabilities across diverse tasks.

cs.LG

A Galton Board Approximation Method for Estimating Pedestrian Walking Preferences within Crowds

This paper proposes a Galton board approximation method to analyze the potential walking preferences of pedestrians. We employ the binomial distribution to estimate the walking preferences of pedestrians in dynamic crowds. Estimating the probability of the right-side preference ($p$) based on observed data poses the challenge, as statistical measures such as means and variances often lead to divergent results. This study aims to explore this issue.

physics.soc-ph