SearcharxivSearch

arXiv subjects

Ting Peng

Publications and source records attributed to Ting Peng.

At least 19 recordsLinked to original sources

Holmes: Multimodal Agentic Diagnosis for Mixed-Language Mobile Crashes at Industrial Scale

Diagnosing mobile crashes in ultra-large-scale industrial applications is a formidable challenge due to the sheer volume of code, the complexity of mixed-language environments, and the inability to reproduce failures locally. Traditional static analysis struggles with scalability, while existing LLM-based agents often rely on reproducible environments unavailable in post-mortem scenarios. We present Holmes, a multi-agent system that automates root cause analysis by synthesizing multimodal runtime signals--stack traces, logs, and thread states--to reconstruct failure contexts without reproduction. Holmes introduces a hierarchical Retrieve-Explore-Reason architecture that leverages low-level artifacts (e.g., registers, assembly) to bridge the semantic gap between open-source business logic and closed-source system frameworks. By dynamically compressing the search space using runtime clues, Holmes precisely navigates 70-million-line codebases to identify non-local defects. Evaluated on real-world crashes from WeChat, Holmes achieves 87.6% accuracy in function-level fault localization and reduces average investigation time by over 98% (to ~77 seconds), demonstrating its effectiveness in transforming labor-intensive debugging into an efficient verification workflow.

cs.AI

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair

Code-agent RL often receives weak feedback: rollout-time signals are reliable and executable, but capture only necessary or surface conditions for task success rather than the target semantic predicate. Using agentic compile-fix as the setting, we study signal reshaping for standard GRPO under such feedback. Our central claim is that GRPO's within-group comparison is meaningful only after three kinds of signals are reshaped: outcome rewards recover semantic ranking, process signals localize intra-trajectory credit, and rollouts from the same prompt remain execution-comparable. We operationalize these conditions with a minimal signal-reshaping construction that leaves GRPO's group-normalized advantage construction unchanged: compile-and-semantic layered rewards reshape trajectory ranking, step-level process scores outside group reward normalization reshape within-trajectory update strength, and failure-cause-aware rollout governance reshapes within-group comparability. Experiments show a clear end-to-end gain: full signal-reshaped GRPO improves strict compile-and-semantic accuracy from the base model's zero-shot $0.385$ to $0.535$. Controlled comparisons further explain the source of this gain: binary rewards remove the compile-only middle tier and degrade trajectory control; on top of layered rewards, process-score weighting further improves accuracy from $0.48$ to $0.53$ and reduces average evaluation steps from $23.50$ to $17.02$. As a boundary comparison, privileged-prompt token-level distillation mainly optimizes local distributional alignment; in long tool-use trajectories, this signal is diluted by non-critical tokens and cannot replace outcome semantics, process credit, or within-group comparability.

cs.AI

Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code Editing

Code editing constitutes a fundamental practice in software development, wherein developers modify existing codebases according to natural language requirements. Accurate code editing necessitates a comprehensive understanding of both the existing codebase and the modification requirements. Although large language models (LLMs) have demonstrated promising performance in code editing tasks, they suffer from substantial inefficiency by generating entire modified files that largely consist of unchanged code. While smaller models could potentially address this inefficiency, they typically lack the capacity to effectively comprehend long code contexts required for accurate editing. To ensure both effectiveness and efficiency, we propose to decompose code editing into a two-stage cascade: \textbf{edit sketch generation}, wherein a large model first produces concise sketches representing the requisite modifications (the more challenging phase), and \textbf{edit sketch application}, wherein a smaller model integrates these sketches into the original code to produce the final output edited code (the simpler phase). This cascaded design reduces the number of tokens generated by the large model, as the majority of the output is handled by the smaller, more efficient model, thereby enhancing overall efficiency. However, the effectiveness of this approach is constrained by current small models' limited capabilities in handling long-context scenarios and cross-file dependencies, which are essential for accurate sketch application in real-world codebases. To address these limitations and enhance smaller models' sketch application capabilities, ...

cs.SE

Strict Entropy Decrease of Clausius Entropy in an Isolated System with Energy-Form Conversion: Theoretical Proof, Numerical Illustration, and Critical Examination

This paper is accountable only to explicitly stated physical assumptions and strict logical inference. Its goal is to run a rigorous stress test of second-law claims within the Clausius framework. We work directly with \textbf{Clausius's entropy definition} for an isolated composite with energy-form conversion. Heat is withdrawn from a cold releasing subsystem with relatively small heat capacity, converted to electrical energy, and then delivered as heat to a hotter subsystem. In the ideal limit, the electrical leg contributes negligibly to Clausius entropy accounting, so the modeled reservoir Clausius sum is \[ \Delta S_{\mathrm{Cl}} = Q\!\left(\frac{1}{T_B}-\frac{1}{T_A}\right) < 0. \] The paper provides a derivation, numerical illustrations, and a scope analysis; any claimed contradiction should be interpreted as a compatibility issue between different axiom sets, not as an algebraic error in the Clausius bookkeeping above.

cond-mat.stat-mech

Harvest Ambient Heat via Constraint-Shaped Phase-Change Cycles: Micro $\Delta T$, Subcooled Liquid, and Liquid-Only Compression

Conventional heat engines typically require two distinct thermal reservoirs, with their efficiency strictly bounded by the Carnot limit. We present a theoretical design for a phase-change heat engine that utilizes water as the working fluid undergoing state transitions within geometry-constrained flow paths. The proposed cycle operates under a micro-temperature difference (1--2\,$^\circ$C) and relies on liquid-only compression. The system harvests thermal energy via an \textbf{ambient micro-temperature difference} relative to the environment ($q_{\mathrm{in}} \approx 8.37\,\mathrm{kJ}/\mathrm{kg}$ at 24--26\,$^\circ$C). Expansion work is recovered from the enthalpy drop during flash evaporation. Comprehensive numerical analysis using NIST property data confirms that, in the reversible limit, the cycle yields positive net work while maintaining standard thermodynamic consistency. This study illustrates the theoretical potential for ambient energy harvesting via low-pressure phase change, although the extremely small work output per cycle suggests that hardware realization will require exceptional mechanical precision to overcome parasitic losses.

cond-mat.stat-mech

Entropy Has No Direction: A Mirror-State Paradox Against Universal Monotonic Entropy Increase and a First-Principles Proof that Constraints Reshape the Entropy Distribution $P_{\infty}(S;\lambda)$

We revisit textbook claims that entropy must increase and show that, under time-reversal invariant microscopic dynamics, no universal trajectory-wise or statistical assertion that the coarse-grained entropy $S(t)$ is non-decreasing can hold. The core is a mirror-state construction: for any microstate $A$ one constructs its time-reversed partner $B$ (momenta inverted); requiring $S(t)$ to be non-decreasing for both $A$ and $B$ forces every time to be a local minimum of $S$ and hence makes $S(t)$ constant along the trajectory. The consistent picture is that entropy is a stochastic variable described by a probability distribution $P(S)$ whose shape depends on constraints and boundary conditions; entropy-based regularities are emergent summaries of constraint-dependent microscopic dynamics, and in practice it is constraints and boundaries -- not entropy itself -- that one manipulates to achieve mixing, separation, or self-organization. Working with Boltzmann (coarse-grained) entropy on the energy shell, we then derive from first principles how constraints reshape the long-time entropy distribution $P_{\infty}(S;\lambda)$ by altering the invariant measure through changes in the Hamiltonian and/or the accessible phase space. In the microcanonical setting we obtain a sharp criterion: the \emph{only} way $P_{\infty}^{(E)}(S;\lambda)$ can remain the same up to translation is when all accessible macrostate volumes are scaled by a common factor; otherwise the distribution changes structurally. We connect this framework to experiments on asymmetric nanopores and molecular gates, to macroscopic examples from civil engineering (windbreak forests, dikes, vortex suppression, traffic-flow control), and to natural phenomena such as lightning guided to lightning rods, snowflake and mineral-veil growth, and the sudden crystallisation of supercooled water.

cond-mat.stat-mech

Geometry Challenges Entropy: Regime-DependentRectification in Nanofluidic Cascades

Can geometry alone reshape equilibrium? Cascaded nanofluidic chambers show complex accumulation patterns, traditionally attributed to geometric diode effects. We use 3D molecular dynamics to decouple funnel rectification from boundary reflection. Simulations with argon parameters (r = 0.19 nm) reveal a striking "reverse" rectification in a 2-chamber setup: the narrow side accumulates over 5x more particles (N_1/N_0 = 5.37 +/- 0.01, p < 0.0001). In a 10-chamber argon cascade, this effect drives massive downstream accumulation. A symmetric control (w_L = w_R) eliminates the gradient, confirming that funnel asymmetry - not boundary/edge effects - is the primary driver in the ballistic regime. By contrast, the super-atom regime is dominated by boundary reflection. Our results challenge standard entropic transport theory and provide design rules for passive, geometry-driven density gradients - no pump, no drive.

physics.comp-ph

Hierarchical Preemptive Holistic Collaborative Systems for Embodied Multi-Agent Systems: Framework, Hybrid Stability, and Scalability Analysis

The coordination of Embodied Multi-Agent Systems in constrained physical environments requires a rigorous balance between safety, scalability, and efficiency. Traditional decentralized approaches, e.g., reactive collision avoidance, are prone to local minima or reciprocal yielding standoffs due to the lack of future intent awareness. In contrast, centralized planning suffers from intractable computational complexity and single-point-of-failure vulnerabilities. To address these limitations, we propose the Hierarchical Preemptive Holistic Collaborative (Prollect) framework, which generalizes the Preemptive Holistic Collaborative System (PHCS) by decomposing the global coordination problem into topologically connected subspace optimizations. We formalize the system as a Hybrid Automaton and introduce a three-stage receding horizon mechanism (frozen execution, preliminary planning, proactive look-ahead windows) with explicit padding to prevent races between coordination dissemination and intent updates. Notably, we design a robust timing protocol with a mandatory Idle Buffer that acts as a dwell-time constraint to eliminate Zeno behaviors and ensure computational stability under jitter. Furthermore, we formalize a Shadow Agent protocol to guarantee seamless trajectory consistency across subspace boundaries, which we treat as an Input-to-State Stability (ISS) problem.

eess.SY

Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage

Large language models (LLMs) demonstrate strong capabilities across a wide range of complex tasks and are increasingly deployed at scale, placing significant demands on inference efficiency. Prior work typically decomposes inference into prefill and decode stages, with the decode stage dominating total latency. To reduce time and memory complexity in the decode stage, a line of work introduces sparse-attention algorithms. In this paper, we show, both empirically and theoretically, that sparse attention can paradoxically increase end-to-end complexity: information loss often induces significantly longer sequences, a phenomenon we term ``Less is Less'' (Lil). To mitigate the Lil problem, we propose an early-stopping algorithm that detects the threshold where information loss exceeds information gain during sparse decoding. Our early-stopping algorithm reduces token consumption by up to 90% with a marginal accuracy degradation of less than 2% across reasoning-intensive benchmarks.

cs.CL

Seasonal thermal stress analysis of defective mass concrete sidewalls based on the average forming temperature method

Thermal cracking in urban underground sidewalls is frequently observed when structures are cast in summer and enter service in winter, as seasonal temperature gradients act under structural restraint. To quantify the local stress field associated with pre-existing cracks, an orthogonal finite-element simulation matrix of 16 combinations is constructed. Distributions of maximum principal stress () at the surface crack tip and along the upper half of the crack bottom are evaluated using steady-state thermal loading and a linear-elastic constitutive model. Across all cases, pronounced tensile stress concentration occurs at both locations: the maximum ranges from 19.2 to 34.1 MPa at the crack surface end and from 17.2 to 29.4 MPa at the crack bottom. These concentrated values are consistently higher than the stress level at the same locations in an otherwise identical uncracked wall, clarifying how seasonal temperature gradients under restraint amplify local stresses around existing defects. The quantitative ranges reported here provide a basis for risk screening and for formulating practical mitigation measures (e.g., joint spacing and insulation strategies) in the design and operation of urban underground enclosure walls. In addition, three-dimensional simulations of randomly distributed internal voids show that adopting average forming temperature increases the peak tensile stress on void surfaces from 3.42 to 4.40 MPa at 10 deg C and from 5.98 to 6.96 MPa at -5 deg C, further highlighting the risk amplification effect of AFT under cold service conditions.

physics.comp-ph

Preemptive Spatiotemporal Trajectory Adjustment for Heterogeneous Vehicles in Highway Merging Zones

Aiming at the problem of driver's perception lag and low utilization efficiency of space-time resources in expressway ramp confluence area, based on the preemptive spatiotemporal trajectory Adjustment system, from the perspective of coordinating spatiotemporal resources, the reasonable value of safe space-time distance in trajectory pre-preparation is quantitatively analyzed. The minimum safety gap required for ramp vehicles to merge into the mainline is analyzed by introducing double positioning error and spatiotemporal trajectory tracking error. A merging control strategy for autonomous driving heterogeneous vehicles is proposed, which integrates vehicle type, driving intention, and safety spatiotemporal distance. The specific confluence strategies of ramp target vehicles and mainline cooperative vehicles under different vehicle types are systematically expounded. A variety of traffic flow and speed scenarios are used for full combination simulation. By comparing the time-position-speed diagram, the vehicle operation characteristics and the dynamic difference of confluence are qualitatively analyzed, and the average speed and average delay are used as the evaluation indices to quantitatively evaluate the performance advantages of the preemptive cooperative confluence control strategy. The results show that the maximum average delay improvement rates of mainline and ramp vehicles are 90.24 % and 74.24 %, respectively. The proposed strategy can effectively avoid potential vehicle conflicts and emergency braking behaviors, improve driving safety in the confluence area, and show significant advantages in driving stability and overall traffic efficiency optimization.

eess.SY

Autonomous Aggregate Sorting in Construction and Mining via Computer Vision-Aided Robotic Arm Systems

Traditional aggregate sorting methods, whether manual or mechanical, often suffer from low precision, limited flexibility, and poor adaptability to diverse material properties such as size, shape, and lithology. To address these limitations, this study presents a computer vision-aided robotic arm system designed for autonomous aggregate sorting in construction and mining applications. The system integrates a six-degree-of-freedom robotic arm, a binocular stereo camera for 3D perception, and a ROS-based control framework. Core techniques include an attention-augmented YOLOv8 model for aggregate detection, stereo matching for 3D localization, Denavit-Hartenberg kinematic modeling for arm motion control, minimum enclosing rectangle analysis for size estimation, and hand-eye calibration for precise coordinate alignment. Experimental validation with four aggregate types achieved an average grasping and sorting success rate of 97.5%, with comparable classification accuracy. Remaining challenges include the reliable handling of small aggregates and texture-based misclassification. Overall, the proposed system demonstrates significant potential to enhance productivity, reduce operational costs, and improve safety in aggregate handling, while providing a scalable framework for advancing smart automation in construction, mining, and recycling industries.

cs.RO

A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat

Code completion, a crucial task in software engineering that enhances developer productivity, has seen substantial improvements with the rapid advancement of large language models (LLMs). In recent years, retrieval-augmented generation (RAG) has emerged as a promising method to enhance the code completion capabilities of LLMs, which leverages relevant context from codebases without requiring model retraining. While existing studies have demonstrated the effectiveness of RAG on public repositories and benchmarks, the potential distribution shift between open-source and closed-source codebases presents unique challenges that remain unexplored. To mitigate the gap, we conduct an empirical study to investigate the performance of widely-used RAG methods for code completion in the industrial-scale codebase of WeChat, one of the largest proprietary software systems. Specifically, we extensively explore two main types of RAG methods, namely identifier-based RAG and similarity-based RAG, across 26 open-source LLMs ranging from 0.5B to 671B parameters. For a more comprehensive analysis, we employ different retrieval techniques for similarity-based RAG, including lexical and semantic retrieval. Based on 1,669 internal repositories, we achieve several key findings: (1) both RAG methods demonstrate effectiveness in closed-source repositories, with similarity-based RAG showing superior performance, (2) the effectiveness of similarity-based RAG improves with more advanced retrieval techniques, where BM25 (lexical retrieval) and GTE-Qwen (semantic retrieval) achieve superior performance, and (3) the combination of lexical and semantic retrieval techniques yields optimal results, demonstrating complementary strengths. Furthermore, we conduct a developer survey to validate the practical utility of RAG methods in real-world development environments.

cs.SE

RAG or Fine-tuning? A Comparative Study on LCMs-based Code Completion in Industry

Code completion, a crucial practice in industrial settings, helps developers improve programming efficiency by automatically suggesting code snippets during development. With the emergence of Large Code Models (LCMs), this field has witnessed significant advancements. Due to the natural differences between open-source and industrial codebases, such as coding patterns and unique internal dependencies, it is a common practice for developers to conduct domain adaptation when adopting LCMs in industry. There exist multiple adaptation approaches, among which retrieval-augmented generation (RAG) and fine-tuning are the two most popular paradigms. However, no prior research has explored the trade-off of the two approaches in industrial scenarios. To mitigate the gap, we comprehensively compare the two paradigms including Retrieval-Augmented Generation (RAG) and Fine-tuning (FT), for industrial code completion in this paper. In collaboration with Tencent's WXG department, we collect over 160,000 internal C++ files as our codebase. We then compare the two types of adaptation approaches from three dimensions that are concerned by industrial practitioners, including effectiveness, efficiency, and parameter sensitivity, using six LCMs. Our findings reveal that RAG, when implemented with appropriate embedding models that map code snippets into dense vector representations, can achieve higher accuracy than fine-tuning alone. Specifically, BM25 presents superior retrieval effectiveness and efficiency among studied RAG methods. Moreover, RAG and fine-tuning are orthogonal and their combination leads to further improvement. We also observe that RAG demonstrates better scalability than FT, showing more sustained performance gains with larger scales of codebase.

cs.SE

Numerical Analysis of Temperature and Stress Fields in Mass Concrete Based on Average Forming Temperature Method

Mass concrete plays a crucial role in large-scale projects such as water conservancy hubs and transportation infrastructure. Due to its substantial volume and poor thermal conductivity, the accumulation of hydration heat during the curing process can lead to uneven temperature gradients and stress field distribution, which may cause structural cracking. This phenomenon represents one of the critical challenges in quality control for hydraulic dams, bridge piers and abutments, tunnel linings, and similar engineering structures. To ensure structural safety, it is imperative to calculate temperature variations while optimizing and controlling the temperature stress field. In this paper, a novel method for calculating the zero-stress temperature field is proposed, considering the temperature history and hydration heat release increments at various locations within mass concrete during the curing period, the parameter of average forming temperature field is defined to subsequently solve the temperature stress field. Several typical hydration heat release models were selected to calibrate the computational accuracy of the average forming temperature. Based on simulation results, an optimal model was applied to validate the effectiveness of the proposed method through practical engineering case studies. The impacts of casting temperature, ambient temperature during the curing period, and dimensional thickness on temperature-induced stresses were systematically investigated. Additionally, stress variations at different representative points were compared with the overall mean stress distribution. The results demonstrate that this method can more accurately evaluate temperature induced stresses caused by seasonal temperature variations. This study provides a more reliable computational basis for ensuring the long-term service safety of mass concrete structures.

physics.comp-ph

Spatiotemporal Trajectory Tracking Method for Vehicles Incorporating Lead-Lag Judgement

In the domain of intelligent transportation systems, especially within the context of autonomous vehicle control, the preemptive holistic collaborative system has been presented as a promising solution to bring a remarkable enhancement in traffic efficiency and a substantial reduction in the accident rate, demonstrating a great potential of development. In order to ensure this system operates as intended, accurate tracking of the spatiotemporal trajectory is of crucial significance. Moreover, minimizing the tracking error is a necessary step in this process. To this end, a novel lead-lag judgment mechanism is proposed. This mechanism precisely quantifies the longitudinal positional deviation between the vehicle and the target trajectory over time, then the deviation is corrected with a real - time acceleration compensation strategy, as a result, the accuracy and reliability of trajectory tracking are significantly enhanced. Real - vehicle experiments were conducted in a dedicated test field to validate the feasibility of this innovative approach empirically. Subsequently, the obtained tracking data was subsequent processed using the lead-lag judgment mechanism. In this step, we carefully analyzed the spatiotemporal error patterns between the vehicle and the target trajectory under different alignments and speeds. Finally, using real highway speed and alignment data, we conducted comprehensive spatiotemporal trajectory tracking simulations. Through experiments and simulations, tracking errors maintained in an acceptable range and reasonable spatiotemporal distance is given during the preemptive merging process on highway ramps. Overall, this study offers valuable insights for highway ramp emerging safety. Future work can expand on these findings.

eess.SY

Preemptive Holistic Collaborative System and Its Application in Road Transportation

Numerous real-world systems, including manufacturing processes, supply chains, and robotic systems, involve multiple independent entities with diverse objectives. The potential for conflicts arises from the inability of these entities to accurately predict and anticipate each other's actions. To address this challenge, we propose the Preemptive Holistic Collaborative System (PHCS) framework. By enabling information sharing and collaborative planning among independent entities, the PHCS facilitates the preemptive resolution of potential conflicts. We apply the PHCS framework to the specific context of road transportation, resulting in the Preemptive Holistic Collaborative Road Transportation System (PHCRTS). This system leverages shared driving intentions and pre-planned trajectories to optimize traffic flow and enhance safety. Simulation experiments in a two-lane merging scenario demonstrate the effectiveness of PHCRTS, reducing vehicle time delays by 90%, increasing traffic capacity by 300%, and eliminating accidents. The PHCS framework offers a promising approach to optimize the performance and safety of complex systems with multiple independent entities.

eess.SY

Enhancing Expressway Ramp Merge Safety and Efficiency via Spatiotemporal Cooperative Control

In the context of autonomous driving on expressways, the issue of ensuring safe and efficient ramp merging remains a significant challenge. Existing systems often struggle to accurately assess the status and intentions of other vehicles, leading to a persistent occurrence of accidents despite efforts to maintain safe distances. This study proposes a novel spatiotemporal cooperative control approach integrating vehicle-road coordination to address this critical issue. A comprehensive methodology is developed, beginning with the calculation of safe distances under varying spatiotemporal conditions. This involves considering multiple factors, including vehicle speed differentials, positioning errors, and clock synchronization errors. Subsequently, an advanced vehicle conflict risk evaluation model is constructed. By incorporating collision acceleration and emergency acceleration as key parameters, this model offers a more accurate and detailed evaluation of potential risks during the ramp merging process. Based on the calculated safe distances and conflict risk evaluations, a mainline priority coordinated control method is formulated. This method enables the pre-planning of vehicle trajectories, effectively reducing conflicts among vehicles. Through rigorous simulations using diverse traffic volume and speed scenarios, the efficacy of the proposed strategy is validated. The results demonstrate remarkable improvements, with the average delay time reduced by an impressive 97.96% and fuel consumption decreased by 6.01%. These outcomes indicate that the proposed approach not only enhances the speed of vehicle merging but also significantly reduces latency and fuel consumption, thereby enhancing the overall performance of ramp merging operations.

eess.SY