SearcharxivSearch

arXiv subjects

Tong Bai

Publications and source records attributed to Tong Bai.

10 recordsLinked to original sources

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend on, conflict with, specialize, or duplicate one another, a structure invisible to both full enumeration and embedding similarity. We present SkillDAG, which models inter-skill relationships as a typed directed graph and exposes it to an LLM agent as an inference-time, agent-callable structural retrieval interface, queried and evolved during execution rather than baked into a fixed retrieval pipeline: each search returns vector matches, typed-edge neighbors, and conflict signals, and a propose-then-commit protocol lets the agent register execution-backed edges so the graph accumulates structure across episodes. On ALFWorld and SkillsBench with MiniMax-M2.7, SkillDAG reaches 67.1% success and 27.3% reward, exceeding the strongest reported Graph-of-Skills baseline by +12.8 and +8.6 points; the advantage ports to gpt-5.2-codex, and intrinsic SkillsBench Ret@K rises from 65.5 to 78.2 under matched queries. These gains trace to isolable mechanisms: candidate ranking that stays robust as the pool grows 10x where a fixed seeding-diffusion pipeline degrades, and set-monotone online edits that enlarge ground-truth recall without evicting prior hits.

cs.AI

Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention

Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over explicit Chain-of-Thought (CoT). However, the opacity of these continuous thought vectors hinders their reliability and controllability. This paper bridges the gap between mechanistic interpretability and actionable control. We first present a systematic analysis using structural, causal, and geometric probes, revealing that latent vectors encode compressed, faithful representations of reasoning steps, with early vectors acting as critical causal hubs. Building on this, we operationalize these interpretability insights into a suite of training-free, decode-time interventions that refine the latent reasoning process by imposing the identified geometric and semantic priors. Extensive experiments across multiple model scales and diverse task domains demonstrate that our approaches consistently improve reasoning accuracy. Our interpretability-guided interventions consistently unlock latent capabilities and improve reasoning accuracy without any parameter updates.

cs.CL

Towards Iterative End-to-End Software Development: A Feature-Driven Multi-Agent Framework

Recent advances in large language model agents offer the promise of automating end-to-end software development from natural language requirements. However, existing approaches largely adopt linear, waterfall-style pipelines, which oversimplify the iterative nature of real-world development and struggle with complex, large-scale projects. To address these limitations, we propose EvoDev, an iterative software development framework inspired by feature-driven development. EvoDev decomposes user requirements into a set of user-valued features and constructs a Feature Map, a directed acyclic graph that explicitly models dependencies between features. Each feature node in the feature map maintains multi-layer contexts, including business logic, software design, and code implementation, which are propagated along dependencies to provide context for subsequent development iterations. We evaluate EvoDev on challenging Android development tasks and show that it outperforms the best-performing baseline, Claude Code, by 57.3%, while improving single-agent performance by 16.0%-58.5% across different base LLMs, highlighting the importance of feature decomposition, dependency modeling, context propagation, and workflow-aware agent design for end-to-end software development. Moreover, our work summarizes practical insights for designing iterative, LLM-driven development frameworks and informs future training of base LLMs to better support iterative software development.

cs.SE

SAGE: Semantic-Aware Shared Sampling for Efficient Diffusion

Diffusion models manifest evident benefits across diverse domains, yet their high sampling cost, requiring dozens of sequential model evaluations, remains a major limitation. Prior efforts mainly accelerate sampling via optimized solvers or distillation, which treat each query independently. In contrast, we reduce total number of steps by sharing early-stage sampling across semantically similar queries. To enable such efficiency gains without sacrificing quality, we propose SAGE, a semantic-aware shared sampling framework that integrates a shared sampling scheme for efficiency and a tailored training strategy for quality preservation. Extensive experiments show that SAGE reduces sampling cost by 25.5%, while improving generation quality with 5.0% lower FID, 5.4% higher CLIP, and 160% higher diversity over baselines.

cs.LG

Environment-Aware AUV Trajectory Design and Resource Management for Multi-Tier Underwater Computing

The Internet of underwater things (IoUT) is envisioned to be an essential part of maritime activities. Given the IoUT devices' wide-area distribution and constrained transmit power, autonomous underwater vehicles (AUVs) have been widely adopted for collecting and forwarding the data sensed by IoUT devices to the surface-stations. In order to accommodate the diverse requirements of IoUT applications, it is imperative to conceive a multi-tier underwater computing (MTUC) framework by carefully harnessing both the computing and the communications as well as the storage resources of both the surface-station and of the AUVs as well as of the IoUT devices. Furthermore, to meet the stringent energy constraints of the IoUT devices and to reduce the operating cost of the MTUC framework, a joint environment-aware AUV trajectory design and resource management problem is formulated, which is a high-dimensional NP-hard problem. To tackle this challenge, we first transform the problem into a Markov decision process (MDP) and solve it with the aid of the asynchronous advantage actor-critic (A3C) algorithm. Our simulation results demonstrate the superiority of our scheme.

cs.DC

Reconfigurable Intelligent Surface Aided Mobile Edge Computing

Given the proliferation of wireless sensors and smart mobile devices, an explosive escalation of the volume of data is anticipated. However, restricted by their limited physical sizes and low manufacturing costs, these wireless devices tend to have limited computational capabilities and battery lives. To overcome this limitation, wireless devices may offload their computational tasks to the nearby computing nodes at the network edge in mobile edge computing (MEC). At the time of writing, the benefits of MEC systems have not been fully exploited, predominately because the computation offloading link is still far from perfect. In this article, we propose to enhance MEC systems by exploiting the emerging technique of reconfigurable intelligent surfaces (RIS), which are capable of `reconfiguring' the wireless propagation environments, hence enhancing the offloading links. The benefits of RISs can be maximized by jointly optimizing both the RISs as well as the communications and computing resource allocations of MEC systems. Unfortunately, this joint optimization imposes new research challenges on the system design. Against this background, this article provides an overview of RIS-assisted MEC systems and highlights their four use cases as well as their design challenges and solutions. Finally, their performance is characterized with the aid of a specific case study, followed by a range of future research ideas.

eess.SP

Latency Minimization for Intelligent Reflecting Surface Aided Mobile Edge Computing

Computation off-loading in mobile edge computing (MEC) systems constitutes an efficient paradigm of supporting resource-intensive applications on mobile devices. However, the benefit of MEC cannot be fully exploited, when the communications link used for off-loading computational tasks is hostile. Fortunately, the propagation-induced impairments may be mitigated by intelligent reflecting surfaces (IRS), which are capable of enhancing both the spectral- and energy-efficiency. Specifically, an IRS comprises an IRS controller and a large number of passive reflecting elements, each of which may impose a phase shift on the incident signal, thus collaboratively improving the propagation environment. In this paper, the beneficial role of IRSs is investigated in MEC systems, where single-antenna devices may opt for off-loading a fraction of their computational tasks to the edge computing node via a multi-antenna access point with the aid of an IRS. Pertinent latency-minimization problems are formulated for both single-device and multi-device scenarios, subject to practical constraints imposed on both the edge computing capability and the IRS phase shift design. To solve this problem, the block coordinate descent (BCD) technique is invoked to decouple the original problem into two subproblems, and then the computing and communications settings are alternatively optimized using low-complexity iterative algorithms. It is demonstrated that our IRS-aided MEC system is capable of significantly outperforming the conventional MEC system operating without IRSs. Quantitatively, about $20~\%$ computational latency reduction is achieved over the conventional MEC system in a single cell of a $300~\rm{m}$ radius and $5$ active devices, relying on a $5$-antenna access point.

eess.SP

Resource Allocation for Intelligent Reflecting Surface Aided Wireless Powered Mobile Edge Computing in OFDM Systems

Wireless powered mobile edge computing (WP-MEC) has been recognized as a promising technique to provide both enhanced computational capability and sustainable energy supply to massive low-power wireless devices. However, its energy consumption becomes substantial, when the transmission link used for wireless energy transfer (WET) and for computation offloading is hostile. To mitigate this hindrance, we propose to employ the emerging technique of intelligent reflecting surface (IRS) in WP-MEC systems, which is capable of providing an additional link both for WET and for computation offloading. Specifically, we consider a multi-user scenario where both the WET and the computation offloading are based on orthogonal frequency-division multiplexing (OFDM) systems. Built on this model, an innovative framework is developed to minimize the energy consumption of the IRS-aided WP-MEC network, by optimizing the power allocation of the WET signals, the local computing frequencies of wireless devices, both the sub-band-device association and the power allocation used for computation offloading, as well as the IRS reflection coefficients. The major challenges of this optimization lie in the strong coupling between the settings of WET and of computing as well as the unit-modules constraint on IRS reflection coefficients. To tackle these issues, the technique of alternative optimization is invoked for decoupling the WET and computing designs, while two sets of locally optimal IRS reflection coefficients are provided for WET and for computation offloading separately relying on the successive convex approximation method. The numerical results demonstrate that our proposed scheme is capable of monumentally outperforming the conventional WP-MEC network without IRSs.

eess.SP

Energy Efficient Transmission Based on Grouped Spatial Modulation for upstream DSL Systems

The digital Subscriber Line (DSL) remains an important component of heterogeneous networking, especially in historic city-centers, where using optical fibre is less realistic. Recently, the power consumption has become an important performance metric in telecommunication due to the associated environmental issues. In the recent bonding model, customer sites have been equipped with two/four copper pairs, which may be exploited for designing grouped spatial modulation (SM) aiming for reducing the power consumption and mitigating the stubborn crosstalk in DSL communications. Explicitly, we view the two pair copper pairs equipped for each user as a group and propose an energy efficient transmission scheme based on grouped SM strategy for the upstream DSL systems, which is capable of reducing the power consumption of the upstream transmitters by activating a single copper line of each user. More especially, in order to compensate for the potential bit-rate reduction imposed by reducing the number of activated lines, the proposed scheme implicitly delivers ``virtual bits" via activating/deactivating the lines in addition to the classic modulation scheme. This is particularly beneficial in the DSL context, because the cross-talk imposed by activating several lines may swamp the desired signal. Furthermore, a pair of near-optimal soft turbo detection schemes are proposed for exploiting the unique properties of the DSL channel in order to eliminate the error propagation problem of SM detection routinely encountered in wireless channels. Both the attainable energy-efficiency and the achievable Bit Error Ratio (BER) are investigated. Our simulation results demonstrate that the proposed group-based SM is capable of outperforming the vectoring scheme both in terms of its energy efficiency for all the examined loop lengths and transmit powers.

eess.SP

Scale Optimization for Full-Image-CNN Vehicle Detection

Many state-of-the-art general object detection methods make use of shared full-image convolutional features (as in Faster R-CNN). This achieves a reasonable test-phase computation time while enjoys the discriminative power provided by large Convolutional Neural Network (CNN) models. Such designs excel on benchmarks which contain natural images but which have very unnatural distributions, i.e. they have an unnaturally high-frequency of the target classes and a bias towards a "friendly" or "dominant" object scale. In this paper we present further study of the use and adaptation of the Faster R-CNN object detection method for datasets presenting natural scale distribution and unbiased real-world object frequency. In particular, we show that better alignment of the detector scale sensitivity to the extant distribution improves vehicle detection performance. We do this by modifying both the selection of Region Proposals, and through using more scale-appropriate full-image convolution features within the CNN model. By selecting better scales in the region proposal input and by combining feature maps through careful design of the convolutional neural network, we improve performance on smaller objects. We significantly increase detection AP for the KITTI dataset car class from 76.3% on our baseline Faster R-CNN detector to 83.6% in our improved detector.

cs.CV