Searcharxiv⌕ Search

arXiv subjects

Min Tang

Publications and source records attributed to Min Tang.

At least 37 records · Page 2Linked to original sources

LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) incorporates external knowledge into large language models (LLMs), improving their adaptability to downstream tasks and enabling information updates. Surprisingly, recent empirical evidence demonstrates that injecting noise into retrieved relevant documents paradoxically facilitates exploitation of external knowledge and improves generation quality. Although counterintuitive and challenging to apply in practice, this phenomenon enables granular control and rigorous analysis of how LLMs integrate external knowledge. Therefore, in this paper, we intervene on noise injection and establish a layer-specific functional demarcation within the LLM: shallow layers specialize in local context modeling, intermediate layers focus on integrating long-range external factual knowledge, and deeper layers primarily rely on parametric internal knowledge. Building on this insight, we propose Layer Fused Decoding (LFD), a simple decoding strategy that directly combines representations from an intermediate layer with final-layer decoding outputs to fully exploit the external factual knowledge. To identify the optimal intermediate layer, we introduce an internal knowledge score (IKS) criterion that selects the layer with the lowest IKS value in the latter half of layers. Experimental results across multiple benchmarks demonstrate that LFD helps RAG systems more effectively surface retrieved context knowledge with minimal cost.

cs.CL↗

Using GUI Agent for Electronic Design Automation

Graphical User Interface (GUI) agents adopt an end-to-end paradigm that maps a screenshot to an action sequence, thereby automating repetitive tasks in virtual environments. However, existing GUI agents are evaluated almost exclusively on commodity software such as Microsoft Word and Excel. Professional Computer-Aided Design (CAD) suites promise an order-of-magnitude higher economic return, yet remain the weakest performance domain for existing agents and are still far from replacing expert Electronic-Design-Automation (EDA) engineers. We therefore present the first systematic study that deploys GUI agents for EDA workflows. Our contributions are: (1) a large-scale dataset named GUI-EDA, including 5 CAD tools and 5 physical domains, comprising 2,000+ high-quality screenshot-answer-action pairs recorded by EDA scientists and engineers during real-world component design; (2) a comprehensive benchmark that evaluates 30+ mainstream GUI agents, demonstrating that EDA tasks constitute a major, unsolved challenge; and (3) an EDA-specialized metric named EDAgent, equipped with a reflection mechanism that achieves reliable performance on industrial CAD software and, for the first time, outperforms Ph.D. students majored in Electrical Engineering. This work extends GUI agents from generic office automation to specialized, high-value engineering domains and offers a new avenue for advancing EDA productivity. The dataset will be released at: https://github.com/aiben-ch/GUI-EDA.

cs.CV↗

RLCAD: Reinforcement Learning Training Gym for Revolution Involved CAD Command Sequence Generation

A CAD command sequence is a typical parametric design paradigm in 3D CAD systems where a model is constructed by overlaying 2D sketches with operations such as extrusion, revolution, and Boolean operations. Although there is growing academic interest in the automatic generation of command sequences, existing methods and datasets only support operations such as 2D sketching, extrusion,and Boolean operations. This limitation makes it challenging to represent more complex geometries. In this paper, we present a reinforcement learning (RL) training environment (gym) built on a CAD geometric engine. Given an input boundary representation (B-Rep) geometry, the policy network in the RL algorithm generates an action. This action, along with previously generated actions, is processed within the gym to produce the corresponding CAD geometry, which is then fed back into the policy network. The rewards, determined by the difference between the generated and target geometries within the gym, are used to update the RL network. Our method supports operations beyond sketches, Boolean, and extrusion, including revolution operations. With this training gym, we achieve state-of-the-art (SOTA) quality in generating command sequences from B-Rep geometries.

cs.LG↗

DeepRTE: Pre-trained Attention-based Neural Network for Radiative Transfer

In this paper, we propose a novel neural network approach, termed DeepRTE, to address the steady-state Radiative Transfer Equation (RTE). The RTE is a differential-integral equation that governs the propagation of radiation through a participating medium, with applications spanning diverse domains such as neutron transport, atmospheric radiative transfer, heat transfer, and optical imaging. Our DeepRTE framework demonstrates superior computational efficiency for solving the steady-state RTE, surpassing traditional methods and existing neural network approaches. This efficiency is achieved by embedding physical information through derivation of the RTE and mathematically-informed network architecture. Concurrently, DeepRTE achieves high accuracy with significantly fewer parameters, largely due to its incorporation of mechanisms such as multi-head attention. Furthermore, DeepRTE is a mesh-free neural operator framework with inherent zero-shot capability. This is achieved by incorporating Green's function theory and pre-training with delta-function inflow boundary conditions into both its architecture design and training data construction. The efficacy of the proposed approach is substantiated through comprehensive numerical experiments.

cs.LG↗

Random ordinate method for mitigating the ray effect in radiative transport equation simulations

The Discrete Ordinates Method (DOM) is the most widely used velocity discretization method for simulating the radiative transport equation. However, the ray effect is a long-standing drawback of DOM. In benchmark tests that exhibit the ray effect, we observe low regularity in the velocity variable of the solution. To address this issue, we propose a Random Ordinate Method (ROM) to mitigate the ray effect. Compared to other strategies proposed in the literature for mitigating the ray effect, ROM offers several advantages: 1) For benchmark tests that exhibit ray effect, the computational cost is lower than that of the DOM; 2) it is simple and requires minimal changes to existing DOM-based code; 3) it is easily parallelizable and independent of the problem setup. A formal analysis is presented for the convergence orders of the error and bias. Numerical tests demonstrate the reduction in computational cost compared to DOM, as well as its effectiveness in mitigating the ray effect.

math.DS↗

Convergence Analysis of the Random Ordinate Method for Mitigating the Ray Effect

The Discrete Ordinates Method (DOM) is widely used for velocity discretization in radiative transport simulations. However, DOM tends to exhibit the ray effect when the velocity discretization is not sufficiently refined, a limitation that is well documented. To counter this, we have developed the Random Ordinates Method (ROM) by integrating randomness into the velocity discretization, which mitigates the ray effect without incurring additional computational costs. ROM partitions the velocity space into n cells, selects a random ordinate from each cell, and solves a DOM system with these ordinates. It leverages the average of multiple samples to achieve a higher convergence order, especially for solutions with low regularity in the velocity variable. In this work, we provide a detailed convergence analysis for ROM, focusing on bias and single-run errors. This analysis is crucial for determining the necessary mesh size and the optimal number of samples required to attain a specified level of accuracy.

math.NA↗

Privacy Risks of LLM-Empowered Recommender Systems: An Inversion Attack Perspective

The large language model (LLM) powered recommendation paradigm has been proposed to address the limitations of traditional recommender systems, which often struggle to handle cold start users or items with new IDs. Despite its effectiveness, this study uncovers that LLM empowered recommender systems are vulnerable to reconstruction attacks that can expose both system and user privacy. To examine this threat, we present the first systematic study on inversion attacks targeting LLM empowered recommender systems, where adversaries attempt to reconstruct original prompts that contain personal preferences, interaction histories, and demographic attributes by exploiting the output logits of recommendation models. We reproduce the vec2text framework and optimize it using our proposed method called Similarity Guided Refinement, enabling more accurate reconstruction of textual prompts from model generated logits. Extensive experiments across two domains (movies and books) and two representative LLM based recommendation models demonstrate that our method achieves high fidelity reconstructions. Specifically, we can recover nearly 65 percent of the user interacted items and correctly infer age and gender in 87 percent of the cases. The experiments also reveal that privacy leakage is largely insensitive to the victim model's performance but highly dependent on domain consistency and prompt complexity. These findings expose critical privacy vulnerabilities in LLM empowered recommender systems.

cs.IR↗

Generating Feasible and Diverse Synthetic Populations Using Diffusion Models

Population synthesis is a critical task that involves generating synthetic yet realistic representations of populations. It is a fundamental problem in agent-based modeling (ABM), which has become the standard to analyze intelligent transportation systems. The synthetic population serves as the primary input for ABM transportation simulation, with traveling agents represented by population members. However, when the number of attributes describing agents becomes large, survey data often cannot densely support the joint distribution of the attributes in the population due to the curse of dimensionality. This sparsity makes it difficult to accurately model and produce the population. Interestingly, deep generative models trained from available sample data can potentially synthesize possible attribute combinations that present in the actual population but do not exist in the sample data(called sampling zeros). Nevertheless, this comes at the cost of falsely generating the infeasible attribute combinations that do not exist in the population (called structural zeros). In this study, a novel diffusion model-based population synthesis method is proposed to estimate the underlying joint distribution of a population. This approach enables the recovery of numerous missing sampling zeros while keeping the generated structural zeros minimal. Our method is compared with other recently proposed approaches such as Variational Autoencoders (VAE) and Generative Adversarial Network (GAN) approaches, which have shown success in high dimensional tabular population synthesis. We assess the performance of the synthesized outputs using a range of metrics, including marginal distribution similarity, feasibility, and diversity. The results demonstrate that our proposed method outperforms previous approaches in achieving a better balance between the feasibility and diversity of the synthesized population.

cs.LG↗

A Dual Radiomic and Dosiomic Filtering Technique for Locoregional Radiation Pneumonitis Prediction in Breast Cancer Patients

Purpose: Radiation pneumonitis (RP) is a serious complication of intensity-modulated radiation therapy (IMRT) for breast cancer patients, underscoring the need for precise and explainable predictive models. This study presents an Explainable Dual-Omics Filtering (EDOF) model that integrates spatially localized dosiomic and radiomic features for voxel-level RP prediction. Methods: A retrospective cohort of 72 breast cancer patients treated with IMRT was analyzed, including 28 who developed RP. The EDOF model consists of two components: (1) dosiomic filtering, which extracts local dose intensity and spatial distribution features from planning dose maps, and (2) radiomic filtering, which captures texture-based features from pre-treatment CT scans. These features are jointly analyzed using the Explainable Boosting Machine (EBM), a transparent machine learning model that enables feature-specific risk evaluation. Model performance was assessed using five-fold cross-validation, reporting area under the curve (AUC), sensitivity, and specificity. Feature importance was quantified by mean absolute scores, and Partial Dependence Plots (PDPs) were used to visualize nonlinear relationships between RP risk and dual-omic features. Results: The EDOF model achieved strong predictive performance (AUC = 0.95 +- 0.01; sensitivity = 0.81 +- 0.05). The most influential features included dosiomic Intensity Mean, dosiomic Intensity Mean Absolute Deviation, and radiomic SRLGLE. PDPs revealed that RP risk increases beyond 5 Gy and rises sharply between 10-30 Gy, consistent with clinical dose thresholds. SRLGLE also captured structural heterogeneity linked to RP in specific lung regions. Conclusion: The EDOF framework enables spatially resolved, explainable RP prediction and may support personalized radiation planning to mitigate pulmonary toxicity.

physics.med-ph↗

Deriving sub-diffusion equations

Sub-diffusion equations are used in a large range of applications including fluids, plasma physics and biology. Their mathematical analysis is advanced even if a much larger literature addresses super-diffusions. The goal of this paper is to provide the microscopic mechanism and rigorous derivation of sub-diffusions when the waiting time distribution of particles follows an age-structured equation and jumps occur at each renewal. The major difficulty to recover sub-diffusions, unlike normal diffusions, is that the assumption of long waiting time implies lack of integrability for the age equilibrium. This prevents to establish strong a priori estimates. Here, the Laplace transform plays the role that Fourier transform plays for the more traditional case of fast diffusions.

math.AP↗

Flow Matching based Sequential Recommender Model

Generative models, particularly diffusion model, have emerged as powerful tools for sequential recommendation. However, accurately modeling user preferences remains challenging due to the noise perturbations inherent in the forward and reverse processes of diffusion-based methods. Towards this end, this study introduces FMRec, a Flow Matching based model that employs a straight flow trajectory and a modified loss tailored for the recommendation task. Additionally, from the diffusion-model perspective, we integrate a reconstruction loss to improve robustness against noise perturbations, thereby retaining user preferences during the forward process. In the reverse process, we employ a deterministic reverse sampler, specifically an ODE-based updating function, to eliminate unnecessary randomness, thereby ensuring that the generated recommendations closely align with user needs. Extensive evaluations on four benchmark datasets reveal that FMRec achieves an average improvement of 6.53% over state-of-the-art methods. The replication code is available at https://github.com/FengLiu-1/FMRec.

cs.IR↗

A deterministic solver for the linear Boltzmann model of a single mono-directional proton beam

The linear Boltzmann model for proton beams is a six-dimensional partial differential equation (PDE). We propose a deterministic solver for the linear Boltzmann model based on scattering decomposition and depth-splitting methods. The main idea is to first divide the protons into primary protons and scattering protons, whose equations are derived using the source iteration method. We then treat depth as the time variable in classical time-evolutionary problems and apply the depth-splitting method. In the depth-splitting method, the full operator is decomposed into three parts, with each subsystem being easily parallelizable, which is crucial for efficient simulations. The resulting discretization exhibits second-order convergence in both the depth and energy variables. The dose distributions obtained from our solver are compared with those from Monte Carlo simulations for various materials and heterogeneous cases.

math.NA↗

Random Source Iteration Method: Mitigating the Ray Effect in the Discrete Ordinates Method

The commonly used velocity discretization for simulating the radiative transport equation (RTE) is the discrete ordinates method (DOM). One of the long-standing drawbacks of DOM is the phenomenon known as the ray effect. Due to the high dimensionality of the RTE, DOM results in a large algebraic system to solve. The Source Iteration (SI) method is the most standard iterative method for solving this system. In this paper, by introducing randomness into the SI method, we propose a novel random source iteration (RSI) method that offers a new way to mitigate the ray effect without increasing the computational cost. We have rigorously proved that RSI is unbiased with respect to the SI method and that its variance is uniformly bounded across iteration steps; thus, the convergence order with respect to the number of samples is $1/2$. Furthermore, we prove that the RSI iteration process, as a Markov chain, is ergodic under mild assumptions. Numerical examples are presented to demonstrate the convergence of RSI and its effectiveness in mitigating the ray effect.

math.NA↗

AdaSociety: An Adaptive Environment with Social Structures for Multi-Agent Decision-Making

Traditional interactive environments limit agents' intelligence growth with fixed tasks. Recently, single-agent environments address this by generating new tasks based on agent actions, enhancing task diversity. We consider the decision-making problem in multi-agent settings, where tasks are further influenced by social connections, affecting rewards and information access. However, existing multi-agent environments lack a combination of adaptive physical surroundings and social connections, hindering the learning of intelligent behaviors. To address this, we introduce AdaSociety, a customizable multi-agent environment featuring expanding state and action spaces, alongside explicit and alterable social structures. As agents progress, the environment adaptively generates new tasks with social structures for agents to undertake. In AdaSociety, we develop three mini-games showcasing distinct social structures and tasks. Initial results demonstrate that specific social structures can promote both individual and collective benefits, though current reinforcement learning and LLM-based algorithms show limited effectiveness in leveraging social structures to enhance performance. Overall, AdaSociety serves as a valuable research platform for exploring intelligence in diverse physical and social settings. The code is available at https://github.com/bigai-ai/AdaSociety.

cs.MA↗

Crossover from ballistic transport to normal diffusion: a kinetic view

The crossover between dispersion patterns has been frequently observed in various systems. Inspired by the pathway-based kinetic model for E. coli chemotaxis that accounts for the intracellular adaptation process and noise, we propose a kinetic model that can exhibit a crossover from ballistic transport to normal diffusion at the population level. At the particle level, this framework aligns with a stochastic individual-based model. Using numerical simulations and rigorous asymptotic analysis, we demonstrate this crossover both analytically and computationally. Notably, under suitable scaling, the model reveals two distinct limits in which the macroscopic densities exhibit either ballistic transport or normal diffusion.

math.AP↗

Improved conditional gradient method for the generalized cone order optimization problem on the local sphere

In this paper, a generalized optimization problem on the local sphere is established by the cone order relation on the tangent space, and solved by an improved conditional gradient method (for short, ICGM). The auxiliary subproblems are constructed by the directed distance function on the tangent space, the iteration step size is updated by the Armijo rule, and the convergence of the ICGM is proved without the convexity of the objective function. Under the assumption of convexity, the clusters of the sequence generated by the ICGM are proved to be the spherical weakly Pareto solutions (also known as weakly efficient solutions) of this problem .

math.OC↗

Image Space Analysis to the generalized optimization problem on a local sphere

This paper introduces and studies the generalized optimization problem (for short, GOP) defined by the conic order relation on a local sphere. The existence of solution to this problem is studied by using image space analysis (for short, ISA), and a class of regular weak separation functions on the local sphere is established. Moreover, a Lagrangian-type sufficient optimality condition and a saddle-point-type necessary optimality condition for GOP is obtained by a second class of weak separation functions, which are based on the Gerstewitz function and the directional distance function. The problem is transformed into a solvable real-valued optimization problem using scalarization methods.

math.OC↗

Reconstructing the kinetic chemotaxis kernel using macroscopic data: well-posedness and ill-posedness

Bacterial motion is steered by external stimuli (chemotaxis), and the motion described on the mesoscopic scale is uniquely determined by a parameter $K$ that models velocity change response from the bacteria. This parameter is called chemotaxis kernel. In a practical setting, it is inferred by experimental data. We deploy a PDE-constrained optimization framework to perform this reconstruction using velocity-averaged, localized data taken in the interior of the domain. The problem can be well-posed or ill-posed depending on the data preparation and the experimental setup. In particular, we propose one specific design that guarantees numerical reconstructability and local convergence. This design is adapted to the discretization of $K$ in space and decouples the reconstruction of local values of $K$ into smaller cell problems, opening up parallelization opportunities. Numerical evidences support the theoretical findings.

math.NA↗