Searcharxiv⌕ Search

arXiv subjects

Chang Yang

Publications and source records attributed to Chang Yang.

At least 37 records · Page 2Linked to original sources

Each Graph is a New Language: Graph Learning with LLMs

Recent efforts leverage Large Language Models (LLMs) for modeling text-attributed graph structures in node classification tasks. These approaches describe graph structures for LLMs to understand or aggregate LLM-generated textual attribute embeddings through graph structure. However, these approaches face two main limitations in modeling graph structures with LLMs. (i) Graph descriptions become verbose in describing high-order graph structure. (ii) Textual attributes alone do not contain adequate graph structure information. It is challenging to model graph structure concisely and adequately with LLMs. LLMs lack built-in mechanisms to model graph structures directly. They also struggle with complex long-range dependencies between high-order nodes and target nodes. Inspired by the observation that LLMs pre-trained on one language can achieve exceptional performance on another with minimal additional training, we propose \textbf{G}raph-\textbf{D}efined \textbf{L}anguage for \textbf{L}arge \textbf{L}anguage \textbf{M}odel (GDL4LLM). This novel framework enables LLMs to transfer their powerful language understanding capabilities to graph-structured data. GDL4LLM translates graphs into a graph language corpus instead of graph descriptions and pre-trains LLMs on this corpus to adequately understand graph structures. During fine-tuning, this corpus describes the structural information of target nodes concisely with only a few tokens. By treating graphs as a new language, GDL4LLM enables LLMs to model graph structures adequately and concisely for node classification tasks. Extensive experiments on three real-world datasets demonstrate that GDL4LLM outperforms description-based and textual attribute embeddings-based baselines by efficiently modeling different orders of graph structure with LLMs.

cs.CL↗

Resolving Latency and Inventory Risk in Market Making with Reinforcement Learning

The latency of the exchanges in Market Making (MM) is inevitable due to hardware limitations, system processing times, delays in receiving data from exchanges, the time required for order transmission to reach the market, etc. Existing reinforcement learning (RL) methods for Market Making (MM) overlook the impact of these latency, which can lead to unintended order cancellations due to price discrepancies between decision and execution times and result in undesired inventory accumulation, exposing MM traders to increased market risk. Therefore, these methods cannot be applied in real MM scenarios. To address these issues, we first build a realistic MM environment with random delays of 30-100 milliseconds for order placement and market information reception, and implement a batch matching mechanism that collects orders within every 500 milliseconds before matching them all at once, simulating the batch auction mechanisms adopted by some exchanges. Then, we propose Relaver, an RL-based method for MM to tackle the latency and inventory risk issues. The three main contributions of Relaver are: i) we introduce an augmented state-action space that incorporates order hold time alongside price and volume, enabling Relaver to optimize execution strategies under latency constraints and time-priority matching mechanisms, ii) we leverage dynamic programming (DP) to guide the exploration of RL training for better policies, iii) we train a market trend predictor, which can guide the agent to intelligently adjust the inventory to reduce the risk. Extensive experiments and ablation studies on four real-world datasets demonstrate that \textsc{Relaver} significantly improves the performance of state-of-the-art RL-based MM strategies across multiple metrics.

cs.LG↗

FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs

Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, and tool limitations. Although large language models (LLMs) have succeeded in various natural language processing tasks and have shown potential in automating workflows through reasoning and contextual understanding, current benchmarks for evaluating LLMs in finance lack sufficient domain-specific data, have simplistic task design, and incomplete evaluation frameworks. To address these gaps, this article presents FinMaster, a comprehensive financial benchmark designed to systematically assess the capabilities of LLM in financial literacy, accounting, auditing, and consulting. Specifically, FinMaster comprises three main modules: i) FinSim, which builds simulators that generate synthetic, privacy-compliant financial data for companies to replicate market dynamics; ii) FinSuite, which provides tasks in core financial domains, spanning 183 tasks of various types and difficulty levels; and iii) FinEval, which develops a unified interface for evaluation. Extensive experiments over state-of-the-art LLMs reveal critical capability gaps in financial reasoning, with accuracy dropping from over 90% on basic tasks to merely 40% on complex scenarios requiring multi-step reasoning. This degradation exhibits the propagation of computational errors, where single-metric calculations initially demonstrating 58% accuracy decreased to 37% in multimetric scenarios. To the best of our knowledge, FinMaster is the first benchmark that covers full-pipeline financial workflows with challenging tasks. We hope that FinMaster can bridge the gap between research and industry practitioners, driving the adoption of LLMs in real-world financial practices to enhance efficiency and accuracy.

cs.AI↗

Why Regression? Binary Encoding Classification Brings Confidence to Stock Market Index Price Prediction

Stock market indices serve as fundamental market measurement that quantify systematic market dynamics. However, accurate index price prediction remains challenging, primarily because existing approaches treat indices as isolated time series and frame the prediction as a simple regression task. These methods fail to capture indices' inherent nature as aggregations of constituent stocks with complex, time-varying interdependencies. To address these limitations, we propose Cubic, a novel end-to-end framework that explicitly models the adaptive fusion of constituent stocks for index price prediction. Our main contributions are threefold. i) Fusion in the latent space: we introduce the fusion mechanism over the latent embedding of the stocks to extract the information from the vast number of stocks. ii) Binary encoding classification: since regression tasks are challenging due to continuous value estimation, we reformulate the regression into the classification task, where the target value is converted to binary and we optimize the prediction of the value of each digit with cross-entropy loss. iii) Confidence-guided prediction and trading: we introduce the regularization loss to address market prediction uncertainty for the index prediction and design the rule-based trading policies based on the confidence. Extensive experiments across multiple stock markets and indices demonstrate that Cubic consistently outperforms state-of-the-art baselines in stock index prediction tasks, achieving superior performance on both forecasting accuracy metrics and downstream trading profitability.

q-fin.ST↗

Precise measurement of CP violating $τ$ EDM through $e^+ e^- \to γ^*, ψ(2s) \to τ^+ τ^-$

A nonzero electric dipole moment of a tauon, $d_τ$, signals CP violation and provides an important probe for new physics. We study methods to measure $d_τ$ at low energy $e^+ e^-$ colliders through the processes $e^+e^- \to γ^*, ψ(2S) \to τ^+τ^-$ with $τ^\pm$ decays into a charged hadron and a tau neutrino. We point out that, with measuring energies of the charged hadron, Im$(d_τ)$ can be measured. On the other hand, selecting events of $τ$ decays after traveling more than the detector resolution distance, Re$(d_τ)$ can also be determined. We find that the precision at Super Tau-Charm Facility (STCF) running at the center energy of $m_{ψ(2S)}$ for 10 year data accumulation, the precision of Im$(d_τ)$ and Re$(d_τ)$ are found to be 1.8 and 11 in unit of $ 10^{-18}~e\,\text{cm}$, respectively. The sensitivity for $d_τ$ measurement precision at the STCF can be reached its optimum at a central energy of $6.3~\text{GeV}$, achieving a precision of $0.7$ for Im$(d_τ)$ and $2.8$ for Re$(d_τ)$ in unit of $ 10^{-18}~e\,\text{cm}$.

hep-ph↗

Multiscale numerical methods for isothermal fluid models of confined plasmas

The aim of this work is to introduce a numerical method to cope with the multiscale nature of confined plasma physics. These investigations are focused on fluid plasma description under large magnetic field. The difficulties in this context stem from intense magnetization of the plasma, inducing a severe anisotropy, possible quasi-neutrality breakdowns, which may occur locally in the plasma and, eventually, the drift regime which prevails for the description of the electrons. These characteristics bring small parameters compared to the scale of the studied device. This work is therefore devoted to highlighting the difficulties specific to this context and to developing numerical methods efficient to cope with this multiscale nature of the physics within the framework of asymptotic-preserving methods.

physics.plasm-ph↗

The Scaling Law for LoRA Base on Mutual Information Upper Bound

LoRA (Low-Rank Adaptation) is a widely used model fine-tuning method. In fine-tuning, the law among model performance, model parameters, and data complexity has been a focal issue in the field. Existing methods often leverage external metrics (such as cross-entropy or perplexity) to evaluate model performance. In the fine-tuning process for large models, two types of knowledge are typically involved: the frozen, general knowledge acquired by the model during pre-training and the new knowledge learned through the LoRA module from the current data. Generally, the less LoRA's learned knowledge relies on the large model, the more it captures the specific knowledge of new data, thereby enhancing its adaptability to new tasks. However, external metrics do not readily capture the dependency relationship between these two types of knowledge. Therefore, we designed an internal metric based on the Mutual Information Upper Bound (MIUB) theory to investigate the scaling law of large-model LoRA fine-tuning. In our experiments, we validated this approach on benchmark datasets, using the Llama3-8B and Phi3-3B models. The results show that the proposed MIUB metric aligns more accurately and stably with the scaling law of LoRA fine-tuning compared to cross-entropy and perplexity.

cs.LG↗

In-Context Exploiter for Extensive-Form Games

Nash equilibrium (NE) is a widely adopted solution concept in game theory due to its stability property. However, we observe that the NE strategy might not always yield the best results, especially against opponents who do not adhere to NE strategies. Based on this observation, we pose a new game-solving question: Can we learn a model that can exploit any, even NE, opponent to maximize their own utility? In this work, we make the first attempt to investigate this problem through in-context learning. Specifically, we introduce a novel method, In-Context Exploiter (ICE), to train a single model that can act as any player in the game and adaptively exploit opponents entirely by in-context learning. Our ICE algorithm involves generating diverse opponent strategies, collecting interactive history training data by a reinforcement learning algorithm, and training a transformer-based agent within a well-designed curriculum learning framework. Finally, comprehensive experimental results validate the effectiveness of our ICE algorithm, showcasing its in-context learning ability to exploit any unknown opponent, thereby positively answering our initial game-solving question.

cs.AI↗

The sign of linear periods

Let $G$ be a group with subgroup $H$, and let $(π,V)$ be a complex representation of $G$. The natural action of the normalizer $N$ of $H$ in $G$ on the space $\mathrm{Hom}_H(π,\mathbb{C})$ of $H$-invariant linear forms on $V$, provides a representation $χ_π$ of $N$ trivial on $H$, which is a character when $\mathrm{Hom}_H(π,\mathbb{C})$ is one dimensional. If moreover $G$ is a reductive group over a local field, and $π$ is smooth irreducible, it is an interesting problem to express $χ_π$ in terms of the possibly conjectural Langlands parameter $ϕ_π$ of $π$. In this paper we consider the following situation: $G=\mathrm{GL}_m(D)$ for $D$ a central division algebra of dimension $d^2$ over a local field $F$ of characteristic zero, $H$ is the centralizer of a non central element $δ\in G$ such that $δ^2$ is in the center of $G$, and $π$ has generic Jacquet-Langlands transfer to $\mathrm{GL}_{md}(F)$. In this setting the space $\mathrm{Hom}_H(π,\mathbb{C})$ is at most one dimensional. When $\mathrm{Hom}_H(π,\mathbb{C})\simeq \mathbb{C}$ and $H\neq N$, we prove that the value of the $χ_π$ on the non trivial class of $\frac{N}{H}$ is $(-1)^mε(ϕ_π)$ where $ε(ϕ_π)$ is the root number of $ϕ_π$. Along the way we extend many useful multiplicity one results for linear and Shalika models to the case of non split $G$. When $F$ is $p$-adic we also classify standard modules with linear periods and Shalika models, which are new results even when $D=F$.

math.RT↗

Configurable Mirror Descent: Towards a Unification of Decision Making

Decision-making problems, categorized as single-agent, e.g., Atari, cooperative multi-agent, e.g., Hanabi, competitive multi-agent, e.g., Hold'em poker, and mixed cooperative and competitive, e.g., football, are ubiquitous in the real world. Various methods are proposed to address the specific decision-making problems. Despite the successes in specific categories, these methods typically evolve independently and cannot generalize to other categories. Therefore, a fundamental question for decision-making is: \emph{Can we develop \textbf{a single algorithm} to tackle \textbf{ALL} categories of decision-making problems?} There are several main challenges to address this question: i) different decision-making categories involve different numbers of agents and different relationships between agents, ii) different categories have different solution concepts and evaluation measures, and iii) there lacks a comprehensive benchmark covering all the categories. This work presents a preliminary attempt to address the question with three main contributions. i) We propose the generalized mirror descent (GMD), a generalization of MD variants, which considers multiple historical policies and works with a broader class of Bregman divergences. ii) We propose the configurable mirror descent (CMD) where a meta-controller is introduced to dynamically adjust the hyper-parameters in GMD conditional on the evaluation measures. iii) We construct the \textsc{GameBench} with 15 academic-friendly games across different decision-making categories. Extensive experiments demonstrate that CMD achieves empirically competitive or better outcomes compared to baselines while providing the capability of exploring diverse dimensions of decision making.

cs.AI↗

Reinforcement Nash Equilibrium Solver

Nash Equilibrium (NE) is the canonical solution concept of game theory, which provides an elegant tool to understand the rationalities. Though mixed strategy NE exists in any game with finite players and actions, computing NE in two- or multi-player general-sum games is PPAD-Complete. Various alternative solutions, e.g., Correlated Equilibrium (CE), and learning methods, e.g., fictitious play (FP), are proposed to approximate NE. For convenience, we call these methods as "inexact solvers", or "solvers" for short. However, the alternative solutions differ from NE and the learning methods generally fail to converge to NE. Therefore, in this work, we propose REinforcement Nash Equilibrium Solver (RENES), which trains a single policy to modify the games with different sizes and applies the solvers on the modified games where the obtained solution is evaluated on the original games. Specifically, our contributions are threefold. i) We represent the games as $α$-rank response graphs and leverage graph neural network (GNN) to handle the games with different sizes as inputs; ii) We use tensor decomposition, e.g., canonical polyadic (CP), to make the dimension of modifying actions fixed for games with different sizes; iii) We train the modifying strategy for games with the widely-used proximal policy optimization (PPO) and apply the solvers to solve the modified games, where the obtained solution is evaluated on original games. Extensive experiments on large-scale normal-form games show that our method can further improve the approximation of NE of different solvers, i.e., $α$-rank, CE, FP and PRD, and can be generalized to unseen games.

cs.GT↗

Self-adaptive PSRO: Towards an Automatic Population-based Game Solver

Policy-Space Response Oracles (PSRO) as a general algorithmic framework has achieved state-of-the-art performance in learning equilibrium policies of two-player zero-sum games. However, the hand-crafted hyperparameter value selection in most of the existing works requires extensive domain knowledge, forming the main barrier to applying PSRO to different games. In this work, we make the first attempt to investigate the possibility of self-adaptively determining the optimal hyperparameter values in the PSRO framework. Our contributions are three-fold: (1) Using several hyperparameters, we propose a parametric PSRO that unifies the gradient descent ascent (GDA) and different PSRO variants. (2) We propose the self-adaptive PSRO (SPSRO) by casting the hyperparameter value selection of the parametric PSRO as a hyperparameter optimization (HPO) problem where our objective is to learn an HPO policy that can self-adaptively determine the optimal hyperparameter values during the running of the parametric PSRO. (3) To overcome the poor performance of online HPO methods, we propose a novel offline HPO approach to optimize the HPO policy based on the Transformer architecture. Experiments on various two-player zero-sum games demonstrate the superiority of SPSRO over different baselines.

cs.AI↗

Complete determination of $SU(3)_F$ amplitudes and strong phase in $Λ_c^+ \to Ξ^0 K^+$

The BESIII collaboration has recently reported the first time measurement of the decay asymmetry $α(Λ_c^+ \to Ξ^0 K^+) = 0.01 \pm 0.16(stat.) \pm 0.03(syst.)$ and also a sizable phase shift of $δ_P-δ_S = -1.55 \pm 0.25$ or $1.59\pm 0.25$ between S- and P-wave amplitudes. This implies significant strong phase shifts in the decay amplitudes. The strong phases indicate the existence of rescattering or loop effects, which are challenging to calculate due to non-perturbative effects. By employing the flavor $SU(3)_F$ symmetry and applying the Körner-Pati-Woo theorem to reduce the number of parameters, we find that the current data already allow us to obtain, for the first time, model-independent decay amplitudes and their strong phases. The establishment of the existence of sizable strong phases opens a window for future investigations into CP violation. In our fit, a notable discrepancy emerges in the branching ratio of $Ξ_c^0 \to Ξ^- π^+$. The direct relationship between $Γ(Λ_c^+ \to Λe^+ν_e)$ and $Γ(Ξ_c^0 \to Ξ^- e^+ν_e)$, along with newly discovered $SU(3)_F$ relations, collectively suggests an underestimation of $\mathcal{B}(Ξ_c^0 \to Ξ^- π^+)$ in experimental findings.

hep-ph↗

Ext-distinction for $p$-adic symmetric spaces

Let $G/H$ be a $p$-adic symmetric space. We compute explicitly the higher relative extension groups for all discrete series representations of $G$ in two examples: the symplectic case and the linear case. The results have immediate applications to the computation of the Euler-Poincaré pairing, the alternating sum of the dimensions of the Ext-groups. In the linear case we confirm a conjecture of Wan which asserts the equality of the Euler-Poincaré pairing with the geometric multiplicity for any irreducible representation of $G$. In the symplectic case we reduce the verification of Wan's conjecture to the case of discrete series representations. We also determine all relatively supercuspidal representations in both cases.

math.RT↗

Global analysis of measured and unmeasured hadronic two-body weak decays of antitriplet charmed baryons

A large amount of data on hadronic two body weak decays of anti-triplet charmed baryons $T_{c\bar 3}$ to an octet baryon $T_8$ and an octet or singlet pseudoscalar meson $P$, $T_{c \bar 3} \to T_8 P$, have been measured. The SU(3) flavor symmetry has been applied to study these decays to obtain insights about weak interactions for charm physics. However not all such decays needed to determine the SU(3) irreducible amplitudes have been measured forbidding a complete global analysis. Previously, it has been shown that data from measured decays can be used to do a global fit to determine all except one parity violating and one parity conserving amplitudes of the relevant SU(3) irreducible amplitudes causing 8 hadronic two body weak decay channels involving $Ξ^0_c$ to $η$ or $η'$ transitions undetermined. It is important to obtain information about these decays in order to guide experimental searches. In this work using newly measured decay modes by BESIII and Belle in 2022, we carry out a global analysis and parameterize the unknown amplitudes to provide the ranges for the branching ratios of the 8 undetermined decays. Our results indicate that the SU(3) flavor symmetry can explain the measured data exceptionally well, with a remarkable minimal $χ^2/d.o.f.$ of 1.21 and predict 80 observables in 45 decays for future experimental data to test. We then vary the unknown SU(3) amplitudes to obtain the allowed range of branching ratios for the 8 undetermined decays. We find that some of them are within reach of near future experimental capabilities. We urge our experimental colleagues to carry out related searches.

hep-ph↗

On local intertwining periods

We prove the absolute convergence, functional equations and meromorphic continuation of local intertwining periods on parabolically induced representations of finite length for certain symmetric spaces over local fields of characteristic zero, including Galois pairs as well as pairs of Prasad and Takloo-Bighash type. Furthermore, for a general symmetric space we prove a sufficient condition for distinction of an induced representation in terms of distinction of its inducing data. Both results generalize previous results of the first two named authors. In particular, for both we remove a boundedness assumption on the inducing data and for the second we further remove any assumption on the symmetric space. Moreover, when the inducing representation is uniformly bounded, we extend the field of cofficients from p-adic to any local field of characteristic zero. In fact this extension holds for all finite length representations under a natural generic irreducibility assumption for parabolic induction. In the case of p-adic symmetric spaces, combined with the necessary conditions for distinction that follow from the geometric lemma, this provides a necessary and sufficient condition for distinction of representations induced from cuspidal.

math.RT↗

Two-body hadronic decays of Xi^{0}_c in light front approach

In this study, we investigate the nonleptonic decays of the charmed-baryon Xi^ {0} _ c induced by the c -> u(d\bar{d})/(s\bar{s}) transition. Utilizing the factorization assumption, we decompose the decay amplitudes in terms of transition form factors which are then calculated within the light-front quark model. We employ helicity amplitudes to analyze the nonleptonic decay modes of the charmed-baryon Xi^ {0} _c and derive benchmark results for decay widths and branching fractions. Our calculations suggest that the branching fractions for some of these rare nonleptonic decays are at the order of 10^ {-4} - 10^ {-3}, which are likely to be detectable at experiments such as LHCb or BESIII. The potential data accumulated in the future may help to further our understanding of the decay mechanism in the presence of charm quarks.

hep-ph↗

Residual spectrum of $\mathrm{GL}_{2n}$ distinguished by $\mathrm{GL}_n \times \mathrm{GL}_n$

Following the regularization method presented by Zydor, we study in this paper the regularized linear periods of square-integrable automormphic forms on $\mathrm{GL}_{2n}(\mathbb{A}_F)$, where $F$ is a number field and $\mathbb{A}_F$ its ring of adeles. We obtain a formula that expresses the regularized period of a noncuspidal, square-integrable automorphic form in terms of degenerate Whittaker functions in an inductive manner. As a consequence we characterize irreducible automorphic representations in the discrete spectrum of $\mathrm{GL}_{2n}(\mathbb{A})$ that are distinguished by $\mathrm{GL}_n(\mathbb{A}) \times \mathrm{GL}_n(\mathbb{A})$. We also show the vanishing of the regularized periods of square-integrable automorphic forms on $\mathrm{GL}_n(\mathbb{A})$ over $\mathrm{GL}_p(\mathbb{A}) \times \mathrm{GL}_q(\mathbb{A})$ when $p$ is not equal to $q$.

math.NT↗