SearcharxivSearch

arXiv subjects

Pan Yang

Publications and source records attributed to Pan Yang.

12 recordsLinked to original sources

CAMS: Towards Compositional Zero-Shot Learning via Gated Cross-Attention and Multi-Space Disentanglement

Compositional zero-shot learning (CZSL) aims to learn the concepts of attributes and objects in seen compositions and to recognize their unseen compositions. Most Contrastive Language-Image Pre-training (CLIP)-based CZSL methods focus on disentangling attributes and objects by leveraging the global semantic representation obtained from the image encoder. However, this representation has limited representational capacity and do not allow for complete disentanglement of the two. To this end, we propose CAMS, which aims to extract semantic features from visual features and perform semantic disentanglement in multidimensional spaces, thereby improving generalization over unseen attribute-object compositions. Specifically, CAMS designs a Gated Cross-Attention that captures fine-grained semantic features from the high-level image encoding blocks of CLIP through a set of latent units, while adaptively suppressing background and other irrelevant information. Subsequently, it conducts Multi-Space Disentanglement to achieve disentanglement of attribute and object semantics. Experiments on three popular benchmarks (MIT-States, UT-Zappos, and C-GQA) demonstrate that CAMS achieves state-of-the-art performance in both closed-world and open-world settings. The code is available at https://github.com/ybyangjing/CAMS.

cs.CV

KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction

In this paper, we propose KnowCoder, a Large Language Model (LLM) to conduct Universal Information Extraction (UIE) via code generation. KnowCoder aims to develop a kind of unified schema representation that LLMs can easily understand and an effective learning framework that encourages LLMs to follow schemas and extract structured knowledge accurately. To achieve these, KnowCoder introduces a code-style schema representation method to uniformly transform different schemas into Python classes, with which complex schema information, such as constraints among tasks in UIE, can be captured in an LLM-friendly manner. We further construct a code-style schema library covering over $\textbf{30,000}$ types of knowledge, which is the largest one for UIE, to the best of our knowledge. To ease the learning process of LLMs, KnowCoder contains a two-phase learning framework that enhances its schema understanding ability via code pretraining and its schema following ability via instruction tuning. After code pretraining on around $1.5$B automatically constructed data, KnowCoder already attains remarkable generalization ability and achieves relative improvements by $\textbf{49.8%}$ F1, compared to LLaMA2, under the few-shot setting. After instruction tuning, KnowCoder further exhibits strong generalization ability on unseen schemas and achieves up to $\textbf{12.5%}$ and $\textbf{21.9%}$, compared to sota baselines, under the zero-shot setting and the low resource setting, respectively. Additionally, based on our unified schema representations, various human-annotated datasets can simultaneously be utilized to refine KnowCoder, which achieves significant improvements up to $\textbf{7.5%}$ under the supervised setting.

cs.LG

Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub

Large Language Models (LLMs) excel in traditional natural language processing tasks but struggle with problems that require complex domain-specific calculations or simulations. While equipping LLMs with external tools to build LLM-based agents can enhance their capabilities, existing approaches lack the flexibility to address diverse and ever-evolving user queries in open domains. Currently, there is also no existing dataset that evaluates LLMs on open-domain knowledge that requires tools to solve. To this end, we introduce OpenAct benchmark to evaluate the open-domain task-solving capability, which is built on human expert consultation and repositories in GitHub. It comprises 339 questions spanning 7 diverse domains that need to be solved with domain-specific methods. In our experiments, even state-of-the-art LLMs and LLM-based agents demonstrate unsatisfactory success rates, underscoring the need for a novel approach. Furthermore, we present OpenAgent, a novel LLM-based agent system that can tackle evolving queries in open domains through autonomously integrating specialized tools from GitHub. OpenAgent employs 1) a hierarchical framework where specialized agents handle specific tasks and can assign tasks to inferior agents, 2) a bi-level experience learning mechanism to learn from both humans' and its own experiences to tackle tool flaws. Experiments demonstrate its superior effectiveness and efficiency, which significantly outperforms baselines. Our data and code are open-source at https://github.com/OpenBMB/OpenAct.

cs.SE

Retrieval-Augmented Code Generation for Universal Information Extraction

Information Extraction (IE) aims to extract structural knowledge (e.g., entities, relations, events) from natural language texts, which brings challenges to existing methods due to task-specific schemas and complex text expressions. Code, as a typical kind of formalized language, is capable of describing structural knowledge under various schemas in a universal way. On the other hand, Large Language Models (LLMs) trained on both codes and texts have demonstrated powerful capabilities of transforming texts into codes, which provides a feasible solution to IE tasks. Therefore, in this paper, we propose a universal retrieval-augmented code generation framework based on LLMs, called Code4UIE, for IE tasks. Specifically, Code4UIE adopts Python classes to define task-specific schemas of various structural knowledge in a universal way. By so doing, extracting knowledge under these schemas can be transformed into generating codes that instantiate the predefined Python classes with the information in texts. To generate these codes more precisely, Code4UIE adopts the in-context learning mechanism to instruct LLMs with examples. In order to obtain appropriate examples for different tasks, Code4UIE explores several example retrieval strategies, which can retrieve examples semantically similar to the given texts. Extensive experiments on five representative IE tasks across nine datasets demonstrate the effectiveness of the Code4UIE framework.

cs.AI

A Variational Auto-Encoder Enabled Multi-Band Channel Prediction Scheme for Indoor Localization

Indoor localization is getting increasing demands for various cutting-edged technologies, like Virtual/Augmented reality and smart home. Traditional model-based localization suffers from significant computational overhead, so fingerprint localization is getting increasing attention, which needs lower computation cost after the fingerprint database is built. However, the accuracy of indoor localization is limited by the complicated indoor environment which brings the multipath signal refraction. In this paper, we provided a scheme to improve the accuracy of indoor fingerprint localization from the frequency domain by predicting the channel state information (CSI) values from another transmitting channel and spliced the multi-band information together to get more precise localization results. We tested our proposed scheme on COST 2100 simulation data and real time orthogonal frequency division multiplexing (OFDM) WiFi data collected from an office scenario.

eess.SP

Enhanced Language Representation with Label Knowledge for Span Extraction

Span extraction, aiming to extract text spans (such as words or phrases) from plain texts, is a fundamental process in Information Extraction. Recent works introduce the label knowledge to enhance the text representation by formalizing the span extraction task into a question answering problem (QA Formalization), which achieves state-of-the-art performance. However, QA Formalization does not fully exploit the label knowledge and suffers from low efficiency in training/inference. To address those problems, we introduce a new paradigm to integrate label knowledge and further propose a novel model to explicitly and efficiently integrate label knowledge into text representations. Specifically, it encodes texts and label annotations independently and then integrates label knowledge into text representation with an elaborate-designed semantics fusion module. We conduct extensive experiments on three typical span extraction tasks: flat NER, nested NER, and event detection. The empirical results show that 1) our method achieves state-of-the-art performance on four benchmarks, and 2) reduces training time and inference time by 76% and 77% on average, respectively, compared with the QA Formalization paradigm. Our code and data are available at https://github.com/Akeepers/LEAR.

cs.CL

Detection of bistable structures via the Conley index and applications to biological systems

Bistability is a ubiquitous phenomenon in life sciences. In this paper, two kinds of bistable structures in dynamical systems are studied: One is two one-point attractors, another is a one-point attractor accompanied by a cycle attractor. By the Conley index theory, we prove that there exist other isolated invariant sets besides the two attractors, and also obtain the possible components and their configuration. Moreover, we find that there is always a separatrix or cycle separatrix, which separates the two attractors. Finally, the biological meanings and implications of these structures are given and discussed.

math.DS

Feedback pinning control of collective behaviors aroused by epidemic spread on complex networks

This paper investigates epidemic control behavioral synchronization for a class of complex networks resulting from spread of epidemic diseases via pinning feedback control strategy. Based on the quenched mean field theory, epidemic control synchronization models with inhibition of contact behavior is constructed, combining with the epidemic transmission system and the complex dynamical network carrying extra controllers. By the properties of convex functions and Gerschgorin theorem, the epidemic threshold of the model is obtained, and the global stability of disease-free equilibrium is analyzed. For individual's infected situation, when epidemic spreads, two types of feedback control strategies depended on the diseases' information are designed: the one only adds controllers to infected individuals, the other adds controllers both to infected and susceptible ones. And by using Lyapunov stability theory, under designed controllers, some criteria that guarantee epidemic control synchronization system achieving behavior synchronization are also derived. Several numerical simulations are performed to show the effectiveness of our theoretical results. As far as we know, this is the first work to address the controlling behavioral synchronization induced by epidemic spreading under the pinning feedback mechanism. It is hopeful that we may have more deeper insight into the essence between disease's spreading and collective behavior controlling in complex dynamical networks.

physics.soc-ph

Two novel immunization strategies for epidemic control in directed scale-free networks with nonlinear infectivity

In this paper, we propose two novel immunization strategies, i.e., combined immunization and duplex immunization, for SIS model in directed scale-free networks, and obtain the epidemic thresholds for them with linear and nonlinear infectivities. With the suggested two new strategies, the epidemic thresholds after immunization are greatly increased. For duplex immunization, we demonstrate that its performance is the best among all usual immunization schemes with respect to degree distribution. And for combined immunization scheme, we show that it is more effective than active immunization. Besides, we give a comprehensive theoretical analysis on applying targeted immunization to directed networks. For targeted immunization strategy, we prove that immunizing nodes with large out-degrees are more effective than immunizing nodes with large in-degrees, and nodes with both large out-degrees and large in-degrees are more worthy to be immunized than nodes with only large out-degrees or large in-degrees. Finally, some numerical analysis are performed to verify and complement our theoretical results. This work is the first to divide the whole population into different types and embed appropriate immunization scheme according to the characteristics of the population, and it will benefit the study of immunization and control of infectious diseases on complex networks.

physics.soc-ph

General neck condition for the limit shape of budding vesicles

The shape equation and linking conditions for a vesicle with two-phase domains are derived. We refine the conjecture on the general neck condition for the limit shape of a budding vesicle proposed by J\"{u}licher and Lipowsky [Phys. Rev. Lett. \textbf{70}, 2964 (1993); Phys. Rev. E \textbf{53}, 2670 (1996)], and then we use the shape equation and linking conditions to prove that this conjecture holds not only for axisymmetric budding vesicles, but also for asymmetric ones. Our study reveals that the mean curvature at any point on the membrane segments adjacent to the neck satisfies the general neck condition for the limit shape of a budding vesicle when the length scale of the membrane segments is much larger than the characteristic size of the neck but still much smaller than the characteristic size of the vesicle.

cond-mat.soft

A Deep Hashing Learning Network

Hashing-based methods seek compact and efficient binary codes that preserve the neighborhood structure in the original data space. For most existing hashing methods, an image is first encoded as a vector of hand-crafted visual feature, followed by a hash projection and quantization step to get the compact binary vector. Most of the hand-crafted features just encode the low-level information of the input, the feature may not preserve the semantic similarities of images pairs. Meanwhile, the hashing function learning process is independent with the feature representation, so the feature may not be optimal for the hashing projection. In this paper, we propose a supervised hashing method based on a well designed deep convolutional neural network, which tries to learn hashing code and compact representations of data simultaneously. The proposed model learn the binary codes by adding a compact sigmoid layer before the loss layer. Experiments on several image data sets show that the proposed model outperforms other state-of-the-art methods.

cs.CV

Coefficient of performance at maximum \chi-criterion for Feynman ratchet as a refrigerator

The \chi-criterion is defined as the product of the energy conversion efficiency and the heat absorbed per-unit-time by the working substance [de Tom\'as et al., Phys. Rev. E, 85 (2012) 010104(R)]. The \chi-criterion for Feynman ratchet as a refrigerator operating between two heat baths is optimized. Asymptotic solutions of the coefficient of performance at maximum \chi-criterion for Feynman ratchet are investigated at both large and small temperature difference. An interpolation formula, which fits the numerical solution very well, is proposed. Besides, the sufficient condition for the universality of the coefficient of performance at maximum \chi is investigated.

cond-mat.stat-mech