SearcharxivSearch

arXiv subjects

David Allen

Publications and source records attributed to David Allen.

11 recordsLinked to original sources

Non-detectable patterns hidden within sequences of bits

In this paper we construct families of bit sequences using combinatorial methods. Each sequence is derived by con- verting a collection of numbers encoding certain combinatorial nu- merics from objects exhibiting symmetry in various dimensions. Using the algorithms first described in [1] we show that the NIST testing suite described in publication 800-22 does not detect these symmetries hidden within these sequences.

math.CO

Volatility and irregularity Capturing in stock price indices using time series Generative adversarial networks (TimeGAN)

This paper captures irregularities in financial time series data, particularly stock prices, in the presence of COVID-19 shock. We conjectured that jumps and irregularities are embedded in stock data due to the pandemic shock, which brings forth irregular trends in the time series data. We put forward that efficient and robust forecasting methods are needed to predict stock closing prices in the presence of the pandemic shock. This piece of information is helpful to investors as far as confidence risk and return boost are concerned. Generative adversarial networks of a time series nature are used to provide new ways of modeling and learning the proper and suitable distribution for the financial time series data under complex setups. Ideally, these traditional models are liable to producing high forecasting errors, and they need to be more robust to capture dependency structures and other stylized facts like volatility in stock markets. The TimeGAN model is used, effectively dealing with this risk of poor forecasts. Using the DAX stock index from January 2010 to November 2022, we trained the LSTM, GRU, WGAN, and TimeGAN models as benchmarks and forecasting errors were noted, and our TimeGAN outperformed them all as indicated by a small forecasting error.

cs.CE

Leveraging Large Language Models in Conversational Recommender Systems

A Conversational Recommender System (CRS) offers increased transparency and control to users by enabling them to engage with the system through a real-time multi-turn dialogue. Recently, Large Language Models (LLMs) have exhibited an unprecedented ability to converse naturally and incorporate world knowledge and common-sense reasoning into language understanding, unlocking the potential of this paradigm. However, effectively leveraging LLMs within a CRS introduces new technical challenges, including properly understanding and controlling a complex conversation and retrieving from external sources of information. These issues are exacerbated by a large, evolving item corpus and a lack of conversational data for training. In this paper, we provide a roadmap for building an end-to-end large-scale CRS using LLMs. In particular, we propose new implementations for user preference understanding, flexible dialogue management and explainable recommendations as part of an integrated architecture powered by LLMs. For improved personalization, we describe how an LLM can consume interpretable natural language user profiles and use them to modulate session-level context. To overcome conversational data limitations in the absence of an existing production CRS, we propose techniques for building a controllable LLM-based user simulator to generate synthetic conversations. As a proof of concept we introduce RecLLM, a large-scale CRS for YouTube videos built on LaMDA, and demonstrate its fluency and diverse functionality through some illustrative example conversations.

cs.IR

Efficient Modelling & Forecasting with range based volatility models and application

This paper considers an alternative method for fitting CARR models using combined estimating functions (CEF) by showing its usefulness in applications in economics and quantitative finance. The associated information matrix for corresponding new estimates is derived to calculate the standard errors. A simulation study is carried out to demonstrate its superiority relative to other two competitors: linear estimating functions (LEF) and the maximum likelihood (ML). Results show that CEF estimates are more efficient than LEF and ML estimates when the error distribution is mis-specified. Taking a real data set from financial economics, we illustrate the usefulness and applicability of the CEF method in practice and report reliable forecast values to minimize the risk in the decision making process.

stat.AP

Geotagging One Hundred Million Twitter Accounts with Total Variation Minimization

Geographically annotated social media is extremely valuable for modern information retrieval. However, when researchers can only access publicly-visible data, one quickly finds that social media users rarely publish location information. In this work, we provide a method which can geolocate the overwhelming majority of active Twitter users, independent of their location sharing preferences, using only publicly-visible Twitter data. Our method infers an unknown user's location by examining their friend's locations. We frame the geotagging problem as an optimization over a social network with a total variation-based objective and provide a scalable and distributed algorithm for its solution. Furthermore, we show how a robust estimate of the geographic dispersion of each user's ego network can be used as a per-user accuracy measure which is effective at removing outlying errors. Leave-many-out evaluation shows that our method is able to infer location for 101,846,236 Twitter users at a median error of 6.38 km, allowing us to geotag over 80\% of public tweets.

cs.SI

New Advances in Inference by Recursive Conditioning

Recursive Conditioning (RC) was introduced recently as the first any-space algorithm for inference in Bayesian networks which can trade time for space by varying the size of its cache at the increment needed to store a floating point number. Under full caching, RC has an asymptotic time and space complexity which is comparable to mainstream algorithms based on variable elimination and clustering (exponential in the network treewidth and linear in its size). We show two main results about RC in this paper. First, we show that its actual space requirements under full caching are much more modest than those needed by mainstream methods and study the implications of this finding. Second, we show that RC can effectively deal with determinism in Bayesian networks by employing standard logical techniques, such as unit resolution, allowing a significant reduction in its time requirements in certain cases. We illustrate our results using a number of benchmark networks, including the very challenging ones that arise in genetic linkage analysis.

cs.AI

Exploiting Evidence in Probabilistic Inference

We define the notion of compiling a Bayesian network with evidence and provide a specific approach for evidence-based compilation, which makes use of logical processing. The approach is practical and advantageous in a number of application areas-including maximum likelihood estimation, sensitivity analysis, and MAP computations-and we provide specific empirical results in the domain of genetic linkage analysis. We also show that the approach is applicable for networks that do not contain determinism, and show that it empirically subsumes the performance of the quickscore algorithm when applied to noisy-or networks.

cs.AI

The Role of Schema Matching in Large Enterprises

To date, the principal use case for schema matching research has been as a precursor for code generation, i.e., constructing mappings between schema elements with the end goal of data transfer. In this paper, we argue that schema matching plays valuable roles independent of mapping construction, especially as schemata grow to industrial scales. Specifically, in large enterprises human decision makers and planners are often the immediate consumer of information derived from schema matchers, instead of schema mapping tools. We list a set of real application areas illustrating this role for schema matching, and then present our experiences tackling a customer problem in one of these areas. We describe the matcher used, where the tool was effective, where it fell short, and our lessons learned about how well current schema matching technology is suited for use in large enterprises. Finally, we suggest a new agenda for schema matching research based on these experiences.

cs.DB

A Counterexample to a conjecture of Bosio and Meersseman

In a paper of Bosio and Meersseman (Real quadrics in Cn, complex manifolds and convex polytopes) the following is conjectured: If P is dual neighborly, then Zp is diffeomorphic to the connected sum of products of spheres. In this paper a counterexample is provided.

math.SG

Policy for access: Framing the question

Five years after the '96 Telecommunications Act, we still find precious little local facilities-based competition. In response there are calls in Congress and even from the FCC for new legislation to "free the Bells." However, the same ideology drove policy, not just five years ago, but also almost twenty years back with the first modern push for "freedom," namely divestiture. How might we frame the question of policy for local access to engender a more fruitful approach? The starting point for this analysis is the network--not bits and bytes, but the human network. With the human network as starting point, the unit of analysis is the community--specifically, the individual in a tension with community. There are two core ideas. The first takes a behavioral approach to the economics--and the relative share between beneficial chaos and order, in economic affairs, becomes explicit. If the first main idea provides a conceptual base for open source, the second core idea distinguishes open source from open design, ie at the information 'frontier' we push forward. The resulting policy frame for access is worked out in the detailed, concrete steps of an extended thought experiment. A small town setting (Concord, Massachusetts) grounds the discussion in the real world. The purpose overall is to stimulate new thinking which may break out of the conundrum where periodic rounds to legislate 'freedom' produce the opposite, recursively. The ultimate aim is better fit between our analytically-driven expectations and economic outcomes.

cs.CY