SearcharxivSearch

arXiv subjects

Benjamin Nguyen

Publications and source records attributed to Benjamin Nguyen.

16 recordsLinked to original sources

Refining Word-Based Grammatical Error Annotation for L2 Korean

Korean grammatical error correction (K-GEC) presents a structural mismatch between word-based evaluation and the morpheme-level locus of many learner errors. Postpositions and verbal endings are bound to lexical hosts, but they encode grammatical relations that must be represented in correction and evaluation. This paper refines word-based grammatical error annotation for L2 Korean by addressing three connected problems in existing resources: surface target realization, Korean-specific edit annotation, and single-reference evaluation. We reconstruct target sentences from the National Institute of Korean Language (NIKL) L2 corpus under morphologically constrained realization rules and convert its morpheme-level annotations into word-level \texttt{m2} edits. We then define a Korean ERRANT-style annotation scheme that preserves the MRU core while distinguishing functional morpheme errors, spelling errors, word boundary errors, and word order errors. We also augment the KoLLA corpus with an additional reference correction, yielding a multi-reference evaluation setting for Korean GEC. Empirical validation shows that the refined NIKL targets yield lower perplexity, the converted \texttt{m2} files achieve higher agreement with source-target edit representations, and the refined resources improve KoBART-based correction under the same model setting. Multi-reference KoLLA evaluation further reduces the penalty imposed on valid corrections that diverge from a single reference, especially for neural and prompted GEC systems. These results show that Korean GEC evaluation depends not only on correction models, but also on reference data and edit annotations that reflect Korean morphology, spacing, and correction variability.

cs.CL

Onto-DP: Constructing Neighborhoods for Differential Privacy on Ontological Databases

In this paper, we investigate how attackers can discover sensitive information embedded within databases by exploiting inference rules. We demonstrate the inadequacy of naively applied existing state of the art differential privacy (DP) models in safeguarding against such attacks. We introduce ontology aware differential privacy (Onto-DP), a novel extension of differential privacy paradigms built on top of any classical DP model by enriching it with semantic awareness. We show that this extension is a sufficient condition to adequately protect against attackers aware of inference rules.

cs.CR

Distributed Transition System with Tags and Value-wise Metric, for Privacy Analysis

We introduce a logical framework named Distributed Labeled Tagged Transition System (DLTTS), using concepts from Probabilistic Automata, Probabilistic Concurrent Systems, and Probabilistic labelled transition systems. We show that DLTTS can be used to formally model how a given piece of private information P (e.g., a set of tuples) stored in a given database D can get captured progressively by an adversary A repeatedly querying D, enhancing the knowledge acquired from the answers to these queries with relational deductions using certain additional non-private data. The database D is assumed protected with generalization mechanisms. We also show that, on a large class of databases, metrics can be defined 'value-wise', and more general notions of adjacency between data bases can be defined, based on these metrics. These notions can also play a role in differentially private protection mechanisms.

cs.LO

Distributed Transition Systems with Tags for Privacy Analysis

We present a logical framework that formally models how a given private information P stored on a given database D, can get captured progressively, by an agent/adversary querying the database repeatedly. Named DLTTS (Distributed Labeled Tagged Transition System), the framework borrows ideas from several domains: Probabilistic Automata of Segala, Probabilistic Concurrent Systems, and Probabilistic labelled transition systems. To every node on a DLTTS is attached a tag that represents the 'current' knowledge of the adversary, acquired from the responses of the answering mechanism of the DBMS to his/her queries, at the nodes traversed earlier, along any given run; this knowledge is completed at the same node, with further relational deductions, possibly in combination with 'public' information from other databases given in advance. A 'blackbox' mechanism is also part of a DLTTS, and it is meant as an oracle; its role is to tell if the private information has been deduced by the adversary at the current node, and if so terminate the run. An additional special feature is that the blackbox also gives information on how 'close', or how 'far', the knowledge of the adversary is, from the private information P , at the current node. A metric is defined for that purpose, on the set of all 'type compatible' tuples from the given database, the data themselves being typed with the headers of the base. Despite the transition systems flavor of our framework, this metric is not 'behavioral' in the sense presented in some other works. It is exclusively database oriented, and allows to define new notions of adjacency and of indistinguishabilty between databases, more generally than those usually based on the Hamming metric (and a restricted notion of adjacency). Examples are given all along to illustrate how our framework works. Keywords:Database, Privacy, Transition System, Probability, Distribution.

cs.CL

A targeted search for repeating fast radio bursts associated with gamma-ray bursts

The origin of fast radio bursts (FRBs) still remains a mystery, even with the increased number of discoveries in the last three years. Growing evidence suggests that some FRBs may originate from magnetars. Large, single-dish telescopes such as Arecibo Observatory (AO) and Green Bank Telescope (GBT) have the sensitivity to detect FRB~121102-like bursts at gigaparsec distances. Here we present searches using AO and GBT that aimed to find potential radio bursts at 11 sites of past $\gamma$--ray bursts that show evidence for the birth of a magnetar. We also performed a search towards GW170817, which has a merger remnant whose nature remains uncertain. We place $10\,\sigma$ fluence upper limits of $\approx 0.036$ Jy ms at 1.4 GHz and $\approx 0.063$ Jy ms at 4.5 GHz for AO data and fluence upper limits of $\approx 0.085$ Jy ms at 1.4 GHz and $\approx 0.098$ Jy ms at 1.9 GHz for GBT data, for a maximum pulse width of $\approx 42$ ms. The AO observations had sufficient sensitivity to detect any FRB of similar luminosity to the one recently detected from the Galactic magnetar SGR 1935+2154. Assuming a Schechter function for the luminosity function of FRBs, we find that our non-detections favor a steep power--law index ($\alpha\lesssim-1.0$) and a large cut--off luminosity ($L_0 \gtrsim 10^{42}$ erg/s).

astro-ph.HE

Techniques d'anonymisation tabulaire : concepts et mise en oeuvre

In this document, we present a state of the art of anonymization techniques for classical tabular datasets. This article is geared towards a general public having some knowledge of mathematics and computer science, but with no need for specific knowledge in anonymization. The objective of this document it to explain anonymization concepts in order to be able to sanitize a dataset and compute reindentification risk. The document contains a large number of examples to help understand the calculations. ----- Dans ce document, nous pr\'esentons l'\'etat de l'art des techniques d'anonymisation pour des bases de donn\'ees classiques (i.e. des tables), \`a destination d'un public technique ayant une formation universitaire de base en math\'ematiques et informatique, mais non sp\'ecialiste. L'objectif de ce document est d'expliquer les concepts permettant de r\'ealiser une anonymisation de donn\'ees tabulaires, et de calculer les risques de r\'eidentification. Le document est largement compos\'e d'exemples permettant au lecteur de comprendre comment mettre en oeuvre les calculs.

cs.CR

Key Exchange Protocol in the Trusted Data Servers Context

The aim of this technical report is to complement the work in [To et al. 2014] by proposing a Group Key Exchange protocol so that the Querier and TDSs (and TDSs themselves) can securely create and exchange the shared key. Then, the security of this protocol is formally proved using the game-based model. Finally, we perform the comparison between this protocol and other related works.

cs.CR

XML content warehousing: Improving sociological studies of mailing lists and web data

In this paper, we present the guidelines for an XML-based approach for the sociological study of Web data such as the analysis of mailing lists or databases available online. The use of an XML warehouse is a flexible solution for storing and processing this kind of data. We propose an implemented solution and show possible applications with our case study of profiles of experts involved in W3C standard-setting activity. We illustrate the sociological use of semi-structured databases by presenting our XML Schema for mailing-list warehousing. An XML Schema allows many adjunctions or crossings of data sources, without modifying existing data sets, while allowing possible structural evolution. We also show that the existence of hidden data implies increased complexity for traditional SQL users. XML content warehousing allows altogether exhaustive warehousing and recursive queries through contents, with far less dependence on the initial storage. We finally present the possibility of exporting the data stored in the warehouse to commonly-used advanced software devoted to sociological analysis.

cs.DB

XQ2P: Efficient XQuery P2P Time Series Processing

In this demonstration, we propose a model for the management of XML time series (TS), using the new XQuery 1.1 window operator. We argue that centralized computation is slow, and demonstrate XQ2P, our prototype of efficient XQuery P2P TS computation in the context of financial analysis of large data sets (>1M values).

cs.DB

Gestion efficace de séries temporelles en P2P: Application à l'analyse technique et l'étude des objets mobiles

In this paper, we propose a simple generic model to manage time series. A time series is composed of a calendar with a typed value for each calendar entry. Although the model could support any kind of XML typed values, in this paper we focus on real numbers, which are the usual application. We define basic vector space operations (plus, minus, scale), and also relational-like and application oriented operators to manage time series. We show the interest of this generic model on two applications: (i) a stock investment helper; (ii) an ecological transport management system. Stock investment requires window-based operations while trip management requires complex queries. The model has been implemented and tested in PHP, Java, and XQuery. We show benchmark results illustrating that the computing of 5000 series of over 100.000 entries in length - common requirements for both applications - is difficult on classical centralized PCs. In order to serve a community of users sharing time series, we propose a P2P implementation of time series by dividing them in segments and providing optimized algorithms for operator expression computation.

cs.DB

The WebStand Project

In this paper we present the state of advancement of the French ANR WebStand project. The objective of this project is to construct a customizable XML based warehouse platform to acquire, transform, analyze, store, query and export data from the web, in particular mailing lists, with the final intension of using this data to perform sociological studies focused on social groups of World Wide Web, with a specific emphasis on the temporal aspects of this data. We are currently using this system to analyze the standardization process of the W3C, through its social network of standard setters.

cs.DB

The WebContent XML Store

In this article, we describe the XML storage system used in the WebContent project. We begin by advocating the use of an XML database in order to store WebContent documents, and we present two different ways of storing and querying these documents : the use of a centralized XML database and the use of a P2P XML database.

cs.DB

Janus: Automatic Ontology Builder from XSD Files

The construction of a reference ontology for a large domain still remains an hard human task. The process is sometimes assisted by software tools that facilitate the information extraction from a textual corpus. Despite of the great use of XML Schema files on the internet and especially in the B2B domain, tools that offer a complete semantic analysis of XML schemas are really rare. In this paper we introduce Janus, a tool for automatically building a reference knowledge base starting from XML Schema files. Janus also provides different useful views to simplify B2B application integration.

cs.DB

Deriving Ontologies from XML Schema

In this paper, we present a method and a tool for deriving a skeleton of an ontology from XML schema files. We first recall what an is ontology and its relationships with XML schemas. Next, we focus on ontology building methodology and associated tool requirements. Then, we introduce Janus, a tool for building an ontology from various XML schemas in a given domain. We summarize the main features of Janus and illustrate its functionalities through a simple example. Finally, we compare our approach to other existing ontology building tools.

cs.PL

Applying an XML Warehouse to Social Network Analysis, Lessons from the WebStand Project

In this paper we present the state of advancement of the French ANR WebStand project. The objective of this project is to construct a customizable XML based warehouse platform to acquire, transform, analyze, store, query and export data from the web, in particular mailing lists, with the final intension of using this data to perform sociological studies focused on social groups of World Wide Web, with a specific emphasis on the temporal aspects of this data. We are currently using this system to analyze the standardization process of the W3C, through its social network of standard setters.

cs.DB

Gouverner la standardisation par les changements d'arene. Le cas du XML

In this paper, we discuss the available approches of the new governance structures of standardization, in order to propose new hypothesis on the way computer sciences languages are dealt with. We consider the example of the XML language and its applications in order to propose a dynamic analysis of this governance, focusing on the coordination that is done by companies, and the strategic usage they have of these arenas to further their goals. We advocate the development of more of such empirical analysis in order to cover all the perspectives of possible international policies in this area.

cs.CY