Searcharxiv⌕ Search

arXiv subjects

Cécile Favre

Publications and source records attributed to Cécile Favre.

At least 19 recordsLinked to original sources

FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity

Large Language Models (LLMs) are increasingly leveraged for translation tasks but often fall short when translating inclusive language -- such as texts containing the singular 'they' pronoun or otherwise reflecting fair linguistic protocols. Because these challenges span both computational and societal domains, it is imperative to critically evaluate how well LLMs handle inclusive translation with a well-founded framework. This paper presents FairTranslate, a novel, fully human-annotated dataset designed to evaluate non-binary gender biases in machine translation systems from English to French. FairTranslate includes 2418 English-French sentence pairs related to occupations, annotated with rich metadata such as the stereotypical alignment of the occupation, grammatical gender indicator ambiguity, and the ground-truth gender label (male, female, or inclusive). We evaluate four leading LLMs (Gemma2-2B, Mistral-7B, Llama3.1-8B, Llama3.3-70B) on this dataset under different prompting procedures. Our results reveal substantial biases in gender representation across LLMs, highlighting persistent challenges in achieving equitable outcomes in machine translation. These findings underscore the need for focused strategies and interventions aimed at ensuring fair and inclusive language usage in LLM-based translation systems. We make the FairTranslate dataset publicly available on Hugging Face, and disclose the code for all experiments on GitHub.

cs.CL↗

A Reference Model for Collaborative Business Intelligence Virtual Assistants

Collaborative Business Analysis (CBA) is a methodology that involves bringing together different stakeholders, including business users, analysts, and technical specialists, to collaboratively analyze data and gain insights into business operations. The primary objective of CBA is to encourage knowledge sharing and collaboration between the different groups involved in business analysis, as this can lead to a more comprehensive understanding of the data and better decision-making. CBA typically involves a range of activities, including data gathering and analysis, brainstorming, problem-solving, decision-making and knowledge sharing. These activities may take place through various channels, such as in-person meetings, virtual collaboration tools or online forums. This paper deals with virtual collaboration tools as an important part of Business Intelligence (BI) platform. Collaborative Business Intelligence (CBI) tools are becoming more user-friendly, accessible, and flexible, allowing users to customize their experience and adapt to their specific needs. The goal of a virtual assistant is to make data exploration more accessible to a wider range of users and to reduce the time and effort required for data analysis. It describes the unified business intelligence semantic model, coupled with a data warehouse and collaborative unit to employ data mining technology. Moreover, we propose a virtual assistant for CBI and a reference model of virtual tools for CBI, which consists of three components: conversational, data exploration and recommendation agents. We believe that the allocation of these three functional tasks allows you to structure the CBI issue and apply relevant and productive models for human-like dialogue, text-to-command transferring, and recommendations simultaneously. The complex approach based on these three points gives the basis for virtual tool for collaboration. CBI encourages people, processes, and technology to enable everyone sharing and leveraging collective expertise, knowledge and data to gain valuable insights for making better decisions. This allows to respond more quickly and effectively to changes in the market or internal operations and improve the progress.

cs.HC↗

Rumor Classification through a Multimodal Fusion Framework and Ensemble Learning

The proliferation of rumors on social media has become a major concern due to its ability to create a devastating impact. Manually assessing the veracity of social media messages is a very time-consuming task that can be much helped by machine learning. Most message veracity verification methods only exploit textual contents and metadata. Very few take both textual and visual contents, and more particularly images, into account. Moreover, prior works have used many classical machine learning models to detect rumors. However, although recent studies have proven the effectiveness of ensemble machine learning approaches, such models have seldom been applied. Thus, in this paper, we propose a set of advanced image features that are inspired from the field of image quality assessment, and introduce the Multimodal fusiON framework to assess message veracIty in social neTwORks (MONITOR), which exploits all message features by exploring various machine learning models. Moreover, we demonstrate the effectiveness of ensemble learning algorithms for rumor detection by using five metalearning models. Eventually, we conduct extensive experiments on two real-world datasets. Results show that MONITOR outperforms state-of-the-art machine learning baselines and that ensemble models significantly increase MONITOR's performance.

cs.CV↗

The Collaborative Business Intelligence Ontology (CBIOnt)

In the current era, many disciplines are seen devoted towards ontology development for their domains with the intention of creating, disseminating and managing resource descriptions of their domain knowledge into machine understandable and processable manner. Ontology construction is a difficult group activity that involves many people with the different expertise. Generally, domain experts are not familiar with the ontology implementation environments and implementation experts do not have all the domain knowledge. We have designed Collaborative Business Intelligence Ontology (CBIOnt) for BI4People project. In this paper, we present CBIOnt that is OWL 2 DL ontology for the description of collaborative session between different collaborators working together on the business intelligent platform. As the collaborative session between various collaborators belongs to some collaborative form, phase and research aspect, therefore CBIOnt captures this knowledge along with the collaborative session content (comments, questions, answers, etc.) so that one can inference various types of information stored on ontologies when required. In addition, it stores the location and temporal-spatial information about the collaboration held between collaborators. We believe CBIOnt serves as a formal framework for dealing with the collaborative session taken place among collaborators on the semantic Web.

cs.DB↗

CH$_3$CN deuteration in the SVS13-A Class I hot-corino. SOLIS XV

We studied the line emission from CH3CN and its deuterated isotopologue CH$_2$DCN towards the prototypical Class I object SVS13-A, where the deuteration of a large number of species has already been reported. Our goal is to measure the CH$_3$CN deuteration in a Class I protostar, for the first time, in order to constrain the CH$_3$CN formation pathways and the chemical evolution from the early prestellar core and Class 0 to the evolved Class I stages. We imaged CH2DCN towards SVS13-A using the IRAM NOEMA interferometer at 3mm in the context of the Large Program SOLIS (with a spatial resolution of 1.8"x1.2"). The NOEMA images have been complemented by the CH$_3$CN and CH$_2$DCN spectra collected by the IRAM-30m Large Program ASAI, that provided an unbiased spectral survey at 3mm, 2mm, and 1.3mm. The observed line emission has been analysed using LTE and non-LTE LVG approaches. The NOEMA/SOLIS images of CH2DCN show that this species emits in an unresolved area centered towards the SVS13-A continuum emission peak, suggesting that methyl cyanide and its isotopologues are associated with the hot corino of SVS13-A, previously imaged via other iCOMs. In addition, we detected 41 and 11 ASAI transitions of CH$_3$CN and CH2DCN, respectively, which cover upper level energies (Eup) from 13 to 442 K and from 18 K to 200 K, respectively. The derived [CH2DCN]/[CH3CN] ratio is $\sim$9\%. This value is consistent with those measured towards prestellar cores and a factor 2-3 higher than those measured in Class 0 protostars. Contrarily to what expected for other molecular species, the CH3CN deuteration does not show a decrease in SVS13-A with respect to measurements in younger prestellar cores and Class 0 protostars. Finally, we discuss why our new results suggest that CH3CN was likely synthesised via gas-phase reactions and frozen onto the dust grain mantles during the cold prestellar phase.

astro-ph.SR↗

FAUST III. Misaligned rotations of the envelope, outflow, and disks in the multiple protostellar system of VLA 1623$-$2417

We report a study of the low-mass Class-0 multiple system VLA 1623AB in the Ophiuchus star-forming region, using H$^{13}$CO$^+$ ($J=3-2$), CS ($J=5-4$), and CCH ($N=3-2$) lines as part of the ALMA Large Program FAUST. The analysis of the velocity fields revealed the rotation motion in the envelope and the velocity gradients in the outflows (about 2000 au down to 50 au). We further investigated the rotation of the circum-binary VLA 1623A disk as well as the VLA 1623B disk. We found that the minor axis of the circum-binary disk of VLA 1623A is misaligned by about 12 degrees with respect to the large-scale outflow and the rotation axis of the envelope. In contrast, the minor axis of the circum-binary disk is parallel to the large-scale magnetic field according to previous dust polarization observations, suggesting that the misalignment may be caused by the different directions of the envelope rotation and the magnetic field. If the velocity gradient of the outflow is caused by rotation, the outflow has a constant angular momentum and the launching radius is estimated to be $5-16$ au, although it cannot be ruled out that the velocity gradient is driven by entrainments of the two high-velocity outflows. Furthermore, we detected for the first time a velocity gradient associated with rotation toward the VLA 16293B disk. The velocity gradient is opposite to the one from the large-scale envelope, outflow, and circum-binary disk. The origin of its opposite gradient is also discussed.

astro-ph.GA↗

Calling to CNN-LSTM for Rumor Detection: A Deep Multi-channel Model for Message Veracity Classification in Microblogs

Reputed by their low-cost, easy-access, real-time and valuable information, social media also wildly spread unverified or fake news. Rumors can notably cause severe damage on individuals and the society. Therefore, rumor detection on social media has recently attracted tremendous attention. Most rumor detection approaches focus on rumor feature analysis and social features, i.e., metadata in social media. Unfortunately, these features are data-specific and may not always be available, e.g., when the rumor has just popped up and not yet propagated. In contrast, post contents (including images or videos) play an important role and can indicate the diffusion purpose of a rumor. Furthermore, rumor classification is also closely related to opinion mining and sentiment analysis. Yet, to the best of our knowledge, exploiting images and sentiments is little investigated.Considering the available multimodal features from microblogs, notably, we propose in this paper an end-to-end model called deepMONITOR that is based on deep neural networks and allows quite accurate automated rumor verification, by utilizing all three characteristics: post textual and image contents, as well as sentiment. deepMONITOR concatenates image features with the joint text and sentiment features to produce a reliable, fused classification. We conduct extensive experiments on two large-scale, real-world datasets. The results show that deepMONITOR achieves a higher accuracy than state-of-the-art methods.

cs.CL↗

MONITOR: A Multimodal Fusion Framework to Assess Message Veracity in Social Networks

Users of social networks tend to post and share content with little restraint. Hence, rumors and fake news can quickly spread on a huge scale. This may pose a threat to the credibility of social media and can cause serious consequences in real life. Therefore, the task of rumor detection and verification has become extremely important. Assessing the veracity of a social media message (e.g., by fact checkers) involves analyzing the text of the message, its context and any multimedia attachment. This is a very time-consuming task that can be much helped by machine learning. In the literature, most message veracity verification methods only exploit textual contents and metadata. Very few take both textual and visual contents, and more particularly images, into account. In this paper, we second the hypothesis that exploiting all of the components of a social media post enhances the accuracy of veracity detection. To further the state of the art, we first propose using a set of advanced image features that are inspired from the field of image quality assessment, which effectively contributes to rumor detection. These metrics are good indicators for the detection of fake images, even for those generated by advanced techniques like generative adversarial networks (GANs). Then, we introduce the Multimodal fusiON framework to assess message veracIty in social neTwORks (MONITOR), which exploits all message features (i.e., text, social context, and image features) by supervised machine learning. Such algorithms provide interpretability and explainability in the decisions taken, which we believe is particularly important in the context of rumor verification. Experimental results show that MONITOR can detect rumors with an accuracy of 96% and 89% on the MediaEval benchmark and the FakeNewsNet dataset, respectively. These results are significantly better than those of state-of-the-art machine learning baselines.

cs.SI↗

Coining goldMEDAL: A New Contribution to Data Lake Generic Metadata Modeling

The rise of big data has revolutionized data exploitation practices and led to the emergence of new concepts. Among them, data lakes have emerged as large heterogeneous data repositories that can be analyzed by various methods. An efficient data lake requires a metadata system that addresses the many problems arising when dealing with big data. In consequence, the study of data lake metadata models is currently an active research topic and many proposals have been made in this regard. However, existing metadata models are either tailored for a specific use case or insufficiently generic to manage different types of data lakes, including our previous model MEDAL. In this paper, we generalize MEDAL's concepts in a new metadata model called goldMEDAL. Moreover, we compare goldMEDAL with the most recent state-of-the-art metadata models aiming at genericity and show that we can reproduce these metadata models with goldMEDAL's concepts. As a proof of concept, we also illustrate that goldMEDAL allows the design of various data lakes by presenting three different use cases.

cs.DB↗

FAUST II. Discovery of a Secondary Outflow in IRAS 15398-3359: Variability in Outflow Direction during the Earliest Stage of Star Formation?

We have observed the very low-mass Class 0 protostar IRAS 15398-3359 at scales ranging from 50 au to 1800 au, as part of the ALMA Large Program FAUST. We uncover a linear feature, visible in H2CO, SO, and C18O line emission, which extends from the source along a direction almost perpendicular to the known active outflow. Molecular line emission from H2CO, SO, SiO, and CH3OH further reveals an arc-like structure connected to the outer end of the linear feature and separated from the protostar, IRAS 15398-3359, by 1200 au. The arc-like structure is blue-shifted with respect to the systemic velocity. A velocity gradient of 1.2 km/s over 1200 au along the linear feature seen in the H2CO emission connects the protostar and the arc-like structure kinematically. SO, SiO, and CH3OH are known to trace shocks, and we interpret the arc-like structure as a relic shock region produced by an outflow previously launched by IRAS 15398-3359. The velocity gradient along the linear structure can be explained as relic outflow motion. The origins of the newly observed arc-like structure and extended linear feature are discussed in relation to turbulent motions within the protostellar core and episodic accretion events during the earliest stage of protostellar evolution.

astro-ph.SR↗

Data Lakes for Digital Humanities

Traditional data in Digital Humanities projects bear various formats (structured, semi-structured, textual) and need substantial transformations (encoding and tagging, stemming, lemmatization, etc.) to be managed and analyzed. To fully master this process, we propose the use of data lakes as a solution to data siloing and big data variety problems. We describe data lake projects we currently run in close collaboration with researchers in humanities and social sciences and discuss the lessons learned running these projects.

cs.DB↗

Including Images into Message Veracity Assessment in Social Media

The extensive use of social media in the diffusion of information has also laid a fertile ground for the spread of rumors, which could significantly affect the credibility of social media. An ever-increasing number of users post news including, in addition to text, multimedia data such as images and videos. Yet, such multimedia content is easily editable due to the broad availability of simple and effective image and video processing tools. The problem of assessing the veracity of social network posts has attracted a lot of attention from researchers in recent years. However, almost all previous works have focused on analyzing textual contents to determine veracity, while visual contents, and more particularly images, remains ignored or little exploited in the literature. In this position paper, we propose a framework that explores two novel ways to assess the veracity of messages published on social networks by analyzing the credibility of both their textual and visual contents.

cs.IR↗

A 3mm chemical exploration of small organics in Class I YSOs

There is mounting evidence that the composition and structure of planetary systems are intimately linked to their birth environments. During the past decade, several spectral surveys probed the chemistry of the earliest stages of star formation and of late planet-forming disks. However, very little is known about the chemistry of intermediate protostellar stages, i.e. Class I Young Stellar Objects (YSOs), where planet formation may have already begun. We present here the first results of a 3mm spectral survey performed with the IRAM-30m telescope to investigate the chemistry of a sample of seven Class I YSOs located in the Taurus star-forming region. These sources were selected to embrace the wide diversity identified for low-mass protostellar envelope and disk systems. We present detections and upper limits of thirteen small ($N_{\rm atoms}\leq3$) C, N, O, and S carriers - namely CO, HCO$^+$, HCN, HNC, CN, N$_2$H$^+$, C$_2$H, CS, SO, HCS$^+$, C$_2$S, SO$_2$, OCS - and some of their D, $^{13}$C, $^{15}$N, $^{18}$O, $^{17}$O, and $^{34}$S isotopologues. Together, these species provide constraints on gas-phase C/N/O ratios, D- and $^{15}$N-fractionation, source temperature and UV exposure, as well as the overall S-chemistry. We find substantial evidence of chemical differentiation among our source sample, some of which can be traced back to Class I physical parameters, such as the disk-to-envelope mass ratio (proxy for Class I evolutionary stage), the source luminosity, and the UV-field strength. Overall, these first results allow us to start investigating the astrochemistry of Class I objects, however, interferometric observations are needed to differentiate envelope versus disk chemistry.

astro-ph.SR↗

Metadata Systems for Data Lakes: Models and Features

Over the past decade, the data lake concept has emerged as an alternative to data warehouses for storing and analyzing big data. A data lake allows storing data without any predefined schema. Therefore, data querying and analysis depend on a metadata system that must be efficient and comprehensive. However, metadata management in data lakes remains a current issue and the criteria for evaluating its effectiveness are more or less nonexistent.In this paper, we introduce MEDAL, a generic, graph-based model for metadata management in data lakes. We also propose evaluation criteria for data lake metadata systems through a list of expected features. Eventually, we show that our approach is more comprehensive than existing metadata systems.

cs.DB↗

Impact of nonconvergence and various approximations of the partition function on the molecular column densities in the interstellar medium

We emphasize that the completeness of the partition function, that is, the use of a converged partition function at the typical temperature range of the survey, is very important to decrease the uncertainty on this quantity and thus to derive reliable interstellar molecular densities. In that context, we show how the use of different approximations for the rovibrational partition function together with some interpolation and/or extrapolation procedures may affect the estimate of the interstellar molecular column density. For that purpose, we apply the partition function calculations to astronomical observations performed with the IRAM-30m telescope towards the NGC7538-IRS1 source of two N-bearing molecules: isocyanic acid (HNCO, a quasilinear molecule) and methyl cyanide (CH$_3$CN, a symmetric top molecule). The case of methyl formate (HCOOCH$_3$), which is an asymmetric top O-bearing molecule containing an internal rotor is also discussed. Our analysis shows that the use of different partition function approximations leads to relative differences in the resulting column densities in the range 9 to 43\%. Thus, we expect this work to be relevant for surveys of sources with temperatures higher than 300~K and to observations in the infrared.

astro-ph.GA↗

Deuterated methanol toward NGC 7538-IRS1

We investigate the deuteration of methanol towards the high-mass star forming region NGC 7538-IRS1. We have carried out a multi-transition study of CH$_3$OH, $^{13}$CH$_3$OH and of the deuterated fllavors, CH$_2$DOH and CH$_3$OD, between 1.0--1.4 mm with the IRAM-30~m antenna. In total, 34 $^{13}$CH$_3$OH, 13 CH$_2$DOH lines and 20 CH$_3$OD lines spanning a wide range of upper-state energies (E$_{up}$) were detected. From the detected transitions, we estimate that the measured D/H does not exceed 1$\%$, with a measured CH$_2$DOH/CH$_3$OH and CH$_3$OD/CH$_3$OH of about (32$\pm$8)$\times$10$^{-4}$ and (10$\pm$4)$\times$10$^{-4}$, respectively. This finding is consistent with the hypothesis of a short-time scale formation during the pre-stellar phase. We find a relative abundance ratio CH$_2$DOH/CH$_3$OD of 3.2 $\pm$ 1.5. This result is consistent with a statistical deuteration. We cannot exclude H/D exchanges between water and methanol if water deuteration is of the order 0.1$\%$, as suggested by recent Herschel observations.

astro-ph.GA↗

Gas density perturbations induced by forming planet(s) in the AS 209 protoplanetary disk as seen with ALMA

The formation of planets occurs within protoplanetary disks surrounding young stars, resulting in perturbation of the gas and dust surface densities. Here, we report the first evidence of spatially resolved gas surface density ($Σ_{g}$) perturbation towards the AS~209 protoplanetary disk from the optically thin C$^{18}$O ($J=2-1$) emission. The observations were carried out at 1.3~mm with ALMA at a spatial resolution of about 0.3$\arcsec$ $\times$ 0.2$\arcsec$ (corresponding to $\sim$ 38 $\times$ 25 au). The C$^{18}$O emission shows a compact ($\le$60~au), centrally peaked emission and an outer ring peaking at 140~au, consistent with that observed in the continuum emission and, its azimuthally averaged radial intensity profile presents a deficit that is spatially coincident with the previously reported dust map. This deficit can only be reproduced with our physico-thermochemical disk model by lowering $Σ_{gas}$ by nearly an order of magnitude in the dust gaps. Another salient result is that contrary to C$^{18}$O, the DCO$^{+}$ ($J=3-2$) emission peaks between the two dust gaps. We infer that the best scenario to explain our observations (C$^{18}$O deficit and DCO$^{+}$ enhancement) is a gas perturbation due to forming-planet(s), that is commensurate with previous continuum observations of the source along with hydrodynamical simulations. Our findings confirm that the previously observed dust gaps are very likely due to perturbation of the gas surface density that is induced by a planet of at least 0.2~M$\rm_{Jupiter}$ in formation. Finally, our observations also show the potential of using CO isotopologues to probe the presence of saturn mass planet(s).

astro-ph.EP↗