SearcharxivSearch

arXiv subjects

Mohamed Chaouch

Publications and source records attributed to Mohamed Chaouch.

17 recordsLinked to original sources

Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their "black-box" nature hinders deployment in critical domains. Grounded Visual Reasoning (GVR) approaches attempt to improve interpretability by explicitly couple textual rationales with visual grounding information, which are typically textual coordinates. This mechanism lacks a learnable semantic link to the visual features, often resulting in a semantic-spatial gap where the model hallucinates coordinates that do not correspond to image evidences. In this work, we introduce Composer, a MLLM that leverages a novel visual grounding mechanism based on learned proxy-tokens to promote faithful interpretability. These discrete symbolic pointers explicitly index the image latent space, allowing the model to manipulate visual regions as addressable, semantically manipulable sets. To rigorously validate our novel grounding mechanism, we constructed ComposerGCoT, a dataset synthesized to enable holistic assessment of reasoning consistency and grounding accuracy. Experimental results indicate that Composer achieves performance parity with its coordinate-based counterpart in final answer accuracy, while improving visual grounding accuracy by +9.0 points. By demonstrating that discrete proxy-tokens capture spatial semantics more effectively than typical textual coordinates, we establish that visual grounding mechanisms with learnable semantic links represent a promising path toward trustworthy and reliable MLLMs.

cs.CV

Improving Controllable Generation: Faster Training and Better Performance via $x_0$-Supervision

Text-to-Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisely control image layouts, something that natural language alone cannot reliably express. Controllable generation methods augment the initial T2I model with additional conditions that more easily describe the scene. Prior works straightforwardly train the augmented network with the same loss as the initial network. Although natural at first glance, this can lead to very long training times in some cases before convergence. In this work, we revisit the training objective of controllable diffusion models through a detailed analysis of their denoising dynamics. We show that direct supervision on the clean target image, dubbed $x_0$-supervision, or an equivalent re-weighting of the diffusion loss, yields faster convergence. Experiments on multiple control settings demonstrate that our formulation accelerates convergence by up to 2$\times$ according to our novel metric (mean Area Under the Convergence Curve - mAUCC), while also improving both visual quality and conditioning accuracy. Our code is available at https://github.com/CEA-LIST/x0-supervision

cs.CV

Modeling the Happiness-Sustainability Nexus via Graphical Lasso and Quantile-on-Quantile Regression

This paper investigates the nexus between subjective well-being and sustainability, proxied by the Sustainable Development Goals (SDG) Index, using cross-country data from 126 nations in 2022. While prior research has highlighted a positive association between happiness and sustainable development, existing approaches largely rely on linear regressions or correlation-based measures that mask distributional heterogeneity, multicollinearity, and potential nonlinear dependence. To address these limitations, we employ a two methodological framework combining Graphical Lasso, and Quantile-on-Quantile Regression (QQR). The Graphical Lasso identifies a direct conditional link between happiness and sustainability after controlling for governance, income, and life expectancy, with a partial correlation of about 0.21. On the other hand, QQR reveals heterogeneous effects across the joint distribution: sustainability gains are positively associated with happiness for low-happiness but high-sustainability countries, negatively associated in high-happiness but low-sustainability contexts, and essentially neutral elsewhere. These findings suggest that the happiness-sustainability link is modest, asymmetric, and context-dependent, underscoring the importance of moving beyond mean-based regressions. From a policy perspective, our results highlight that institutional quality, income, and demographic factors remain the dominant drivers of both happiness and sustainability, while the interplay between the two dimensions is most pronounced in distributional extremes.

econ.EM

Market competition and poverty dynamics: Short and long run effects across financial development levels

This paper investigates how market competition influences poverty dynamics using a functional econometric framework that captures both contemporaneous and lagged effects. Using annual data for 48 countries from 1991-2017, we estimate function-on-function regressions linking poverty headcount ratios to market concentration and other macroeconomic indicators. The results show that, based on the entire sample, stronger competition initially increased poverty during structural adjustment phases, but its adverse impact weakened after 2010 as economies adapted and efficiency gains emerged. The estimated bivariate surfaces reveal that the effect of competition on poverty often persists over multiple years (around 5 years), highlighting the importance of intertemporal transmission. Then, functional clustering based on market capitalization (MCAP) uncovers strong heterogeneity: pro-poor 5-years lagged effect of competition in low- and medium-MCAP economies, while it remains insignificant to weakly negative in high-MCAP countries. Overall, the findings underscore the value of functional data methods in uncovering evolving and lag-dependent poverty-competition linkages that static panel models fail to capture.

econ.EM

3D-COCO: extension of MS-COCO dataset for image detection and 3D reconstruction modules

We introduce 3D-COCO, an extension of the original MS-COCO dataset providing 3D models and 2D-3D alignment annotations. 3D-COCO was designed to achieve computer vision tasks such as 3D reconstruction or image detection configurable with textual, 2D image, and 3D CAD model queries. We complete the existing MS-COCO dataset with 28K 3D models collected on ShapeNet and Objaverse. By using an IoU-based method, we match each MS-COCO annotation with the best 3D models to provide a 2D-3D alignment. The open-source nature of 3D-COCO is a premiere that should pave the way for new research on 3D-related topics. The dataset and its source codes is available at https://kalisteo.cea.fr/index.php/coco3d-object-detection-and-reconstruction/

cs.CV

Online Nonparametric Supervised Learning for Massive Data

Despite their benefits in terms of simplicity, low computational cost and data requirement, parametric machine learning algorithms, such as linear discriminant analysis, quadratic discriminant analysis or logistic regression, suffer from serious drawbacks including linearity, poor fit of features to the usually imposed normal distribution and high dimensionality. Batch kernel-based nonparametric classifier, which overcomes the linearity and normality of features constraints, represent an interesting alternative for supervised classification problem. However, it suffers from the ``curse of dimension". The problem can be alleviated by the explosive sample size in the era of big data, while large-scale data size presents some challenges in the storage of data and the calculation of the classifier. These challenges make the classical batch nonparametric classifier no longer applicable. This motivates us to develop a fast algorithm adapted to the real-time calculation of the nonparametric classifier in massive as well as streaming data frameworks. This online classifier includes two steps. First, we consider an online principle components analysis to reduce the dimension of the features with a very low computation cost. Then, a stochastic approximation algorithm is deployed to obtain a real-time calculation of the nonparametric classifier. The proposed methods are evaluated and compared to some commonly used machine learning algorithms for real-time fetal well-being monitoring. The study revealed that, in terms of accuracy, the offline (or Batch), as well as, the online classifiers are good competitors to the random forest algorithm. Moreover, we show that the online classifier gives the best trade-off accuracy/computation cost compared to the offline classifier.

stat.ML

Functional conditional volatility modeling with missing data: inference and application to energy commodities

This paper explores the nonparametric estimation of the volatility component in a heteroscedastic scalar-on-function regression model, where the underlying discrete-time process is ergodic and subject to a missing-at-random mechanism. We first propose a simplified estimator for the regression and volatility operators, constructed solely from the observed data. The asymptotic properties of these estimators, including the almost sure uniform consistency rate and asymptotic distribution, are rigorously analyzed. Subsequently, the simplified estimators are employed to impute the missing data in the original process, enhancing the estimation of the regression and volatility components. The asymptotic behavior of these imputed estimators is also thoroughly investigated. A numerical comparison of the simplified and imputed estimators is presented using simulated data. Finally, the methodology is applied to real-world data to model the volatility of daily natural gas returns, utilizing intraday EU/USD exchange rate return curves sampled at a 1-hour frequency.

stat.ME

Generalized regression operator estimation for continuous time functional data processes with missing at random response

In this paper, we are interested in nonparametric kernel estimation of a generalized regression function, including conditional cumulative distribution and conditional quantile functions, based on an incomplete sample $(X_t, Y_t, ζ_t)_{t\in \mathbb{ R}^+}$ copies of a continuous-time stationary ergodic process $(X, Y, ζ)$. The predictor $X$ is valued in some infinite-dimensional space, whereas the real-valued process $Y$ is observed when $ζ= 1$ and missing whenever $ζ= 0$. Pointwise and uniform consistency (with rates) of these estimators as well as a central limit theorem are established. Conditional bias and asymptotic quadratic error are also provided. Asymptotic and bootstrap-based confidence intervals for the generalized regression function are also discussed. A first simulation study is performed to compare the discrete-time to the continuous-time estimations. A second simulation is also conducted to discuss the selection of the optimal sampling mesh in the continuous-time case. Finally, it is worth noting that our results are stated under ergodic assumption without assuming any classical mixing conditions.

math.ST

Joint parametric specification checking of conditional mean and volatility in time series models with martingale difference innovations

Using cumulative residual processes, we propose joint goodness-of-fit tests for conditional means and variances functions in the context of nonlinear time series with martingale difference innovations. The main challenge comes from the fact the cumulative residual process no longer admits, under the null hypothesis, a distribution-free limit. To obtain a practical solution one either transforms the process in order to achieve a distribution-free limit or approximates the non-distribution free limit using a numerical or a re-sampling technique. Here the three solutions will be considered.It is shown that the proposed tests have nontrivial power against a class of root-n local alternatives, and are suitable when the conditioning information set is infinite-dimensional, which allows including models like autoregressive conditional heteroscedastic stochastic models with dependent innovations. The approach presented assumes only certain conditions on the first- and second-order conditional moments, without imposing any autoregression model. The test procedures introduced are compared with each other and with other competitors in terms of their power using a simulation study and a real data application. These simulations have shown that the statistical powers of tests based on re-sampling or numerical approximation of the original statistics are in general slightly better than those based on a martingale transformation of the original process.

stat.ME

LapNet : Automatic Balanced Loss and Optimal Assignment for Real-Time Dense Object Detection

Real-time single-stage object detectors based on deep learning still remain less accurate than more complex ones. The trade-off between model performance and computational speed is a major challenge. In this paper, we propose a new way to efficiently learn a single-shot detector which offers a very good compromise between these two objectives. To this end, we introduce LapNet, an anchor based detector, trained end-to-end without any sampling strategy. Our approach aims to overcome two important problems encountered in training an anchor based detector: (1) ambiguity in the assignment of anchor to ground truth and (2) class and object size imbalance. To address the first limitation, we propose a soft positive/negative anchor assignment procedure based on a new overlapping function called "Per-Object Normalized Overlap" (PONO). This soft assignment can be self-corrected by the network itself to avoid ambiguity between close objects. To cope with the second limitation, we propose to learn additional weights, that are not used at inference, to efficiently manage sample imbalance. These two contributions make the detector learning more generic whatever the training dataset. Various experiments show the effectiveness of the proposed approach.

cs.CV

Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image

In this paper, we present a novel approach, called Deep MANTA (Deep Many-Tasks), for many-task vehicle analysis from a given image. A robust convolutional network is introduced for simultaneous vehicle detection, part localization, visibility characterization and 3D dimension estimation. Its architecture is based on a new coarse-to-fine object proposal that boosts the vehicle detection. Moreover, the Deep MANTA network is able to localize vehicle parts even if these parts are not visible. In the inference, the network's outputs are used by a real time robust pose estimation algorithm for fine orientation estimation and 3D vehicle localization. We show in experiments that our method outperforms monocular state-of-the-art approaches on vehicle detection, orientation and 3D location tasks on the very challenging KITTI benchmark.

cs.CV

Nonparametric M-estimation for right censored regression model with stationary ergodic data

The present paper deals with a nonparametric M-estimation for right censored regression model with stationary ergodic data. Defined as an implicit function, a kernel type estimator of a family of robust regression is considered when the covariate take its values in R^d (d >= 1) and the data are sampled from stationary ergodic process. The strong consistency (with rate) and the asymptotic distribution of the estimator are established under mild assumptions. Moreover, a usable confidence interval is provided which does not depend on any unknown quantity. Our results hold without any mixing condition and do not require the existence of marginal densities. A comparison study based on simulated data is also provided.

stat.ME

Rate of uniform consistency for a class of mode regression on functional stationary ergodic data. Application to electricity consumption

The aim of this paper is to study the asymptotic properties of a class of kernel conditional mode estimates whenever functional stationary ergodic data are considered. To be more precise on the matter, in the ergodic data setting, we consider a random element $(X, Z)$ taking values in some semi-metric abstract space $E\times F$. For a real function $φ$ defined on the space $F$ and $x\in E$, we consider the conditional mode of the real random variable $φ(Z)$ given the event $``X=x"$. While estimating the conditional mode function, say $θ_φ(x)$, using the well-known kernel estimator, we establish the strong consistency with rate of this estimate uniformly over Vapnik-Chervonenkis classes of functions $φ$. Notice that the ergodic setting offers a more general framework than the usual mixing structure. Two applications to energy data are provided to illustrate some examples of the proposed approach in time series forecasting framework. The first one consists in forecasting the {\it daily peak} of electricity demand in France (measured in Giga-Watt). Whereas the second one deals with the short-term forecasting of the electrical {\it energy} (measured in Giga-Watt per Hour) that may be consumed over some time intervals that cover the peak demand.

stat.ME

Nonparametric Multivariate L1-median Regression Estimation with Functional Covariates

In this paper, a nonparametric estimator is proposed for estimating the L1-median for multivariate conditional distribution when the covariates take values in an infinite dimensional space. The multivariate case is more appropriate to predict the components of a vector of random variables simultaneously rather than predicting each of them separately. While estimating the conditional L1-median function using the well-known Nadarya-Waston estimator, we establish the strong consistency of this estimator as well as the asymptotic normality. We also present some simulations and provide how to built conditional con?fidence ellipsoids for the multivariate L1-median regression in practice. Some numerical study in chemiometrical real data are carried out to compare the multivariate L1-median regression with the vector of marginal median regression when the covariate X is a curve as well as X is a random vector.

math.ST

Kernel-smoothed conditional quantiles of randomly censored functional stationary ergodic data

This paper, investigates the conditional quantile estimation of a scalar random response and a functional random covariate (i.e. valued in some infinite-dimensional space) whenever {\it functional stationary ergodic data with random censorship} are considered. We introduce a kernel type estimator of the conditional quantile function. We establish the strong consistency with rate of this estimator as well as the asymptotic normality which induces a confidence interval that is usable in practice since it does not depend on any unknown quantity. An application to electricity peak demand interval prediction with censored smart meter data is carried out to show the performance of the proposed estimator.

math.ST

Using complex surveys to estimate the $L_1$-median of a functional variable: application to electricity load curves

Mean profiles are widely used as indicators of the electricity consumption habits of customers. Currently, in Électricité De France (EDF), class load profiles are estimated using point-wise mean function. Unfortunately, it is well known that the mean is highly sensitive to the presence of outliers, such as one or more consumers with unusually high-levels of consumption. In this paper, we propose an alternative to the mean profile: the $L_1$-median profile which is more robust. When dealing with large datasets of functional data (load curves for example), survey sampling approaches are useful for estimating the median profile avoiding storing the whole data. We propose here estimators of the median trajectory using several sampling strategies and estimators. A comparison between them is illustrated by means of a test population. We develop a stratification based on the linearized variable which substantially improves the accuracy of the estimator compared to simple random sampling without replacement. We suggest also an improved estimator that takes into account auxiliary information. Some potential areas for future research are also highlighted.

stat.OT

Properties of Design-Based Functional Principal Components Analysis

This work aims at performing Functional Principal Components Analysis (FPCA) with Horvitz-Thompson estimators when the observations are curves collected with survey sampling techniques. One important motivation for this study is that FPCA is a dimension reduction tool which is the first step to develop model assisted approaches that can take auxiliary information into account. FPCA relies on the estimation of the eigenelements of the covariance operator which can be seen as nonlinear functionals. Adapting to our functional context the linearization technique based on the influence function developed by Deville (1999), we prove that these estimators are asymptotically design unbiased and consistent. Under mild assumptions, asymptotic variances are derived for the FPCA' estimators and consistent estimators of them are proposed. Our approach is illustrated with a simulation study and we check the good properties of the proposed estimators of the eigenelements as well as their variance estimators obtained with the linearization approach.

math.ST