Searcharxiv⌕ Search

arXiv subjects

João Silva

Publications and source records attributed to João Silva.

13 recordsLinked to original sources

Copula Transformations for Data-Consistent Inversion

Data-consistent inversion (DCI) constructs probability measures whose push-forward distributions agree with observed data, while iterative data-consistent inversion (iDCI) extends this framework to generalized stochastic inverse problems by enforcing multiple push-forward constraints sequentially. Although iDCI avoids the direct approximation of high-dimensional joint densities, its relationship to the original joint DCI solution has remained unclear. In this work, we establish this relationship through copula theory. Using Sklar's theorem, we derive a factorization of the DCI update into separate marginal and dependence transformations and show that the discrepancy remaining after convergence of the iDCI algorithm is entirely characterized by the copulas associated with the observed and predicted joint distributions. This characterization motivates a copula-transformed iDCI solution, and we prove that an exact copula transformation recovers the original DCI solution. We further establish convergence results for approximate copula transformations under converging sequences of reference measures and progressively enriched feasible sets. Numerical examples demonstrate how the geometry induced by the quantity-of-interest map governs the importance of the copula transformation, illustrate an adaptive reference-measure refinement strategy for improving computational accuracy under a fixed sampling budget, and demonstrate the progressive refinement of generalized stochastic inverse problems through heterogeneous, asynchronously acquired experiments.

stat.ML↗

Progressing beyond Art Masterpieces or Touristic Clichés: how to assess your LLMs for cultural alignment?

Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- until recently there has been limited work on the design and development of datasets for cultural assessment. Here, we review existing approaches to such datasets and identify their main limitations. To address these issues, we propose design guidelines for annotators and report on the construction of a dataset built according to these principles. We further present a series of contrastive experiments conducted with this dataset. The results demonstrate that our design yields test sets with greater discriminative power, effectively distinguishing between models specialized for a given culture and those that are not, ceteris paribus.

cs.CL↗

CLARIN-PT-LDB: An Open LLM Leaderboard for Portuguese to assess Language, Culture and Civility

This paper reports on the development of a leaderboard of Open Large Language Models (LLM) for European Portuguese (PT-PT), and on its associated benchmarks. This leaderboard comes as a way to address a gap in the evaluation of LLM for European Portuguese, which so far had no leaderboard dedicated to this variant of the language. The paper also reports on novel benchmarks, including some that address aspects of performance that so far have not been available in benchmarks for European Portuguese, namely model safeguards and alignment to Portuguese culture. The leaderboard is available at https://huggingface.co/spaces/PORTULAN/portuguese-llm-leaderboard.

cs.CL↗

Sovereign AI-based Public Services are Viable and Affordable

The rapid expansion of AI-based remote services has intensified debates about the long-term implications of growing structural concentration in infrastructure and expertise. As AI capabilities become increasingly intertwined with geopolitical interests, the availability and reliability of foundational AI services can no longer be taken for granted. This issue is particularly pressing for AI-enabled public services for citizens, as governments and public agencies are progressively adopting 24/7 AI-driven support systems typically operated through commercial offerings from a small oligopoly of global technology providers. This paper challenges the prevailing assumption that general-purpose architectures, offered by these providers, are the optimal choice for all application contexts. Through practical experimentation, we demonstrate that viable and cost-effective alternatives exist. Alternatives that align with principles of digital and cultural sovereignty. Our findings provide an empirical illustration that sovereign AI-based public services are both technically feasible and economically sustainable, capable of operating effectively on premises with modest computational and financial resources while maintaining cultural and digital autonomy. The technical insights and deployment lessons reported here are intended to inform the adoption of similar sovereign AI public services by national agencies and governments worldwide.

cs.CL↗

Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training

Instruction-guided image editing consists in taking an image and an instruction and deliverring that image altered according to that instruction. State-of-the-art approaches to this task suffer from the typical scaling up and domain adaptation hindrances related to supervision as they eventually resort to some kind of task-specific labelling, masking or training. We propose a novel approach that does without any such task-specific supervision and offers thus a better potential for improvement. Its assessment demonstrates that it is highly effective, achieving very competitive performance.

cs.CL↗

Leveraging LLMs for On-the-Fly Instruction Guided Image Editing

The combination of language processing and image processing keeps attracting increased interest given recent impressive advances that leverage the combined strengths of both domains of research. Among these advances, the task of editing an image on the basis solely of a natural language instruction stands out as a most challenging endeavour. While recent approaches for this task resort, in one way or other, to some form of preliminary preparation, training or fine-tuning, this paper explores a novel approach: We propose a preparation-free method that permits instruction-guided image editing on the fly. This approach is organized along three steps properly orchestrated that resort to image captioning and DDIM inversion, followed by obtaining the edit direction embedding, followed by image editing proper. While dispensing with preliminary preparation, our approach demonstrates to be effective and competitive, outperforming recent, state of the art models for this task when evaluated on the MAGICBRUSH dataset.

cs.CL↗

Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family

Sentence encoder encode the semantics of their input, enabling key downstream applications such as classification, clustering, or retrieval. In this paper, we present Serafim PT*, a family of open-source sentence encoders for Portuguese with various sizes, suited to different hardware/compute budgets. Each model exhibits state-of-the-art performance and is made openly available under a permissive license, allowing its use for both commercial and research purposes. Besides the sentence encoders, this paper contributes a systematic study and lessons learned concerning the selection criteria of learning objectives and parameters that support top-performing encoders.

cs.CL↗

Rigidity of free boundary minimal disks in mean convex three-manifolds

The purpose of this article is study rigidity of free boundary minimal two-disks that locally maximize the modified Hawking mass on a Riemannian three-manifold with positive lower bound on its scalar curvature and mean convex boundary. Assuming the strict stability of Σ, we prove that a neighborhood of it in M is isometric to one of the half de Sitter-Schwarzschild space.

math.DG↗

Advancing Generative AI for Portuguese with Open Decoder Gervásio PT*

To advance the neural decoding of Portuguese, in this paper we present a fully open Transformer-based, instruction-tuned decoder model that sets a new state of the art in this respect. To develop this decoder, which we named Gervásio PT*, a strong LLaMA~2 7B model was used as a starting point, and its further improvement through additional training was done over language resources that include new instruction data sets of Portuguese prepared for this purpose, which are also contributed in this paper. All versions of Gervásio are open source and distributed for free under an open license, including for either research or commercial usage, and can be run on consumer-grade hardware, thus seeking to contribute to the advancement of research and innovation in language technology for Portuguese.

cs.CL↗

Fostering the Ecosystem of Open Neural Encoders for Portuguese with Albertina PT* Family

To foster the neural encoding of Portuguese, this paper contributes foundation encoder models that represent an expansion of the still very scarce ecosystem of large language models specifically developed for this language that are fully open, in the sense that they are open source and openly distributed for free under an open license for any purpose, thus including research and commercial usages. Like most languages other than English, Portuguese is low-resourced in terms of these foundational language resources, there being the inaugural 900 million parameter Albertina and 335 million Bertimbau. Taking this couple of models as an inaugural set, we present the extension of the ecosystem of state-of-the-art open encoders for Portuguese with a larger, top performance-driven model with 1.5 billion parameters, and a smaller, efficiency-driven model with 100 million parameters. While achieving this primary goal, further results that are relevant for this ecosystem were obtained as well, namely new datasets for Portuguese based on the SuperGLUE benchmark, which we also distribute openly.

cs.CL↗

Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*

To advance the neural encoding of Portuguese (PT), and a fortiori the technological preparation of this language for the digital age, we developed a Transformer-based foundation model that sets a new state of the art in this respect for two of its variants, namely European Portuguese from Portugal (PT-PT) and American Portuguese from Brazil (PT-BR). To develop this encoder, which we named Albertina PT-*, a strong model was used as a starting point, DeBERTa, and its pre-training was done over data sets of Portuguese, namely over data sets we gathered for PT-PT and PT-BR, and over the brWaC corpus for PT-BR. The performance of Albertina and competing models was assessed by evaluating them on prominent downstream language processing tasks adapted for Portuguese. Both Albertina PT-PT and PT-BR versions are distributed free of charge and under the most permissive license possible and can be run on consumer-grade hardware, thus seeking to contribute to the advancement of research and innovation in language technology for Portuguese.

cs.CL↗

Delocalization of topological surface states by diagonal disorder in nodal loop semimetals

The effect of Anderson diagonal disorder on the topological surface (``drumhead'') states of a Weyl nodal loop semimetal is addressed. Since diagonal disorder breaks chiral symmetry, a winding number cannot be defined. Seen as a perturbation, the weak random potential mixes the clean exponentially localized drumhead states of the semimetal, thereby producing two effects: (i) the algebraic decay of the surface states into the bulk; (ii) a broadening of the low energy density of surface states of the open system due to degeneracy lifting. This behavior persists with increasing disorder, up to the bulk semimetal-to-metal transition at the critical disorder $W_{c}$. Above $W_{c}$, the surface states hybridize with bulk states and become extended into the bulk.

cond-mat.dis-nn↗

An Alternating Direction Algorithm for Hybrid Precoding and Combining in Millimeter Wave MIMO Systems

Millimeter-wave (mmWave) technology is one of the most promising candidates for future wireless communication systems as it can offer large underutilized bandwidths and eases the implementation of large antenna arrays which are required to help overcome the severe signal attenuation that occurs at these frequencies. To reduce the high cost and power consumption of a fully digital mmWave precoder and combiner, hybrid analog/digital designs based on analog phase shifters are often adopted. In this work we derive an iterative algorithm for the hybrid precoding and combining design for spatial multiplexing in mmWave massive multiple-input multiple-output (MIMO) systems. To cope with the difficulty of handling the hardware constraint imposed by the analog phase shifters we use the alternating direction method of the multipliers (ADMM) to split the hybrid design problem into a sequence of smaller subproblems. This results in an iterative algorithm where the design of the analog precoder/combiner consists of a closed form solution followed by a simple projection over the set of matrices with equal magnitude elements. It is initially developed for the fully-connected structure and then extended to the partially-connected architecture which allows simpler hardware implementation. Furthermore, to cope with the more likely wideband scenarios where the channel is frequency selective, we also extend the algorithm to an orthogonal frequency division multiplexing (OFDM) based mmWave system. Simulation results in different scenarios show that the proposed design algorithms are capable of achieving performances close to the optimal fully digital solution and can work with a broad range of configuration of antennas, RF chains and data streams.

eess.SP↗