SearcharxivSearch

arXiv subjects

Luis Lara

Publications and source records attributed to Luis Lara.

11 recordsLinked to original sources

Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards

An AI system for professional floor plan design must precisely control room dimensions and areas while respecting the desired connectivity between rooms and maintaining functional and aesthetic quality. Existing generative approaches focus primarily on respecting the requested connectivity between rooms, but do not support generating floor plans that respect numerical constraints. We introduce a text-based floor plan generation approach that fine-tunes a large language model (LLM) on real plans and then applies reinforcement learning with verifiable rewards (RLVR) to improve adherence to topological and numerical constraints while discouraging invalid or overlapping outputs. Furthermore, we design a set of constraint adherence metrics to systematically measure how generated floor plans align with user-defined constraints. Our model generates floor plans that satisfy user-defined connectivity and numerical constraints and outperforms existing methods on Realism, Compatibility, and Diversity metrics. Across all tasks, our approach achieves at least a 94% relative reduction in Compatibility compared with existing methods. Our results demonstrate that LLMs can effectively handle constraints in this setting, suggesting broader applications for text-based generative modeling.

cs.CL

The Invisibility Hypothesis: Promises of AGI and the Future of the Global South

Discussions surrounding Artificial General Intelligence have largely focused on technical feasibility, timelines, and existential risk, often treating its social impact as being the same across different populations. Less attention has been paid to how advanced AI systems may interact with existing global inequalities. This paper examines the implications of AGI for people in the Global South, arguing that the availability of highly autonomous, general-purpose cognitive systems does not guarantee equitable outcomes. We establish that, as scientific discovery, economic coordination, and governance become increasingly automated, the relevance of human individuals may become conditional on access to infrastructure, institutional inclusion, and geopolitical circumstances rather than skills or intelligence. Under this setting, the Global South faces different pathways: in the best case, geographic location is no longer relevant as AGI fully democratizes access to knowledge and essential services for everyone in the globe; in the worst case, existing structural constraints are severely amplified, rendering already marginalized populations not merely economically invisible, but functionally irrelevant to global systems. We ground this analysis in empirical signals from contemporary AI deployment and extend to potential trajectories, highlighting both risk and opportunity pathways for Latin America, Africa, and South Asia.

cs.CY

Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes

Video diffusion techniques have advanced significantly in recent years; however, they struggle to generate realistic imagery of car crashes due to the scarcity of accident events in most driving datasets. Improving traffic safety requires realistic and controllable accident simulations. To tackle the problem, we propose Ctrl-Crash, a controllable car crash video generation model that conditions on signals such as bounding boxes, crash types, and an initial image frame. Our approach enables counterfactual scenario generation where minor variations in input can lead to dramatically different crash outcomes. To support fine-grained control at inference time, we leverage classifier-free guidance with independently tunable scales for each conditioning signal. Ctrl-Crash achieves state-of-the-art performance across quantitative video quality metrics (e.g., FVD and JEDi) and qualitative measurements based on a human-evaluation of physical realism and video quality compared to prior diffusion-based methods.

cs.CV

Diagnosing COVID-19 Severity from Chest X-Ray Images Using ViT and CNN Architectures

The COVID-19 pandemic strained healthcare resources and prompted discussion about how machine learning can alleviate physician burdens and contribute to diagnosis. Chest x-rays (CXRs) are used for diagnosis of COVID-19, but few studies predict the severity of a patient's condition from CXRs. In this study, we produce a large COVID severity dataset by merging three sources and investigate the efficacy of transfer learning using ImageNet- and CXR-pretrained models and vision transformers (ViTs) in both severity regression and classification tasks. A pretrained DenseNet161 model performed the best on the three class severity prediction problem, reaching 80% accuracy overall and 77.3%, 83.9%, and 70% on mild, moderate and severe cases, respectively. The ViT had the best regression results, with a mean absolute error of 0.5676 compared to radiologist-predicted severity scores. The project's source code is publicly available.

eess.IV

DStruct2Design: Data and Benchmarks for Data Structure Driven Generative Floor Plan Design

Text conditioned generative models for images have yielded impressive results. Text conditioned floorplan generation as a special type of raster image generation task also received particular attention. However there are many use cases in floorpla generation where numerical properties of the generated result are more important than the aesthetics. For instance, one might want to specify sizes for certain rooms in a floorplan and compare the generated floorplan with given specifications Current approaches, datasets and commonly used evaluations do not support these kinds of constraints. As such, an attractive strategy is to generate an intermediate data structure that contains numerical properties of a floorplan which can be used to generate the final floorplan image. To explore this setting we (1) construct a new dataset for this data-structure to data-structure formulation of floorplan generation using two popular image based floorplan datasets RPLAN and ProcTHOR-10k, and provide the tools to convert further procedurally generated ProcTHOR floorplan data into our format. (2) We explore the task of floorplan generation given a partial or complete set of constraints and we design a series of metrics and benchmarks to enable evaluating how well samples generated from models respect the constraints. (3) We create multiple baselines by finetuning a large language model (LLM), Llama3, and demonstrate the feasibility of using floorplan data structure conditioned LLMs for the problem of floorplan generation respecting numerical constraints. We hope that our new datasets and benchmarks will encourage further research on different ways to improve the performance of LLMs and other generative modelling techniques for generating designs where quantitative constraints are only partially specified, but must be respected.

cs.CL

Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

An image editing model should be able to perform diverse edits, ranging from object replacement, changing attributes or style, to performing actions or movement, which require many forms of reasoning. Current general instruction-guided editing models have significant shortcomings with action and reasoning-centric edits. Object, attribute or stylistic changes can be learned from visually static datasets. On the other hand, high-quality data for action and reasoning-centric edits is scarce and has to come from entirely different sources that cover e.g. physical dynamics, temporality and spatial reasoning. To this end, we meticulously curate the AURORA Dataset (Action-Reasoning-Object-Attribute), a collection of high-quality training data, human-annotated and curated from videos and simulation engines. We focus on a key aspect of quality training data: triplets (source image, prompt, target image) contain a single meaningful visual change described by the prompt, i.e., truly minimal changes between source and target images. To demonstrate the value of our dataset, we evaluate an AURORA-finetuned model on a new expert-curated benchmark (AURORA-Bench) covering 8 diverse editing tasks. Our model significantly outperforms previous editing models as judged by human raters. For automatic evaluations, we find important flaws in previous metrics and caution their use for semantically hard editing tasks. Instead, we propose a new automatic metric that focuses on discriminative understanding. We hope that our efforts : (1) curating a quality training dataset and an evaluation benchmark, (2) developing critical evaluations, and (3) releasing a state-of-the-art model, will fuel further progress on general image editing.

cs.CV

Approximate solutions of one dimensional systems with fractional derivative

The fractional calculus is useful to model non-local phenomena. We construct a method to evaluate the fractional Caputo derivative by means of a simple explicit quadratic segmentary interpolation. This method yields to numerical resolution of ordinary fractional differential equations. Due to the non-locality of the fractional derivative, we may establish an equivalence between fractional oscillators and ordinary oscillators with a dissipative term.

math.NA

The direction of time: from the global arrow to the local arrow

In this paper we discuss the traditional approaches to the problem of the arrow of time. On the basis of this discussion we adopt a global and non-entropic approach, according to which the arrow of time has a global origin and is an intrinsic, geometrical feature of space-time. Finally, we show how the global arrow is translated into local terms as a local time-asymmetric flux of energy

quant-ph

The cosmological origin of time-asymmetry

In this paper we address the problem of the arrow of time from a cosmological point of view, rejecting the traditional entropic approach that defines the future direction of time as the direction of the entropy increase: from our perspective, the arrow of time has a global origin and it is an intrinsic, geometrical feature of space-time. Time orientability and existence of a cosmic time are necessary conditions for defining an arrow of time, which is manifested globally as the time-asymmetry of the universe as a whole, and locally as a time-asymmetric energy flux. We also consider arrows of time of different origins (quantum, electromagnetic, thermodynamic, etc.) showing that they can be non-conventionally defined only if the geometrical arrow is previously defined.

quant-ph

Rigged Hilbert spaces and time-asymmetry: the case of the upside-down simple harmonic oscillator

The upside-down simple harmonic oscillator system is studied in the contexts of quantum mechanics and classical statistical mechanics. It is shown that in order to study in a simple manner the creation and decay of a physical system by ways of Gamow vectors we must formulate the theory in a time-asymmetric fashion, namely using two different rigged Hilbert spaces to describe states evolving towards the past and the future. The spaces defined in the contexts of quantum and classical statistical mechanics are shown to be directly related by the Wigner function.

quant-ph