SearcharxivSearch

arXiv subjects

Aaron Haag

Publications and source records attributed to Aaron Haag.

4 recordsLinked to original sources

Test-Time Scaling for CAD Generation via Verifier-Free Consensus Selection

Large language models can write parametric CAD programs from a natural-language description (text-to-CAD generation), but a single sample is often wrong. Increasing test-time compute by sampling multiple candidates only helps if a good candidate can be identified, yet no ground-truth model is available at generation time. Existing systems often require a separate verifier, such as a vision-language judge, to select among candidates. We investigate whether the candidate pool itself provides enough signal for effective selection and a verifier-free alternative. We introduce 3D CAD consensus selection, hereafter consensus selection: sample $N$ parametric CAD programs, compile them to 3D models, and return the candidate that agrees most with the rest of the pool. The method is training-free and compatible with existing CAD agents. We investigate geometric and topological notions of agreement, each of which improves its corresponding evaluation metric. On the exact candidate pools of a state-of-the-art CAD generation method, geometric consensus improves all three geometric metrics over the method's verifier, while topological consensus matches it on topology. Across every tested LLM and prompt variant, geometric consensus also improves geometric accuracy over random selection from the same pool, reducing Chamfer distance by $1-10\%$.

cs.CE

FeatureFox: Sample-Efficient Panoptic Graph Segmentation for Machining Feature Recognition in B-Rep 3D-CAD Models

Automatic feature recognition (AFR) on B-Rep 3D-CAD models is central to CAD/CAM automation, yet most learning-based methods are complex, data-hungry, and evaluate instance grouping and semantic labeling separately. We present FeatureFox, a panoptic AFR pipeline that outputs machining instances with semantic labels: a calibrated binary edge classifier on enriched edge attributes localizes feature boundaries, instances are recovered as connected components in a pruned face-adjacency graph, and a per-instance classifier predicts the machining class from aggregated subgraph attributes. We evaluate on MFInstSeg using Panoptic Quality (PQ), which jointly scores instance separation and semantic correctness. FeatureFox is substantially more sample- and compute-efficient than the deep baseline AAGNet, reaching $\mathrm{PQ}>0.9$ with $\sim250$ training parts versus $\sim5{,}000$ for AAGNet, and training on the full MFInstSeg set takes seconds on a GPU. On the full training set, AAGNet surpasses FeatureFox marginally in PQ, while FeatureFox remains slightly ahead in feature-level recognition and localization accuracy. Finally, leveraging its low data requirement, we train FeatureFox on $270$ manually labeled industrial CAD parts and show qualitative generalization to an unseen real industrial part, indicating practical real-world applicability.

cs.CE

Training LLMs for Generating IEC 61131-3 Structured Text with Online Feedback

IEC 61131-3 Structured Text (ST) is a widely used programming language for programmable logic controllers (PLCs) in automation systems. However, generating ST code with LLMs poses unique challenges due to limited data in public training datasets and the complexity of ST language syntax. This paper proposes an approach to fine-tune LLMs for the generation of ST code that leverages a preference-based learning method through an online process involving compiler feedback and evaluation from an LLM-based ST expert. In this framework, the model is iteratively refined and generates new training samples, which are subsequently evaluated by a compiler for syntactical correctness and by a specialized LLM that excels at assessing semantic accuracy, though it is not optimized for code generation itself. This approach results in marked improvements for the trained LLM, leading to higher compilation success rates and better semantic precision. As a result, the framework proves highly suitable for industrial automation applications and outperforms state-of-the-art models.

cs.SE

Joint Embeddings for Graph Instruction Tuning

Large Language Models (LLMs) have achieved impressive performance in text understanding and have become an essential tool for building smart assistants. Originally focusing on text, they have been enhanced with multimodal capabilities in recent works that successfully built visual instruction following assistants. As far as the graph modality goes, however, no such assistants have yet been developed. Graph structures are complex in that they represent relation between different features and are permutation invariant. Moreover, representing them in purely textual form does not always lead to good LLM performance even for finetuned models. As a result, there is a need to develop a new method to integrate graphs in LLMs for general graph understanding. This work explores the integration of the graph modality in LLM for general graph instruction following tasks. It aims at producing a deep learning model that enhances an underlying LLM with graph embeddings and trains it to understand them and to produce, given an instruction, an answer grounded in the graph representation. The approach performs significantly better than a graph to text approach and remains consistent even for larger graphs.

cs.SE