SearcharxivSearch

arXiv subjects

Jeremy Guntoro

Publications and source records attributed to Jeremy Guntoro.

2 recordsLinked to original sources

Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes

Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored. We ask how much biosecurity-relevant signal is linearly accessible in these representations by training minimal linear and attention probes on frozen Evo 2 layer-26 activations, without fine-tuning the underlying model. Across held-out metagenomic test sets, the probes detect antimicrobial resistance (AMR) with strong discrimination: a linear probe reaches a region-level ROC-AUC of 0.888 (mean-pool), rising to 0.977 with a single-head attention probe. The probes resolve finer-grained AMR drug-class subcategories and separate them from unrelated functional genes, providing additional evidence that the learned signal is not explained solely by generic functional-gene status. Bacterial virulence is also decodable, though more weakly (region-level ROC-AUC 0.833). The AMR probe retains comparable ranking performance on simulated short reads without retraining, enabling evaluation before assembly in settings where assembly is computationally costly or unreliable. It achieves a read-level ROC-AUC of 0.898 (mean-pool), comparable to the mean-pooled full-region result. Within SynGenome, AMR-associated prompt labels are only weakly recoverable from Evo 1.5-generated sequences; these prompt-derived labels do not establish the function of the generated response sequences. A complementary sparse-autoencoder analysis recovers interpretable resistance-associated features but proves less consistent than the supervised probes. Together, these results position lightweight embedding-based probes as a fast, inexpensive first-pass detection layer for metagenomic biosurveillance and map both strengths and current limits of the approach. This work was conducted as part of the AIxBio Hackathon 2026 hosted by BlueDot Impact, Apart Research, and Cambridge Biosecurity Hub.

q-bio.GN

The Role of Sequence Information in Minimal Models of Molecular Assembly

Sequence-directed assembly processes - such as protein folding - allow the assembly of a large number of structures with high accuracy from only a small handful of fundamental building blocks. We aim to explore how efficiently sequence information can be used to direct assembly by studying variants of the temperature-1 abstract tile assembly model (aTAM). We ask whether, for each variant, their exists a finite set of tile types that can deterministically assemble any shape producible by a given assembly model; we call such tile type sets "universal assembly kits". Our first model, which we call the "backboned aTAM", generates backbone-assisted assembly by forcing tiles to be added to lattice positions neighbouring the immediately preceding tile, using a predetermined sequence of tile types. We demonstrate the existence of universal assembly kit for the backboned aTAM, and show that the existence of this set is maintained even under stringent restrictions to the rules of assembly. We compare these results to a less constrained model that we call sequenced aTAM, which also uses a predetermined sequence of tiles, but does not constrain a tile to neighbour the immediately preceding tiles. We prove that this model has no universal assembly kit in the stringent case. The lack of such a kit is surprising, given that the number of tile sequences of length N scales faster than both the number and worst-case Kolmogorov complexity of producible shapes of size N for a sufficiently large - but finite - set of tiles. Our results demonstrate the importance of physical mechanisms, and specifically geometric constraints, in facilitating efficient use of the information in molecular programs for structure assembly.

physics.bio-ph