SearcharxivSearch

arXiv subjects

Simon Lee

Publications and source records attributed to Simon Lee.

7 recordsLinked to original sources

Case study of a national-level academic conference organised in hybrid mode at low cost

In July 2025, the University of Adelaide hosted the Astronomical Society of Australia's Annual Scientific Meeting on its North Terrace campus. We ran the conference in a hybrid mode, with options for in-person and online attendance. This report details the procedures that we used to enable the online mode of the conference at minimal cost and minimal inconvenience to the in-person attendees. We discuss our choices of hardware and software and how we integrated these systems together. We summarise our experience of organising a local AV team and the procedures that we set for running the AV in each session. We present statistics of the online attendance numbers and post-conference survey feedback, and discuss the lessons we feel other organisers may particularly be able to learn from.

astro-ph.IM

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level reconstruction over high-level functional context. We introduce JEPA-DNA, a model-agnostic continual training framework that integrates a Joint-Embedding Predictive Architecture (JEPA) with traditional generative objectives. By supervising global sequence embeddings in a latent space, JEPA-DNA forces models to predict the functional representations of masked genomic segments, shifting the learning signal from token recovery to semantic alignment. We evaluate JEPA-DNA on 17 diverse genomic benchmark tasks, demonstrating consistent gains in linear probing and zero-shot performance regardless of the underlying GFM architecture or generative objective. Our framework establishes a new state-of-the-art for GFMs, surpassing the best existing models by bridging generative precision with latent semantic grounding. Through extensive ablation studies, we further characterize the synergistic interplay between generative and latent objectives. Our code is publicly available at https://github.com/NVIDIA-Digital-Bio/JEPA-DNA.

cs.AI

LFM2 Technical Report

We present LFM2, a family of Liquid Foundation Models designed for efficient on-device deployment and strong task capabilities. Using hardware-in-the-loop architecture search under edge latency and memory constraints, we obtain a compact hybrid backbone that combines gated short convolutions with a small number of grouped query attention blocks, delivering up to 2x faster prefill and decode on CPUs compared to similarly sized models. The LFM2 family covers 350M-8.3B parameters, including dense models (350M, 700M, 1.2B, 2.6B) and a mixture-of-experts variant (8.3B total, 1.5B active), all with 32K context length. LFM2's training pipeline includes a tempered, decoupled Top-K knowledge distillation objective that avoids support mismatch; curriculum learning with difficulty-ordered data; and a three-stage post-training recipe of supervised fine-tuning, length-normalized preference optimization, and model merging. Pre-trained on 10-12T tokens, LFM2 models achieve strong results across diverse benchmarks; for example, LFM2-2.6B reaches 79.56% on IFEval and 82.41% on GSM8K. We further build multimodal and retrieval variants: LFM2-VL for vision-language tasks, LFM2-Audio for speech, and LFM2-ColBERT for retrieval. LFM2-VL supports tunable accuracy-latency tradeoffs via token-efficient visual processing, while LFM2-Audio separates audio input and output pathways to enable real-time speech-to-speech interaction competitive with models 3x larger. LFM2-ColBERT provides a low-latency encoder for queries and documents, enabling high-performance retrieval across multiple languages. All models are released with open weights and deployment packages for ExecuTorch, llama.cpp, and vLLM, making LFM2 a practical base for edge applications that need fast, memory-efficient inference and strong task capabilities.

cs.LG

Reflections from Research Roundtables at the Conference on Health, Inference, and Learning (CHIL) 2025

The 6th Annual Conference on Health, Inference, and Learning (CHIL 2025), hosted by the Association for Health Learning and Inference (AHLI), was held in person on June 25-27, 2025, at the University of California, Berkeley, in Berkeley, California, USA. As part of this year's program, we hosted Research Roundtables to catalyze collaborative, small-group dialogue around critical, timely topics at the intersection of machine learning and healthcare. Each roundtable was moderated by a team of senior and junior chairs who fostered open exchange, intellectual curiosity, and inclusive engagement. The sessions emphasized rigorous discussion of key challenges, exploration of emerging opportunities, and collective ideation toward actionable directions in the field. In total, eight roundtables were held by 19 roundtable chairs on topics of "Explainability, Interpretability, and Transparency," "Uncertainty, Bias, and Fairness," "Causality," "Domain Adaptation," "Foundation Models," "Learning from Small Medical Data," "Multimodal Methods," and "Scalable, Translational Healthcare Solutions."

cs.LG

A Case Study Exploring the Current Landscape of Synthetic Medical Record Generation with Commercial LLMs

Synthetic Electronic Health Records (EHRs) offer a valuable opportunity to create privacy preserving and harmonized structured data, supporting numerous applications in healthcare. Key benefits of synthetic data include precise control over the data schema, improved fairness and representation of patient populations, and the ability to share datasets without concerns about compromising real individuals privacy. Consequently, the AI community has increasingly turned to Large Language Models (LLMs) to generate synthetic data across various domains. However, a significant challenge in healthcare is ensuring that synthetic health records reliably generalize across different hospitals, a long standing issue in the field. In this work, we evaluate the current state of commercial LLMs for generating synthetic data and investigate multiple aspects of the generation process to identify areas where these models excel and where they fall short. Our main finding from this work is that while LLMs can reliably generate synthetic health records for smaller subsets of features, they struggle to preserve realistic distributions and correlations as the dimensionality of the data increases, ultimately limiting their ability to generalize across diverse hospital settings.

cs.CL

Optimising an Array of Cherenkov Telescopes in Australia for the Detection of TeV Gamma-Ray Transients

As TeV gamma-ray astronomy progresses into the era of the Cherenkov Telescope Array (CTA), instantaneously following up on gamma-ray transients is becoming more important than ever. To this end, a worldwide network of Imaging Atmospheric Cherenkov Telescopes has been proposed. Australia is ideally suited to provide coverage of part of the Southern Hemisphere sky inaccessible to H.E.S.S. in Namibia and the upcoming CTA-South in Chile. This study assesses the sources detectable by a small, transient-focused array in Australia based on CTA telescope designs. The TeV emission of extragalactic sources (including the majority of gamma-ray transients) can suffer significant absorption by the extragalactic background light. As such, we explored the improvements possible by implementing stereoscopic and topological triggers, as well as lowered image cleaning thresholds, to access lower energies. We modelled flaring gamma-ray sources based on past measurements from the satellite-based gamma-ray telescope Fermi-LAT. We estimate that an array of four Medium-Sized Telescopes (MSTs) would detect $\sim$24 active galactic nucleus flares >5$\sigma$ per year, up to a redshift of $z\approx1.5$. Two MSTs achieved $\sim$80-90% of the detections of four MSTs. The modelled Galactic transients were detectable within the observation time of one night, 11 of the 21 modelled gamma-ray bursts were detectable, as were $\sim$10% of unidentified transients. An array of MST-class telescopes would thus be a valuable complementary telescope array for transient TeV gamma-ray astronomy.

astro-ph.IM

Performance of a Small Array of Imaging Air Cherenkov Telescopes sited in Australia

As TeV gamma-ray astronomy progresses into the era of the Cherenkov Telescope Array (CTA), there is a desire for the capacity to instantaneously follow up on transient phenomena and continuously monitor gamma-ray flux at energies above $10^{12}$ eV. To this end, a worldwide network of Imaging Air Cherenkov Telescopes (IACTs) is required to provide triggers for CTA observations and complementary continuous monitoring. An IACT array sited in Australia would contribute significant coverage of the Southern Hemisphere sky. Here, we investigate the suitability of a small IACT array and how different design factors influence its performance. Monte Carlo simulations were produced based on the Small-Sized Telescope (SST) and Medium-Sized Telescope (MST) designs from CTA. Angular resolution improved with larger baseline distances up to 277m between telescopes, and energy thresholds were lower at 1000m altitude than at 0m. The $\sim$300 GeV energy threshold of MSTs proved more suitable for observing transients than the $\sim$1.2 TeV threshold of SSTs. An array of four MSTs at 1000m was estimated to give a 5.7$\sigma$ detection of an RS Ophiuchi-like nova eruption from a 4-hour observation. We conclude that an array of four MST-class IACTs at an Australian site would ideally complement the capabilities of CTA.

astro-ph.IM