SearcharxivSearch

arXiv subjects

Joseph Cho

Publications and source records attributed to Joseph Cho.

At least 19 recordsLinked to original sources

Ruled zero mean curvature surfaces in the three-dimensional light cone

We obtain a complete classification of ruled zero mean curvature surfaces in the three-dimensional light cone. En route, we examine geodesics and screw motions in the space form, allowing us to discover helicoids. We also consider their relationship to catenoids using Weierstrass representations of zero mean curvature surfaces in the three-dimensional light cone.

math.DG

Monodromy of Darboux transformations of polarised curves

We show that every finite type polarised curve in the conformal $2$-sphere with a polynomial conserved quantity admits a resonance point, under a non-orthogonality assumption on the conserved quantity. Using this fact, we deduce that every finite type curve polarised by space form arc-length in the conformal $2$-sphere admits a resonance point, possibly on a multiple cover.

math.DG

Zero mean curvature surfaces in isotropic space with planar curvature lines

We give a comprehensive account of zero mean curvature surfaces in isotropic 3-space with planar curvature lines. After giving a complete classification all such surfaces, we show that they belong to a 1-parameter family of surfaces. We then investigate their relationship to Thomsen-type surfaces in isotropic 3-space, those zero mean curvature surfaces in isotropic 3-space that are also affine minimal.

math.DG

SurGen: Text-Guided Diffusion Model for Surgical Video Generation

Diffusion-based video generation models have made significant strides, producing outputs with improved visual fidelity, temporal coherence, and user control. These advancements hold great promise for improving surgical education by enabling more realistic, diverse, and interactive simulation environments. In this study, we introduce SurGen, a text-guided diffusion model tailored for surgical video synthesis. SurGen produces videos with the highest resolution and longest duration among existing surgical video generation models. We validate the visual and temporal quality of the outputs using standard image and video generation metrics. Additionally, we assess their alignment to the corresponding text prompts through a deep learning classifier trained on surgical data. Our results demonstrate the potential of diffusion models to serve as valuable educational tools for surgical trainees.

cs.CV

GP-VLS: A general-purpose vision language model for surgery

Surgery requires comprehensive medical knowledge, visual assessment skills, and procedural expertise. While recent surgical AI models have focused on solving task-specific problems, there is a need for general-purpose systems that can understand surgical scenes and interact through natural language. This paper introduces GP-VLS, a general-purpose vision language model for surgery that integrates medical and surgical knowledge with visual scene understanding. For comprehensively evaluating general-purpose surgical models, we propose SurgiQual, which evaluates across medical and surgical knowledge benchmarks as well as surgical vision-language questions. To train GP-VLS, we develop six new datasets spanning medical knowledge, surgical textbooks, and vision-language pairs for tasks like phase recognition and tool identification. We show that GP-VLS significantly outperforms existing open- and closed-source models on surgical vision-language tasks, with 8-21% improvements in accuracy across SurgiQual benchmarks. GP-VLS also demonstrates strong performance on medical and surgical knowledge tests compared to open-source alternatives. Overall, GP-VLS provides an open-source foundation for developing AI assistants to support surgeons across a wide range of tasks and scenarios. The code and data for this work is publicly available at gpvls-surgery-vlm.github.io.

cs.CV

A Generalist Model for Diverse Text-Guided Medical Image Synthesis

Deep learning algorithms require extensive data to achieve robust performance. However, data availability is often restricted in the medical domain due to patient privacy concerns. Synthetic data presents a possible solution to these challenges. Image generative models have found increasing use for medical applications, but are often task-specific, thus limiting their scalability. Moreover, existing models frequently rely on private datasets for training, which constrain their reproducibility. To address this, we introduce MediSyn: an open-access, generalist, text-guided latent diffusion model capable of generating synthetic images across 6 medical specialties and 10 imaging modalities, while being trained exclusively on publicly available data. Through extensive experimentation, we provide several key contributions. First, we demonstrate that training a generative model on visually diverse medical images does not degrade synthetic image quality. Second, we show that this generalist approach is substantially more computationally efficient than a coordinated suite of task-specific models. Third, we establish that a generalist model can produce realistic, text-aligned synthetic images across visually and medically distinct modalities, as validated by expert physicians. Fourth, we provide empirical evidence that these synthetic images are visually distinct from their corresponding real patient images, alleviating concerns about data memorization in image generative models. Finally, we demonstrate that a generalist model can produce synthetic images that improve classifier performance in data-limited settings across multiple medical specialties. Altogether, our findings highlight the immense potential of generalist image generative models to accelerate algorithmic research and development in medicine.

cs.CV

Almanac Copilot: Towards Autonomous Electronic Health Record Navigation

Clinicians spend large amounts of time on clinical documentation, and inefficiencies impact quality of care and increase clinician burnout. Despite the promise of electronic medical records (EMR), the transition from paper-based records has been negatively associated with clinician wellness, in part due to poor user experience, increased burden of documentation, and alert fatigue. In this study, we present Almanac Copilot, an autonomous agent capable of assisting clinicians with EMR-specific tasks such as information retrieval and order placement. On EHR-QA, a synthetic evaluation dataset of 300 common EHR queries based on real patient data, Almanac Copilot obtains a successful task completion rate of 74% (n = 221 tasks) with a mean score of 2.45 over 3 (95% CI:2.34-2.56). By automating routine tasks and streamlining the documentation process, our findings highlight the significant potential of autonomous agents to mitigate the cognitive load imposed on clinicians by current EMR systems.

cs.AI

Sora as a World Model? A Complete Survey on Text-to-Video Generation

The evolution of video generation from text, from animating MNIST to simulating the world with Sora, has progressed at a breakneck speed. Here, we systematically discuss how far text-to-video generation technology supports essential requirements in world modeling. We curate 250+ studies on text-based video synthesis and world modeling. We then observe that recent models increasingly support spatial, action, and strategic intelligences in world modeling through adherence to completeness, consistency, invention, as well as human interaction and control. We conclude that text-to-video generation is adept at world modeling, although homework in several aspects, such as the diversity-consistency trade-offs, remains to be addressed.

cs.AI

Discrete constant mean curvature cylinders and isothermic tori

We consider the monodromy problem of Darboux transforms of discrete isothermic surfaces using the integrable theory of discrete polarised curves. Then we provide, for the first time, closed-form discrete parametrisations of discrete isothermic cylinders, discrete constant mean curvature cylinders, and discrete isothermic tori.

math.DG

A Generalizable Deep Learning System for Cardiac MRI

Cardiac MRI allows for a comprehensive assessment of myocardial structure, function and tissue characteristics. Here we describe a foundational vision system for cardiac MRI, capable of representing the breadth of human cardiovascular disease and health. Our deep-learning model is trained via self-supervised contrastive learning, in which visual concepts in cine-sequence cardiac MRI scans are learned from the raw text of the accompanying radiology reports. We train and evaluate our model on data from four large academic clinical institutions in the United States. We additionally showcase the performance of our models on the UK BioBank and two additional publicly available external datasets. We explore emergent capabilities of our system and demonstrate remarkable performance across a range of tasks, including the problem of left-ventricular ejection fraction regression and the diagnosis of 39 different conditions such as cardiac amyloidosis and hypertrophic cardiomyopathy. We show that our deep-learning system is capable of not only contextualizing the staggering complexity of human cardiovascular disease but can be directed towards clinical problems of interest, yielding impressive, clinical-grade diagnostic accuracy with a fraction of the training data typically required for such tasks.

eess.IV

Lie minimal Weingarten surfaces

We consider Lie minimal surfaces, the critical points of the simplest Lie sphere invariant energy, in Riemannian space forms. These surfaces can be characterized via their Euler-Lagrange equations, which take the form of differential equations of the principal curvatures. Surfaces with constant mean curvature that satisfy these equations turn out to be rotational in their space form. We generalize in flat ambient space: here surfaces where the principal curvatures satisfy an affine relationship as well as elliptic linear Weingarten surfaces are rotational as well.

math.DG

Periodic discrete Darboux transforms

We express Darboux transformations of discrete polarised curves as parallel sections of discrete connections in the quaternionic formalism. This immediately leads to the linearisation of the monodromy of the transformation. We also consider the integrable reduction to the case of discrete bicycle correspondence. Applying our method to the case of discrete circles, we obtain closed-form discrete parametrisations of all (closed) Darboux transforms and (closed) bicycle correspondences.

math.DG

A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering

The Segment Anything Model (SAM), developed by Meta AI Research, represents a significant breakthrough in computer vision, offering a robust framework for image and video segmentation. This survey provides a comprehensive exploration of the SAM family, including SAM and SAM 2, highlighting their advancements in granularity and contextual understanding. Our study demonstrates SAM's versatility across a wide range of applications while identifying areas where improvements are needed, particularly in scenarios requiring high granularity and in the absence of explicit prompts. By mapping the evolution and capabilities of SAM models, we offer insights into their strengths and limitations and suggest future research directions, including domain-specific adaptations and enhanced memory and propagation mechanisms. We believe that this survey comprehensively covers the breadth of SAM's applications and challenges, setting the stage for ongoing advancements in segmentation technology.

cs.CV

Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era

Generative AI has made significant progress in recent years, with text-guided content generation being the most practical as it facilitates interaction between human instructions and AI-generated content (AIGC). Thanks to advancements in text-to-image and 3D modeling technologies, like neural radiance field (NeRF), text-to-3D has emerged as a nascent yet highly active research field. Our work conducts a comprehensive survey on this topic and follows up on subsequent research progress in the overall field, aiming to help readers interested in this direction quickly catch up with its rapid development. First, we introduce 3D data representations, including both Structured and non-Structured data. Building on this pre-requisite, we introduce various core technologies to achieve satisfactory text-to-3D results. Additionally, we present mainstream baselines and research directions in recent text-to-3D technology, including fidelity, efficiency, consistency, controllability, diversity, and applicability. Furthermore, we summarize the usage of text-to-3D technology in various applications, including avatar generation, texture generation, scene generation and 3D editing. Finally, we discuss the agenda for the future development of text-to-3D.

cs.CV

Spinor representation in isotropic 3-space via Laguerre geometry

We give a detailed description of the geometry of isotropic space, in parallel to those of Euclidean space within the realm of Laguerre geometry. After developing basic surface theory in isotropic space, we define spin transformations, directly leading to the spinor representation of conformal surfaces in isotropic space. As an application, we obtain the Weierstrass-type representation for zero mean curvature surfaces, and the Kenmotsu-type representation for constant mean curvature surfaces, allowing us to construct many explicit examples.

math.DG