SearcharxivSearch

arXiv subjects

Junhan Zhao

Publications and source records attributed to Junhan Zhao.

17 recordsLinked to original sources

Quantum Sampling Architecture for Protein Structure Reconstruction on Utility-Scale Hardware

Predicting the structure of short peptides in protein binding pockets remains difficult because this regime requires physics-based conformational search, yet existing methods do not provide a practical way to carry out that search on current hardware. We present QSAD, a quantum-classical framework that reformulates peptide structure prediction as amino-acid-level Hamiltonian sampling and replaces iterative optimization with non-iterative Hamiltonian evolution. Executed entirely on IBM Heron R2 across 101 binding-pocket peptides (5-18 residues), QSAD improves prediction accuracy by 27-71% over all evaluated AI and quantum baselines while maintaining the lowest variance across tested lengths. QSAD also tolerates noise levels 3-5x beyond typical hardware error rates, where iterative methods fail, and reduces mean quantum execution time by 27x relative to VQE. The sampled ensemble further supports approximate reconstruction of protein energy landscapes. These results establish coarse-grained quantum sampling as a practical computational path for structure prediction in regimes where data-driven methods lack sufficient signal.

cs.ET

Discovery of a Featureless Tidal Disruption Event at z~1 with the Wide Field Survey Telescope

We report the discovery of tidal disruption event (TDE) WFST250820mmsw/AT2025wet by the 2.5-meter Wide Field Survey Telescope (WFST). It exhibits a blue nuclear flare throughout the observed evolution with a g-band peak magnitude ~22, which is about 3 magnitudes brighter than its host galaxy. A Keck/LRIS spectrum taken near the optical peak reveals a featureless blue continuum, with no discernible emission lines. However, its redshift can be accurately determined to be 1.037 by its host galaxy absorption lines. Blackbody fits to the multiband spectral energy distribution (SED) of AT2025wet yield a constant temperature of ~19,000K and a peak luminosity of (8.27 +0.92 -0.71)*10^44 erg s^-1 while actually the SED likely peaks at a much shorter wavelength than a 19,000K blackbody. The SED modeling of the host galaxy implies a stellar mass of ~10^11.2 M_odot and an estimated central black hole mass of ~10^8 M_odot, with no evidence of significant active galactic nucleus activity prior to the flare. All of these observations are well consistent with a featureless TDE scenario, making it the highest-redshift non-jetted TDE known to date. TDEs at such high redshift provide us a unique opportunity to explore the intrinsic SEDs of TDEs, particularly to test whether they peak in the extreme-UV regime, thereby addressing the missing energy puzzle and the origin of optical emission in TDEs. Ongoing surveys represented by WFST and the Legacy Survey of Space and Time (LSST) are expected to discover an increasing number of TDEs at higher redshifts, which will extend our census of SMBHs across redshift space and help unravel the mysteries of optical TDEs through direct probes of their UV emission.

astro-ph.HE

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challenges: a lack of decision transparency, the difficulty of aligning response timing with visual evidence, and the need to maintain a global, causally consistent understanding under tight computational budgets. To address these issues, we propose a novel framework that decouples reasoning control from memory integration. We introduce \textbf{\model{}}, an instantiation of this framework with two core components. First, the \emph{Active Thinking Decision Maker (ATDM)} is a transparent reasoning controller that externalizes its decision process using observable progress ($\boldsymbol{\rho}$) and confidence ($\boldsymbol{c}$) metrics. This allows it to precisely time its response $t_r$ to match the first-sufficient-evidence timestamp $t^\star$ while streaming its reasoning to the user. Second, the \emph{Hierarchical Progressive Semantic Integration (HPSI)} module acts as an efficient memory system. It employs a set of learnable, multi-level aggregation tokens that are propagated across clips to build a rich, global cognitive state without exceeding token budgets. %Our approach sets a new standard on key online video understanding benchmarks, achieving strong performance of \textbf{71.6\%} on StreamingBench and \textbf{46.9\%} on OVOBench, demonstrating a robust solution for evidence-aligned and transparent online video analysis. Extensive experiments demonstrate the effectiveness of ATDM and HPSI, e.g., Thinking-QwenVL improves the accuracy of the previous state-of-the-art from 67.63\% to 71.60\% on the StreamingBench benchmark.

cs.CV

Computational Pathology in the Era of Emerging Foundation and Agentic AI -- International Expert Perspectives on Clinical Integration and Translational Readiness

Recent breakthroughs in artificial intelligence through foundation models and agents have accelerated the evolution of computational pathology. Demonstrated performance gains reported across academia in benchmarking datasets in predictive tasks such as diagnosis, prognosis, and treatment response have ignited substantial enthusiasm for clinical application. Despite this development momentum, real world adoption has lagged, as implementation faces economic, technical, and administrative challenges. Beyond existing discussions of technical architectures and comparative performance, this review considers how these emerging AI systems can be responsibly integrated into medical practice by connecting deployable clinical relevance with downstream analytical capabilities and their technical maturity, operational readiness, and economic and regulatory context. Drawing on perspectives from an international group, we provide a practical assessment of current capabilities and barriers to adoption in patient care settings.

cs.CE

WFST Supernovae in the First Year: I. Statistical Study of 16 Early-phase Type Ia Supernovae from the Pilot Survey

In this paper we present 16 early-phase type Ia supernovae (SNe Ia) discovered during the pilot survey of the 2.5-meter Wide Field Survey Telescope (WFST-PS) from March 4 to July 10, 2024, including three SNe Ia with early-excess emission features (EExSNe Ia). The discovery magnitude of the 16 WFST-PS early-phase SNe is at least 3 mag fainter than their peak brightness. A large scatter of color indices is found in approximately the first 10 days of supernova explosions, indicating diverse photometric behaviors in the early phase. Three EExSNe Ia show relatively brighter peak luminosities and longer rise time compared to those of non-EExSNe Ia. The results indicate that current theoretical models require further refinement to fully capture the early photometric evolution of SNe Ia. Based on the initial high-cadence ugr-band data from the WFST-PS survey, we emphasize that early near-ultraviolet (NUV) observations are indispensable for placing tight constraints on the explosion mechanisms and progenitor systems of SNe Ia.

astro-ph.HE

WFST Supernovae in the First Year: II. SN 2024aedt: Systematical Study of a Transitional Type Ia Supernova

We present comprehensive photometric and spectroscopic observations of a transitional type Ia SN 2024aedt, discovered by the 2.5-meter Wide Field Survey Telescope (WFST) within one day of the explosion. Its light curve is characterized by a peak absolute magnitude of $M_B = -18.49 \pm 0.03$ mag and a decline rate of $\Delta m_{15}(B) = 1.53 \pm 0.36$ mag, placing the object on the $\Delta m_{15}(B)$--$M_B$ diagram in the transition region between normal and subluminous SNe Ia. Furthermore, the early-color evolution and host galaxy environment of SN 2024aedt underscore its transitional nature, sharing properties with both normal and 91bg-like SNe Ia. Light-curve modeling with MOSFiT yields a synthesized $^{56}\mathrm{Ni}$ mass of $0.414 \pm 0.042\,M_{\odot}$ and a total ejecta mass of $0.548 \pm 0.108\,M_{\odot}$. A comparison with theoretical models suggests that the evolutionary trend can be broadly explained by both delayed-detonation (DDT) and double-detonation (DDet) scenarios while possible early-excess emissions predicted by DDet cannot be identified given the limited detections soon after the SN explosion. Although the overall spectral evolution of SN 2024aedt is similar to that of other transitional SNe Ia, the spectroscopic comparison reveals diversity in the early-phase blue-end features, which becomes more homogeneous at later phases. The result indicates the importance of early-time observations in understanding the origin of SN Ia diversity.

astro-ph.HE

WFST Supernovae in the First Year: III. Systematical Study of the Photometric Behavior of Early-phase Core-collapse Supernovae

We investigate the multiband photometric properties of seven supernovae (SNe) showing double-peaked light-curve evolution and prominent shock-cooling emission, observed by the Wide Field Survey Telescope (WFST) during its first year of operation. By jointly employing an analytic early shock-cooling model and the Arnett radioactive-diffusion model, we fit the bolometric light curves and infer ejecta masses in the range $1.1$-$2.6 M_\odot$, consistent with a transitional population between ultra-stripped supernovae (USSNe) and normal stripped-envelope supernovae (SESNe). The envelope masses are estimated to be $M_{\rm env}=0.1$-$0.4 M_\odot$, while the progenitors are constrained to be yellow or blue supergiants (YSGs/BSGs) with radii of $R=120$-$300 R_\odot$. Using empirical relations, we estimate progenitor luminosities of $L=10^{4.6}$-$10^{4.9} L_\odot$, corresponding to zero-age main-sequence (ZAMS) masses of $8$-$20 M_\odot$. Theoretical models suggest that such progenitors are more naturally produced through binary evolution channels, as single-star evolutionary pathways are unable to yield ejecta masses this low.

astro-ph.HE

Illuminating the Mass Gap Through Deep Optical Constraint on a Neutron Star Merger Candidate S250206dm

The gravitational wave (GW) event S250206dm, as the first well-localized neutron star merger candidate potentially located in the mass gap, presented a unique opportunity to probe the electromagnetic signatures from such a system. Here we report a deep, multiband search with the new 2.5-meter Wide Field Survey Telescope (WFST), covering about 64% of the localization region up to a 5-sigma limiting magnitude of 23 mag. In total, 12 potential candidates have been identified while none of them are likely related to S250206dm. This non-detection provides the most stringent constraint to date on any associated kilonova. Crucially, an AT 2017gfo-like event at 269 Mpc can be excluded by WFST observations alone. Based on ejecta mass limits, a neutron star-black hole with a large mass ratio (Q >= 3.2) is disfavored. This optical-derived constraint on the mass ratio reaches, for the first time, a precision comparable to that inferred from the GW signal. This work presents the best observation of this type of events until now, and demonstrates the power of rapid, deep follow-up observations to constrain the properties of compact binary progenitors, offering key insights into the constituents of the mass gap.

astro-ph.HE

HyperADRs: A Hierarchical Hypergraph Framework for Drug-Gene-ADR Prediction

Adverse drug reactions (ADRs) are a major barrier to safe and effective pharmacotherapy and increasingly reflect higher order interactions between drugs, genetic background, and clinical phenotypes. Existing graph based approaches usually predict ADRs as properties of drugs or drug pairs, leaving the causal gene implicit and limiting their value for pharmacogenomic decision making. We introduce HyperADRs, a hierarchical hypergraph framework that predicts ADR risk at the level of drug-gene-ADR triads. Starting from curated pharmacogenomic annotations in PharmGKB and the pharmacogenomics subdatabase of DrugBank, we construct high confidence triplets and integrate them with auxiliary molecular, functional, and disease relations from precision-medicine-oriented knowledge graphs. Drugs, genes, and ADR concepts are embedded with modality appropriate pretrained models (UniMol, ESM2, SapBERT) and propagated through a hypergraph convolutional network. A FiLM based, query conditioned contrastive learning module learns context specific representations so that, given any two entities, the model retrieves the correct third entity against many candidates. To improve robustness and interpretability, we propose a nine category ADR macro system scheme that reduces large heterogeneous "other" bins while aligning with organ system reasoning in clinical pharmacology. Across drug-, gene-, and ADR-held-out evaluations on PharmGKB, HyperADRs matches or exceeds strong baselines on ranking based metrics. When trained on PharmGKB and tested on unseen DrugBank triplets, HyperADRs maintains its ranking advantage, indicating that the learned representations capture transferable biological mechanisms and can support mechanistically grounded pharmacogenomic hypothesis generation.

q-bio.QM

A Hybrid Quantum-AI Framework for Protein Structure Prediction on NISQ Devices

Variational quantum algorithms provide a direct, physics-based approach to protein structure prediction, but their accuracy is limited by the coarse resolution of the energy landscapes generated on current noisy devices. We propose a hybrid framework that combines quantum computation with deep learning, formulating structure prediction as a problem of energy fusion. Candidate conformations are obtained through the Variational Quantum Eigensolver (VQE) executed on IBM's 127-qubit superconducting processor, which defines a global yet low-resolution quantum energy surface. To refine these basins, secondary structure probabilities and dihedral angle distributions predicted by the NSP3 neural network are incorporated as statistical potentials. These additional terms sharpen the valleys of the quantum landscape, resulting in a fused energy function that enhances effective resolution and better distinguishes native-like structures. Evaluation on 375 conformations from 75 protein fragments shows consistent improvements over AlphaFold3, ColabFold, and quantum-only predictions, achieving a mean RMSD of 4.9 {\AA} with statistical significance (p < 0.001). The findings demonstrate that energy fusion offers a systematic method for combining data-driven models with quantum algorithms, improving the practical applicability of near-term quantum computing to molecular and structural biology.

cs.ET

A Generative Foundation Model for Chest Radiography

The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. Here we develop `ChexGen', a generative vision-language foundation model that introduces a unified framework for text-, mask-, and bounding box-guided synthesis of chest radiographs. Built upon the latent diffusion transformer architecture, ChexGen was pretrained on the largest curated chest X-ray dataset to date, consisting of 960,000 radiograph-report pairs. ChexGen achieves accurate synthesis of radiographs through expert evaluations and quantitative metrics. We demonstrate the utility of ChexGen for training data augmentation and supervised pretraining, which led to performance improvements across disease classification, detection, and segmentation tasks using a small fraction of training data. Further, our model enables the creation of diverse patient cohorts that enhance model fairness by detecting and mitigating demographic biases. Our study supports the transformative role of generative foundation models in building more accurate, data-efficient, and equitable medical AI systems.

cs.CV

DISPROTBENCH: Uncovering the Functional Limits of Protein Structure Prediction Models in Intrinsically Disordered Regions

Intrinsically disordered regions (IDRs) play central roles in cellular function, yet remain poorly evaluated by existing protein structure prediction benchmarks. Current evaluations largely focus on well-folded domains, overlooking three fundamental challenges in realistic biological settings: the structural complexity of proteins, the resulting low availability of reliable ground truth, and prediction uncertainty that can propagate into high-risk downstream failures, such as in drug discovery, protein-protein interaction modeling, and functional annotation. We present DisProtBench, an IDR-centric benchmark that explicitly incorporates prediction uncertainty into the evaluation of protein structure prediction models (PSPMs). To address structural complexity and ground-truth scarcity, we curate and unify a large-scale, multi-modal dataset spanning disease-relevant IDRs, GPCR-ligand interactions, and multimeric protein complexes. To assess predictive uncertainty, we introduce Functional Uncertainty Sensitivity (FUS), a novel prediction uncertainty-stratified metric that quantifies downstream task performance under prediction uncertainty. Using this benchmark, we conduct a systematic evaluation of state-of-the-art PSPMs and reveal clear, task-dependent failure modes. Protein-protein interaction prediction degrades sharply in IDRs, while structure-based drug discovery remains comparatively robust. These effects are largely invisible to standard global accuracy metrics, which overestimate functional reliability under prediction uncertainty. We have open-sourced our benchmark and the codebase at https://github.com/Susan571/DisProtBench.

q-bio.BM

Radiance Field Learners As UAV First-Person Viewers

First-Person-View (FPV) holds immense potential for revolutionizing the trajectory of Unmanned Aerial Vehicles (UAVs), offering an exhilarating avenue for navigating complex building structures. Yet, traditional Neural Radiance Field (NeRF) methods face challenges such as sampling single points per iteration and requiring an extensive array of views for supervision. UAV videos exacerbate these issues with limited viewpoints and significant spatial scale variations, resulting in inadequate detail rendering across diverse scales. In response, we introduce FPV-NeRF, addressing these challenges through three key facets: (1) Temporal consistency. Leveraging spatio-temporal continuity ensures seamless coherence between frames; (2) Global structure. Incorporating various global features during point sampling preserves space integrity; (3) Local granularity. Employing a comprehensive framework and multi-resolution supervision for multi-scale scene feature representation tackles the intricacies of UAV video spatial scales. Additionally, due to the scarcity of publicly available FPV videos, we introduce an innovative view synthesis method using NeRF to generate FPV perspectives from UAV footage, enhancing spatial perception for drones. Our novel dataset spans diverse trajectories, from outdoor to indoor environments, in the UAV domain, differing significantly from traditional NeRF scenarios. Through extensive experiments encompassing both interior and exterior building structures, FPV-NeRF demonstrates a superior understanding of the UAV flying space, outperforming state-of-the-art methods in our curated UAV dataset. Explore our project page for further insights: https://fpv-nerf.github.io/.

cs.CV

Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation

Recently, large-scale text-to-image (T2I) diffusion models have emerged as a powerful tool for image-to-image translation (I2I), allowing open-domain image translation via user-provided text prompts. This paper proposes frequency-controlled diffusion model (FCDiffusion), an end-to-end diffusion-based framework that contributes a novel solution to text-guided I2I from a frequency-domain perspective. At the heart of our framework is a feature-space frequency-domain filtering module based on Discrete Cosine Transform, which filters the latent features of the source image in the DCT domain, yielding filtered image features bearing different DCT spectral bands as different control signals to the pre-trained Latent Diffusion Model. We reveal that control signals of different DCT spectral bands bridge the source image and the T2I generated image in different correlations (e.g., style, structure, layout, contour, etc.), and thus enable versatile I2I applications emphasizing different I2I correlations, including style-guided content creation, image semantic manipulation, image scene translation, and image style translation. Different from related approaches, FCDiffusion establishes a unified text-guided I2I framework suitable for diverse image translation tasks simply by switching among different frequency control branches at inference time. The effectiveness and superiority of our method for text-guided I2I are demonstrated with extensive experiments both qualitatively and quantitatively. Our project is publicly available at: https://xianggao1102.github.io/FCDiffusion/.

cs.CV

Federated attention consistent learning models for prostate cancer diagnosis and Gleason grading

Artificial intelligence (AI) holds significant promise in transforming medical imaging, enhancing diagnostics, and refining treatment strategies. However, the reliance on extensive multicenter datasets for training AI models poses challenges due to privacy concerns. Federated learning provides a solution by facilitating collaborative model training across multiple centers without sharing raw data. This study introduces a federated attention-consistent learning (FACL) framework to address challenges associated with large-scale pathological images and data heterogeneity. FACL enhances model generalization by maximizing attention consistency between local clients and the server model. To ensure privacy and validate robustness, we incorporated differential privacy by introducing noise during parameter transfer. We assessed the effectiveness of FACL in cancer diagnosis and Gleason grading tasks using 19,461 whole-slide images of prostate cancer from multiple centers. In the diagnosis task, FACL achieved an area under the curve (AUC) of 0.9718, outperforming seven centers with an average AUC of 0.9499 when categories are relatively balanced. For the Gleason grading task, FACL attained a Kappa score of 0.8463, surpassing the average Kappa score of 0.7379 from six centers. In conclusion, FACL offers a robust, accurate, and cost-effective AI training model for prostate cancer pathology while maintaining effective data safeguards.

cs.CV

Phoenixmap: An Abstract Approach to Visualize 2D Spatial Distributions

The multidimensional nature of spatial data poses a challenge for visualization. In this paper, we introduce Phoenixmap, a simple abstract visualization method to address the issue of visualizing multiple spatial distributions at once. The Phoenixmap approach starts by identifying the enclosed outline of the point collection, then assigns different widths to outline segments according to the segments' corresponding inside regions. Thus, one 2D distribution is represented as an outline with varied thicknesses. Phoenixmap is capable of overlaying multiple outlines and comparing them across categories of objects in a 2D space. We chose heatmap as a benchmark spatial visualization method and conducted user studies to compare performances among Phoenixmap, heatmap, and dot distribution map. Based on the analysis and participant feedback, we demonstrate that Phoenixmap 1) allows users to perceive and compare spatial distribution data efficiently; 2) frees up graphics space with a concise form that can provide visualization design possibilities like overlapping; and 3) provides a good quantitative perceptual estimating capability given the proper legends. Finally, we discuss several possible applications of Phoenixmap and present one visualization of multiple species of birds' active regions in a nature preserve.

cs.HC

Memorizing All for Implicit Discourse Relation Recognition

Implicit discourse relation recognition is a challenging task due to the absence of the necessary informative clue from explicit connectives. The prediction of relations requires a deep understanding of the semantic meanings of sentence pairs. As implicit discourse relation recognizer has to carefully tackle the semantic similarity of the given sentence pairs and the severe data sparsity issue exists in the meantime, it is supposed to be beneficial from mastering the entire training data. Thus in this paper, we propose a novel memory mechanism to tackle the challenges for further performance improvement. The memory mechanism is adequately memorizing information by pairing representations and discourse relations of all training instances, which right fills the slot of the data-hungry issue in the current implicit discourse relation recognizer. Our experiments show that our full model with memorizing the entire training set reaches new state-of-the-art against strong baselines, which especially for the first time exceeds the milestone of 60% accuracy in the 4-way task.

cs.CL