SearcharxivSearch

arXiv subjects

Hui Shi

Publications and source records attributed to Hui Shi.

At least 19 recordsLinked to original sources

Decreasing Digital Distraction in College Students: Associated Online Learning Strategies Identified by Unsupervised Data Mining Approaches

The proliferation of digital tools in education offers numerous benefits but also introduces significant challenges, notably digital distractions that hinder academic performance, especially in online learning contexts. This study employed unsupervised data mining techniques, specifically association rule mining and clustering analysis, to identify effective learning strategies associated with lower levels of digital distractions among college students. Data from 530 participants revealed that self-regulated learning strategies (i.e., goal setting, environment structuring, and time management) co-occurred most consistently with lower digital distractions. Additionally, learner-instructor and learner-content engagement strategies, as well as technical competencies, also tended to appear in the same profiles as lower distraction. Interestingly, reliance on peer help-seeking and learner-learner engagement strategies appeared less often in those lower distraction profiles. These findings offer actionable implications for educators to design targeted interventions that foster focused and productive online learning environments.

cs.CY

A Spatially Resolved HI Survey of Seyfert Galaxies: the Role of AGN Feedback in Shaping Atomic Gas Reservoirs

Active galactic nucleus (AGN) feedback is a key ingredient in galaxy evolution, yet its impact on the cold atomic gas reservoir -- the neutral hydrogen (HI) phase -- remains poorly constrained. We present the most extensive spatially resolved HI 21-cm survey of Seyfert AGN hosts to date, based on observations with the Giant Metrewave Radio Telescope (GMRT). Our high-resolution HI maps of eight Seyfert galaxies reveal detailed kinematics and surface density distributions of their atomic gas disks. We find that AGN-host galaxies exhibit a slightly shallower HI mass-size relation than the canonical relation or the SIMBA simulation predictions; however, the measured slope remains consistent with the canonical value within $2\sigma$ uncertainties. This result suggests that AGN feedback does not significantly disrupt the global extent or large-scale structure of atomic gas reservoirs. To investigate the internal HI kinematics in greater detail, we perform a 3D kinematic forward modeling of the HI disk in UGC 4503. Our analysis reveals an elevated intrinsic velocity dispersion of $\sigma = 14.9^{+6.1}_{-3.8}$ km/s and a reduced level of rotational support, with $V/\sigma = 14.28_{-4.17}^{+4.97}$, compared to large-sample star-forming spirals. These kinematic signatures, together with localized residuals in the velocity field, indicate that AGN-driven outflows or jets may inject or indirectly affect the turbulence in the atomic gas disk, potentially regulating the cold gas reservoir. Future GMRT observations, combined with optical integral-field spectroscopy from MaNGA, will enable quantitative constraints on the role of AGN feedback in regulating star formation efficiency across a larger and more representative galaxy sample.

astro-ph.GA

Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation

Medical imaging quality control (QC) is essential for accurate diagnosis, yet traditional QC methods remain labor-intensive and subjective. To address this challenge, in this study, we establish a standardized dataset and evaluation framework for medical imaging QC, systematically assessing large language models (LLMs) in image quality assessment and report standardization. Specifically, we first constructed and anonymized a dataset of 161 chest X-ray (CXR) radiographs and 219 CT reports for evaluation. Then, multiple LLMs, including Gemini 2.0-Flash, GPT-4o, and DeepSeek-R1, were evaluated based on recall, precision, and F1 score to detect technical errors and inconsistencies. Experimental results show that Gemini 2.0-Flash achieved a Macro F1 score of 90 in CXR tasks, demonstrating strong generalization but limited fine-grained performance. DeepSeek-R1 excelled in CT report auditing with a 62.23\% recall rate, outperforming other models. However, its distilled variants performed poorly, while InternLM2.5-7B-chat exhibited the highest additional discovery rate, indicating broader but less precise error detection. These findings highlight the potential of LLMs in medical imaging QC, with DeepSeek-R1 and Gemini 2.0-Flash demonstrating superior performance.

cs.CL

The Variability of Persistent Radio Sources of Fast Radio Bursts

Over 700 bright millisecond-duration radio transients, known as Fast Radio Bursts (FRBs), have been identified to date. Nevertheless, the origin of FRBs remains unknown. The two repeating FRBs (FRB 20121102A and FRB 20190520B) have been verified to be associated with persistent radio sources (PRSs), making them the best candidates to study the nature of FRBs. Monitoring the variability in PRSs is essential for understanding their physical nature. We conducted 22 observations of the PRSs linked to FRB 20121102A and FRB 20190520B using the Karl G. Jansky Very Large Array (VLA), to study their variability. We have observed significant flux variability for the PRSs of FRB 20121102A and FRB 20190520B, with a confidence level exceeding 99.99%, based on the observations covering the longest timescale recorded to date. The observed variability of the two PRSs exhibits no significant difference in amplitude across both short and long timescales. We found that the radio-derived star formation rates of the two FRB hosts are significantly higher than those measured by the optical $H_{\alpha}$ emissions, indicating that their host galaxies are highly obscured or most radio emissions are not from star formation processes. The observed timescale of PRS flux evolution constrained the magnetic field of FRB 20121102A with $B_\parallel\gtrsim1~{\rm mG}$ and FRB 20190520B with $B_\parallel\gtrsim0.1~{\rm mG}$.

astro-ph.HE

Discovery of novel antimicrobial peptides with notable antibacterial potency by a LLM-based foundation model

Large language models (LLMs) have shown remarkable advancements in chemistry and biomedical research, acting as versatile foundation models for various tasks. We introduce AMP-Designer, an LLM-based approach for swiftly designing novel antimicrobial peptides (AMPs) with desired properties. Within 11 days, AMP-Designer achieved the de novo design of 18 AMPs with broad-spectrum activity against Gram-negative bacteria. In vitro validation revealed a 94.4% success rate, with two candidates demonstrating exceptional antibacterial efficacy, minimal hemotoxicity, stability in human plasma, and low potential to induce resistance, as evidenced by significant bacterial load reduction in murine lung infection experiments. The entire process, from design to validation, concluded in 48 days. AMP-Designer excels in creating AMPs targeting specific strains despite limited data availability, with a top candidate displaying a minimum inhibitory concentration of 2.0 {\mu}g/ml against Propionibacterium acnes. Integrating advanced machine learning techniques, AMP-Designer demonstrates remarkable efficiency, paving the way for innovative solutions to antibiotic resistance.

q-bio.BM

AAVDiff: Experimental Validation of Enhanced Viability and Diversity in Recombinant Adeno-Associated Virus (AAV) Capsids through Diffusion Generation

Recombinant adeno-associated virus (rAAV) vectors have revolutionized gene therapy, but their broad tropism and suboptimal transduction efficiency limit their clinical applications. To overcome these limitations, researchers have focused on designing and screening capsid libraries to identify improved vectors. However, the large sequence space and limited resources present challenges in identifying viable capsid variants. In this study, we propose an end-to-end diffusion model to generate capsid sequences with enhanced viability. Using publicly available AAV2 data, we generated 38,000 diverse AAV2 viral protein (VP) sequences, and evaluated 8,000 for viral selection. The results attested the superiority of our model compared to traditional methods. Additionally, in the absence of AAV9 capsid data, apart from one wild-type sequence, we used the same model to directly generate a number of viable sequences with up to 9 mutations. we transferred the remaining 30,000 samples to the AAV9 domain. Furthermore, we conducted mutagenesis on AAV9 VP hypervariable regions VI and V, contributing to the continuous improvement of the AAV9 VP sequence. This research represents a significant advancement in the design and functional validation of rAAV vectors, offering innovative solutions to enhance specificity and transduction efficiency in gene therapy applications.

cs.AI

Layer-by-layer connection for large area single crystal boron nitride multilayer films

Boron nitride (BN) is today considered as one of the most promising materials for many novel applications including bright single photon emission, deep UV opto-electronics, small sized solid-state neutron detector, and high-performance two-dimensional materials, etc. Despite the recent successful fabrication of large-area BN single-crystals (typically <= 5 atomic layers), the scalable growth of thicker single-crystalline BN films still constitutes a great challenge. In this work, we demonstrate an approach to grow large-area multilayer single-crystal BN films by chemical vapor deposition on face-centered cubic Fe-Ni (111) single crystal alloy thin films with different stoichiometric phases. We show that the BN growth is greatly tunable and improved by increasing the Fe content in single-crystal Fe-Ni (111). The formation of pyramid-shaped multilayer BN domains with aligned orientation enables a continuous connection following a layer-by-layer, 'first-meet-first-connect', mosaic stitching mechanism. By means of selected area electron diffraction, micro-photoluminescence spectroscopy in the deep UV and high-resolution transmission electron microscopy, the layer-by-layer connection mechanism is unambiguously evidenced, and the stacking order has been verified to occur as unidirectional AB and ABC stackings, i.e., in the Bernal and rhombohedral BN phase.

cond-mat.mtrl-sci

A High-Mass Young Star-forming Core Escaping from Its Parental Filament

We studied the unique kinematic properties in massive filament G352.63-1.07 at $10^3$-AU spatial scale with the dense molecular tracers observed with the Atacama Large Millimeter/submillimeter Array (ALMA). We find the central massive core M1 (12 $M_\odot$) being separated from the surrounding filament with a velocity difference of $v- {v}_{sys}=-2$ km/s and a transverse separation within 3 arcsec. Meanwhile, as shown in multiple dense-gas tracers, M1 has a spatial extension closely aligned with the main filament and is connected to the filament towards its both ends. M1 thus represents a very beginning state for a massive young star-forming core escaping from the parental filament, within a time scale of $\sim 4000$ years. Based on its kinetic energy ($3.5\times10^{44}$ erg), the core escape is unlikely solely due to the original filament motion or magnetic field, but requires more energetic events such as a rapid intense anisotropic collapse. The released energy also seems to noticeably increase the environmental turbulence. This may help the filament to become stabilized again.

astro-ph.GA

Batch-less stochastic gradient descent for compressive learning of deep regularization for image denoising

We consider the problem of denoising with the help of prior information taken from a database of clean signals or images. Denoising with variational methods is very efficient if a regularizer well adapted to the nature of the data is available. Thanks to the maximum a posteriori Bayesian framework, such regularizer can be systematically linked with the distribution of the data. With deep neural networks (DNN), complex distributions can be recovered from a large training database.To reduce the computational burden of this task, we adapt the compressive learning framework to the learning of regularizers parametrized by DNN. We propose two variants of stochastic gradient descent (SGD) for the recovery of deep regularization parameters from a heavily compressed database. These algorithms outperform the initially proposed method that was limited to low-dimensional signals, each iteration using information from the whole database. They also benefit from classical SGD convergence guarantees. Thanks to these improvements we show that this method can be applied for patch based image denoising.}

cs.LG

Contrastive Learning with Bidirectional Transformers for Sequential Recommendation

Contrastive learning with Transformer-based sequence encoder has gained predominance for sequential recommendation. It maximizes the agreements between paired sequence augmentations that share similar semantics. However, existing contrastive learning approaches in sequential recommendation mainly center upon left-to-right unidirectional Transformers as base encoders, which are suboptimal for sequential recommendation because user behaviors may not be a rigid left-to-right sequence. To tackle that, we propose a novel framework named \textbf{C}ontrastive learning with \textbf{Bi}directional \textbf{T}ransformers for sequential recommendation (\textbf{CBiT}). Specifically, we first apply the slide window technique for long user sequences in bidirectional Transformers, which allows for a more fine-grained division of user sequences. Then we combine the cloze task mask and the dropout mask to generate high-quality positive samples and perform multi-pair contrastive learning, which demonstrates better performance and adaptability compared with the normal one-pair contrastive learning. Moreover, we introduce a novel dynamic loss reweighting strategy to balance between the cloze task loss and the contrastive loss. Experiment results on three public benchmark datasets show that our model outperforms state-of-the-art models for sequential recommendation.

cs.IR

Everyone's Preference Changes Differently: Weighted Multi-Interest Retrieval Model

User embeddings (vectorized representations of a user) are essential in recommendation systems. Numerous approaches have been proposed to construct a representation for the user in order to find similar items for retrieval tasks, and they have been proven effective in industrial recommendation systems as well. Recently people have discovered the power of using multiple embeddings to represent a user, with the hope that each embedding represents the user's interest in a certain topic. With multi-interest representation, it's important to model the user's preference over the different topics and how the preference change with time. However, existing approaches either fail to estimate the user's affinity to each interest or unreasonably assume every interest of every user fades with an equal rate with time, thus hurting the recall of candidate retrieval. In this paper, we propose the Multi-Interest Preference (MIP) model, an approach that not only produces multi-interest for users by using the user's sequential engagement more effectively but also automatically learns a set of weights to represent the preference over each embedding so that the candidates can be retrieved from each interest proportionally. Extensive experiments have been done on various industrial-scale datasets to demonstrate the effectiveness of our approach.

cs.IR

Learning Bounded Context-Free-Grammar via LSTM and the Transformer:Difference and Explanations

Long Short-Term Memory (LSTM) and Transformers are two popular neural architectures used for natural language processing tasks. Theoretical results show that both are Turing-complete and can represent any context-free language (CFL).In practice, it is often observed that Transformer models have better representation power than LSTM. But the reason is barely understood. We study such practical differences between LSTM and Transformer and propose an explanation based on their latent space decomposition patterns. To achieve this goal, we introduce an oracle training paradigm, which forces the decomposition of the latent representation of LSTM and the Transformer and supervises with the transitions of the Pushdown Automaton (PDA) of the corresponding CFL. With the forced decomposition, we show that the performance upper bounds of LSTM and Transformer in learning CFL are close: both of them can simulate a stack and perform stack operation along with state transitions. However, the absence of forced decomposition leads to the failure of LSTM models to capture the stack and stack operations, while having a marginal impact on the Transformer model. Lastly, we connect the experiment on the prototypical PDA to a real-world parsing task to re-verify the conclusions

cs.CL

Convergent Filaments Contracting Towards an Intermediate-mass Prestellar Core

Filamentary structures are closely associated with star-forming cores, but their detailed physical connections are still not clear. We studied the dense gas in the region of OMC-3 MMS-7 in Orion A molecular cloud using the molecular lines observed with the Atacama Large Millimeter/submillimeter Array (ALMA) and the Submillimeter Array (SMA). The ALMA N$_2$H$^+$ (1-0) emission has revealed three dense filaments intersected at the center, coincident with the central core MMS-7, which has a mass of $3.6\,M_\odot$. The filaments and cores are embedded in a parental clump with total mass of $29\,M_\odot$. The N$_2$H$^+$ velocity field exhibits a noticeable increasing trend along the filaments towards the central core MMS-7 with a scale of $v-v_{\rm lsr} \simeq 1.5$ ${\rm km\, s^{-1}}$ over a spatial range of $\sim$20 arcsec ($8\times 10^3$ AU), corresponding to a gradient of $40\,{\rm km\, s^{-1}}\,{\rm pc}^{-1}$. This feature is most likely to indicate an infall motion towards the center. The derived infall rate ($8\times 10^{-5}\,M_\odot$ year$^{-1}$) and timescale ($3.6\times 10^5$ years) are much lower than that in a spherical free-fall collapse and more consistent with the contraction of filament structures. The filaments also exhibit a possible fragmentation, but it does not seem to largely interrupt the gas structure or the infall motion towards the center. MMS-7 thus provides an example of filamentary infall into an individual prestellar core. The filament contraction could be less intense but more steady than the global spherical collapse, and may help generate an intermediate- or even high-mass star.

astro-ph.GA

Hyperfine Group Ratio: A Recipe for Deriving Kinetic Temperature from Ammonia Inversion Lines

Although ammonia is a widely used interstellar thermometer, the estimation of its rotational and kinetic temperatures can be affected by the blended Hyperfine Components (HFCs). We developed a new recipe, referred to as the HyperFine Group Ratio (HFGR), which utilizes only direct observables, namely the intensity ratios between the grouped HFCs. As tested on the model spectra, the empirical formulae in HFGR can derive the rotational temperature ($T_{\rm rot}$) from the HFC group ratios in an unambiguous manner. We compared HFGR with two other classical methods, intensity ratio and hyperfine fitting, based on both simulated spectra and real data. HFGR has three major improvements. First, HFGR does not require modeling the HFC or fitting the line profiles, thus is more robust against the effect of HFC blending. Second, the simulation-enabled empirical formulae are much faster than fitting the spectra over the parameter space, so the computer time and human time can be both largely saved. Third, the statistical uncertainty of the temperature $\Delta T_{\rm rot}$ as a function of the signal-to-noise ratio (SNR) is a natural product of the HFGR recipe. The internal error of HFGR is $\Delta T_{\rm rot}\leq0.5$ K over a broad parameter space of rotational temperature (10 to 60 K), line width (0.3 to 4 km/s), and optical depth (0 to 5). When there is a spectral noise, HFGR can also maintain a reasonable uncertainty level at $\Delta T_{\rm rot}\leq 1.0$ K (1 $\sigma$) when SNR > 4.

astro-ph.GA

Filament Intersections and Cold Dense Cores in Orion A North

We studied the filament structures and dense cores in OMC-2,3 region in Orion A North molecular cloud using the high-resolution N2H+ (1-0) spectral cube observed with the Atacama Large Millimeter/Submillimeter Array (ALMA). The filament network over a total length of 2 pc is found to contain 170 intersections and 128 candidate dense cores. The dense cores are all displaced from the infrared point sources (possible young stars), and the major fraction of cores (103) are located around the intersections. Towards the intersections, there is also an increasing trend for the total column density Ntot as well as the the power-law index of the column-density Probability Distribution Function (N-PDF), suggesting that the intersections would in general have more significant gas assembly than the other part of the filament paths. The virial analysis shows that the dense cores mostly have virial mass ratio of alpha_vir=M_vir/M_gas<1.0, suggesting them to be bounded by the self gravity. In the mean time, only about 23 percent of the cores have critical mass ratio of alpha_crit=M_crit/M_gas<1.0, suggesting them to be unstable against core collapse. Combining these results, it shows that the major fraction of the cold starless and possible prestellar cores in OMC-2,3 are being assembled around the intersections, and currently in a gravitationally bound state. But more extensive core collapse and star formation may still require continuous core-mass growth or other perturbatio

astro-ph.GA

Sulfur-bearing molecules in Orion KL

We present an observational study of the sulfur (S)-bearing species towards Orion KL at 1.3 mm by combining ALMA and IRAM-30\,m single-dish data. At a linear resolution of $\sim$800 au and a velocity resolution of 1 $\mathrm{km\, s^{-1}\, }$, we have identified 79 molecular lines from 6 S-bearing species. In these S-bearing species, we found a clear dichotomy between carbon-sulfur compounds and carbon-free S-bearing species in various characteristics, e.g., line profiles, spatial morphology, and molecular abundances with respect to $\rm H_2$. Lines from the carbon-sulfur compounds (i.e., OCS, $^{13}$CS, H$_2$CS) exhibit spatial distributions concentrated around the continuum peaks and extended to the south ridge. The full width at half maximum (FWHM) linewidth of these molecular lines is in the range of 2 $\sim$ 11 $\mathrm{km\, s^{-1}\, }$. The molecular abundances of OCS and H$_2$CS decrease slightly from the cold ($\sim$68 K) to the hot ($\sim$176 K) regions. In contrast, lines from the carbon-free S-bearing species (i.e., SO$_2$, $^{34}$SO, H$_2$S) are spatially more extended to the northeast of mm4, exhibiting broader FWHM linewidths (15 $\sim$ 26 $\mathrm{km\, s^{-1}\, }$). The molecular abundances of carbon-free S-bearing species increase by over an order of magnitude as the temperature increase from 50 K to 100 K. In particular, $\mathrm{^{34}SO/^{34}SO_2}$ and $\mathrm{OCS/SO_2}$ are enhanced from the warmer regions ($>$100 K) to the colder regions ($\sim$50 K). Such enhancements are consistent with the transformation of SO$_2$ at warmer regions and the influence of shocks.

astro-ph.SR

Application of a two dipole model to PSR J1640-4631, a pulsar with an anomalous braking index

Recent timing observation provides an intriguing result for the braking index of the X-ray pulsar PSR J1640-4631, which has a measured braking index $n=3.15\pm0.03$. The decrease of the inclination angle between between the spin axis and the magnetic axis can be responsible for such a high braking index. However, the physical mechanisms causing the change of the magnetic inclination angle have not been fully understood. In this Letter, we apply a two-dipole model given by Hamil et al (2016) to explain the decrease of the magnetic inclination angle of PSR J1640-4631. The rotation effect of a charged sphere and the magnetization of ferromagnetically ordered material produce magnetic moments $M_{1}$ and $M_{2}$, respectively. There exist a minimum of the potential energy for the magnetic moment $M_{2}$ in the magnetic field of $M_{1}$, hence the $M_{2}$ will freely rotate around the minimum energy position (i. e equilibrium position), similar to a simple pendulum. Our calculation indicate the magnetic moment $M_{2}$ would evolve towards alignment with the spin axis for PSR J1640-4631, and cause the magnetic inclination angle to decrease. The single peak in the pulse profile favors a relatively low change rate of the magnetic inclination angle.

astro-ph.HE

Towards Safety-Aware Computing System Design in Autonomous Vehicles

Recently, autonomous driving development ignited competition among car makers and technical corporations. Low-level automation cars are already commercially available. But high automated vehicles where the vehicle drives by itself without human monitoring is still at infancy. Such autonomous vehicles (AVs) rely on the computing system in the car to to interpret the environment and make driving decisions. Therefore, computing system design is essential particularly in enhancing the attainment of driving safety. However, to our knowledge, no clear guideline exists so far regarding safety-aware AV computing system and architecture design. To understand the safety requirement of AV computing system, we performed a field study by running industrial Level-4 autonomous driving fleets in various locations, road conditions, and traffic patterns. The field study indicates that traditional computing system performance metrics, such as tail latency, average latency, maximum latency, and timeout, cannot fully satisfy the safety requirement for AV computing system design. To address this issue, we propose a `safety score' as a primary metric for measuring the level of safety in AV computing system design. Furthermore, we propose a perception latency model, which helps architects estimate the safety score of given architecture and system design without physically testing them in an AV. We demonstrate the use of our safety score and latency model, by developing and evaluating a safety-aware AV computing system computation hardware resource management scheme.

cs.RO