Searcharxiv⌕ Search

arXiv subjects

Bo Peng

Publications and source records attributed to Bo Peng.

At least 163 records · Page 9Linked to original sources

FAST Ultra-Deep Survey (FUDS): the star formation histories of FUDS0 galaxies

We present the ultraviolet, optical and infrared counterparts of 128 galaxies detected in neutral hydrogen (HI) in the FAST Ultra-Deep Survey (FUDS) field 0 (FUDS0). HI mass upper limits are also calculated for 134 non-detections in the field. Stellar masses ($M_*$), star formation rates (SFRs) and star formation histories are computed by fitting spectral energy distributions (SEDs) using ProSpect. The results show that HI-selected galaxies prefer recent long-lasting, but mild star formation activity, while HI non-detections have earlier and more intense star formation activity. Based on their distribution on the SFR versus $M_*$ diagram, the typical evolution of HI-selected galaxies follows three distinct stages: (i) Early stage: the total SFR increases, though the specific SFR (sSFR) decreases from 10$^{-8}$ to 10$^{-9}$ yr$^{-1}$; (ii) Mass accumulation stage: the SFR is steady, and stellar mass increase linearly with time; (iii) Quenching stage: star formation activity quenches on a rapid timescale and at constant stellar mass. 37 non-detections are located on star-forming main sequence, but are not detected in HI due to low sensitivity close to field edges or close to strong radio frequency interference. Comparisons with the existing optical, optically-selected HI, and HI catalogs show a good agreement with respect to measured $M_*$ and SFR, with minor discrepancies due to selection effects. The ongoing full FUDS survey will help us better explore the evolutionary stages of HI galaxies through a larger sample.

astro-ph.GA↗

log-RRIM: Yield Prediction via Local-to-global Reaction Representation Learning and Interaction Modeling

Accurate prediction of chemical reaction yields is crucial for optimizing organic synthesis, potentially reducing time and resources spent on experimentation. With the rise of artificial intelligence (AI), there is growing interest in leveraging AI-based methods to accelerate yield predictions without conducting in vitro experiments. We present log-RRIM, an innovative graph transformer-based framework designed for predicting chemical reaction yields. A key feature of log-RRIM is its integration of a cross-attention mechanism that focuses on the interplay between reagents and reaction centers. This design reflects a fundamental principle in chemical reactions: the crucial role of reagents in influencing bond-breaking and formation processes, which ultimately affect reaction yields. log-RRIM also implements a local-to-global reaction representation learning strategy. This approach initially captures detailed molecule-level information and then models and aggregates intermolecular interactions. Through this hierarchical process, log-RRIM effectively captures how different molecular fragments contribute to and influence the overall reaction yield, regardless of their size variations. log-RRIM shows superior performance in our experiments, especially for medium to high-yielding reactions, proving its reliability as a predictor. The framework's sophisticated modeling of reactant-reagent interactions and precise capture of molecular fragment contributions make it a valuable tool for reaction planning and optimization in chemical synthesis. The data and codes of log-RRIM are accessible through https://github.com/ninglab/Yield_log_RRIM.

q-bio.BM↗

Revisiting Ranking for Online Bipartite Matching with Random Arrivals: the Primal-Dual Analysis

We revisit the celebrated Ranking algorithm by Karp, Vazirani, and Vazirani (STOC 1990) for online bipartite matching under the random arrival model, that is shown to be $0.696$-competitive for unweighted graphs by Mahdian and Yan (STOC 2011) and $0.662$-competitive for vertex-weighted graphs by Jin and Williamson (WINE 2021). In this work, we explore the limitation of the primal-dual analysis of Ranking and aim to bridge the gap between unweighted and vertex-weighted graphs. We show that the competitive ratio of Ranking is between $0.686$ and $0.703$, under our current knowledge of Ranking and the framework of primal-dual analysis. This confirms a conjecture by Huang, Tang, Wu, and Zhang (TALG 2019), stating that the primal-dual analysis could lead to a competitive ratio that is very close to $0.696$. Our analysis involves proper discretizations of a variational problem and uses LP solver to pin down the numerical number. As a bonus of our discretization approach, our competitive analysis of Ranking applies to a more relaxed random arrival model. E.g., we show that even when each online vertex arrives independently at an early or late stage, the Ranking algorithm is at least $0.665$-competitive, beating the $1-1/e \approx 0.632$ competitive ratio under the adversarial arrival model.

cs.DS↗

Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models

In order to deeply understand the capability of pretrained language models in text generation and conduct a diagnostic evaluation, we propose TGEA, an error-annotated dataset with multiple benchmark tasks for text generation from pretrained language models (PLMs). We use carefully selected prompt words to guide GPT-2 to generate candidate sentences, from which we select 47K for error annotation. Crowdsourced workers manually check each of these sentences and detect 12k erroneous sentences. We create an error taxonomy to cover 24 types of errors occurring in these erroneous sentences according to the nature of errors with respect to linguistics and knowledge (eg, common sense). For each erroneous span in PLM-generated sentences, we also detect another span that is closely associated with it. Each error is hence manually labeled with comprehensive annotations, including the span of the error, the associated span, minimal correction to the error, the type of the error, and rationale behind the error. Apart from the fully annotated dataset, we also present a detailed description of the data collection procedure, statistics and analysis of the dataset. This is the first dataset with comprehensive annotations for PLM-generated texts, which facilitates the diagnostic evaluation of PLM-based text generation. Furthermore, we use TGEA as a benchmark dataset and propose a series of automatic diagnosis tasks, including error detection, error type classification, associated span detection, error rationale generation, to further promote future study on the automatic error detection and correction on texts generated by pretrained language models.

cs.CL↗

PA-CLIP: Enhancing Zero-Shot Anomaly Detection through Pseudo-Anomaly Awareness

In industrial anomaly detection (IAD), accurately identifying defects amidst diverse anomalies and under varying imaging conditions remains a significant challenge. Traditional approaches often struggle with high false-positive rates, frequently misclassifying normal shadows and surface deformations as defects, an issue that becomes particularly pronounced in products with complex and intricate surface features. To address these challenges, we introduce PA-CLIP, a zero-shot anomaly detection method that reduces background noise and enhances defect detection through a pseudo-anomaly-based framework. The proposed method integrates a multiscale feature aggregation strategy for capturing detailed global and local information, two memory banks for distinguishing background information, including normal patterns and pseudo-anomalies, from true anomaly features, and a decision-making module designed to minimize false positives caused by environmental variations while maintaining high defect sensitivity. Demonstrated on the MVTec AD and VisA datasets, PA-CLIP outperforms existing zero-shot methods, providing a robust solution for industrial defect detection.

cs.CV↗

Quantum time dynamics mediated by the Yang-Baxter equation and artificial neural networks

Quantum computing shows great potential, but errors pose a significant challenge. This study explores new strategies for mitigating quantum errors using artificial neural networks (ANN) and the Yang-Baxter equation (YBE). Unlike traditional error mitigation methods, which are computationally intensive, we investigate artificial error mitigation. We developed a novel method that combines ANN for noise mitigation combined with the YBE to generate noisy data. This approach effectively reduces noise in quantum simulations, enhancing the accuracy of the results. The YBE rigorously preserves quantum correlations and symmetries in spin chain simulations in certain classes of integrable lattice models, enabling effective compression of quantum circuits while retaining linear scalability with the number of qubits. This compression facilitates both full and partial implementations, allowing the generation of noisy quantum data on hardware alongside noiseless simulations using classical platforms. By introducing controlled noise through the YBE, we enhance the dataset for error mitigation. We train an ANN model on partial data from quantum simulations, demonstrating its effectiveness in mitigating errors in time-evolving quantum states, providing a scalable framework to enhance quantum computation fidelity, particularly in noisy intermediate-scale quantum (NISQ) systems. We demonstrate the efficacy of this approach by performing quantum time dynamics simulations using the Heisenberg XY Hamiltonian on real quantum devices.

quant-ph↗

Quantum Simulation of Boson-Related Hamiltonians: Techniques, Effective Hamiltonian Construction, and Error Analysis

A broad spectrum of physical systems in condensed-matter and high-energy physics, vibrational spectroscopy, and circuit and cavity QED necessitates the incorporation of bosonic degrees of freedom, such as phonons, photons, and gluons, into optimized fermion algorithms for near-future quantum simulations. In particular, when a quantum system is surrounded by an external environment, its basic physics can usually be simplified to a spin or fermionic system interacting with bosonic modes. Nevertheless, troublesome factors such as the magnitude of the bosonic degrees of freedom typically complicate the direct quantum simulation of these interacting models, necessitating the consideration of a comprehensive plan. This strategy should specifically include a suitable fermion/boson-to-qubit mapping scheme to encode sufficiently large yet manageable bosonic modes, and a method for truncating and/or downfolding the Hamiltonian to the defined subspace for performing an approximate but highly accurate simulation, guided by rigorous error analysis. In this pedagogical tutorial review, we aim to provide such an exhaustive strategy, focusing on encoding and simulating certain bosonic-related model Hamiltonians, inclusive of their static properties and time evolutions. Specifically, we emphasize two aspects: (1) the discussion of recently developed quantum algorithms for these interacting models and the construction of effective Hamiltonians, and (2) a detailed analysis regarding a tightened error bound for truncating the bosonic modes for a class of fermion-boson interacting Hamiltonians.

quant-ph↗

SWaT: Statistical Modeling of Video Watch Time through User Behavior Analysis

The significance of estimating video watch time has been highlighted by the rising importance of (short) video recommendation, which has become a core product of mainstream social media platforms. Modeling video watch time, however, has been challenged by the complexity of user-video interaction, such as different user behavior modes in watching the recommended videos and varying watching probability over the video progress bar. Despite the importance and challenges, existing literature on modeling video watch time mostly focuses on relatively black-box mechanical enhancement of the classical regression/classification losses, without factoring in user behavior in a principled manner. In this paper, we for the first time take on a user-centric perspective to model video watch time, from which we propose a white-box statistical framework that directly translates various user behavior assumptions in watching (short) videos into statistical watch time models. These behavior assumptions are portrayed by our domain knowledge on users' behavior modes in video watching. We further employ bucketization to cope with user's non-stationary watching probability over the video progress bar, which additionally helps to respect the constraint of video length and facilitate the practical compatibility between the continuous regression event of watch time and other binary classification events. We test our models extensively on two public datasets, a large-scale offline industrial dataset, and an online A/B test on a short video platform with hundreds of millions of daily-active users. On all experiments, our models perform competitively against strong relevant baselines, demonstrating the efficacy of our user-centric perspective and proposed framework.

cs.IR↗

Stereo Image Coding for Machines with Joint Visual Feature Compression

2D image coding for machines (ICM) has achieved great success in coding efficiency, while less effort has been devoted to stereo image fields. To promote the efficiency of stereo image compression (SIC) and intelligent analysis, the stereo image coding for machines (SICM) is formulated and explored in this paper. More specifically, a machine vision-oriented stereo feature compression network (MVSFC-Net) is proposed for SICM, where the stereo visual features are effectively extracted, compressed, and transmitted for 3D visual task. To efficiently compress stereo visual features in MVSFC-Net, a stereo multi-scale feature compression (SMFC) module is designed to gradually transform sparse stereo multi-scale features into compact joint visual representations by removing spatial, inter-view, and cross-scale redundancies simultaneously. Experimental results show that the proposed MVSFC-Net obtains superior compression efficiency as well as 3D visual task performance, when compared with the existing ICM anchors recommended by MPEG and the state-of-the-art SIC method.

cs.CV↗

Generating 3D Binding Molecules Using Shape-Conditioned Diffusion Models with Guidance

Drug development is a critical but notoriously resource- and time-consuming process. In this manuscript, we develop a novel generative artificial intelligence (genAI) method DiffSMol to facilitate drug development. DiffSmol generates 3D binding molecules based on the shapes of known ligands. DiffSMol encapsulates geometric details of ligand shapes within pre-trained, expressive shape embeddings and then generates new binding molecules through a diffusion model. DiffSMol further modifies the generated 3D structures iteratively via shape guidance to better resemble the ligand shapes. It also tailors the generated molecules toward optimal binding affinities under the guidance of protein pockets. Here, we show that DiffSMol outperforms the state-of-the-art methods on benchmark datasets. When generating binding molecules resembling ligand shapes, DiffSMol with shape guidance achieves a success rate 61.4%, substantially outperforming the best baseline (11.2%), meanwhile producing molecules with novel molecular graph structures. DiffSMol with pocket guidance also outperforms the best baseline in binding affinities by 13.2%, and even by 17.7% when combined with shape guidance. Case studies for two critical drug targets demonstrate very favorable physicochemical and pharmacokinetic properties of the generated molecules, thus, the potential of DiffSMol in developing promising drug candidates.

cs.LG↗

An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation

Embodied manipulation is a fundamental ability in the realm of embodied artificial intelligence. Although current embodied manipulation models show certain generalizations in specific settings, they struggle in new environments and tasks due to the complexity and diversity of real-world scenarios. The traditional end-to-end data collection and training manner leads to significant data demands. Decomposing end-to-end tasks into atomic skills helps reduce data requirements and improves the task success rate. However, existing methods are limited by predefined skill sets that cannot be dynamically updated. To address the issue, we introduce a three-wheeled data-driven method to build an atomic skill library. We divide tasks into subtasks using the Vision-Language-Planning (VLP). Then, atomic skill definitions are formed by abstracting the subtasks. Finally, an atomic skill library is constructed via data collection and Vision-Language-Action (VLA) fine-tuning. As the atomic skill library expands dynamically with the three-wheel update strategy, the range of tasks it can cover grows naturally. In this way, our method shifts focus from end-to-end tasks to atomic skills, significantly reducing data costs while maintaining high performance and enabling efficient adaptation to new tasks. Extensive experiments in real-world settings demonstrate the effectiveness and efficiency of our approach.

cs.RO↗

Direct high-resolution observation of feedback and chemical enrichment in the circumgalactic medium at redshift z ~ 2.8

The circumgalactic medium (CGM) plays a vital role in galaxy evolution, however, studying the emission from CGM is challenging due to its low surface brightness and the complexities involved in interpreting resonant lines such as Ly$α$. The near-infrared coverage, unprecedented sensitivity, and high spatial resolution of JWST enable us to study the optical strong lines associated with the extended Ly$α$ "nebulae" at redshifts of 2--3. These lines serve as diagnostic tools to infer the physical conditions in the CGM gas reservoir of these systems. In deep medium-band images taken by the JWST, we serendipitously discovered the [O III] emission from the CGM around a massive interacting galaxy system at a redshift z~2.8, known to be embedded in a bright extended (100 kpc) Ly$α$ "nebula." This is the first time that the [O III] lines have been detected from a Ly$α$ "nebula." The JWST images reveal that the CGM gas actually resides in narrow (~ 2.5 kpc) filamentary structures with strong [O III] emission, tracing the same extent as the Ly$α$ emission. An analysis of the [O III] suggests that the emitting CGM is fully ionized and is energetically dominated by mechanical heating. We also find that the density and pressure are higher than those commonly predicted by simulations of the CGM. We conclude that the observed CGM emission originates from the gas expelled by the episodic feedback processes, cooling down and enriching the CGM, while traveling a distance of at least 60 kpc. These observations demonstrate how intensive feedback processes shape gas distribution and properties in the CGM around massive halos. While access to such deep, high-resolution imaging opens up a new discovery space for investigating the CGM, it also challenges numerical simulations with respect to explaining and reproducing the exquisitely complex structures revealed by the observations.

astro-ph.GA↗

Bowen's Problem 32 and the conjugacy problem for systems with specification

We show that Rufus Bowen's Problem 32 on the classification of symbolic systems with the specification property does not admit a solution that would use concrete invariants. To this end, we construct a class of symbolic systems with the specification property and show that the conjugacy relation on this class is too complicated to admit such a classification. More generally, we gauge the complexity of the classification problem for symbolic systems with the specification property. Along the way, we also provide answers to two questions related to the classification of pointed systems with the specification property: to a question of Ding and Gu related to the complexity of the classification of pointed Cantor systems with the specification property and to a question of Bruin and Vejnar related to the complexity of the classification of pointed Hilbert cube systems with the specification property.

math.LO↗

D$^3$-Human: Dynamic Disentangled Digital Human from Monocular Video

We introduce D$^3$-Human, a method for reconstructing Dynamic Disentangled Digital Human geometry from monocular videos. Past monocular video human reconstruction primarily focuses on reconstructing undecoupled clothed human bodies or only reconstructing clothing, making it difficult to apply directly in applications such as animation production. The challenge in reconstructing decoupled clothing and body lies in the occlusion caused by clothing over the body. To this end, the details of the visible area and the plausibility of the invisible area must be ensured during the reconstruction process. Our proposed method combines explicit and implicit representations to model the decoupled clothed human body, leveraging the robustness of explicit representations and the flexibility of implicit representations. Specifically, we reconstruct the visible region as SDF and propose a novel human manifold signed distance field (hmSDF) to segment the visible clothing and visible body, and then merge the visible and invisible body. Extensive experimental results demonstrate that, compared with existing reconstruction schemes, D$^3$-Human can achieve high-quality decoupled reconstruction of the human body wearing different clothing, and can be directly applied to clothing transfer and animation.

cs.CV↗

C$_{60}$ building blocks with tuneable structures for tailored functionalities

We show that C$_{60}$ fullerene molecules can serve as promising building blocks in the construction of versatile crystal structures with unique symmetries using first-principles calculations. These phases include quasi-2D layered structures and 3D van der Waals crystals where the molecules adopt varied orientations. The interplay of molecular arrangement and lattice symmetry results in a variety of tuneable crystal structures with distinct properties. Specifically, the electronic structures of these phases vary significantly, offering potential for fine-tuning the band gap for electronics and optoelectronics. Additionally, the optical properties of these materials are strongly influenced by their crystalline symmetry and molecular alignment, providing avenues for tailoring optical responses for photonics. Our findings highlight the potential of fullerene-based building blocks in the rational design of functional materials.

cond-mat.mes-hall↗

Tuning electronic and optical properties of 2D polymeric C$_{60}$ by stacking two layers

Benefiting from improved stability due to stronger interlayer van der Waals interactions, few-layer fullerene networks are experimentally more accessible compared to monolayer polymeric C$_{60}$. However, there is a lack of systematic theoretical studies on the material properties of few-layer C$_{60}$ networks. Here, we compare the structural, electronic and optical properties of bilayer and monolayer fullerene networks. The band gap and band-edge positions remain mostly unchanged after stacking two layers into a bilayer, enabling the bilayer to be almost as efficient a photocatalyst as the monolayer. The effective mass ratio along different directions is varied for conduction band states due to interlayer interactions,leading to enhanced anisotropy in carrier transport. Additionally, stronger exciton absorption is found in the bilayer than that in the monolayer over the entire visible light range, rendering the bilayer a more promising candidate for photovoltaics. Moreoever, the polarisation dependence of optical absorption in the bilayer is increased in the red-yellow light range, offering unique opportunities in photonics and display technologies with tailored optical properties over specific directions. Our study provides strategies to tune electronic and optical properties of 2D polymeric C$_{60}$ via the introduction of stacking degrees of freedom.

cond-mat.mtrl-sci↗