Searcharxiv⌕ Search

arXiv subjects

Bo Peng

Publications and source records attributed to Bo Peng.

At least 91 records · Page 5Linked to original sources

HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming

Content-aware streaming requires dynamic, chunk-level importance weights to optimize subjective quality of experience (QoE). However, direct human annotation is prohibitively expensive while vision-saliency models generalize poorly. We introduce HiVid, the first framework to leverage Large Language Models (LLMs) as a scalable human proxy to generate high-fidelity weights for both Video-on-Demand (VOD) and live streaming. We address 3 non-trivial challenges: (1) To extend LLMs' limited modality and circumvent token limits, we propose a perception module to assess frames in a local context window, autoregressively building a coherent understanding of the video. (2) For VOD with rating inconsistency across local windows, we propose a ranking module to perform global re-ranking with a novel LLM-guided merge-sort algorithm. (3) For live streaming which requires low-latency, online inference without future knowledge, we propose a prediction module to predict future weights with a multi-modal time series model, which comprises a content-aware attention and adaptive horizon to accommodate asynchronous LLM inference. Extensive experiments show HiVid improves weight prediction accuracy by up to 11.5\% for VOD and 26\% for live streaming over SOTA baselines. Real-world user study validates HiVid boosts streaming QoE correlation by 14.7\%.

cs.CV↗

Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era

Industrial Anomaly Detection (IAD) is a cornerstone for ensuring operational safety, maintaining product quality, and optimizing manufacturing efficiency. However, the advancement of IAD algorithms is severely hindered by the limitations of existing public benchmarks. Current datasets often suffer from restricted category diversity and insufficient scale, leading to performance saturation and poor model transferability in complex, real-world scenarios. To bridge this gap, we introduce Real-IAD Variety, the largest and most diverse IAD benchmark. It comprises 198,950 high-resolution images across 160 distinct object categories. The dataset ensures unprecedented diversity by covering 28 industries, 24 material types, 22 color variations, and 27 defect types. Our extensive experimental analysis highlights the substantial challenges posed by this benchmark: state-of-the-art multi-class unsupervised anomaly detection methods suffer significant performance degradation (ranging from 10% to 20%) when scaled from 30 to 160 categories. Conversely, we demonstrate that zero-shot and few-shot IAD models exhibit remarkable robustness to category scale-up, maintaining consistent performance and significantly enhancing generalization across diverse industrial contexts. This unprecedented scale positions Real-IAD Variety as an essential resource for training and evaluating next-generation foundation IAD models.

cs.CV↗

Searching for Optimal Prices in Two-Sided Markets

We investigate online pricing in two-sided markets where a platform repeatedly posts prices based on binary accept/reject feedback to maximize gains-from-trade (GFT) or profit. We characterize the regret achievable across three mechanism classes: Single-Price, Two-Price, and Segmented-Price. For profit maximization, we design an algorithm using Two-Price Mechanisms that achieves $O(n^2 \log\log T)$ regret, where $n$ is the number of traders. For GFT maximization, the optimal regret depends critically on both market size and mechanism expressiveness. Constant regret is achievable in bilateral trade, but this guarantee breaks down as the market grows: even in a one-seller, two-buyer market, any algorithm using Single-Price Mechanisms suffers regret at least $Ω\!\big(\frac{\log\log T}{\log\log\log\log T}\big)$, and we provide a nearly matching $O(\log\log T)$ upper bound for general one-to-many markets. In full many-to-many markets, we prove that Two-Price Mechanisms inevitably incur linear regret $Ω(T)$ due to a \emph{mismatch phenomenon}, wherein inefficient pairings prevent near-optimal trade. To overcome this barrier, we introduce \emph{Segmented-Price Mechanisms}, which partition traders into groups and assign distinct prices per group. Using this richer mechanism, we design an algorithm achieving $O(n^2 \log\log T + n^3)$ regret for GFT maximization. Finally, we extend our results to the contextual setting, where traders' costs and values depend linearly on observed $d$-dimensional features that vary across rounds, obtaining regret bounds of $O(n^2 d \log\log T + n^2 d \log d)$ for profit and $O(n^2 d^2 \log T)$ for GFT. Our work delineates sharp boundaries between learnable and unlearnable regimes in two-sided dynamic pricing and demonstrates how modest increases in pricing expressiveness can circumvent fundamental hardness barriers.

cs.GT↗

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

Recent attempts to transfer features from 2D Vision-Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields noisy and fragmented predictions, whereas enforcing geometric coherence necessitates costly training pipelines and large-scale annotated 3D data. We argue that this limitation stems from the dominant segmentation-and-matching paradigm, which fails to reconcile 2D semantics with 3D geometric structure. The geometric cues are not eliminated during the 2D-to-3D transfer but remain latent within the noisy and view-aggregated features. To exploit this property, we propose GeoPurify that applies a small Student Affinity Network to purify 2D VLM-generated 3D point features using geometric priors distilled from a 3D self-supervised teacher model. During inference, we devise a Geometry-Guided Pooling module to further denoise the point cloud and ensure the semantic and structural consistency. Benefiting from latent geometric information and the learned affinity network, GeoPurify effectively mitigates the trade-off and achieves superior data efficiency. Extensive experiments on major 3D benchmarks demonstrate that GeoPurify achieves or surpasses state-of-the-art performance while utilizing only about 1.5% of the training data.

cs.CV↗

DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing

Text-guided image editing faces significant challenges when considering training and inference flexibility. Much literature collects large amounts of annotated image-text pairs to train text-conditioned generative models from scratch, which is expensive and not efficient. After that, some approaches that leverage pre-trained vision-language models have been proposed to avoid data collection, but they are limited by either per text-prompt optimization or inference-time hyper-parameters tuning. To address these issues, we investigate and identify a specific space, referred to as CLIP DeltaSpace, where the CLIP visual feature difference of two images is semantically aligned with the CLIP textual feature difference of their corresponding text descriptions. Based on DeltaSpace, we propose a novel framework called DeltaEdit, which maps the CLIP visual feature differences to the latent space directions of a generative model during the training phase, and predicts the latent space directions from the CLIP textual feature differences during the inference phase. And this design endows DeltaEdit with two advantages: (1) text-free training; (2) generalization to various text prompts for zero-shot inference. Extensive experiments validate the effectiveness and versatility of DeltaEdit with different generative models, including both the GAN model and the diffusion model, in achieving flexible text-guided image editing. Code is available at https://github.com/Yueming6568/DeltaEdit.

cs.CV↗

GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization

The placement of normalization layers, specifically Pre-Norm and Post-Norm, remains an open question in Transformer architecture design. In this work, we rethink these approaches through the lens of manifold optimization, interpreting the outputs of the Feed-Forward Network (FFN) and attention layers as update directions in optimization. Building on this perspective, we introduce GeoNorm, a novel method that replaces standard normalization with geodesic updates on the manifold. Furthermore, analogous to learning rate schedules, we propose a layer-wise update decay for the FFN and attention components. Comprehensive experiments demonstrate that GeoNorm consistently outperforms existing normalization methods in Transformer models. Crucially, GeoNorm can be seamlessly integrated into standard Transformer architectures, achieving performance improvements with negligible additional computational cost.

cs.LG↗

Endogenous Reprompting: Self-Evolving Cognitive Alignment for Unified Multimodal Models

Unified Multimodal Models (UMMs) exhibit strong understanding, yet this capability often fails to effectively guide generation. We identify this as a Cognitive Gap: the model lacks the understanding of how to enhance its own generation process. To bridge this gap, we propose Endogenous Reprompting, a mechanism that transforms the model's understanding from a passive encoding process into an explicit generative reasoning step by generating self-aligned descriptors during generation. To achieve this, we introduce SEER (Self-Evolving Evaluator and Reprompter), a training framework that establishes a two-stage endogenous loop using only 300 samples from a compact proxy task, Visual Instruction Elaboration. First, Reinforcement Learning with Verifiable Rewards (RLVR) activates the model's latent evaluation ability via curriculum learning, producing a high-fidelity endogenous reward signal. Second, Reinforcement Learning with Model-rewarded Thinking (RLMT) leverages this signal to optimize the generative reasoning policy. Experiments show that SEER consistently outperforms state-of-the-art baselines in evaluation accuracy, reprompting efficiency, and generation quality, without sacrificing general multimodal capabilities.

cs.AI↗

Voltage-controlled topological spin textures in the monolayer limit

The physics of phase transitions in low-dimensional systems has long been a subject of significant research interest. Long-range magnetic order in the strict two-dimensional limit, whose discovery circumvented the Mermin-Wagner theorem, has rapidly emerged as a research focus. However, the demonstration of a non-trivial topological spin textures in two-dimensional limit has remained elusive. Here, we demonstrate the out-of-plane electric field breaks inversion symmetry while simultaneously modulating the electronic band structure, enabling electrically tunable spin-orbit interaction for creation and manipulation of topological spin textures in monolayer CrI3. The realization of ideal two-dimensional topological spin textures may offer not only an experimental testbed for probing the Berezinskii-Kosterlitz-Thouless mechanism, but also potential insights into unresolved quantum phenomena including superconductivity and superfluidity. Moreover, voltage-controlled spin-orbit interaction offers a novel pathway to engineer two-dimensional spin textures with tailored symmetries and topologies, while opening avenues for skyrmion-based next-generation information technologies.

cond-mat.mes-hall↗

CauScientist: Teaching LLMs to Respect Data for Causal Discovery

Causal discovery is fundamental to scientific understanding and reliable decision-making. Existing approaches face critical limitations: purely data-driven methods suffer from statistical indistinguishability and modeling assumptions, while recent LLM-based methods either ignore statistical evidence or incorporate unverified priors that can mislead result. To this end, we propose CauScientist, a collaborative framework that synergizes LLMs as hypothesis-generating "data scientists" with probabilistic statistics as rigorous "verifiers". CauScientist employs hybrid initialization to select superior starting graphs, iteratively refines structures through LLM-proposed modifications validated by statistical criteria, and maintains error memory to guide efficient search space. Experiments demonstrate that CauScientist substantially outperforms purely data-driven baselines, achieving up to 53.8% F1 score improvement and enhancing recall from 35.0% to 100.0%. Notably, while standalone LLM performance degrades with graph complexity, CauScientist reduces structural hamming distance (SHD) by 44.0% compared to Qwen3-32B on 37-node graphs. Our project page is at https://github.com/OpenCausaLab/CauScientist.

cs.CL↗

Chemically decisive benchmarks on the path to quantum utility

Progress towards quantum utility in chemistry requires not only algorithmic advances, but also the identification of chemically meaningful problems whose electronic structure fundamentally challenges classical methods. Here, we introduce a curated hierarchy of chemically decisive benchmark systems designed to probe distinct regimes of electronic correlation relevant to molecular, bioinorganic, and heavy-element chemistry. Moving beyond minimal toy models, our benchmark set spans multireference bond breaking (N$_2$), high-spin transition-metal chemistry (FeS), biologically relevant iron-sulfur clusters ([2Fe-2S]), and actinide-actinide bonding (U$_2$), which exhibits extreme sensitivity to active-space choice, relativistic treatment, and correlation hierarchy even within advanced multireference frameworks. As a concrete realization, we benchmark a recently developed automated and adaptive quantum algorithm based on generator-coordinate-inspired subspace expansion,ADAPT-GCIM, using a black-box workflow that integrates entropy-based active-space selection via the ActiveSpaceFinder tool. Across this chemically diverse problem set, ADAPT-GCIM achieves high accuracy in challenging correlation regimes. Equally importantly, these benchmarks expose general failure modes and design constraints-independent of any specific algorithm-highlighting the necessity of problem-aware and correlation-specific strategies for treating strongly correlated chemistry on quantum computers. To support systematic benchmarking and reproducible comparisons, the Hamiltonians for all systems studied are made openly available.

physics.chem-ph↗

Magneto-optical-electric joint-measurement scanning imaging system for identification of two-dimensional vdW multiferroic

As an advanced imaging system, the magneto-optical-electric joint-measurement scanning imaging system (MOEJSI) brings spectroscopic techniques with unmatched spatial resolution to very low temperature, high magnetic field and high electric field measurements. It was developed for investigating the magnetic and ferroelectric properties and their mutual control through magneto-optical-electric joint-measurements, besides Raman and photoluminescence features. In particular, the reflective magnetic circular dichroism (RMCD) loops and imaging, linear dichroism (LD) imaging and polarization-electric field hysteresis loop can be achieved when simultaneously applied high magnetic field (7 T) and electric field (100 V) at low temperature of 10 K.

physics.app-ph↗

A Fractional Calculus Framework for Open Quantum Dynamics: From Liouville to Lindblad to Memory Kernels

Open quantum systems exhibit dynamics ranging from unitary evolution to irreversible dissipation. While the Gorini--Kossakowski--Sudarshan--Lindblad (GKSL) equation uniquely characterizes Markovian CPTP evolution, many physical platforms display non-Markovian features such as algebraic relaxation and coherence backflow. Fractional calculus provides a natural way to model such long-memory behavior through power-law temporal kernels introduced by fractional time derivatives. Here we develop a unified framework that embeds fractional master equations within the broader hierarchy of open-system formalisms. The fractional equation forms a structured subclass of memory-kernel models, reduces to the Lindblad form at unit order, and, through Bochner--Phillips subordination, admits a CPTP representation as an average over Lindblad semigroups. Its resolvent structure further connects fractional dynamics to established non-Markovian approaches, including Nakajima--Zwanzig kernels and hierarchical equations of motion, providing a compact surrogate for long-memory effects. This formulation positions fractional calculus as a rigorous and practical language for quantum dynamics with intrinsic memory, supporting both analytical insight and efficient quantum simulation.

quant-ph↗

A Thick Volatile Atmosphere on the Ultrahot Super-Earth TOI-561 b

Ultrashort-period (USP) exoplanets -- with $R_p \leq 2~$R$_{\oplus}$ and periods $\leq$1 day -- are expected to be stripped of volatile atmospheres by intense host star irradiation, which is corroborated by their nominal bulk densities and previous eclipse observations consistent with bare rock surfaces. However, a few USP planets appear anomalously under-dense relative to an Earth-like composition, suggesting an exotic interior structure (e.g., core-less) or a volatile-rich secondary atmosphere increasing their apparent radius. Here we present the first dayside emission spectrum of the low-density (4.3$\pm$0.4 g~cm$^{-3}$) USP planet TOI-561 b, which orbits an iron-poor, alpha-rich, $\sim$10 Gyr old thick disk star. Our 3-5 $μ$m JWST/NIRSpec observations demonstrate the dayside of TOI-561 b is inconsistent with a bare-rock surface at high statistical significance, suggesting instead a thick volatile envelope that is cooling the dayside to well below the $\sim$3000 K expected in the bare-rock or thin-atmosphere case. These results reject the popular hypothesis of complete atmospheric desiccation for highly irradiated exoplanets and support predictions that planetary-scale magma oceans can retain substantial reservoirs of volatiles, opening the geophysical study of ultrahot super-Earths through the lenses of their atmospheres.

astro-ph.EP↗

CauSight: Learning to Supersense for Visual Causal Discovery

Causal thinking enables humans to understand not just what is seen, but why it happens. To replicate this capability in modern AI systems, we introduce the task of visual causal discovery. It requires models to infer cause-and-effect relations among visual entities across diverse scenarios instead of merely perceiving their presence. To this end, we first construct the Visual Causal Graph dataset (VCG-32K), a large-scale collection of over 32,000 images annotated with entity-level causal graphs, and further develop CauSight, a novel vision-language model to perform visual causal discovery through causally aware reasoning. Our training recipe integrates three components: (1) training data curation from VCG-32K, (2) Tree-of-Causal-Thought (ToCT) for synthesizing reasoning trajectories, and (3) reinforcement learning with a designed causal reward to refine the reasoning policy. Experiments show that CauSight outperforms GPT-4.1 on visual causal discovery, achieving over a threefold performance boost (21% absolute gain). Our code, model, and dataset are fully open-sourced at project page: https://github.com/OpenCausaLab/CauSight.

cs.CV↗

The diverse morphology of gravitational wave signals from merging neutron-star white-dwarf binaries

In sufficiently compact neutron star-white dwarf (NSWD) binary systems, orbital decay means the white dwarf eventually fills its shrinking Roche lobe, initiating a phase of mass transfer. The exchange of angular momentum-both internal and external-plays a critical role in determining the binary's evolutionary outcome. For neutron stars with relatively low magnetic fields and spin frequencies, whether the orbital separation continues to shrink depends on the interplay between gravitational wave (GW) radiation and mass transfer dynamics. We compute the orbital evolution of NSWD binaries across a broad parameter space, incorporating four key variables. Our results reveal distinct boundaries in the NS-WD mass-mass diagram: binaries with white dwarf masses above these thresholds undergo rapid orbital decay and direct coalescence. The dependence of these boundaries on system parameters indicates that Roche-lobe-filling NSWD binaries can follow multiple evolutionary pathways -- a phenomenon we refer to as branched or polymorphic evolution. NSWD binary systems emit strong and diverse GW signals, many of which would be detectable by space-based GW observatories. The morphology of the evolving GW waveform provides a direct diagnostic for the NSWD binary configuration, including any contribution from an accretion disk. Our models can provide critical waveform templates for identifying merging binary signals in real-time GW data.

astro-ph.HE↗

HunyuanVideo 1.5 Technical Report

We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference on consumer-grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architecture featuring selective and sliding tile attention (SSTA), enhanced bilingual understanding through glyph-aware text encoding, progressive pre-training and post-training, and an efficient video super-resolution network. Leveraging these designs, we developed a unified framework capable of high-quality text-to-video and image-to-video generation across multiple durations and resolutions. Extensive experiments demonstrate that this compact and proficient model establishes a new state-of-the-art among open-source video generation models. By releasing the code and model weights, we provide the community with a high-performance foundation that lowers the barrier to video creation and research, making advanced video generation accessible to a broader audience. All open-source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5.

cs.CV↗