SearcharxivSearch

arXiv subjects

Bowen Fu

Publications and source records attributed to Bowen Fu.

At least 19 recordsLinked to original sources

One-Dimensional Simulations of the Topological Defects in a 3:1 $U(1)$ Model

The domain wall is a kind of topological defect that can appear when a discrete symmetry is broken. If the discrete symmetry appears as an intermediate symmetry during a $U(1)$ symmetry breaking, the domain walls are connected to cosmic strings, forming walls bounded by strings. Intuitively, the domain wall disappears if the breaking scale of the discrete symmetry is comparable to that of the $U(1)$ symmetry. In this paper, relying on a 3:1 $U(1)$ model, we show the detailed processes of the disappearance of the domain wall. Due to the existence of the non-negligible ``bias angle'' $\beta$, the relevance of the ``$Z_3$ symmetry'' and the domain wall is blurred, and thereby the evaluations of the string profiles in a hybrid wall-string network should be revised. We also made some preliminary calculations of the gravitational waves generated by the wall-string network created in the early universe.

hep-ph

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming

Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities. Satellite-street scene pairs, with their complex contexts and extreme viewpoint variations, provide an ideal testbed. Motivated by this, we introduce CVSBench, a large-scale benchmark for evaluating cross-view spatial reasoning through satellite-street pairs. This benchmark supports multiple tasks, including cross-view VQA, cross-view grounding, and viewpoint identification. CVSBench comprises 3,297 cross-view image groups with 9,468 object-level annotations and 40,679 question-answer (QA) pairs, enabling systematic and controlled evaluation of cross-view spatial reasoning. Extensive evaluations reveal that advanced VLMs struggle to maintain object-level and layout consistency under drastic viewpoint changes. To bridge this gap towards human-like spatial cognition, we investigate two categories of approaches: spatially grounded reasoning and the incorporation of cognitive map inputs. Our findings demonstrate that language-only reasoning yields marginal improvements, while incorporating visual spatial imagination via a 3D scene imagination pipeline substantially improves cross-view reasoning. These results highlight the necessity of explicit visual-spatial representations for robust spatial cognition in VLMs. Our data and code are released at https://huggingface.co/datasets/zlyzlyzly/CVSBench.

cs.CV

MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop

Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, existing benchmarks, including those for macOS, evaluate agents without framework augmentation and rely on binary evaluation. As a result, they fail to capture both the framework capabilities leveraged by modern CUAs and the partial progress on long-horizon, multi-application tasks. We present MacAgentBench, a comprehensive macOS agent benchmark comprising 676 tasks across 25 applications, with nearly 60% involving both GUI and CLI interaction. The benchmark adopts deterministic rule-based evaluation and introduces fine-grained multi-checkpoint scoring with capability annotations for multi-application tasks. Experiments across three frameworks and 16 models show that the best configuration, Claude Opus 4.6 on OpenClaw, attains 73.7% Pass@1, while this advantage is primarily driven by the skill library rather than by framework design. Fine-grained metrics further reveal that models with similar Pass@1 can differ substantially in sub-goal completion. Our code and data are publicly available at https://github.com/JetAstra/MacAgentBench.

cs.AI

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution from research assistants into autonomous research agents, these systems still exhibit significant limitations in field sensitivity, research ethics, and nuanced scientific judgment. Consequently, frontier agents remain unable to fully replace human researchers. To bridge this gap, we conceptualize the AARR (Act As a Real Researcher) benchmark series. Unlike existing benchmarks that primarily assess macro-level execution capabilities, AARR focuses on whether agents can emulate the professionalism, thoroughness, and nuanced reasoning that characterize human researchers in granular research scenarios. In this work, we propose AARRI-Bench (Act As a Real Research Intern), the first benchmark in this series. We conduct extensive experiments across frontier models and agentic systems, revealing that even the best-performing configuration (Mini-SWE-Agent with Claude Opus 4.7) achieves only 68.3\% success rate, frequently overlooking subtle yet critical details that are obvious to real human researchers. Our results indicate that developing researcher-like AI requires further exploration of research behavior, rather than merely complex scaffolding. Our data is released at https://github.com/AARR-bench/AARRI-bench.

cs.AI

The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results

Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across distinct domains. As part of NTIRE 2026, we hosted the second CD-FSOD Challenge to systematically evaluate and promote progress in detecting objects in unseen target domains under limited annotation conditions. The challenge received strong community interest, with 128 registered participants and a total of 696 submissions. Among them, 31 teams actively participated, and 19 teams submitted valid final results. Participants explored a wide range of strategies, introducing innovative methods that push the performance frontier under both open-source and closed-source tracks. This report presents a detailed overview of the NTIRE 2026 CD-FSOD Challenge, including a summary of the submitted approaches and an analysis of the final results across all participating teams. Challenge Codes: https://github.com/ohMargin/NTIRE2026_CDFSOD.

cs.CV

Domain Walls in $A_4$ Flavour Models

The spontaneous breaking of an $A_4$ flavour symmetry, often used to predict leptonic mixing, can lead to the formation of domain walls which can annihilate and generate a stochastic gravitational wave background. We study this phenomenon in three scenarios where the nature of the scalar field responsible for breaking the $A_4$ symmetry spontaneously differs: real, complex, and supersymmetric. For the real scalar, a biased potential produces metastable walls that decay into oscillating two-wall systems with important consequences for gravitational wave signals. In the complex scalar case, we discuss the interplay between domain walls and global strings and classify the types of domain walls that form in terms of the $A_4$ group symmetries. We investigate the properties of supersymmetric $A_4$ domain walls, and highlight the BPS walls. Through a detailed analysis of these models with non-Abelian symmetries, we discover new kinds of domain walls, which we denote as ``oreo''-type composite domain walls, CP-violating domain walls and SUSY non-Abelian domain walls. Finally we show how these results may be achieved in leptonic $A_4$ flavour models, with and without supersymmetry, and discuss their distinctive gravitational wave signatures.

hep-ph

Ovis-Image Technical Report

We introduce $\textbf{Ovis-Image}$, a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational constraints. Built upon our previous Ovis-U1 framework, Ovis-Image integrates a diffusion-based visual decoder with the stronger Ovis 2.5 multimodal backbone, leveraging a text-centric training pipeline that combines large-scale pre-training with carefully tailored post-training refinements. Despite its compact architecture, Ovis-Image achieves text rendering performance on par with significantly larger open models such as Qwen-Image and approaches closed-source systems like Seedream and GPT4o. Crucially, the model remains deployable on a single high-end GPU with moderate memory, narrowing the gap between frontier-level text rendering and practical deployment. Our results indicate that combining a strong multimodal backbone with a carefully designed, text-focused training recipe is sufficient to achieve reliable bilingual text rendering without resorting to oversized or proprietary models.

cs.CV

ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a passive perception paradigm, suffering from increased redundancy when obtaining finer visual inputs. In this work, we explore a new active perception paradigm that enables models to revisit information-rich regions. First, we present LRS-GRO, a large-scale benchmark dataset tailored for active perception in UHR RS processing, encompassing 17 question types across global, region, and object levels, annotated via a semi-automatic pipeline. Building on LRS-GRO, we propose ZoomEarth, an adaptive cropping-zooming framework with a novel Region-Guided reward that provides fine-grained guidance. Trained via supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO), ZoomEarth achieves state-of-the-art performance on LRS-GRO and, in the zero-shot setting, on three public UHR remote sensing benchmarks. Furthermore, ZoomEarth can be seamlessly integrated with downstream models for tasks such as cloud removal, denoising, segmentation, and image editing through simple tool interfaces, demonstrating strong versatility and extensibility.

cs.CV

Assessing the Effects of Monetary Shocks on Macroeconomic Stars: A SMUC-IV Framework

This paper proposes a structural multivariate unobserved components model with external instrument (SMUC-IV) to investigate the effects of monetary policy shocks on key U.S. macroeconomic "stars"-namely, the level of potential output, the growth rate of potential output, trend inflation, and the neutral interest rate. A key feature of our approach is the use of an external instrument to identify monetary policy shocks within the multivariate unobserved components modeling framework. We develop an MCMC estimation method to facilitate posterior inference within our proposed SMUC-IV framework. In addition, we propose an marginal likelihood estimator to enable model comparison across alternative specifications. Our empirical analysis shows that contractionary monetary policy shocks have significant negative effects on the macroeconomic stars, highlighting the nonzero long-run effects of transitory monetary policy shocks.

econ.EM

Minimal Multi-Majoron Model

In order to provide a natural framework for hierarchical right-handed neutrinos, we propose a realistic ultraviolet complete minimal multi-Majoron model (MMMM). We consider two right-handed neutrinos for simplicity, although the model is readily extendable to more. The minimal model introduces two complex scalar Majoron fields $\phi_1$ and $\phi_2$, whose couplings to the two respective right-handed neutrinos are controlled by an extra global $U(1)_N$ symmetry. We show that a flavon field is required to facilitate the effective Yukawa couplings, in order to implement the type I seesaw mechanism. We analyse the resulting phenomenology related to neutrino masses, flavour mixing and cosmological predictions concerning the formation and decay of topological defects like the global cosmic strings and the domain walls when the $U(1)_N\times U(1)_{B-L}$ symmetry is broken. The resulting gravitational wave spectrum is a distinctive combination of the spectrum from the global cosmic string and strong first-order phase transitions when the symmetries are broken, the strength of the latter being enhanced by the second Majoron field. The resulting characteristic spectrum determines the two right-handed neutrino mass scales within the considered framework.

hep-ph

Non-Abelian Domain Walls and Gravitational Waves

We investigate the properties of domain walls arising from non-Abelian discrete symmetries, which we refer to as non-Abelian domain walls. We focus on $S_4$, one of the most commonly used groups in lepton flavour mixing models. The spontaneous breaking of $S_4$ leads to distinct vacua preserving a residual $Z_2$ or $Z_3$ symmetry. Five types of domain walls are found, labelled as SI, SII, TI, TII, and TIII, respectively, the former two separating $Z_2$ vacua and the latter three separating $Z_3$ vacua. We highlight that SI, TI and TIII may be unstable for some regions of the parameter space and decay to stable domain walls. Stable domain walls can collapse and release gravitational radiation for a suitable size of explicit symmetry breaking. A symmetry-breaking scale of order 100 TeV may explain the recent discovery of nanohertz gravitational waves by PTA experiments. For the first time, we investigate the properties of these domain walls, which we obtain numerically with semi-analytical formulas applied to compute the tension and thickness across a wide range of parameter space. We estimate the resulting gravitational wave spectrum and find that, thanks to their rich vacuum structure, non-Abelian domain walls manifest in a very interesting and complex phenomenology.

hep-ph

Type-I two-Higgs-doublet model and gravitational waves from domain walls bounded by strings

The spontaneous breaking of a $U(1)$ symmetry via an intermediate discrete symmetry may yield a hybrid topological defect of \emph{domain walls bounded by cosmic strings}. The decay of this defect network leads to a unique gravitational wave signal spanning many orders in observable frequencies, that can be distinguished from signals generated by other sources. We investigate the production of gravitational waves from this mechanism in the context of the type-I two-Higgs-doublet model extended by a $U(1)_R$ symmetry, that simultaneously accommodates the seesaw mechanism, anomaly cancellation, and eliminates flavour-changing neutral currents. The gravitational wave spectrum produced by the string-bounded-wall network can be detected for $U(1)_R$ breaking scale from $10^{12}$ to $10^{15}$ GeV in forthcoming interferometers including LISA and Einstein Telescope, with a distinctive $f^{3}$ slope and inflexion in the frequency range between microhertz and hertz.

hep-ph

On the cosmological abundance of magnetic monopoles

We demonstrate that Debye shielding cannot be employed to constrain the cosmological abundance of magnetic monopoles, contrary to what is stated in the previous literature. Current model-independent bounds on the monopole abundance are then revisited for unit Dirac magnetic charge. We find that the Andromeda Parker bound can be employed to set an upper limit on the monopole flux at the level of $F_M\lesssim 5.3\times 10^{-19}\,\text{cm}^{-2}\text{s}^{-1}\text{sr}^{-1}$ for a monopole mass $10^{13}\,\text{GeV}/c^2\lesssim m\lesssim 10^{16}\,\text{GeV}/c^2$, which is more stringent than the MACRO direct search limit by two orders of magnitude. This translates into stringent constraints on the monopole density parameter $\Omega_M$ at the level of $10^{-7}-10^{-4}$ depending on the mass. For larger monopole masses the scenarios in which magnetic monopoles account for all or the majority of dark matter are disfavored.

hep-ph

Non-orthogonal cavity modes near exceptional points in the far field

Non-orthogonal eigenstates are a fundamental feature of non-Hermitian systems and are accompanied by the emergence of nontrivial features. However, the platforms to explore non-Hermitian mode couplings mainly measure near-field effects, and the far-field behaviour remain mostly unexplored. Here, we study how a microcavity with non-Hermitian mode coupling exhibits eigenstate non-orthogonality by investigating the spatial field and the far-field polarization of cavity modes. The non-Hermiticity arises from asymmetric backscattering, which is controlled by integrating two scatterers of different size and location into a microdisk. We observe that the spatial field overlaps of two modes increases abruptly to its maximum value, whilst different far-field elliptical polarizations of two modes coalesce when approaching an exceptional point. We demonstrate such features experimentally by measuring the far-field polarization from the fabricated microdisks. Our work reveals the non-orthogonality in the far-field degree of freedom, and the integrability of the microdisks paves a way to integrate more non-Hermitian optical properties into nanophotonic systems.

physics.optics

D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction

Reconstructing hand-held objects from a single RGB image is a challenging task in computer vision. In contrast to prior works that utilize deterministic modeling paradigms, we employ a point cloud denoising diffusion model to account for the probabilistic nature of this problem. In the core, we introduce centroid-fixed dual-stream conditional diffusion for monocular hand-held object reconstruction (D-SCo), tackling two predominant challenges. First, to avoid the object centroid from deviating, we utilize a novel hand-constrained centroid fixing paradigm, enhancing the stability of diffusion and reverse processes and the precision of feature projection. Second, we introduce a dual-stream denoiser to semantically and geometrically model hand-object interactions with a novel unified hand-object semantic embedding, enhancing the reconstruction performance of the hand-occluded region of the object. Experiments on the synthetic ObMan dataset and three real-world datasets HO3D, MOW and DexYCB demonstrate that our approach can surpass all other state-of-the-art methods.

cs.CV

ShapeMatcher: Self-Supervised Joint Shape Canonicalization, Segmentation, Retrieval and Deformation

In this paper, we present ShapeMatcher, a unified self-supervised learning framework for joint shape canonicalization, segmentation, retrieval and deformation. Given a partially-observed object in an arbitrary pose, we first canonicalize the object by extracting point-wise affine-invariant features, disentangling inherent structure of the object with its pose and size. These learned features are then leveraged to predict semantically consistent part segmentation and corresponding part centers. Next, our lightweight retrieval module aggregates the features within each part as its retrieval token and compare all the tokens with source shapes from a pre-established database to identify the most geometrically similar shape. Finally, we deform the retrieved shape in the deformation module to tightly fit the input object by harnessing part center guided neural cage deformation. The key insight of ShapeMaker is the simultaneous training of the four highly-associated processes: canonicalization, segmentation, retrieval, and deformation, leveraging cross-task consistency losses for mutual supervision. Extensive experiments on synthetic datasets PartNet, ComplementMe, and real-world dataset Scan2CAD demonstrate that ShapeMaker surpasses competitors by a large margin.

cs.CV

LanPose: Language-Instructed 6D Object Pose Estimation for Robotic Assembly

Comprehending natural language instructions is a critical skill for robots to cooperate effectively with humans. In this paper, we aim to learn 6D poses for roboticassembly by natural language instructions. For this purpose, Language-Instructed 6D Pose Regression Network (LanPose) is proposed to jointly predict the 6D poses of the observed object and the corresponding assembly position. Our proposed approach is based on the fusion of geometric and linguistic features, which allows us to finely integrate multi-modality input and map it to the 6D pose in SE(3) space by the cross-attention mechanism and the language-integrated 6D pose mapping module, respectively. To validate the effectiveness of our approach, an integrated robotic system is established to precisely and robustly perceive, grasp, manipulate and assemble blocks by language commands. 98.09 and 93.55 in ADD(-S)-0.1d are derived for the prediction of 6D object pose and 6D assembly pose, respectively. Both quantitative and qualitative results demonstrate the effectiveness of our proposed language-instructed 6D pose estimation methodology and its potential to enable robots to better understand and execute natural language instructions.

cs.RO

MOHO: Learning Single-view Hand-held Object Reconstruction with Multi-view Occlusion-Aware Supervision

Previous works concerning single-view hand-held object reconstruction typically rely on supervision from 3D ground-truth models, which are hard to collect in real world. In contrast, readily accessible hand-object videos offer a promising training data source, but they only give heavily occluded object observations. In this paper, we present a novel synthetic-to-real framework to exploit Multi-view Occlusion-aware supervision from hand-object videos for Hand-held Object reconstruction (MOHO) from a single image, tackling two predominant challenges in such setting: hand-induced occlusion and object's self-occlusion. First, in the synthetic pre-training stage, we render a large-scaled synthetic dataset SOMVideo with hand-object images and multi-view occlusion-free supervisions, adopted to address hand-induced occlusion in both 2D and 3D spaces. Second, in the real-world finetuning stage, MOHO leverages the amodal-mask-weighted geometric supervision to mitigate the unfaithful guidance caused by the hand-occluded supervising views in real world. Moreover, domain-consistent occlusion-aware features are amalgamated in MOHO to resist object's self-occlusion for inferring the complete object shape. Extensive experiments on HO3D and DexYCB datasets demonstrate 2D-supervised MOHO gains superior results against 3D-supervised methods by a large margin.

cs.CV