SearcharxivSearch

arXiv subjects

Bowen Li

Publications and source records attributed to Bowen Li.

At least 19 recordsLinked to original sources

Classification of finite-dimensional Nichols algebras of rank two in twisted Yetter--Drinfeld categories

Let \(G\) be a finite non-abelian group, let \(\Phi\in Z^3(G,\mathbb C^\times)\) be normalized, and let \(V,W\) be finite-dimensional simple objects of \({}_G^G\mathcal{YD}^{\Phi}\). We classify, up to interchange, the braided-indecomposable pairs \((V,W)\) whose supports generate \(G\) and for which \(\mathcal B(V\oplus W)\) is finite-dimensional. The classification consists of eight cases with five possible support quandles. In every case the Cartan graph is standard of type \(A_2\), \(B_2\), or \(G_2\), and the dimension is determined explicitly. A new phenomenon occurs for \(\Gamma _2\): the twisted setting admits a family of type \(G_2\) absent from the ordinary \(\Gamma _2\) classification and we construct an explicit example over a non-abelian group of order \(16\).

math.QA

Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection

Graphic designs, such as posters, advertisements, and infographics, are an important medium for communicating information and shaping understanding. Unlike natural images, they consist of layered elements with explicit compositional order. However, existing object detection models treat these elements as an unordered set, leaving compositional order unexploited. To address this limitation, we present Detect Anything in Graphic Design (DAD), a model that formulates graphic design detection as compositional deconstruction. It decodes elements in compositional order, using lower-layer elements to better detect higher-layer ones. The key feature of DAD is amodal detection, which predicts the full bounding box of each element, including regions occluded by elements placed above it. Building on this formulation, we propose Element Relative Policy Optimization (EleRPO), which extends GRPO from sequence-level supervision to element-level optimization. EleRPO provides fine-grained training signals that capture how each detected element contributes to overall detection quality, and works synergistically with compositional order to improve detection performance. To support training and evaluation, we build a dataset of 10 million graphic designs. Experiments show that DAD outperforms all baselines and achieves human-level performance in amodal detection, supporting effective image-to-layer decomposition. EleRPO consistently improves over GRPO across nine detection benchmarks.

cs.CV

EEG-Driven Decoding Framework for Passenger Hazard Perception in Highly Automated Vehicles

Reliable risk assessment remains a central challenge for Autonomous Vehicles (AVs). Despite advances in automation, passenger cognition provides a non-intrusive auxiliary signal that improves both objective and perceived safety without requiring active human intervention. We introduce an Electroencephalogram (EEG)-based Brain-Computer Interface (BCI) that decodes passenger neural responses for both Risk Prediction (RP) and Danger Identification (DI), explicitly modeling humans as passengers to match real-world AV use. To achieve this, we propose the Passenger Cognitive Model (PCM), Risk-aware Sequential Labeling (RSL), and the Passenger EEG Decoding Strategy (PEDS), which integrates a 3D Convolutional Recurrent Neural Network (3D-CRNN) model for joint EEG decoding. Experimental results show that 3D-CRNN achieves a Balanced Accuracy (BA) of $95.3\% \pm 2.7\%$ in RP and improves single-subject DI from $80.9\% \pm 3.9\%$ to $85.0\% \pm 3.2\%$ with RSL. Event-wise analyses further show that 3D-CRNN consistently outperforms other models across different event types in RP and DI. In generalization experiments, 3D-CRNN achieves $77.0\% \pm 5.3\%$ BA in cross-session DI and $77.4\% \pm 1.1\%$ BA on seen subjects in cross-subject evaluation, while maintaining a $64.9\% \pm 8.5\%$ BA on unseen subjects, demonstrating promising generalizability and transferability across both intra-subject and inter-subject variability. These findings establish an Electroencephalogram (EEG) decoding framework for AV passenger hazard perception and suggest that passenger cognitive signals can provide auxiliary supervision for future AV decision-making and Safety of the Intended Functionality (SOTIF) support.

cs.AI

Waveguiding in systems of high contrast resonators: Theory and fast computations

In this work, we study guided modes in systems of high-contrast resonators near nonzero interior Neumann frequencies, beyond the subwavelength regime. In the regular exterior regime, where the exterior Dirichlet problem is well-posed at the reference wavenumber, we introduce an infinite-dimensional frequency-dependent capacitance operator obtained by compressing the exterior Helmholtz Dirichlet-to-Neumann map to the traces of the interior resonant Neumann eigenspaces. We prove the norm-resolvent convergence of the continuous problem to this discrete effective operator as the contrast $\delta\to0$, and derive first-order asymptotic formulas for compact-defect frequencies and line-defect band functions. We then establish exponential off-diagonal decay of the capacitance coefficients by a Combes--Thomas argument, yielding an exponentially accurate truncation of the discrete operator, and show that its retained coefficients can be computed from local Helmholtz problems. At the physical frequency, this local approximation converges exponentially under a uniform stability assumption for the growing finite-cluster problems. The stability assumption can be removed by introducing a vanishing complex absorption together with a Hermitian symmetrization. In particular, an absorption parameter of order $\sqrt\delta$, together with interaction truncation and patch radii of order $|\log\delta|$, suffices to preserve the $O(\delta^2)$ accuracy of the first-order high-contrast expansion of the defect eigenfrequencies, yielding a fast computational method. Numerical experiments for dipole and quadrupole resonances illustrate the accuracy, exponential locality, and applicability of the discrete model to straight and bent waveguides generated by material or geometric detuning.

math.NA

Nonlinear Modal Reduction for Subwavelength Dielectric Scattering

We study three-dimensional wave scattering by high-index dielectric resonators with Kerr-type nonlinearity under plane-wave incidence, together with the associated nonlinear dielectric scattering resonances, through a nonlinear Lippmann-Schwinger equation. Using a Lyapunov-Schmidt reduction near a simple eigenmode of the Newtonian potential, we obtain a local decomposition of the scattered wave into a resonant contribution and a controlled remainder and derive an explicit nonlinear equation for the resonant modal coefficient for sufficiently small contrast parameter and locally small resonant amplitude and incident field. A second reduction covers the regime of incident fields of order one and weak nonlinearity. Our results extend for the first time the linear modal decomposition to nonlinear wave-scattering problems. We complement them by developing a numerical framework based on Nystr\"om discretization, real Newton iteration, and pseudo-arclength continuation. For single resonators of several geometries, our computations confirm the predicted high-contrast resonance scaling and the nonlinear modal approximation, and exhibit the multivalued incident-wave response. For a mirror-symmetric resonator dimer, we track symmetric, antisymmetric, and symmetry-broken resonance branches over a broad range of separations. Our computations show that the symmetry-breaking threshold increases as the gap between the resonators decreases, in agreement with the leading-order bifurcation theory.

math.AP

Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance

Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within policy. We introduce Fiducia-bench, a benchmark for the governability of financial agents---whether they escalate when obligated, abstain when required, and leave an auditable trail---and use it to study a question no prior benchmark addresses: does decomposing an agent into components degrade its governance? It does, and the mechanism is specific. Policy-relevant facts discovered by one component are attenuated at the handoff boundary before reaching the component that must act on them. In a 626-episode experiment across 100 KYC/AML task variants, two models, and three architectures, a 32B open-weights model attenuated 0% of discovered facts under a single-loop baseline, 56% under a fixed pipeline, and 85% under an orchestrator-subagent architecture (all at constraint distance 2). A stronger model (gpt-4.1-mini) attenuated 3-6% under the same conditions, suggesting the governance cost of decomposition is partly a function of model capability. Critically, the same mechanism produces both under-escalation and over-escalation, depending on whether the dropped fact was a risk signal or an exculpating one. The benchmark, all tasks, and the verification harness are open-source

cs.AI

Reaction-Transformation-Aware Flow Matching for Generalizable Transition State Generation

Transition-state (TS) structures define the energetic barriers and mechanistic pathways of elementary chemical reactions, yet their identification remains computationally demanding because conventional saddle-point searches require expensive quantum-mechanical calculations. Recent machine-learning approaches have accelerated TS generation by predicting structures from reaction endpoint information, but they primarily learn geometric correspondence between endpoints and TSs, leaving the structural transformations underlying elementary reactions implicitly represented. To address this limitation, we introduce TransTS, a reaction-transformation-aware framework for generalizable TS generation from atom-mapped reactant-product pairs. TransTS explicitly learns atom-level structural transformations between reaction endpoints and integrates them with a unified atom-aligned geometric representation of reactants, TSs and products, enabling reaction-aware equivariant generation of TS geometries. TransTS is designed to provide reliable TS initial guesses for subsequent quantum-chemical refinement, where generated structures are evaluated not only by geometric similarity but also by their ability to converge to validated saddle points and recover the intended reaction pathways. Across IID and zero-shot OOD benchmarks, TransTS demonstrates improved TS initialization quality, with particularly strong generalization to unseen reaction distributions. On the challenging GDB-10-rxn and GDB-17-rxn OOD benchmarks, TransTS generates TS candidates that more frequently converge to validated saddle points and recover the intended elementary reactions after refinement than existing approaches under the same training regime. Scaling reaction coverage and model capacity further improves both geometric fidelity and refinement outcomes.

physics.chem-ph

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact solution by compressing multimodal trajectories into a few continuous tokens. Existing methods, however, usually map each trajectory to one fixed memory block and train it mainly through next-action supervision. This creates three practical problems: important details may be lost during compression, the same memory block must serve different decision stages, and irrelevant retrieved trajectories may still mislead the agent. We introduce FocusMem, which separates these responsibilities within a compact latent-memory interface. A role-aware content basis encourages episodic memory to retain reusable experience and working memory to retain task progress. A state-conditioned readout generates a decision-specific view of the same stored evidence, while a lightweight trust gate can suppress memory blocks that appear irrelevant to the current step. All components are trained while the GUI policy remains frozen. Across five GUI-agent benchmarks, FocusMem consistently outperforms a fully matched action-only fixed-memory baseline and prior latent memory adaptations. Further analysis shows that semantic and functional supervision preserve complementary information, state-conditioned readout is more robust as surrounding trajectory context grows, and the trust gate reduces the harm caused by injected irrelevant episodic evidence. These results show that effective latent memory depends not only on compressing past interaction, but also on what is retained, what is exposed, and what is allowed.

cs.CV

Long Directed Cycles in Vertex-Transitive Digraphs

The search for Hamiltonian cycles in vertex-transitive graphs and digraphs is a classical problem at the interface of graph theory and group theory. In the undirected setting, this goes back to the well-known conjectures of Lov\'asz and Thomassen concerning Hamiltonian paths and cycles in connected vertex-transitive graphs. Dating back to Rankin's 1946 work, the directed analogue has an even longer history, linking the search for long cycles to classical group-rearrangement problems. Trotter and Erd\H{o}s showed in 1978 that connected vertex-transitive digraphs need not be Hamiltonian. In light of this result, Alspach asked in 1981 whether there exist connected vertex-transitive digraphs whose longest directed cycle misses arbitrarily many vertices. This question was only recently resolved by Buci\'c, Hendrey, Mohar, Steiner and Yepremyan, who constructed connected vertex-transitive digraphs on $n$ vertices whose longest directed cycle omits $(1-o(1))\log n$ vertices. They conjectured that the number of omitted vertices can grow linearly with $n$, remarking that it would already be interesting to improve their logarithmic lower bound to a polynomial bound. In this paper, we confirm their conjecture in a strong form by constructing infinitely many connected vertex-transitive digraphs on $n$ vertices whose longest directed cycle omits at least $n/12$ vertices. In the same work, Buci\'c, Hendrey, Mohar, Steiner and Yepremyan also proved that every connected vertex-transitive digraph on $n$ vertices contains a directed cycle of length $\Omega(n^{1/3})$, giving the first lower bound for this problem that grows with $n$. We improve this to $\Omega(\sqrt n)$, matching the order of Babai's classical theorem from 1979 for undirected vertex-transitive graphs.

math.CO

On the Cartan Graphs of Nichols Algebras over Coquasi-Hopf Algebras

Over an algebraically closed field of characteristic zero, let $H$ be a coquasi-Hopf algebra with bijective antipode, and let $M$ be a tuple of finite-dimensional simple Yetter--Drinfeld modules over $H$. We prove that, if $M$ admits all reflections, then its associated semi-Cartan graph is a Cartan graph. We characterize the finiteness of this Cartan graph by tensor decomposability of $\mathcal B(M)$ and obtain a finite-dimensionality criterion of Nichols algebras. We also show that braided monoidal equivalences preserve reflections and the associated Cartan graphs. As applications, we prove that every Cartan graph associated with a diagonal type tuple over a finite abelian group equipped with an abelian \(3\)-cocycle is covered by one arising from a diagonal type tuple over some finite abelian group $G$ with trivial associator, and that their real-root sets agree at corresponding objects. We also construct a Cartan graph of the former kind that cannot be obtained from any diagonal tuple in \({}_G^G\mathcal{YD}\).

math.QA

Rethinking Scientific Discovery in the Agentic Era

Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem formulation, literature grounding, model use, simulation, validation, and knowledge reuse. This paper presents \textbf{SCION (Scientific Collaborative Innovation with Agentic Organizational Nexus)}, an agentic scientific operating system that acts as an \textbf{organizational nexus}. Through a Science Agent serving as a \textbf{Meta-Harness}, SCION connects scientific tasks, tools, agents, artifacts, and memory, transforming research into an executable, auditable, and reusable operational process. At its core is the \textbf{Research Execution Plan (REP)}, which compiles high-level scientific intent into staged objectives, dependencies, verification checkpoints, tool requirements, expected artifacts, and fallback conditions. SCION further integrates hierarchical multi-agent execution, profile-driven specialization, selective context construction, governed delegation, and layered epistemic memory to support long-horizon scientific work. We formulate discovery under SCION as \textbf{Target-conditioned Inverse Search} and extend it to hidden-target settings through batch active search under finite experimental budgets. Applications in materials analysis, molecule design, and protein or antibody screening, together with experiments on scientific reading, idea generation, molecule generation, and antibody screening, show that SCION outperforms existing autonomous research-agent baselines, especially in decomposition, verification, refinement, and memory reuse. Overall, SCION shifts AI from isolated tools toward a coordinated operational layer for traceable and reusable scientific innovation.

cs.CL

Recover, Discover, Plan: Learning Skills and Concepts from Robot Failures

Intelligent robots should not only recover from failures, but also acquire the abstract knowledge needed to avoid them in the future. While reinforcement learning (RL) can learn reactive recovery behaviors, training a separate policy for every distinct failure mode is highly inefficient. We introduce Recovery-Driven Synthesis of Relational Concepts (ReSYNC), the first approach that progressively discovers and refines state abstractions (relational predicates) from failure-recovery experience to support abstract planning. Unlike purely reactive methods, ReSYNC jointly learns skills and concepts through an incremental dual-learning process. In the skill-learning phase, the robot uses RL to learn to recover from failures seen in training tasks. In the concept-learning phase, the robot discovers new relational predicates and refines its abstract planning model to explain and generalize the learned recovery behaviors. This interaction enables ReSYNC to convert local recoveries seen during training into global failure avoidance at test time. Across four simulated domains, we show that ReSYNC's ability to continually expand and refine its abstraction library allows it to solve long-horizon, previously unseen problems, outperforming strong baselines by over 50%. Additionally, we demonstrate sim-to-real transfer of ReSYNC, where it performs real-world non-prehensile manipulation skills and generalizes to unseen scenarios through abstract planning. Overall, ReSYNC represents a significant step toward robots that autonomously acquire abstractions for scalable, failure-aware planning in the physical world.

cs.RO

Neuro-Symbolic Learning for Long-Horizon Task Planning Under Complex Logical Constraints

Task planning often suffers from severe efficiency bottlenecks when robots must reason over long-horizon action sequences under complex logical constraints, including object affordances, spatial relationships, and sequential action dependencies. Recent neuro-symbolic methods improve planning efficiency by learning object-importance scores to prune task-irrelevant objects, but they typically rely on fixed offline supervision generated from full search spaces. This creates a train-test mismatch: at deployment, the planner operates in pruned search spaces induced by the model's own imperfect predictions, leading to exposure bias and degraded planning performance. To address this challenge, we formulate object-importance learning for task planning as an imperative learning-based bilevel optimization problem. The upper level optimizes a neural scorer, while the lower level solves a symbolic planning problem in the score-pruned search space. To stabilize this learning process, we introduce a 3R strategy into the lower-level planning, using parallel Repair, Restart, and Rollback recovery to provide reliable and adaptive feedback for upper-level learning. Experiments on three challenging benchmarks demonstrate state-of-the-art performance, including an 80.04% reduction in failure rate and a 57.14% reduction in planning time. We further validate the framework on a quadruped-based mobile manipulator in simulation and the real world, demonstrating its potential for efficient and deployable neuro-symbolic task planning.

cs.RO

Answer Presence Drives RAG Rewriting Gains

Retrieval-augmented QA pipelines often route retrieved passages through an LLM \emph{rewriter} before a smaller reader, lifting F1 by tens of points on multi-hop benchmarks; this gain is typically credited to improved evidence quality. We ask whether that lift is causally driven by the gold answer string appearing in the rewritten context rather than by curation per se, using a controlled intervention audit. For each rewritten context we re-run the reader after one of four controlled edits to the compile output: removing the gold answer span, replacing a length-matched random non-answer span (placebo), or injecting the gold into rewrites where it was absent (at the prefix or at a midpoint sentence boundary). Across twelve completed (cell, baseline) intervention runs spanning three reader families (Qwen2.5-7B, Qwen3.5-35B, GLM-4.7), two datasets (HotpotQA, 2WikiMultihopQA), and three compiler arrangements (MA-only, MB-only, MA$+$verify), removing the gold answer drops reader F1 by $28$ to $64$ points beyond the length-matched placebo on paired \texttt{answer-in-compile} strata, and prepending the gold into rewrites that lacked it raises F1 by $+0.7$ to $+9.7$ points in $10$ of $12$ (cell, baseline) combinations. A companion five-sentinel audit shows the conventional single-\texttt{[MASK]} probe is itself sentinel-fragile: on 2Wiki it reports a $+4.12$~F1 ``non-leakage residual'' that flips to $-3.33$ to $-7.81$~F1 under four alternative sentinels and fails an equivalence test for three of those four ($1/4$~pass). We do not propose a new rewriter or mitigation; we release the intervention runner and the sentinel panel so that other rewriter-gain claims can be tested against the same standard.

cs.AI

Passively synchronized dual-color mode-locked fiber lasers based on nonlinear amplifying loop mirrors

We have proposed and implemented a novel scheme for passive all-optical synchronization between erbium and ytterbium mode-locked fiber lasers. The passive locking of repetition rates for the dual-color pulses was realized by cross-phase modulation within phase-biased nonlinear amplifying loop mirrors. In contrast to previous demonstrations, the synchronization system was configured in an all-polarization-maintaining structure, thus gaining substantially improved stability and robustness. Consequently, the maximum tolerance of cavity-length mismatch of 16.2 mm was achieved unprecedentedly, which was at least one order of magnitude longer than previously reported results for comparable temporal durations of involved pulses. The corresponding relative timing jitter was measured to be 31 fs within 1-MHz bandwidth. Such tight and robust synchronization fiber laser system offers a great potential for various applications, such as pump-probe microscopy, Raman scattering spectroscopy and nonlinear frequency generation.

physics.optics

Coincidence-pumping upconversion detector based on passively synchronized fiber laser system

We experimentally demonstrated a high-performance frequency upconversion detector for telecom-band photons based on a passively synchronized fiber laser system. The involved coincidence pumping technique enabled to spectrally convert the pulsed infrared photons into the visible regime with a conversion efficiency of 72\%. The overall detection efficiency of the upconversion detector reached to 30\% with a low noise equivalent power of $3\times10^{-17}\ \text{W/Hz}^{1/2}$. In contrast to previous demonstrations, the whole upconversion detection system was constructed in an all-polarization-maintaining fiber structure, thus favoring substantial improvement of compactness and robustness. Moreover, the long-term stability was manifested by at least ten-hour operation with a relative fluctuation of count rates as small as 0.26\%. The achieved features here would be desirable in many practical applications requiring efficient and robust coherent manipulation of pulsed optical fields by nonlinear frequency conversion.

physics.optics

HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete human personalities where emotional dimensions hold equal importance. In this paper, we introduce a novel benchmark to systematically assess whether LLM agents can simulate coherent, human-like psychology. Specifically, our benchmark constructs 11 diverse human characters grounded in orthogonal Big Five personality traits, with each profile deeply integrated with 1,000 structured autobiographical-style episodic memories distributed across theory-grounded developmental life stages. To rigorously evaluate the psychological manifestations of LLMs, we designed a curated suite of 64 decision-making scenarios, guided by the DIAMONDS taxonomy, a psychological framework that characterizes situations along eight dimensions: Duty, Intellect, Adversity, Mating, pOsitivity, Negativity, Deception, and Sociality. By subjecting agents to varying scenarios, the benchmark evaluates whether they can consolidate their innate personality traits and autobiographical memories to make behavioral decisions that are consistent with their specific psychological profiles. After systematic human validation and filtering, we obtained a benchmark consisting of 673 multiple-choice questions (MCQs). We believe this benchmark provides a principled and scalable testbed for studying human-like emotions, personality consistency, and value-consistent behavioural decision-making in LLM-based agents.

cs.CL

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

Building state-of-the-art text-to-speech (TTS) systems typically demands millions of hours of proprietary data and complex multi-stage architectures, creating substantial barriers for resource-constrained research teams. In this report, we present PilotTTS, a lightweight autoregressive TTS system that achieves competitive performance through minimalist architecture and rigorous data engineering. PilotTTS is trained on only 200K hours of data processed entirely with open-source tools. Specifically, our contributions are: (1) a reproducible multi-stage data processing pipeline covering quality assessment, label annotation, and filtering, and (2) a compact model architecture that employs Q-Former-based conditioning to decouple speaker identity from speaking style via cross-sample paired training. Within a unified framework, PilotTTS supports zero-shot voice cloning, emotion synthesis (11 categories), paralinguistic synthesis (4 categories), and Chinese dialect synthesis (14 dialects). On the Seed-TTS Eval benchmark, PilotTTS achieves the lowest WER of 1.50% on test-en, a CER of 0.87% on test-zh, and the highest speaker similarity on both test sets (0.862 and 0.815), outperforming systems trained on significantly larger datasets. We release the complete data pipeline recipe, pretrained weights, and code at https://github.com/AMAPVOICE/PilotTTS.

cs.SD