Searcharxiv⌕ Search

arXiv subjects

Rui Chen

Publications and source records attributed to Rui Chen.

At least 37 records · Page 2Linked to original sources

On Faber-Krahn inequality for Dirichlet Fractional-Logarithmic Laplacian

We study the Faber--Krahn problem for the fractional--logarithmic Laplacian $(-Δ)^{s+\ln}$, $0<s<1$, with Fourier symbol $|ξ|^{2s}\ln|ξ|^2$. We prove a small-volume and large-volume dichotomy. For sufficiently small volume, balls minimize the first Dirichlet eigenvalue, with rigidity up to translation.

math.AP↗

Doppler-Robust Vortex Wavefront Design for Integrated Sensing and Communication

Integrated sensing and communication (ISAC) is a promising paradigm for future wireless systems due to spectrum reuse, hardware sharing, and joint waveform design. In dynamic scenes, Doppler shifts degrade both sensing and communication, which is particularly critical for beam-sensitive orbital angular momentum (OAM) wavefronts. To address this, we propose a Doppler-robust ISAC framework, which first senses and then communicates. Specifically, in the sensing phase, multiple vortex modes are simultaneously transmitted via code-division mode-multiplexing (CDMM). To solve Doppler-induced inter-mode interference, we propose a velocity-consistency matching (VCM)-expectation maximization (EM) algorithm that jointly decodes the sensing matrix and estimates range, azimuth, elevation, and velocity for multiple moving targets. In the communication phase, the joint transmitter (Tx) beamforming and receiver (Rx) beam steering are configured from the estimated channel state information (CSI). We further quantify the sensing-communication allocation trade-off by evaluating how pilot length affects estimation accuracy, beam alignment, and spectral efficiency (SE). Simulation results show that the proposed VCM-EM and ISAC designs achieve higher sensing accuracy and communication SE than baseline schemes in dynamic scenarios.

eess.SP↗

Fully Coupled Nonlinear Forward-Backward Stochastic Difference Equations: Spectral Contraction and Infinite-Horizon Maximum Principle

This paper develops an explicit spectral-contraction approach to studying fully coupled nonlinear forward--backward stochastic difference equations on infinite horizon and their applications to stochastic control. The estimates for this fully coupled system are assembled into an explicit two-dimensional nonnegative matrix. Its spectral radius yields an explicit sufficient discount threshold, and a corresponding equivalent weighted product norm is constructed to establish the contraction property. Moreover, for an infinite-horizon control problem with convex control constraints and an accumulated discounted running cost, we derive a Pontryagin-type stochastic maximum principle, its equivalent pointwise normal-cone formulation, and a verification theorem. Finally, a recursive risk-adjusted portfolio example is given to demonstrate the applications of our theoretical results, and a projection-type sufficient optimality condition for this example is derived.

math.OC↗

Two-point estimates for the logarithmic p-flux of the first Dirichlet p-eigenfunction

Let \(u>0\) be the first Dirichlet \(p\)-eigenfunction on a bounded convex domain \(Ω\subset\mathbb R^N\), and set \[ X_Ω:=|\nabla\log u|^{p-2}\nabla\log u. \] We study whether the sharp one-dimensional two-point modulus for \(X_Ω\), which for \(p=2\) reduces to the logarithmic-gradient estimate of Andrews and Clutterbuck, persists for \(p\neq2\). We prove sharp estimates on intervals and balls for every \(p>1\), with the radial modulus on balls strictly larger than the one-dimensional one. In dimensions \(N\ge2\), however, the one-dimensional modulus fails on general convex domains for every \(p\neq2\). For \(1 2\), on thin domains \(Ω_\varepsilon=D\times(-\varepsilon,\varepsilon)\), with \(D\subset\mathbb R^{N-1}\) bounded and convex, the normalized first eigenfunctions satisfy \[ u_\varepsilon(x,\varepsilon z)\longrightarrow \frac{ϕ(z)}{ϕ(0)} \left(\frac{G_D(x)}{G_D(x_0)}\right)^{2/p} \] locally uniformly in \(D\times(-1,1)\), where \(ϕ\) and \(G_D\) are the first Dirichlet eigenfunctions of the \(p\)-Laplacian on \((-1,1)\) and of the Laplacian on \(D\), respectively. This yields the failure for \(p>2\). Finally, for arbitrary \(C^2\) functions, positivity of the symmetric differential of the \(p\)-gradient implies convexity, and the converse holds universally if and only if \(p=2\).

math.AP↗

Editable Visual Design

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the ``creative brain'' for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand ``visual world simulator'' to synthesize standalone visual assets. Operating under an ``imagine first, then act'' closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.

cs.CV↗

Learning 3D Editing without Paired Supervision via Generative Prior Distillation

Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the severe scarcity of high-quality paired training data. Existing approaches attempt to bypass this by either relying on slow test-time optimization or training on pseudo-pairs constructed via complex pipelines, which often introduce structural drift and geometric artifacts. In this paper, we propose a novel framework that learns feed-forward 3D editing without paired 3D supervision via Generative Prior Distillation. Instead of relying on ground-truth 3D pairs, our core idea is to distill visual, semantic, and geometric knowledge from powerful foundation models directly into a 3D editing model. Specifically, through a differentiable rendering pipeline, we supervise the 3D representation using two complementary signals: a 2D visual prior from an image editing model at the main editing view, and a semantic prior from a Vision-Language Model at novel views to ensure strict instruction following and source identity preservation. Crucially, to address the geometric collapse and multi-view inconsistencies inherent in 2D projection supervision, we introduce a 3D-aware Distribution Matching regularization. Acting as a geometric prior, this term operates in the 3D latent space, constraining the edited output to remain within the manifold of realistic 3D assets defined by a pretrained image to 3D teacher model. Extensive experiments demonstrate that our method achieves superior instruction fidelity and cross-view consistency, significantly outperforming state-of-the-art baselines. Our project is available at: https://github.com/thiamine128/PriorEdit3D.

cs.CV↗

Analysis of Electromagnetic Scattering from Semiconductor Nanostructures by Solving Coupled Volume Integral and Two-fluid Hydrodynamic Equations

Semiconductor-based plasmonic nanostructures support localized surface plasmon modes in the infrared region. Unlike metallic nanostructures, they support both free electrons and holes, requiring a two-fluid hydrodynamic Drude equation (HDE) to accurately capture spatial dispersion effects and low-frequency acoustic plasmon modes that cannot be described by single-fluid models. In this work, a volume integral equation (VIE)-based solver is proposed for the analysis of electromagnetic scattering from semiconductor nanostructures. The proposed approach couples the VIE, formulated in terms of the electric flux density and the free-electron and hole polarization currents, with the two-fluid HDE. The coupled system is discretized using a tetrahedral mesh and solved efficiently using a two-level iterative solver. In contrast to finite-element-based methods, the proposed VIE-based approach does not require domain-wide meshing and inherently satisfies the radiation condition, thereby eliminating artificial absorbing boundaries. Numerical results for InSb-type semiconductor nanostructures demonstrate the accuracy and efficiency of the proposed VIE-based solver and its ability to capture unique optical phenomena, such as acoustic plasmon resonances and the blueshift of localized surface plasmon resonances, that cannot be described by the single-fluid HDE or classical Drude-based models.

physics.optics↗

Pion radiative decays of excited hidden-charm pentaquark molecules: from $Σ_c^{(*)}\bar{D}^{(*)}(2S)$ molecules to the reported $P_c$ states

The discovery of the hidden-charm pentaquarks \(P_c(4312)\), \(P_c(4440)\) and \(P_c(4457)\) by the LHCb Collaboration are very likely to identify as the \(Σ_c^{(*)}\bar{D}^{(*)}\) molecules. A natural and crucial extension is the existence of excited molecular partners built from a ground-state charmed baryon and a radially excited anti-charmed meson, namely \(Σ_c^{(*)}\bar{D}^{(*)}(2S)\) molecules. In a framework of chiral quark model, we systematic study pion-emission decays of such excited molecules into the known ground-state \(P_c\) molecules. Our results show that the decay widths are sensitive to the spin structures and the coupled-channel interferences, i.e., the \(Σ_c\bar{D}(2S)/Σ_c\bar{D}^*(2S)/Σ_c^*\bar{D}^*(2S)[1/2(1/2^-)]\) state decays to \(P_c(4440)\) with a width of several MeV, while the width to \(P_c(4457)\) is suppressed below \(0.3\) MeV due to destructive interference. The pion-emission decay can be the key to unveiling the excited molecular spectrum of hidden-charm pentaquarks and provides decisive experimental signatures. We expect the future experiments such as the LHCb and PANDA can verify our predictions.

hep-ph↗

Optical conductivity signature of Van Hove singularity in altermagnetic topological systems

We investigate the topological phases, joint density of states (JDOS), optical conductivities, and magneto-optical responses of a two-dimensional $d$-wave altermagnet with spin-orbit coupling and Zeeman splitting. The system hosts gapped Dirac points at the high-symmetry points $Γ$, $\textrm{M}$, $\textrm{X}$, and $\textrm{Y}$. We show that the JDOS exhibits kinks at the corresponding Dirac gap frequencies and pronounced peaks at Van Hove singularities, whose positions can be tuned by the altermagnetic order. These features are reflected in the optical conductivities, with complementary signatures in their real and imaginary parts. In particular, the Van Hove signatures in the transverse optical conductivity disappear in the absence of $d$-wave altermagnetism, revealing an altermagnet-induced optical signature of the Van Hove singularity. Finally, the Faraday and Kerr rotations exhibit characteristic features inherited from the optical conductivity. Our results establish optical and magneto-optical spectroscopy as sensitive probes of Dirac gaps and Van Hove singularities in altermagnetic topological systems.

cond-mat.mes-hall↗

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

Explorable image-to-scene generation is essential for applications in gaming, robotics, and virtual reality. Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as sparse point clouds or 2D panoramas, leading to stochastic hallucinations, long-term drifts and suboptimal 3D consistency. We present SpatialCrafter, a novel two-stage framework that addresses these issues by introducing a global 3D proxy for high-fidelity image-to-scene generation. Specifically, we decompose the generation process into global proxy generation and appearance refinement. For proxy generation, we propose a Point-anchored Sparse Structure~(PaSS) Flow module that predicts a spatially aligned and geometrically consistent 3D proxy. For appearance refinement, we re-frame the VDM as a Generative Deferred Refiner which synthesizes high-frequency photorealistic details upon proxy-defined scene geometry. To better integrate the proxy with the pre-trained VDM, we introduce Parallel Geometry Injection and Proxy-Aware Corruption training strategies, which improve robustness to proxy artifacts without disrupting the pretrained generative manifold. Furthermore, as no suitable dataset exists for this explorable scene generation task, we construct a new large-scale dataset of 115K scenes. To the best of our knowledge, it is the first hybrid dataset for image-to-scene generation. Extensive experiments on both synthetic and real-world datasets show that SpatialCrafter outperforms state-of-the-art methods, mitigates long-term drift, and remains robust and consistent under rapid camera motion and extreme viewpoint changes. Our project page: \href{https://fangchuan.github.io/SpatialCrafter/}{fangchuan.github.io/SpatialCrafter/}

cs.CV↗

Sharp Lifespan Dichotomies and Threshold Phenomena for Semilinear Heat Equations Driven by the Logarithmic Laplacian

We investigate nonnegative mild solutions of $\partial_t u+(-Δ)^{\ln}u=f(u)$ in $(0,T)\times\mathbb R^N$, with initial datum $u(0,\cdot)=μu_0$, $μ>0$. Unlike the classical and fractional heat semigroups, the positive logarithmic heat kernel exists only for $0<t<N/2$, and the corresponding linear evolution may become singular at its terminal time, with lifespan and growth depending on the spatial decay of $u_0$. The behavior of $f$ near zero determines local solvability: if $\int_{0^+} dσ/f(σ)<\infty$, then no finite nonnegative solution exists on any positive time interval. Under suitable assumptions on $u_0$ and $f$, we establish well-posedness for the nonintegrable logarithmic heat kernel. If $f$ has at most global linear growth, the nonlinear solution attains the full linear lifespan, while the Osgood condition at infinity implies that the maximal existence time tends to zero as $μ\to\infty$. We further distinguish slow-decay, fast-decay, and critical-tail initial data. In the noncritical regimes, a weighted Osgood tail condition yields blow-up strictly before the linear terminal time; if it fails, an amplitude threshold occurs under additional assumptions on $f$. In the critical regime, the dividing power is $3/2$: a square-root weighted Osgood condition yields premature blow-up, while its failure again leads to an amplitude threshold. Finally, we obtain terminal-time blow-up estimates and sharp rates for power nonlinearities.

math.AP↗

Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification

Automated tea leaf disease classification supports precision agriculture, yet deploying accurate models on edge devices remains challenging under tight compute budgets. Self-supervised vision foundation models such as DINOv2 provide strong features but are too large for field deployment, while lightweight models trained from scratch on small agricultural datasets often underfit. We study cross-architecture knowledge distillation (KD) from a fine-tuned DINOv2 teacher (Vision Transformer) to a compact bidirectional Visual State Space Model (LVSSM) student, an underexplored direction because the architectures use fundamentally different token-mixing mechanisms. We identify and fix two training-stability problems that prevent the from-scratch SSM student from learning on limited data: a single large patch-embedding convolution and a fusion layer that severs the residual path. With a progressive convolutional stem and gated bidirectional selective-scan block, the 4.45M-parameter student trains stably. Across three seeds, temperature-scaled logit distillation raises test accuracy from 92.32+/-2.14% to 95.41+/-1.17% (best single run: 96.20%; macro-F1: 94.45%), a +3.09 percentage-point mean gain. The student uses 5.0 times fewer parameters than the 22M-parameter teacher while retaining 98.3% of its accuracy. Ablations show that intermediate feature-alignment losses reduce accuracy, making simple logit-level KD the strongest configuration. A fair from-scratch comparison shows the gain is specific to students that start below the teacher. We report per-class metrics, confusion matrices, bootstrap confidence intervals, and FLOPs/latency measurements, and discuss limitations including the single-dataset scope and simplified non-official SSM implementation.

cs.CV↗

R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models

High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little. This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion. We introduce \emph{R2M-Bench} (\textbf{R}elative \textbf{R}evisit \textbf{M}emory Benchmark), a benchmark of observable revisit-selective consistency. For every detected return, R2M-Bench compares the revisit pair with two controls from the same rollout: a gap-matched non-revisit pair that measures generic temporal stability and a short-range pair that estimates short-horizon consistency. These comparisons produce \emph{MemoryGain} (MG), the revisit advantage over the temporal baseline, and the \emph{Normalized Memory Ratio} (NMR), which normalizes this advantage by the short-to-baseline dynamic range. R2M-Bench combines 100 reference scenes with three leave-and-return trajectories to form 300 instances and evaluates appearance fidelity, scene and object identity, local geometry, and persistent state. Across seven action-conditioned video world models, Overall NMR correlates with human consistency judgments at Spearman's $ρ=0.547$ (95\% CI $[0.45,0.63]$). Its within-model correlation magnitude with generated motion is $0.072$, compared with $0.207$ for raw revisit similarity, indicating that relative calibration substantially reduces the slow-motion shortcut. DreamX-World-Memo achieves the highest Overall NMR among the evaluated video models. Together, these results support same-rollout relative calibration as a practical way to distinguish revisit-specific consistency from generic temporal stability.

cs.CV↗

STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification

Responding to rising global food security needs, precision agriculture and deep learning-based plant disease diagnosis have become crucial. Yet, deploying high-precision models on edge devices is challenging. Most lightweight networks use attention mechanisms designed for generic object recognition, which poorly capture subtle pathological features like irregular lesion shapes and complex textures. To overcome this, we propose a twofold solution: first, using a training-free neural architecture search method (DeepMAD) to create an efficient network backbone for edge devices; second, introducing the Shape-Texture Attention Module (STAM). STAM splits attention into two branches -- one using deformable convolutions (DCNv4) for shape awareness and the other using a Gabor filter bank for texture awareness. On the public CCMT plant disease dataset, our STA-Net model (with 401K parameters and 51.1M FLOPs) reached 89.00% accuracy and an F1 score of 88.96%. Ablation studies confirm STAM significantly improves performance over baseline and standard attention models. Integrating domain knowledge via decoupled attention thus presents a promising path for edge-deployed precision agriculture AI. The source code is available at https://github.com/RzMY/STA-Net.

cs.CV↗

PRISM: Precision and contact-rich Real-world Industrial Skill dataset with Multimodal sensing

Recent progress in robotic learning has been fueled by large-scale datasets collected in everyday environments. However, most existing datasets emphasize short-horizon, low-contact tasks such as pick-and-place, and therefore do not capture the precision control, force/torque or tactile regulation, and multimodal feedback required for industrial assembly. To address this gap, we introduce PRISM, a large-scale multimodal dataset for contact-rich industrial operations. The dataset spans more than 25 manipulation tasks (e.g., electronic components plug/unplug, conveyor-based sorting) and covers diverse mechanical constraints. PRISM includes more than 5,000 trajectories totaling 45 hours of teleoperated demonstrations, recorded using synchronized multi-view RGB-D, force/torque, tactile, and robot-state measurements. In contrast to datasets collected in household or laboratory settings, PRISM provides a realistic benchmark for multimodal perception and control under high-precision industrial constraints, and serves as a foundation for contact-rich, generalizable manipulation in real-world manufacturing environments. The dataset is open-sourced at: https://tengbo-yu.github.io/PRISM/

cs.RO↗

Extending the Constituent Gluon Model to Heavy-Flavour Hybrids: A Unified Study of $c\bar{c}g$ Mesons

We investigate the mass spectra and two-body strong decay properties of ground charmonium hybrids within the framework of a constituent gluon model. Based on the assumption that non-perturbative QCD endows the gluon with an effective mass, we extend the chiral quark model by introducing a single new parameter, the constituent gluon mass $m_g=450$~MeV, which is fixed from previous studies of light hybrids, while other parameters are taken directly from successful descriptions of ordinary meson spectra. We systematically compute the spectra for various quantum numbers and find good agreement with results from lattice QCD, potential models, and other approaches. The corresponding decay widths are also reasonable. For experimental searches, we recommend focusing on the exotic $1^{-+}$ and $2^{+-}$ states, which decay prominently into $D\bar{D}_1$ and $D\bar{D}_2^*$ channels, respectively. Among ordinary quantum numbers, the $0^{-+}$, $2^{-+}$, and $1^{+-}$ states with significant decays into orbitally excited charm mesons are also suggested. Our results provide a unified and consistent description of charmonium hybrids and offer clear guidance for future experimental identification.

hep-ph↗

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm $\mathrm{SE}(3)$ transformations into attention via \textbf{PRoPE-style geometric encoding}, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight \textbf{depth branch} for scene-level geometry and use \textbf{SAM3 masks} with a frozen \textbf{V-JEPA teacher} to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, \model{} achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.

cs.CV↗

High second Chern number induced by long-range hopping in a four-dimensional Dirac model

Four-dimensional (4D) topological systems provide a promising platform for exploring topological phenomena beyond three dimensions. So far, extensive recent studies on 4D topological insulators have focused on the 4D Dirac model, while its second Chern number is restricted to a limited set of values. In this work, we demonstrate that introducing long-range hopping into the 4D Dirac model induces topological phases with high second Chern numbers. Furthermore, we show that the long-range hopping can transform a trivial insulator into a topological insulator with a nonzero second Chern number. Our work establishes long-range hopping as a powerful route for engineering 4D topological states and reveals new possibilities for realizing unconventional topological phases beyond minimal models.

cond-mat.mes-hall↗