SearcharxivSearch

arXiv subjects

Can Huang

Publications and source records attributed to Can Huang.

At least 19 recordsLinked to original sources

Transition of Photonic Dissipative Dynamics through the Exceptional Point

The decay of light in an optical structure depends not only on the intrinsic properties of the material but also on the surrounding electromagnetic environment. This principle has laid the foundation for the engineering of dissipation in photonic emitters. In the conventional wisdom, dissipation is governed by a fixed set of decay channels, each defined by the eigenstates of the structure, and energy leaks through and interacts with these channels. Here we provide experimental evidence that this paradigm fails in non-Hermitian systems. Specifically, we observe an accelerated transient decay in a pair of coupled microcavities tuned near the exceptional point, revealing that photonic dissipation can be governed not by reconfiguring existing loss channels, but rather by restructuring the underlying state space. The universality of this phenomenon is corroborated through two independent control parameters.The finding provides a new perspective on dissipative dynamics in open optical systems and offers a distinct mechanism for controlling transient decay in ultrafast photonic systems.

physics.optics

Rare Earth Ion Coupling Implements Attention-Like Reservoir Computing

We present a physical computing paradigm that harnesses the intrinsic nonlinear dynamics of rare earth doped core shell nanoparticles as a computational substrate. By directly exploiting cross relaxation and energy transfer upconversion processes, the system realizes a state dependent transfer function whose effective decay rate evolves with the instantaneous Er3+ population, which mathematically analogous to gating and attention mechanisms in recurrent neural networks. The three spectrally resolved emission channels inherently span disparate timescales, endowing the reservoir with native multitimescale feature extraction without auxiliary engineering. Under the reservoir computing framework, the coupled three channel system achieves a total memory capacity exceeding fourfold that of a single ion reservoir; capacity decomposition further reveals that the nonzero cross memory capacity is a direct signature of many body Tm3+@Er3+ coupling. On the Mackey Glass and Santa Fe chaotic benchmarks, the system attains normalized mean squared errors of 1.2x10-3 and 2.1x10-2, respectively, with only 125 virtual nodes. These results establish rare earth nanoparticles as a compelling platform for compact and hardware integrable neuromorphic computing, and introduce "inward evolution", the deliberate exploitation of intra material quantum dynamics, as a generalizable design principle for next generation physical computing systems.

physics.optics

From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory

Large language model (LLM) agents are increasingly deployed in long-running settings where improving through experience at test time becomes important. A common approach is to update an explicit memory after each interaction to guide future decisions. However, most existing methods rely on hand-designed prompting rules, making it difficult to align memory updates with downstream objectives over multi-step horizons consistently. We propose MemoPilot, a plug-in memory copilot that explicitly trains the memory update process to improve a frozen LLM's performance across sequential interactions. We formulate memory updating as a multi-turn decision problem and optimize it end-to-end with multi-turn GRPO. Our training recipe introduces (i) a turn-wise reward signal and (ii) a context-independent, turn-level advantage estimation across rollouts, enabling finer-grained credit assignment and more stable training in multi-turn settings. We evaluate MemoPilot on two testbeds: multi-round Rock-Paper-Scissors (RPS) and Limit Texas Hold'em (LHE). Across both environments, MemoPilot substantially improves test-time learning of a frozen player over strong baselines, ranking first in Elo ratings on both games (1762 on LHE and 1590 on RPS) and outperforming all baseline memory methods and proprietary models, including DeepSeek-V3.2.

cs.CL

Bound state in the continuum induced room-temperature superfluorescence

Superfluorescence is a collective emission from several quantum emitters that initially have random phases and are then synchronized through vacuum field interactions. Despite its fascinating prospects in quantum information processing, optical computing and advanced photonic devices, a key challenge in harnessing superfluorescence is alleviating its reliance on cryogenic conditions. Recently, room-temperature superfluorescence has been successfully achieved using upconverted nanoparticles and quasi two-dimensional lead halide perovskites. These approaches, however, are restricted to a few specific material designs and unsuitable for wide promotion. Here, we report a universal strategy to elevate the operating temperature of superfluorescence. We reveal that the symmetry-protected optical bound state in the continuum (BIC) can break the size limitation of superfluorescence ({\lambda}^3) and correlate distant but similar emitters without violating the selection rules, significantly accelerating synchronization process and promoting the possibility of room-temperature superfluorescence. This effect has been experimentally verified using a series of BIC metasurfaces made of different lead halide perovskites. Key features such as the quadratic increase in transient peak intensity and the reduction in pulse width and build-up time at the BIC wavelength confirm the realization of room-temperature superfluorescence that is absent in the pristine material. A theoretical model is also built to explain the experimental observations. This research demonstrates that the operating temperatures of coherent macroscopic states can be effectively improved by artificial field, paving a critical step towards constructing building blocks for optical and quantum applications.

physics.optics

AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards

Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a fundamental challenge: the generation process is text-centric, whereas its quality is governed by visual aesthetics. This modality gap leads current models to frequently produce slides with aesthetically suboptimal layouts. Existing solutions typically rely either on heavy visual reflection, which incurs high inference cost yet yields limited gains; or on fine-tuning with large-scale datasets, which still provides weak and indirect aesthetic supervision. In contrast, the explicit use of aesthetic principles as supervision remains unexplored. In this work, we present AeSlides, a reinforcement learning framework with verifiable rewards for Aesthetic layout supervision in Slide generation. We introduce a suite of meticulously designed verifiable metrics to quantify slide layout quality, capturing key layout issues in an accurate, efficient, and low-cost manner. Leveraging these verifiable metrics, we develop a GRPO-based reinforcement learning method that directly optimizes slide generation models for aesthetically coherent layouts. With only 5K training prompts on GLM-4.7-Flash, AeSlides improves aspect ratio compliance from 36% to 85%, while reducing whitespace by 44%, element collisions by 43%, and visual imbalance by 28%. Human evaluation further shows a substantial improvement in overall quality, increasing scores from 3.31 to 3.56 (+7.6%), outperforming both model-based reward optimization and reflection-based agentic approaches, and even edging out Claude-Sonnet-4.5. These results demonstrate that such a verifiable aesthetic paradigm provides an efficient and scalable approach to aligning slide generation with human aesthetic preferences. Our repository is available at https://github.com/ympan0508/aeslides.

cs.CV

Non-asymptotic uniform in time error bounds for new and old numerical schemes for SPDEs

We study numerical schemes for Stochastic Partial Differential Equations (SPDEs). We introduce a general method of proof of non-asymptotic uniform in time error bounds on numerical integrators for SPDEs, ensuring the schemes capture both the transient and the long term dynamics faithfully. We then consider SPDEs with non-globally Lipshitz nonlinearities, which include for example the stochastic Allen-Cahn equation and some stochastic advection-diffusion equations. For the case of Allen-Cahn type SPDEs we show that the classic semi-implicit Euler time-discretization can exhibit finite time blow up. This motivates analysing other schemes which do not suffer from this blow-up problem. We consider three numerical schemes for SPDEs with non globally Lipshitz nonlinearity: a fully implicit scheme and two tamed schemes. For these schemes we prove non-asymptotic uniform in time error bounds by leveraging our general criterion, and provide numerical comparisons. While the main emphasis in this paper is on the properties of the time-discretization, the schemes we consider are full space-time discretization of the SPDE.

math.NA

Photonic Neuromorphic Computing enabled by a BIC Metasurface

Photonic neuromorphic computing promises revolutionary advances in parallel and high-speed processing, yet a key challenge persists: co-integrating nonlinearity, dense connectivity, and intrinsic memory monolithically to enable brain-inspired, spatiotemporal information processing. Here, we overcome this challenge by introducing a monolithic photonic recurrent network based on an active metasurface operating at bound state in the continuum (BIC). The BIC mode mediates strong,long-range coupling across the lattice, creating a reconfigurable recurrent network topology in hardware. Concurrently, the gain medium provides both optical nonlinearity for neuronal activation and a finite carrier lifetime that serves as a built in, analog temporal memory. This synergy enables computation to emerge directly from the collective spatiotemporal dynamics of the driven-dissipative photonic system, effectively realizing a physical reservoir computer on a chip. We experimentally validate a minimal yet physically complete system on benchmark tasks: brain MRI image classification and human action recognition, achieving 92.16% and 85.36% accuracies, respectively. This work establishes a scalable pathway toward ultrafast, energy-efficient neuromorphic intelligence where processing is an inherent property of tailored light matter interaction.

physics.optics

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

Visual Text Rendering (VTR) remains a critical challenge in text-to-image generation, where even advanced models frequently produce text with structural anomalies such as distortion, blurriness, and misalignment. However, we find that leading MLLMs and specialist OCR models largely fail to perceive these structural anomalies, creating a critical bottleneck for both VTR evaluation and RL-based optimization. As a result, even state-of-the-art generators (e.g., Seedream4.0, Qwen-Image) still struggle to render structurally faithful text. To address this, we propose TextPecker, a plug-and-play structural anomaly perceptive RL strategy that mitigates noisy reward signals and works with any textto-image generator. To enable this capability, we construct a recognition dataset with character-level structural-anomaly annotations and develop a stroke-editing synthesis engine to expand structural-error coverage. Experiments show that TextPecker consistently improves diverse text-to-image models; even on the well-optimized Qwen-Image, it significantly yields average gains of 4% in structural fidelity and 8.7% in semantic alignment for Chinese text rendering, establishing a new state-of-the-art in high-fidelity VTR. Our work fills a gap in VTR optimization, providing a foundational step towards reliable and structural faithful visual text generation.

cs.CV

GLM-5: from Vibe Coding to Agentic Engineering

We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to significantly reduce training and inference costs while maintaining long-context fidelity. To advance model alignment and autonomy, we implement a new asynchronous reinforcement learning infrastructure that drastically improves post-training efficiency by decoupling generation from training. Furthermore, we propose novel asynchronous agent RL algorithms that further improve RL quality, enabling the model to learn from complex, long-horizon interactions more effectively. Through these innovations, GLM-5 achieves state-of-the-art performance on major open benchmarks. Most critically, GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges. Code, models, and more information are available at https://github.com/zai-org/GLM-5.

cs.LG

Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting

Document parsing has garnered widespread attention as vision-language models (VLMs) advance OCR capabilities. However, the field remains fragmented across dozens of specialized models with varying strengths, forcing users to navigate complex model selection and limiting system scalability. Moreover, existing two-stage approaches depend on axis-aligned bounding boxes for layout detection, failing to handle distorted or photographed documents effectively. To this end, we present Dolphin-v2, a two-stage document image parsing model that substantially improves upon the original Dolphin. In the first stage, Dolphin-v2 jointly performs document type classification (digital-born versus photographed) alongside layout analysis. For digital-born documents, it conducts finer-grained element detection with reading order prediction. In the second stage, we employ a hybrid parsing strategy: photographed documents are parsed holistically as complete pages to handle geometric distortions, while digital-born documents undergo element-wise parallel parsing guided by the detected layout anchors, enabling efficient content extraction. Compared with the original Dolphin, Dolphin-v2 introduces several crucial enhancements: (1) robust parsing of photographed documents via holistic page-level understanding, (2) finer-grained element detection (21 categories) with semantic attribute extraction such as author information and document metadata, and (3) code block recognition with indentation preservation, which existing systems typically lack. Comprehensive evaluations are conducted on DocPTBench, OmniDocBench, and our self-constructed RealDoc-160 benchmark. The results demonstrate substantial improvements: +14.78 points overall on the challenging OmniDocBench and 91% error reduction on photographed documents, while maintaining efficient inference through parallel processing.

cs.CV

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

Text and formulas constitute the core informational components of many documents. Accurately and efficiently recognizing both is crucial for developing robust and generalizable document parsing systems. Recently, vision-language models (VLMs) have achieved impressive unified recognition of text and formulas. However, they are large-sized and computationally demanding, restricting their usage in many applications. In this paper, we propose UniRec-0.1B, a unified recognition model with only 0.1B parameters. It is capable of performing text and formula recognition at multiple levels, including characters, words, lines, paragraphs, and documents. To implement this task, we first establish UniRec40M, a large-scale dataset comprises 40 million text, formula and mixed samples, enabling the training of a powerful yet lightweight model. Secondly, we identify two challenges when building such a lightweight but unified expert model. They are: structural variability across levels and semantic entanglement between textual and formulaic content. To tackle these, we introduce a hierarchical supervision training that explicitly guides structural comprehension, and a semantic-decoupled tokenizer that separates text and formula representations. Finally, we develop a comprehensive evaluation benchmark covering Chinese and English documents from multiple domains and with multiple levels. Experimental results on this and public benchmarks demonstrate that UniRec-0.1B outperforms both general-purpose VLMs and leading document parsing expert models, while achieving 2-9x speedup, validating its effectiveness and efficiency. Codebase and Dataset: https://github.com/Topdu/OpenOCR.

cs.CV

ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering

This paper introduces ChineseVideoBench, a pioneering benchmark specifically designed for evaluating Multimodal Large Language Models (MLLMs) in Chinese Video Question Answering. The growing demand for sophisticated video analysis capabilities highlights the critical need for comprehensive, culturally-aware evaluation frameworks. ChineseVideoBench addresses this gap by providing a robust dataset and tailored evaluation metrics, enabling rigorous assessment of state-of-the-art MLLMs on complex Chinese video content. Specifically, ChineseVideoBench comprises 8 main classes and 12 sub-classes, encompassing tasks that demand both deep video understanding and nuanced Chinese linguistic and cultural awareness. Our empirical evaluations reveal that ChineseVideoBench presents a significant challenge to current MLLMs. Among the models assessed, Gemini 2.5 Pro achieves the highest performance with an overall score of 77.9%, while InternVL-38B emerges as the most competitive open-source model.

cs.CV

An efficient fully explicit scheme for stochastic Navier-Stokes equations driven by multiplicative noise

This work proposes an efficient, linear, and fully decoupled pressure-correction scheme for the 2D stochastic Navier-Stokes equations with multiplicative noise and Dirichlet boundary condition. Leveraging the auxiliary variable approach, the scheme is fully explicit yet unconditionally stable. At each time step, it only requires solving Poisson-type equations with constant coefficients. To the best of our knowledge, this is the first application of the auxiliary variable method to stochastic Navier-Stokes equations. We provide a detailed strong convergence analysis for the linearized equation under standard assumptions.

math.NA

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while current Vision-Language Models (VLMs) struggle with their visual and linguistic complexity. Existing document benchmarks focus on English printed texts or simplified Chinese, leaving a gap for evaluating VLMs on ancient Chinese documents. To address this, we present AncientDoc, the first benchmark for Chinese ancient documents, designed to assess VLMs from OCR to knowledge reasoning. AncientDoc includes five tasks (page-level OCR, vernacular translation, reasoning-based QA, knowledge-based QA, linguistic variant QA) and covers 14 document types, over 100 books, and about 3,000 pages. Based on AncientDoc, we evaluate mainstream VLMs using multiple metrics, supplemented by a human-aligned large language model for scoring.

cs.CL

Compact All optical Reservoir Computing via Luminescence Dynamics in Rare-earth Ions-doped Nanocrystals

Optical neuromorphic computing offers a promising route to high speed, energy efficient information processing. However, photonic neurons, as the critical components for enhancing computational expressivity, still face significant bottlenecks in nonlinear mapping and memory capacity. Here, we demonstrate an all optical reservoir computing system based on rare earth ions doped nanocrystals for the first time, leveraging their intrinsic nonlinear luminescence dynamics and multitimescale memory. Unlike traditional schemes that require bulky optical delays or intricate resonant structures, our platform exploits the material's inherent properties: nonlinear cross-relaxation processes enable nonlinear mapping while millisecond-scale metastable energy levels provide fading memory. As a proof of concept, we achieve 90.7% accuracy in MNIST digit classification and low-error chaotic time-series prediction (NRMSE < 0.1) using the rare-earth ions based system. Our work significantly reduce system footprint and complexity, offering a scalable, fully optical solution for edge computing and real-time neuromorphic applications.

physics.optics

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this, we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model's performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods.

cs.AI

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning, GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks. We release both GLM-4.5 (355B parameters) and a compact version, GLM-4.5-Air (106B parameters), to advance research in reasoning and agentic AI systems. Code, models, and more information are available at https://github.com/zai-org/GLM-4.5.

cs.CL

Post-Completion Learning for Language Models

Current language model training paradigms typically terminate learning upon reaching the end-of-sequence ( ) token, overlooking the potential learning opportunities in the post-completion space. We propose Post-Completion Learning (PCL), a novel training framework that systematically utilizes the sequence space after model output completion, to enhance both the reasoning and self-evaluation abilities. PCL enables models to continue generating self-assessments and reward predictions during training, while maintaining efficient inference by stopping at the completion point. To fully utilize this post-completion space, we design a white-box reinforcement learning method: let the model evaluate the output content according to the reward rules, then calculate and align the score with the reward functions for supervision. We implement dual-track SFT to optimize both reasoning and evaluation capabilities, and mixed it with RL training to achieve multi-objective hybrid optimization. Experimental results on different datasets and models demonstrate consistent improvements over traditional SFT and RL methods. Our method provides a new technical path for language model training that enhances output quality while preserving deployment efficiency.

cs.CL