SearcharxivSearch

arXiv subjects

Rong Wang

Publications and source records attributed to Rong Wang.

At least 19 recordsLinked to original sources

Mirror Chern insulators in two-dimensional altermagnetic Tc$_2$Cl$_2$O and Tc$_2$Br$_2$O

The interplay between altermagnetism and crystalline band topology provides an intriguing avenue for realizing unconventional topological phases with distinctive spin-dependent properties. Here, based on first-principles calculations and theoretical analysis, we identify monolayer $\mathrm{Tc}_2X_2\mathrm{O}$ ($X$ = Cl, Br) as a family of two-dimensional altermagnetic mirror Chern insulators. In the absence of spin--orbit coupling (SOC), both monolayers exhibit robust altermagnetism with mirror-spin coupling and host two symmetry-protected Weyl points in each spin channel near the Fermi level. The Weyl points in opposite spin channels carry distinct mirror-symmetry eigenvalues, $m_z=\pm i$. Upon inclusion of SOC, the Weyl points are gapped, and the two mirror sectors acquire opposite Chern numbers, ${\cal {C}}_{+}=1$ and ${\cal {C}}_{-}=-1$, resulting in a nonzero mirror Chern number ${\cal {C}}_m=1$. A low-energy $k\cdot p$ model captures the symmetry protection of the Weyl points and elucidates their SOC-induced mass gaps and topological character. Furthermore, the resulting mirror Chern insulating phases host helical edge states within the bulk band gap and exhibit a quantized spin Hall conductivity. Our work establishes a direct connection between altermagnetism and mirror Chern topology and provides a promising platform for exploring unconventional topological and spin-dependent phenomena in two-dimensional altermagnetic materials.

cond-mat.mtrl-sci

The 10th AI City Challenge

The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with vehicle detection, classification, and tracking, the challenge has grown into a broad benchmark suite for multi-camera perception, multimodal reasoning, synthetic-to-real learning, generative forecasting, and privacy-preserving evaluation. The 2026 edition continued this growth with 325 registered teams, up from 245 in 2025, and participation from 26 countries and regions, up from 15. Its six primary tracks cover multi-camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. Track 3 further includes two out-of-domain leaderboards, submitted as Tracks 7 and 8, for fisheye traffic-violation understanding and pedestrian situated-intent VQA. This paper summarizes the challenge setup, datasets, evaluation protocols, leaderboard results, and workshop papers. Across tracks, successful systems combine foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference.

cs.CV

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-level interventions can conflate one rollout's unreliability with low expected training value of its prompt. We define prompt-level teacher continuation reliability $R$ as the teacher's probability of reaching a correct answer from a student prefix, averaged over prefixes and trajectories induced by the current student. Oracle experiments show that high-$R$ prompts yield larger OPD gains and that descending-$R$ training outperforms random and ascending orders on a fixed prompt pool. Because estimating $R$ requires many teacher continuations, we use the maximum ROUGE-5 F1 between one independent student rollout and verifier-correct same-prompt teacher trajectories. Across ten equal-frequency bins of this actual score, mean $R$ rises monotonically, showing that the proxy separates coarse reliability levels. ReOrder-OPD sorts prompts by the proxy, then draws independent on-policy training trajectories for vanilla OPD. It improves every matched aggregate comparison across Qwen3 and Gemma4 mathematics settings and Qwen3 code settings. Gains in all six FiRe-OPD and ExOPD settings show that prompt ordering complements within-trajectory supervision.

cs.LG

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech

Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text. Existing mitigation mainly relies on architectural changes or additional training, while decoding-time control remains underexplored. We present a conditional information view that distinguishes text-derived alignment information from experience information supplied by acoustic context and learned speech regularities. We hypothesize that an important class of hallucinations begins when alignment support is insufficiently reflected in the selected token at a vulnerable transition. Using predictions from the same speech LM with and without text conditions, we propose Experience-Calibrated Contrastive Decoding (ECCD), a training-free method that strengthens alignment support while preserving useful experience information. ECCD preserves the original expert distribution, applies only positive alignment enhancement, and calibrates its strength using set-level experience compatibility. Across four models, ECCD reduces WER/CER by up to 55.6% in all SeedTTS-Eval settings and 24 of 25 multilingual CV3-Eval settings. A listening test yields a CMOS gain of $+0.644$ while retaining strong speaker similarity. Further analysis shows that alignment influence and decision-level gain vary within linguistic units and are lower at first-error boundaries than at matched correct boundaries. Overall, these extensive experiments and analyses identify conditional information control as a promising decoding-time direction for mitigating speech hallucination.

eess.AS

Contextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech Comprehension

Naturalistic language comprehension requires listeners to process both local probabilistic expectations and contextual semantic relations. This study tested whether contextual semantic relevance, measuring how strongly a target word relates to its recent semantic context, is associated with fMRI BOLD responses independently of word surprisal and lexical, timing, acoustic, and prosodic controls. We analyzed two public datasets: Alice (23 participants, one narrative) and Narratives (47 participants, 185 runs, four stories) using FIR/deconvolution and generalized additive mixed models. In Alice, semantic relevance was significant across all ROIs in FIR analyses, whereas surprisal was not. In GAMMs, both predictors showed broad significance. In Narratives, both predictors showed comparable spatial prevalence across ROIs. Semantic relevance showed robust BOLD associations across both datasets, with a particularly strong advantage over surprisal in the timing-sensitive Alice FIR analysis. The regionally heterogeneous direction of semantic relevance effects, with negative effects in posterior semantic regions and positive effects in frontal integration regions, suggests involvement of functionally distinct neural processes rather than a single uniform mechanism. These findings indicate that contextual semantic fit and local probabilistic expectation make partially distinct, dataset-dependent contributions to hemodynamic responses during naturalistic listening.

cs.CL

Contextual Semantic Relevance and Word Surprisal Predict N400 and P600 Dynamics During Naturalistic Reading

Word surprisal is a well-established computational predictor of human neural responses during language comprehension, but it remains less clear whether local semantic fit explains neural response variation beyond lexical expectation during naturalistic reading. Using the Dublin EEG-based Reading Experiment Corpus (DERCo), this study examined whether contextual semantic relevance predicts word-locked EEG activity in the N400 and P600 windows. Contextual semantic relevance was computed as an attention-aware measure of how strongly a target word is semantically connected to its recent discourse context, and it was compared with GPT-based word surprisal. Across 22 participants and 32 EEG channels, we tested both predictors using regression-based ERP analyses and generalized additive mixed models while controlling for lexical variables and repeated observations. Both predictors were reliably associated with EEG responses, but they showed partly different temporal and scalp-level patterns. Surprisal captured expectancy-related variation, whereas contextual semantic relevance showed robust effects across N400- and P600-window mean voltages, with particularly strong explanatory support in the P600 window. Model comparisons indicated that contextual semantic relevance contributed explanatory value beyond lexical controls and surprisal. These findings suggest that naturalistic reading depends on both lexical expectation and local semantic integration, and that contextual semantic relevance offers an interpretable computational link between discourse semantic fit and ERP dynamics.

cs.CL

From Gentlemen to Frontiermen: Masculine Formations in English-Language Fiction (1771--1930)

Masculinity in nineteenth-century fiction is not a single ideal but a field of competing scripts. Drawing on 150 British and American canonical novels from the txtLAB Novel450 corpus, published between 1771 and 1930, this paper examines the changing relative prominence of competing models of masculine authority. To focus the analysis on masculine characterisation, the study extracts male-character-centred text windows by using coreference resolution to group names, nominal mentions, and pronouns into character-specific reference chains. It then fits an unsupervised structural topic model with publication year and author gender as topic-prevalence covariates. The model identifies six distinct masculine formations: aristocratic-chivalric, Christian manhood, gentlemanly respectability, country squire, professional-commercial, and imperial/adventure. Across the corpus, formations tied to inherited rank and sacred authority decline, while those organised around paid work and adventure rise. The largest increase occurs not in professional-commercial breadwinning but in imperial/adventure masculinity, particularly the frontier-wilderness register. The trajectory points to a reallocation from inherited and sacred status towards achieved, commercial, and expansionary forms of masculine authority. Adventurous and commercial formations are also more prevalent in novels by authors recorded as male. Because these formations emerge without a seeded vocabulary yet align with categories established in independent scholarship, the article offers a reproducible method for measuring the reorganisation of gendered authority across the long nineteenth century.

cs.CL

GKDT: General Keypoint Detection Transformer

With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful yet general keypoint detection (GKD) model to support diverse tasks has become increasingly important in the field. To this end, we firstly present a large-scale unified keypoint dataset called MegaKPT. The dataset is composed of over 1.3 million diverse object instances from twenty-nine existing datasets, and enjoys high-quality unified annotations with keypoint text descriptions. Based on MegaKPT, we develop GKDT, a simple, flexible and powerful DINOv3 based Transformer model for General Keypoint Detection. Our GKDT supports visual prompts, text prompts, or both. To enhance model training, we also propose a suite of useful strategies such as mix-modal prompted training and dynamic importance sampling. By testing over 22 test sets with seen or unseen objects, our single GKDT model shows strong performance and generality in detecting keypoints on broad categories, with most categories over 90\% PCK@0.1 accuracy, offering high practical applicability to real-world problems. The dataset, models, and codes will be released at https://github.com/AlanLuSun/General-Keypoint-Detection.

cs.CV

Prior over Evidence: Stereotype-Driven Diagnosis in LLM-Based L2 Pronunciation Feedback

Large language models are increasingly deployed for written pronunciation feedback in second-language (L2) English learning, under the assumption that their diagnoses are grounded in the supplied speech evidence rather than in priors from pretraining. This assumption is tested on 1,800 L2-Arctic utterances spanning six L1 backgrounds, three audio-capable LLMs, four pronunciation dimensions, and five evidence conditions ranging from a text-only baseline to numeric acoustic features and raw audio. Each (utterance x model x condition x dimension) cell is scored on three metrics: Rating Accuracy (RA) against gold labels, Evidence Coherence (EC) assessing internal consistency without ground truth, and Grounded Correctness (GC) evaluated against gold evidence. Results show three findings across models. First, rating accuracy and grounded reasoning decouple: 39.6% of judged cells contain internally coherent reasoning that supports a wrong rating, against only 15.8% where the reasoning supports a correct rating. Second, phoneme-level feedback converges to a fixed inventory of L2-English difficulty phones that recurs across all six L1 backgrounds and all evidence conditions. Third, acoustic evidence improves the rating only when the supplied feature directly probes the target dimension: textualised F0 range raises pitch-variation grounding from (0.18-0.19) to (0.45-0.62) across all three models, while stress and phoneme correctness, which require target-to-realisation alignment, remain ungrounded. The same audio waveform without textualised F0 values does not reproduce this improvement. These findings indicate that current general-purpose LLMs are more reliable as verbalisers of externally computed pronunciation evidence than as standalone diagnostic engines.

cs.CL

Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generation works mostly rely on pixel-wise attention to enforce view consistency across frames. However, such mechanism does not impose sufficient constraints for long-range extrapolation, e.g. rear-view synthesis, in which pixel correspondences to the input image are limited. Consequently, these works often fail to produce results with a plausible and coherent structure. To tackle this issue, we propose to leverage rich shape priors from a 3D foundational generative model as an auxiliary constraint, motivated by its capability of modeling realistic object shape distributions learned from large 3D asset corpora. Specifically, we prompt the video generation with two scales of latent features encoded by the 3D foundation model: (i) a denoised global latent vector as an overall structural guidance, and (ii) a set of latent images projected from volumetric features to provide view-dependent and fine-grained geometry details. In contrast to commonly used 2.5D representations such as depth or normal maps, these compact features can model complete object shapes, and help to improve inference efficiency by avoiding explicit mesh extraction. To achieve effective shape conditioning, we introduce a multi-scale 3D adapter to inject feature tokens to the base video model via cross-attention, which retains its capabilities from general video pretraining and enables a simple and model-agonistic fine-tuning process. Extensive experiments on multiple benchmarks show that our method achieves superior visual quality, shape realism and multi-view consistency compared to state-of-the-art methods, and robustly generalizes to complex camera trajectories and in-the-wild images.

cs.CV

Fully-Passive Twin-Field Quantum Key Distribution

We propose a fully passive twin-field quantum key distribution (QKD) setup where basis choice, decoy-state preparation and encoding are all implemented entirely by post-processing without any active modulation. Our protocol can remove the potential side-channels from both source modulators and detectors, and additionally retain the high key rate advantage offered by twin-field QKD, thus offering great implementation security and good performance. Importantly, we also propose a post-processing strategy that uses mismatched phase slices and minimizes the effect of sifting. We show with numerical simulation that the new protocol can still beat the repeaterless bound and provide satisfactory key rate.

quant-ph

FSMC-Pose: Frequency and Spatial Fusion with Multiscale Self-calibration for Cattle Mounting Pose Estimation

Mounting posture is an important visual indicator of estrus in dairy cattle. However, achieving reliable mounting pose estimation in real-world environments remains challenging due to cluttered backgrounds and frequent inter-animal occlusion. We present FSMC-Pose, a top-down framework that integrates a lightweight frequency-spatial fusion backbone, CattleMountNet, and a multiscale self-calibration head, SC2Head. Specifically, we design two algorithmic components for CattleMountNet: the Spatial Frequency Enhancement Block (SFEBlock) and the Receptive Aggregation Block (RABlock). SFEBlock separates cattle from cluttered backgrounds, while RABlock captures multiscale contextual information. The Spatial-Channel Self-Calibration Head (SC2Head) attends to spatial and channel dependencies and introduces a self-calibration branch to mitigate structural misalignment under inter-animal overlap. We construct a mounting dataset, MOUNT-Cattle, covering 1176 mounting instances, which follows the COCO format and supports drop-in training across pose estimation models. Using a comprehensive dataset that combines MOUNT-Cattle with the public NWAFU-Cattle dataset, FSMC-Pose achieves higher accuracy than strong baselines, with markedly lower computational and parameter costs, while maintaining real-time inference on commodity GPUs. Extensive experiments and qualitative analyses show that FSMC-Pose effectively captures and estimates cattle mounting pose in complex and cluttered environments. Dataset and code are available at https://github.com/elianafang/FSMC-Pose.

cs.CV

AI assisted optimization of integrated waveguide polarizers containing 2D reduced graphene oxide

Reduced graphene oxide (rGO) exhibits strong anisotropic light absorption and high compatibility with photonic integrated chips, making it a promising material for implementing high performance onchip polarization selective devices. The performance of rGO integrated waveguide polarizers is highly dependent on the waveguide geometry, and achieving optimal performance requires exploring a large parameter space, making conventional mode simulation methods computationally demanding. Here, we propose and demonstrate a machine learning framework based on fully connected neural networks (FCNNs) to map the dependence of the polarizer figure of merit (FOM) on the waveguide geometry. Once trained by using a small dataset of low resolution mode simulation results, the FCNN framework can rapidly and accurately predict FOM values across a large structural parameter space with high resolution. Results show that this method can reduce overall computing time by more than 4 orders of magnitude as compared to the mode simulation methods, and achieve high prediction accuracy with an average deviation (AD) below 0.05. These results highlight the FCNN based machine learning framework as an efficient tool for the design and optimization of rGO integrated waveguide polarizers.

physics.optics

AI based design of 2D material integrated optical polarizers

On-chip integration of highly anisotropic two-dimensional (2D) materials offers new opportunities for realizing high performance polarization selective devices. Obtaining optimized designs for such devices requires extensively sweeping large parameter spaces, which in conventional approaches relies on massive mode simulations that demand considerable computational resources. Here, we address this limitation by developing a machine learning (ML) model based on fully connected neural networks (FCNNs). Trained by using mode simulation results for low resolution structural parameters, the FCNN model can accurately predict polarizer figures of merits (FOMs) for high resolution parameters and rapidly map the global variation trend across the entire parameter space. We test the performance of the FCNN model using two types of polarizers with 2D graphene oxide (GO) and molybdenum disulfide (MoS2). Results show that, compared to conventional mode simulation approach, our approach can not only reduce the overall computing time by about 4 orders of magnitude, but also achieve highly accurate FOM predictions with an average deviation of less than 0.04. In addition, the measured FOM values for the fabricated devices show good agreement with the predicted ones, with discrepancies remaining below 0.2. These results validate artificial intelligence (AI) as an effective approach for designing and optimizing 2D-material based optical polarizers with high efficiency.

physics.optics

Dynamics of a nonlocal epidemic model with a new free boundary condition, part 1: Spreading-vanishing dichotomy

This paper investigates the long-time dynamics of a nonlocal epidemic model with free boundaries, where a pathogen with density $u(t,x)$ and the infected humans with density $v(t,x)$ evolve according to a reaction-diffusion system with nonlocal diffusion over a one dimensional interval $[g(t), h(t)]$, which represents the epidemic region expanding through its boundaries $x=g(t)$ and $x=h(t)$, known as free boundaries. Such a model with free boundary conditions based on those of Cao et al. \cite{fb27} was considered by several works. Inspired by recent works of Feng et al. \cite{fb20} and Long et al. \cite{fb5}, we propose a new free boundary condition, where the expansion rate of the epidemic region, determined by $h'(t)$ and $g'(t)$, is proportional to a linear combination of the outward flux of the pathogen \(u\) through the range boundary (as in \cite{fb27}) and the weighted total population of infected individuals \(v\) within the region (as in \cite{fb5}). We prove that the system under this new free boundary condition is well-posed, and its long-time dynamical behavior is characterized by a spreading-vanishing dichotomy. Moreover, we obtain sharp criteria for this dichotomy, including a sharp threshold in terms of the initial data $(u_0,v_0)$; and by studying a related eigenvalue problem, we also find a sharp threshold in terms of the diffusion rate, which complements related results in Nguyen and Vo \cite{fb7}. This is Part $1$ of a two part series. In Part $2$, we will determine the spreading speed of the model when spreading occurs, and for some typical classes of kernel functions, we will obtain the precise rates of accelerated spreading.

math.AP

Certifying optimal device-independent quantum randomness in quantum networks

Bell nonlocality provides a device-independent (DI) way to certify quantum randomness, based on which true random numbers can be extracted from the observed correlations without detail characterizations on devices for quantum state preparation and measurement. However, the efficiency of current strategies for DI randomness certification is still heavily constrained when it comes to non-maximal Bell values, especially for multiple parties. Here, we present a family of multipartite Bell inequalities that allows to certify optimal quantum randomness and self-test GHZ (Greenberger-Horne-Zeilinger) states, which are inspired from the stabilizer group of the GHZ state. Due to the simple representation of stabilizer group for GHZ states, this family of Bell inequalities is of simple structure and can be easily expanded to more parties. Compared with the Mermin-type inequalities, this family of Bell inequality is more efficient in certifying quantum randomness when non-maximal Bell values achieved. Meanwhile, the general analytical upper bound for the Holevo quantity is presented, and achieves better performance compared with the MABK (Mermin-Ardehali-Belinskii-Klyshko) inequality, Parity-CHSH (Clauser-Horne-Shimony-Holt) inequality and Holz inequality at $N=3$, which is of particular interests for experimental researches on DI quantum cryptography in quantum networks.

quant-ph

Multi-class Support Vector Machine with Maximizing Minimum Margin

Support Vector Machine (SVM) stands out as a prominent machine learning technique widely applied in practical pattern recognition tasks. It achieves binary classification by maximizing the "margin", which represents the minimum distance between instances and the decision boundary. Although many efforts have been dedicated to expanding SVM for multi-class case through strategies such as one versus one and one versus the rest, satisfactory solutions remain to be developed. In this paper, we propose a novel method for multi-class SVM that incorporates pairwise class loss considerations and maximizes the minimum margin. Adhering to this concept, we embrace a new formulation that imparts heightened flexibility to multi-class SVM. Furthermore, the correlations between the proposed method and multiple forms of multi-class SVM are analyzed. The proposed regularizer, akin to the concept of "margin", can serve as a seamless enhancement over the softmax in deep learning, providing guidance for network parameter learning. Empirical evaluations demonstrate the effectiveness and superiority of our proposed method over existing multi-classification methods.

cs.LG

Data-driven inference of brain dynamical states from the r-spectrum of correlation matrices

We present a data-driven framework to characterize large-scale brain dynamical states directly from correlation matrices at the single-subject level. By treating correlation thresholding as a percolation-like probe of connectivity, the approach tracks multiple cluster- and network-level observables and identifies a characteristic percolation threshold, rc, at which these signatures converge. We use $r_c$ as an operational and physically interpretable descriptor of large-scale brain dynamical state. Applied to resting-state fMRI data from a large cohort of healthy individuals (N = 996), the method yields stable, subject-specific estimates that covary systematically with established dynamical indicators such as temporal autocorrelations. Numerical simulations of a whole-brain model with a known critical regime further show that $r_c$ tracks changes in collective dynamics under controlled variations of excitability. By replacing arbitrary threshold selection with a criterion intrinsic to correlation structure, the r-spectra provides a physically grounded approach for comparing brain dynamical states across individuals.

q-bio.NC