SearcharxivSearch

arXiv subjects

Jingyuan Zhao

Publications and source records attributed to Jingyuan Zhao.

18 recordsLinked to original sources

Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents

Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-the-box capability for extracting the key-values, but their verbalized confidence signals are unreliable and weakly track field correctness. This paper introduces a decomposed confidence layer along three interpretable channels, including perception, layout, and validation. Together with a final conformal risk control, the score can be used for reliable STP of financial documents. The method is validated on three public datasets covering real invoices, synthetic invoices, and ad-buy forms, using two different VLM families (Qwen3.6-27B and Gemini-3.1-Flash-Lite). Our decomposed score consistently improves the separation of correct from incorrect extractions, substantially raising the AUROC from 0.54-0.74 for VLM verbalized signals to 0.90-0.99 with contributions from all three designed channels. Crucially for industrial deployment, this enables usable STP. The native VLM confidence signals could clear only 0.1%-7.0% of fields under risk control at a target error of <10%. In contrast, the proposed method auto-approves 49-72% of fields while holding the empirical error of the accepted tier at or below the target.

cs.AI

Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification

Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data. We present an end-to-end pipeline for customer-level mule detection comprising three stages: (1) a LightGBM classifier trained on 280 engineered features spanning transaction patterns, account demographics, network topology, and temporal behaviour; (2) a TreeSHAP attribution layer that decomposes each prediction into feature contributions; and (3) a large language model (LLM) module that converts SHAP attributions into analyst-facing natural-language narratives. We evaluate across three open-weight LLM families and assess explanation quality through analyst feedback. In a live production deployment, the system achieves a yield rate of 89%, up from 61% under the incumbent rule-based system, with monthly alert volume expanding from 211 to 302, reflecting broader true-positive coverage rather than increased noise. This corresponds to a 60% incremental adverse detection beyond existing review workflows, substantially outperforming the rule-based approach. Qualitative feedback from analysts indicates that LLM-generated narratives reduce cognitive load during alert triage. We further discuss implications of deploying LLM-augmented explainability in regulated financial environments.

cs.CR

Weak-Strong Steady-State Microbunching Accelerator Light Source

We propose a phase space manipulation involving one energy modulation sandwiched by two dispersion sections which converts a bunched particle beam or bunch train to ultra-high-harmonic density modulation, while the energy modulation in principle can be arbitrarily weak. The same scheme can also be used for energy bunching, creating energy levels in a bunched beam. We further propose a mechanism invoking three laser modulators in a storage ring to longitudinally focus the electron beam both weakly and strongly, such that a microbunch train and its high-density-harmonics or energy bunching form and sustain turn-by-turn. We call this mechanism weak-strong steady-state microbunching (Weak-Strong SSMB). The longitudinal beta function can vary by seven orders of magnitude along such a ring, with the minimal value squeezed to 10 nm. An example application of Weak-Strong SSMB for kW coherent EUV radiation is presented. Extension to X-ray can be anticipated. An energy-leveled electron beam enables $γ$-ray frequency comb production. The ideas can be scaled to wavelengths like RF and THz, for bunch length and energy spread control, ultrashort X-ray and coherent THz generation. Our work establishes a new paradigm for longitudinal dynamics study, accelerator light source development, and opens great potential for accelerator physics and technology.

physics.acc-ph

Echo Enhanced Strong Focusing for Coherent Short-Wavelength Radiation

Storage-ring-based fully coherent light sources, including steady-state microbunching (SSMB), as well as compact seeded FELs driven by laser plasma accelerators, typically have relatively large intrinsic energy spreads. Extending the spectral reach of these facilities toward the X-ray regime represents a major challenge, as existing seeded schemes require rather extreme parameters to generate appreciable microbunching at high harmonics. In this Letter, we propose an echo enhanced strong focusing scheme that employs transverse-longitudinal coupling together with the beam echo effect to simultaneously resolve the energy spread bottleneck and enable efficient high-harmonic generation. This approach substantially relaxes the requirements on both the intrinsic energy spread and the transverse emittance, paving the way for soft X-ray production using relatively weak laser modulation. Based on this scheme, we further present an SSMB storage ring capable of generating kW-level average power 6.7 nm soft X-ray radiation.

physics.acc-ph

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows

Structured information extraction from long, multilingual scanned financial documents is a core requirement in industrial KYC and compliance workflows. These documents are typically non machine readable, noisy, and visually heterogeneous. They usually span dozens of pages while containing only sparse task relevant information. Although recent vision-language models achieve strong benchmark performance, directly applying them end to end to full financial reports often leads to unreliable extraction under real world conditions. We present a multistage extraction framework that integrates image preprocessing, multilingual OCR, hybrid page-level retrieval, and compact VLM-based structured extraction. The design separates page localization from multimodal reasoning, enabling more accurate extraction from complex multipage documents. We evaluated the framework on 120 production KYC documents comprising about 3000 multilingual scanned pages. Across multiple OCR-VLM combinations, the proposed pipeline consistently outperforms direct PDF-to-VLM baselines, improving field-level accuracy by up to 31.9 percentage points. The best configuration, PaddleOCR with MiniCPM2.6, achieves 87.27 percent accuracy. Ablation studies show that page-level retrieval is the dominant factor in performance improvements, particularly for complex financial statements and non-English documents.

cs.CV

PsychAgent: An Experience-Driven Lifelong Learning Agent for Self-Evolving Psychological Counselor

Existing methods for AI psychological counselors predominantly rely on supervised fine-tuning using static dialogue datasets. However, this contrasts with human experts, who continuously refine their proficiency through clinical practice and accumulated experience. To bridge this gap, we propose an Experience-Driven Lifelong Learning Agent (\texttt{PsychAgent}) for psychological counseling. First, we establish a Memory-Augmented Planning Engine tailored for longitudinal multi-session interactions, which ensures therapeutic continuity through persistent memory and strategic planning. Second, to support self-evolution, we design a Skill Evolution Engine that extracts new practice-grounded skills from historical counseling trajectories. Finally, we introduce a Reinforced Internalization Engine that integrates the evolved skills into the model via rejection fine-tuning, aiming to improve performance across diverse scenarios. Comparative analysis shows that our approach achieves higher scores than strong general LLMs (e.g., GPT-5.4, Gemini-3) and domain-specific baselines across all reported evaluation dimensions. These results suggest that lifelong learning can improve the consistency and overall quality of multi-session counseling responses.

cs.AI

PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor

To develop a reliable AI for psychological assessment, we introduce \texttt{PsychEval}, a multi-session, multi-therapy, and highly realistic benchmark designed to address three key challenges: \textbf{1) Can we train a highly realistic AI counselor?} Realistic counseling is a longitudinal task requiring sustained memory and dynamic goal tracking. We propose a multi-session benchmark (spanning 6-10 sessions across three distinct stages) that demands critical capabilities such as memory continuity, adaptive reasoning, and longitudinal planning. The dataset is annotated with extensive professional skills, comprising over 677 meta-skills and 4577 atomic skills. \textbf{2) How to train a multi-therapy AI counselor?} While existing models often focus on a single therapy, complex cases frequently require flexible strategies among various therapies. We construct a diverse dataset covering five therapeutic modalities (Psychodynamic, Behaviorism, CBT, Humanistic Existentialist, and Postmodernist) alongside an integrative therapy with a unified three-stage clinical framework across six core psychological topics. \textbf{3) How to systematically evaluate an AI counselor?} We establish a holistic evaluation framework with 18 therapy-specific and therapy-shared metrics across Client-Level and Counselor-Level dimensions. To support this, we also construct over 2,000 diverse client profiles. Extensive experimental analysis fully validates the superior quality and clinical fidelity of our dataset. Crucially, \texttt{PsychEval} transcends static benchmarking to serve as a high-fidelity reinforcement learning environment that enables the self-evolutionary training of clinically responsible and adaptive AI counselors.

cs.AI

Merging Physics-Based Synthetic Data and Machine Learning for Thermal Monitoring of Lithium-ion Batteries: The Role of Data Fidelity

Since the internal temperature is less accessible than surface temperature, there is an urgent need to develop accurate and real-time estimation algorithms for better thermal management and safety. This work presents a novel framework for resource-efficient and scalable development of accurate, robust, and adaptive internal temperature estimation algorithms by blending physics-based modeling with machine learning, in order to address the key challenges in data collection, model parameterization, and estimator design that traditionally hinder both approaches. In this framework, a physics-based model is leveraged to generate simulation data that includes different operating scenarios by sweeping the model parameters and input profiles. Such a cheap simulation dataset can be used to pre-train the machine learning algorithm to capture the underlying mapping relationship. To bridge the simulation-to-reality gap resulting from imperfect modeling, transfer learning with unsupervised domain adaptation is applied to fine-tune the pre-trained machine learning model, by using limited operational data (without internal temperature values) from target batteries. The proposed framework is validated under different operating conditions and across multiple cylindrical batteries with convective air cooling, achieving a root mean square error of 0.5 °C when relying solely on prior knowledge of battery thermal properties, and less than 0.1 °C when using thermal parameters close to the ground truth. Furthermore, the role of the simulation data quality in the proposed framework has been comprehensively investigated to identify promising ways of synthetic data generation to guarantee the performance of the machine learning model.

eess.SY

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

The proliferation of multi-modal fake news on social media poses a significant threat to public trust and social stability. Traditional detection methods, primarily text-based, often fall short due to the deceptive interplay between misleading text and images. While Large Vision-Language Models (LVLMs) offer promising avenues for multi-modal understanding, effectively fusing diverse modal information, especially when their importance is imbalanced or contradictory, remains a critical challenge. This paper introduces MM-FusionNet, an innovative framework leveraging LVLMs for robust multi-modal fake news detection. Our core contribution is the Context-Aware Dynamic Fusion Module (CADFM), which employs bi-directional cross-modal attention and a novel dynamic modal gating network. This mechanism adaptively learns and assigns importance weights to textual and visual features based on their contextual relevance, enabling intelligent prioritization of information. Evaluated on the large-scale Multi-modal Fake News Dataset (LMFND) comprising 80,000 samples, MM-FusionNet achieves a state-of-the-art F1-score of 0.938, surpassing existing multi-modal baselines by approximately 0.5% and significantly outperforming single-modal approaches. Further analysis demonstrates the model's dynamic weighting capabilities, its robustness to modality perturbations, and performance remarkably close to human-level, underscoring its practical efficacy and interpretability for real-world fake news detection.

cs.CR

Discovery of Two New Eruptions of the Ultrashort Recurrence Time Nova M31N 2017-01e

We report the recent discovery of two new eruptions of the recurrent nova M31N 2017-01e in the Andromeda galaxy. The latest eruption, M31N 2024-08c, reached $R=17.8$ on 2024 August 06.85 UT, $\sim2$ months earlier than predicted. In addition to this recent eruption, a search of archival PTF data has revealed a previously unreported eruption on 2014 June 18.46 UT that reached a peak brightness of $R\sim17.9$ approximately a day later. The addition of these two eruption timings has allowed us to update the mean recurrence time of the nova. We find $\langle T_\mathrm{rec} \rangle = 924.0\pm7.0$ days ($2.53\pm0.02$ yr), which is slightly shorter than our previous determination. Thus, M31N 2017-01e remains the nova with the second shortest recurrence time known, with only M31N 2008-12a being shorter. We also present a low-resolution spectrum of the likely quiescent counterpart of the nova, a $\sim20.5$ mag evolved B star displaying an $\sim14.3$ d photometric modulation.

astro-ph.SR

A Shock Flash Breaking Out of a Dusty Red Supergiant

Shock breakout emission is light that arises when a shockwave, generated by core-collapse explosion of a massive star, passes through its outer envelope. Hitherto, the earliest detection of such a signal was at several hours after the explosion, though a few others had been reported. The temporal evolution of early light curves should reveal insights into the shock propagation, including explosion asymmetry and environment in the vicinity, but this has been hampered by the lack of multiwavelength observations. Here we report the instant multiband observations of a type II supernova (SN 2023ixf) in the galaxy M101 (at a distance of 6.85+/-0.15 Mpc), beginning at about 1.4 hours after the explosion. The exploding star was a red supergiant with a radius of about 440 solar radii. The light curves evolved rapidly, on timescales of 1-2 hours, and appeared unusually fainter and redder than predicted by models within the first few hours, which we attribute to an optically thick dust shell before it was disrupted by the shockwave. We infer that the breakout and perhaps the distribution of the surrounding dust were not spherically symmetric.

astro-ph.HE

M31N 2013-10c: A Newly Identified Recurrent Nova in M31

The nova M31N 2023-11f (2023yoa) has been recently identified as the second eruption of a previously recognized nova, M31N 2013-10c, establishing the latter object as the 21st recurrent nova system thus far identified in M31. Here we present well sampled $R$-band lightcurves of both the 2013 and 2023 eruptions of this system. The photometric evolution of each eruption was quite similar as expected for the same progenitor system. The 2013 and 2023 eruptions each reached peak magnitudes just brighter than $R\sim16$, with fits to the declining branches of the eruptions yielding times to decline by two magnitudes of $t_2(R)=5.5\pm1.7$ and $t_2(R)=3.4\pm1.5$ days, respectively. M31N 2013-10c has an absolute magnitude at peak, $M_R=-8.8\pm0.2$, making it the most luminous known recurrent nova in M31.

astro-ph.SR

TransientViT: A novel CNN - Vision Transformer hybrid real/bogus transient classifier for the Kilodegree Automatic Transient Survey

The detection and analysis of transient astronomical sources is of great importance to understand their time evolution. Traditional pipelines identify transient sources from difference (D) images derived by subtracting prior-observed reference images (R) from new science images (N), a process that involves extensive manual inspection. In this study, we present TransientViT, a hybrid convolutional neural network (CNN) - vision transformer (ViT) model to differentiate between transients and image artifacts for the Kilodegree Automatic Transient Survey (KATS). TransientViT utilizes CNNs to reduce the image resolution and a hierarchical attention mechanism to model features globally. We propose a novel KATS-T 200K dataset that combines the difference images with both long- and short-term images, providing a temporally continuous, multidimensional dataset. Using this dataset as the input, TransientViT achieved a superior performance in comparison to other transformer- and CNN-based models, with an overall area under the curve (AUC) of 0.97 and an accuracy of 99.44%. Ablation studies demonstrated the impact of different input channels, multi-input fusion methods, and cross-inference strategies on the model performance. As a final step, a voting-based ensemble to combine the inference results of three NRD images further improved the model's prediction reliability and robustness. This hybrid model will act as a crucial reference for future studies on real/bogus transient classification.

astro-ph.IM

M31N 2017-01e: Discovery of a Previous Eruption in this Enigmatic Recurrent Nova

We report the discovery of a previously unknown eruption of the recurrent nova M31N 2017-01e that took place on 11 January 2012. The earlier eruption was detected by Pan-STARRS and occurred 1847 days (5.06 yr) prior to the eruption on 31 January 2017 (M31N 2017-01e). The nova has now been seen to have had a total of four recorded eruptions (M31N 2012-01c, 2017-01e, 2019-09d, and 2022-03d) with a mean time between outbursts of just $929.5\pm6.8$ days ($2.545\pm0.019$ yr), the second shortest recurrence time known for any nova. We also show that there is a blue variable source ($\langle V \rangle = 20.56\pm0.17$, $B-V\simeq0.045$), apparently coincident with the position of the nova, that exhibits a 14.3 d periodicity. Possible models of the system are proposed, but none are entirely satisfactory.

astro-ph.SR

M31N 1926-07c: A Recurrent Nova in M31 with a 2.8 Year Recurrence Time

The M31 recurrent nova M31N 1926-07c has had five recorded eruptions. Well-sampled light curves of the two most recent outbursts, in January of 2020 (M31N 2020-01b) and September 2022 (M31N 2022-09a), are presented showing that the photometric evolution of the two events were quite similar, with peak magnitudes of $R=17.2\pm0.1$ and $R=17.1\pm0.1$, and $t_2$ times of $9.7\pm0.9$ and $8.1\pm0.5$ days for the 2020 and 2022 eruptions, respectively. After considering the dates of the four most recent eruptions (where the cycle count is believed to be known), a mean recurrence interval of $\langle P_\mathrm{rec}\rangle=2.78\pm0.03$ years is found, establishing that M31N 1926-07c has one of the shortest recurrence times known.

astro-ph.SR

Ultrasound-Guided Assistive Robots for Scoliosis Assessment with Optimization-based Control and Variable Impedance

Assistive robots for healthcare have seen a growing demand due to the great potential of relieving medical practitioners from routine jobs. In this paper, we investigate the development of an optimization-based control framework for an ultrasound-guided assistive robot to perform scoliosis assessment. A conventional procedure for scoliosis assessment with ultrasound imaging typically requires a medical practitioner to slide an ultrasound probe along a patient's back. To automate this type of procedure, we need to consider multiple objectives, such as contact force, position, orientation, energy, posture, etc. To address the aforementioned components, we propose to formulate the control framework design as a quadratic programming problem with each objective weighed by its task priority subject to a set of equality and inequality constraints. In addition, as the robot needs to establish constant contact with the patient during spine scanning, we incorporate variable impedance regulation of the end-effector position and orientation in the control architecture to enhance safety and stability during the physical human-robot interaction. Wherein, the variable impedance gains are retrieved by learning from the medical expert's demonstrations. The proposed methodology is evaluated by conducting real-world experiments of autonomous scoliosis assessment with a robot manipulator xArm. The effectiveness is verified by the obtained coronal spinal images of both a phantom and a human subject.

cs.RO

The Unusual Eruption of the Extragalactic Classical Nova M31N 2017-09a

M31N 2017-09a is a classical nova and was observed for some 160 days following its initial eruption, during which time it underwent a number of bright secondary outbursts. The light-curve is characterized by continual variation with excursions of at least 0.5 magnitudes on a daily time-scale. The lower envelope of the eruption suggests that a single power-law can describe the decline rate. The eruption is relatively long with $t_2 = 111$, and $t_3 = 153$ days.

astro-ph.SR