SearcharxivSearch

arXiv subjects

Jiangbo Zhang

Publications and source records attributed to Jiangbo Zhang.

5 recordsLinked to original sources

ElectriQ: A Benchmark for Assessing the Response Capability of Large Language Models in Power Marketing

As power systems decarbonise and digitalise, high penetrations of distributed energy resources and flexible tariffs make electric power marketing (EPM) a key interface between regulation, system operation and sustainable-energy deployment. Many utilities still rely on human agents and rule- or intent-based chatbots with fragmented knowledge bases that struggle with long, cross-scenario dialogues and fall short of requirements for compliant, verifiable and DR-ready interactions. Meanwhile, frontier large language models (LLMs) show strong conversational ability but are evaluated on generic benchmarks that underweight sector-specific terminology, regulatory reasoning and multi-turn process stability. To address this gap, we present ElectriQ, a large-scale benchmark and evaluation framework for LLMs in EPM. ElectriQ contains over 550k dialogues across six service domains and 24 sub-scenarios and defines a unified protocol that combines human ratings, automatic metrics and two compliance stress tests-Statutory Citation Correctness and Long-Dialogue Consistency. Building on ElectriQ, we propose SEEK-RAG, a retrieval-augmented method that injects policy and domain knowledge during finetuning and inference. Experiments on 13 LLMs show that domain-aligned 7B models with SEEK-RAG match or surpass much larger models while reducing computational cost, providing an auditable, regulation-aware basis for deploying LLM-based EPM assistants that support demand-side management, renewable integration and resilient grid operation.

cs.CL

OrthoInsight: Rib Fracture Diagnosis and Report Generation Based on Multi-Modal Large Models

The growing volume of medical imaging data has increased the need for automated diagnostic tools, especially for musculoskeletal injuries like rib fractures, commonly detected via CT scans. Manual interpretation is time-consuming and error-prone. We propose OrthoInsight, a multi-modal deep learning framework for rib fracture diagnosis and report generation. It integrates a YOLOv9 model for fracture detection, a medical knowledge graph for retrieving clinical context, and a fine-tuned LLaVA language model for generating diagnostic reports. OrthoInsight combines visual features from CT images with expert textual data to deliver clinically useful outputs. Evaluated on 28,675 annotated CT images and expert reports, it achieves high performance across Diagnostic Accuracy, Content Completeness, Logical Coherence, and Clinical Guidance Value, with an average score of 4.28, outperforming models like GPT-4 and Claude-3. This study demonstrates the potential of multi-modal learning in transforming medical image analysis and providing effective support for radiologists.

eess.IV

LumiCRS: Asymmetric Contrastive Prototype Learning for Long-Tail Conversational Recommender Systems

Conversational recommender systems (CRSs) often suffer from an extreme long-tail distribution of dialogue data, causing a strong bias toward head-frequency blockbusters that sacrifices diversity and exacerbates the cold-start problem. An empirical analysis of DCRS and statistics on the REDIAL corpus show that only 10% of head movies account for nearly half of all mentions, whereas about 70% of tail movies receive merely 26% of the attention. This imbalance gives rise to three critical challenges: head over-fitting, body representation drift, and tail sparsity. To address these issues, we propose LumiCRS, an end-to-end framework that mitigates long-tail imbalance through three mutually reinforcing layers: (i) an Adaptive Comprehensive Focal Loss (ACFL) that dynamically adjusts class weights and focusing factors to curb head over-fitting and reduce popularity bias; (ii) Prototype Learning for Long-Tail Recommendation, which selects semantic, affective, and contextual prototypes to guide clustering and stabilize body and tail representations; and (iii) a GPT-4o-driven prototype-guided dialogue augmentation module that automatically generates diverse long-tail conversational snippets to alleviate tail sparsity and distribution shift. Together, these strategies enable LumiCRS to markedly improve recommendation accuracy, diversity, and fairness: on the REDIAL and INSPIRED benchmarks, LumiCRS boosts Recall@10 and Tail-Recall@10 by 7-15% over fifteen strong baselines, while human evaluations confirm superior fluency, informativeness, and long-tail relevance. These results demonstrate the effectiveness of multi-layer collaboration in building an efficient and fair long-tail conversational recommender.

cs.AI

SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults

The reliability of substation equipment is crucial to the stability of power systems, but traditional fault analysis methods heavily rely on manual expertise, limiting their effectiveness in handling complex and large-scale data. This paper proposes a substation equipment fault analysis method based on a multimodal large language model (MLLM). We developed a database containing 40,000 entries, including images, defect labels, and analysis reports, and used an image-to-video generation model for data augmentation. Detailed fault analysis reports were generated using GPT-4. Based on this database, we developed SubstationAI, the first model dedicated to substation fault analysis, and designed a fault diagnosis knowledge base along with knowledge enhancement methods. Experimental results show that SubstationAI significantly outperforms existing models, such as GPT-4, across various evaluation metrics, demonstrating higher accuracy and practicality in fault cause analysis, repair suggestions, and preventive measures, providing a more advanced solution for substation equipment fault analysis.

cs.AI

Dynamics of Opinions with Bounded Confidence in Social Cliques: Emergence of Fluctuations

In this paper, we study the evolution of opinions over social networks with bounded confidence in social cliques. Node initial opinions are independently and identically distributed; at each time step, nodes review the average opinions of a randomly selected local clique. The clique averages may represent local group pressures on peers. Then nodes update their opinions under bounded confidence: only when the difference between an agent individual opinion and the corresponding local clique pressure is below a threshold, this agent opinion is updated according to the DeGroot rule as a weighted average of the two values. As a result, this opinion dynamics is a generalization of the classical Deffuant-Weisbuch model in which only pairwise interactions take place. First of all, we prove conditions under which all node opinions converge to finite limits. We show that in the limits the event that all nodes achieve a consensus, and the event that all nodes achieve pairwise distinct limits, i.e., social disagreements, are both nontrivial events. Next, we show that opinion fluctuations may take place in the sense that at least one agent in the network fails to hold a converging opinion trajectory. In fact, we prove that this fluctuation event happens with a strictly positive probability, and also constructively present an initial value event under which the fluctuation event arises with probability one. These results add to the understanding of the role of bounded confidence in social opinion dynamics, and the possibility of fluctuation reveals that bringing in cliques in Deffuant-Weisbuch models have fundamentally changed the behavior of such opinion dynamical processes.

cs.SI