SearcharxivSearch

arXiv subjects

Yang Dong

Publications and source records attributed to Yang Dong.

At least 19 recordsLinked to original sources

CoeusBI: A Comprehensive Interactive Business Intelligence System Powered by LLMs at Baidu [Extended Version]

The advent of Large Language Models has catalyzed the emergence of interactive Business Intelligence (BI) systems. Although commercial BI products increasingly adopt semantic layers paired with natural language interfaces, they predominantly rely on manual configurations to define metrics and dimensions. Real-world deployments face critical challenges: (a) frequent JOIN operations degrade the accuracy of SQL generation; (b) wide schemas exacerbate the challenge of schema linking; and (c) the generation of dialect-specific queries and the accurate support for multi-round dialogues incur high computational costs and yield limited accuracy. We introduce CoeusBI, an industrial-scale interactive BI system that addresses these barriers through a novel Dual-Agent Architecture paired with a Hierarchical Schema Linking module: (1) an offline View Generation Agent that utilizes error-feedback to autonomously convert complex JOIN queries into simple single-view queries, which eliminates the need for manual semantic modeling; (2) a Hierarchical Schema Linking module that leverages vector retrieval over views to handle exceptionally wide schemas efficiently; and (3) a dynamic Routing Agent that evaluates dialogue contexts to route queries, dynamically invoking either the synthesis of new intermediate representations or targeted modifications of existing ones, before compiling the unified representation via a deterministic SQL compiler that is agnostic to dialects. Extensive experiments on both public datasets and production datasets demonstrate that CoeusBI achieves significant improvements in query accuracy, token efficiency, and user satisfaction relative to existing methods. CoeusBI is deployed as a standalone service on the data platform of Baidu and is widely used across multiple business lines supporting thousands of users daily, thereby evidencing strong practicality and scalability.

cs.DB

ChatBI: Towards Natural Language to Complex Business Intelligence SQL

The Natural Language to SQL (NL2SQL) technology provides non-expert users who are unfamiliar with databases the opportunity to use SQL for data analysis.Converting Natural Language to Business Intelligence (NL2BI) is a popular practical scenario for NL2SQL in actual production systems. Compared to NL2SQL, NL2BI introduces more challenges. In this paper, we propose ChatBI, a comprehensive and efficient technology for solving the NL2BI task. First, we analyze the interaction mode, an important module where NL2SQL and NL2BI differ in use, and design a smaller and cheaper model to match this interaction mode. In BI scenarios, tables contain a huge number of columns, making it impossible for existing NL2SQL methods that rely on Large Language Models (LLMs) for schema linking to proceed due to token limitations. The higher proportion of ambiguous columns in BI scenarios also makes schema linking difficult. ChatBI combines existing view technology in the database community to first decompose the schema linking problem into a Single View Selection problem and then uses a smaller and cheaper machine learning model to select the single view with a significantly reduced number of columns. The columns of this single view are then passed as the required columns for schema linking into the LLM. Finally, ChatBI proposes a phased process flow different from existing process flows, which allows ChatBI to generate SQL containing complex semantics and comparison relations more accurately. We have deployed ChatBI on Baidu's data platform and integrated it into multiple product lines for large-scale production task evaluation. The obtained results highlight its superiority in practicality, versatility, and efficiency. At the same time, compared with the current mainstream NL2SQL technology under our real BI scenario data tables and queries, it also achieved the best results.

cs.DB

RCBEVDet: Radar-camera Fusion in Bird's Eye View for 3D Object Detection

Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely on cameras is difficult to achieve highly accurate and robust 3D object detection. An effective solution to this issue is combining multi-view cameras with the economical millimeter-wave radar sensor to achieve more reliable multi-modal 3D object detection. In this paper, we introduce RCBEVDet, a radar-camera fusion 3D object detection method in the bird's eye view (BEV). Specifically, we first design RadarBEVNet for radar BEV feature extraction. RadarBEVNet consists of a dual-stream radar backbone and a Radar Cross-Section (RCS) aware BEV encoder. In the dual-stream radar backbone, a point-based encoder and a transformer-based encoder are proposed to extract radar features, with an injection and extraction module to facilitate communication between the two encoders. The RCS-aware BEV encoder takes RCS as the object size prior to scattering the point feature in BEV. Besides, we present the Cross-Attention Multi-layer Fusion module to automatically align the multi-modal BEV feature from radar and camera with the deformable attention mechanism, and then fuse the feature with channel and spatial fusion layers. Experimental results show that RCBEVDet achieves new state-of-the-art radar-camera fusion results on nuScenes and view-of-delft (VoD) 3D object detection benchmarks. Furthermore, RCBEVDet achieves better 3D detection results than all real-time camera-only and radar-camera 3D object detectors with a faster inference speed at 21~28 FPS. The source code will be released at https://github.com/VDIGPKU/RCBEVDet.

cs.CV

From Pixel to Slide image: Polarization Modality-based Pathological Diagnosis Using Representation Learning

Thyroid cancer is the most common endocrine malignancy, and accurately distinguishing between benign and malignant thyroid tumors is crucial for developing effective treatment plans in clinical practice. Pathologically, thyroid tumors pose diagnostic challenges due to improper specimen sampling. In this study, we have designed a three-stage model using representation learning to integrate pixel-level and slice-level annotations for distinguishing thyroid tumors. This structure includes a pathology structure recognition method to predict structures related to thyroid tumors, an encoder-decoder network to extract pixel-level annotation information by learning the feature representations of image blocks, and an attention-based learning mechanism for the final classification task. This mechanism learns the importance of different image blocks in a pathological region, globally considering the information from each block. In the third stage, all information from the image blocks in a region is aggregated using attention mechanisms, followed by classification to determine the category of the region. Experimental results demonstrate that our proposed method can predict microscopic structures more accurately. After color-coding, the method achieves results on unstained pathology slides that approximate the quality of Hematoxylin and eosin staining, reducing the need for stained pathology slides. Furthermore, by leveraging the concept of indirect measurement and extracting polarized features from structures correlated with lesions, the proposed method can also classify samples where membrane structures cannot be obtained through sampling, providing a potential objective and highly accurate indirect diagnostic technique for thyroid tumors.

eess.IV

A Polarization and Radiomics Feature Fusion Network for the Classification of Hepatocellular Carcinoma and Intrahepatic Cholangiocarcinoma

Classifying hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma (ICC) is a critical step in treatment selection and prognosis evaluation for patients with liver diseases. Traditional histopathological diagnosis poses challenges in this context. In this study, we introduce a novel polarization and radiomics feature fusion network, which combines polarization features obtained from Mueller matrix images of liver pathological samples with radiomics features derived from corresponding pathological images to classify HCC and ICC. Our fusion network integrates a two-tier fusion approach, comprising early feature-level fusion and late classification-level fusion. By harnessing the strengths of polarization imaging techniques and image feature-based machine learning, our proposed fusion network significantly enhances classification accuracy. Notably, even at reduced imaging resolutions, the fusion network maintains robust performance due to the additional information provided by polarization features, which may not align with human visual perception. Our experimental results underscore the potential of this fusion network as a powerful tool for computer-aided diagnosis of HCC and ICC, showcasing the benefits and prospects of integrating polarization imaging techniques into the current image-intensive digital pathological diagnosis. We aim to contribute this innovative approach to top-tier journals, offering fresh insights and valuable tools in the fields of medical imaging and cancer diagnosis. By introducing polarization imaging into liver cancer classification, we demonstrate its interdisciplinary potential in addressing challenges in medical image analysis, promising advancements in medical imaging and cancer diagnosis.

eess.IV

Quantum enhanced radio detection and ranging with solid spins

The accurate radio frequency (RF) ranging and localizing of objects has benefited the researches including autonomous driving, the Internet of Things, and manufacturing. Quantum receivers have been proposed to detect the radio signal with ability that can outperform conventional measurement. As one of the most promising candidates, solid spin shows superior robustness, high spatial resolution and miniaturization. However, challenges arise from the moderate response to a high frequency RF signal. Here, by exploiting the coherent interaction between quantum sensor and RF field, we demonstrate quantum enhanced radio detection and ranging. The RF magnetic sensitivity is improved by three orders to 21 $pT/\sqrt{Hz}$, based on nanoscale quantum sensing and RF focusing. Further enhancing the response of spins to the target's position through multi-photon excitation, a ranging accuracy of 16 $\mu m$ is realized with a GHz RF signal. The results pave the way for exploring quantum enhanced radar and communications with solid spins.

quant-ph

Learn from Yesterday: A Semi-Supervised Continual Learning Method for Supervision-Limited Text-to-SQL Task Streams

Conventional text-to-SQL studies are limited to a single task with a fixed-size training and test set. When confronted with a stream of tasks common in real-world applications, existing methods struggle with the problems of insufficient supervised data and high retraining costs. The former tends to cause overfitting on unseen databases for the new task, while the latter makes a full review of instances from past tasks impractical for the model, resulting in forgetting of learned SQL structures and database schemas. To address the problems, this paper proposes integrating semi-supervised learning (SSL) and continual learning (CL) in a stream of text-to-SQL tasks and offers two promising solutions in turn. The first solution Vanilla is to perform self-training, augmenting the supervised training data with predicted pseudo-labeled instances of the current task, while replacing the full volume retraining with episodic memory replay to balance the training efficiency with the performance of previous tasks. The improved solution SFNet takes advantage of the intrinsic connection between CL and SSL. It uses in-memory past information to help current SSL, while adding high-quality pseudo instances in memory to improve future replay. The experiments on two datasets shows that SFNet outperforms the widely-used SSL-only and CL-only baselines on multiple metrics.

cs.CL

HeteroQA: Learning towards Question-and-Answering through Multiple Information Sources via Heterogeneous Graph Modeling

Community Question Answering (CQA) is a well-defined task that can be used in many scenarios, such as E-Commerce and online user community for special interests. In these communities, users can post articles, give comment, raise a question and answer it. These data form the heterogeneous information sources where each information source have their own special structure and context (comments attached to an article or related question with answers). Most of the CQA methods only incorporate articles or Wikipedia to extract knowledge and answer the user's question. However, various types of information sources in the community are not fully explored by these CQA methods and these multiple information sources (MIS) can provide more related knowledge to user's questions. Thus, we propose a question-aware heterogeneous graph transformer to incorporate the MIS in the user community to automatically generate the answer. To evaluate our proposed method, we conduct the experiments on two datasets: $\text{MSM}^{\text{plus}}$ the modified version of benchmark dataset MS-MARCO and the AntQA dataset which is the first large-scale CQA dataset with four types of MIS. Extensive experiments on two datasets show that our model outperforms all the baselines in terms of all the metrics.

cs.CL

S+PAGE: A Speaker and Position-Aware Graph Neural Network Model for Emotion Recognition in Conversation

Emotion recognition in conversation (ERC) has attracted much attention in recent years for its necessity in widespread applications. Existing ERC methods mostly model the self and inter-speaker context separately, posing a major issue for lacking enough interaction between them. In this paper, we propose a novel Speaker and Position-Aware Graph neural network model for ERC (S+PAGE), which contains three stages to combine the benefits of both Transformer and relational graph convolution network (R-GCN) for better contextual modeling. Firstly, a two-stream conversational Transformer is presented to extract the coarse self and inter-speaker contextual features for each utterance. Then, a speaker and position-aware conversation graph is constructed, and we propose an enhanced R-GCN model, called PAG, to refine the coarse features guided by a relative positional encoding. Finally, both of the features from the former two stages are input into a conditional random field layer to model the emotion transfer.

cs.CL

Relation Aware Semi-autoregressive Semantic Parsing for NL2SQL

Natural language to SQL (NL2SQL) aims to parse a natural language with a given database into a SQL query, which widely appears in practical Internet applications. Jointly encode database schema and question utterance is a difficult but important task in NL2SQL. One solution is to treat the input as a heterogeneous graph. However, it failed to learn good word representation in question utterance. Learning better word representation is important for constructing a well-designed NL2SQL system. To solve the challenging task, we present a Relation aware Semi-autogressive Semantic Parsing (\MODN) ~framework, which is more adaptable for NL2SQL. It first learns relation embedding over the schema entities and question words with predefined schema relations with ELECTRA and relation aware transformer layer as backbone. Then we decode the query SQL with a semi-autoregressive parser and predefined SQL syntax. From empirical results and case study, our model shows its effectiveness in learning better word representation in NL2SQL.

cs.CL

Revealing complex optical phenomena through vectorial metrics

Advances in vectorial polarisation-resolved imaging are bringing new capabilities to applications ranging from fundamental physics through to clinical diagnosis. Imaging polarimetry requires determination of the Mueller matrix (MM) at every point, providing a complete description of an object's vectorial properties. Despite forming a comprehensive representation, the MM does not usually provide easily-interpretable information about the object's internal structure. Certain simpler vectorial metrics are derived from subsets of the MM elements. These metrics permit extraction of signatures that provide direct indicators of hidden optical properties of complex systems, while featuring an intriguing asymmetry about what information can or cannot be inferred via these metrics. We harness such characteristics to reveal the spin-Hall effect of light, infer microscopic structure within laser-written photonic waveguides, and conduct rapid pathological diagnosis through analysis of healthy and cancerous tissue. This provides new insight for the broader usage of such asymmetric inferred vectorial information.

physics.optics

SeaD: End-to-end Text-to-SQL Generation with Schema-aware Denoising

In text-to-SQL task, seq-to-seq models often lead to sub-optimal performance due to limitations in their architecture. In this paper, we present a simple yet effective approach that adapts transformer-based seq-to-seq model to robust text-to-SQL generation. Instead of inducing constraint to decoder or reformat the task as slot-filling, we propose to train seq-to-seq model with Schema aware Denoising (SeaD), which consists of two denoising objectives that train model to either recover input or predict output from two novel erosion and shuffle noises. These denoising objectives acts as the auxiliary tasks for better modeling the structural data in S2S generation. In addition, we improve and propose a clause-sensitive execution guided (EG) decoding strategy to overcome the limitation of EG decoding for generative model. The experiments show that the proposed method improves the performance of seq-to-seq model in both schema linking and grammar correctness and establishes new state-of-the-art on WikiSQL benchmark. The results indicate that the capacity of vanilla seq-to-seq architecture for text-to-SQL may have been under-estimated.

cs.CL

Heisenberg-Limited Waveform Estimation with Solid-State Spins in Diamond

The newly established Heisenberg limit in arbitrary waveform estimation is quite different with parameter estimation and shows a unique characteristic of a future quantum version of oscilloscope. However, it is still a non-trivial challenge to generate a large number of exotic quantum entangled states to achieve this quantum limit. Here, by employing the time-domain quantum difference detection method, we demonstrate Heisenberg-limited waveform quantum estimation with diamond spins under ambient condition in the experiment. Periodic dynamical decoupling is applied to enhance both the dynamic range and sensitivity by one order of magnitude. Using this quantum-enhanced estimation scheme, the estimation error of an unknown waveform is reduced by more than $5$ dB below the standard quantum limit with $N\sim{\text{2}} \times {\text{1}}{{\text{0}}^3}$ resources, where more than ${1 \times {\text{1}}{{\text{0}}^5}}$ resources would be required to achieve a similar error level using classical detection. This work provides an essential step towards realizing quantum-enhanced structure recognition in a continuous space and time.

quant-ph

Focus the electromagnetic field to $10^{-6} \lambda$ for ultra-high enhancement of field-matter interaction

Focusing electromagnetic field to enhance the interaction with matter has been promoting researches and applications of nano electronics and photonics. Usually, the evanescent-wave coupling is adopted in various nano structures and materials to confine the electromagnetic field into a subwavelength space. Here, based on the direct coupling with confined electron oscillations in a nanowire, we demonstrate an extreme localization of microwave field down to 10$^{-6}\lambda$. A hybrid nanowire-bowtie antenna is further designed to focus the free-space microwave to this deep-subwavelength space. Detected by the nitrogen vacancy center in diamond, the field intensity and microwave-spin interaction strength are enhanced by 2.0$\times$10$^{8}$ and 1.4$\times$10$^{4}$ times, respectively. Such an extreme concentration of microwave field will further promote integrated quantum information processing, sensing and microwave photonics in a nanoscale system.

physics.optics

Fast high-fidelity geometric quantum control with quantum brachistochrones

We experimentally demonstrate fast and high-fidelity geometric control of a quantum system with the most brachistochrone method on hybrid spin registers in diamond. Based on the time-optimal universal geometric control, single geometric gates with the fidelities over 99.2% on the spin state of nitrogen-vacancy center are realized with average durations shortened by 74.9%, comparing with conventional geometric method. The fidelity of the fast geometric two-qubit gate exceeds 96.5% on the hybrid spin registers. With these fast high-fidelity gates available, we implement quantum entanglement-enhanced phase estimation algorithm and demonstrate the Heisenberg quantum limit at room-temperature. By comparing with the conventional geometric circuit, the measurement bandwidth and sensitivity is enhanced by 3.5 and 2.9 times. Hence, our results show that high-fidelity quantum control based on a fast geometric route will be a versatile tool for broad applications of quantum information processing in practice.

quant-ph

A high-sensitivity fiber-coupled diamond magnetometer with surface coating

Nitrogen-vacancy quantum defects in diamond offer a promising platform for magnetometry because of their remarkable optical and spin properties. In this Letter, we present a high-sensitivity and wide-bandwidth fiber-based quantum magnetometer for practical applications. By coating the diamond surface with silver reflective film, both the fluorescence collection and excitation efficiency are enhanced. Additionally, tracking pulsed optically detected magnetic resonance spectrum allowed a magnetic field sensitivity of $35$ pT$/\sqrt{\rm{Hz}}$ and a bandwidth of $4.1$ KHz. Finally, this magnetometer was successfully applied to map the magnetic field induced by the current-carrying copper-wire mesh. Such a stable and compact magnetometry can provide a powerful tool in many areas of physical, chemical, and biological researches.

physics.app-ph

Experimental implementation of universal holonomic quantum computation on solid-state spins with optimal control

Experimental realization of a universal set of quantum logic gates with high-fidelity is critical to quantum information processing, which is always challenging by inevitable interaction between the quantum system and environment. Geometric quantum computation is noise immune, and thus offers a robust way to enhance the control fidelity. Here, we experimentally implement the recently proposed extensible nonadiabatic holonomic quantum computation with solid spins in diamond at room-temperature, which maintains both flexibility and resilience against decoherence and system control errors. Compared with previous geometric method, the fidelities of a universal set of holonomic single-qubit and two-qubit quantum logic gates are improved in experiment. Therefore, this work makes an important step towards fault-tolerant scalable geometric quantum computation in realistic systems.

quant-ph

A robust fiber-based quantum thermometer coupled with nitrogen-vacancy centers

The nitrogen-vacancy center in diamond has been broadly applied in quantum sensing since it is sensitive to different physical quantities. Meanwhile, it is difficult to isolate disturbances from unwanted physical quantities in practical applications. Here, we present a robust fiber-based quantum thermometer which can significantly isolate the magnetic field noise and microwave power shift. With a frequency modulation scheme, we realize the temperature measurement by detecting the variation of the sharp-dip in the zero-field optically detected magnetic resonance spectrum in a high-density nitrogen-vacancy ensemble. Thanks to its simplicity and compatibility in implementation and robustness in the isolation of magnetic and microwave noise, this quantum thermometer is then applied to the surface temperature imaging of an electronic chip with a sensitivity of $18$ $\rm{mK}/\sqrt{\rm{Hz}}$. It paves the way to high sensitive temperature measurement in ambiguous environments.

physics.app-ph