SearcharxivSearch

arXiv subjects

Xinyu Shen

Publications and source records attributed to Xinyu Shen.

10 recordsLinked to original sources

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authentic professional domains. XpertBench consists of 1,346 meticulously curated tasks across 80 categories, spanning finance, healthcare, legal services, education, and dual-track research (STEM and Humanities). These tasks are derived from over 1,000 submissions by domain experts--including researchers from elite institutions and practitioners with extensive clinical or industrial experience--ensuring superior ecological validity. Each task uses detailed rubrics with mostly 15-40 weighted checkpoints to assess professional rigor. To facilitate scalable yet human-aligned assessment, we introduce ShotJudge, a novel evaluation paradigm that employs LLM judges calibrated with expert few-shot exemplars to mitigate self-rewarding biases. Our empirical evaluation of state-of-the-art LLMs reveals a pronounced performance ceiling: even leading models achieve a peak success rate of only ~66%, with a mean score around 55%. Models also exhibit domain-specific divergence, showing non-overlapping strengths in quantitative reasoning versus linguistic synthesis.. These findings underscore a significant "expert-gap" in current AI systems and establish XpertBench as a critical instrument for navigating the transition from general-purpose assistants to specialized professional collaborators.

cs.AI

Ligand Engineering for Precise Control of Ultrathin CsPbI3 Nanoplatelet Superlattices for Efficient Light-Emitting Diodes

Strongly-confined perovskite nanoplatelets (PeNPLs) offer opportunities not found in conventional isotropic nanocubes, especially in producing linearly polarized light, as well as enhancing outcoupling through control over the transition dipole moment. But this requires ultrathin nanoplatelets with three or fewer monolayers of PbI6 octahedra across the thickness, which are challenging to synthesise uniformly, and their luminescence is strongly affected by surface defects. Together, these limit the performance of ultrathin PeNPLs in light-emitting diodes (LEDs). Here, we address these challenges with an ancillary ligand engineering strategy. We demonstrate that ligands with phosphoryl functional groups strongly bind to the perovskite surface, while having an organic backbone that is not sterically bulky ensures high ligand density. By modulating nucleation and growth, these ancillary ligands lead to monodisperse PeNPLs that stack more uniformly when self-assembled into superlattices, with suppressed agglomeration. As a result, from edge-up PeNPL superlattices, we achieve enhanced degree of polarization, while from face-down PeNPL superlattices, we achieve enhanced outcoupling that results in LEDs with 13.1% external quantum efficiency, the highest reported for ultrathin PeNPL LEDs. This work establishes ancillary ligand-induced synthesis as a decisive route to achieve uniform nanoplatelets with robust orientation control, enabling full utilization of the multifunctionality of anisotropic PeNPLs.

physics.optics

Enhanced Stability and Linearly Polarized Emission from CsPbI$_3$ Perovskite Nanoplatelets through A-site Cation Engineering

The anisotropy of perovskite nanoplatelets (PeNPLs) opens up many opportunities in optoelectronics, including enabling the emission of linearly polarized light. But the limited stability of PeNPLs is a pressing challenge, especially for red-emitting CsPbI$_3$. Herein, we address this limitation by alloying FA into the perovskite cuboctahedral site. Unlike Cs/FA alloying in bulk thin films or nonconfined nanocubes, FA incorporation in nanoplatelets requires meticulous control over the reaction conditions, given that nanoplatelets are obtained in kinetically-driven growth regimes instead of thermodynamically-driven conditions. Through in-situ photoluminescence (PL) measurements, we find that excess FA leads to uncontrolled growth, where phase-impurities and nanoplatelets of multiple thicknesses co-exist. Restricting the FA content to up to 25% Cs substitution enables monodisperse PeNPLs, and increases the PL quantum yield (from 53% to 61%), exciton lifetime (from 18 ns to 27 ns), and stability in ambient air (from ~2 days to >7 days) compared to CsPbI$_3$. This arises due to hydrogen bonding between FA and the oleate and oleylammonium ligands, anchoring them to the surface to improve optoelectronic properties and stability. The reduction in non-radiative recombination, improvement in the nanoplatelet aspect ratio, and higher ligand density lead to FA-containing PeNPLs more effectively forming edge-up superlattices, enhancing the PL degree of linear polarization from 5.1% (CsPbI$_3$) to 9.4% (Cs$_{0.75}$FA$_{0.25}$PbI$_3$). These fundamental insights show how the stability limitations of PeNPLs could be addressed, and these materials grown more precisely to improve their performance as polarized light emitters, critical for utilizing them in next-generation display, bioimaging and communications applications.

cond-mat.mtrl-sci

Interpretable Credit Default Prediction with Ensemble Learning and SHAP

This study focuses on the problem of credit default prediction, builds a modeling framework based on machine learning, and conducts comparative experiments on a variety of mainstream classification algorithms. Through preprocessing, feature engineering, and model training of the Home Credit dataset, the performance of multiple models including logistic regression, random forest, XGBoost, LightGBM, etc. in terms of accuracy, precision, and recall is evaluated. The results show that the ensemble learning method has obvious advantages in predictive performance, especially in dealing with complex nonlinear relationships between features and data imbalance problems. It shows strong robustness. At the same time, the SHAP method is used to analyze the importance and dependency of features, and it is found that the external credit score variable plays a dominant role in model decision making, which helps to improve the model's interpretability and practical application value. The research results provide effective reference and technical support for the intelligent development of credit risk control systems.

cs.LG

CU-Net: a U-Net architecture for efficient brain-tumor segmentation on BraTS 2019 dataset

Accurately segmenting brain tumors from MRI scans is important for developing effective treatment plans and improving patient outcomes. This study introduces a new implementation of the Columbia-University-Net (CU-Net) architecture for brain tumor segmentation using the BraTS 2019 dataset. The CU-Net model has a symmetrical U-shaped structure and uses convolutional layers, max pooling, and upsampling operations to achieve high-resolution segmentation. Our CU-Net model achieved a Dice score of 82.41%, surpassing two other state-of-the-art models. This improvement in segmentation accuracy highlights the robustness and effectiveness of the model, which helps to accurately delineate tumor boundaries, which is crucial for surgical planning and radiation therapy, and ultimately has the potential to improve patient outcomes.

cs.CV

Harnessing XGBoost for Robust Biomarker Selection of Obsessive-Compulsive Disorder (OCD) from Adolescent Brain Cognitive Development (ABCD) data

This study evaluates the performance of various supervised machine learning models in analyzing highly correlated neural signaling data from the Adolescent Brain Cognitive Development (ABCD) Study, with a focus on predicting obsessive-compulsive disorder scales. We simulated a dataset to mimic the correlation structures commonly found in imaging data and evaluated logistic regression, elastic networks, random forests, and XGBoost on their ability to handle multicollinearity and accurately identify predictive features. Our study aims to guide the selection of appropriate machine learning methods for processing neuroimaging data, highlighting models that best capture underlying signals in high feature correlations and prioritize clinically relevant features associated with Obsessive-Compulsive Disorder (OCD).

q-bio.NC

A Protein Structure Prediction Approach Leveraging Transformer and CNN Integration

Proteins are essential for life, and their structure determines their function. The protein secondary structure is formed by the folding of the protein primary structure, and the protein tertiary structure is formed by the bending and folding of the secondary structure. Therefore, the study of protein secondary structure is very helpful to the overall understanding of protein structure. Although the accuracy of protein secondary structure prediction has continuously improved with the development of machine learning and deep learning, progress in the field of protein structure prediction, unfortunately, remains insufficient to meet the large demand for protein information. Therefore, based on the advantages of deep learning-based methods in feature extraction and learning ability, this paper adopts a two-dimensional fusion deep neural network model, DstruCCN, which uses Convolutional Neural Networks (CCN) and a supervised Transformer protein language model for single-sequence protein structure prediction. The training features of the two are combined to predict the protein Transformer binding site matrix, and then the three-dimensional structure is reconstructed using energy minimization.

q-bio.BM

Construction and application of artificial intelligence crowdsourcing map based on multi-track GPS data

In recent years, the rapid development of high-precision map technology combined with artificial intelligence has ushered in a new development opportunity in the field of intelligent vehicles. High-precision map technology is an important guarantee for intelligent vehicles to achieve autonomous driving. However, due to the lack of research on high-precision map technology, it is difficult to rationally use this technology in the field of intelligent vehicles. Therefore, relevant researchers studied a fast and effective algorithm to generate high-precision GPS data from a large number of low-precision GPS trajectory data fusion, and generated several key data points to simplify the description of GPS trajectory, and realized the "crowdsourced update" model based on a large number of social vehicles for map data collection came into being. This kind of algorithm has the important significance to improve the data accuracy, reduce the measurement cost and reduce the data storage space. On this basis, this paper analyzes the implementation form of crowdsourcing map, so as to improve the various information data in the high-precision map according to the actual situation, and promote the high-precision map can be reasonably applied to the intelligent car.

cs.AI

DeepGI: An Automated Approach for Gastrointestinal Tract Segmentation in MRI Scans

Gastrointestinal (GI) tract cancers pose a global health challenge, demanding precise radiotherapy planning for optimal treatment outcomes. This paper introduces a cutting-edge approach to automate the segmentation of GI tract regions in magnetic resonance imaging (MRI) scans. Leveraging advanced deep learning architectures, the proposed model integrates Inception-V4 for initial classification, UNet++ with a VGG19 encoder for 2.5D data, and Edge UNet for grayscale data segmentation. Meticulous data preprocessing, including innovative 2.5D processing, is employed to enhance adaptability, robustness, and accuracy. This work addresses the manual and time-consuming segmentation process in current radiotherapy planning, presenting a unified model that captures intricate anatomical details. The integration of diverse architectures, each specializing in unique aspects of the segmentation task, signifies a novel and comprehensive solution. This model emerges as an efficient and accurate tool for clinicians, marking a significant advancement in the field of GI tract image segmentation for radiotherapy planning.

eess.IV

Characterizing Datasets for Social Visual Question Answering, and the New TinySocial Dataset

Modern social intelligence includes the ability to watch videos and answer questions about social and theory-of-mind-related content, e.g., for a scene in Harry Potter, "Is the father really upset about the boys flying the car?" Social visual question answering (social VQA) is emerging as a valuable methodology for studying social reasoning in both humans (e.g., children with autism) and AI agents. However, this problem space spans enormous variations in both videos and questions. We discuss methods for creating and characterizing social VQA datasets, including 1) crowdsourcing versus in-house authoring, including sample comparisons of two new datasets that we created (TinySocial-Crowd and TinySocial-InHouse) and the previously existing Social-IQ dataset; 2) a new rubric for characterizing the difficulty and content of a given video; and 3) a new rubric for characterizing question types. We close by describing how having well-characterized social VQA datasets will enhance the explainability of AI agents and can also inform assessments and educational interventions for people.

cs.HC