SearcharxivSearch

arXiv subjects

Siyu Tian

Publications and source records attributed to Siyu Tian.

11 recordsLinked to original sources

GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action representations grounded by robot-action supervision. Specifically, VGGT-$Ω$ jointly aggregates multi-view observations at each timestep to construct geometry-aware multi-view states. The latent world model predicts the next-step patch tokens of a selected target view using patch and register tokens obtained after multi-view aggregation, thereby retaining multi-view geometric information without predicting the complete multi-view state. We use the wrist view as the target in our experiments, placing greater emphasis on end-effector motion and local gripper-object interactions. Finally, the shared latent-action representations condition both the latent world model and the flow-matching action head, allowing latent-prediction supervision and ground-truth robot-action supervision to jointly shape the same latent-action representations. Experiments across both simulation and real-world environments demonstrate the effectiveness and robustness of GWM-VLA.

cs.RO

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches

Recent progress in large language models (LLMs) has outpaced the development of effective evaluation methods. Traditional benchmarks rely on task-specific metrics and static datasets, which often suffer from fairness issues, limited scalability, and contamination risks. In this paper, we introduce Teach2Eval, an indirect evaluation framework inspired by the Feynman Technique. Instead of directly testing LLMs on predefined tasks, our method evaluates a model's multiple abilities to teach weaker student models to perform tasks effectively. By converting open-ended tasks into standardized multiple-choice questions (MCQs) through teacher-generated feedback, Teach2Eval enables scalable, automated, and multi-dimensional assessment. Our approach not only avoids data leakage and memorization but also captures a broad range of cognitive abilities that are orthogonal to current benchmarks. Experimental results across 26 leading LLMs show strong alignment with existing human and model-based dynamic rankings, while offering additional interpretability for training guidance.

cs.CL

$R^3$-NL2GQL: A Model Coordination and Knowledge Graph Alignment Approach for NL2GQL

While current tasks of converting natural language to SQL (NL2SQL) using Foundation Models have shown impressive achievements, adapting these approaches for converting natural language to Graph Query Language (NL2GQL) encounters hurdles due to the distinct nature of GQL compared to SQL, alongside the diverse forms of GQL. Moving away from traditional rule-based and slot-filling methodologies, we introduce a novel approach, $R^3$-NL2GQL, integrating both small and large Foundation Models for ranking, rewriting, and refining tasks. This method leverages the interpretative strengths of smaller models for initial ranking and rewriting stages, while capitalizing on the superior generalization and query generation prowess of larger models for the final transformation of natural language queries into GQL formats. Addressing the scarcity of datasets in this emerging field, we have developed a bilingual dataset, sourced from graph database manuals and selected open-source Knowledge Graphs (KGs). Our evaluation of this methodology on this dataset demonstrates its promising efficacy and robustness.

cs.CL

DIAS: A Dataset and Benchmark for Intracranial Artery Segmentation in DSA sequences

The automated segmentation of Intracranial Arteries (IA) in Digital Subtraction Angiography (DSA) plays a crucial role in the quantification of vascular morphology, significantly contributing to computer-assisted stroke research and clinical practice. Current research primarily focuses on the segmentation of single-frame DSA using proprietary datasets. However, these methods face challenges due to the inherent limitation of single-frame DSA, which only partially displays vascular contrast, thereby hindering accurate vascular structure representation. In this work, we introduce DIAS, a dataset specifically developed for IA segmentation in DSA sequences. We establish a comprehensive benchmark for evaluating DIAS, covering full, weak, and semi-supervised segmentation methods. Specifically, we propose the vessel sequence segmentation network, in which the sequence feature extraction module effectively captures spatiotemporal representations of intravascular contrast, achieving intracranial artery segmentation in 2D+Time DSA sequences. For weakly-supervised IA segmentation, we propose a novel scribble learning-based image segmentation framework, which, under the guidance of scribble labels, employs cross pseudo-supervision and consistency regularization to improve the performance of the segmentation network. Furthermore, we introduce the random patch-based self-training framework, aimed at alleviating the performance constraints encountered in IA segmentation due to the limited availability of annotated DSA data. Our extensive experiments on the DIAS dataset demonstrate the effectiveness of these methods as potential baselines for future research and clinical applications. The dataset and code are publicly available at https://doi.org/10.5281/zenodo.11396520 and https://github.com/lseventeen/DIAS.

eess.IV

SilverSight: A Multi-Task Chinese Financial Large Language Model Based on Adaptive Semantic Space Learning

Large language models (LLMs) are increasingly being applied across various specialized fields, leveraging their extensive knowledge to empower a multitude of scenarios within these domains. However, each field encompasses a variety of specific tasks that require learning, and the diverse, heterogeneous data across these domains can lead to conflicts during model task transfer. In response to this challenge, our study introduces an Adaptive Semantic Space Learning (ASSL) framework, which utilizes the adaptive reorganization of data distributions within the semantic space to enhance the performance and selection efficacy of multi-expert models. Utilizing this framework, we trained a financial multi-task LLM named "SilverSight". Our research findings demonstrate that our framework can achieve results close to those obtained with full data training using only 10% of the data, while also exhibiting strong generalization capabilities.

cs.CL

UDCR: Unsupervised Aortic DSA/CTA Rigid Registration Using Deep Reinforcement Learning and Overlap Degree Calculation

The rigid registration of aortic Digital Subtraction Angiography (DSA) and Computed Tomography Angiography (CTA) can provide 3D anatomical details of the vasculature for the interventional surgical treatment of conditions such as aortic dissection and aortic aneurysms, holding significant value for clinical research. However, the current methods for 2D/3D image registration are dependent on manual annotations or synthetic data, as well as the extraction of landmarks, which is not suitable for cross-modal registration of aortic DSA/CTA. In this paper, we propose an unsupervised method, UDCR, for aortic DSA/CTA rigid registration based on deep reinforcement learning. Leveraging the imaging principles and characteristics of DSA and CTA, we have constructed a cross-dimensional registration environment based on spatial transformations. Specifically, we propose an overlap degree calculation reward function that measures the intensity difference between the foreground and background, aimed at assessing the accuracy of registration between segmentation maps and DSA images. This method is highly flexible, allowing for the loading of pre-trained models to perform registration directly or to seek the optimal spatial transformation parameters through online learning. We manually annotated 61 pairs of aortic DSA/CTA for algorithm evaluation. The results indicate that the proposed UDCR achieved a Mean Absolute Error (MAE) of 2.85 mm in translation and 4.35° in rotation, showing significant potential for clinical applications.

eess.IV

Tuning Optical Properties of Metamaterials by Mie Scattering for Efficient Sub-ambient Daytime Radiative Cooling

The management of the abundant eggshell biowaste produced worldwide has become a problematic issue due to the generated odor and microorganisms after directly disposing eggshell biowaste in landfills. Herein, we propose a novel method to convert the hazardous eggshell biowaste to valuable resources for energy management applications. Eggshell-based films are fabricated by embedding eggshell powders into polymer matrix to achieve highly efficient sub-ambient daytime radiative cooling. Benefiting from the Mie scattering of eggshell particles/air pores in the solar spectrum and strong emission of eggshell in the mid-infrared (mid-IR) range, the eggshell-based films present high reflection of 0.96 in the solar spectrum and high emission of 0.95 in the mid-IR range, resulting in significant average temperature drop of 5°C and 12°C below the ambient temperature during daytime and nighttime, respectively. Moreover, the eggshell-based films exhibit excellent flexibility and self-cleaning properties, which are beneficial for practical long-term outdoor applications. Our proposed design provides a novel means for an environmentally friendly and sustainable management of the eggshell biowaste.

physics.optics

Molecular Understanding of the Effect of Hydrogen on Graphene Growth by Plasma-Enhanced Chemical Vapor Deposition

Plasma-enhanced chemical vapor deposition (PECVD) provides a low-temperature, highly-efficient, and catalyst-free route to fabricate graphene materials by virtue of the unique properties of plasma. In this paper, we conduct reactive molecular dynamics simulations to theoretically study the detailed growth process of graphene by PECVD at the atomic scale. Hydrocarbon radicals with different carbon/hydrogen (C/H) ratios are employed as dissociated precursors in the plasma environment during the growth process. The simulation results show that hydrogen content in the precursors significantly affects the growth behavior and critical properties of graphene. The highest number of hexagonal carbon rings formed in the graphene sheets, which is an indicator of their quality, is achieved for a C/H ratio of 1:1 in the precursors. Moreover, increasing the content of hydrogen in the precursors is shown to reduce the growth rate of carbon clusters, and prevent the formation of curved carbon structures during the growth process. The findings provide a detailed understanding of the fundamental mechanisms regarding the effects of hydrogen on the growth of graphene in a PECVD process.

physics.plasm-ph

Enhanced Thermal Transport across the Interface between Charged Graphene Electrodes and Poly(ethylene oxide) Electrolytes by Non-covalent Functionalization

Interfacial thermal transport between electrodes and polymer electrolytes can play a crucial role in the thermal management of solid-state lithium-ion batteries (SLIBs). Modifying the electrode surface with functional molecules can effectively increase the interfacial thermal conductance (ITC) between electrodes and polymers (e.g., electrolytes, separators); however, how they influence the interfacial thermal transport in SLIBs during charge/discharge remains unknown. In this work, we conduct molecular dynamics (MD) simulations to investigate the ITC between charged electrodes and solid-state polymer electrolytes (SPEs) mixed with ionic liquids (ILs). We find that ILs could self assemble at the electrode surface and act as non-covalent functional molecules that could significantly enhance the interfacial thermal transport during charge/discharge because of the formation of a densely packed cationic or anionic layer at the interface. While the electrostatic interactions between the charged electrode and the IL ions are responsible for forming these dense interfacial layers, the enhancement of ITC is mainly contributed by the increased Lennard-Jones (LJ) interactions between the charged electrodes and ILs. This work may provide useful insights into the understanding of interfacial thermal transport between electrodes and electrolytes of SLIBs during charge/discharge.

cond-mat.mes-hall

Machine learning and high-throughput robust design of P3HT-CNT composite thin films for high electrical conductivity

Combining high-throughput experiments with machine learning allows quick optimization of parameter spaces towards achieving target properties. In this study, we demonstrate that machine learning, combined with multi-labeled datasets, can additionally be used for scientific understanding and hypothesis testing. We introduce an automated flow system with high-throughput drop-casting for thin film preparation, followed by fast characterization of optical and electrical properties, with the capability to complete one cycle of learning of fully labeled ~160 samples in a single day. We combine regio-regular poly-3-hexylthiophene with various carbon nanotubes to achieve electrical conductivities as high as 1200 S/cm. Interestingly, a non-intuitive local optimum emerges when 10% of double-walled carbon nanotubes are added with long single wall carbon nanotubes, where the conductivity is seen to be as high as 700 S/cm, which we subsequently explain with high fidelity optical characterization. Employing dataset resampling strategies and graph-based regressions allows us to account for experimental cost and uncertainty estimation of correlated multi-outputs, and supports the proving of the hypothesis linking charge delocalization to electrical conductivity. We therefore present a robust machine-learning driven high-throughput experimental scheme that can be applied to optimize and understand properties of composites, or hybrid organic-inorganic materials.

physics.app-ph

Big enterprise registration data imputation: Supporting spatiotemporal analysis of industries in China

Big, fine-grained enterprise registration data that includes time and location information enables us to quantitatively analyze, visualize, and understand the patterns of industries at multiple scales across time and space. However, data quality issues like incompleteness and ambiguity, hinder such analysis and application. These issues become more challenging when the volume of data is immense and constantly growing. High Performance Computing (HPC) frameworks can tackle big data computational issues, but few studies have systematically investigated imputation methods for enterprise registration data in this type of computing environment. In this paper, we propose a big data imputation workflow based on Apache Spark as well as a bare-metal computing cluster, to impute enterprise registration data. We integrated external data sources, employed Natural Language Processing (NLP), and compared several machine-learning methods to address incompleteness and ambiguity problems found in enterprise registration data. Experimental results illustrate the feasibility, efficiency, and scalability of the proposed HPC-based imputation framework, which also provides a reference for other big georeferenced text data processing. Using these imputation results, we visualize and briefly discuss the spatiotemporal distribution of industries in China, demonstrating the potential applications of such data when quality issues are resolved.

cs.CY