SearcharxivSearch

arXiv subjects

Zhengyi Zhou

Publications and source records attributed to Zhengyi Zhou.

At least 19 recordsLinked to original sources

The Landscape of problematic papers in the field of non-coding RNA

Retractions have increased sharply in recent years, alongside a growing number of papers that receive post-publication comments questioning their reliability (commented papers). Together, retracted and commented papers undermine the credibility of scientific research and may also threaten public health. In this study, we examine problematic papers in the field of non-coding RNA (ncRNA) from multiple perspectives to identify common patterns and inform strategies for addressing large-scale fraudulent publications. We find that studies on under-investigated ncRNAs are more likely to become problematic papers. These papers often show substantial textual similarity, and many additional papers with similar text also display suspicious image duplication. Healthcare institutions, particularly those with lower publication output, appear especially vulnerable to producing such papers. Most problematic papers are concentrated in a small set of journals, many of which do not adequately address concerns raised after publication. Overall, our findings indicate that a substantial number of problematic papers may remain undetected and that their shared characteristics can support more effective strategies for identifying and curbing large-scale fraudulent publications.

cs.DL

Mapping Academic Integrity: Global Retraction Trends Explored through a Topic Lens

Scientific publications have long served as the cornerstone of innovation, exhibiting stable growth over the years. Recently, however, retractions have surged dramatically, driven largely by the proliferation of low-quality and fraudulent articles, posing a substantial threat to research integrity. By integrating annual publication and retraction data, this study employs the relative retraction rate (R3) to systematically examine disparities and evolving trends from a topical perspective. Our analysis reveals that the number of retractions has grown significantly faster than that of global publications, yielding an overall retraction rate of 0.12%. While retractions occur across all disciplines, substantial disparities exist, ranging from 0.035% in Physics to 0.34% in Computer Science. This gap widens at finer levels of granularity, reaching roughly 8.99% in Human-Computer Interaction. Moreover, unusually high R3 values frequently coincide with rapid publication growth in specific fields. We also developed Retraction Monitor, a web application for monitoring retraction dynamics across diverse fields, enabling stakeholders to visualize these trends and assess risks to research integrity. These findings provide valuable insights for identifying high-risk fields and developing tailored governance policies to strengthen research rigor and mitigate field-specific retraction risks.

cs.DL

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification

Knowledge-Based Visual Question Answering (KB-VQA) requires grounding visual queries to external knowledge beyond directly observable content in images. While recent multi modal large language models (MLLMs) show strong perceptual abilities, they struggle on KB-VQA tasks requiring groundings from both fine-grained entity and evidence levels. Most existing multi-modal retrieval augmented generation (MM-RAG) methods tightly couple entity discrimination and section-level evidence ranking into a single re-ranking stage, leading to high cost and limited generalization. In this work, we revisit existing MM-RAG solutions from a workflow perspective and argue both entity-level and fact-level groundings are key bottlenecks. We observe that although MLLMs often fail under open-ended entity naming, they can better identify the correct entity when selecting from a small set of candidate names. Based on this insight, we propose a simple and training-free identify-before-answer IBA framework that decouples entity identification from section-level re-ranking. Our approach prompts an MLLM to select high-confidence entities using only candidate names, followed by an off-the-shelf textual re-ranker for evidence selection. Experiments on Encyclopedic-VQA and InfoSeek show that our method consistently outperforms fine-tuned multi-modal re-ranking baselines while reducing training and inference complexity. Additional analyses reveal that the improvements arise not only from better entity identification, but also from selecting more informative evidence once correct entity is fixed. Our implementation is made public to ease reproducibility.

cs.CL

Differentially Private Hierarchical Heavy Hitters

The task of finding _Hierarchical_ Heavy Hitters (HHH) was introduced by Cormode et al. [VLDB 2003] as a generalisation of the heavy hitter problem. While finding HHH in data streams has been studied extensively, the question of releasing HHH when the underlying data is private remains unexplored. In this paper, we study differentially private HHH release in both the streaming and non-streaming setting. In the non-streaming setting, we show the surprising result that the relative error in estimating the residual count for any prefix is independent of the height of the hierarchy and the number of heavy hitters in the stream. Meanwhile, in the streaming setting, although the exact version of HHH has low global sensitivity (as counting queries are 1-sensitive), the approximation functions due to streaming have high global sensitivity, linear in the available space. Despite this obstacle, we show that the absolute error for estimating frequencies in the steaming setting is independent of the available space.

cs.CR

Unifying and Optimizing Data Values for Selection via Sequential Decision-Making

Data selection has emerged as a crucial downstream application of data valuation, yet the theoretical foundations for using data values in selection remain underexplored. We reformulate data selection as a sequential decision-making problem where the optimal selection sequence arises from dynamic programming, and data values can be understood as encodings of this optimal sequence. This framework unifies and reinterprets existing methods like Data Shapley through the lens of approximate dynamic programming, revealing them as myopic linear approximations to the sequential problem. We further analyze how selection optimality degrades with utility curvature under submodularity, explaining when and why these approximations fail. To bridge theory and practice, we propose an efficient bipartite graph-based surrogate that preserves submodular structure while enabling scalable greedy selection with provable guarantees. Experiments on classical ML benchmarks and large-scale LLM fine-tuning data selection demonstrate substantial improvements over existing methods. Code is publicly available at https://github.com/frankhlchi/SeqDataVal

cs.AI

SOMA: Efficient Multi-turn LLM Serving via Small Language Model

Large Language Models (LLMs) are increasingly deployed in multi-turn dialogue settings where preserving conversational context across turns is essential. A standard serving practice concatenates the full dialogue history at every turn, which reliably maintains coherence but incurs substantial cost in latency, memory, and API expenditure, especially when queries are routed to large proprietary models. Existing approaches often struggle to balance the trade-off between response quality and efficiency. We propose a framework that exploits the early turns of a session to estimate a local response manifold and then adapt a smaller surrogate model to this local region for the remainder of the conversation. Concretely, we learn soft prompts that maximize semantic divergence between the large and surrogate small language models' responses to surface least-aligned local directions, stabilize training with anti-degeneration control, and distill the mined cases into localized LoRA fine-tuning so the surrogate runs without prompts at inference. A simple gate enables a one-time switch with rollback on drift. We further provide a theoretical analysis for key components in SOMA. Extensive experiments show the effectiveness of SOMA. The source code is provided at: https://github.com/LabRAI/SOMA.

cs.CL

SortingHat: Redefining Operating Systems Education with a Tailored Digital Teaching Assistant

Operating Systems (OS) courses are among the most challenging in computer science education due to the complexity of internal structures and the diversity of running environments. Traditional teaching methods often fail to address the diverse backgrounds, learning speeds, and practical needs of students. To tackle these challenges, we present SortingHat, a personalized digital teaching assistant tailored specifically for OS education. SortingHat integrates advanced AI technologies, including a retrieval augmented generation (RAG) framework and multi agent reinforcement learning (MARL), to deliver adaptive, scalable, and effective educational support. SortingHat features a 3D digital human interface powered by large language models (LLMs) to provide personalized, empathetic, and context aware guidance. It generates tailored exercises based on each student's learning history and academic performance, reinforcing weak areas and challenging advanced concepts. Additionally, the system incorporates a robust evaluation pipeline that ensures fair, consistent, and unbiased grading of student submissions while delivering personalized, actionable feedback for improvement. By combining personalized guidance, adaptive content creation, and automated assessment, SortingHat transforms OS education into an engaging, immersive, and scalable experience.

cs.HC

Contact $(+1)$-surgeries and algebraic overtwistedness

We show that a contact $(+1)$-surgery along a Legendrian sphere in a flexibly fillable contact manifold ($c_1=0$ if not subcritical) yields a contact manifold that is algebraically overtwisted if the Legendrian's homology class is not annihilated in the filling. Our construction can also be implemented in more general contact manifolds yielding algebraically overtwisted manifolds through $(+1)$-surgeries. This gives new proof of the vanishing of contact homology for overtwisted contact manifolds. Our result can be viewed as the symplectic field theory analog in any dimension of the vanishing of contact Ozsváth-Szabó invariant for $(+1)$-surgeries on two-component Legendrian links proved by Ding, Li, and Wu.

math.SG

Tight contact structures without symplectic fillings are everywhere

We show that for all $n \ge 3$, any $(2n+1)$-dimensional manifold that admits a tight contact structure, also admits a tight but non-fillable contact structure, in the same almost contact class. For $n=2$, we obtain the same result, provided that the first Chern class vanishes. We further construct Liouville but not Weinstein fillable contact structures on any Weinstein fillable contact manifold of dimension at least $7$ with torsion first Chern class.

math.SG

Algebraic planar torsion in contact manifolds

We demonstrate that the functorial properties of the symplectic field theory under strong cobordisms and surgery cobordisms can produce finite algebraic (planar) torsions from simple examples, which gives a unified treatment of most of the known computations of algebraic (planar) torsions. In addition, we obtain many families of new examples, notably including (1) stably fillable examples in all dimensions $\ge 5$ with algebraic (planar) torsion precisely $k$ for any given $k\in \mathbb{N}_+$, confirming a conjecture of Latschev and Wendl; (2) contact structures on spheres of all dimensions at least $5$ with finite algebraic planar torsion at least $1$, which implies that tight not weakly fillable contact structures are ubiquitous in higher dimensions. We also explain that all known examples of contact manifolds without strong/weak fillings in dimension $\ge 5$ have algebraic planar torsion.

math.SG

StackPilot: Autonomous Function Agents for Scalable and Environment-Free Code Execution

Recent advances in large language models (LLMs) have substantially enhanced automated code generation across a wide range of programming languages. Nonetheless, verifying the correctness and executability of LLM-generated code remains a significant challenge, as traditional methods rely on language-specific compilers and environment-dependent runtimes. To overcome these limitations, we introduce StackPilot, an LLM-native, multi-agent framework designed for language-agnostic code verification and execution, which operates independently of conventional toolchains. StackPilot offers three principal innovations: (1) a Function-as-Agents paradigm, in which each function is modeled as an autonomous agent capable of fine-grained reasoning and collaborative verification; (2) an LLM-as-Executor strategy, which enables scalable verification via stack-based scheduling; and (3) a novel snapshot mechanism that preserves complete execution contexts, facilitating deterministic and lossless context switching during verification. Empirical evaluations demonstrate that StackPilot achieves framework reliability rates between 89% and 97%, substantially outperforming baseline approaches. These results indicate that StackPilot can reliably verify and execute a significantly larger proportion of LLM-generated code across diverse programming tasks compared to existing methods.

cs.PL

RSFT functors for strong cobordisms and applications

We extend the hierarchy functors of [33] to the case of strong symplectic cobordisms, via deformations with Maurer--Cartan elements. In particular, we prove that the concave boundary of a strong cobordism has finite algebraic planar torsion if the convex boundary does, which yields a functorial proof of finite algebraic planar torsion for contact manifolds admitting strong cobordisms to overtwisted contact manifolds. We also show the existence of contact $3$-folds without strong cobordisms to the standard contact $3$-sphere, that are not cofillable. We also include generalizations of the theory relating our notion of algebraic planar torsion to Latschev--Wendl's notion of algebraic torsion, discussing variations from counting holomorphic curves with general constraints and invariants extracted from higher genera holomorphic curves from an algebraic perspective.

math.SG

LLmFPCA-detect: LLM-powered Multivariate Functional PCA for Anomaly Detection in Sparse Longitudinal Texts

Sparse longitudinal (SL) textual data arises when individuals generate text repeatedly over time (e.g., customer reviews, occasional social media posts, electronic medical records across visits), but the frequency and timing of observations vary across individuals. These complex textual data sets have immense potential to inform future policy and targeted recommendations. However, because SL text data lack dedicated methods and are noisy, heterogeneous, and prone to anomalies, detecting and inferring key patterns is challenging. We introduce LLmFPCA-detect, a flexible framework that pairs LLM-based text embeddings with functional data analysis to detect clusters and infer anomalies in large SL text datasets. First, LLmFPCA-detect embeds each piece of text into an application-specific numeric space using LLM prompts. Sparse multivariate functional principal component analysis (mFPCA) conducted in the numeric space forms the workhorse to recover primary population characteristics, and produces subject-level scores which, together with baseline static covariates, facilitate data segmentation, unsupervised anomaly detection and inference, and enable other downstream tasks. In particular, we leverage LLMs to perform dynamic keyword profiling guided by the data segments and anomalies discovered by LLmFPCA-detect, and we show that cluster-specific functional PC scores from LLmFPCA-detect, used as features in existing pipelines, help boost prediction performance. We support the stability of LLmFPCA-detect with experiments and evaluate it on two different applications using public datasets, Amazon customer-review trajectories, and Wikipedia talk-page comment streams, demonstrating utility across domains and outperforming state-of-the-art baselines.

stat.ML

UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation

Articulated 3D objects play a vital role in realistic simulation and embodied robotics, yet manually constructing such assets remains costly and difficult to scale. In this paper, we present UniArt, a diffusion-based framework that directly synthesizes fully articulated 3D objects from a single image in an end-to-end manner. Unlike prior multi-stage techniques, UniArt establishes a unified latent representation that jointly encodes geometry, texture, part segmentation, and kinematic parameters. We introduce a reversible joint-to-voxel embedding, which spatially aligns articulation features with volumetric geometry, enabling the model to learn coherent motion behaviors alongside structural formation. Furthermore, we formulate articulation type prediction as an open-set problem, removing the need for fixed joint semantics and allowing generalization to novel joint categories and unseen object types. Experiments on the PartNet-Mobility benchmark demonstrate that UniArt achieves state-of-the-art mesh quality and articulation accuracy.

cs.CV

Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking

Knowledge Graph Question Answering (KGQA) systems rely on high-quality benchmarks to evaluate complex multi-hop reasoning. However, despite their widespread use, popular datasets such as WebQSP and CWQ suffer from critical quality issues, including inaccurate or incomplete ground-truth annotations, poorly constructed questions that are ambiguous, trivial, or unanswerable, and outdated or inconsistent knowledge. Through a manual audit of 16 popular KGQA datasets, including WebQSP and CWQ, we find that the average factual correctness rate is only 57 %. To address these issues, we introduce KGQAGen, an LLM-in-the-loop framework that systematically resolves these pitfalls. KGQAGen combines structured knowledge grounding, LLM-guided generation, and symbolic verification to produce challenging and verifiable QA instances. Using KGQAGen, we construct KGQAGen-10k, a ten-thousand scale benchmark grounded in Wikidata, and evaluate a diverse set of KG-RAG models. Experimental results demonstrate that even state-of-the-art systems struggle on this benchmark, highlighting its ability to expose limitations of existing models. Our findings advocate for more rigorous benchmark construction and position KGQAGen as a scalable framework for advancing KGQA evaluation.

cs.CL

Kähler compactification of $\mathbb{C}^n$ and Reeb dynamics

Let $X$ be a smooth complex manifold. Assume that $Y\subset X$ is a Kähler submanifold such that $X\setminus Y$ is biholomorphic to $\mathbb{C}^n$. We prove that $(X, Y)$ is biholomorphic to the standard example $(\mathbb{P}^n, \mathbb{P}^{n-1})$. We then study certain Kähler orbifold compactifications of $\mathbb{C}^n$ and, as an application, prove that on $\mathbb{C}^3$ the flat metric is the only asymptotically conical Ricci-flat Kähler metric whose metric cone at infinity has a smooth link. As a key technical ingredient, we derive a new characterization of minimal discrepancy of isolated Fano cone singularities by using $S^1$-equivariant positive symplectic homology.

math.DG

Unknottedness of symplectic submanifold fillings

We show that any symplectic filling of the standard contact submanifold $(\mathbb{S}^{2n-1},ξ_{\mathrm{std}})$ of $(\mathbb{S}^{2n+1},ξ_{\mathrm{std}})$ in $(\mathbb{D}^{n+1},ω_{\mathrm{std}})$ is smoothly unknotted if $n\ge 2$. We also give a self-contained proof of the Siefring intersection formula between punctured holomorphic curves and holomorphic hypersurfaces used in the proof using the $L$-simple setup of Bao-Honda.

math.SG