Searcharxiv⌕ Search

subject

cs.DL

cs.DL: explore 44 source-linked works published from 2025 to 2026, with original documents and citations.

This collection is a preview while coverage and quality are evaluated.

Search within this collection

Coverage and selection

Includes records with this source-supplied label or an explicit phrase match in their metadata. Matches indicate a mention, not proof that a paper uses a method or tests a material. Source versions are consolidated by DOI.

Sources: arxiv. Collection updated 2026-09-16. Counts describe this index, not the complete source archives.

Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture

The Indonesian Digital Library of Culture (Perpustakaan Digital Budaya Indonesia, PDBI; budaya-indonesia.org) is a participatory platform that has collected tens of thousands of entries on Nusantara cultural heritage through public contribution since 2007. Manual contribution faces three structural barriers: coverage (knowledge is scattered across languages and sites), integrity (open sources mix authentic documentation with noise), and completeness (subjects are recorded but their data remain shallow). This paper presents a methodological framework for autonomous, AI-based harvesting of cultural knowledge from the open web, designed to expand corpus coverage while intensifying per-entry data depth. The methodology is organised as a five-stage economic funnel: focused crawling, multilingual extraction and canonicalisation, vector encoding with blocking, agentic decision-making, and idempotent publication, under the principle of deterministic orchestration, agentic decisions. Each stage is formalised: funnel economics and optimal filter ordering; crawl-frontier dynamics as a subcritical branching process that explains the necessity of recurrent re-seeding; fact-level novelty via a containment measure; Bayesian multi-source evidence fusion with elevated publication thresholds for sacred categories; exactly-once effects via idempotent upserts and the transactional outbox; sliding-window inference budgeting with a reservation protocol; statistical quality auditing; and seed selection as submodular coverage maximisation. The framework retains four high-value human roles: curator of direction, escalation approver, quality auditor, and guardian of meaning, while machine autonomy is raised in stages. Ethical, legal, and cultural-sensitivity implications are discussed, including the architectural guarantee that the machine never overwrites human contributions.

cs.AI↗

The conservative turn in science: The changing character of knowledge recombination

This study examines how the dominant mode of cross-disciplinary knowledge combination has changed over time. It addresses a limitation of existing studies that document declining disruptiveness at the level of individual works but do not reveal whether the mechanisms of cross-disciplinary knowledge flow have themselves shifted. Drawing on 56 million publications and 816 million citations from OpenAlex (1960-2025), the study classifies cross-disciplinary citations into five types based on the set relationship between the subfield portfolios of citing and cited publications and introduces a semantic-distance measure derived from SPECTER-v2 embeddings of publication abstracts. It subsequently demonstrates that cross-disciplinary citations have shifted markedly away from configurations with no shared disciplinary ground toward types built on partially overlapping knowledge bases, a reorientation further reflected in a steady decline in average semantic distance between citing and cited fields. This shift co-occurs with a diffusion of brokerage positions across the network and a greater reliance on older literature, consistent with a system that has become broader and more decentralized in its interdisciplinary reach. Finally, this study discusses the implications of these findings for science policy and for the design of cross-disciplinary funding mechanisms.

cs.SI↗

Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews

Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850 machine reviews of ICLR submissions from 2018 to 2025, all generated in one February-April 2025 window with one model family and one prompt, so that its year-to-year coefficients track submission composition alone and the human-minus-frozen trend difference identifies reviewer preference drift. On 32,638 submissions with 124,615 human reviews, the human coefficient on non-domain lexical complexity falls from +0.142 to -0.015 while the frozen rater moves from +0.080 to +0.082; the three-way difference-in-differences is -0.0100 (q=0.013), and forty random-wordlist placebos through the same specification centre on zero. Humans still reward sentence-length variability, which the frozen rater never registers, while the frozen rater still pays for lexical complexity at its earlier rate. Every claim is held to a double gate of false-discovery control and interval exclusion, and the findings that failed adversarial re-testing are reported. Reviewers discounted a cue whose production cost collapsed, as models of manipulable signals prescribe; an LLM judge calibrated to historical human preferences inherits the earlier schedule and drifts out of alignment while its agreement with humans on totals stays ordinary.

cs.CL↗

Recent Advances and Trends in Research Paper Recommender Systems: A Comprehensive Survey

As the volume of scientific publications grows exponentially, researchers increasingly face difficulties in locating relevant literature. Research Paper Recommender Systems have become vital tools to mitigate this information overload by delivering personalized suggestions. This survey provides a comprehensive analysis of Research Paper Recommender Systems developed between November 2021 and December 2024, building upon prior reviews in the field. It presents an extensive overview of the techniques and approaches employed, the datasets utilized, the evaluation metrics and procedures applied, and the status of both enduring and emerging challenges observed during the research. Unlike prior surveys, this survey goes beyond merely cataloguing techniques and models, providing a thorough examination of how these methods are implemented across different stages of the recommendation process. By furnishing a detailed and structured reference, this work aims to function as a consultative resource for the research community, supporting informed decision-making and guiding future investigations in the advances of effective Research Paper Recommender Systems.

cs.IR↗

Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions

Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on language priors. Comparing open-weight VLMs with traditional OCR baselines on low-resource Ancient Greek critical editions, we show that VLM errors often remain fluent even when wrong, producing plausible Greek substitutions where traditional engines produce local recognition noise. To analyze visual evidence during decoding, we introduce controlled image perturbations and token-level grounding measures based on conditional versus image-free decoding distributions. Under character-level perturbations, VLMs diverge sharply from the perturbed ground truth while traditional OCR remains comparatively faithful; however, token-level analysis shows that prior reliance is model-specific: in an OCR-specialist model, fluent lexical errors are produced with little reliance on the image, whereas general-purpose VLMs remain conditioned on the visual input even when wrong. Decode-time interventions fail to reliably restore grounding, while post-OCR language-model correction improves several systems only by repairing text after generation. Our results extend prior evidence of OCR language-prior reliance to low-resource historical documents and a broader set of models, showing that fluent output is not necessarily visually grounded and motivating interpretability-driven evaluation beyond aggregate accuracy.

cs.CL↗

An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026

Excess vocabulary, a word's frequency above its pre-2023 trend, is how the change in scholarly English after 2022 has been measured. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018-August 2026), with 47,165 Vietnamese abstracts for comparison. Placebo floors are 0.1-2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: sisahada "suggest" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like araboda "look into" fall to a quarter of trend. Under stated assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5%, 10.5% and 16.1% for 2024-2026 and a split-half set bound 7.8%, 20.6% and 33.0%. Holzwarth et al.'s estimator under the same discipline gives 41.9% and 72.1% for 2025-2026. Subject-matter controls reduce but do not remove it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 of the 33.0 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries none, the Korean shift persists at 30 to 66% of the rate where it does. Control abstracts from three providers reproduce the rising words, with marker turnover consistent with model generations; implied prevalences are scenario-dependent.

cs.CL↗

Same Problem, Different Field: Cross-Domain Solution Import via Domain-Stripped Computational Fingerprints

The same underlying computational problem is solved across unrelated fields under different names: recursive Bayesian state estimation appears as a "Kalman filter" in control, "Bayesian forecasting" in pharmacokinetics, and "data assimilation" in geoscience. Topical and citation-based scientific embeddings cannot see this shared problem. We distill each paper once into a domain- and method-name-stripped faceted computational fingerprint, a free-text mechanism skeleton plus controlled computational facets. We define a tunable, facet-selectable similarity over it. The goal is solution import: surface cross-field pairs solving the same problem, so a bespoke implementation can be swapped for another field's standard, specialized solver. On a benchmark of 18 method families across 109 papers, the skeleton lifts cross-domain retrieval average precision over the abstract from 0.222 to 0.513, and the whole fingerprint reaches 0.557. Strikingly, four trained scientific embedders all fall below plain abstract+TF-IDF: they encode topical and citation similarity, the wrong signal for this task. The gain is the representation: the abstract-to-skeleton swap lifts every embedder, and the pipeline is one cached LLM call per paper plus a cheap embedder. An interventional re-skin / math-edit test shows the fingerprint tracks the computation, not the field. On a 501-paper wild corpus, known twins dominate the top of the ranking (23 of the top 30); with planted pairs excluded from the results, three blind LLM judges rate 3 of the top 5 and 8 of the top 30 pairs genuine import candidates, and 0 of 30 random ones. The human verification is the four executed imports: in one, an open standard solver reproduces a bespoke clinical dosing engine's output. We release the benchmark, the code, and the distillation prompt.

cs.DL↗

The Art of Hierarchical Competing Patterns: Gaussian Process Optimization of Hyphenation

Hyphenation patterns remain a compact and widely deployed solution for word breaking in typesetting systems, text processors, and web rendering engines, but their generation still depends on manually tuned patgen program parameter profiles. We formulate patgen profile selection as a black-box hyperparameter optimization problem and evaluate Gaussian-process Bayesian optimization for this task. The search objective combines a precision-oriented F_{1/7}-score with an explicit trie size-accuracy trade-off using a normalized trie-size penalty. We evaluate the method on 17 hyphenated word-list datasets covering 14 languages and multiple scripts. Against two strong hand-tuned profiles regenerated from the same 8/10 training split and evaluated on the same 1/10 held-out test split, the GP-optimized profiles improve F_{1/7} on 16 of 17 datasets and reduce trie size on all 17. The median optimized/baseline trie ratio is 0.407. A dataset-level sign test gives p = 1.37e-4; a separate budget-matched comparison on five representative datasets shows that systematic search is competitive and usually improves over the best hand-tuned profile under the fixed comparison objective. The results show that model-based optimization can make pattern generation more reproducible and less dependent on expert trial-and-error while keeping the accuracy-compactness trade-off explicit.

cs.CL↗

From Citations to Contributions: LLM-Assisted Credit Scoring of Research Articles

Citation-based measures of scientific influence typically treat citations as uniform signals, ignoring the different roles that cited works play in a paper's contribution. We introduce contribution-based credit scoring for research articles: a structured citation analysis that decomposes a paper's credit between its own original contribution and the prior work it builds on. Motivated by a cooperative-game view of scientific credit, we propose the contribution tree, a hierarchical framework that conserves importance across the document structure and separates original from citation-derived contribution. To make this framework scalable, we use LLMs as noisy comparative estimators of local importance. We further extend the model to article collections by propagating contributions through weighted citation graphs, yielding corpus-level contributions and normalized influence scores. Our experiments suggest that our framework captures contribution signals beyond surface-level heuristics. Our code is available at https://github.com/sanaebrahimi/Importance_Scoring/

cs.DL↗

X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna

We demonstrate X-DigCheck, a domain-independent environment for building and maintaining application profiles as they co-evolve with the data they describe. Profiles developed against a fixed ontology quickly drift from the schema they were meant to capture. X-DigCheck treats profile construction as a continuous ontology-data co-evolution loop: data are lifted into RDF against the profile, checked through competency questions and SHACL, and the resulting reports jointly drive revisions of the ontology, mappings, constraints, and graph. The loop is agnostic to the domain and to the pipeline that produces the graph. We validate and demonstrate the tool in the cultural heritage domain, on the construction of RupeMagna-RTI, the first Reflectance Transformation Imaging (RTI) specialisation of the Cultural Heritage Survey ODP (CHS-ODP), aligned with CIDOC-CRM/CRMdig, ArCo, CHAD-KG, and Getty AAT, with semRTI as the lifting pipeline of this use case. The demonstration lets visitors run one full turn of the loop -on the shipped Rupe Magna (Grosio, Italy) RTI survey, or on a profile and graph of their own -executing the competency-question and SHACL checks live and reading the bidirectional coverage report that flags modelling gaps and stale assumptions. The result is a portable co-evolution environment for profile engineering, together with a reusable RTI application profile produced through it. A screencast of the demonstration is available at https://zenodo.org/records/22210609.

cs.DB↗

URL Extraction from Scholarly Documents: A Cross-Format Comparative Analysis

URLs in scholarly documents link to rich external resources such as datasets, software, publications, and websites. Extracting these URLs is crucial in the data preparation stage of many downstream tasks, such as link rot analysis, web crawling, and building knowledge graphs. However, existing studies often downplay this phase, simply extracting URLs from a single format, usually text directly converted from PDFs. We present a systematic study evaluating URL extraction across six input formats (text with annotation layer, LaTeX, HTML, XML, Markdown, and PNG converted from PDF). To support the evaluation, we compiled a benchmark dataset consisting of 2,338 manually annotated URLs from 200 arXiv papers spanning a wide range of domains over a 33-year period. In addition to evaluating individual file formats, we also compared 63 composite input-format combinations. Our extensive evaluations indicate that TEXTWAL achieves the best performance among single-format inputs, while TEXTWAL+LaTeX achieves the best overall URL extraction performance. The same trend is observed for URLs linking to open-access datasets and software. To further validate these findings, we apply our format-specific URL extraction pipelines to a longitudinal random sample of 364,744 arXiv papers spanning 33 years. We observe a sharp increase in URL density after 2015, along with remarkable differences in URL extraction across file formats over time. Overall, our study highlights the importance of selecting an appropriate format for URL extraction from scholarly documents. The dataset and code are publicly available at: https://github.com/lamps-lab/arxiv-url-bench .

cs.DL↗

AI-Assisted Writing Is Growing Fastest Among Less Established Scientists in Non-English-Speaking Countries

The recent emergence of AI-assisted writing raises an important question: how is this new technology being adopted across the scientific community, and how does adoption vary across linguistic and professional contexts? We analyze over two million full-text biomedical publications from PubMed Central from 2021 to 2024 using a distribution-based framework to estimate AI-generated content. We found that, in biomedical publications, AI-generated content increased substantially after ChatGPT, with larger increases in publications from countries with lower English proficiency. Increases were also greater among scientists with fewer publications and citations, those at earlier career stages, and those at lower-ranked institutions. Prior AI research experience was associated with greater increases in AI-assisted writing, which were also modestly associated with greater increases in publication productivity. These findings show that AI-assisted writing is growing fastest among biomedical scientists who may have historically faced barriers, a pattern with potentially positive implications for equity in science.

cs.DL↗

Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI

Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not equal valid explanation. The deepest risk is not factual error alone but the appearance that an explanation is already established without clear sources, page numbers, editions, or evidence. We liken the page anchor to Ariadne's thread: within the labyrinth of generative fluency, it is the thread that leads the scholar back to the source. This paper proposes Traceable Scholarship as the minimum normative condition for AI-assisted humanistic research, situating it across the three revolutions of knowledge infrastructure: print, digital, and generative AI. We introduce page anchors, dual page numbers, citation-first generation, NO_EVIDENCE, human verification, four-level compliance, and Scope Contract, and present AIH-Infra as a three-layer reference implementation: Contexture (document structuring), Open WebUI AIH-Infra (traceable knowledge base), and AIH-Infra MCP Server (agent gateway). A case study on a 29-volume Kant Akademie-Ausgabe knowledge base illustrates how traceability supports retrieval correction, evidence grading, and judgment downgrading. Traceability is not a software feature; it is the condition under which humanistic research can remain public and refutable in the age of generative AI.

cs.AI↗

SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation

Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to limited reference coverage and, more importantly, fails to replicate the expert-driven revision process that is crucial for writing high-quality surveys. In this paper, we introduce SurveyAgent-HKA, a multi-agent framework that improves end-to-end scientific survey generation by incorporating knowledge derived from published surveys and peer-review comments. The framework decomposes survey generation into well-defined sub-tasks handled by LLM-powered agent. It first retrieves relevant papers from multiple sources and identifies key topics through clustering to construct an initial outline, which is then refined using outlines from related human-written surveys. Based on the refined outline, topic-focused papers are retrieved and re-ranked to select for drafting a well-grounded survey. Then, we identify common issues raised by experts in peer-review comments from published surveys to guide the revisions and finalize the survey. Experiments on two domains show that our approach outperforms mainstream baselines in citation quality, structural consistency, and content quality. Furthermore, our framework is efficient in both time and cost, making it a practical solution for broader AI-assisted scientific writing applications.

cs.CL↗

VietProfs: A Public Directory of the Vietnamese Academic and Research Diaspora

The Vietnamese academic diaspora spans hundreds of universities and research institutes worldwide, yet there has never been a shared, searchable directory for this community. Prospective graduate students looking for advisors who share their background, researchers seeking collaborators, and organizers looking for speakers have had to rely on manual, one-off web searches. To address this need, we built VietProfs (https://vietprofs.roars.dev), an open, searchable directory of Vietnamese and Vietnamese-diaspora faculty and permanent research scientists outside Vietnam. Each profile pairs a portrait and authoritative diacritic name with verified appointment, degree, and honors data, browsable through a fast client-side web interface. This paper covers two topics: the design of the directory itself and the engineering required to keep it current. Part I describes the directory: its goals, eligibility and appointment track rules, alignment with standard U.S. federal classifications (NCES, NSF, and NIH), profile card design, and current roster insights. Part II reports on automated directory maintenance. Keeping over a thousand records accurate against a constantly shifting web is notoriously difficult. We describe an LLM-assisted maintenance system combining automated search and periodic revalidation with deterministic validation gates, operating under the principle: AI proposes, evidence decides. As of September 2026, the directory includes 1,152 verified records across 492 institutions in 23 countries, covering doctoral cohorts from 1939 to 2026. Aggregate statistics are computed dynamically at runtime, and all changes are audited through public Git history. VietProfs provides a practical community resource while offering a concrete case study in building reliable, agent-maintained public datasets.

cs.DL↗

Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin

Medieval documentary sources remain inadequately served by existing natural language processing tools. None of the five readily available Latin treebank models attains usable performance on a collection of 160 inventories compiled in Marseille between 1258 and 1446. The best labelled attachment score is 0.62 and the best morphology-aware score is 0.24. Performance does not correlate with either genre or period proximity. To address this shortfall, in-domain training data was generated as a by-product of using these inadequate models. In each of nine iterations, a model pre-annotated 200 sentences; an expert corrected the annotations; and the corrected sentences were used to train the subsequent model, with batches sampled independently of model state, without active-learning selection. Thirty-three hours of annotation effort over 1,804 sentences increased universal part-of-speech accuracy from 0.80 to 0.98 and labelled attachment from 0.48 to 0.92, outperforming all baselines on the reported metrics while using 97% less training data than the largest one of them. Annotator effort declined from 54% of tokens to a plateau of 14-18%, an operational progress metric that requires no separate gold standard and can serve as a stopping criterion.

cs.CL↗

The Generative AI Gold Rush in Theoretical and Computational Research

Generative AI is changing the production conditions of theoretical and computational research, but its sys tem level effects require measures that separate plat form growth, field specific divergence, and production structure. We assemble 2,080 monthly observations for twenty arXiv archives from January 2018 through Au gust 2026 and a separate pseudonymized Mathematics author panel. A regularized convex synthetic control fitted through December 2025 identifies the January August 2026 anomaly, while spatial placebos, prior year pseudo holdouts, donor refits, and alternative preperiods assess comparative robustness. Mathematics recorded 47,127 list entries, 33.5% above 2025 and 11.9% above a synthetic counterfactual of 42,113 entries. Qualified donor and preperiod designs yield 9.6% to 14.9%, and Mathematics has the largest RMSPE ratio among fifteen eligible placebo archives. Subfield growth is broad, with 29 of 30 primary math.* categories expanding. The author panel shows a marked thickening of the repeated output tail. The share of active author units produc ing at least five submissions rose from 2.45% to 3.80%, while the ten submission tail rose from 0.21% to 0.49%. These results document a new and unusually large 2026 Mathematics production regime shift. Its timing and production structure, combined with independent evi dence on AI diffusion and verifiable research tasks, are consistent with delayed diffusion and capability thresh old mechanisms. The comparative design identifies the anomaly, and separate triangulation evaluates AI related explanations. The findings locate verification, selection, and attention as central constraints for research gover nance.

cs.DL↗

Westlake Scholar: AI-Enhanced Scholarly Discovery over an Institutional Repository

Institutional repositories (IRs) provide mature infrastructure for preserving and disseminating research outputs, but conventional record- and document-centric interfaces provide limited support for connecting deposited papers to related research and people. We present Westlake Scholar, an open-source, institution-grounded platform that adds four complementary artificial intelligence (AI) services to repository infrastructure: contextual paper reading, research-direction-guided paper discovery, publication-grounded expert discovery, and AI-generated research chronologies for scholars. The services draw on a shared institutional knowledge layer connecting approved publication records, paper content, and scholar--publication relationships. This allows the same paper to support contextual reading, cross-paper discovery, expert matching, and longitudinal views of scholarly work. Westlake Scholar provides an open and governable implementation of an institution-controlled AI layer that connects repository content, scholarly discovery, and researcher relationships while preserving provenance, human review, and institutional governance. A deployment at Westlake University, in operation since April 2026, demonstrates that the integrated system can operate in a live institutional setting.

cs.DL↗
Compare source metadata on this page
WorkPublishedSource identifierSource
Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture2026-09-082609.08105arxiv
The conservative turn in science: The changing character of knowledge recombination2026-09-082609.08468arxiv
Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews2026-09-082609.08475arxiv
Recent Advances and Trends in Research Paper Recommender Systems: A Comprehensive Survey2026-09-072508.08828arxiv
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions2026-09-072605.27750arxiv
An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-20262026-09-072609.07447arxiv
Same Problem, Different Field: Cross-Domain Solution Import via Domain-Stripped Computational Fingerprints2026-09-152609.07595arxiv
The Art of Hierarchical Competing Patterns: Gaussian Process Optimization of Hyphenation2026-09-072609.07638arxiv
From Citations to Contributions: LLM-Assisted Credit Scoring of Research Articles2026-09-072609.07673arxiv
X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna2026-09-292609.07694arxiv
URL Extraction from Scholarly Documents: A Cross-Format Comparative Analysis2026-09-072609.08019arxiv
AI-Assisted Writing Is Growing Fastest Among Less Established Scientists in Non-English-Speaking Countries2026-09-062511.15872arxiv
Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI2026-09-052607.20916arxiv
SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation2026-09-052609.05938arxiv
VietProfs: A Public Directory of the Vietnamese Academic and Research Diaspora2026-09-052609.06091arxiv
Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin2026-09-052609.06266arxiv
The Generative AI Gold Rush in Theoretical and Computational Research2026-09-042609.04872arxiv
Westlake Scholar: AI-Enhanced Scholarly Discovery over an Institutional Repository2026-09-042609.05072arxiv

These are bibliographic comparisons, not experimental rankings. Follow the original document for methods and conditions.