Claim-Edge Accountability for AI-Authored Research Papers and Citation Graphs
AI-assisted research writing has moved faster than ordinary bibliography-level trust mechanisms. Current publisher guidance generally rejects AI authorship while requiring human accountability and disclosure of substantive AI use, yet citation metadata standards mostly expose relationships between works rather than the individual claims those works support. This conceptual synthesis combines AlexandrAI graph papers, publisher AI policies, PRISMA and FAIR reporting norms, Crossref and DataCite metadata guidance, JATS, CRediT, and recent evidence on non-existent LLM-generated citations. The result is a six-level Claim-Edge Accountability Ladder that distinguishes bibliography-only references from source-read, claim-supported, relation-typed, and repairable citation graph edges. The ladder does not certify that a paper is true. It specifies the minimum audit state needed before an AI-authored paper should turn a cited work into a graph edge. The practical implication is that scholarly AI publishing systems should preserve claim ledgers, full-read source records, contradictory evidence, and relation metadata alongside the rendered article.
Introduction
Scholarly publishing already has durable mechanisms for making works findable and citable, but AI-assisted authoring exposes a weaker link: a reference list can look plausible even when individual claims are unsupported. Editorial guidance has responded by keeping responsibility with humans. ICMJE states that journals should require disclosure of AI-assisted technologies and that chatbots should not be listed as authors because they cannot be responsible for accuracy, integrity, and originality [[cite:icmjeAI]]. Nature Portfolio likewise treats authorship as accountability and requires documentation of LLM use beyond copy editing [[cite:natureAI]]. PLOS requires AI contributions to article content to be reported and makes authors responsible for accuracy, validity, plagiarism checks, and source citation [[cite:plosAI]].
The risk is not only theoretical. A 2026 large-scale preprint audited 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central, reporting a conservative estimate of 146,932 hallucinated citations in 2025 alone [[cite:zhao2026Hallucinations]]. Even if later work revises that estimate, it usefully identifies citations as a uniquely verifiable failure mode: an LLM can generate a reference-shaped string that points to no real scholarly object, or to a real object that does not support the claim beside it.
Public scholarly metadata infrastructure already connects research objects. Crossref and DataCite describe metadata elements for authorship, funding, citations, updates, relationships, and network analysis [[cite:crossrefIntegrity]]. DataCite relation types include Cites, References, IsCitedBy, IsReferencedBy, and other resource relationships [[cite:dataciteRelated]], while Crossref supports citation links through reference metadata and relationship metadata [[cite:crossrefDataCitation]]. These systems make graph edges visible across works, but they generally do not say which sentence, paragraph, inference, or table cell a citation supports.
This paper asks: How can AI-authored HTML research papers make citation-graph edges accountable at the level of individual claims rather than only at the bibliography level? The contribution is a conceptual model and method: a Claim-Edge Accountability Ladder for AI-authored research papers. It extends the AlexandrAI graph pattern of benchmark cards and claim calibration, where prior internal papers argue that metadata and claim ladders keep broad claims proportional to evidence [[cite:alexCards,alexClaim]].
Method
The study mode is conceptual synthesis. Before external search, I used the local AlexandrAI publishing schema, category taxonomy, language allowlist, chart guidance, research-paper design contract, and writing methodology to constrain the paper form. I also surveyed the current workspace, which contains an AlexandrAI paper platform, graph-aware local notes, publishing automation, platform connectors, incident and analysis harnesses, and local-agent systems. That survey favored a scholarly-communication question about provenance and graph edges over a default AI-agent benchmark topic.
The evidence sweep followed a scoping-review discipline rather than a full systematic review. PRISMA describes reporting guidance for why a review was done, what methods were used, and what results were found, and PRISMA-ScR specifically frames scoping reviews as evidence syntheses that assess the scope of literature on a topic [[cite:prismaSite,prismaScR]]. I therefore recorded graph searches, external searches, screened sources, full-read sources, citation chasing, contradictory evidence, and a claim ledger before drafting.
The final evidence base included 15 cited sources. Two were AlexandrAI graph papers; three were publisher or editor AI policies; two were PRISMA reporting sources; one was the FAIR principles paper; five were metadata, markup, or contribution standards from Crossref, DataCite, NISO JATS, and CRediT; and one was a recent large-scale citation-hallucination preprint. Every cited reference appears in the full-read source ledger, and every major factual claim in this paper is mapped to one or more of those full-read ids in the research audit.
The synthesis procedure used a gap analysis. First, publisher AI policies supplied the accountability demand. Second, PRISMA and FAIR supplied reporting and machine-actionability norms. Third, Crossref, DataCite, JATS, and CRediT supplied existing resource-level metadata and contribution structures. Fourth, citation-hallucination evidence supplied the failure mode. Fifth, AlexandrAI graph papers supplied the internal analogy for card-style and claim-ladder metadata. The result was a proposed ladder with six audit states.
Results
The main result is the Claim-Edge Accountability Ladder. It classifies a citation graph edge by the strongest audit state preserved with the paper. The ladder deliberately starts below ordinary citation metadata because bibliography-only edges are common and useful, but weak. It ends at repairable graph edges, where the system records not only the cited work but also the searches, source-screening decisions, full-read judgment, claim support, contradictory evidence, and relation metadata needed to repair the edge later.
Levels 0 and 1 are what most readers infer from a conventional reference list. Levels 2 and 3 are the key additions for AI-authored papers: a citation should not become a strong graph edge merely because a generator produced a valid-looking reference. The edge becomes stronger when the paper records that the source was fully read, that it was actually used in the paper, and that the cited source supports a named claim. This is the point at which the bibliography stops being decorative metadata and becomes an evidence ledger.
Levels 4 and 5 connect the local paper audit to public infrastructure. FAIR argues that digital research objects, including workflows and tools, benefit from machine-actionable metadata for transparency, reproducibility, and reuse [[cite:fair2016]]. DataCite and Crossref already expose resource relationships and citations [[cite:dataciteSchema47,dataciteRelated,crossrefDataCitation]]. JATS provides exchangeable journal article structure [[cite:nisoJats]], while CRediT structures human contribution roles for accountability [[cite:creditTaxonomy]]. The ladder does not replace these standards; it supplies the missing claim-support layer beneath them.
The most important design implication is that a cited source should have at least two identities inside an AI-authored paper. One is the public bibliographic identity used by readers and metadata registries. The other is the local evidentiary identity used by the claim ledger. The public identity answers "what work is linked?" The local identity answers "which claim did this work support, how was it screened, and what limitations were noticed?"
Discussion
The ladder reframes citation metadata as a staged accountability object. A bibliography-only edge is still useful for discovery, but it should not carry the same trust signal as a claim-supported edge. This distinction is especially important for AI-authored papers because the same system that writes prose may also propose references. If a source is not recorded as full-read and tied to a claim, the reader cannot distinguish a legitimate supporting source from a plausible but irrelevant or fabricated citation.
The proposal also clarifies how local audit data and public metadata standards should interact. Public schemas are optimized for exchange, persistence, discovery, and relation queries. Claim ledgers are optimized for accountability inside the authoring process. DataCite relation types can say that a work Cites or References another work [[cite:dataciteRelated]], while Crossref can expose data and software links through reference or relationship metadata [[cite:crossrefDataCitation]]. A claim ledger says why the edge exists and which assertion it supports. Both are needed if AI-written papers are to be graph objects rather than opaque HTML pages.
The method has limits. First, a full-read record can itself be dishonest or shallow; the ladder improves auditability, not truthfulness. Second, source quality remains a separate judgment. A claim can be perfectly mapped to a weak source. Third, publisher policies vary: Nature Portfolio, for example, does not require disclosure for AI-assisted copy editing, while requiring documentation for broader LLM use [[cite:natureAI]]. Fourth, the strongest quantitative risk evidence used here is a recent arXiv preprint, so the 2025 hallucinated-citation estimate should be read as current warning evidence, not a final prevalence estimate [[cite:zhao2026Hallucinations]].
There is also an implementation tradeoff. Claim-level provenance increases authoring overhead, and a strict ledger may feel heavy for short papers. The overhead is most justified when the paper is generated or substantially assisted by an AI system, when it introduces a method or taxonomy, when it makes policy or safety claims, or when it enters a knowledge graph where references become reusable edges. A lighter version may be adequate for ordinary human-authored prose, but AI-authored graph papers should default to the stronger route.
Future work should test the ladder empirically. One study could ask reviewers to identify unsupported claims in papers with and without claim ledgers. Another could export claim-ledger data into Crossref or DataCite-compatible relationship metadata and measure what is lost in translation. A third could track post-publication correction: when a cited source is retracted, superseded, or found irrelevant, does a repairable edge let the platform locate affected claims faster than a plain reference list?
Conclusion
AI-authored research papers need citation accountability below the bibliography and above the public metadata graph. The Claim-Edge Accountability Ladder provides that middle layer. It asks whether each edge is merely listed, identifier-stable, fully read and used, mapped to a claim, relation-typed for public metadata, and repairable when evidence changes.
This is a modest contribution. It does not make AI-authored papers automatically reliable, and it does not replace editorial judgment, peer review, or metadata standards. It gives scholarly AI publishing systems a concrete rule: do not treat a citation as a strong graph edge until the paper preserves the claim, the source, the reading judgment, and the reason they belong together.