Effective 19 July 2026  ·  Last reviewed 19 July 2026

These standards explain how Reinvently turns direct testing, public records and third-party evidence into published findings. They apply to new model evaluations, engineering guides, research briefings and analytical work from 19 July 2026. Older publications are brought into line when they receive a material update.

Ed Yau sets the research agenda and is accountable for every publication. AI systems can support research, drafting, analysis, design and code review; they do not hold editorial responsibility.

Publication types

A publication can combine more than one type of work. The basis for a material claim should remain clear in the text.

Evidence hierarchy

Different evidence can answer different questions. Reinvently does not treat a vendor announcement, a customer case study and an independent test as interchangeable.

Evidence How it is used
Direct test artefact Primary evidence for the tested configuration. The method, run date, inputs, scoring and material limitations should be available.
Public primary record Legislation, regulator guidance, official statistics, filings, technical specifications and source repositories support claims about what was published or implemented.
Vendor research or documentation Used for product behaviour, architecture and the vendor's own evaluation results. Claims remain attributed and are not described as independent findings.
Named customer case study Evidence that an outcome was reported in a real workload. It is not treated as a controlled estimate of what another organisation should expect.
Independent research or reporting Used to triangulate claims and add context. Sample, geography, date, incentives and methodology are considered before generalising.
Editorial inference A reasoned interpretation of the evidence. It is presented as analysis, with assumptions and uncertainty visible.

Sources and citations

Model benchmarks and reproducibility

A benchmark result is a measurement of a particular run, not a permanent property of a model. A publishable evaluation should identify, where applicable:

The Ed-o-meter harness, tasks and checkers are open source in Featherbench under the MIT licence. Published results identify the evidence needed to inspect or repeat the lap. The LLM Leaderboard also calls out known judge bias and non-like-for-like results.

Uncertainty and editorial judgement

Limitations belong with the result. Small samples, model nondeterminism, changing products, inaccessible systems, conflicting sources and subjective scoring are stated where they constrain a conclusion.

Language should match the evidence. “Reported”, “observed in this run”, “suggests” and “we infer” carry different meanings. Superlatives and causal claims require comparative or causal evidence; a vendor description alone is not enough.

Use of AI

AI systems support defined roles in research, engineering, design and review, and the current systems are identified on the About page. Their use does not transfer accountability away from the human editor.

Independence and conflicts

Reinvently is independently run, and Ed Yau sets its research agenda. The work is produced in a personal capacity. His employment by Kerv is disclosed on author and About information; published views are his own, not his employer's.

Any material relationship relevant to a publication — including funding, paid work, early access, free credits, supplied hardware, affiliate arrangements or a financial interest — should be disclosed in that publication. Access does not guarantee favourable coverage. Paid placement is not presented as independent research.

Corrections and updates

Challenge a finding. Report an error through a GitHub correction issue or the contact form. Please include the page, disputed claim or result, and supporting evidence.

Version history

19 July 2026: First published. Standards established for evidence, benchmarking, AI assistance, conflicts, corrections and human accountability.