These standards explain how Reinvently turns direct testing, public records and third-party evidence into published findings. They apply to new model evaluations, engineering guides, research briefings and analytical work from 19 July 2026. Older publications are brought into line when they receive a material update.
Ed Yau sets the research agenda and is accountable for every publication. AI systems can support research, drafting, analysis, design and code review; they do not hold editorial responsibility.
Publication types
A publication can combine more than one type of work. The basis for a material claim should remain clear in the text.
- Original evaluation. A model, tool or system is run against a stated task, dataset or decision framework. The result is limited to the tested configuration.
- Engineering comparison. Products or architectures are compared against explicit requirements. Documentation and direct use are preferred to feature-list marketing.
- Research briefing. A development in policy, markets or technology is explained through attributable primary and secondary sources.
- Analysis. Evidence is synthesised into an interpretation or decision guide. Inference is not presented as observed fact.
Evidence hierarchy
Different evidence can answer different questions. Reinvently does not treat a vendor announcement, a customer case study and an independent test as interchangeable.
| Evidence | How it is used |
|---|---|
| Direct test artefact | Primary evidence for the tested configuration. The method, run date, inputs, scoring and material limitations should be available. |
| Public primary record | Legislation, regulator guidance, official statistics, filings, technical specifications and source repositories support claims about what was published or implemented. |
| Vendor research or documentation | Used for product behaviour, architecture and the vendor's own evaluation results. Claims remain attributed and are not described as independent findings. |
| Named customer case study | Evidence that an outcome was reported in a real workload. It is not treated as a controlled estimate of what another organisation should expect. |
| Independent research or reporting | Used to triangulate claims and add context. Sample, geography, date, incentives and methodology are considered before generalising. |
| Editorial inference | A reasoned interpretation of the evidence. It is presented as analysis, with assumptions and uncertainty visible. |
Sources and citations
- Material factual claims link to the most direct accessible source available. A link must support the claim beside it, not merely discuss the same subject.
- Time-sensitive facts are checked at publication or material update. Dates, product versions and preview or general-availability status are stated where they affect interpretation.
- Statistics are reported with the population, sample, geography and period when those details are available and material.
- Vendor and customer evidence is attributed in the sentence that uses it. Lack of independent replication is stated when it changes how the result should be read.
- If a primary source is inaccessible, deleted or insufficient, the publication says what substitute evidence was used.
- AI-generated prose, summaries and search results are not evidence. The underlying source is inspected before it is cited.
Model benchmarks and reproducibility
A benchmark result is a measurement of a particular run, not a permanent property of a model. A publishable evaluation should identify, where applicable:
- the research question, tasks, dataset and exclusions;
- the model name and version, provider or API route, run date and material settings;
- the harness version, prompts, tools, environment and timeout or retry policy;
- the scoring rubric, judge model or human review process, including any self-judging result;
- failures, refusals, missing data and deviations from the planned run;
- sample size, run-to-run variability and limits on generalising to production workloads; and
- cost and latency definitions when models are compared on either measure.
The Ed-o-meter harness, tasks and checkers are open source in Featherbench under the MIT licence. Published results identify the evidence needed to inspect or repeat the lap. The LLM Leaderboard also calls out known judge bias and non-like-for-like results.
Uncertainty and editorial judgement
Limitations belong with the result. Small samples, model nondeterminism, changing products, inaccessible systems, conflicting sources and subjective scoring are stated where they constrain a conclusion.
Language should match the evidence. “Reported”, “observed in this run”, “suggests” and “we infer” carry different meanings. Superlatives and causal claims require comparative or causal evidence; a vendor description alone is not enough.
Use of AI
AI systems support defined roles in research, engineering, design and review, and the current systems are identified on the About page. Their use does not transfer accountability away from the human editor.
- Ed Yau chooses the question, approves the method and makes the final editorial decision.
- Factual claims and citations are reviewed against their underlying sources before publication.
- When a model scores benchmark outputs, the judge and any known self-judging or positional bias are disclosed.
- AI-generated illustrations are editorial presentation, not research evidence, unless a caption explicitly identifies a data visualisation and its source.
- AI assistance is never used to invent a quotation, test run, source or first-hand experience.
Independence and conflicts
Reinvently is independently run, and Ed Yau sets its research agenda. The work is produced in a personal capacity. His employment by Kerv is disclosed on author and About information; published views are his own, not his employer's.
Any material relationship relevant to a publication — including funding, paid work, early access, free credits, supplied hardware, affiliate arrangements or a financial interest — should be disclosed in that publication. Access does not guarantee favourable coverage. Paid placement is not presented as independent research.
Corrections and updates
- Minor edit. Spelling, grammar, formatting and broken-link repairs can be made without a correction note when they do not change meaning.
- Material update. New evidence, a changed conclusion or a substantial product update changes the page's modified date. Where the interpretation changes materially, the publication should explain what changed.
- Correction. A factual, methodological or attribution error is corrected promptly and accompanied by a visible note describing the error and correction.
- Benchmark rerun. A new model or harness version produces a new attributable result. Earlier runs remain inspectable through published artefacts or the public repository history.
Challenge a finding. Report an error through a GitHub correction issue or the contact form. Please include the page, disputed claim or result, and supporting evidence.
Version history
19 July 2026: First published. Standards established for evidence, benchmarking, AI assistance, conflicts, corrections and human accountability.