Reinvently is an independent AI research organisation. We evaluate models and engineering tools, publish original benchmark data, and analyse the governance and adoption questions shaping how AI is used in the UK.
The work is designed to make difficult technology decisions more legible. Findings are tied to a test run, an attributable source or a clearly labelled judgement. Where evidence is incomplete, the limitation belongs in the result rather than in a footnote no one reads.
Research areas
Model evaluation and AI engineering
Reproducible model comparisons, benchmark methodology, developer tooling and the engineering controls needed to put AI systems into production.
Explore engineering research →
UK AI governance
Regulation, public infrastructure and policy interpreted as an operating environment for technology leaders—not a stream of announcements.
Explore governance research →
AI adoption and value
Evidence about where AI creates measurable value, where programmes stall, and which organisational conditions change the outcome.
Explore adoption research →How we work
- Test before concluding. Tool and model comparisons begin with direct use, controlled runs or reproducible artefacts wherever access allows.
- Show the workings. The Ed-o-meter LLM Leaderboard, its Featherbench harness and the underlying methodology are published so results can be inspected rather than merely trusted.
- Trace the evidence. Statistics and policy claims link to their source; citations are checked against the material they are used to support.
- State uncertainty. Small samples, access constraints, judgement calls and conflicting evidence are part of the finding and are reported as such.
The full research and editorial standards cover the evidence hierarchy, benchmark reproducibility, AI assistance, conflicts, corrections and material updates.
Research and editorial team
Ed Yau sets the research agenda and is accountable for every published finding. AI systems support defined research, engineering, design and review roles. They are credited here to make the production process transparent; editorial responsibility remains human.
Ed Yau
Research Lead & Editor-in-Chief · Applied AI Architect, Kerv
Sets the research agenda, defines the methodology and signs off every published finding. The work is produced in a personal capacity; the views are his own, not his employer's.
Sonnet Anthropic
Designer · Claude Sonnet 5
Designs how the research is presented, keeping dense evidence legible across formats and screen sizes.
Fable Anthropic
Researcher · Claude Fable 5
Maps the evidence base, gathers source material and statistics, and checks each citation against the material it is used to support.
Sonnet Anthropic
Benchmark Engineer · Claude Sonnet 5
Builds and runs the evaluation harness, processes the results and produces the charts. Published benchmark numbers remain tied to reproducible runs.
Fable Anthropic
Research Editor · Claude Fable 5
Challenges claims, tightens methods and prose, applies British English and removes hype that the evidence cannot support.
Codex OpenAI
Code & PR Reviewer · OpenAI Codex
Maintains the publication code, reviews implementation against the brief, and validates changes before research updates are published.
Follow the research
New benchmarks, analysis and engineering guides are available by email — subscribe for updates. Prefer a feed? Atom and RSS are available too.