Engineering & Evaluation
Original model benchmarks, architecture decisions, tools and delivery practices.
Explore the research →UK AI Governance
Policy, regulation and the practical controls UK organisations need to track.
Follow the landscape →AI Adoption & Value
Evidence on ROI, organisational readiness and moving from pilots to useful systems.
See what works →Browse by format: research briefings or engineering guides.
Latest research
How Much Does AI Improve Software Development Productivity?
Strong evidence supports 1.1–1.3× release output; highest-automation cases report a low-confidence 4.5× median and 10×+ upper end.
Read more
Building the Ed-o-meter: Notes on Writing My Own LLM Benchmark
The harness, the tasks, the decisions that turned out to matter, and the mistakes — written down mostly because the mistakes were more informative than the results.
Read more
Locked-Down Fable Disappoints, Cheap-as-Chips GLM-5.2 Unsafe, GPT-5.5 Wins as the All-Rounder
GPT-5.5 passed 28 of 28, GLM-5.2 passed 26 at an eighth of the cost, and Fable 5 refused nine tasks — four of them entirely benign. Our data, and a deep dive into the refusals.
Read more
The UK AI Policy Landscape: What Enterprise Leaders Need to Track in 2026
There is no UK AI Act, and none is coming. What exists instead is an estate — regulators, compute money, sandboxes, institutes — that only makes sense viewed whole. The map, with the moving parts marked.
Read more
AI Growth Labs: What the UK's Regulatory Sandboxes Mean for Your Sector
Regulatory sandboxes for AI open to legal services this summer, with other sectors to follow. What direct access to your regulator buys, and what to do before applications open.
Read more
The Agent Governance Gap: $202m Budgets, 26 Percent Cost Visibility
Q2 2026 data shows agent deployment plateauing, orchestration doubling and employee resistance quadrupling — with the money committed before the controls exist.
Read more
Is Claude Fable 5 Worth It for Enterprise Coding?
Fable 5 leads every coding benchmark here by a wide margin, at double the price of Anthropic's own Opus 4.8. Whether the premium is worth it, and where GLM-5.2 does the job just as well.
Read more
How to Sandbox AI Agent Code: Firecracker, OpenSandbox, Docker, SmolVM and nono Compared
The brief said microVMs, but only two of the five are. A decision guide to the four isolation models for running AI agent code — and how to pick the right one.
Read more
Multi-Agent Orchestration Frameworks Compared: ruflo, Aperant, Sandcastle, Mission Control and Maestro
An evidence-based comparison of five multi-agent orchestration tools — ruflo, Aperant, Sandcastle, Mission Control and Maestro — by popularity, use case, hosting and community feedback.
Read more
GSD, BMAD, OpenSpec, or GitHub Spec Kit: Choosing the Right AI Development Framework
Four spec-driven AI development frameworks have emerged as the serious options for structured AI coding. They share a starting principle — and diverge sharply on everything else.
Read more
Claude Memory and Dream: Evidence, Architecture and Risk
Reported deployment gains, system architecture, and the governance risks of persistent, self-editing agents.
Read more
GPT-5.5 and the Age of Autonomous Task Completion
OpenAI's GPT-5.5 is the first mainstream model designed for AI that acts rather than assists. Here is what the governance and workflow implications mean for enterprise teams.
Read more
UK AI Regulation Tightens: What the ICO Guidance and Parliament's Inquiry Mean for Employers
The ICO, Parliament, and the DRCF are reshaping what UK employers must do when AI touches decisions that affect people. Here is what has changed and what to do about it.
Read more
The Enterprise AI Reality Check: Why 79% of Organisations Aren't Seeing ROI
AI funding hit $300 billion in Q1 2026. Meanwhile, new research shows 79% of organisations face significant adoption challenges. Here is what the successful 29% are doing differently.
Read more
UK Sovereign AI: What the Government's £500m Bet Means for Your Organisation
Liz Kendall's £500m Sovereign AI programme frames compute concentration as a national security risk. Here is what it signals for enterprise technology strategy and procurement.
Read more
A2A Protocol: An Enterprise Architecture Decision Guide
What A2A standardises, where interoperability still stops, and how to evaluate vendor support in architecture and procurement.
Read more
Global AI Adoption: How Different Countries Are Embracing Artificial Intelligence
From US velocity to EU caution, Chinese state strategy to Middle Eastern sovereign ambition — a country-by-country breakdown of AI attitudes, intensity and what it means for global business.
Read more
Claude Mythos and Project Glasswing: What the Security Evidence Shows
Reported vulnerability findings, limits of the available data, and implications for security teams.
Read more
UK AI Companies: A Field Guide to Ten Organisations
An unranked selection used to examine what the UK ecosystem produces and where it remains constrained.
Read more
RAG vs GraphRAG: Which Retrieval Architecture Is Right for Your AI Application?
Retrieval-augmented generation transformed enterprise AI. Now GraphRAG is challenging the assumptions it was built on.
Read more
Copilot Studio vs Azure AI Foundry vs AWS Bedrock
Three platforms, three layers of the AI stack. A practical guide to choosing the right enterprise AI platform for your organisation.
Read more
The UK's AI Skills Gap: Why Closing It Is One of the Biggest Opportunities of the Decade
73% of UK adults have had no AI training at all. Here's where the opportunity lies — and where to start.
Read more
GitHub Copilot vs Claude Code vs Cursor
A practical comparison of three AI coding tools — with real-world pricing models, costs at scale, and a decision guide for engineering teams in 2026.
Read more
The UK's Public AI Infrastructure: Government Institutes Shaping the National AI Ecosystem
Behind the UK's AI startups and research labs sits a layer of publicly funded institutions most people rarely encounter — from the Alan Turing Institute to the AI Security Institute.
Read more
Microsoft Foundry: Architecture, Governance and Trade-offs
A decision guide to Foundry's control plane, model and agent services, Azure advantages, operational costs and trade-offs.
Read more
Generative AI Adoption by Industry: Use Cases, Data and What Comes Next
A sector-by-sector breakdown of where generative AI is actually being used — from financial services and healthcare to charities and education.
Read more
AI-Assisted Web Development: Where It Helps and Where It Fails
Where coding and content tools save time, where they create risk, and which review controls matter.
Read more