AI engineering decisions are difficult to make from vendor benchmarks alone. This collection focuses on evidence that can be inspected: identical tasks, explicit assumptions, reproducible methods and trade-offs that remain visible rather than being compressed into a single score.

Flagship research

The Ed-o-meter LLM Leaderboard

Ten leading models run through the same 28-task Featherbench evaluation. Compare pass rate, answer quality, security behaviour, latency and cost — with the limitations shown alongside the numbers.

Explore the LLM Leaderboard →
Abstract illustration of three wireframe forms representing models under evaluation

Engineering research and guides

Building the Ed-o-meter: Notes on Writing My Own LLM Benchmark

The harness, tasks and evaluation decisions behind Featherbench — including the mistakes that proved most useful.

Read more

GLM-5.2, Fable 5 and GPT-5.5 on the Same 28 Tasks

What identical prompts reveal about reliability, refusals, quality and cost.

Read more

Is Claude Fable 5 Worth It for Enterprise Coding?

How to interpret its benchmark lead, price premium and refusal behaviour.

Read more

How to Sandbox AI Agent Code

Firecracker, OpenSandbox, Docker, SmolVM and nono compared by isolation model.

Read more

Multi-Agent Orchestration Frameworks Compared

Five orchestration tools compared by use case, hosting, maturity and community evidence.

Read more

RAG vs GraphRAG

Which retrieval architecture fits your data, questions and operational constraints.

Read more