AI engineering decisions are difficult to make from vendor benchmarks alone. This collection focuses on evidence that can be inspected: identical tasks, explicit assumptions, reproducible methods and trade-offs that remain visible rather than being compressed into a single score.
Flagship research
The Ed-o-meter LLM Leaderboard
Ten leading models run through the same 28-task Featherbench evaluation. Compare pass rate, answer quality, security behaviour, latency and cost — with the limitations shown alongside the numbers.
Explore the LLM Leaderboard →
Engineering research and guides
Building the Ed-o-meter: Notes on Writing My Own LLM Benchmark
The harness, tasks and evaluation decisions behind Featherbench — including the mistakes that proved most useful.
Read more
GLM-5.2, Fable 5 and GPT-5.5 on the Same 28 Tasks
What identical prompts reveal about reliability, refusals, quality and cost.
Read more
Is Claude Fable 5 Worth It for Enterprise Coding?
How to interpret its benchmark lead, price premium and refusal behaviour.
Read more
How to Sandbox AI Agent Code
Firecracker, OpenSandbox, Docker, SmolVM and nono compared by isolation model.
Read more
Multi-Agent Orchestration Frameworks Compared
Five orchestration tools compared by use case, hosting, maturity and community evidence.
Read more
RAG vs GraphRAG
Which retrieval architecture fits your data, questions and operational constraints.
Read more