Working on AI in the open,
and showing our working
We're interested in one question: what does it take to make AI answer from a real model of the world rather than improvise? We test it, we build with it, and we publish what we find — benchmarks, field notes, the things that didn't work. All free to read.
AI Curriculum: Intuition & The Ontology
Structured 1-Day, 2-Day, and 3-Day courses across education and enterprise automation. 100% free and open-access.
Intuition Track
For students, tutors, teachers, and academic operations. Moving from passive AI answering machines to Socratic intellectual partners and high-agency skills.
The Ontology Track
For founders, ops managers, and engineers. Automating messy document workflows, CLI pipelines, token unit economics, and enterprise RAG.
Featured field notes
Practical guides to the choices behind the evals: where an AI system can be changed, why the strongest model does not automatically make the strongest product, and what happens when a panel of models makes falsifiable predictions.
What an AI exam-marking system taught us about local evals, designed workflows and giving production volume to workhorse models.
A practical guide to the layers you can change: skills, context and tools above the model; generation settings, weights and training below it.
Six frontier models predict HSC exam questions from official NESA evidence, backtested blind against the real 2025 papers — and pre-registered for public scoring in November.
How we test these ideas in practice
We use Envoy and Envoy Warehouse as a working system for testing ideas about grounded AI. Read the technical walkthrough of the architecture, skills and data model.
Come argue with us
Disagree with a result? Working on something similar? Want a method explained, or a model added to the next eval? We like the conversation more than the conclusion — write to us.
Connect