Working on AI in the open,
and showing our working

We're interested in one question: what does it take to make AI answer from a real model of the world rather than improvise? We test it, we build with it, and we publish what we find — benchmarks, field notes, the things that didn't work. All free to read.

The testbed

How we test these ideas in practice

We use Envoy and Envoy Warehouse as a working system for testing ideas about grounded AI. Read the technical walkthrough of the architecture, skills and data model.

Come argue with us

Disagree with a result? Working on something similar? Want a method explained, or a model added to the next eval? We like the conversation more than the conclusion — write to us.

Connect