Evaluating models on languages they weren't built for
What our benchmarks look like when the test set isn't English — and why the gap matters more than the average score.
Read →What our research team is reading, testing, and arguing about — and what it means for building here.
What our benchmarks look like when the test set isn't English — and why the gap matters more than the average score.
Read →Latency budgets, fallbacks, and the evaluation harness nobody puts in the demo video.
Read →Notes from conversations with public institutions about procurement, sovereignty, and maintenance.
Read →How we decide an idea deserves its own cap table — and what it leaves the studio with.
Read →Where a smaller model on cheaper hardware beats the frontier — and how often that's the right trade.
Read →The argument, written out properly, for the place as well as the plan.
Read →