Skip to content

BEN EBSWORTH

INFRA | SOFTWARE | HARDWARE

BLOGPROJECTABOUT{ }↳ NEWTwelvefree…⌘

Latest writing

All posts →
Engineering08 AUG 2026· 11 min read

Twelve free models just walked into our benchmark — three of them beat the frontier

We wired OpenRouter's free tier into our 7-task LLM harness, registered 13 models with full metadata, and ran a fair 5-iteration sweep across all of them. Ling 3.0 Tiny, Laguna XS 2.1 and Gemma 4 26B posted averages above 98 on a board that Kimi K3 leads at 90.5 — and the entire run cost us nothing.

→
Software19 JULY 2026· 8 min read

The delta rule: linear attention for a million-token context

Full attention pays an n² bill that a 1M-token context can't afford. Linear attention swaps the bill for a memory you write to — and the delta rule is what makes that memory smart. Kimi calls K3's KDA a 'hybrid linear attention mechanism'; this is the family it belongs to, from the kernel trick to gated delta updates.

The delta rule: linear attention for a million-token context
Software19 JULY 2026· 7 min read

We pointed our own benchmark at Kimi K3 on launch week

Our 7-task harness renders (or shows) what models actually generate, live and sandboxed. Running Kimi K3 against K2.7, Gemini and Codex broke the harness three different ways before it produced a fair table — here's the data, and what K3 is actually good at.

We pointed our own benchmark at Kimi K3 on launch week

From the lab

All 30+ effects →

Small, working simulations. Drag the controls on the full pages, or just watch the previews cycle here. The selection rotates daily.