05
Writing Dispatches from the deep.
Technical posts and essays.
2026 · 07 · 14
RL Post-Training on Macs
Multi-turn RL of an 8B mixture-of-experts model with 14 Macs in four countries generating rollouts and one B200 training it. Held-out pass@1 on PaperSearchQA went from 29% to 63%.
technical · 30 min
Pluralis blog
2026 · 01 · 05 The Act of Creation
On moving from passive absorption to active curation, and finding the self through creation.
essay · 12 min
Substack · Liminal
2025 · 08 · 14 Introducing q Evaluation Harness
The first open-source evaluation framework for LLMs on q/kdb+. Top models score 96.2% on Python's HumanEval; on the same problems in q, the best one, Grok 4, scores 43.4%.
technical · 7 min
Medium · KX Systems