ErfanMiahi
About The longer version.
I 'm a research scientist at Pluralis, working on decentralized RL post-training, where the rollouts come from machines spread across the internet instead of one datacenter. In July we post-trained an 8B mixture-of-experts model with 14 Macs in four countries generating rollouts and one B200 training it. Held-out pass@1 on PaperSearchQA went from 29% to 63%.
Before Pluralis I was at Covenant AI, where I helped build Covenant-72B, the largest permissionless pre-training run at its release, and Grail, an open RL post-training system for LLM reasoning. There I also wrote PULSE, which cuts weight-sync traffic in distributed RL by over 100×. It was accepted at NeurIPS 2026, and Fireworks cites it in their write-up of the infrastructure behind Cursor's Composer 2 RL run.
Before that I was the founding ML research engineer at DeepR Analytics, a Toronto proprietary-trading firm, where I built its RL trading pipeline. I did my M.Sc. at the University of Alberta's RLAI Lab (Richard Sutton's lab), supervised by Martha White and Marlos C. Machado, and worked with Google DeepMind on how neural network representations behave in RL.
Outside work I skydive (60 jumps so far), fly in wind tunnels, do parkour and run long distances. I keep going back to Nietzsche's philosophy and Jung's analytical psychology.
- 268
- Citations, Oct 2026
- 11
- Papers
- 100×
- Less weight-sync traffic
- 29→63%
- Pass@1 with 14 Macs
- 72B
- Parameters in Covenant-72B
- 10+
- Mentees since 2017
Research Publications & preprints.
NeurIPS · EMNLP · EACL · AIJ · MLJ · TMLR · IEEE T-Cyb · 268 citations on Google Scholar, October 2026.
Selected
-
Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
Joel Lidin, Amir Sarfi, Erfan Miahi, Quentin Anthony, Shivam Chauhan, Evangelos Pappas, Benjamin Thérien, Eugene Belilovsky, Samuel Dare
The largest permissionless, globally distributed pre-training run at its release (72B parameters, ~1.1T tokens).
arXivpreprint -
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
Erfan Miahi, Eugene Belilovsky
About 99% of per-step weight updates are invisible after the BF16 cast. PULSE uses that to cut weight-sync traffic by over 100× with bit-identical weights. Cited by Fireworks.
NeurIPSaccepted -
Investigating the Properties of Neural Network Representations in Reinforcement Learning
Han Wang, Erfan Miahi, Martha White, Marlos C. Machado, Zaheer Abbas, Raksha Kumaraswamy
Joint work with Google DeepMind, published in the Artificial Intelligence Journal.
AIJjournal
All publications
-
How Reliable are Confidence Estimators for Large Reasoning Models?
Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Ivan Brugere, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi
European Chapter of the ACL
EACLmain -
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Charese H. Smiley, Ivan Brugere, Kundan Thind, Mohammad M. Ghassemi
Confidence estimation for vision-language models
arXivpreprint -
Calibrating LLM Confidence by Probing Perturbed Representation Stability
Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi
Empirical Methods in Natural Language Processing
EMNLPmain -
GVFs in the Real World: Making Predictions Online for Water Treatment
Muhammad K. Janjua, Haseeb Shah, Martha White, Erfan Miahi, Marlos C. Machado, Adam White
Machine Learning Journal
MLJjournal -
ResMax: An Alternative Soft-Greedy Operator for Reinforcement Learning
Erfan Miahi, Revan MacQueen, Alex Ayoub, Abbas Masoumzadeh, Martha White
Transactions on Machine Learning Research
TMLRjournal -
Genetic Neural Architecture Search for Automatic Assessment of Human Sperm Images
Erfan Miahi, S.A. Mirroshandel, A. Nasr
First author · NAS for medical imaging
Expert Syst. Appl.journal -
Scalable Transfer Evolutionary Optimization: Coping with Big Task Instances
Mojtaba Shakeri, Erfan Miahi, Abhishek Gupta, Yew-Soon Ong
NTU × A*STAR collaboration
IEEE T-Cybjournal -
Effect of Deep Transfer and Multi-task Learning on Sperm Abnormality Detection
A. Abbasi, Erfan Miahi, S.A. Mirroshandel
Most-cited paper · Computers in Biology and Medicine
Comput. Biol. Med.journal
Work Projects in the open.
stoa
RL post-training with Apple-silicon Macs as the rollout fleet. The Macs generate rollouts in int8 with MLX, a GPU trains, and weight deltas sync through Cloudflare R2. Our July run used 14 Macs in four countries.
Covenant-72B
Pre-trained a 72B LLM on ~1.1T tokens with permissionless peers over the open internet, the largest run of its kind at its release.
Grail
Geo-distributed, incentivized RL post-training for LLM reasoning, where untrusted nodes generate rollouts that are verified before training. Covenant AI's open research stack.
q Evaluation Harness
The first open-source evaluation framework for LLMs on q/kdb+, built with KX. Top models score 96.2% on Python's HumanEval; on the same problems in q, the best one, Grok 4, scores 43.4%.
Draw Your Circuit
A Meta Quest 3 app that turns a hand-drawn circuit sketch into a schematic.
Deep-RL-CS285-Pytorch
PyTorch solutions to Berkeley's CS285 deep RL assignments, my most-starred repo (144 stars).
Writing Dispatches from the deep.
Technical posts and essays.
RL Post-Training on Macs
Multi-turn RL of an 8B mixture-of-experts model with 14 Macs in four countries generating rollouts and one B200 training it. Held-out pass@1 on PaperSearchQA went from 29% to 63%.
The Act of Creation
On moving from passive absorption to active curation, and finding the self through creation.
Introducing q Evaluation Harness
The first open-source evaluation framework for LLMs on q/kdb+. Top models score 96.2% on Python's HumanEval; on the same problems in q, the best one, Grok 4, scores 43.4%.
Reading Books at depth.
Books I keep going back to. The rest of the shelf is on Goodreads.
The secret for harvesting from existence the greatest fruitfulness and the greatest enjoyment is: to live dangerously.
Mentorship Mentoring since 2017.
I've mentored 10+ students in monthly or bi-weekly calls, mostly on AI research and careers. One went from a regional Iranian university to a PhD at Michigan State.
Email me with 'Mentorship Program' in the subject line. The address is in Contact below.
- 01
You want to do AI research or engineering and are looking for an honest sparring partner.
- 02
You're willing to show up consistently, even when the work stops being exciting.
- 03
You'd rather be told the truth than something comfortable.
Off-desk When I surface.
Skydiving, wind tunnels, parkour and other adventures.