N 43°39′ · W 79°23′ · Toronto

ErfanMiahi

Ex-founder, now working on LLM post-training at Pluralis. I've collaborated with Google DeepMind and Rich Sutton's RLAI Lab. Outside work I read philosophy and do extreme sports.

[email protected]CVGitHub

now
Pluralis
before
Covenant-72B
latest paper
PULSE · NeurIPS 2026
m.sc.
RLAI · UAlberta
02

About The longer version.

Illustrated portrait of Erfan Miahi, Research Scientist · Post-training & RL

I 'm a research scientist at Pluralis, working on decentralized RL post-training, where the rollouts come from machines spread across the internet instead of one datacenter. In July we post-trained an 8B mixture-of-experts model with 14 Macs in four countries generating rollouts and one B200 training it. Held-out pass@1 on PaperSearchQA went from 29% to 63%.

Before Pluralis I was at Covenant AI, where I helped build Covenant-72B, the largest permissionless pre-training run at its release, and Grail, an open RL post-training system for LLM reasoning. There I also wrote PULSE, which cuts weight-sync traffic in distributed RL by over 100×. It was accepted at NeurIPS 2026, and Fireworks cites it in their write-up of the infrastructure behind Cursor's Composer 2 RL run.

Before that I was the founding ML research engineer at DeepR Analytics, a Toronto proprietary-trading firm, where I built its RL trading pipeline. I did my M.Sc. at the University of Alberta's RLAI Lab (Richard Sutton's lab), supervised by Martha White and Marlos C. Machado, and worked with Google DeepMind on how neural network representations behave in RL.

Outside work I skydive (60 jumps so far), fly in wind tunnels, do parkour and run long distances. I keep going back to Nietzsche's philosophy and Jung's analytical psychology.

268
Citations, Oct 2026
11
Papers
100×
Less weight-sync traffic
29→63%
Pass@1 with 14 Macs
72B
Parameters in Covenant-72B
10+
Mentees since 2017
03

Research Publications & preprints.

NeurIPS · EMNLP · EACL · AIJ · MLJ · TMLR · IEEE T-Cyb · 268 citations on Google Scholar, October 2026.

Selected

  1. Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet

    Joel Lidin, Amir Sarfi, Erfan Miahi, Quentin Anthony, Shivam Chauhan, Evangelos Pappas, Benjamin Thérien, Eugene Belilovsky, Samuel Dare

    The largest permissionless, globally distributed pre-training run at its release (72B parameters, ~1.1T tokens).

    arXiv
    preprint
  2. Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL

    Erfan Miahi, Eugene Belilovsky

    About 99% of per-step weight updates are invisible after the BF16 cast. PULSE uses that to cut weight-sync traffic by over 100× with bit-identical weights. Cited by Fireworks.

    NeurIPS
    accepted
  3. Investigating the Properties of Neural Network Representations in Reinforcement Learning

    Han Wang, Erfan Miahi, Martha White, Marlos C. Machado, Zaheer Abbas, Raksha Kumaraswamy

    Joint work with Google DeepMind, published in the Artificial Intelligence Journal.

    AIJ
    journal

All publications

  1. How Reliable are Confidence Estimators for Large Reasoning Models?

    Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Ivan Brugere, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi

    European Chapter of the ACL

    EACL
    main
  2. Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking

    Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Charese H. Smiley, Ivan Brugere, Kundan Thind, Mohammad M. Ghassemi

    Confidence estimation for vision-language models

    arXiv
    preprint
  3. Calibrating LLM Confidence by Probing Perturbed Representation Stability

    Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi

    Empirical Methods in Natural Language Processing

    EMNLP
    main
  4. GVFs in the Real World: Making Predictions Online for Water Treatment

    Muhammad K. Janjua, Haseeb Shah, Martha White, Erfan Miahi, Marlos C. Machado, Adam White

    Machine Learning Journal

    MLJ
    journal
  5. ResMax: An Alternative Soft-Greedy Operator for Reinforcement Learning

    Erfan Miahi, Revan MacQueen, Alex Ayoub, Abbas Masoumzadeh, Martha White

    Transactions on Machine Learning Research

    TMLR
    journal
  6. Genetic Neural Architecture Search for Automatic Assessment of Human Sperm Images

    Erfan Miahi, S.A. Mirroshandel, A. Nasr

    First author · NAS for medical imaging

    Expert Syst. Appl.
    journal
  7. Scalable Transfer Evolutionary Optimization: Coping with Big Task Instances

    Mojtaba Shakeri, Erfan Miahi, Abhishek Gupta, Yew-Soon Ong

    NTU × A*STAR collaboration

    IEEE T-Cyb
    journal
  8. Effect of Deep Transfer and Multi-task Learning on Sperm Abnormality Detection

    A. Abbasi, Erfan Miahi, S.A. Mirroshandel

    Most-cited paper · Computers in Biology and Medicine

    Comput. Biol. Med.
    journal
06

Reading Books at depth.

Books I keep going back to. The rest of the shelf is on Goodreads.

Thus Spoke Zarathustra
Friedrich Nietzsche
On Intelligence
Jeff Hawkins
Five Dialogues
Plato
A Little History of the World
E.H. Gombrich
The Denial of Death
Ernest Becker
Memories, Dreams, Reflections
C.G. Jung
Good to Great
Jim Collins
The Art of Being
Erich Fromm
The Last Lecture
Randy Pausch
Art History
Dana Arnold
How to Win Friends…
Dale Carnegie
See full shelf on Goodreads ↗
The secret for harvesting from existence the greatest fruitfulness and the greatest enjoyment is: to live dangerously.
Friedrich Nietzsche · Die fröhliche Wissenschaft · §283
07

Mentorship Mentoring since 2017.

I've mentored 10+ students in monthly or bi-weekly calls, mostly on AI research and careers. One went from a regional Iranian university to a PhD at Michigan State.

How to apply

Email me with 'Mentorship Program' in the subject line. The address is in Contact below.

You might be a fit if
  1. 01

    You want to do AI research or engineering and are looking for an honest sparring partner.

  2. 02

    You're willing to show up consistently, even when the work stops being exciting.

  3. 03

    You'd rather be told the truth than something comfortable.