Mastering EpistemologyGuide · Map · Audio فا
His Master's Voice by Francis Barraud (1898–99)
The audio edition

How the audio was made

Francis Barraud, His Master's Voice, 1898–99

Every chapter of the guide is also available as a narrated audio track. Listen in the audio player, which has section markers, playback speed and resume. You can also open the MP3 files directly: each one carries chapter markers that podcast and audiobook apps can use.

To take the audiobook with you, the player page offers three ways: save the chapters in your browser for offline listening, download them all as one ZIP of MP3 files with a playlist, or subscribe in a podcast app to the feed https://epis.duckdns.org/guide/audio/feed.xml.

The chapters marked ElevenLabs are read by three synthetic ElevenLabs voices: Arthur narrates, Jane reads every third section and asks the quiz questions, and Adam Stone reads the quotations. The other chapters are narrated by the open-source Kokoro-82M text-to-speech model, and will be replaced one by one. None of it was read by a person.

Total: 14 h 37 min.


The ElevenLabs edition

The chapters are being re-recorded one at a time with ElevenLabs (the eleven_v4_turbo model), from the same narration scripts described below. Three voices share the reading:

Instead of chimes between sections, each heading is followed by a pause (a second after a section heading, a little less after smaller ones). A chime is kept at the opening, before the quiz and before the last section. Each chapter is checked by transcribing it back with a speech-recognition model and comparing the result with the script. The tool is guide/fa/audio/tools/eleven_narrate.py --edition en, shared with the Persian edition.

The sections below describe how the scripts are written, and how the Kokoro chapters were voiced.

What changes when the guide is read aloud

A page is scanned by the eye, but a recording has to be followed in order. So the audio does not read the markdown word for word. Each chapter is first turned into a narration script (scripts/), a plain-text version written for the ear:

How the reading is made less robotic

Monotone text-to-speech comes from three things: every sentence has the same shape, every pause is the same length, and one voice does everything. The narration works against all three.

  1. Several voices. The narrator (Kokoro’s af_heart voice) reads the main text. Quotations from philosophers and scientists are read in a second voice (am_michael), and the narrator then says who said it. In dialogues, each speaker has their own voice (am_puck, and the British bf_emma, which also gets British pronunciation). The argument between two friends in Chapter 1 and the social-media thread in Chapter 16 sound like people talking.
  2. Pacing and structure you can hear. A soft two-note chime and a pause mark each new section. Headings are read a little more slowly. Example arguments and other set-off passages are read slightly slower, with space around them. List items and table rows have shorter gaps than paragraphs. Sentences and clauses get real pauses. A quieter single bell introduces each “critical-thinking lesson” box, and the narrator names the box: “Here’s the critical-thinking lesson.”
  3. Accurate pronunciation. Words are converted to sounds by misaki, the grapheme-to-phoneme system Kokoro was trained with. It gets part-of-speech tags from a bundled tagger, so that “a”, “the”, “read” and “live” come out right in context. A pronunciation lexicon (tools/lexicon.txt) covers about 350 names and terms that dictionaries get wrong, including Gettier, Peirce, Nagel, Lakatos, Frege, Semmelweis, Theaetetus, a priori, tu quoque, Nyāya, pramāṇa, al-Ghazālī, Suhrawardī and Zhuangzi.
  4. Consistent sound. Each track starts and ends with silence and is normalized to −18 LUFS, a comfortable speech level that makes chapters equally loud, with a limiter keeping peaks below full scale. Tracks are encoded as 40 kbps mono MP3 and tagged with title, album and track number. Each section is a chapter marker in the MP3.

The voice is synthetic, so expect occasional oddities: a stress on the wrong syllable, a flat reading of a joke, or a rare mispronounced name. The written chapter is always the reference.

Files

Path What it is
NN-*.mp3 One track per chapter, with chapter markers
index.html, tracks.js The player page and its track list with section times
scripts/NN-*.txt The narration scripts: what is read, in which voice, with which pauses
tools/make_script.py Converts a chapter’s markdown into a narration script
tools/script_extras.py Hand-written intros and outros, table readings, and spoken rewrites of formulas
tools/lexicon.txt Pronunciations for names and foreign terms
tools/narrate.py Renders scripts to MP3: pronunciation, voices, pacing, chimes, loudness, tags

Regenerating the audio

After editing a chapter, rebuild its script and audio. Rendered sentences are cached, so only changed lines are synthesized again.

cd guide/audio/tools
python3 make_script.py              # rebuild all scripts from the chapters
python3 narrate.py 05               # re-render chapter 5 (or --all --jobs 4)
python3 narrate.py --phonemes "Gettier's paper"   # check a pronunciation

Requirements: Python 3.10 or later with onnxruntime, kokoro-onnx, misaki[en], spacy, textblob, num2words, soundfile and numpy, plus espeak-ng and ffmpeg. Put the Kokoro v1.0 ONNX model (kokoro-v1.0.onnx, or a quantized version) and the voice file (voices-v1.0.bin) in tools/kokoro/. Both are published with kokoro-onnx. These tracks were rendered with the full-precision (fp32) v1.0 model, which on a CPU is both faster and better-sounding than the 8-bit quantized one.

To fix a mispronunciation, add a line to tools/lexicon.txt and re-render the chapter. To change what is read, edit the chapter and run make_script.py, or add a rewrite to script_extras.py. Edits made by hand to a script are overwritten the next time make_script.py runs.

Credits and licenses