Skip to the content.

Level 2 Intermediate

Speech from text on your own machine

Local Text to Speech reads text you paste into the terminal and writes a single MP3. The engine is Kokoro, run locally through KPipeline. The voice is af_bella, language code a (American English), speed 1. Each chunk of audio comes back as a NumPy array. The script concatenates those arrays, writes temp_full.wav at 24,000 Hz, and exports output_full.mp3 at 192 kbps with pydub.

It is for someone who wants spoken audio from their own text on a machine they administer. A draft, a study note, or a long passage goes in through standard input. The file that does the work is the script in this repository. Kokoro’s weights load in the same process that reads the text.

SamWiki places this project at Level 2, intermediate. The control flow is linear: read until end of input, generate, join, export, play, and delete the temporary WAV. A finished MP3 depends on a local Python environment with Kokoro and its weights, NumPy, soundfile, and pydub, plus ffmpeg for the MP3 encode. Playback opens Windows Media Player with start /min wmplayer. The file the script is built to leave behind is output_full.mp3.

Empty input stops the script at the prompt. If the generator returns no chunks, it exits before writing audio. After a successful export, temp_full.wav is removed and the MP3 remains.

The script

import sys
import os
import soundfile as sf
import numpy as np
from kokoro import KPipeline
from pydub import AudioSegment

# --- TEXT INPUT LOGIC ---
print("Enter/Paste your text. Press Ctrl+D (Unix) or Ctrl+Z (Windows) then Enter to save:")
# This reads everything including newlines, paragraphs, and punctuation until EOF
text = sys.stdin.read().strip()

if not text:
    print("Error: No text provided.")
    sys.exit()
# --------------------------------

print("--- Kokoro TTS ---")
print("Status: Loading Model...")
pipeline = KPipeline(lang_code='a') 

print("Status: Processing Text...")
generator = pipeline(text, voice='af_bella', speed=1)

# Create an empty list to hold all the generated audio chunks
audio_chunks = []

# Loop through the generator and collect the audio chunks
for i, (gs, ps, audio) in enumerate(generator):
    print(f"Status: Generating Part {i+1}...")
    audio_chunks.append(audio)

if not audio_chunks:
    print("Error: No audio was generated.")
    sys.exit()

print("Status: Concatenating audio...")
# Join all the numpy arrays into one single continuous array
full_audio = np.concatenate(audio_chunks)

# Define file names
temp_wav = 'temp_full.wav'
final_mp3 = 'output_full.mp3'

print("Status: Writing temporary WAV...")
# Save the combined audio. Note: 24000 is the standard sample rate for Kokoro
sf.write(temp_wav, full_audio, 24000)

print(f"Status: Exporting to {final_mp3}...")
audio_segment = AudioSegment.from_wav(temp_wav)
audio_segment.export(final_mp3, format="mp3", bitrate="192k")

print("Status: Playing...")
# Play the final combined file
os.system(f'start /min wmplayer "{os.path.abspath(final_mp3)}"')

# Clean up the temporary WAV file
if os.path.exists(temp_wav):
    os.remove(temp_wav)

print("--- Finished! ---")

Contribute

Corrections and small improvements belong in a pull request on this repository.

View on GitHub Contribute / Open a PR