Text-to-Speech

On Device AI can read text aloud using three on-device engines: Apple's built-in voices, Kokoro neural TTS, and PocketTTS for low-latency streaming speech. Pick the one that fits your needs in Settings → Voice.

TTS Engines

Three text-to-speech engines are available:

Which engine to pick

Apple voices are lightweight and cover the most languages. Kokoro sounds more natural and still supports multiple languages. PocketTTS has the lowest playback latency but only works with English text.

PocketTTS settings

When PocketTTS is the active engine, two extra controls appear in Settings → Voice:

PocketTTS supports English only. Non-English text will still synthesize on a best-effort basis, but the results are unpredictable. For other languages, switch to Kokoro or Apple.

Using TTS

There are several ways to use text-to-speech:

Capture text from an image

Text-to-Speech can also read a passage out of a photo or an imported image. Capture pages with the camera or import images in order, and the app crops the document and then runs on-device Apple Vision OCR. Review and edit the recognized text before it replaces or is added to your draft, and select multiple separate regions when a page has more than one block worth reading. Progress is visible and the pass can be cancelled, with actionable messages when recognition fails. Source images used for OCR are transient and are not kept.

One passage or a whole book?

Use image capture here for a single passage. For a full book of photographed pages, use the scanned-book workflow in Listen Library, which keeps each page beside its spoken text.

Queued playback and the sleep timer

This screen handles one block of text at a time. When you want a queue, Listen Library and RSS Feeds share a durable Up Next list, an app-global mini-player, resume-at-position tracking, word or sentence highlighting, and a sleep timer that stops playback after 5 to 60 minutes. Listen playback speed is a separate setting from the speech speed below.

Generate & Save

From the Text to Speech tab, you can export speech to a WAV file. Tap "Generate & Save" to write the audio to your device's recordings folder with a metadata sidecar (text, engine, voice, speed). Cancelling mid-export removes any partial files automatically.

Save Text to a Knowledge Library

If generated or pasted text should become reusable project context, import it into a Knowledge Library from TTS history or from the library's Add History flow. The text becomes an editable document in the destination library. Exported audio stays in Text-to-Speech history; only the text is copied into the library.

Speech Speed

Adjust the speech speed in Settings → Voice:

Auto-Play Responses

When enabled, AI responses in chat are automatically converted to speech and played back. This creates a hands-free conversational experience, especially useful when combined with voice input.

Toggle auto-play in Settings → Voice → "Auto-play voice responses".