Listen Library
The Listen Library turns long-form documents into something you can read and play back. Import a PDF, a Markdown file, a web article, or a scanned book, then follow the text on screen while an on-device voice narrates it. Open it from the Side Menu or the Menu hub on iPhone, iPad, Mac, and Apple Vision Pro.
Listen Library is a reading and listening queue. Knowledge Libraries are for retrieval, where AI Chat answers questions from passages in the documents you selected. The two workspaces stay independent, so reading progress and deletions in one do not affect the other.
Add items
Use the plus button on the Listen shelf to choose an import method:
- Import File: PDF, Markdown, plain text, HTML, or RTF.
- Readable URL: the article is extracted with its source provenance preserved.
- Paste Text: enter text directly and give it an optional document title.
- Create from Page Images: the scanned-book workflow for photographed or scanned pages.
- Share sheet: supported text, URLs, and documents shared from other apps offer to add to Listen.
- Saved RSS entries: choosing Save to Shelf in RSS Feeds copies an article into Listen Library permanently.
Imported files are copied into app-managed storage, so an item keeps working after the original file is moved or deleted. If a source has no extractable text, is encrypted without a password, or is an image-only PDF, the import is rejected with an explanation instead of creating an unplayable item. Use the scanned-book workflow for image-only pages.
When an import matches an existing item by source identity or content hash, you can open the existing item, replace it, or create an independent copy.
The shelf and the reader
The shelf lists saved documents with a format badge (PDF, Scanned Book, Text, RSS), sentence-based resume context such as "Resume at Sentence 4 of 120", audio-ready status, and a last-listened timestamp. Search matches titles, extracted text, OCR text, and metadata. You can filter by format, audio-ready state, unread or played state, and favorites, and sort by title, date added, last listened, progress, or file size.
Selecting a document opens the reader at the last valid reading position. The reader preserves the source presentation rather than replacing it with plain text:
- PDFs render as native pages with single or two-page spreads, zoom, table-of-contents navigation, in-document search, and text selection. Narration skips repetitive running headers, footers, and page numbers.
- Markdown renders headings, paragraphs, emphasis, lists, and links.
- Plain text renders with reader typography and adjustable font size, line spacing, and dark mode.
- Scanned books render the retained page images in final page order.
Transport controls cover play, pause, restart the current sentence, skip backward or forward, jump between sentences, and scrubbing. Follow Playback scrolls the text automatically and highlights the active passage. Playback speed runs from 0.5x to 2.0x and is independent of standalone Text-to-Speech speed.
Row actions and transfers
Each shelf row has an overflow menu with Document Info, Rename, Copy to Knowledge, Move to Knowledge, and Delete. Document Info shows the source, format, reading progress, retained asset size, generated-audio size, page and sentence counts, and the TTS engine in use.
Copy to Knowledge creates an independent snapshot of the readable text in a Knowledge Library and leaves the Listen item and its progress untouched. Move to Knowledge transfers the document out of Listen and removes it from the shelf. Deleting a Listen document does not delete a separate Knowledge Library copy, and deleting the Knowledge copy does not affect the Listen item.
Transfers work in both directions. Copying a supported Knowledge document into Listen creates an independent Listen item and recreates a renderable source asset when one is available. If the Knowledge document only has extracted text, Listen imports it as formatted text and says so rather than implying the original PDF layout came with it. Copying a scanned book to Knowledge carries the corrected readable text and scanned-source provenance; the page images and split-view layout stay in Listen.
Deleting a Listen item removes its library entry, extracted text, page images, managed source asset, and reader preprocessing data, and applies the generated-audio deletion choice you select.
Scanned books and OCR
Choose Create from Page Images to build a book from photographed or scanned pages. Everything runs locally with Apple Vision OCR, and the workflow is free.
- Capture or import pages with the multi-page camera flow or an ordered image import.
- Organize pages: reorder, crop, rotate, or split facing-page scans in a visual editor.
- Recognize text on every retained page, with recoverable progress if the pass is interrupted.
- Review and correct the OCR text per page. Corrected text is authoritative for playback, and the correction preview uses Apple's built-in voices.
- Save the book as a shelf item with an ordered, versioned page bundle.
Opening a saved scanned book shows a split view with the original page images beside the extracted text, and View Full Page opens a full-resolution viewer for zooming, panning, and rotating. Spoken sentences are highlighted on both sides. You can keep the page images for visual reading or delete them to reclaim disk space. Partial work is recoverable as a draft, and a save that fails never leaves a shelf item whose page assets are missing.
Text-to-Speech can also capture a passage from a photo or imported image with on-device OCR, crop the document, and let you edit the recognized text before speaking it. Use that path for a single passage and the scanned-book workflow for a whole book. Image-only PDFs are outside scanned-book import.
Up Next and progress
Up Next is one durable, ordered queue shared by Listen Library documents and RSS entries. Use Play Next to insert an item at the front of the pending queue or Add to Queue to append it. Reorder pending items by dragging, remove individual items, or clear the queue. Bulk actions are available, and the queue survives app relaunches. Removing a queued item never deletes its source document or subscription.
Playback is app-global. Leave the reader while audio is playing and the same session continues, with a mini-player available across the app for pause, resume, skip, seek, speed, previous, next, and close. Previous and Next walk the queue history without losing progress.
Each item stores its current sentence, source character range, last-read date, completion state, and document-specific reader preferences, so sentence-based resume stays exact across screen sizes, font sizes, and TTS engines. If a source is replaced or re-imported with changed content, progress is remapped to a safe matching position or reset with a visible notice.
Background and lock-screen playback requires Pro; without it, playback pauses when the app moves to the background. Queue scale and feed subscription counts also follow the free, trial, and Pro tiers. Reading, foreground playback, scanned-book OCR, and word-level highlighting are free.
Word highlighting
When the active narration engine reports word-exact timing, the reader highlights the currently spoken word in rendered text and in PDFs. When only sentence-boundary timing is available, it highlights the active sentence instead. If word mapping would be unreliable because of text normalization, PDF layout, or scanned-page geometry, the reader falls back to the nearest reliable sentence, line, or page rather than highlighting the wrong text. Highlighting advances from audible playback, not from generation progress, so the visual position matches what you hear. Word-level Listen highlighting is free.
Sleep timer
Open the sleep-timer menu in the Listen mini-player to stop playback after 5, 10, 15, 30, 45, or 60 minutes, or cancel a running timer. The timer belongs to the app: it keeps counting while playback is paused and resumes its countdown after a relaunch. The Listen widget shows the remaining countdown alongside playback progress when there is room, but widgets cannot change the timer.
Voices and narration
Listen uses one synchronized preferred narration profile with engine-aware voice and instruction controls, and voice preferences can be scoped to an individual document. Available options depend on the selected engine, its language support, and any speech assets you downloaded. Shared Listening Settings serves Listen Library, RSS Feeds, the readers, narration preferences, and generated-audio management, and you can open it from the shelf, the RSS Feeds header, or a reader. An active session may keep its original settings until you choose to apply new ones.
Narration consistency is always on, reusing the same synthesis seed so tone and pronunciation stay stable across sessions. Regenerate Narration Variation re-synthesizes with a fresh seed when you want a different prosody. Engines that cannot safely apply narration instructions read the exact source text instead, which keeps article and document wording accurate.
Generated audio and storage
Generated narration audio is cached locally for offline listening up to a default limit of 2.15 GB, and the oldest unpinned files are evicted automatically when the limit is reached. Open Manage Generated Audio in Listening Settings to inspect cached files grouped by article, feed, voice, or model, review sizes and last-played dates, and delete what you no longer need. A dedicated action removes audio left behind by previous voices or models.
See also
- RSS Feeds — follow sites and turn new articles into a listening queue.
- Text-to-Speech — engines, voices, speech speed, and image passage capture.
- Knowledge Libraries — retrieval for AI Chat, and the destination for Copy or Move to Knowledge.
- Widgets — resume playback and reach Up Next from the Home Screen.