Read-along audio timing connects a recorded passage to the words a reader sees. The difficult part is keeping that connection accurate after text edits, regenerated sentences, and layout changes. Start with paragraph-level alignment unless the product specifically needs smaller units. A reliable paragraph highlight is more useful than word highlights that repeatedly point to the wrong phrase.
Freeze the text before marking boundaries
Choose the exact text edition that will accompany the recording. Give every paragraph a stable ID and keep headings, captions, and footnotes distinguishable from the main passage. The display copy and the spoken copy may need different punctuation, but any meaningful difference should be documented.
For example, a displayed abbreviation might need its expanded phrase in speech. Store both versions under the same paragraph ID. Do not quietly replace the public spelling with phonetic text used to improve pronunciation.
Create a small manifest containing paragraph ID, display text revision, audio filename, and timing status. When a paragraph changes, mark its timing as needing review. Treat a regenerated file as a new recording even if the words are unchanged, because its duration and internal rhythm may differ.
For a learning activity, coordinate the passage with its listening practice task before recording. Question changes can affect which parts deserve replay controls.
Mark timings against the final recording
Listen from the start and mark the beginning and end of each spoken paragraph. Include enough context to know whether a short pause belongs to the preceding sentence or introduces the next section. Choose one convention and apply it consistently.
An audio editor can help keep a visible set of boundaries. The Audacity label-track documentation describes labels for points and regions. Use your editor's current controls to record review markers; the labels themselves do not implement highlighting on a website.
Keep timing data outside the prose so editors can update text without accidentally deleting timestamps. A simple row might say P030, 00:24.8 start, 00:39.2 end, approved. Avoid pretending that these example values are universal targets.
If you need sentence-level highlighting, subdivide only after paragraph boundaries work. Each additional cue is another connection that can drift when the recording is revised.
Check navigation as well as continuous playback
Test a reader who starts at the beginning, pauses halfway through, and resumes. Then test jumping directly to a later paragraph. The highlighted text should reflect what is currently audible, including after seeking backward.
Check the gap between paragraphs. If the previous highlight remains visible during a long pause, that may be acceptable; if the next paragraph highlights before its audio begins, it can look like a skipped phrase. Decide the intended behavior and document it for the developer.
Try the layout at narrow and wide widths. Visual wrapping can change without changing paragraph IDs. Timing should attach to content units, not to the third displayed line, because line breaks depend on font size and screen width.
Keep readable text available independently of highlighting. The text-alternative workflow helps distinguish the transcript from the optional synchronized presentation.
Revalidate changed sections and their neighbors
When audio changes, remeasure the affected paragraph and all later offsets if the recording is one continuous file. If paragraphs are separate files, confirm their playback order and any transition gaps instead. Neither storage arrangement removes the need to listen to the join.
Record what was checked: text revision, audio revision, cue revision, and destination page. This prevents a correct cue sheet from being attached to an older audio export. For several languages, maintain a version matrix instead of copying the original-language timestamps.
Generate a short passage in the voice studio, download the approved audio, and build a three-paragraph alignment test in your playback tool. The success criterion is simple: a reader can pause, replay, and jump between sections while the displayed text continues to describe the words they hear.
