soundWave
soundWave is an NVDA add-on that renders text to an audio file using installed speech engines.
It can export audio from the clipboard, a text file, or text entered directly in the add-on.
Quick start
- Press NVDA+Ctrl+= to open soundWave.
- Choose a synthesizer.
- Optionally set the record base folder. This is the default save folder for renders; soundWave creates output folders under it using your naming templates.
- Choose the input source: clipboard, a text file, a folder of text files, or typed/pasted text.
- Configure the synthesizer. Use Test to preview the selected settings.
- Choose an output filename and format.
- Wait for the render to complete, then review the summary.
In soundWave render dialogs, press F1 to open this manual.
In the soundWave panel inside NVDA Settings, tab to the Help button or press Alt+H.
Dialog controls include keyboard accelerators where practical, such as Alt+T for Test.
Input and output
- Input: clipboard text, a text file, a folder of text files, or typed/pasted text.
- Output: WAV is always supported. MP3, FLAC, and M4A are available when ffmpeg is installed and available on the system PATH.
- Batch input: choosing a folder renders each readable text-like file as a separate audio file in sorted order. By default, batch renders are written to a synthesizer and voice subfolder, and filenames are automatically prefixed with ordered numbers.
- Large inputs: long text is rendered in chunks. If a single WAV would become too large, soundWave writes numbered part files.
- Folders and filenames: soundWave can name single-render folders, single-render files, batch folders, and batch files separately.
- Voice names: when soundWave can identify the selected voice, it can include that voice in folder or file templates so multiple renders with the same input and synthesizer are easier to distinguish.
Render progress
During rendering, soundWave shows a progress dialog with a Cancel button and optional line-by-line details.
Use Show details to review current file, text chunk, elapsed time, rendered audio length, and conversion rate where available.
Long durations are shown as minutes or hours rather than raw seconds.
For long renders, use Minimize or press Alt+M to hide the progress dialog.
The render continues in the background and the hidden dialog is kept out of Alt+Tab.
Press the soundWave command again, NVDA+Ctrl+=, to bring the progress dialog back.
Settings
Open NVDA Settings and choose soundWave to set defaults used by future renders.
- Default save folder: the record base folder used by the first soundWave dialog. Leaving it blank uses Documents\SoundWave, so rendered files stay somewhere most users can find.
- Folder and file naming: separate templates control single-render folders, single-render files, batch folders, and batch files. The tables below show the defaults, tokens, and examples.
- Default output format: choose WAV, or MP3, FLAC, or M4A when ffmpeg is available. Optional checkboxes let you skip the single-render Save As dialog and the batch format picker.
- Pitch post-processing: optionally apply a Sonic-style pitch shift after the synthesizer has rendered the audio. Sonic pitch uses a 0 to 100 scale where 50 is unchanged. Lower values deepen the finished file and higher values raise it. The full range maps to roughly -12 to +12 semitones and requires ffmpeg with rubberband support.
- After rendering: choose whether soundWave opens the output folder once the whole render finishes, opens rendered audio in your default media player, or shows the render summary. Batch auto-play opens the rendered files as a playlist. Automatic folder opening does not reopen a folder soundWave has already opened in the current NVDA session.
Naming templates
Naming fields are templates. soundWave replaces each token with information from the current render before suggesting the folder or filename.
You can remove tokens, move them, add your own words, or change the punctuation between them.
| Setting | Default | Example output | What it controls |
| Single render folder pattern |
%engine% - %voice% |
Sonata - Kirsten Two En Medium |
The folder offered by the Save As dialog for clipboard, typed text, or one text file. |
| Single file name pattern |
%source% |
Clipboard.wav |
The suggested filename for a single render. |
| Batch folder pattern |
%engine% - %voice% |
Keynote Gold - Fred |
The output folder created when rendering a folder of text files. |
| Batch file pattern |
%number% - %source% |
03 - chapter three.wav |
The filename for each rendered file in a batch. |
Naming examples
| Pattern | Example output | When it is useful |
%engine% - %voice% |
Orpheus - Synthetic Dave |
Groups output by synthesizer and voice. |
%source% - %engine% |
chapter one - SAPI5.wav |
Keeps the source first while still identifying the synthesizer. |
%number% - %source% |
01 - opening credits.wav |
Keeps batch output in a predictable order. |
Filename token meanings
| Token | Meaning | Example value |
%source% |
The clipboard, typed text, source file name, or source folder item name. |
Clipboard, Typed, chapter one, or meeting notes |
%engine% |
The synthesizer name. |
Orpheus, Sonata, SAPI5, or Pocket TTS |
%voice% |
The selected voice when soundWave can identify it. If no voice is known, this part is left out cleanly. |
Synthetic Dave, Kirsten Two En Medium, or Fred |
%number% |
An ordered number. Folder renders add this automatically; the token is available for custom naming patterns. |
01, 02, or 03 |
Synthesizer Options
Synthesizers expose settings such as voice, language, variant, speed, rate, pitch, or volume where those options are available.
Most settings dialogs include:
- Auto-speak when changing settings to preview changes automatically.
- Test to play a short sample using the current settings.
- Persistent settings saved in NVDA configuration.
Supported Synthesizers
| Synthesizer | Notes |
| SAPI5 |
Uses installed Microsoft SAPI5 voices. 32-bit SAPI5 voices are supported through a separate 32-bit helper path when available. SAPI5 rate, pitch, and volume can be adjusted in the render dialog. soundWave asks SAPI for the selected voice's default output format instead of forcing a fixed sample rate, falling back to 22.05 kHz 16-bit mono only if SAPI does not expose a format. |
| googleTtsForNvda |
Requires the Google TTS For NVDA add-on and installed Google voice packages. soundWave captures the add-on's generated PCM audio directly when rendering. The Test button renders a short sample to a temporary WAV and plays it back, avoiding NVDA's live speech path. soundWave only lists Google voices that are installed and reliable through this offline render path, so its list may be shorter than the Google add-on's full voice catalogue. |
| Pocket TTS |
Requires the Pocket TTS add-on. soundWave renders in smaller segments when needed so longer text is less likely to hit model limits. The options dialog includes Pocket TTS's end-of-sentence sensitivity setting when the installed synth exposes it. |
| Supertonic |
Requires the Supertonic add-on. soundWave uses smaller render chunks for Supertonic to avoid model-size failures on longer input. |
| Sonata |
Requires the Sonata Neural Voices add-on and available voice configuration files. |
| Eloquence and IBMTTS |
Available when the relevant NVDA add-on is installed. soundWave detects the speech engine automatically and exposes voice and speed options before rendering. |
| Orpheus |
Requires Orpheus to be the current NVDA synthesizer before rendering starts. Language, voice, speed, pitch, and volume can be adjusted before rendering. During rendering, soundWave temporarily switches NVDA to an available fallback synthesizer. |
| Orpheus Classic |
Available when the Orpheus Classic add-on is installed. soundWave uses a dedicated capture path for this older synthesizer so longer pauses and punctuation do not produce truncated output. |
| DECtalk and Keynote Gold |
Available when the required add-on is installed and discoverable. Keynote Gold exposes voice, speed, pitch, volume, and rate boost in its render dialog. |
| Other NVDA synthesizers |
soundWave can render many other installed NVDA synthesizers when they use NVDA's normal audio output path. The generic dialog applies voice, variant, rate, pitch, and volume where the selected synthesizer exposes those settings. |
Render Summary
After rendering, soundWave reports the synthesizer, saved path, elapsed time, audio length, and realtime speed factor when available.
Updates
soundWave uses NVDA's add-on update channel for store-compatible updates.
Changes
1.2.0
- Added optional post-render pitch processing. It works on the completed WAV before final WAV/MP3/FLAC/M4A output, so existing synthesizer render paths stay unchanged. The setting is off by default, uses a 0 to 100 Sonic pitch scale, and maps the full range to roughly -12 to +12 semitones.
- Added Polish and Slovak interface localization, based on #10 by Kazimierz Parzych and DJ Graco.
- Added Polish and Slovak manuals, based on #9 by Kazimierz Parzych.
- Added an output-folder write check before rendering starts, so inaccessible or broken folders fail quickly with a clear SoundWave error instead of rendering first and failing during final conversion.
- Closes #7.
- Closes #8.
1.1.2
- Added a dedicated Orpheus Classic capture path. soundWave now renders Orpheus Classic through its normal NVDA driver flow while capturing generated audio directly, avoiding the very short output that could occur through generic NVDA capture.
- Improved Orpheus Classic rendering around larger punctuation pauses, including lines with exclamation marks and longer pauses inside Mastodon-style posts.
1.1.1
- Improved long googleTtsForNvda renders by reusing a single Google bridge instance across the whole render, using smaller Google-specific chunks, and retrying recoverable Chrome DevTools bridge failures.
- Changed the progress dialog so it can be minimized with Alt+M. Pressing the soundWave command again restores the active render progress instead of opening a second render dialog.
- Improved elapsed time and audio length reporting so long renders are shown in minutes or hours instead of only raw seconds.
- Restored generic NVDA synthesizer voice, variant, language, rate, pitch, and volume settings after capture to reduce the chance of settings leaking back into the user's running NVDA speech.
- Closes #5.
1.1.0
- Greatly improved googleTtsForNvda support. SoundWave now renders through Google's bridge more reliably, lists only installed voices that work through the offline render path, retries once if the bridge page is not ready, and refuses to silently fall back to a different voice.
- Improved googleTtsForNvda previews by rendering the test phrase to a temporary WAV before playback, allowing automatic preview to stay enabled without using NVDA's live speech path.
- Improved googleTtsForNvda render progress text while SoundWave is waiting for Google to deliver audio.
- Improved Pocket TTS rendering by skipping blank input lines and splitting text into safer segments for the model.
- Added Pocket TTS end-of-sentence sensitivity to the synth options dialog when available.
- Added Supertonic-specific chunking so longer input can render without hitting the model's input-size limit.
- Added FLAC and M4A output when ffmpeg is installed and available on the system PATH.
- Changed output format selection so single renders and batch renders remember the last selected format.
- Added opt-in settings to use the default output format without showing the single-render Save As dialog or batch format picker.
- Changed batch auto-play so completed folder renders open as a playlist rather than playing only the first rendered file.
- Changed automatic output folder opening so soundWave does not reopen the same folder repeatedly during the same NVDA session.
- Added a soundWave settings panel in NVDA Settings for the default save folder, naming templates, and post-render actions.
- Added post-render actions to open the output folder, play the completed audio file, and control whether the render summary is shown.
- Added an Open Folder button to the render completion summary.
- Added folder input for batch rendering.
- Improved output naming with separate templates for single-render folders, single-render files, batch folders, and batch files.
- Improved suggested output filenames so selected voice names are included for more render paths.
- Split specialist synthesizer support into separate modules for SAPI5, DECtalk, BestSpeech, Sonata, Pocket TTS, googleTtsForNvda, Eloquence/IBMTTS, Orpheus, and generic NVDA capture paths.
- Added automatic detection for Eloquence and IBMTTS so they render through the dedicated engine path instead of the generic NVDA capture path.
1.0.4
- Added googleTtsForNvda rendering support, based on PR #3 by Ruslan Dolovaniuk.
1.0.3
- Added SAPI5 pitch support using SAPI's native XML pitch tag.
- Changed SAPI5 WAV rendering to use the selected voice's default SAPI output format where available, rather than always forcing 22.05 kHz 16-bit mono.
1.0.2
- Added pitch and volume controls for Orpheus, Keynote Gold, and generic NVDA synth rendering.
- Added volume support for SAPI5 rendering.
- Improved keyboard adjustment for numeric options. Page Up and Page Down now make larger changes, and Home and End jump to minimum and maximum values.
- Changed render details to a line-by-line list so progress information is easier to review with a screen reader.
- Remembered whether render details are shown or hidden between renders.
- Added selected voice names to suggested output filenames when SoundWave can identify the voice.
- Added F1 manual access in soundWave render dialogs and tabbable Help buttons.
- Improved dialog keyboard accelerators, including Test buttons and synthesizer option controls.
1.0.1
- Aligned update handling with NVDA Add-on Store distribution.
1.0.0
Troubleshooting
- No audio is produced: the selected synthesizer may not expose an audio path soundWave can render. Try another synthesizer and check the NVDA log.
- Eloquence or IBMTTS is not available: install or update the relevant NVDA add-on, reload NVDA, and try again.
- Orpheus will not start: switch NVDA's current synthesizer to Orpheus, then open soundWave again.
- WorldVoice does not initialize: check the NVDA log for missing files under WorldVoice-workspace. soundWave cannot render WorldVoice until WorldVoice itself loads cleanly.
- MP3, FLAC, or M4A export fails: install ffmpeg and ensure the ffmpeg command works from a Command Prompt.
- Rendering appears stuck: use Cancel. If the issue repeats, include the synthesizer name, settings, input type, and relevant NVDA log entries in the report.