Eve
The default endpoint voice and the quickest neutral starting point for an initial casting pass.
Fish Voice Studio / Replicate endpoint
Grok text to speech turns a written performance into downloadable audio through the xai/grok-text-to-speech endpoint on Replicate. Direct five voices with speech tags, choose from 20 languages plus automatic detection, and export in one of five formats.
015 voices
0220 languages + auto
035 output formats
Precision audio workstation
Write and direct the take on the left, then set its voice and delivery format on the right.
Endpoint / 01
Grok text to speech converts a written script into natural-sounding audio. This page operates the xai/grok-text-to-speech model exposed by Replicate, then saves successful output to Fish Voice storage for account-owned playback and download.
The callable endpoint currently gives this studio five named voices, 20 specified languages plus auto detection, speech-direction tags, five output encodings, sample-rate control, MP3 bit-rate control, and optional text normalization.
This is not an official xAI website. xAI's broader Voice product describes additional voices, languages, streaming, timestamps, and voice cloning; those broader capabilities are not promised by this Replicate-backed studio.
Cast / 02
The endpoint documents five voice IDs. Audition the same revealing sentence before a full render because names alone do not guarantee a fit for your subject, language, or mix.
The default endpoint voice and the quickest neutral starting point for an initial casting pass.
A distinct documented option to compare when Eve does not match the intended delivery.
A documented alternative worth testing on direct narration, product lines, and concise prompts.
A separate voice choice for side-by-side auditions with identical text and language settings.
The fifth documented endpoint voice for completing a controlled five-way casting test.
Workflow / 03
A good Grok TTS workflow separates writing, performance direction, technical delivery, and final review.
Start with the exact spoken words and include the name, number, or long sentence most likely to reveal a problem.
Insert restrained speech tags, choose one of the five voices, and keep the same line when comparing candidates.
Use MP3 for convenient sharing, WAV or PCM for editing, and telephony codecs only when the receiving phone system requires them.
Sign in, confirm the credit estimate, review playable output in the browser, then download the owned file for production.
Speech tags place short performance cues directly in the script. Use only the cues you can hear and verify in a test render.
[pause]The result is ready. [pause] Let us review it.
Create an intentional break between thoughts.
[laugh]That was not in the brief. [laugh]
Test a brief laugh where the line genuinely calls for it.
[sigh][sigh] We need one more take.
Signal a visible change of attitude before the line.
[breath]Hold on. [breath] Start from the top.
Add a small human beat without rewriting the sentence.
<whisper><whisper> Keep this between us.
Audition a quieter direction on a short, isolated phrase.
Choose the format from its next destination, not from a universal quality ranking.
Choose MP3 for compact previews, review links, podcast assembly, and video drafts.
Bit-rate control applies only to MP3 in this endpoint.
Choose WAV when an editor or audio workstation needs a browser-playable uncompressed container.
Keep the sample rate aligned with the production timeline.
Choose PCM when downstream software explicitly requests raw pulse-code audio.
Raw PCM is download-first and may need import settings in the target application.
Choose μ-law for a phone or IVR system whose technical specification names that codec.
Confirm the required sample rate with the receiving system.
Choose A-law for telephony infrastructure that specifically requires A-law audio.
Do not substitute it for μ-law without checking the deployment specification.
Production / 06
Explore three practical deliverables—directed narration, multilingual assets, and telephony prompts—before choosing the voice and format for your production.
CASE 01Cast the hook with one controlled line, direct pauses around the key claim, export a review file, and replace only the take that fails picture or mix.
CASE 02Keep source and translated scripts paired, select the declared target language, and review names and numbers before exporting each locale's owned asset.
CASE 03Render complete branches, inspect duration and clarity, then deliver μ-law or A-law only after the phone platform's codec and sample-rate requirements are confirmed.
Decision / 07
There is no universal winner: choose from the controls, voice sourcing, and integration model your actual workflow requires.
Choose Grok when…
you want to operate Replicate's documented five-voice endpoint in a browser, direct speech tags, choose MP3, WAV, raw, or telephony output, and keep the result in Fish Voice history.
Choose ElevenLabs when…
your project depends on its larger documented voice library, model selection, or an ElevenLabs voice-cloning workflow. Verify the current plan and model documentation before committing.
Choose OpenAI when…
your application is already built around the OpenAI API and instruction-driven speech generation is the control model you need. Confirm current model and voice availability in OpenAI's documentation.
No. Fish Voice provides a browser workflow around the xai/grok-text-to-speech endpoint available through Replicate. xAI and Replicate remain the sources for their own product and endpoint documentation.
The current endpoint options in this studio are Eve, Ara, Rex, Sal, and Leo. Audition identical copy across candidates before choosing a production voice.
The endpoint selector lists 20 language choices plus automatic detection. Pick the explicit target language when reviewing localized names, numbers, and terminology.
Use MP3 for compact sharing, WAV or PCM when an editing or software workflow requests them, and μ-law or A-law only when a phone system specifies the codec.
The result player is available for MP3 and WAV. PCM, μ-law, and A-law remain downloadable, but their playback depends on the target system's raw-audio settings.
Yes. Signing in protects paid provider usage and lets Fish Voice apply the account's credits, text limit, saved asset ownership, and generation history.
The editor keeps your script and settings. The server follows the existing failed-generation credit reconciliation path and returns an actionable, non-sensitive error category.
No. Those capabilities are outside the Replicate endpoint controls implemented here, even when broader xAI product pages describe additional voice functionality.
Evidence / 10
Official documentation defines the claims and boundaries used on this page; availability can change, so check the linked source before planning a production integration.
Ready / Studio
Start with the sentence that carries the most pronunciation or timing risk, then review the output before scaling the script.