Fish Voice Studio / Replicate endpoint

Grok Text to Speech: Create Expressive AI Voice

Grok text to speech turns a written performance into downloadable audio through the xai/grok-text-to-speech endpoint on Replicate. Direct five voices with speech tags, choose from 20 languages plus automatic detection, and export in one of five formats.

Precision audio workstation

Grok TTS Studio

Write and direct the take on the left, then set its voice and delivery format on the right.

Script and performance direction
0 / 120 characters

Free accounts can submit up to 120 characters per generation.

Estimated cost: 0 credits

Insert performance direction

Voice and delivery
Advanced audio settings+
Checking account and credits…

Endpoint / 01

What is Grok text to speech?

Grok text to speech converts a written script into natural-sounding audio. This page operates the xai/grok-text-to-speech model exposed by Replicate, then saves successful output to Fish Voice storage for account-owned playback and download.

The callable endpoint currently gives this studio five named voices, 20 specified languages plus auto detection, speech-direction tags, five output encodings, sample-rate control, MP3 bit-rate control, and optional text normalization.

This is not an official xAI website. xAI's broader Voice product describes additional voices, languages, streaming, timestamps, and voice cloning; those broader capabilities are not promised by this Replicate-backed studio.

Cast / 02

Choose the right Grok TTS voice

The endpoint documents five voice IDs. Audition the same revealing sentence before a full render because names alone do not guarantee a fit for your subject, language, or mix.

V-01

Eve

The default endpoint voice and the quickest neutral starting point for an initial casting pass.

V-02

Ara

A distinct documented option to compare when Eve does not match the intended delivery.

V-03

Rex

A documented alternative worth testing on direct narration, product lines, and concise prompts.

V-04

Sal

A separate voice choice for side-by-side auditions with identical text and language settings.

V-05

Leo

The fifth documented endpoint voice for completing a controlled five-way casting test.

Workflow / 03

How to use Grok text to speech

A good Grok TTS workflow separates writing, performance direction, technical delivery, and final review.

  1. 01

    Write one reviewable script

    Start with the exact spoken words and include the name, number, or long sentence most likely to reveal a problem.

  2. 02

    Direct and cast the take

    Insert restrained speech tags, choose one of the five voices, and keep the same line when comparing candidates.

  3. 03

    Choose the destination format

    Use MP3 for convenient sharing, WAV or PCM for editing, and telephony codecs only when the receiving phone system requires them.

  4. 04

    Generate, listen, and download

    Sign in, confirm the credit estimate, review playable output in the browser, then download the owned file for production.

Direction / 04

Speech tags for expressive Grok TTS

Speech tags place short performance cues directly in the script. Use only the cues you can hear and verify in a test render.

[pause]

The result is ready. [pause] Let us review it.

Create an intentional break between thoughts.

[laugh]

That was not in the brief. [laugh]

Test a brief laugh where the line genuinely calls for it.

[sigh]

[sigh] We need one more take.

Signal a visible change of attitude before the line.

[breath]

Hold on. [breath] Start from the top.

Add a small human beat without rewriting the sentence.

<whisper>

<whisper> Keep this between us.

Audition a quieter direction on a short, isolated phrase.

Delivery / 05

Audio formats for creators, apps, and phone systems

Choose the format from its next destination, not from a universal quality ranking.

MP3

Choose MP3 for compact previews, review links, podcast assembly, and video drafts.

Bit-rate control applies only to MP3 in this endpoint.

WAV

Choose WAV when an editor or audio workstation needs a browser-playable uncompressed container.

Keep the sample rate aligned with the production timeline.

PCM

Choose PCM when downstream software explicitly requests raw pulse-code audio.

Raw PCM is download-first and may need import settings in the target application.

μ-law

Choose μ-law for a phone or IVR system whose technical specification names that codec.

Confirm the required sample rate with the receiving system.

A-law

Choose A-law for telephony infrastructure that specifically requires A-law audio.

Do not substitute it for μ-law without checking the deployment specification.

Production / 06

Where a directed Grok TTS take fits

Explore three practical deliverables—directed narration, multilingual assets, and telephony prompts—before choosing the voice and format for your production.

Grok TTS podcast narration being reviewed beside a video timeline, waveform, directed script, and approved audio takeCASE 01

Podcast and video narration

Cast the hook with one controlled line, direct pauses around the key claim, export a review file, and replace only the take that fails picture or mix.

Grok TTS multilingual localization delivery with paired scripts, language versions, waveforms, and final audio exportsCASE 02

Multilingual localization

Keep source and translated scripts paired, select the declared target language, and review names and numbers before exporting each locale's owned asset.

Grok TTS telephony package showing a completed IVR call flow, prompt waveforms, and verified codec delivery filesCASE 03

IVR and support prompts

Render complete branches, inspect duration and clarity, then deliver μ-law or A-law only after the phone platform's codec and sample-rate requirements are confirmed.

Decision / 07

Grok TTS vs ElevenLabs vs OpenAI TTS

There is no universal winner: choose from the controls, voice sourcing, and integration model your actual workflow requires.

01

Grok TTS here

Choose Grok when…

you want to operate Replicate's documented five-voice endpoint in a browser, direct speech tags, choose MP3, WAV, raw, or telephony output, and keep the result in Fish Voice history.

02

ElevenLabs

Choose ElevenLabs when…

your project depends on its larger documented voice library, model selection, or an ElevenLabs voice-cloning workflow. Verify the current plan and model documentation before committing.

03

OpenAI TTS

Choose OpenAI when…

your application is already built around the OpenAI API and instruction-driven speech generation is the control model you need. Confirm current model and voice availability in OpenAI's documentation.

Boundaries / 08

Endpoint limits and evidence boundaries

  • Generation requires a signed-in Fish Voice account and credits; free accounts are limited to 120 characters per request and active paid accounts to 1,000.
  • This studio exposes the five voices, 20 languages plus auto, formats, sample rates, MP3 bit rates, normalization switch, and speech-tag workflow supported by its Replicate integration.
  • It does not claim xAI's broader voice catalog, streaming, timestamps, pronunciation overrides, or voice cloning, and it does not publish unverifiable quality, speed, or price rankings.
  • Raw PCM, μ-law, and A-law files are offered as downloads because normal browser audio players do not reliably interpret every deployment-specific raw format.
Answers / 09

Grok text to speech FAQ

Is this an official xAI Grok TTS site?+

No. Fish Voice provides a browser workflow around the xai/grok-text-to-speech endpoint available through Replicate. xAI and Replicate remain the sources for their own product and endpoint documentation.

Which Grok TTS voices can I use here?+

The current endpoint options in this studio are Eve, Ara, Rex, Sal, and Leo. Audition identical copy across candidates before choosing a production voice.

How many languages does this Grok text to speech studio support?+

The endpoint selector lists 20 language choices plus automatic detection. Pick the explicit target language when reviewing localized names, numbers, and terminology.

Which audio format should I choose?+

Use MP3 for compact sharing, WAV or PCM when an editing or software workflow requests them, and μ-law or A-law only when a phone system specifies the codec.

Can I play every output in my browser?+

The result player is available for MP3 and WAV. PCM, μ-law, and A-law remain downloadable, but their playback depends on the target system's raw-audio settings.

Does generation require an account?+

Yes. Signing in protects paid provider usage and lets Fish Voice apply the account's credits, text limit, saved asset ownership, and generation history.

What happens if generation fails?+

The editor keeps your script and settings. The server follows the existing failed-generation credit reconciliation path and returns an actionable, non-sensitive error category.

Does this page include xAI voice cloning or streaming?+

No. Those capabilities are outside the Replicate endpoint controls implemented here, even when broader xAI product pages describe additional voice functionality.

Evidence / 10

Capability sources

Official documentation defines the claims and boundaries used on this page; availability can change, so check the linked source before planning a production integration.

Ready / Studio

Direct a Grok TTS take from script to owned audio

Start with the sentence that carries the most pronunciation or timing risk, then review the output before scaling the script.