{"id":260971,"date":"2026-08-03T16:24:50","date_gmt":"2026-08-03T23:24:50","guid":{"rendered":"https:\/\/picsart.com\/blog\/?p=260971"},"modified":"2026-08-03T16:24:50","modified_gmt":"2026-08-03T23:24:50","slug":"all-about-grok-tts","status":"publish","type":"post","link":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/","title":{"rendered":"All about the Grok text to speech model and how to use it"},"content":{"rendered":"<figure><\/figure>\r\n<p>Grok text to speech turns written text into spoken audio with a single API call, and the feature that separates it from a standard voice generator is control. Inline speech tags let you place a laugh, a whisper, or a pause exactly where you want one, so delivery stops being something you hope for and becomes something you write. xAI released it on April 17, 2026 alongside Grok Speech to Text, built on the same stack that runs Grok Voice, Tesla vehicles, and Starlink customer support.<\/p>\r\n<p>Grok TTS is available in the <a href=\"https:\/\/picsart.com\/ai-playground\/\">Picsart AI Playground<\/a>, so trying it takes a prompt rather than an API key. Every technical detail below comes from xAI&#8217;s own documentation.<\/p>\r\n<h2><span id=\"Grok_text_to_speech_at_a_glance\">Grok text to speech at a glance<\/span><\/h2>\r\n\r\n<figure class=\"wp-block-table\">\r\n<table style=\"border-collapse: collapse; width: 100%; font-size: 12px; table-layout: fixed;\">\r\n<thead>\r\n<tr>\r\n<th style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Spec<\/th>\r\n<th style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Detail<\/th>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Developer<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">xAI<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Released<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">April 17, 2026<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Two ways in<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">A standard request for finished audio, or streaming for real time<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Max text<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">15,000 characters per standard request<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Languages<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">20, plus auto-detect<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Voices on Picsart<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">5<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Price via xAI<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">$15.00 per 1 million characters<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Price via Picsart<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">1 credit per 1,000 characters<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n<h2><span id=\"What_is_Grok_text_to_speech\">What is Grok text to speech?<\/span><\/h2>\r\n<p>Grok text to speech, also called Grok TTS, is a standalone audio API from xAI that converts text into spoken audio. It arrived as one half of a pair, with Grok Speech to Text handling transcription in the other direction, and it is aimed at voice agents, accessibility tools, podcasts, and interactive audio.<\/p>\r\n<p>There are two ways in. The standard one takes a block of text and hands back a finished audio file. The streaming one sends audio back while the text is still arriving, which is what live voice assistants need.<\/p>\r\n<p>Either way, the request is short. You supply the words, the language, and the voice, and audio comes back. Everything else has a sensible default.<\/p>\r\n<p>Two things are required: the words you want spoken, up to 15,000 characters, and the language they should be read in. Everything else is optional.<\/p>\r\n<ul>\r\n<li><strong>Voice.<\/strong> Which of the built-in voices reads the script, Eve by default.<\/li>\r\n<li><strong>Output format.<\/strong> The file type and audio quality, MP3 at 24 kHz by default.<\/li>\r\n<li><strong>Speed.<\/strong> Anywhere from 0.7 to 1.5, with 1.0 as normal.<\/li>\r\n<li><strong>Symbols as words.<\/strong> Numbers, abbreviations, and symbols get spoken out in full. Off by default.<\/li>\r\n<li><strong>Timestamps.<\/strong> The exact timing of every character comes back with the audio. Off by default.<\/li>\r\n<\/ul>\r\n<h2><span id=\"Speech_tags_are_the_real_feature\">Speech tags are the real feature<\/span><\/h2>\r\n<p>Most text-to-speech systems give you a voice and a speed slider. Grok TTS gives you a markup layer inside the text itself, and it comes in two forms.<\/p>\r\n<p>Write &#8220;So I walked in and [pause] there it was. [laugh] I honestly could not believe it!&#8221; and the beat and the laugh land exactly where the sentence needs them.<\/p>\r\n<p><strong>Inline tags<\/strong> sit in square brackets and fire a single expression at the exact point you drop them in. They cover pauses, laughter and crying, mouth sounds, and breathing, with tags such as pause, long-pause, and laugh.<\/p>\r\n<p><strong>Wrapping tags<\/strong> go around a stretch of text and change how that whole section is delivered, opening and closing around it. They cover volume and intensity, pitch and speed, and vocal style, with tags such as whisper, slow, and soft. Wrap one sentence in a whisper and the volume drops for that line alone, then returns to normal.<\/p>\r\n<p>xAI&#8217;s guidance is specific. Place inline tags where the expression would naturally happen in speech rather than scattering them. Combine them with punctuation, since &#8220;Really? [laugh] That&#8217;s incredible!&#8221; beats stacking tags together. Reach for a pause to let a thought land. Wrap complete phrases rather than single words. Tags can also be nested, so a slow tag and a soft tag together give one line both qualities.<\/p>\r\n<h2><span id=\"Voices_and_voice_cloning\">Voices and voice cloning<\/span><\/h2>\r\n<p>Every built-in voice has its own personality, and Picsart exposes five of them in the AI Playground. xAI keeps the full roster behind a list-voices endpoint rather than publishing it in the docs, so the lineup can grow without breaking anything already built on it.<\/p>\r\n<p>Custom voices go further. A voice can be cloned from a short reference clip through the Custom Voices API, or created for free in the xAI console, then picked from the list exactly like a built-in one. Cloning a voice you do not have permission to use is the obvious thing to avoid here, and consent from the speaker matters as much as the technical setup.<\/p>\r\n<h2><span id=\"Supported_languages\">Supported languages<\/span><\/h2>\r\n<p>Grok TTS reads 20 languages, and it can detect which one you have written in rather than making you declare it.<\/p>\r\n<p>The list covers English, Arabic in Egyptian, Saudi, and Emirati variants, Bengali, Chinese (Simplified), French, German, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese in Brazilian and European variants, Russian, Spanish in Mexican and Castilian variants, Turkish, and Vietnamese.<\/p>\r\n<p>xAI notes that the model can produce speech in languages beyond that list with varying accuracy, so treat anything outside the 20 as worth testing before it ships. In the Picsart AI Playground the language sits next to a separate accent control, so English can be set to American rather than left generic.<\/p>\r\n<h2><span id=\"Output_formats_and_audio_quality\">Output formats and audio quality<\/span><\/h2>\r\n<p>Output settings decide what kind of file comes back. A podcast, a website, and a phone system each want something different, and the defaults are tuned for the web.<\/p>\r\n<h3>Choose a file type<\/h3>\r\n<ul>\r\n<li><strong>MP3<\/strong> for general use and web playback, the widest compatibility<\/li>\r\n<li><strong>WAV<\/strong> for lossless audio heading into an edit<\/li>\r\n<li><strong>PCM<\/strong> for raw audio in real-time processing<\/li>\r\n<li><strong>mulaw<\/strong> and <strong>alaw<\/strong> for phone systems<\/li>\r\n<\/ul>\r\n<h3>Choose the audio quality<\/h3>\r\n<ul>\r\n<li><strong>Sample rate:<\/strong> 8, 16, 22.05, 24, 44.1, or 48 kHz, defaulting to 24 kHz<\/li>\r\n<li><strong>Bit rate<\/strong>, MP3 only: 32, 64, 96, 128, or 192 kbps, defaulting to 128 kbps<\/li>\r\n<\/ul>\r\n<p>Higher numbers mean better sound and bigger files. xAI&#8217;s rule of thumb pairs 8 kHz with phone systems, 24 kHz with the web, and 44.1 kHz or higher with anything heading into an edit.<\/p>\r\n<h2><span id=\"Timestamps_for_captions_and_lip_sync\">Timestamps for captions and lip sync<\/span><\/h2>\r\n<p>Switch timestamps on and Grok TTS reports the exact moment every single character is spoken. That is the difference between captions that drift and captions that land on the right word, and it is what makes karaoke highlights and lip sync possible without guesswork.<\/p>\r\n<p>The response changes shape when you do this. Instead of a plain audio file, you get a data package with the audio tucked inside it, so saving the file takes one extra step in code.<\/p>\r\n<p>One quirk is worth knowing. The timing list mirrors your text character for character, including spaces, punctuation, and the speech tags themselves. A token spoken as several words puts its whole span on the first character, so $5 read aloud as &#8220;five dollars&#8221; hangs all of that time on the $. Read the list in order rather than matching it against positions in your original text.<\/p>\r\n<h2><span id=\"Streaming_for_live_voice_agents\">Streaming for live voice agents<\/span><\/h2>\r\n<p>The streaming endpoint is for anything that talks back. Text goes in piece by piece and audio comes back the same way, so the voice starts speaking before the sentence is finished. That removes the pause that makes an assistant feel slow, and it lifts the length cap that applies to a standard request.<\/p>\r\n\r\n<figure class=\"wp-block-table\">\r\n<table style=\"border-collapse: collapse; width: 100%; font-size: 12px; table-layout: fixed;\">\r\n<thead>\r\n<tr>\r\n<th style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Limit<\/th>\r\n<th style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Standard endpoint<\/th>\r\n<th style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Streaming endpoint<\/th>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Max text length<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">15,000 characters per request<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">No total limit, 15,000 per chunk<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Request timeout<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">15 minutes<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">None, the connection stays open<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Concurrent sessions<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">Not applicable<\/td>\r\n<td style=\"word-break: normal; overflow-wrap: normal; hyphens: none; padding: 6px 8px; border: 1px solid #e5e5e5; text-align: left; vertical-align: top;\">50 per team<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/figure>\r\n\r\n<p>Two behaviors make it work for real conversations. The connection stays open between turns, so a back-and-forth exchange never reconnects, and nothing from one answer bleeds into the next.<\/p>\r\n<p>The second is interruption. A single cancel message stops the voice mid-sentence and readies it for whatever the user says instead, saving roughly 600 milliseconds every time somebody cuts in. Anyone who has talked over a voice assistant and waited for it to finish knows why that matters.<\/p>\r\n<h2><span id=\"Grok_TTS_pricing\">Grok TTS pricing<\/span><\/h2>\r\n<p>Pricing is usage-based on both routes, with no tiers to compare and no surcharge for picking a particular voice. A 1,000-word script runs somewhere near 6,000 characters, so a million characters covers roughly 160 scripts of that length.<\/p>\r\n<ul>\r\n<li><strong>Through xAI:<\/strong> $15.00 per 1 million characters, with streaming billed on total input characters.<\/li>\r\n<li><strong>Through the Picsart AI Playground:<\/strong> 1 credit per 1,000 characters, on the same balance as the image and video models.<\/li>\r\n<\/ul>\r\n<h2><span id=\"Getting_better_output\">Getting better output<\/span><\/h2>\r\n<p>A few habits raise quality more than any parameter change.<\/p>\r\n<ul>\r\n<li><strong>Punctuate naturally.<\/strong> Commas, periods, and question marks drive pacing and intonation. &#8220;Wait, really?&#8221; lands better than &#8220;Wait really&#8221;.<\/li>\r\n<li><strong>Let punctuation carry emotion.<\/strong> &#8220;That&#8217;s amazing!&#8221; reads enthusiastic while &#8220;That&#8217;s amazing.&#8221; reads flat, with no tag required.<\/li>\r\n<li><strong>Break long text into paragraphs.<\/strong> Paragraph breaks create natural pauses and hold quality steady across longer scripts.<\/li>\r\n<li><strong>Split long scripts.<\/strong> Past the character cap, stream instead, or segment by paragraph and join the audio.<\/li>\r\n<li><strong>Keep the API key server-side.<\/strong> Calling the endpoint from a browser exposes the key, so proxy requests through a backend.<\/li>\r\n<li><strong>Cache repeated audio.<\/strong> Text that gets requested more than once should be stored rather than regenerated.<\/li>\r\n<\/ul>\r\n<h2><span id=\"How_to_use_Grok_text_to_speech_on_Picsart\">How to use Grok text to speech on Picsart<\/span><\/h2>\r\n<p>Grok TTS is an API first, which means a key, a backend, and code to handle the response. The <a href=\"https:\/\/picsart.com\/ai-playground\/\">Picsart AI Playground<\/a> skips all of that. Grok TTS sits in the Audio mode there with five voices and multilingual support, carrying a Fast label, so a script becomes a voiceover without a single line of code.<\/p>\r\n<p>To generate one:<\/p>\r\n<ol>\r\n<li>Open the AI Playground and switch the mode selector to Audio.<\/li>\r\n<li>Choose Grok TTS from the model dropdown.<\/li>\r\n<li>Set the language and the accent, such as English and American.<\/li>\r\n<li>Pick a voice, such as Eve.<\/li>\r\n<li>Type or paste the script into the prompt bar and generate.<\/li>\r\n<\/ol>\r\n<p>The accent selector is worth pausing on, because it sits alongside the language rather than inside it. Picking English still leaves the choice of how that English sounds, which is the difference between a voiceover that fits an audience and one that merely speaks their language.<\/p>\r\n<p>Everything then lands in the same workspace as the editing tools, so narration goes onto a timeline without a download and a re-import. The model browser also has a Compare toggle, and the Audio tab holds eleven voice models, so hearing the same line read two ways takes one toggle rather than two accounts.<\/p>\r\n<section class=\"section_faq\" id=\"faq-faq-6a7269593db0a\">\n            <h2 class=\"faq_title\" id=\"Get_answers_to_common_questions\">Get answers to common questions<\/h2>\n    \n    <div class=\"faq_items\">\n                    <div class=\"faq_item faq_item--active\">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"true\">\n                    <span class=\"faq_question_text\">What is Grok text to speech?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"false\">\n                    <div class=\"faq_answer_content\"><p>Grok text to speech is xAI&#8217;s model for turning written text into spoken audio, released on April 17, 2026. Its distinguishing feature is speech tags, which let you script laughter, pauses, and whispers into the text itself.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">How much does Grok TTS cost?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>$15.00 per 1 million characters through xAI, or 1 credit per 1,000 characters in the Picsart AI Playground.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">How many languages does Grok TTS support?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Twenty, plus auto-detect. The model can handle additional languages with varying accuracy.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">Can Grok TTS clone a voice?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Yes. A custom voice can be cloned from a short reference clip through the Custom Voices API or created in the xAI console, then used exactly like a built-in voice. Use it only with the speaker&#8217;s permission.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">What is the character limit for Grok TTS?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>15,000 characters for a standard request. Streaming has no total limit, though each chunk sent is capped at the same figure.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n                    <div class=\"faq_item \">\n                <button type=\"button\" class=\"faq_question\" aria-expanded=\"false\">\n                    <span class=\"faq_question_text\">Can I use Grok TTS without writing code?<\/span>\n                    <svg class=\"faq_chevron\" width=\"24\" height=\"24\" viewBox=\"0 0 24 24\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n                        <path d=\"M6 9L12 15L18 9\" stroke=\"currentColor\" stroke-width=\"1.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n                    <\/svg>\n                <\/button>\n                <div class=\"faq_answer\" aria-hidden=\"true\" data-collapsed>\n                    <div class=\"faq_answer_content\"><p>Yes. Grok TTS is available in the Picsart AI Playground under the Audio tab, where a pasted script becomes a voiceover without an API key.<\/p>\n<\/div>\n                <\/div>\n                <div class=\"faq_divider\"><\/div>\n            <\/div>\n            <\/div>\n<\/section>\n\n<script type=\"application\/ld+json\">\n{\n    \"@context\": \"https:\/\/schema.org\",\n    \"@type\": \"FAQPage\",\n    \"mainEntity\": [\n        {\n            \"@type\": \"Question\",\n            \"name\": \"What is Grok text to speech?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Grok text to speech is xAI&#8217;s model for turning written text into spoken audio, released on April 17, 2026. Its distinguishing feature is speech tags, which let you script laughter, pauses, and whispers into the text itself.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How much does Grok TTS cost?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"$15.00 per 1 million characters through xAI, or 1 credit per 1,000 characters in the Picsart AI Playground.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How many languages does Grok TTS support?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Twenty, plus auto-detect. The model can handle additional languages with varying accuracy.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can Grok TTS clone a voice?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. A custom voice can be cloned from a short reference clip through the Custom Voices API or created in the xAI console, then used exactly like a built-in voice. Use it only with the speaker&#8217;s permission.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"What is the character limit for Grok TTS?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"15,000 characters for a standard request. Streaming has no total limit, though each chunk sent is capped at the same figure.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can I use Grok TTS without writing code?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. Grok TTS is available in the Picsart AI Playground under the Audio tab, where a pasted script becomes a voiceover without an API key.\"\n            }\n        }\n    ]\n}<\/script>\n\n<script>\n(function() {\n    var container = document.getElementById('faq-faq-6a7269593db0a');\n    if (!container) return;\n\n    var items = container.querySelectorAll('.faq_item');\n    items.forEach(function(item) {\n        var button = item.querySelector('.faq_question');\n        var answer = item.querySelector('.faq_answer');\n        if (!button || !answer) return;\n\n        button.addEventListener('click', function() {\n            var isActive = item.classList.contains('faq_item--active');\n\n            if (isActive) {\n                item.classList.remove('faq_item--active');\n                button.setAttribute('aria-expanded', 'false');\n                answer.setAttribute('aria-hidden', 'true');\n                answer.setAttribute('data-collapsed', '');\n            } else {\n                items.forEach(function(other) {\n                    var otherBtn = other.querySelector('.faq_question');\n                    var otherAnswer = other.querySelector('.faq_answer');\n                    other.classList.remove('faq_item--active');\n                    if (otherBtn) otherBtn.setAttribute('aria-expanded', 'false');\n                    if (otherAnswer) {\n                        otherAnswer.setAttribute('aria-hidden', 'true');\n                        otherAnswer.setAttribute('data-collapsed', '');\n                    }\n                });\n                item.classList.add('faq_item--active');\n                button.setAttribute('aria-expanded', 'true');\n                answer.removeAttribute('data-collapsed');\n                answer.setAttribute('aria-hidden', 'false');\n            }\n        });\n    });\n})();\n<\/script>\n\r\n<h2><span id=\"Turn_a_script_into_a_voiceover\">Turn a script into a voiceover<\/span><\/h2>\r\n<p>Grok text to speech goes well past basic narration, with speech tags, custom voices, and character-level timestamps that give you control over how a line actually lands. Open the <a href=\"https:\/\/picsart.com\/ai-playground\/\">Picsart AI Playground<\/a>, switch to Audio, pick Grok TTS, and hear your script read back in a voice you chose.<\/p>","protected":false},"excerpt":{"rendered":"<p>Grok text to speech turns written text into spoken audio with a single API call, and the feature that separates it from a standard voice generator is control. Inline speech tags let you place a laugh, a whisper, or a pause exactly where you want one, so delivery stops being something you hope for and &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/picsart.com\/blog\/all-about-grok-tts\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;All about the Grok text to speech model and how to use it&#8221;<\/span><\/a><\/p>\n","protected":false},"author":146,"featured_media":260950,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_yoast_wpseo_title":"Grok TTS: all about xAI's text to speech model","_yoast_wpseo_metadesc":"Grok text to speech explained: speech tags, voices, 20 languages, output formats, and pricing, plus how to run it in the Picsart AI Playground.","faq_show":true,"faq_enable_schema":true,"how_to_show":false,"how_to_show_on_single":false,"how_to_enable_schema":false,"how_to_is_upload":false,"faq_title":"Get answers to common questions","how_to_title":"","how_to_layout":"","how_to_cta_text":"","how_to_cta_url":"","how_to_image_alt":"","how_to_display_image":0,"faq_items":null,"how_to_steps":[],"prompt_box_show":false,"prompt_box_placeholder":"","prompt_box_deeplink":"","prompt_box_submit_label":"","footnotes":""},"categories":[3181,1669],"tags":[3695],"class_list":["post-260971","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai","category-inspiration","tag-audio-generation","entry"],"acf":{"footer_banner_name":"Start your design in Picsart","footer_banner_link_":"\/","footer_banner_button_text_":"Get Started","faq_show":true,"faq_title":"Get answers to common questions","faq_enable_schema":true,"faq_items":[{"question":"What is Grok text to speech?","answer":"Grok text to speech is xAI's model for turning written text into spoken audio, released on April 17, 2026. Its distinguishing feature is speech tags, which let you script laughter, pauses, and whispers into the text itself."},{"question":"How much does Grok TTS cost?","answer":"$15.00 per 1 million characters through xAI, or 1 credit per 1,000 characters in the Picsart AI Playground."},{"question":"How many languages does Grok TTS support?","answer":"Twenty, plus auto-detect. The model can handle additional languages with varying accuracy."},{"question":"Can Grok TTS clone a voice?","answer":"Yes. A custom voice can be cloned from a short reference clip through the Custom Voices API or created in the xAI console, then used exactly like a built-in voice. Use it only with the speaker's permission."},{"question":"What is the character limit for Grok TTS?","answer":"15,000 characters for a standard request. Streaming has no total limit, though each chunk sent is capped at the same figure."},{"question":"Can I use Grok TTS without writing code?","answer":"Yes. Grok TTS is available in the Picsart AI Playground under the Audio tab, where a pasted script becomes a voiceover without an API key."}],"how_to_show":false,"how_to_show_on_single":false,"how_to_title":"","how_to_layout":"default","how_to_steps":null,"how_to_enable_schema":true,"how_to_is_upload":true,"how_to_cta_text":"","how_to_cta_url":"https:\/\/picsart.com\/create\/editor","how_to_display_image":null,"how_to_image_alt":"","prompt_box_show":false,"prompt_box_placeholder":"","prompt_box_deeplink":"https:\/\/picsart.com\/create\/editor?category=miniapps&app=com.picsart.aura","prompt_box_submit_label":"Create","try_prompt_show":false,"try_prompt_title":"Try this prompt","try_prompt_text":"","try_prompt_deeplink":"","tips_show":false,"tips_title":"Tips for best results","tips_items":null,"cta_banner_show":false,"cta_banner_title":"Need more space?","cta_banner_subtitle":"Extend any image in any direction with AI.","cta_banner_button_label":"Expand image","cta_banner_button_url":"","related_tools_title":"Related tools","related_tools_items":null,"post_level":""},"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.5 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Grok TTS: all about xAI&#039;s text to speech model<\/title>\n<meta name=\"description\" content=\"Grok text to speech explained: speech tags, voices, 20 languages, output formats, and pricing, plus how to run it in the Picsart AI Playground.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/picsart.com\/blog\/all-about-grok-tts\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Grok TTS: all about xAI&#039;s text to speech model\" \/>\n<meta property=\"og:description\" content=\"Grok text to speech explained: speech tags, voices, 20 languages, output formats, and pricing, plus how to run it in the Picsart AI Playground.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/picsart.com\/blog\/all-about-grok-tts\/\" \/>\n<meta property=\"og:site_name\" content=\"Picsart Blog\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/picsart\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-03T23:24:50+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/cdnblog.picsart.com\/2026\/08\/grok-text-to-speech-cover.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"800\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Julia Tovmasyan\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@PicsArtStudio\" \/>\n<meta name=\"twitter:site\" content=\"@PicsArtStudio\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Julia Tovmasyan\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Grok TTS: all about xAI's text to speech model","description":"Grok text to speech explained: speech tags, voices, 20 languages, output formats, and pricing, plus how to run it in the Picsart AI Playground.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/","og_locale":"en_US","og_type":"article","og_title":"Grok TTS: all about xAI's text to speech model","og_description":"Grok text to speech explained: speech tags, voices, 20 languages, output formats, and pricing, plus how to run it in the Picsart AI Playground.","og_url":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/","og_site_name":"Picsart Blog","article_publisher":"https:\/\/www.facebook.com\/picsart","article_published_time":"2026-08-03T23:24:50+00:00","og_image":[{"width":1200,"height":800,"url":"https:\/\/cdnblog.picsart.com\/2026\/08\/grok-text-to-speech-cover.png","type":"image\/png"}],"author":"Julia Tovmasyan","twitter_card":"summary_large_image","twitter_creator":"@PicsArtStudio","twitter_site":"@PicsArtStudio","twitter_misc":{"Written by":"Julia Tovmasyan","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/#article","isPartOf":{"@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/"},"author":{"name":"Julia Tovmasyan","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/74b70f3125250c23596a5306775b702d"},"headline":"All about the Grok text to speech model and how to use it","datePublished":"2026-08-03T23:24:50+00:00","mainEntityOfPage":{"@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/"},"wordCount":1736,"publisher":{"@id":"https:\/\/picsart.com\/blog\/ko\/#organization"},"image":{"@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/#primaryimage"},"thumbnailUrl":"https:\/\/cdnblog.picsart.com\/2026\/08\/grok-text-to-speech-cover.png","keywords":["Audio Generation"],"articleSection":["AI","Inspirational"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/","url":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/","name":"Grok TTS: all about xAI's text to speech model","isPartOf":{"@id":"https:\/\/picsart.com\/blog\/ko\/#website"},"primaryImageOfPage":{"@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/#primaryimage"},"image":{"@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/#primaryimage"},"thumbnailUrl":"https:\/\/cdnblog.picsart.com\/2026\/08\/grok-text-to-speech-cover.png","datePublished":"2026-08-03T23:24:50+00:00","description":"Grok text to speech explained: speech tags, voices, 20 languages, output formats, and pricing, plus how to run it in the Picsart AI Playground.","breadcrumb":{"@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/picsart.com\/blog\/all-about-grok-tts\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/#primaryimage","url":"https:\/\/cdnblog.picsart.com\/2026\/08\/grok-text-to-speech-cover.png","contentUrl":"https:\/\/cdnblog.picsart.com\/2026\/08\/grok-text-to-speech-cover.png","width":1200,"height":800,"caption":"Text to speech interface in the Picsart AI Audio Generator with a voiceover script, voice list, and waveform player"},{"@type":"BreadcrumbList","@id":"https:\/\/picsart.com\/blog\/all-about-grok-tts\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/picsart.com\/blog\/"},{"@type":"ListItem","position":2,"name":"All about the Grok text to speech model and how to use it"}]},{"@type":"WebSite","@id":"https:\/\/picsart.com\/blog\/ko\/#website","url":"https:\/\/picsart.com\/blog\/ko\/","name":"Picsart Blog","description":"Keep up with the latest news in photo editing, digital photography, and art trends.","publisher":{"@id":"https:\/\/picsart.com\/blog\/ko\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/picsart.com\/blog\/ko\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/picsart.com\/blog\/ko\/#organization","name":"PicsArt Inc.","url":"https:\/\/picsart.com\/blog\/ko\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/logo\/image\/","url":"https:\/\/cdnblog.picsart.com\/2016\/02\/PicsArt-logo.png","contentUrl":"https:\/\/cdnblog.picsart.com\/2016\/02\/PicsArt-logo.png","width":195,"height":43,"caption":"PicsArt Inc."},"image":{"@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/picsart","https:\/\/x.com\/PicsArtStudio","https:\/\/www.instagram.com\/picsart","https:\/\/www.linkedin.com\/company\/picsart-photo-studio","https:\/\/www.pinterest.com\/picsart"]},{"@type":"Person","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/74b70f3125250c23596a5306775b702d","name":"Julia Tovmasyan","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/picsart.com\/blog\/ko\/#\/schema\/person\/image\/","url":"https:\/\/cdnblog.picsart.com\/2026\/03\/3285C16C-FD87-4868-A2F0-04B6A0815CE1-150x150.jpg","contentUrl":"https:\/\/cdnblog.picsart.com\/2026\/03\/3285C16C-FD87-4868-A2F0-04B6A0815CE1-150x150.jpg","caption":"Julia Tovmasyan"}}]}},"featured_image":{"url":"https:\/\/cdnblog.picsart.com\/2026\/08\/grok-text-to-speech-cover.png","dimensions":[]},"_links":{"self":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260971","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/users\/146"}],"replies":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/comments?post=260971"}],"version-history":[{"count":11,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260971\/revisions"}],"predecessor-version":[{"id":260982,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/posts\/260971\/revisions\/260982"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/media\/260950"}],"wp:attachment":[{"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/media?parent=260971"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/categories?post=260971"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/picsart.com\/blog\/wp-json\/wp\/v2\/tags?post=260971"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}