Get API key

Uncensored OpenAI-compatible API for developers.

How to Integrate VibeVoice with the Suno API

The Suno API generates audio from text prompts, but it lacks a dedicated endpoint for lyric generation, requiring developers to source text externally. By integrating VibeVoice, you can programmatically generate uncensored, high-quality lyrics that feed directly into Suno’s audio pipeline, ensuring precise control over the final track’s content.

Updated

Why Pair VibeVoice with Suno?

The Suno API is powerful for audio generation, but it is primarily a text-to-audio model. While it can generate lyrics internally, users often find the output unpredictable, repetitive, or constrained by internal safety filters. For developers building automated music pipelines, relying on Suno’s internal lyric generation is like letting a black box decide your song’s narrative.

VibeVoice solves this by serving as a dedicated, uncensored text generator. It acts as the creative engine for your pipeline, producing lyrics that are structurally sound and thematically precise. Because VibeVoice is an OpenAI-compatible API, it integrates seamlessly into existing Python or Node.js workflows that already handle Suno requests.

This separation of concerns allows you to iterate on lyrics rapidly without consuming Suno’s generation credits. You can refine the text until it is perfect, then send it to Suno for audio synthesis. This is particularly useful for niche genres or adult themes where Suno’s default filters might introduce unwanted refusals or generic phrasing. VibeVoice’s uncensored nature ensures that your lyrical intent is preserved exactly as written, making it the ideal text layer for specialized audio pipelines.

Step 1: Set Up Your VibeVoice Key

Before you can generate lyrics, you need an API key from VibeVoice. The signup process is designed for developers who value speed and privacy. You can sign up using Google or an email address, with no phone number required. Once registered, your API key is displayed immediately, allowing you to start testing without delay.

VibeVoice operates on a prepaid credit model charged by real token usage. Errors and refusals are free, so you can experiment with prompt variations without worrying about wasted credits. There are no monthly subscriptions or hidden fees. Credit never expires, and you can top up using USDT (TRC20) or USDC (Base) with amounts ranging from $10 to $500. Higher top-ups even come with bonus credits, making it cost-effective for high-volume pipelines.

Ensure you store your API key securely in your environment variables. VibeVoice limits you to one active key per account, which simplifies usage tracking and billing. This straightforward setup ensures that your text generation layer is ready to feed into Suno’s audio engine with minimal configuration overhead.

Step 2: Generate Lyrics with VibeVoice

Generating lyrics with VibeVoice is straightforward thanks to its OpenAI-compatible interface. You send a POST request to the /v1/chat/completions endpoint, specifying the model as uncensored. This model is tuned to answer without content refusals, making it ideal for creative writing tasks that might trigger standard filters.

Use the system role to define the structure of your lyrics. For example, you can instruct the model to output verses, choruses, and bridges in a specific format. VibeVoice supports streaming via SSE, which is useful if you want to display progress in a UI, though for API pipelines, batch responses are often more efficient.

Be mindful of the context window, which is 64,000 tokens. This allows you to pass detailed genre instructions, character backgrounds, or previous verse context to ensure consistency. The maximum output per request is 16,000 tokens, which is more than enough for most single-song lyrics. If you need longer outputs, you can split the generation into parts, but keeping it within a single request ensures better coherence.

Step 3: Format Prompts for Suno

Once you have your lyrics from VibeVoice, you need to format them for Suno. Suno’s API expects text inputs that can be passed as either a prompt or a lyrics field, depending on the specific endpoint version you are using. The key is to ensure the text is clean and properly structured.

Strip any markdown formatting from VibeVoice’s output before sending it to Suno. Suno works best with raw text that clearly delineates sections like [Verse], [Chorus], and [Bridge]. If VibeVoice includes extra commentary, use a simple regex or string split to isolate the lyrical content.

Consider using VibeVoice’s JSON mode (response_format: {"type": "json_object"}) to get structured output. This allows you to programmatically extract verses and choruses, ensuring that your Suno API call receives exactly the text you intend. This level of control is difficult to achieve when relying on Suno’s internal generation, where the text is generated alongside the audio in a single step.

Step 4: Handle Suno API Responses

After sending your VibeVoice-generated lyrics to Suno, you will receive a job ID. Suno processes audio generation asynchronously, so you need to poll the status endpoint or use webhooks if available. The response will include the audio URL once generation is complete.

Since you are controlling the text input, the audio output will closely match your lyrical intent. However, Suno may still apply its own audio-specific filters or stylistic choices. If the audio doesn’t match the vibe, you can regenerate the text using VibeVoice with adjusted parameters (e.g., lower temperature for more deterministic output) and try again.

This two-step process adds a bit of latency compared to Suno’s native generation, but the trade-off is worth it for content-specific pipelines. You get precise control over the narrative, which is crucial for applications like automated music journalism, themed playlists, or adult-oriented audio content where Suno’s default filters might be too restrictive.

Advanced: Streaming for Real-Time Generation

For applications that require real-time feedback, VibeVoice supports Server-Sent Events (SSE). This allows you to stream lyrics as they are generated, which can be useful for UI updates or progressive audio rendering. However, for most Suno integration workflows, batch processing is more efficient.

If you do use streaming, ensure your code handles the SSE stream correctly, parsing each chunk and assembling the final text. VibeVoice includes token usage in the last chunk, which helps you monitor costs in real-time. Remember that streaming adds complexity to your code, so only use it if your use case demands it.

For example, if you are building a live music creation tool, you might stream lyrics to a user interface while Suno prepares the audio queue. This creates a smoother user experience, but it requires more robust error handling and state management. For simple batch jobs, stick to standard POST requests for reliability and simplicity.

Error Handling & Rate Limits

VibeVoice enforces a rate limit of 300 requests per minute per key, with a concurrent request limit of 8. This is generally sufficient for most development pipelines, but you should implement retry logic with exponential backoff to handle transient errors. VibeVoice’s API follows standard HTTP error codes, making it easy to integrate with common HTTP clients.

Common errors include rate limits, invalid API keys, or input token limits. Since VibeVoice is uncensored, you are less likely to encounter content-based refusals, but you may still hit token limits if your prompt is too long. Monitor your token usage closely, as costs are charged per 1M input and output tokens.

For Suno, errors are typically related to audio generation failures or quota limits. Always check the status of your Suno jobs before assuming failure. If VibeVoice generates a lyric that Suno rejects, you can quickly regenerate the text without losing your audio credits, as the text generation is a separate step.

Cost Optimization Tips

VibeVoice is priced at $0.25 per 1M input tokens and $1.00 per 1M output tokens. This is significantly cheaper than many commercial LLM APIs, making it ideal for high-volume lyric generation. Since errors are free, you can afford to experiment with different prompts until you get the perfect lyrics.

To optimize costs, keep your prompts concise but descriptive. Use the system prompt to define the style, and keep the user prompt focused on the specific song structure. Avoid repeating information that can be inferred. VibeVoice’s uncensored model is efficient, so you don’t need to send massive context windows unless necessary.

Top up with crypto for bonus credits. $50 top-ups give you +5% bonus credit, and $100 top-ups give you +10%. This bonus applies to all future usage, effectively reducing your cost per token. Since credit never expires, you can top up once and use it over time, making it a cost-effective solution for long-term projects.

Conclusion

Integrating VibeVoice with the Suno API gives you unparalleled control over the lyrical content of your generated audio. By separating text generation from audio synthesis, you can fine-tune your lyrics without the constraints of Suno’s internal filters. VibeVoice’s uncensored model, low cost, and OpenAI-compatible interface make it the ideal text layer for modern AI music pipelines.

Whether you are building a niche music service, automating content creation, or experimenting with adult-themed audio, VibeVoice provides the reliability and flexibility you need. Start by generating a few lyrics, format them for Suno, and watch your custom audio pipelines come to life. The combination of precise text control and powerful audio generation opens up new possibilities for AI-driven music creation.

Questions and answers

Does VibeVoice generate audio or just text?

VibeVoice is a text-only API. It generates lyrics and other textual content. It does not produce audio, images, or video. You need a separate service like Suno or ElevenLabs to convert the text into audio.

Can I use VibeVoice for commercial songs?

Yes, VibeVoice provides the text output which you can use commercially. You should check Suno’s specific terms regarding the ownership of generated audio, but VibeVoice’s text output is yours to use freely.

How do I top up my VibeVoice credit?

You can top up using USDT (TRC20) or USDC (Base) via crypto. There are no credit card or PayPal options. Top-ups range from $10 to $500, with bonus credits for larger amounts.

Is VibeVoice uncensored?

Yes, VibeVoice is uncensored for lawful adult use. It does not refuse controversial, fictional, or adult themes. The only hard limit is on sexual content involving minors, which is always blocked.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key