Get API key

Uncensored OpenAI-compatible API for developers.

Quickstart for Vibevoice pipelines

Integrate VibeVoice, the uncensored OpenAI-compatible chat completions API, into your pipeline with this quickstart guide. Use the official OpenAI SDKs to send raw text requests and receive refusal-free outputs for prompt generation or script writing.

Prerequisites and Base URL

Before starting, ensure you have an API key. Sign up via Google or email on the Get API key page to receive your key immediately. No phone number is required. You will use the official OpenAI SDKs for your preferred language. Set your base URL to https://api.vibevoice.top/v1. The model ID is uncensored. This endpoint supports standard chat completions, function calling, and JSON mode. It does not support embeddings, image generation, or audio outputs. Keep your key secure; each account allows only one active key at a time.

Sending Your First Request

Send a POST request to the chat completions endpoint. Include your API key in the Authorization header. The request body must specify the model uncensored and your messages array. The API returns structured JSON with the model's response. You can control output length with the max_tokens parameter. If omitted, the default output limit is 2,048 tokens. The context window supports up to 64,000 tokens total.

Python SDK Integration

Use the official openai Python package to interact with the API. Initialize the client with your API key and the custom base URL. Pass the model name and your system or user messages. The SDK handles serialization and error parsing automatically. This approach is ideal for backend services or data processing scripts where you need reliable, uncensored text generation for downstream tasks.

Node.js SDK Integration

Initialize the Node.js OpenAI client with your credentials and the VibeVoice base URL. Pass the required parameters to the chat.completions.create method. This method supports asynchronous execution, making it suitable for high-throughput pipelines. You can adjust parameters like temperature and top_p to control randomness. The response object contains the generated text and token usage statistics.

Streaming Responses

Enable streaming by setting the stream parameter to true. The API returns a Server-Sent Events (SSE) stream. Each chunk contains partial token data. The final chunk includes the complete token usage statistics for the request. Streaming reduces perceived latency for long-form text generation. It is useful for real-time UI updates or progressive text assembly in voice or video pipelines.

Rate Limits, Errors, and JSON Mode

The API enforces a limit of 300 requests per minute and 8 concurrent requests per key. Requests exceeding 8 MB are rejected. A 401 error indicates an invalid or missing API key. A 402 error means your prepaid credit is exhausted; top up via crypto to continue. A 429 error signals rate limiting. You can enforce structured output using response_format: {"type": "json_object"}. The model refuses sexual content involving minors, but other lawful adult content is allowed.

cURL

curl https://api.vibevoice.top/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python

from openai import OpenAI

client = OpenAI(base_url="https://api.vibevoice.top/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.vibevoice.top/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Questions and answers

What is the cost per token?

Input tokens cost $0.25 per 1M tokens. Output tokens cost $1.00 per 1M tokens. Errors and refusals do not consume credit. Prepaid credit never expires.

How do I top up my credit?

Tops are accepted via crypto only: USDT (TRC20) or USDC (Base). Minimum top-up is $10, maximum is $500. You receive a 5% bonus for tops of $50 and a 10% bonus for tops of $100.

Does the model refuse all adult content?

No. The model is uncensored for lawful adult, fictional, and controversial topics. The only hard limit is sexual content involving minors, which is always refused.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key