API Documentation: Chat Completions for Text Pipelines
Get your uncensored LLM working in your speech to text pipeline with this quickstart. Use the OpenAI-compatible endpoint to post text, stream results, or call tools via standard SDKs.
Base URL & Authentication
Start by pointing your client to the base URL https://api.speechtotextapis.com/v1. This endpoint is fully OpenAI-compatible, meaning you can use any standard openai SDK by simply overriding the base_url parameter. You do not need to rewrite your request logic; just swap the URL and inject your API key. Your key is generated immediately upon signup on the key page. It is tied to your account, and you can regenerate it at any time, which revokes the old key instantly. There is no phone number requirement, and the trial credit activates without a card. This setup keeps your pipeline simple while ensuring you have access to an uncensored model that handles lawful adult content without softening outputs.
First Request
Send a standard POST request to /v1/chat/completions. The model identifier is always uncensored. This model is an open-weight variant tuned for minimal refusal on lawful topics. It runs on our dedicated GPU servers. You do not get GPT or Claude; you get a consistent, single-model experience. The request body includes the model, messages, and stream flags. Below is a basic cURL example to verify connectivity and get a text response.
curl https://api.speechtotextapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
This request returns a JSON object with the generated text. If the key is invalid, you receive a 401 error. If your prepaid credit is exhausted, you receive a 402 error. These are the primary blockers to check before debugging logic.
Python SDK Integration
For Python developers, the official openai library works out of the box. Initialize the client with your key and the custom base URL. Pass the uncensored model ID in your chat completion call. This approach is ideal for post-processing ASR output, cleaning up transcription noise, or extracting structured data from raw text. The uncensored nature ensures that if your transcript contains adult dialogue or controversial terms, the LLM will process it without blocking or rewriting for propriety. This is critical for creative pipelines where fidelity to the source audio is paramount. The SDK handles tokenization and retry logic automatically.
from openai import OpenAI
client = OpenAI(base_url="https://api.speechtotextapis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Remember that the model ID is strictly uncensored. Do not substitute it with gpt-4 or other generic identifiers, as those may not route to our specific uncensored weights.
Node SDK Integration
Node.js developers can use the openai npm package similarly. Set the apiKey and baseURL in the client constructor. Then call chat.completions.create with the uncensored model. This is useful for real-time captioning scripts or server-side text generation tasks. The API supports standard JSON payloads. You can include system messages to define the tone or role of the assistant. The uncensored model responds directly to these prompts without extra safety layers that might alter creative output. Ensure you handle the response stream if you are processing large documents. The Node SDK abstracts the HTTP details, letting you focus on the text pipeline logic.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.speechtotextapis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Rate limits apply per key. If you exceed 300 requests per minute, the API returns a 429 error. Implement exponential backoff in your retry logic to handle this gracefully.
Streaming Responses (SSE)
For low-latency applications, enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) chunks. Each chunk contains a partial text delta. This is essential for live captioning or interactive chat interfaces. The uncensored model supports streaming natively. You do not need to wait for the full completion to start displaying text. This reduces perceived latency significantly. Configure your SDK to handle the stream events. In Python, iterate over the stream object. In Node, use the async iterator. The stream ends with a final chunk containing the finish reason. This method is efficient for bandwidth and user experience.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Streaming does not change the token billing. You are charged for the total input and output tokens regardless of whether you stream or request a full response.
Limits & Constraints
Your account has a hard limit of 300 requests per minute per API key. The request body cannot exceed 8 MB. The context window is 100,000 tokens, covering both prompt and completion. This is sufficient for most document processing and long-context transcription tasks. If you send a prompt that exceeds the window, the API returns an error. The uncensored model blocks only one specific hard limit: sexual content involving minors. All other lawful adult content is permitted. This makes it ideal for media pipelines where content variety is expected. There are no subscriptions; you pay as you go. Prepaid credit does not expire. Use the GET /v1/models endpoint to verify availability if needed, though the model ID is always uncensored.
Under the hood: specs
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Parameter | Details |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model | uncensored |
| Base URL | https://api.speechtotextapis.com/v1 |
| Authentication | Authorization: Bearer YOUR_KEY |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Max context | 100,000 tokens, input and output combined |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| JSON mode | JSON object mode via response_format json_object |
| Rate limit | 300/min per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Parallel requests | up to 8 in parallel per key |
| Request size | up to 8 MB per request |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Free trial | $0.50 for 7 days, no card |
| Bonus credit | +5% from $50, +10% from $100 |
| How you pay | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Credit expiry | no monthly fee; paid credit does not expire |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Key management | one key per account, regenerate any time (the old one stops working) |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Account | Google or e-mail and password |
Error reference
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Is this the official OpenAI API?
No, this is an independent service. We host our own uncensored model that is compatible with the OpenAI chat completions interface. It is not GPT-4 or GPT-3.5. Check OpenAI's documentation for their specific pricing and limits, as ours are different.
What happens if I run out of credit?
Your API key remains active, but requests will return a 402 Payment Required error. You must top up your prepaid account via crypto (USDT or USDC) to resume usage. Credits never expire, so you can add funds whenever convenient.
Does the uncensored model filter adult content?
It filters sexual content involving minors, which is a hard block. For all other lawful adult content, fictional scenarios, or controversial topics, the model does not refuse or soften the output. This ensures your pipeline preserves the original tone and meaning of your source text.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.