Flux API: Prompts and scripts over one endpoint
Integrate our uncensored LLM into your media pipeline with a single OpenAI-compatible endpoint. Write scripts, extract metadata, or post-process text without content refusals.
uncensoredhttps://api.fluxapikey.com/v1
Base URL and Authentication
Our API follows the standard OpenAI chat-completions structure. To begin, you need a base URL and an API key. The base URL for all requests is https://api.fluxapikey.com/v1. You can obtain an API key by signing up on the Get API key page with just an email and password. No credit card is required for the trial.
Pass your key in the Authorization header as a Bearer token. Each account is limited to one key at a time; regenerating it invalidates the previous one. Prompts are not used for training, and the only hard content limit is no sexual content involving minors.
Sending Your First Request
Make a POST request to /v1/chat/completions with your prompt. The model ID is always uncensored. This endpoint accepts standard JSON payloads and returns text responses. It supports streaming via Server-Sent Events (SSE) and tool/function calling.
Here is a basic example using curl:
curl https://api.fluxapikey.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'This request sends your text to our GPU servers. The model is tuned to answer without refusals for lawful adult, fictional, or controversial topics. You can adjust the max_tokens and temperature parameters as needed.
Python SDK Integration
Use the official OpenAI Python SDK to interact with our API. Set the base URL to our endpoint and provide your API key. This allows you to use familiar methods like client.chat.completions.create.
from openai import OpenAI
client = OpenAI(base_url="https://api.fluxapikey.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)This setup ensures compatibility with existing codebases. The response object contains the generated text in response.choices[0].message.content. You can parse this text for scripts, captions, or metadata extraction in your video or audio workflows.
Node.js SDK Usage
For JavaScript environments, use the OpenAI Node.js SDK. Configure the baseURL and apiKey properties to point to our service. This enables seamless integration into Node.js applications for real-time text generation.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.fluxapikey.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);The response structure matches the OpenAI format. You can access the generated content and handle errors using standard try-catch blocks. This approach is ideal for server-side processing in web applications or CLI tools.
Streaming Responses via SSE
Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) with incremental chunks of text. This is useful for real-time applications where you want to display text as it is generated.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Each chunk contains a partial response. You can accumulate these chunks to reconstruct the full text or display them in real-time. Streaming reduces the perceived latency for users waiting for long scripts or descriptions.
Limits, Errors, and Context Window
Our API has specific limits to ensure reliability. You can send up to 300 requests per minute per key. The maximum request body size is 8 MB. The context window supports 100,000 tokens for both prompt and completion combined.
Common errors include 401 for invalid keys, 402 for insufficient credit, and 429 for rate limits. If you exceed the rate limit, wait and retry. Our pricing is transparent: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Prepaid credit never expires, and you can top up from $10.
Capabilities and limits
Use this table to decide whether the API fits your project before you buy credit.
| Spec | Value |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Model ID | uncensored |
| Authentication | Bearer token in the Authorization header |
| Base URL | https://api.fluxapikey.com/v1 |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| JSON mode | JSON object mode via response_format json_object |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| Context window | 100,000 tokens, input and output combined |
| Max body | 8 MB request body |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Concurrency | 8 requests at the same time per key |
| Rate limit | 300/min per key |
| Subscription | paid credit never expires, no subscription |
| Volume bonus | +5% from $50, +10% from $100 |
| How you pay | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Trial credit | $0.50 for 7 days, no card |
| Key management | one active key per account; a new key replaces the old one |
| Account | Google or e-mail and password |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
HTTP errors
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What is the context window size?
The context window supports 100,000 tokens for both prompt and completion combined. This allows for long conversations or large document processing within a single request.
How do I handle errors?
Common errors include 401 for invalid keys, 402 for insufficient credit, and 429 for rate limits. You can check the response status code and message to determine the issue and take appropriate action.
Can I use this API for image or video generation?
No, this API is for text-only chat completions. We provide the uncensored brain for prompt engineering, script writing, and post-processing. For image or video generation, you would need to use a dedicated model or service.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.