OpenAI-compatible API for uncensored code generation

Get API key

First request in five minutes

Get your first response in under five minutes. This guide covers the essential steps to authenticate, configure parameters, and execute your initial API call using our uncensored DeepSeek-compatible endpoint.

Authentication and Base URL

Our API is fully OpenAI-compatible. You interact with the service by pointing your client library to https://api.deepseekcoder.cc/v1 and providing a valid API key. Unlike the official deepseek v4 api or services like the mistral api, we do not require complex OAuth flows. Simply sign up to receive your key, which you will pass in the Authorization header as a Bearer token.

Authentication is straightforward: include the header Authorization: Bearer YOUR_API_KEY in every request. If the key is missing or invalid, the API returns a 401 error. If your prepaid credit is exhausted, you receive a 402 error. You can generate a new key at any time via the dashboard.

Your First Request

Start by sending a simple chat completion request. This confirms your key works and demonstrates the model's behavior. The endpoint accepts standard OpenAI JSON payloads.

The response includes the generated text, token usage metrics, and the model identifier. Because we focus exclusively on code generation and text output, you will not see embeddings or image generation fields in the response object. This simplicity ensures your integration remains lightweight and predictable.

Begin with a basic prompt to verify connectivity before moving to more complex parameter configurations.

  • Method: POST
  • Endpoint: /v1/chat/completions
  • Content-Type: application/json
curl https://api.deepseekcoder.cc/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

For Python developers, the official OpenAI SDK works seamlessly with our endpoint. You only need to override the base URL and supply your API key. This approach allows you to use the same code structure you would with other providers, making migration or parallel testing trivial.

Define the base URL explicitly to ensure requests route to our servers. The model identifier uncensored triggers our specific open-weight model. This setup ensures you get consistent results without vendor lock-in. You can adjust parameters like temperature and top_p directly in the client constructor or per-request to fine-tune the output style.

This method is ideal for scripts, backend services, or automated code generation pipelines where reliability and speed are paramount.

from openai import OpenAI

client = OpenAI(base_url="https://api.deepseekcoder.cc/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js SDK Usage

JavaScript and TypeScript developers can use the OpenAI Node SDK to interact with our API. Similar to Python, you must configure the base URL to point to our service. This ensures that all requests, including streaming and function calling, route correctly.

The Node SDK handles pagination and response parsing efficiently. You can define the model as uncensored and set your API key in the environment or configuration object. This setup is particularly useful for serverless functions or edge computing environments where low latency is critical.

Ensure you handle potential network errors gracefully, especially when dealing with large payloads or high-concurrency scenarios. The SDK abstracts much of the complexity, allowing you to focus on logic.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.deepseekcoder.cc/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

For real-time applications, enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream, delivering tokens as they are generated. This reduces perceived latency significantly, providing immediate feedback to users.

Each chunk in the stream contains partial data. The final chunk includes the complete token usage statistics, allowing you to track costs accurately. Streaming is supported for all standard chat completions. It is particularly useful for code generation, where developers prefer to see output incrementally rather than waiting for the full response.

Handle stream termination and errors appropriately to ensure a smooth user experience. The stream ends naturally when the model completes the response or when a limit is reached.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Limits, Errors, and Context

Understand the operational limits to design robust integrations. You are limited to 300 requests per minute per key and can have up to 8 concurrent requests. The maximum request body size is 8 MB. The context window supports 64,000 tokens total, with a maximum output of 16,000 tokens per request.

Common errors include 401 (invalid key), 402 (insufficient credit), and 429 (rate limit exceeded). If you exceed the rate limit, wait and retry. Errors do not consume credit, ensuring you only pay for successful completions. The context window includes both the prompt and the completion. Plan your token usage carefully to avoid truncation or unexpected behavior.

Questions and answers

How do I handle API key rotation?

You can generate a new API key at any time from your dashboard. When you create a new key, it becomes active immediately. You can revoke or replace the old key as needed. Each account is limited to one active key at a time, ensuring secure and straightforward access management.

What happens if I exceed the rate limit?

If you exceed 300 requests per minute or 8 concurrent requests, the API returns a 429 status code. Your requests are not charged credit during this period. You should implement exponential backoff in your client to handle retries efficiently without overwhelming the service.

Is the context window 64k per request or total?

The 64,000-token limit applies to the total context window, which includes both the prompt tokens and the completion tokens generated in a single request. The maximum output for any single request is 16,000 tokens. Ensure your inputs fit within these constraints to avoid truncation.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.