api.demimonde.dev · version 1

One key, every dialect.

The Demimonde API speaks OpenAI and Anthropic natively on the same hostname, the same key, and the same balance. Point the tools you already use at the gateway and pick your flavor.

Endpointurl
POST https://api.demimonde.dev/v1/chat/completions

Introduction

What this is

A single inference gateway in front of magnetar-1, built on Qwen 3.6 27B and served on our infrastructure. It implements the OpenAI Chat Completions protocol and the Anthropic Messages protocol in full, including streaming event grammars, tool calling, and native error shapes. Nothing is translated in your client.

Message content is never stored. Logs carry metadata only: timestamps, token counts, latency, status. What you send and what the model returns exist in memory for the duration of the request and nowhere else.

Web chat and API calls draw from the same balance and the same usage pools. There is no separate billing wall between them.

Quickstart

Your first stream

Create a key in settings → API (it is shown once), then pick the dialect you already have an SDK for.

curl https://api.demimonde.dev/v1/chat/completions \
  -H "Authorization: Bearer dm-live-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "magnetar-1",
    "messages": [{"role": "user", "content": "…"}],
    "stream": true
  }'

Authentication

Keys

One key space serves every endpoint. Send it as a bearer token or, on Anthropic routes, as the x-api-key header. Either is accepted everywhere.

Headershttp
Authorization: Bearer dm-live-…

# or, equivalently on /v1/messages routes
x-api-key: dm-live-…

Keys are displayed once at creation and stored hashed. You can name them, set a monthly spend cap per key, and revoke them from the dashboard. Last used metadata is visible per key.

Endpoints

The surface

RouteMethodDialectAuth header
/v1/chat/completionsPOSTOpenAIBearer
/v1/messagesPOSTAnthropicx-api-key or Bearer
/v1/completionsPOSTOpenAI legacyBearer
/v1/modelsGETOpenAIBearer
/v1/models/{id}GETOpenAIBearer
/v1/meGETbotheither
/v1/usageGETbotheither
/v1/messages/count_tokensPOSTAnthropicx-api-key
/.well-known/openapi.yamlGET—public

GET /v1/me returns your tier, the binding usage window as a percentage with its reset time, and your prepaid balance. The full machine readable spec lives at /.well-known/openapi.yaml.

Streaming

Server sent events, both ways

Set stream: true (OpenAI) or use the messages streaming flow (Anthropic) and the gateway emits the native event grammar for the dialect you chose. There is no intermediate format.

Event grammarsse
# OpenAI dialect
data: {"type":"chat.completion.chunk","choices":[{"delta":{"content":"…"}}]}

# Anthropic dialect
event: message_start
event: content_block_start
event: content_block_delta   → input_text_delta
event: content_block_stop
event: message_delta         → stop_reason, usage
event: message_stop

Disconnect mid stream and you are billed only for tokens actually emitted, measured from upstream usage. The abort propagates to the GPU worker within half a second.

Models

magnetar-1

One model, served on our own infrastructure.

magnetar-1

stable alias
parameters
50B
context
262,144 tokens
modalities
text in, text out
retention
zero
origin
Qwen 3.6 27B
snapshots
magnetar-1-YYYY.MM

Dated snapshots are immutable, so you can pin a version and never be surprised by a behavior change. Snapshots are deprecated with 90 days of notice. The stable alias always points at the current production model.

Errors

Native shapes, honest codes

Errors arrive in the error shape of the dialect you called, with the correct HTTP status. There is no cross dialect error soup.

Error shapesjson
// OpenAI dialect
{"error": {"message": "…", "type": "invalid_request_error", "code": "…"}}

// Anthropic dialect
{"type": "error", "error": {"type": "invalid_request_error", "message": "…"}}
StatusCodeMeaning
400invalid_requestMalformed body, unknown parameter
400floor_refusalMatched The Floor, the legal floor (scope at /floor, patterns private)
401invalid_keyMissing, revoked, or malformed key
402insufficient_balanceBody includes the exact deficit
429rate_limit_exceededRolling window exhausted, retry after included
429stream_limitConcurrent stream cap for your tier reached
503upstream_unavailableGPU worker cold boot, retry with backoff

Cold starts return a retry_after hint and streaming requests may begin with a warming status event while the worker comes up. Subscribers ride the priority queue.

Limits

The guardrails

LimitValue
Request body cap1 MB
Context window262,144 tokens, max_tokens clamped to the remainder
IdempotencyIdempotency-Key header honored on POSTs
Concurrent streams · Pro2
Concurrent streams · Max4

Rolling usage windows are tracked server side and exposed only as a percentage through /v1/me, never as raw token counts.

Pricing and balance

One rate card

Token typePrice
Input$2.00 per million
Input, prefix cache hit$0.20 per million
Output$7.00 per million

Prepaid balance is shared across web chat and API. Deposits in crypto (BTC, XMR, USDT, ETH) earn +10% credit. Subscriptions include usage pools measured against the same rate card, and overage meters seamlessly to your balance instead of hard blocking mid stream.

Tool configs

Copy paste setups

The gateway was designed for tools like these. Same key, same balance, pick your dialect.

{
  "provider": "openai",
  "baseURL": "https://api.demimonde.dev/v1",
  "apiKey": "dm-live-…",
  "model": "magnetar-1"
}
Cline · VS Code settingstext
Provider:       OpenAI Compatible
Base URL:       https://api.demimonde.dev/v1
API Key:        dm-live-…
Model ID:       magnetar-1
LibreChat · librechat.yamlyaml
# librechat.yaml
endpoints:
  custom:
    name: "Demimonde"
    baseURL: "https://api.demimonde.dev/v1"
    apiKey: "dm-live-…"
    model: "magnetar-1"
SillyTavern · chat completion sourcetext
Chat Completion Source:  Custom (OpenAI compatible)
Custom Endpoint:         https://api.demimonde.dev/v1
API Key:                 dm-live-…
Model:                   magnetar-1

Ready to try it? Open the chat or grab a key in settings.