# LLM Publisher Guide (/docs/publisher-guide/llm)



Publish an AI model as a Void App. One model per app, billed per token. Your server just needs to speak OpenAI-compatible `POST /v1/chat/completions`.

## Quick start [#quick-start]

Use any OpenAI-compatible server — vLLM, Ollama, or similar:

```bash
# Ollama
ollama pull smollm2:135m
curl http://localhost:11434/v1/chat/completions -H "Content-Type: application/json" \
  -d '{"model":"smollm2:135m","messages":[{"role":"user","content":"hi"}],"stream":false}'

# vLLM
vllm serve HuggingFaceTB/SmolLM2-135M-Instruct --host 0.0.0.0 --port 9000 --served-model-name smollm2:135m
curl http://localhost:9000/v1/chat/completions -H "Content-Type: application/json" \
  -d '{"model":"smollm2:135m","messages":[{"role":"user","content":"hi"}]}'
```

Both return:

```json
{"choices":[{"message":{"content":"Hello"}}],"usage":{"prompt_tokens":10,"completion_tokens":5,"total_tokens":15}}
```

## Publish (Console) [#publish-console]

1. Verify a domain: **Console → Domains** → follow the verification steps.
2. Go to **Publish Apps → LLM** and fill in:
   * **App Name** — unique per publisher `a-z, 0-9, -, _` (3–50 chars)
   * **Model** — the exact model ID your server serves (for example `HuggingFaceTB/SmolLM-135M-Instruct`)
   * **Server URL** — `https://your-host/v1` (must end with `/v1`; the gateway will call `/chat/completions` on it)
3. Click **Create Draft** — you will receive a publisher API key. The gateway uses it to authenticate to your server, so your server should accept any `Authorization: Bearer <token>` header.
4. In the overview click **Publish** to make your app live in the Marketplace.

You can configure free and paid tiers with token limits and pricing during publishing.

## Your server contract [#your-server-contract]

Your server must expose a single endpoint:

**`POST {serverUrl}/v1/chat/completions`**

**Request — the gateway forwards the buyer's JSON verbatim:**

```json
{"model":"your-model-id","messages":[{"role":"user","content":"hi"}],"temperature":0.7,"max_tokens":64,"stream":false}
```

`model` and `messages` are required. Multimodal content (`image_url`, `audio`, `video`) is not supported in V1 and will be rejected before reaching your server.

**Response — you must return valid OpenAI shape:**

```json
{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 123,
  "model": "your-model-id",
  "choices": [{"index":0,"message":{"role":"assistant","content":"Hello"},"finish_reason":"stop"}],
  "usage": {"prompt_tokens":10,"completion_tokens":5,"total_tokens":15}
}
```

`total_tokens` must be present and greater than zero. Return `4xx`/`5xx` with `{"error":{"message":"..."}}` on failure.

**Auth:** Accept any `Authorization: Bearer <token>` header.

**Limits:** Response must be under 16MB.

## Billing [#billing]

Apps are metered per token (`total_tokens` from your response). Your free and paid tier limits are token limits. The platform fee is 20%.

## Test via gateway [#test-via-gateway]

After your app is published and a buyer has purchased it (free or paid):

```bash
curl -X POST "https://api.openvoidnet.com/v1-beta/llm/{publisher}/{app}" \
  -H "Authorization: Bearer vnb-sk-..." \
  -H "Content-Type: application/json" \
  -d '{"model":"your-model-id","messages":[{"role":"user","content":"hello"}]}'
```

Or with the OpenAI SDK — just change the `base_url`:

```python
from openai import OpenAI
client = OpenAI(base_url="https://api.openvoidnet.com/v1-beta/llm/<publisher>/<app>", api_key="vnb-sk-...")
response = client.chat.completions.create(model="your-model-id", messages=[{"role":"user","content":"hi"}])
print(response.choices[0].message.content, response.usage.total_tokens)
```

## Limits V1 [#limits-v1]

TEXT only, single model per app, no embeddings. `stream:true` is supported (SSE) — your stream must end with a final usage chunk. Expose `GET /v1/models` so buyers can list your model. See [Buyer Guide (LLM)](https://docs.openvoidnet.com/docs/buyer-guide/llm).
