# LLM Overview (/docs/llm/overview)



LLM apps serve exactly **one AI model** each through an OpenAI-compatible endpoint. Billed per token from `usage.total_tokens`.

## At a glance [#at-a-glance]

|              |                                                                                            |
| ------------ | ------------------------------------------------------------------------------------------ |
| Endpoint     | `POST /v1-beta/llm/{username}/{appname}`                                                   |
| Body         | `{"model":"<published-model-id>","messages":[{...}],"stream":false}`                       |
| Billing unit | Tokens (`total_tokens / 1M × price`)                                                       |
| Start here   | [Buyer Guide](https://docs.openvoidnet.com/docs/llm/buyer) to call · [Publisher Guide](https://docs.openvoidnet.com/docs/llm/publisher) to publish |

## Rules that bite [#rules-that-bite]

* The `model` you send must match the app's published model exactly — a read-only listing is available at `GET /v1-beta/llm/{username}/{appname}/models` (see [Buyer Guide](https://docs.openvoidnet.com/docs/llm/buyer)).
* TEXT only: `image_url`, `audio`, `video` are rejected before reaching the publisher.
* `stream:true` is supported (SSE); the gateway injects `stream_options.include_usage:true` when missing. Streams that never send a final usage chunk fail `502` instead of going unbilled.
* Publishers: serve `POST {serverUrl}/chat/completions` (`{serverUrl}` is your configured Server URL, which already ends in `/v1`), return `usage.total_tokens` greater than zero on every success.
