Models
Every model this platform serves, read live from the gateway — not a list someone wrote down and forgot to update.
GET /v1/models is the source of truth — call it before you hardcode an id anywhere, because the table below is a snapshot of that same call and this page can go stale between deploys the way any documentation can.
GET
/v1/modelsResponse
{
"object": "list",
"data": [
{"id": "data-auto", "object": "model", "owned_by": "openai"},
{"id": "data-general", "object": "model", "owned_by": "openai"},
{"id": "data-law-ir", "object": "model", "owned_by": "openai"},
{"id": "data-law-es", "object": "model", "owned_by": "openai"},
{"id": "data-stt", "object": "model", "owned_by": "openai"},
{"id": "data-tts", "object": "model", "owned_by": "openai"},
{"id": "data-embed", "object": "model", "owned_by": "openai"}
]
}owned_by: "openai" is a compatibility artifact of the OpenAI wire format, not a claim about who trained the weights — every id above runs on this platform's own hardware. See Introduction.Deprecated aliases
The previous model ids — every
larsa-* and lardad-* name, for example data-general or data-law-ir — still route to the same models for now, but they are deprecated aliases and no longer listed by GET /v1/models. They will be removed after a notice period, so write new code against the data-* ids above and migrate existing code when convenient.Reference
| Model | For | Context | Modalities | Tools |
|---|---|---|---|---|
data-auto | One id for everything — classifies your message from its text and content and dispatches to the right specialist below. See Chat. | 262,144 | text, image → text | Yes |
data-general | The default chat and vision model — llama.cpp serving Qwen3.6-35B-A3B. Reads images and documents. See Chat and Vision. | 262,144 | text, image → text | Yes |
data-law-ir | Iranian law, retrieval-grounded over 117,499 provisions before the model answers. | 262,144 | text → text | No |
data-law-es | Spanish and EU law (BOE consolidated), retrieval-grounded over 835,951 provisions. | 262,144 | text → text | No |
data-stt | Speech to text — Whisper large-v3, silence-based chunking for long audio. See Speech to text and Realtime transcription. | — | audio → text | — |
data-tts | Text to speech — Aava for Persian, Chatterbox for 23 other languages, with a round-trip check against data-stt available on request. See Text to speech. | — | text → audio | — |
data-embed | Embeddings — BAAI's bge-m3, 1024 dimensions, multilingual. The same server the retrieval corpora above are indexed with. See Embeddings. | 8,192 | text → vector | — |
Two ceilings that apply to all of them
- Context, 262,144 tokens. Both chat backends were started with the same window, and the retrieval models generate on the same weights after their search step — so it is one number across the board, not a per-model detail to look up.
- Output, 8,192 tokens per request, platform-wide. A generation with no ceiling can run to the context limit and hold a GPU slot for everyone else —
max_tokensmay ask for less than 8,192, never more. Requesting more is clamped, not refused.
The public price list at
/pricing also carries a handful of rows — gpt-4o, unknown, the wildcard * — used internally for cost estimation. They are not callable model ids; GET /v1/models is the list that matters.