> ## Documentation Index
> Fetch the complete documentation index at: https://docs.aiql.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat completions

> Query a workspace through the OpenAI Chat Completions API

AiQL exposes OpenAI-compatible `POST /v1/chat/completions` and `GET /v1/models` endpoints. Point an OpenAI SDK, LiteLLM, or any other OpenAI client at `https://api.aiql.io/v1` and the request runs as a workspace query.

The last user message is the prompt. AiQL runs the same agent chat as `POST /queries`. Tool use stays internal; the client only sees assistant text.

## Prerequisites

* An AiQL API key in the `AIQL_API_KEY` environment variable
* A workspace ID with a finished ingestion (see [Run a full ingestion](/cookbooks/ingestion))
* The official OpenAI client for your language

## Authenticate

Send the API key as a Bearer token. Completions also need the workspace in `AiQL-Workspace`. OpenAI clients pass that as a default header.

<CodeGroup>
  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.aiql.io/v1",
      api_key=os.environ["AIQL_API_KEY"],
      default_headers={"AiQL-Workspace": "YOUR_WORKSPACE_ID"},
  )
  ```

  ```typescript TypeScript theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.aiql.io/v1",
    apiKey: process.env.AIQL_API_KEY,
    defaultHeaders: { "AiQL-Workspace": "YOUR_WORKSPACE_ID" },
  });
  ```

  ```bash curl theme={null}
  curl https://api.aiql.io/v1/chat/completions \
    -H "Authorization: Bearer $AIQL_API_KEY" \
    -H "AiQL-Workspace: YOUR_WORKSPACE_ID" \
    -H "Content-Type: application/json"
  ```
</CodeGroup>

## Models

`GET /v1/models` lists the interaction modes. Use `ask`, `reason`, or `build`. Any other model id still runs the chat in `ask` mode, so examples that send `gpt-4o` work.

<CodeGroup>
  ```python Python theme={null}
  models = client.models.list()
  print([model.id for model in models.data])
  ```

  ```typescript TypeScript theme={null}
  const models = await client.models.list();
  console.log(models.data.map((model) => model.id));
  ```

  ```bash curl theme={null}
  curl https://api.aiql.io/v1/models \
    -H "Authorization: Bearer $AIQL_API_KEY"
  ```
</CodeGroup>

## Create a completion

Pass `messages` the way you would for OpenAI. The last `user` message becomes the query. Earlier turns are given to the agent as conversation history.

<CodeGroup>
  ```python Python theme={null}
  completion = client.chat.completions.create(
      model="ask",
      messages=[
          {
              "role": "user",
              "content": "list every invoice over $1,000 from last quarter",
          }
      ],
  )
  print(completion.choices[0].message.content)
  ```

  ```typescript TypeScript theme={null}
  const completion = await client.chat.completions.create({
    model: "ask",
    messages: [
      {
        role: "user",
        content: "list every invoice over $1,000 from last quarter",
      },
    ],
  });
  console.log(completion.choices[0].message.content);
  ```

  ```bash curl theme={null}
  curl -X POST https://api.aiql.io/v1/chat/completions \
    -H "Authorization: Bearer $AIQL_API_KEY" \
    -H "AiQL-Workspace: YOUR_WORKSPACE_ID" \
    -H "Content-Type: application/json" \
    -d '{"model":"ask","messages":[{"role":"user","content":"list every invoice over $1,000 from last quarter"}]}'
  ```
</CodeGroup>

The JSON response matches OpenAI's `chat.completion` object. `id` is `chatcmpl-<query_id>`, so you can fetch the same turn later with `GET /queries/{id}`.

```json theme={null}
{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1710000000,
  "model": "ask",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "There are 12 invoices over $1,000 from last quarter."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  }
}
```

JSON mode waits until the query finishes. `temperature`, `tools`, and similar OpenAI fields are accepted and ignored. Only `n=1` is supported.

## Stream tokens

Set `stream: true` to receive OpenAI SSE chunks (`chat.completion.chunk`) as the agent writes text. The stream ends with `data: [DONE]`.

<CodeGroup>
  ```python Python theme={null}
  stream = client.chat.completions.create(
      model="ask",
      messages=[
          {
              "role": "user",
              "content": "list every invoice over $1,000 from last quarter",
          }
      ],
      stream=True,
  )
  for chunk in stream:
      delta = chunk.choices[0].delta.content
      if delta:
          print(delta, end="", flush=True)
  ```

  ```typescript TypeScript theme={null}
  const stream = await client.chat.completions.create({
    model: "ask",
    messages: [
      {
        role: "user",
        content: "list every invoice over $1,000 from last quarter",
      },
    ],
    stream: true,
  });
  for await (const chunk of stream) {
    const delta = chunk.choices[0]?.delta?.content;
    if (delta) process.stdout.write(delta);
  }
  ```

  ```bash curl theme={null}
  curl -X POST https://api.aiql.io/v1/chat/completions \
    -H "Authorization: Bearer $AIQL_API_KEY" \
    -H "AiQL-Workspace: YOUR_WORKSPACE_ID" \
    -H "Content-Type: application/json" \
    -d '{"model":"ask","stream":true,"messages":[{"role":"user","content":"list every invoice over $1,000 from last quarter"}]}'
  ```
</CodeGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.