POST /v1/chat/completions and GET /v1/models endpoints. Point an OpenAI SDK, LiteLLM, or any other OpenAI client at https://api.aiql.io/v1 and the request runs as a workspace query.
The last user message is the prompt. AiQL runs the same agent chat as POST /queries. Tool use stays internal; the client only sees assistant text.
Prerequisites
- An AiQL API key in the
AIQL_API_KEYenvironment variable - A workspace ID with a finished ingestion (see Run a full ingestion)
- The official OpenAI client for your language
Authenticate
Send the API key as a Bearer token. Completions also need the workspace inAiQL-Workspace. OpenAI clients pass that as a default header.
Models
GET /v1/models lists the interaction modes. Use ask, reason, or build. Any other model id still runs the chat in ask mode, so examples that send gpt-4o work.
Create a completion
Passmessages the way you would for OpenAI. The last user message becomes the query. Earlier turns are given to the agent as conversation history.
chat.completion object. id is chatcmpl-<query_id>, so you can fetch the same turn later with GET /queries/{id}.
temperature, tools, and similar OpenAI fields are accepted and ignored. Only n=1 is supported.
Stream tokens
Setstream: true to receive OpenAI SSE chunks (chat.completion.chunk) as the agent writes text. The stream ends with data: [DONE].