OpenAI-compatible API

Make your first request in minutes.

Keep your existing OpenAI client. Change the base URL, use a Neviri inference key, and choose a configured model.

Quickstart

  1. 1

    Create an inference key

    Sign in, open API keys, and create an nvai_… key. The full secret is shown once, so store it securely.

  2. 2

    Set your environment variable

    Keep the key outside source control.

    terminal
    export NEVIRI_AI_KEY="nvai_your_api_key_here"
  3. 3

    Send a chat completion

    Use the base URL https://api.ai.neviri.com/v1 with a model available in your environment.

Make your first request

Use raw HTTP or the official OpenAI SDK. Install the SDK first with pip install openai or npm install openai.

curl
curl https://api.ai.neviri.com/v1/chat/completions \
  -H "Authorization: Bearer $NEVIRI_AI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "neviri/flagship-chat",
    "messages": [
      { "role": "user", "content": "Explain edge computing in one sentence." }
    ]
  }'

Expected response

Read generated text from choices[0].message.content. Token counts are returned in usage.

json
{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "model": "neviri/flagship-chat",
  "choices": [{
    "index": 0,
    "message": { "role": "assistant", "content": "Edge computing..." },
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 14,
    "completion_tokens": 18,
    "total_tokens": 32
  }
}

Discover available models

Model availability is deployment-specific. Query the catalogue before relying on a slug; the endpoint returns configured model ids, context windows, modalities, and reference pricing metadata.

terminal
curl https://api.ai.neviri.com/v1/models
Browse the visual catalogue

Stream a response

Set stream: true to receive incremental OpenAI-compatible chunks. Raw HTTP clients receive Server-Sent Events ending with data: [DONE].

typescript
const stream = await client.chat.completions.create({
  model: "neviri/flagship-chat",
  messages: [{ role: "user", content: "Write a two-line poem." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Handle errors

Errors include a machine-readable code and a message. Do not log API keys or complete request payloads while diagnosing failures.

HTTPMeaningWhat to do
401Authentication failedCheck that the nvai_ key is present, active, and sent as a bearer token.
402Insufficient balanceReduce the request's maximum completion size or ask the workspace owner to fund the wallet.
404Model not foundUse a model id returned by GET /v1/models.
422Invalid requestCheck the request body and ensure messages is not empty.
502Upstream errorRetry with backoff; the configured upstream provider failed.
504Upstream timeoutRetry or reduce the work requested from the model.

OpenAI compatibility

Neviri currently exposes OpenAI-compatible chat completions and model discovery. Chat requests accept messages, streaming, sampling controls, stop sequences, penalties, seeds, JSON response format, and tool definitions. Upstream model support can vary.

Available now

  • POST /v1/chat/completions
  • GET /v1/models
  • Streaming chat completions
  • Official OpenAI Python and Node clients

Not yet documented as available

  • Embeddings
  • Image, audio, and file APIs
  • Dedicated Neviri SDKs
  • Automatic fallback and routing controls

Documentation and platform roadmap

These items are planned, not currently available commitments. They will move into the main documentation only after their behavior is implemented and verified.

  • Embeddings API and integration examples
  • Published rate limits, response headers, and 429 guidance
  • Verified framework recipes for LangChain, LlamaIndex, AutoGen, and CrewAI
  • Interactive request builder and downloadable OpenAPI specification
  • Dedicated client and agent SDKs
  • AI-assistant access through an MCP documentation server
  • Provider controls, automatic fallbacks, presets, and routing policies