Quickstart
- 1
Create an inference key
Sign in, open API keys, and create an
nvai_…key. The full secret is shown once, so store it securely. - 2
Set your environment variable
Keep the key outside source control.
terminalexport NEVIRI_AI_KEY="nvai_your_api_key_here" - 3
Send a chat completion
Use the base URL
https://api.ai.neviri.com/v1with a model available in your environment.
Make your first request
Use raw HTTP or the official OpenAI SDK. Install the SDK first with pip install openai or npm install openai.
curl https://api.ai.neviri.com/v1/chat/completions \
-H "Authorization: Bearer $NEVIRI_AI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "neviri/flagship-chat",
"messages": [
{ "role": "user", "content": "Explain edge computing in one sentence." }
]
}'Expected response
Read generated text from choices[0].message.content. Token counts are returned in usage.
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "neviri/flagship-chat",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "Edge computing..." },
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 14,
"completion_tokens": 18,
"total_tokens": 32
}
}Discover available models
Model availability is deployment-specific. Query the catalogue before relying on a slug; the endpoint returns configured model ids, context windows, modalities, and reference pricing metadata.
curl https://api.ai.neviri.com/v1/modelsStream a response
Set stream: true to receive incremental OpenAI-compatible chunks. Raw HTTP clients receive Server-Sent Events ending with data: [DONE].
const stream = await client.chat.completions.create({
model: "neviri/flagship-chat",
messages: [{ role: "user", content: "Write a two-line poem." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Handle errors
Errors include a machine-readable code and a message. Do not log API keys or complete request payloads while diagnosing failures.
| HTTP | Meaning | What to do |
|---|---|---|
| 401 | Authentication failed | Check that the nvai_ key is present, active, and sent as a bearer token. |
| 402 | Insufficient balance | Reduce the request's maximum completion size or ask the workspace owner to fund the wallet. |
| 404 | Model not found | Use a model id returned by GET /v1/models. |
| 422 | Invalid request | Check the request body and ensure messages is not empty. |
| 502 | Upstream error | Retry with backoff; the configured upstream provider failed. |
| 504 | Upstream timeout | Retry or reduce the work requested from the model. |
OpenAI compatibility
Neviri currently exposes OpenAI-compatible chat completions and model discovery. Chat requests accept messages, streaming, sampling controls, stop sequences, penalties, seeds, JSON response format, and tool definitions. Upstream model support can vary.
Available now
- POST /v1/chat/completions
- GET /v1/models
- Streaming chat completions
- Official OpenAI Python and Node clients
Not yet documented as available
- Embeddings
- Image, audio, and file APIs
- Dedicated Neviri SDKs
- Automatic fallback and routing controls
Documentation and platform roadmap
These items are planned, not currently available commitments. They will move into the main documentation only after their behavior is implemented and verified.
- Embeddings API and integration examples
- Published rate limits, response headers, and 429 guidance
- Verified framework recipes for LangChain, LlamaIndex, AutoGen, and CrewAI
- Interactive request builder and downloadable OpenAPI specification
- Dedicated client and agent SDKs
- AI-assistant access through an MCP documentation server
- Provider controls, automatic fallbacks, presets, and routing policies