Inference API (OpenAI-compatible)

One endpoint speaks the protocol every AI library already knows. If your code works with OpenAI, it works here — change base_url, keep everything else.

from openai import OpenAI

client = OpenAI(base_url="https://platform.senaiy.ai/v1", api_key="sk-sf-...")

What is supported

Capability Endpoint Status
Chat completions POST /v1/chat/completions ✅ verified in production
Streaming (SSE) same, "stream": true ✅ verified
Vision (image input) same, image_url content on a vision model ✅ verified — see below
Embeddings POST /v1/embeddings ✅ verified
Model listing GET /v1/models ✅ verified — scoped to your key
Speech-to-text POST /v1/audio/transcriptions ✅ verified — availability follows your key's allowlist
Image generation / Assistants / Files / Fine-tuning — ❌ not offered — use Agents & Knowledge instead

One interface, many providers

The catalog spans OpenAI (gpt-4o, gpt-4o-mini), Anthropic (claude-sonnet, claude-haiku), Google (gemini-flash), regional and local engines — all through the same API shape, all rated in USD against the same wallet. Your key's allowlist decides which of them you see and may call.

for model in client.models.list().data:
    print(model.id)

Streaming

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "اكتب فقرة عن الرياض"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Vision (image input)

Vision-capable models (gpt-4o, gpt-4o-mini, claude-sonnet, claude-haiku, gemini-flash, …) accept images inline using the standard image_url content part — pass a public URL or a data: URI. Check a model's supports_vision flag in the portal's Developer → Models directory before relying on it.

resp = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "image_url", "image_url": {"url": "https://…/photo.png"}},
    ]}],
)

Speech-to-text

resp = client.audio.transcriptions.create(
    model="whisper-large-v3",
    file=open("meeting.mp3", "rb"),
)
print(resp.text)

Which speech models you can call is decided by your key's allowlist, same as any other model.

Residency

Organizations under a local_only residency policy may only call local-badged engines; a cloud model returns residency_policy_violation. That refusal is the feature: data governed to stay in-country stays in-country.

Structured output

Pass response_format with a JSON schema on models that support it and validate the result server-side — or better, put the schema on an agent or workflow node, where the platform validates and retries for you.

Discard
Save
This page has been updated since your last edit. Your draft may contain outdated content. Load Latest Version

On this page

Review Changes ← Back to Content
Message Status Space Raised By Last update on