Inference API (OpenAI-compatible)
One endpoint speaks the protocol every AI library already knows. If your code
works with OpenAI, it works here — change base_url, keep everything else.
from openai import OpenAI
client = OpenAI(base_url="https://platform.senaiy.ai/v1", api_key="sk-sf-...")
What is supported
| Capability | Endpoint | Status |
|---|---|---|
| Chat completions | POST /v1/chat/completions |
✅ verified in production |
| Streaming (SSE) | same, "stream": true |
✅ verified |
| Vision (image input) | same, image_url content on a vision model |
✅ verified — see below |
| Embeddings | POST /v1/embeddings |
✅ verified |
| Model listing | GET /v1/models |
✅ verified — scoped to your key |
| Speech-to-text | POST /v1/audio/transcriptions |
✅ verified — availability follows your key's allowlist |
| Image generation / Assistants / Files / Fine-tuning | — | ❌ not offered — use Agents & Knowledge instead |
One interface, many providers
The catalog spans OpenAI (gpt-4o, gpt-4o-mini), Anthropic (claude-sonnet,
claude-haiku), Google (gemini-flash), regional and local engines — all through
the same API shape, all rated in USD against the same wallet. Your key's allowlist
decides which of them you see and may call.
for model in client.models.list().data:
print(model.id)
Streaming
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "اكتب فقرة عن الرياض"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Vision (image input)
Vision-capable models (gpt-4o, gpt-4o-mini, claude-sonnet, claude-haiku,
gemini-flash, …) accept images inline using the standard image_url content part —
pass a public URL or a data: URI. Check a model's supports_vision flag in the portal's
Developer → Models directory before relying on it.
resp = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "https://…/photo.png"}},
]}],
)
Speech-to-text
resp = client.audio.transcriptions.create(
model="whisper-large-v3",
file=open("meeting.mp3", "rb"),
)
print(resp.text)
Which speech models you can call is decided by your key's allowlist, same as any other model.
Residency
Organizations under a local_only residency policy may only call local-badged
engines; a cloud model returns residency_policy_violation. That refusal is the
feature: data governed to stay in-country stays in-country.
Structured output
Pass response_format with a JSON schema on models that support it and validate
the result server-side — or better, put the schema on an agent or workflow
node, where the platform validates and retries for you.