Errors and limits
Two envelopes
Inference endpoints (OpenAI-compatible) return OpenAI-style errors:
{"error": {"message": "key not allowed to access model. …", "type": "key_model_access_denied"}}
Platform REST endpoints return Frappe's envelope — successful results are
wrapped in message, and errors carry exc_type plus a readable message:
{"message": {"run_id": "...", "status": "Queued"}}
Errors you will actually meet
| Error | Meaning | What to do |
|---|---|---|
invalid_api_key / 401 |
Key wrong, rotated or disabled | Check the secret; rotate in the portal |
key_model_access_denied |
Model not on the key's allowlist | Use GET /v1/models; rotate the key to change its allowlist |
Key lacks scope(s): … |
Endpoint outside the key's scopes | Issue a key with the right scopes |
residency_policy_violation |
Cloud model under a local-only policy | Use a local-badged model — this is by design |
budget_exceeded / gateway refusal |
Wallet exhausted | Top up, or enable auto-reload |
429 |
Rate limit | Back off and retry |
413 |
Payload over 256 KB (workflow trigger) | Send a reference, not the blob |
Waiting Approval (a status, not an error) |
A human must approve a sensitive step | Approve in the portal; the run resumes |
Retries, the safe way
- Chat/embeddings: idempotent by nature — retry on 5xx.
- Workflow starts: send an
Idempotency-Keyheader — a replay returns the original run ("replayed": true), never a duplicate. - Agent runs: poll
GET /v1/runs/{id}rather than re-posting; a run that died before its first step is automatically revived by the platform within minutes.