AGENTYK
Developers / API

Build on Agentyk

One OpenAI-compatible API for the whole AgentykLM lineup — chat, transcription, embeddings, and reranking — reachable with a single agk_ key. EU-hosted, no data retention, live docs that can’t drift.

Quickstart

If you already use the OpenAI SDK, you are three lines away: change the base URL, drop in your agk_ key, pick a model.

curl https://api.agentyk.xyz/v1/chat/completions \
  -H "Authorization: Bearer $AGENTYK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agentyk-jaguar-1",
    "messages": [
      { "role": "user", "content": "Summarise the EU AI Act in three points." }
    ]
  }'

Model ids are illustrative — call GET /v1/models or browse the live reference for the current lineup.

The surface

A focused, OpenAI-compatible /v1 surface plus self-serve account and key management.

POST

/v1/chat/completions

Chat & reasoning across the AgentykLM tiers. OpenAI-compatible request and response, including streaming and tool calls — plus a GBNF grammar field and strict JSON schema output when a reply has to be guaranteed, not requested.

POST

/v1/audio/transcriptions

Speech-to-text with Agentyk Scribe. Drop in any audio format; EU-hosted, the audio is never stored. €0.04 per hour of audio, pro-rated to the second, with a 10-second minimum billed length.

WSS

/v1/realtime?model=agentyk-scribe-live-1.0

Live transcription with Agentyk Scribe Live. Stream audio over a WebSocket and read the text while it is being spoken; the transcript at the end matches batch accuracy.

POST

/v1/embeddings

1024-dimension text embeddings (agentyk-embedding-1.0) for search and retrieval.

POST

/v1/rerank

Cross-encoder reranking (agentyk-reranker-1.0) to sharpen retrieval results.

GET

/v1/models

List the available models and their capabilities — the source of truth for model ids.

POST

/api/keys

Programmatically manage your agk_ API keys, plus credits, usage and subscriptions.

Why build on Agentyk

Drop-in

OpenAI-compatible

Point any OpenAI SDK at https://api.agentyk.xyz/v1 and change the model name. Existing chat, embeddings, and transcription code just works.

One key

A single agk_ key

One key reaches every surface — all ten chat tiers, transcription, embeddings, and reranking. No per-model credentials to juggle.

Private

EU-hosted, no retention

Inference runs in EU jurisdiction under GDPR. Prompts and audio are processed and dropped — nothing is persisted or used for training.

Always-correct

Generated, live docs

The Scalar reference at agentyk.xyz/docs renders the API's own generated OpenAPI spec, so the documentation can never drift from the implementation.

Making the output a guarantee

Most of the time you do not want the model's opinion about the format — you want a value from a fixed set, and you want it every time. Three mechanisms can narrow what comes back, and they are not interchangeable. Only one of them is a guarantee.

Prompt instructions

Guarantees: Nothing.

The model usually complies. “Usually” is not a guarantee, and the failures are the interesting ones.

response_format: json_object

Guarantees: The reply parses as JSON.

Not which keys. The same prompt can return cleaned_text one day and {language, text} the next.

response_format: json_schema + strict

Guarantees: Keys, nesting, types, array length, and any enum you declare.

Not the characters inside a plain "type": "string". That field accepts every string there is.

grammar (GBNF)

Guarantees: Exactly the language you wrote, character for character.

Nothing outside the grammar can be emitted — no near-miss, no code fence, no apology.

Grammar-constrained decoding

/v1/chat/completions accepts a grammarfield holding a GBNF grammar, and passes it to the model untouched. The grammar is applied while the answer is being generated: at every step, tokens that would take the output outside your grammar are removed from consideration. So the reply cannot be outside it. This is the only mechanism here that turns “please answer with one of these three labels” into something you never have to validate again.

It is worth being blunt about the trade: this is a llama.cpp extension, not an OpenAI-compatible parameter. Your SDK will not have a field for it, and that is expected — send it in the raw request body (extra_body in the OpenAI Python SDK, an extra key in the JSON body anywhere else). Everything else about the request and the response stays OpenAI-shaped.

tickets.gbnf
# Bind each record's id to its OWN rule, so the label can only land
# on the record it belongs to.
root     ::= "[" rec-1041 "," rec-1042 "]"
rec-1041 ::= "{\"id\":\"T-1041\",\"label\":" label "}"
rec-1042 ::= "{\"id\":\"T-1042\",\"label\":" label "}"
label    ::= "\"billing\"" | "\"bug\"" | "\"feature-request\""
sending it
GRAMMAR = open("tickets.gbnf").read()

resp = client.chat.completions.create(
    model="jaguar-express",
    max_tokens=80,
    messages=[{"role": "user", "content":
        "T-1041: 'I was charged twice this month.' "
        "T-1042: 'The export button does nothing.' Label each ticket."}],
    extra_body={"grammar": GRAMMAR},   # not an OpenAI field — send it raw
)

# [{"id":"T-1041","label":"billing"},{"id":"T-1042","label":"bug"}]
#
# There is no other possible response. The set of strings the model is
# allowed to emit is the set you wrote, and it is checked token by token.

Two traps, each of which costs a whole run

trap 1 — constrain the pair, not the vocabulary

The instinct is to list every permitted string once, as one global alternation, and let the model fill in the rest. That grammar is satisfied by a permitted string in the wrong record — it guarantees that each value exists in your vocabulary, not that it was attached to the right thing. If two fields have to agree with each other, the constraint has to bind the pair.

# TRAP. Every string here is permitted — so every PAIRING is permitted too.
root   ::= "[" record ("," record)* "]"
record ::= "{\"id\":" id ",\"label\":" label "}"
id     ::= "\"T-1041\"" | "\"T-1042\""
label  ::= "\"billing\"" | "\"bug\"" | "\"feature-request\""

# This grammar happily accepts all of the following:
#   [{"id":"T-1041","label":"bug"},{"id":"T-1041","label":"bug"}]   same id twice
#   [{"id":"T-1042","label":"billing"},{"id":"T-1041","label":"bug"}]   swapped
#
# A global list of permitted strings guarantees only that each string is
# permitted somewhere. If two fields have to agree, the rule that emits one
# must be the rule that emits the other.
trap 2 — rule names take no underscores

GBNF rule names accept letters, digits and hyphens only. A single underscore fails the entire grammar, and the request comes back as a bare 400 whose body reports an upstream error without mentioning the grammar at all. There is nothing in the response to point you at the real cause, so it is worth knowing before you meet it: if a grammar that looks correct returns 400, check the rule names first.

root ::= yes_or_no          # ✗ underscore
yes_or_no ::= "yes" | "no"  #   -> HTTP 400, and the body does not mention the grammar

root ::= yes-or-no          # ✓ letters, digits and hyphens only
yes-or-no ::= "yes" | "no"

Structured output, and the thing it does not do

response_format accepts json_schema with strict: true, not only json_object. It is the right tool when you need a predictable object to deserialise — and the wrong tool, on its own, when you need exact string content.

"response_format": {
  "type": "json_schema",
  "json_schema": {
    "name": "ticket",
    "strict": true,
    "schema": {
      "type": "object",
      "properties": {
        "id":    { "type": "string" },
        "label": { "type": "string" }
      },
      "required": ["id", "label"],
      "additionalProperties": false
    }
  }
}

// Returns, reliably:   { "id": "T-1041", "label": "Billing Issue" }
//
// Both keys present, both strings, nothing extra — the schema did its job.
// And "Billing Issue" is not one of your three labels, because you never
// told the schema what the labels were. Declaring
//     "label": { "type": "string", "enum": ["billing","bug","feature-request"] }
// changes that answer to "billing", because an enum becomes a real
// constraint. A bare "type": "string" never is.
the caveat that matters

A strict schema constrains the shape of a field and says nothing about the characters inside it. Every key you declared will be there, correctly typed, with nothing extra — and a field declared "type": "string" will accept any string in the world, including a plausible-looking wrong one. On a task that needs exact string content, a strict schema scores the same as plain JSON mode. If the value must come from a fixed set, say so with enum, which is enforced, or drop to a grammar.

counts and lengths

Array bounds do hold: under "maxItems": 3 you get three elements. But a bound on the count is not a bound on the content — asked for forty items under a limit of three, the reply came back with three elements and the other thirty comma-separated inside the third string. If the number of items is part of what you are relying on, write the repetition into a grammar, where the count and the shape of each item are constrained together.

Defaults worth knowing

We set no sampling parameters on your behalf. A request that sends none is served with the tier's own built-in defaults, and those are tuned for open-ended generation rather than for repeatability.

sampling

Send no sampling parameters and the express tier runs at temperature 1.0, top-k 64, top-p 0.95, min-p 0.05. That is a creative default, not a sober one: the same prompt sent twice returns two different answers, which surprises people who expected near-determinism out of the box. temperature, top_p, top_k and seed all pass through and are honoured, so ask for what you need — temperature: 0 with a fixed seed returns the same answer run after run.

reasoning

Reasoning is off by default, so you are not paying for hidden thinking tokens you did not ask for. Turn it on per request with "chat_template_kwargs": { "enable_thinking": true } and the reply carries a reasoning_content field alongside the usual content; set it explicitly and your setting is respected rather than overridden. The OpenAI reasoning_effort field is accepted but currently has no effect on this tier — a request carrying only reasoning_effort is still served with reasoning off, so use enable_thinking as the switch.

max_tokens

Set one. There is no small implicit cap standing in for you, and a prompt that invites a long answer will be given one — we have watched a single request run to twelve thousand output tokens for a caller who set no limit. An explicit max_tokens bounds both your bill and your latency, and it is a hard requirement rather than a nicety if you use constrained tool calls, below.

Constrained tool calls, and the two headers that tell you the truth

Set agentyk_strict_tools: true on a non-streaming call that forces exactly one function, and the function's parameter schema — enum and pattern included — is enforced on the arguments while they are generated, instead of being a suggestion the model usually follows.

resp = client.chat.completions.with_raw_response.create(
    model="jaguar-express",
    max_tokens=2500,         # ~1.5x the longest output your prompt can produce
    messages=messages,
    tools=tools,             # exactly one function
    tool_choice={"type": "function", "function": {"name": "classify"}},
    extra_body={"agentyk_strict_tools": True},
)

mode = resp.headers.get("X-Agentyk-Strict-Tools")         # enforced | fallback
why  = resp.headers.get("X-Agentyk-Strict-Tools-Reason")  # only when fallback
if mode != "enforced":
    logger.warning("strict tools fell back: %s", why)

completion = resp.parse()
why the headers exist

Enforcement works by constraining the output to your schema. If that generation hits the token cap, the output is a valid JSON prefix and nothing more, so we fall back to an unconstrained attempt rather than fail your request.

The fallback usually still returns well-formed output. From your side nothing looks wrong — you get a tool call with sensible arguments — and the token count you receive belongs to the second attempt, so it always appears comfortably under your cap. A client-side histogram of completion tokens cannot find this. These two response headers are the only way to detect it.

Log X-Agentyk-Strict-Tools on every call and alert when fallback rises above a small threshold of your traffic — a rising rate almost always means max_tokens is now straddling your output length distribution, and raising it restores the guarantee. Read them server-side: they are not exposed to browser JavaScript.

enforced

— none —

The constraint held. The arguments match your schema.

fallback

max_tokens_too_small

The constrained attempt was cut off mid-JSON. Raise max_tokens. This is the actionable one, and the common one.

fallback

empty_content

The constrained attempt produced no JSON.

fallback

invalid_json

The constrained attempt produced something that was not JSON.

fallback

upstream_error / unavailable

A backend error, or no capacity for the constrained attempt.

fallback

no_forced_tool_call

The request does not qualify: it needs exactly one function and a tool_choice that forces that function.

fallback

unsupported_stream

stream: true. The constrained path is non-streaming only, so opting in on a stream does nothing.

What usage costs

Plans are flat monthly; metered usage on the API is billed per unit on top. These are the published rates — GET /api/credits/rates returns the same card, complete and machine-readable, for every tier.

Transcription

€0.04 per hour of audio

Pro-rated to the second and rounded once over the whole request. Minimum billed length 10 seconds: a shorter clip is charged as 10 seconds. A request whose duration cannot be measured is not charged at all, and is never lifted to the minimum.

Express language tier

€0.15 / €0.60 per million tokens

€0.15 per million input tokens, €0.60 per million output. Cached input is billed at half the input rate — €0.075 per million. Tokens are counted from the usage block of your own response.

Volume discounts apply automatically on month-to-date spend, and live transcription sessions are not metered per clip — see plans and pricing.

Live transcription

Agentyk Scribe Live streams text back while the audio is still being spoken. Open one WebSocket per speaker, send raw PCM at real time, and read partials within about half a second; opt in to refine and the transcript you hold at the end is the same one a batch request would have returned.

import asyncio, json, wave, websockets

URL = "wss://api.agentyk.xyz/v1/realtime?model=agentyk-scribe-live-1.0"

async def main():
    async with websockets.connect(URL, additional_headers={"Authorization": "Bearer agk_..."}) as ws:
        await ws.send(json.dumps({"type": "start", "language": "de", "refine": True}))

        async def read():
            async for m in ws:
                f = json.loads(m)
                if f["type"] == "partial":        print("...", f["text"])   # provisional; replaces the previous partial
                elif f["type"] == "final":        print("   ", f["text"])   # committed
                elif f["type"] == "transcript":   print("== ", f["text"])   # the whole session, batch accuracy
                elif f["type"] in ("done", "error", "busy"):
                    print(f); return

        reader = asyncio.create_task(read())
        w = wave.open("meeting.wav", "rb")           # 16 kHz, mono, 16-bit PCM
        while frames := w.readframes(1600):          # 100 ms per frame, sent at real time
            await ws.send(frames)
            await asyncio.sleep(0.1)
        await ws.send(json.dumps({"type": "stop"}))
        await reader

asyncio.run(main())
you send
  • start — language (ISO 639-1 or auto), refine, sample_rate, format, channels, diarize (speaker labels), final_labels (with diarize; default false: the more-than-four-speakers mode, see speakers below)
  • binary — raw PCM frames, 20–100 ms each, at real time
  • stop — flush everything and end the session
you receive
  • ready — session open; echoes the accepted settings
  • partial — provisional text, half a second behind the speaker; replaces the previous partial
  • final — committed text, appended to the utterance
  • utterance — with refine: the finished utterance re-decoded with full context, replacing its finals
  • transcript — with refine, at stop: the whole session re-decoded — batch accuracy
  • done — audio_ms consumed, utterance count, why the session ended, speakers when labels were on, and final_pass (whether the re-check was delivered)
  • speaker — with diarize: a speaker number on partials and finals (finals split where the speaker changes), turns on utterance and transcript; up to four speakers while live
  • speakers — with diarize and final_labels: true, at stop: every speaker turn of the session re-checked with full context (pass: final, supersedes: all). Live labels cover up to four speakers; ask for this when the room has more, and expect stop to take longer (typically under a minute). Below four speakers the live labels are the better ones. Apply it over the labels you hold; the text is unchanged
  • busy / error — capacity reached (retry after retry_after_ms) / an actionable error code

Audio is 16-bit or float PCM, mono or stereo, 8–48 kHz — declare it once in start. Sessions run up to eight hours and end on their own, transcript included, after thirty minutes without speech; keep sending silence during breaks, it costs nothing. Twenty-five European languages are covered by default; other languages, Persian among them, work with an explicit language code. Same agk_ key and the same Scribe plan as batch transcription, but metered per session rather than per clip: usage is counted per audio second and the 10-second minimum billed length that applies to a batch request does not apply here.

SDKs

On-device inference SDKs

Ship sovereign AI inside your own apps. The Agentyk SDKs run the smaller AgentykLM tiers on-device and fall back to the cloud API when a task needs a bigger model — the same agk_ key throughout. SDK access is included from the Pro plan up.

Knowledge API

Ground answers you can verify

The Agentyk Knowledge API turns your documents into a verifiable atomic claim graph and answers against it — every response cites the exact source and ships a verification code.

Start building in minutes

Create a key, point your SDK at api.agentyk.xyz, and ship on European, sovereign infrastructure.