Notices
Last updated: 23 September 2026
Agentyk is built on open-weight AI, on open, openly-licensed data, and on open-source software. This page explains, at a high level, the foundation models our hosted Service draws on, the open knowledge-graph data and ontologies that Agentyk Knowledge builds on, the open-source software libraries we build with, and the licences and attributions that apply. It forms part of, and is incorporated by reference into, our Terms of Service.
1. We build on open-weight models
The language, speech, embedding, reranking, and document-OCR / vision models behind AgentykCloud chat and our APIs (including Agentyk Scribe and Agentyk Knowledge) are built on, and fine-tuned and optimised from, open-weight foundation models— models whose trained parameters are published under an open or source-available licence. We run them on our own European infrastructure. We do not build the Service on closed, proprietary models accessed over someone else's API, and we do not use your prompts, inputs, or outputs to train models.
2. Which models we use
We continuously evaluate, combine, fine-tune, and rotate the models we run as the open-weight ecosystem advances. Depending on the product, the tier, the region, and the time, our deployments may draw on models from families such as:
- Google Gemma
- Meta Llama
- Alibaba Qwen
- Baidu (PaddleOCR, ERNIE)
- Mistral and Mixtral
- Microsoft Phi
- DeepSeek
- IBM Granite
- TII Falcon
- OpenAI open-weight (gpt-oss) models
- Whisper-family and other open speech-to-text models
- other open speech-to-text models, including NVIDIA Parakeet, CTranslate2 / faster-whisper and whisper.cpp conversions of OpenAI Whisper, and OpenMOSS MOSS-Transcribe-Diarize
- open speaker-diarisation and voice-activity-detection models, including NVIDIA Sortformer and Silero VAD
- open text-to-speech and voice models, including Resemble AI Chatterbox and Chatterbox Multilingual, Kokoro, and StyleTTS 2-based models
- open talking-head and avatar models, including SoulX FlashHead, together with Meta wav2vec 2.0 audio features and Google MediaPipe face detection
- open diffusion language models, including diffusion releases from the families above
- open document-OCR and vision-language models (for reading scanned and image-only documents), including PP-OCR / PaddleOCR and vision-language releases from the families above
- open embedding and reranking models (for retrieval and knowledge-graph grounding), including releases from the families above
- Microsoft Harrier and other open embedding models
- open information-extraction and coreference models (for turning documents into a knowledge graph) — including GLiNER-family named-entity and relation extractors, fastcoref coreference, and ReFinED-family entity linkers
- and other open-weight models, including our own fine-tunes of the above
This list is illustrative and non-exhaustive, and the specific models in use change over time. For commercial, security, and operational reasons, we do not disclose which specific model, version, or configuration powers any particular product tier or codename, and a given codename may map to different underlying models over time.
3. How we build and serve our models
We do more than run these models unchanged. To deliver the Service we build derivative and optimised model systems by applying a range of engineering and machine-learning techniques. Depending on the product, the tier, and the time, these may include, among others:
- Fine-tuning and adaptation— supervised fine-tuning, instruction tuning, parameter-efficient methods (such as LoRA and adapters), continued or domain pre-training, and preference or reinforcement-learning alignment (such as RLHF, RLAIF, or DPO).
- Distillation— training smaller or faster models to reproduce the behaviour of larger ones.
- Quantization and pruning— reducing numerical precision or removing redundant parameters to serve models more efficiently.
- Model merging, ensembling, and mixture-of-experts — merging or combining multiple models, and routing or multi-model setups that run several models and return a selected, consensus, or best answer.
- Prompt and system-prompt engineering— instructions, templates, and few-shot examples that shape behaviour.
- Retrieval-augmented generation, knowledge graphs, and search grounding — supplementing a model with retrieved documents, structured knowledge, embeddings, or web and search results.
- Tool use and agentic harnesses— orchestrating multi-step reasoning, tool or function calls, planning, and verification or self-consistency loops.
- Content moderation and safety— classifiers, guardrails, and filtering applied before, during, or after generation.
- Inference and serving optimisation— techniques such as speculative decoding, batching, caching, KV-cache and decoding optimisations, reranking, and constrained or structured decoding.
- Other related techniques— we continuously adopt new methods as the field advances.
The result is a derivative system that may behave differently from, and perform better than, any single underlying model. This description is general, illustrative, and non-binding: the specific techniques and their combination change over time, we apply them differently across products and tiers, and nothing here commits us to using, or continuing to use, any particular technique.
4. Knowledge-graph data & ontologies
Agentyk Knowledge structures documents into a verifiable knowledge graph using open, openly-licensed ontologies and knowledge bases. Each is used under its own licence:
- YAGO(Creative Commons Attribution-ShareAlike) — the general base ontology, taxonomy, and SHACL constraints, adopted version-pinned and unmodified.
- schema.org(Creative Commons Attribution-ShareAlike 3.0) — the general-purpose vocabulary that YAGO builds upon.
- Wikidata(Creative Commons CC0 / public domain) — stable identifiers used to link and disambiguate entities.
- W3C PROV-O and Dublin Core Terms — provenance and document-metadata vocabularies.
We use these resources under their respective licences, and where a licence requires attribution, this page serves as that attribution. Any terms a customer defines to extend the vocabulary for their own knowledge base are stored in the customer's own database and are the customer's own work and data — not part of, and not distributed by, Agentyk.
Beyond these ontologies, the Service is built, trained, and evaluated with other open data, each used under its own licence:
- GDELT — news-event data from The GDELT Project, used for search grounding.
- Open-access scholarly articles published under Creative Commons Attribution licences, discovered through OpenAlex(metadata CC0) — each article is attributed to its authors under its own licence.
- CML-TTS(Creative Commons Attribution 4.0), built on public-domain LibriVox recordings — used to train and fine-tune speech models.
- VoxPopuli (CC0), and the AMI Meeting Corpus, LibriSpeech, FLEURS, Multilingual LibriSpeech, and VoxConverse(Creative Commons Attribution 4.0) — used to evaluate and benchmark speech recognition and speaker diarisation.
5. Licences and attribution
Each foundation model we use is used under its own open-source or source-available licence (for example the Gemma Terms of Use, the Llama Community Licence, and the Apache 2.0 or MIT licences under which several of the families above are released), and the open data and ontologies in section 4 are used under their respective licences (such as Creative Commons Attribution-ShareAlike and CC0). We comply with those licences, including any attribution requirements and use restrictions they impose, and we pass through use restrictions to you via our Acceptable Use Policy. Where a licence requires it, this page and our notices serve as the required attribution (for example, products in this Service that are built with Llama are “Built with Llama”). The respective model names, dataset names, and trademarks belong to their owners; their inclusion here is for attribution and does not imply that those owners endorse, sponsor, or are affiliated with Agentyk or Sylvanity B.V. Agentyk® is a registered trademark of Sylvanity B.V., and Agentyk and the Agentyk logo are Sylvanity B.V.'s own marks.
NVIDIA models. NVIDIA Parakeet is used under the Creative Commons Attribution 4.0 licence, and we have converted and optimised it for serving. NVIDIA Sortformer: Licensed by NVIDIA Corporation under the NVIDIA Open Model License; we have converted and optimised it for serving.
6. Open-source software components
Beyond the models, the Service is built with open-source software libraries — most under a permissive licence (Apache-2.0, MIT, or BSD), and some under the weak-copyleft LGPL or the GPL (see the notices below) — where a licence requires attribution, this page serves as it. Among the more significant, for information extraction and the knowledge graph:
- GLiNER / GLiNER2(Apache-2.0) — zero-shot named-entity and relation extraction.
- fastcoref(MIT) — coreference resolution.
- sentence-transformers (Apache-2.0) with a DeBERTa-v3 NLI cross-encoder(Apache-2.0) — a natural-language-inference faithfulness gate (verifying each claim against its source span).
- ReFinED(Apache-2.0) — entity linking to Wikidata.
- spaCy(MIT) — NLP pipeline utilities.
- rapidfuzz (MIT) and networkx(BSD) — batch entity deduplication (fuzzy matching + clustering).
- quantulum3(MIT) — parsing quantities and units from text, used to validate the numeric effect sizes (odds ratios, confidence intervals, percentages) extracted into the knowledge graph.
- RapidOCR and Tesseract(Apache-2.0) — document OCR.
- PyTorch (BSD-3-Clause) and Hugging Face Transformers(Apache-2.0) — the model runtime.
- and other Apache-2.0, MIT, and BSD-licensed libraries across our services (such as FastAPI, RDFLib, and Oxigraph).
For model inference and serving:
- llama.cpp and ggml(MIT) — model inference across CPUs and GPUs.
- llama2.cby Andrej Karpathy (MIT) — minimal C inference.
- vLLM (Apache-2.0) and Hugging Face Text Embeddings Inference(Apache-2.0) — model serving.
- Ollama(MIT) — model packaging and serving.
- NVIDIA TensorRT-LLM (Apache-2.0) and NVIDIA Triton Inference Server(BSD-3-Clause) — optimised model serving.
- ONNX Runtime(MIT) — running compact models.
- Qdrant (Apache-2.0) and ClickHouse(Apache-2.0) — vector search and analytical storage.
For speech, voice, and avatar:
- whisper.cpp (MIT), faster-whisper and CTranslate2(MIT) — speech recognition.
- NVIDIA NeMo, NeMo-Speech.cpp, and NeMo text processing (Apache-2.0), with Pynini and OpenFst(Apache-2.0) — speech recognition, diarisation, and inverse text normalisation.
- moss-transcribe.cpp(MIT) — transcription with speaker diarisation.
- chatterbox.cpp (MIT), the Chatterbox inference code (MIT), open Chatterbox fine-tuning tooling (Apache-2.0), and S3Tokenizer(Apache-2.0) — speech synthesis and fine-tuning.
- Perthby Resemble AI (MIT) — audio watermarking of synthesised speech.
- kokoro and misaki (Apache-2.0), the StyleTTS 2 code (MIT), and kikiri-tts(Apache-2.0) — speech synthesis, grapheme-to-phoneme conversion, and training.
- MediaPipe(Apache-2.0) — face detection.
This list is illustrative and non-exhaustive and changes as our stack evolves. Each component remains under its own licence, and the respective names and trademarks belong to their owners.
LGPL notice. One transitive dependency, num2words(LGPL-2.1) — pulled in by quantulum3 to spell out numbers — is licensed under the GNU Lesser General Public License rather than a permissive licence. We use it unmodified, as a standard library dependency of our server-side Service. This notice serves as the attribution and notice its licence requires; a copy of the LGPL and of the library's source is available on request, and you may obtain, replace, or relink the library under the terms of the LGPL.
The same library, num2words, is also used directly, unmodified and on the same LGPL terms, to normalise numbers in text before speech synthesis.
GPL notice. A few components are licensed under the GNU General Public License. For grapheme-to-phoneme conversion in speech synthesis we use eSpeak NG (GPL-3.0-or-later), loaded through the phonemizer library (GPL-3.0); and we run Poppler's PDF utilities (GPL) and a GPL-enabled build of FFmpeg as separate programs for document rendering and media encoding. We use each of them unmodified, on our own servers, as part of our server-side Service; we do not distribute them. This notice serves as their attribution; a copy of the GPL and of the corresponding source is available on request.
7. No warranty from upstream
Open-weight models and open data are provided by their authors without warranty. Output generated through the Service is subject to the disclaimers in our Terms of Service — it may be inaccurate and is not professional advice. You remain responsible for reviewing and verifying output before relying on it.
8. Changes
Because we rotate and upgrade models and update the data and ontologies we build on, we revise this page from time to time; the “Last updated” date above reflects the latest revision. Questions about model or data licensing or attribution: info@sylvanity.eu.