Prepare AI Models
Provide inference endpoints from cloud API providers (OpenAI, Anthropic, Gemini, Bedrock, Azure), self-hosted open-weight models with vLLM, or both. After installing Kindo, register each model from the superadmin dashboard’s Models page. This page covers endpoint preparation and the information you need for registration.
How models get registered
Section titled “How models get registered”After the install:
- Open the superadmin dashboard’s Models page and select the Models tab.
- Use the Required models panel to add an embedding or transcription model with its preset. To register another model, open the add-model sheet.
- Choose the provider, enter its credentials, review the model settings, and save.
- Open the Feature flags tab on the Models page to choose which models each feature uses.
Registration makes the model available through the endpoint you supplied. Model Configuration explains model groups and provider fallbacks.
Your job on this page is to:
- Decide which providers/endpoints you will use.
- Make sure each endpoint is reachable from the cluster.
- Collect the facts each model’s registration needs.
What to know about each model before you register it
Section titled “What to know about each model before you register it”Collect the following information before opening the add-model sheet. Credentials depend on the provider.
| Fact | Required | What it is |
|---|---|---|
| Serving name | yes | The unique routing key Kindo and LiteLLM address the model by. Must match --served-model-name for vLLM, and must be the reserved name for a required role. |
| Display name | yes | Human-readable name shown in the UI model picker. |
| Provider | yes | The LiteLLM provider, which sets the connection type and the credential fields the form asks for. |
| Provider display name | yes | The provider your users see for this model: pick an existing one, or choose New provider… and name it. A new name can’t match an existing provider in any letter case. |
| LiteLLM model ID | yes | The provider-prefixed model ID LiteLLM routes with (see the prefix table below). |
| Credential | varies | API key, endpoint URL, API version, AWS keys and region, or a service account, whichever the provider needs. See the playbooks below. |
| Context window | yes for chat | The model’s full advertised context window. Do not low-ball it. Kindo seeds it from the catalog when the model ID matches; a chat model with neither value nor match is rejected. |
| Max output tokens | recommended | The largest response the model is allowed to generate in one call. |
| Extra headers | optional | Enter {"extra_headers": {...}} in Additional litellm_params (JSON, optional) for headers sent with every inference request, such as an x-api-key gateway credential. |
| Cost tier | recommended | LOW, MEDIUM, or HIGH, carried in the model’s metadata. |
| Description / link | optional | Surfaced in the model picker. |
Provider prefix reference
Section titled “Provider prefix reference”| Provider | LiteLLM model ID prefix | Example |
|---|---|---|
| OpenAI | openai/ | openai/gpt-4o |
| Anthropic | anthropic/ | anthropic/claude-sonnet-4-20250514 |
| Google Gemini | gemini/ | gemini/gemini-2.5-pro |
| Azure OpenAI | azure/ | azure/<your-deployment-name> |
| AWS Bedrock | bedrock/ | bedrock/us.anthropic.claude-sonnet-4-20250514-v1:0 |
| Self-hosted (vLLM) | hosted_vllm/ or openai/ | hosted_vllm/deephat-v2, openai/deephat-v2 |
Cloud provider playbooks
Section titled “Cloud provider playbooks”For cloud models, prepare an API key, the model ID with its provider prefix, and network egress.
Have ready:
- LiteLLM model ID:
openai/gpt-4o - Credential: your OpenAI API key
- Context window / max output tokens: the values OpenAI publishes for that model (128000 / 16384 for GPT-4o)
Notes:
- Use the exact model ID OpenAI publishes (e.g.
gpt-4o,gpt-4o-mini). Outdated IDs silently fall back or fail. - Leave the endpoint URL empty. LiteLLM routes to
https://api.openai.com/v1by default. - Organization ID can be supplied by setting
OPENAI_ORG_IDin the LiteLLM environment if needed.
Have ready:
- LiteLLM model ID:
anthropic/claude-sonnet-4-20250514 - Credential: your Anthropic API key
- Context window / max output tokens: 200000 / 16384 for Claude Sonnet 4
Notes:
- Use the dated model ID (e.g.
claude-sonnet-4-20250514), not an alias, for reproducible deployments.
Have ready:
- LiteLLM model ID:
gemini/gemini-2.5-pro - Credential: your Google AI API key
- Context window / max output tokens: 1000000 / 65536 for Gemini 2.5 Pro
Have ready:
- LiteLLM model ID:
bedrock/us.anthropic.claude-sonnet-4-20250514-v1:0 - AWS region and one authentication method
- Context window / max output tokens: the values the underlying model publishes
Notes:
- Copy a model or inference-profile ID from the AWS Bedrock model cards, then prefix it with
bedrock/. - In the add-model sheet, choose API key, IRSA or EKS Pod Identity (no keys), or Static IAM user keys. A static key overrides the pod role.
- Use cross-region inference profiles (e.g.
us.anthropic.claude-*) instead of region-specific ARNs for better availability. IAM must then allow both the profile and the underlying model in every region the profile routes to. - IAM requires
bedrock:InvokeModelandbedrock:InvokeModelWithResponseStream. Chat streams, so it fails without the second action. - Third-party models sold through AWS Marketplace, such as Anthropic and OpenAI, need a one-time subscription per account before they can be called.
Keyless access (IRSA):
LiteLLM calls Bedrock as the service account litellm in the litellm namespace. To use a pod role instead of keys:
-
Create an IAM role with the Bedrock permissions above and a trust policy that lets that service account assume it through the cluster’s OIDC provider:
{"Version": "2012-10-17","Statement": [{"Effect": "Allow","Principal": {"Federated": "arn:aws:iam::<account-id>:oidc-provider/oidc.eks.<cluster-region>.amazonaws.com/id/<oidc-id>"},"Action": "sts:AssumeRoleWithWebIdentity","Condition": {"StringEquals": {"oidc.eks.<cluster-region>.amazonaws.com/id/<oidc-id>:aud": "sts.amazonaws.com","oidc.eks.<cluster-region>.amazonaws.com/id/<oidc-id>:sub": "system:serviceaccount:litellm:litellm"}}}]}<oidc-id>is the last path segment ofaws eks describe-cluster --name <cluster> --query cluster.identity.oidc.issuer. With EKS Pod Identity, associate the role with service accountlitellmin namespacelitellminstead. -
From the install directory, annotate the service account with the role and apply it:
Terminal window kindo config helm-override set litellm \'serviceAccount.annotations.eks\.amazonaws\.com/role-arn=arn:aws:iam::<account-id>:role/<role-name>'kindo config helm-override apply litellm -
Restart LiteLLM so its pods pick up the role, then confirm they carry it:
Terminal window kubectl -n litellm rollout restart deployment/litellmkubectl -n litellm exec deploy/litellm -- printenv AWS_ROLE_ARN
Have ready:
- LiteLLM model ID:
azure/<your-deployment-name> - Credential: the Azure endpoint URL (e.g.
https://my-resource.openai.azure.com), API version (e.g.2024-10-21), and API key - Base model: the underlying OpenAI model (
gpt-4o,gpt-4o-mini, etc.) - Context window / max output tokens: the values the underlying OpenAI model publishes
Notes:
- The LiteLLM model ID uses your Azure deployment name, not the OpenAI model name.
- The endpoint URL, API version, and API key are all required. Azure will not route without them.
- Set
baseModelto the underlying OpenAI model (gpt-4o,gpt-4o-mini, etc.) so LiteLLM can report correct pricing and capabilities.
For an OpenAI-compatible gateway, use the openai/ prefix and the gateway’s endpoint URL. For a vLLM server behind the gateway, hosted_vllm/ also works. If the gateway requires a custom header such as x-api-key, enter it as {"extra_headers": {...}} in Additional litellm_params (JSON, optional).
Have ready:
- LiteLLM model ID:
openai/gpt-4o; the suffix must match the gateway’s served model ID - Credential: the gateway endpoint URL (e.g.
https://gateway.example.com/v1) and its API key - Additional litellm_params (JSON, optional): the gateway credential under
extra_headers, using the header the gateway requires - Context window / max output tokens: the values the upstream model publishes
Notes:
- Check the gateway’s
GET /v1/modelslisting for the exact served model IDs; many gateways prefix or suffix the upstream model name. - Headers in Additional litellm_params (JSON, optional) are sent with every inference request to this model.
- Keep the API key field set when the gateway authenticates through a custom header.
For a gateway that requires x-api-key, enter:
{ "extra_headers": { "x-api-key": "<gateway secret>" } }Self-hosted vLLM patterns
Section titled “Self-hosted vLLM patterns”Self-hosting gets you full data-residency control and unlimited request volume, but you own the inference server. Kindo recommends vLLM, and this section covers it.
Pre-deployment checklist
Section titled “Pre-deployment checklist”- GPUs sized for your target context length (table below).
- NVIDIA drivers +
nvidia-container-runtimeinstalled;nvidia-smishows your GPUs. - Model weights pulled (HuggingFace token exported:
export HF_TOKEN=...). - Official model card read: context window, max output tokens, dtype, and any special flags.
- Correct
--tool-call-parseridentified for your model family. - DNS/hostname planned; the endpoint will be reachable from the cluster.
GPU sizing
Section titled “GPU sizing”Context length is the primary VRAM driver via the KV cache. A 70B model at 4k context needs far less VRAM than the same model at 128k.
| Model size | Quantization | Short (4k) | Medium (32k) | Full (128k+) |
|---|---|---|---|---|
| 7–8B | FP16/BF16 | 1× 24GB | 1× 24GB | 1× 48GB |
| 7–8B | FP8 | 1× 24GB | 1× 24GB | 1× 24GB |
| 13B | FP16/BF16 | 1× 48GB | 1× 80GB | 1× 80GB |
| 13B | FP8 | 1× 24GB | 1× 48GB | 1× 80GB |
| 30–34B | FP16/BF16 | 1× 80GB | 2× 80GB | 2–4× 80GB |
| 30–34B | FP8 | 1× 48GB | 1× 80GB | 1–2× 80GB |
| 70B | FP16/BF16 | 2× 80GB | 4× 80GB | 4–8× 80GB |
| 70B | FP8 | 1× 80GB | 2× 80GB | 2–4× 80GB |
| 70B | AWQ/GPTQ | 1× 80GB | 2× 80GB | 4× 80GB |
Essential vLLM flags
Section titled “Essential vLLM flags”vllm serve <model-id> \ --served-model-name <name> # must match the serving name you register --port 8000 \ --max-model-len <context-length> # the model's full context window --dtype bfloat16 \ --tensor-parallel-size <num-gpus> \ --tool-call-parser <parser> \ # required for agents --enable-auto-tool-choice \ # required alongside tool-call-parser --enable-prefix-caching \ --enable-chunked-prefillCheck these flags if the model fails to start or call tools:
--max-model-len: set to the model’s full context. If vLLM OOMs on startup, it will report the maximum your hardware can support; use that number, don’t guess a lower one.--tool-call-parser: without it, Kindo agents cannot call tools. Always pair with--enable-auto-tool-choice.
Tool-call parser reference
Section titled “Tool-call parser reference”| Model family | --tool-call-parser | Notes |
|---|---|---|
| NVIDIA Nemotron 3 Super | qwen3_coder | Also set --reasoning-parser nemotron_v3. Requires --trust-remote-code. |
| Mistral / Mixtral | mistral | |
| DeepSeek-V3 / R1 | hermes | Verify against latest vLLM release. |
| Qwen 3 / Qwen 3 Coder | qwen3_coder | |
| DeepHat V2 | qwen3_coder | See DeepHat section below. |
| GPT OSS | openai | Also set --reasoning-parser openai_gptoss. |
--served-model-name and the serving name
Section titled “--served-model-name and the serving name”--served-model-name is the string vLLM uses for the model field in its OpenAI-compatible API. Kindo routes requests via LiteLLM, which resolves a registered serving name to its LiteLLM model ID (e.g. openai/my-name or hosted_vllm/my-name), then strips the prefix and sends "model": "my-name" to your inference server.
The rule: --served-model-name in vLLM must match the suffix of the LiteLLM model ID (after openai/ or hosted_vllm/) and the serving name you register with Kindo.
vllm serve Qwen/Qwen3-Coder-30B-Instruct \ --served-model-name qwen3-coder-30b \ ...Register that model with the serving name qwen3-coder-30b, the LiteLLM model ID openai/qwen3-coder-30b, and the endpoint URL http://vllm-qwen.inference.svc.cluster.local:8000/v1.
Full example: NVIDIA Nemotron 3 Super on 4× H100
Section titled “Full example: NVIDIA Nemotron 3 Super on 4× H100”vllm serve nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 \ --served-model-name nemotron \ --port 8000 \ --max-model-len 1000000 \ --kv-cache-dtype fp8 \ --dtype bfloat16 \ --tensor-parallel-size 4 \ --trust-remote-code \ --tool-call-parser qwen3_coder \ --enable-auto-tool-choice \ --reasoning-parser nemotron_v3DeepHat
Section titled “DeepHat”DeepHat V2 is Kindo’s cybersecurity-focused model built for offensive reasoning, long-context analysis, and secure execution. Treat it like any other self-hosted vLLM model; the specifics below just call out its hardware and flag requirements.
Hardware requirements
Section titled “Hardware requirements”DeepHat V2 serves its full 250k-token context on:
- 1× B200 GPU, or
- 2× H100 GPUs.
A single H100 can serve at most ~90,000 tokens due to KV-cache memory.
Dependencies
Section titled “Dependencies”vllm>=0.14.1(brings all implicit dependencies).- A HuggingFace access token provisioned by Kindo:
export HF_TOKEN=<token>.
Option A: Standalone vLLM
Section titled “Option A: Standalone vLLM”vllm serve DeepHat/DeepHat-V2-ext \ --served-model-name deephat-v2 \ --port 8000 \ --max-model-len 250000 \ --dtype bfloat16 \ --tensor-parallel-size 2 \ --tool-call-parser qwen3_coder \ --enable-prefix-caching \ --enable-chunked-prefill \ --enable-auto-tool-choicevllm serve DeepHat/DeepHat-V2-ext \ --served-model-name deephat-v2 \ --port 8000 \ --max-model-len 250000 \ --dtype bfloat16 \ --tensor-parallel-size 1 \ --tool-call-parser qwen3_coder \ --enable-prefix-caching \ --enable-chunked-prefill \ --enable-auto-tool-choiceOption B: Optional Helm chart
Section titled “Option B: Optional Helm chart”Skip this if you already have LLM serving infrastructure; run vLLM wherever you want and point Kindo at it. If you want a turnkey in-cluster deployment, Kindo ships a Helm chart.
values-deephat.yaml for 2× H100:
model: DeepHat/DeepHat-V2-extservedModelName: deephat-v2maxModelLen: 250000dtype: bfloat16tensorParallelSize: '2'enableChunkedPrefill: 'true'enablePrefixCaching: 'true'enableAutoToolChoice: 'true'toolCallParser: 'qwen3_coder'
resources: limits: nvidia.com/gpu: 2 requests: nvidia.com/gpu: 2
hfToken: '<your_huggingface_token>'vllmApiKey: '<your_api_key>'For 1× B200, set tensorParallelSize: "1" and GPU resources to 1.
Deploy:
helm install deephat-v2 . -f values-deephat.yamlRegistering DeepHat
Section titled “Registering DeepHat”Once the endpoint is serving, register it from the superadmin dashboard’s Models page:
- Serving name:
deephat-v2 - LiteLLM model ID:
openai/deephat-v2(orhosted_vllm/deephat-v2) - Display name: DeepHat V2
- Provider: the OpenAI-compatible option
- Provider display name: the provider name your users should see, for example the team or service that hosts it
- Credential: endpoint URL
http://deephat-v2.inference.svc.cluster.local:8000/v1, plus whatever you set asvllmApiKeyabove - Context window / max output tokens: 250000 / 16384
Verify: Is the endpoint reachable from the cluster?
Section titled “Verify: Is the endpoint reachable from the cluster?”Always verify the endpoint from inside the cluster before registering the model. Check DNS and firewall access from a pod in the cluster.
Step 1: DNS from a pod
Section titled “Step 1: DNS from a pod”kubectl run dns-test --rm -it --restart=Never \ --image=busybox:1.36 \ -- nslookup <your-inference-host>Step 2: HTTP /v1/models from a pod
Section titled “Step 2: HTTP /v1/models from a pod”For a cluster-internal vLLM service:
kubectl run net-test --rm -it --restart=Never \ --image=curlimages/curl:8.10.1 \ -- curl -sS -m 5 http://<your-inference-host>:<port>/v1/modelsFor a cloud provider, use the provider’s host with your key. For example, for Anthropic:
kubectl run net-test --rm -it --restart=Never \ --image=curlimages/curl:8.10.1 \ --env ANTHROPIC_API_KEY=sk-ant-... \ -- sh -c 'curl -sS -m 10 https://api.anthropic.com/v1/models \ -H "x-api-key: $ANTHROPIC_API_KEY" \ -H "anthropic-version: 2023-06-01"'Expected: a JSON response listing the model(s). If this fails, fix connectivity before proceeding; Kindo’s own inference requests will fail the same way once the model is registered.
Step 3 (self-hosted only): tool-call smoke test
Section titled “Step 3 (self-hosted only): tool-call smoke test”kubectl run tool-test --rm -it --restart=Never \ --image=curlimages/curl:8.10.1 \ -- curl -sS -m 30 \ -X POST http://<your-inference-host>:<port>/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "<your-served-model-name>", "messages": [{"role": "user", "content": "What is the weather in SF?"}], "tools": [{ "type": "function", "function": { "name": "get_weather", "description": "Get current weather for a location", "parameters": { "type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"] } } }], "tool_choice": "auto", "max_tokens": 256 }'Expected: the response contains a tool_calls array with "name": "get_weather" and "location": "San Francisco" (or similar). A plain text response means your --tool-call-parser is missing or wrong; agents will not work.
Common pitfalls
Section titled “Common pitfalls”| Pitfall | Impact | Prevention |
|---|---|---|
| Context window too small | Conversations truncated, agents lose context | Use the model’s full advertised window |
Missing --tool-call-parser | Agents cannot call tools | Always pair with --enable-auto-tool-choice |
--served-model-name ≠ registered serving name | LiteLLM returns 404 | Keep both strings identical |
| DNS not reachable from cluster | Model registers but every request fails | Verify with kubectl run ... nslookup first |
| Bedrock model not subscribed or not allowed | 403 or AccessDeniedException from LiteLLM | Subscribe to the model and allow it in IAM before registering it |
| Azure missing the API version | LiteLLM rejects requests | Endpoint URL, API version, and API key are all required |
Gateway ignores Authorization: Bearer | Model registers but every request gets 401/403 | Set extra_headers in Additional litellm_params (JSON, optional) |
| Skipped verification | Kindo blamed for inference bugs | Complete endpoint verification before registering |
Once each endpoint passes the kubectl run checks, you’re ready to install; model registration comes after the install completes.
