Skip to content

Prepare AI Models

Provide inference endpoints from cloud API providers (OpenAI, Anthropic, Gemini, Bedrock, Azure), self-hosted open-weight models with vLLM, or both. After installing Kindo, register each model from the superadmin dashboard’s Models page. This page covers endpoint preparation and the information you need for registration.

After the install:

  1. Open the superadmin dashboard’s Models page and select the Models tab.
  2. Use the Required models panel to add an embedding or transcription model with its preset. To register another model, open the add-model sheet.
  3. Choose the provider, enter its credentials, review the model settings, and save.
  4. Open the Feature flags tab on the Models page to choose which models each feature uses.

Registration makes the model available through the endpoint you supplied. Model Configuration explains model groups and provider fallbacks.

Your job on this page is to:

  1. Decide which providers/endpoints you will use.
  2. Make sure each endpoint is reachable from the cluster.
  3. Collect the facts each model’s registration needs.

What to know about each model before you register it

Section titled “What to know about each model before you register it”

Collect the following information before opening the add-model sheet. Credentials depend on the provider.

FactRequiredWhat it is
Serving nameyesThe unique routing key Kindo and LiteLLM address the model by. Must match --served-model-name for vLLM, and must be the reserved name for a required role.
Display nameyesHuman-readable name shown in the UI model picker.
ProvideryesThe LiteLLM provider, which sets the connection type and the credential fields the form asks for.
Provider display nameyesThe provider your users see for this model: pick an existing one, or choose New provider… and name it. A new name can’t match an existing provider in any letter case.
LiteLLM model IDyesThe provider-prefixed model ID LiteLLM routes with (see the prefix table below).
CredentialvariesAPI key, endpoint URL, API version, AWS keys and region, or a service account, whichever the provider needs. See the playbooks below.
Context windowyes for chatThe model’s full advertised context window. Do not low-ball it. Kindo seeds it from the catalog when the model ID matches; a chat model with neither value nor match is rejected.
Max output tokensrecommendedThe largest response the model is allowed to generate in one call.
Extra headersoptionalEnter {"extra_headers": {...}} in Additional litellm_params (JSON, optional) for headers sent with every inference request, such as an x-api-key gateway credential.
Cost tierrecommendedLOW, MEDIUM, or HIGH, carried in the model’s metadata.
Description / linkoptionalSurfaced in the model picker.
ProviderLiteLLM model ID prefixExample
OpenAIopenai/openai/gpt-4o
Anthropicanthropic/anthropic/claude-sonnet-4-20250514
Google Geminigemini/gemini/gemini-2.5-pro
Azure OpenAIazure/azure/<your-deployment-name>
AWS Bedrockbedrock/bedrock/us.anthropic.claude-sonnet-4-20250514-v1:0
Self-hosted (vLLM)hosted_vllm/ or openai/hosted_vllm/deephat-v2, openai/deephat-v2

For cloud models, prepare an API key, the model ID with its provider prefix, and network egress.

Have ready:

  • LiteLLM model ID: openai/gpt-4o
  • Credential: your OpenAI API key
  • Context window / max output tokens: the values OpenAI publishes for that model (128000 / 16384 for GPT-4o)

Notes:

  • Use the exact model ID OpenAI publishes (e.g. gpt-4o, gpt-4o-mini). Outdated IDs silently fall back or fail.
  • Leave the endpoint URL empty. LiteLLM routes to https://api.openai.com/v1 by default.
  • Organization ID can be supplied by setting OPENAI_ORG_ID in the LiteLLM environment if needed.

Self-hosting gets you full data-residency control and unlimited request volume, but you own the inference server. Kindo recommends vLLM, and this section covers it.

  • GPUs sized for your target context length (table below).
  • NVIDIA drivers + nvidia-container-runtime installed; nvidia-smi shows your GPUs.
  • Model weights pulled (HuggingFace token exported: export HF_TOKEN=...).
  • Official model card read: context window, max output tokens, dtype, and any special flags.
  • Correct --tool-call-parser identified for your model family.
  • DNS/hostname planned; the endpoint will be reachable from the cluster.

Context length is the primary VRAM driver via the KV cache. A 70B model at 4k context needs far less VRAM than the same model at 128k.

Model sizeQuantizationShort (4k)Medium (32k)Full (128k+)
7–8BFP16/BF161× 24GB1× 24GB1× 48GB
7–8BFP81× 24GB1× 24GB1× 24GB
13BFP16/BF161× 48GB1× 80GB1× 80GB
13BFP81× 24GB1× 48GB1× 80GB
30–34BFP16/BF161× 80GB2× 80GB2–4× 80GB
30–34BFP81× 48GB1× 80GB1–2× 80GB
70BFP16/BF162× 80GB4× 80GB4–8× 80GB
70BFP81× 80GB2× 80GB2–4× 80GB
70BAWQ/GPTQ1× 80GB2× 80GB4× 80GB
Terminal window
vllm serve <model-id> \
--served-model-name <name> # must match the serving name you register
--port 8000 \
--max-model-len <context-length> # the model's full context window
--dtype bfloat16 \
--tensor-parallel-size <num-gpus> \
--tool-call-parser <parser> \ # required for agents
--enable-auto-tool-choice \ # required alongside tool-call-parser
--enable-prefix-caching \
--enable-chunked-prefill

Check these flags if the model fails to start or call tools:

  1. --max-model-len: set to the model’s full context. If vLLM OOMs on startup, it will report the maximum your hardware can support; use that number, don’t guess a lower one.
  2. --tool-call-parser: without it, Kindo agents cannot call tools. Always pair with --enable-auto-tool-choice.
Model family--tool-call-parserNotes
NVIDIA Nemotron 3 Superqwen3_coderAlso set --reasoning-parser nemotron_v3. Requires --trust-remote-code.
Mistral / Mixtralmistral
DeepSeek-V3 / R1hermesVerify against latest vLLM release.
Qwen 3 / Qwen 3 Coderqwen3_coder
DeepHat V2qwen3_coderSee DeepHat section below.
GPT OSSopenaiAlso set --reasoning-parser openai_gptoss.

--served-model-name is the string vLLM uses for the model field in its OpenAI-compatible API. Kindo routes requests via LiteLLM, which resolves a registered serving name to its LiteLLM model ID (e.g. openai/my-name or hosted_vllm/my-name), then strips the prefix and sends "model": "my-name" to your inference server.

The rule: --served-model-name in vLLM must match the suffix of the LiteLLM model ID (after openai/ or hosted_vllm/) and the serving name you register with Kindo.

Terminal window
vllm serve Qwen/Qwen3-Coder-30B-Instruct \
--served-model-name qwen3-coder-30b \
...

Register that model with the serving name qwen3-coder-30b, the LiteLLM model ID openai/qwen3-coder-30b, and the endpoint URL http://vllm-qwen.inference.svc.cluster.local:8000/v1.

Full example: NVIDIA Nemotron 3 Super on 4× H100

Section titled “Full example: NVIDIA Nemotron 3 Super on 4× H100”
Terminal window
vllm serve nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 \
--served-model-name nemotron \
--port 8000 \
--max-model-len 1000000 \
--kv-cache-dtype fp8 \
--dtype bfloat16 \
--tensor-parallel-size 4 \
--trust-remote-code \
--tool-call-parser qwen3_coder \
--enable-auto-tool-choice \
--reasoning-parser nemotron_v3

DeepHat V2 is Kindo’s cybersecurity-focused model built for offensive reasoning, long-context analysis, and secure execution. Treat it like any other self-hosted vLLM model; the specifics below just call out its hardware and flag requirements.

DeepHat V2 serves its full 250k-token context on:

  • 1× B200 GPU, or
  • 2× H100 GPUs.

A single H100 can serve at most ~90,000 tokens due to KV-cache memory.

  • vllm>=0.14.1 (brings all implicit dependencies).
  • A HuggingFace access token provisioned by Kindo: export HF_TOKEN=<token>.
Terminal window
vllm serve DeepHat/DeepHat-V2-ext \
--served-model-name deephat-v2 \
--port 8000 \
--max-model-len 250000 \
--dtype bfloat16 \
--tensor-parallel-size 2 \
--tool-call-parser qwen3_coder \
--enable-prefix-caching \
--enable-chunked-prefill \
--enable-auto-tool-choice

Skip this if you already have LLM serving infrastructure; run vLLM wherever you want and point Kindo at it. If you want a turnkey in-cluster deployment, Kindo ships a Helm chart.

values-deephat.yaml for 2× H100:

model: DeepHat/DeepHat-V2-ext
servedModelName: deephat-v2
maxModelLen: 250000
dtype: bfloat16
tensorParallelSize: '2'
enableChunkedPrefill: 'true'
enablePrefixCaching: 'true'
enableAutoToolChoice: 'true'
toolCallParser: 'qwen3_coder'
resources:
limits:
nvidia.com/gpu: 2
requests:
nvidia.com/gpu: 2
hfToken: '<your_huggingface_token>'
vllmApiKey: '<your_api_key>'

For 1× B200, set tensorParallelSize: "1" and GPU resources to 1.

Deploy:

Terminal window
helm install deephat-v2 . -f values-deephat.yaml

Once the endpoint is serving, register it from the superadmin dashboard’s Models page:

  • Serving name: deephat-v2
  • LiteLLM model ID: openai/deephat-v2 (or hosted_vllm/deephat-v2)
  • Display name: DeepHat V2
  • Provider: the OpenAI-compatible option
  • Provider display name: the provider name your users should see, for example the team or service that hosts it
  • Credential: endpoint URL http://deephat-v2.inference.svc.cluster.local:8000/v1, plus whatever you set as vllmApiKey above
  • Context window / max output tokens: 250000 / 16384

Verify: Is the endpoint reachable from the cluster?

Section titled “Verify: Is the endpoint reachable from the cluster?”

Always verify the endpoint from inside the cluster before registering the model. Check DNS and firewall access from a pod in the cluster.

Terminal window
kubectl run dns-test --rm -it --restart=Never \
--image=busybox:1.36 \
-- nslookup <your-inference-host>

For a cluster-internal vLLM service:

Terminal window
kubectl run net-test --rm -it --restart=Never \
--image=curlimages/curl:8.10.1 \
-- curl -sS -m 5 http://<your-inference-host>:<port>/v1/models

For a cloud provider, use the provider’s host with your key. For example, for Anthropic:

Terminal window
kubectl run net-test --rm -it --restart=Never \
--image=curlimages/curl:8.10.1 \
--env ANTHROPIC_API_KEY=sk-ant-... \
-- sh -c 'curl -sS -m 10 https://api.anthropic.com/v1/models \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01"'

Expected: a JSON response listing the model(s). If this fails, fix connectivity before proceeding; Kindo’s own inference requests will fail the same way once the model is registered.

Step 3 (self-hosted only): tool-call smoke test

Section titled “Step 3 (self-hosted only): tool-call smoke test”
Terminal window
kubectl run tool-test --rm -it --restart=Never \
--image=curlimages/curl:8.10.1 \
-- curl -sS -m 30 \
-X POST http://<your-inference-host>:<port>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<your-served-model-name>",
"messages": [{"role": "user", "content": "What is the weather in SF?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a location",
"parameters": {
"type": "object",
"properties": {"location": {"type": "string"}},
"required": ["location"]
}
}
}],
"tool_choice": "auto",
"max_tokens": 256
}'

Expected: the response contains a tool_calls array with "name": "get_weather" and "location": "San Francisco" (or similar). A plain text response means your --tool-call-parser is missing or wrong; agents will not work.

PitfallImpactPrevention
Context window too smallConversations truncated, agents lose contextUse the model’s full advertised window
Missing --tool-call-parserAgents cannot call toolsAlways pair with --enable-auto-tool-choice
--served-model-name ≠ registered serving nameLiteLLM returns 404Keep both strings identical
DNS not reachable from clusterModel registers but every request failsVerify with kubectl run ... nslookup first
Bedrock model not subscribed or not allowed403 or AccessDeniedException from LiteLLMSubscribe to the model and allow it in IAM before registering it
Azure missing the API versionLiteLLM rejects requestsEndpoint URL, API version, and API key are all required
Gateway ignores Authorization: BearerModel registers but every request gets 401/403Set extra_headers in Additional litellm_params (JSON, optional)
Skipped verificationKindo blamed for inference bugsComplete endpoint verification before registering

Once each endpoint passes the kubectl run checks, you’re ready to install; model registration comes after the install completes.