Route the work
Configure fast, mid, large, vision, CPU and fallback model roles. Inspect routing decisions, escalation and downgrade activity as requests move through the available capacity.
Performance
A model runtime is one part of the system. Kaldryn adds workload routing, shared inference capacity and the measurements that help administrators understand latency, throughput and hardware pressure.
Explore this part of the platformConfigure fast, mid, large, vision, CPU and fallback model roles. Inspect routing decisions, escalation and downgrade activity as requests move through the available capacity.
Follow time to first token, tokens per second and context-fill percentiles. Compare benchmark runs with live resource readings and inspect request activity across users and model tiers.
See GPU memory, CPU, storage and pending work. Manage model placement and GPU nodes through the cluster console, with secure node enrollment, connectivity checks and engine configuration.
Interactive software preview · illustrative data
Measure inference performance on your deployment.
TTFT p50
180 ms
Tokens / sec
420
Active requests
32
Time to first token benchmarks
Based on the Kaldryn Platform admin console. Available features depend on license tier, role, hardware and configuration.
Throughput and concurrency vary by model, quantisation, context length, hardware and workload. The comparison below illustrates a conservative 9× scenario; confirm performance with your own workload.
See the potential of shared inference with a conservative 9× comparison for 32 simultaneous chats on GB10 hardware.
the per-user throughput in this illustration
Continuous batching · shared GPU
Reference baseline for this illustration
At 6.3 tokens per second per user versus a 0.7 reference baseline, the illustrated throughput is 9×. Across 32 chats, that is 201.6 versus 22.4 tokens per second.
Chat, document search with RAG, agents and model fine-tuning come together in one platform.
Manage SSO, role-based access, DLP and audit logs alongside GPU health, signed updates and backups in one console.
Illustrative comparison, not a new measured benchmark: 0.7 × 9 = 6.3 tokens/s per user. Reference workload: GB10, Qwen2.5-7B, 8k context, 150-token responses, warm engine, 32 concurrent chats. Actual throughput depends on the model, hardware and configuration.
Discuss your deployment
Discuss your deployment