WoolyAI Private Multi-agent Inference Stack for DGX Spark

Private, sovereign & low-cost multi-agent inference stack for agentic AI

WoolyAI Private Inference Stack is a new inference server for DGX Spark clusters, built to serve multi-agentic apps that run as a coordinated workflow across multiple models, modalities, tools, and context windows.

Deploy dedicated

Serve many agent types

Serve multiple model families

Run agents at low cost

How it fits together

Agent workloads

Many agent types, models, and tools

Supervisor

Knowledge

Custom LLMs

Verifier

Wooly Inference Stack

Memory-aware multi-model serving across the cluster — not one static endpoint

Dedicated DGX Spark cluster

Private customer capacity · 4 Sparks

Spark 1

Spark 2

Spark 3

Spark 4

Why DGX Spark?

Accessible hardware, serious private inference.

DGX Spark is smaller and more accessible than large data-center GPU systems, but large enough to support serious private inference workloads when paired with the right serving software.

 

Built for the memory-constrained AI era

Stop deploying a separate expensive GPU stack for every model.

Instead of deploying a separate expensive GPU stack for every model, InferenceStack for DGX Spark Cluster gives each customer a private inference environment that can serve many models of different sizes through one memory-model-residency-aware serving layer.

Benchmark

View performance for multi-model agent workloads

Benchmarks from internal tests on 2× NVIDIA DGX Spark systems — validated single-model throughput and multi-model agent workflow serving across a coordinated InferenceStack environment.

Single-model execution

Large models can run with usable throughput on low-cost unified-memory GPU nodes.

Llama-Benchy • TP=2 • PP=2048 / TG=128 / depth=4096

4,848

tok/s fastest prefill

Gemma 4 26B A4B C1 prefill on the two-node DGX Spark setup.

85.33

tok/s fastest C4 decode

Nemotron 3 Nano Omni C4 aggregate decode across four simultaneous requests.

22.10

tok/s DeepSeek C1 decode

DeepSeek V4 Flash C1 decode on the two-node DGX Spark setup with no speculative decoding.

Multi-model workflow

Switch models safely and keep decoding across an agent pipeline.

Simulated agent workflow • 1 endpoint • 3 models • C4 bursts

93.44

tok/s Nemotron decode

C4 decode throughput in the simulated workflow execution window.

64.65

tok/s Gemma decode

C4 decode throughput after safe activation from DeepSeek to Gemma.

2.34s

fastest activation

Safe-boundary activation from Gemma to Nemotron before the new model batch begins.

16 / 16

requests completed

100% success, zero runtime failures across both ranks in the workflow run.

Internal validation • Metrics from July 2026 two-DGX-Spark Llama-Benchy and multi-model workflow reports.

What it’s not

Our goal is not to replace every H100, GB200, or GB300 deployment. For maximum-throughput frontier-scale inference, those platforms remain the right choice.

Who is this right for

Private enterprise multi-agentic inference: workloads that need privacy, multiple models, predictable cost, and dedicated capacity, but do not need a separate premium data-center GPU stack for every model in the workflow.