AllodialFurnace Internal

Internal · for calls, not for customers

Where each machine stops.

The short version, from the report:

Three desktop-class options sit below the Furnace cabinet, and they are not interchangeable: one or two Sparks, one RTX PRO 6000 in a workstation, and a Furnace 2kW. The Spark pair is the capacity play; the single PRO 6000 is the speed play; and at roughly the same money the PRO 6000 is the better machine for every municipal task in this report except one.

One or two DGX Sparks running a 20–35B open-weight model will do the document work of a mid-size city. That is not a comfortable sentence, but it follows directly from our own Round 9 arithmetic and from published Spark benchmarks, and a county CIO with a Spark on their desk will reach it in an afternoon.

Two paired DGX Sparks (64 GB each)

The Spark pair is the capacity play.

What it is
It is a desktop computer with a very good GPU in it.
What it costs
Two Sparks of either kind are under $10,000 and that is the figure to carry.
What it does

One or two DGX Sparks running a 20–35B open-weight model will do the document work of a mid-size city.

12,000–20,000 document tasks a month per Spark. Two Sparks: 25,000–40,000.

Where it stops

It stops at four things: interactive concurrency past about a dozen simultaneous staff, a hardware guarantee that the vendor cannot read the records, isolation between departments on one box, and anything that has to be run lights-out as a managed appliance.

One RTX PRO 6000 workstation

The single PRO 6000 is the speed play.

What it is
The PRO 6000 Blackwell is a 96 GB GDDR7 card at ~1.8 TB/s — six and a half times the Spark's bandwidth — with 24,064 CUDA cores, ECC memory, and a 600 W TDP in the Workstation Edition.
What it costs
In a workstation with a 1,000 W supply, a current Xeon or Threadripper, 128 GB of system RAM and NVMe, the OEM price is $13,000–$16,000.
What it does

For every task in this report, the PRO 6000 finishes it three to six times sooner, serves two to three times the seats, and handles the long-document cases — the 100-page permit packet, the 400-page records file — that make a Spark pair crawl, because prefill is the bottleneck in Round 9's profile and prefill is where the bandwidth gap shows most.

Where it stops

The security ceiling is identical on both — software hardening, TPM, secure boot, no enclave — so neither changes the "administer without reading" argument, and both are the same distance from a Furnace on that axis.

The pair is quieter, draws less, and runs from two ordinary outlets; a 600 W card wants a 15 A circuit to itself and will warm the room it is in.

Ours

Furnace cabinet · 7kW and up

The 7kW and 70kW tiers are the real hardware business.

What it is
The cabinet is for the point where several departments share one machine, a vendor has to maintain it without being able to read it, or the queue outgrows a desktop.
What it costs
Quoted per configuration
What it does

What does not run on a Spark is the part that depends on the silicon: hardware-enforced "we can administer it and cannot read it," per-department attested isolation, and a management plane we can operate without a person at the desk.

Where it stops

Furnace's capital line has to be justified on what the Spark cannot do, never on per-task cost. Internal reasoning — do not say verbatim.

Those three are the real reason to buy the box, and they only bind for specific customers: anyone with CJI or multi-agency data on one machine, anyone who wants a vendor to run it, and anyone past about 40,000 document tasks or 15 concurrent seats.

Which side of the line is this customer on?

Answer what you know. The first rule that matches decides.

Every threshold below is the report's estimate, not a measurement.

Concurrent interactive seats

Threshold from: Comfortable concurrent seats at 15–20 tok/s each
Report estimate, not a measurement

Document tasks per month

Threshold from: Document tasks per month, derated 3–5×
Report estimate, not a measurement

Longest routine document

Threshold from: Where the arithmetic turns: long documents.
Report estimate, not a measurement

Criminal-justice information, or several agencies on one machine

Threshold from: Hardware isolation
Report estimate, not a measurement

Who administers it

Threshold from: Serviceability
Report estimate, not a measurement

Any of the Spark-pair conditions

Threshold from: The Spark pair is the right buy only when the customer needs more than 96 GB in one address space, or needs two physically separate boxes, or cannot provide a 600 W circuit.
Report estimate, not a measurement

Task by task

Every cell in the report's words. Hover or tap a cell for what decides it.

Figures here are third-party estimates; sources are under The numbers.

County task1–2 Sparks, 20–35B MoEOne RTX PRO 6000 workstationFurnace 2kW / 7kW
Permit and zoning screening, residential and ADU subset, citation per point Yes. Single-document, bounded, checkable. 1,000–2,000/month is a rounding error. Yes, including drawings and 100-page packets at usable latency. Yes
What decides itNothing at city scale; drawings and 100-page packets push latency to minutes
Benefits eligibility preparation (SNAP), cited fields → assembled determination Yes at county volume. ~87 applications per 1,000 residents a year → a 300k county is ~2,200/month. Yes, with headroom for a multi-county COG. Yes
What decides itIsolation from other departments' data; the data is not CJI, so no TEE requirement
Fee, tax and utility billing calculation over the adopted rate schedule Yes — and it does not need the GPU. This is the rule engine, not the model. Yes — no GPU needed. Yes
What decides itNothing. Runs on a laptop.
Records request triage and redaction, exemption mapped per element Yes for current volume (4–55 requests per 1,000 residents a year). No for a backlog: a 20,000-page digitization is weeks of prefill. Yes, and a backlog is days, not weeks. Yes
What decides itPage volume
311 triage and routing, ~700 requests per 1,000 residents a year, short prompts Yes. Short inputs, trivial prefill. A 500k city's 29,000/month fits one Spark. Yes. Yes
What decides itNothing
Interactive assistant, dozens of staff Yes, for one department. ~10–15 concurrent at readable speed. Yes, 30–40 concurrent. Yes, hundreds
What decides itSeat count
Interactive assistant, thousands of seats No. No. 7kW and up
What decides itSeat count, bandwidth
Several departments on one machine, each unable to observe the other No. Software separation only; no MIG, no per-tenant attested key, shared memory. No. Same software-only separation. Yes, five-axis separation
What decides itHardware isolation
Anything touching criminal-justice information (CJIS personnel rule: vendor cannot be able to read it) No, if the vendor administers it. Spark has no enclave; "administer without reading" cannot be enforced by hardware. A county-administered Spark with no vendor access is a different question. No, same reason: no enclave. Yes, by architecture
What decides itTEE and attestation
Police narrative drafting, risk scoring, HR screening Excluded on both by policy Excluded Excluded
What decides it—
Running it lights-out as an appliance someone else maintains No BMC. Updates and restarts need a person at the desk, or an SSH path that is itself the security hole. Mostly no. Some OEM towers have a BMC; it is not an appliance. Yes, management plane
What decides itServiceability
Dense 70B+ models at interactive speed No (4–6 tok/s) Yes (~32 tok/s) Yes
What decides itBandwidth
Heat recovery worth metering No. 240 W is a space heater on low. No. 600–900 W; a space heater on high. 2kW: marginal. 7kW+: real in a year-round hot-water building.
What decides itScale

No task matches that filter.

The numbers

Third-party figures, each row with its source.

Third-party benchmarks on varying software stacks. Allodial has measured none of these.

Spark single-stream benchmarks
From the report section: What a Spark is, in the numbers that matter
ModelPrefill tok/sDecode tok/s, one streamReadable?Source
Llama 3.1 8B q4~7,600~38YesOllama, DGX Spark performance
gpt-oss 20B MXFP4~3,200~58YesOllama, DGX Spark performance
Gemma 3 27B q4~830~11MarginalOllama, DGX Spark performance
Qwen3 32B q4 (dense)~700~9MarginalOllama, DGX Spark performance
gpt-oss 120B MXFP4 (MoE, ~5B active)~1,200–1,700~39–41YesOllama, DGX Spark performanceIntuitionLabs review, Feb 2026
Qwen3.6 35B-A3B FP8 (MoE)—~28–30YesCorti, Qwen3.6-35B on DGX Spark
Llama 3.1 70B q4 (dense)~1,900~4–6NoOllama, DGX Spark performanceExxact, comparing inference engines on DGX Spark

Context lengths behind these figures: the Ollama rows (Llama 3.1, gpt-oss, Gemma 3, Qwen3 32B) used a prompt of roughly 2,000–3,000 tokens of prose and a 500-token output, caching off; the Qwen3.6 35B figure is Corti's own serving-stack measurement, context length not published; the Exxact engine comparison does not publish its prompt length. None of these is the Round 9 profile of 30,000 tokens in, and at that input every decode figure here falls.

Published single-stream figures (Ollama, llama.cpp, vLLM, late 2025 to mid 2026, after NVIDIA's TensorRT-LLM updates): Under concurrency with vLLM, aggregate decode on a mid-size MoE reaches roughly 150–300 tok/s across ten or more simultaneous requests; Ollama tuned for four parallel streams lands near 120. — Exxact, comparing inference engines on DGX Spark, Corti, Qwen3.6-35B on DGX Spark

Two paired Sparks against one PRO workstation
From the report section: One RTX PRO 6000 in a workstation versus two paired Sparks
MeasureTwo paired DGX Sparks (64 GB each)One RTX PRO 6000 workstationSource
Usable model memory128 GB pooled96 GBIntuitionLabs review, Feb 2026Hardware Corner
Memory bandwidth273 GB/s each; ~1.7× effective when paired~1,790 GB/sIntuitionLabs review, Feb 2026Hardware Corner
Dense 70B Q4, one stream4–8 tok/s~32 tok/sOllama, DGX Spark performanceHardware Corner
27B dense Q4, one stream~11 (one box); ~19 paired~61 tok/sOllama, DGX Spark performanceKubesimplifyHardware Corner
30B-class MoE, one stream~30~200 tok/sCorti, Qwen3.6-35B on DGX SparkHardware Corner
gpt-oss 120B MXFP4, one stream~40~210 tok/sOllama, DGX Spark performanceHardware Corner
Prefill, 27B class~800–1,500 tok/s~3,300 tok/sOllama, DGX Spark performanceHardware Corner
Prefill, 120B MoE~1,200–1,700~4,600 tok/sOllama, DGX Spark performanceIntuitionLabs review, Feb 2026Hardware Corner
Batched aggregate, 27B class~150–300 tok/s~670 tok/s at 32 streamsExxact, comparing inference engines on DGX SparkKubesimplify
Comfortable concurrent seats at 15–20 tok/s each10–1530–40Report arithmetic from the rows above
Round 9 task, machine time~30–35 s~8–10 sReport arithmetic
Document tasks per month, derated 3–5×25,000–40,00060,000–100,000Report arithmetic
Largest model that is pleasant to use~35B MoE; 120B MoE workable120B MoE at full speed; dense 70B comfortably; 235B MoE in Q4 fitsCorti, Qwen3.6-35B on DGX SparkExxact, comparing inference engines on DGX SparkHardware Corner
Power at the wall~480 W700–900 W under load (Max-Q: ~450 W)IntuitionLabs review, Feb 2026Hardware Corner
ECCNoYesHardware Corner
Confidential computing / attested enclaveNoNoinference-exchange issue #65NVIDIA Secure AI with Blackwell and Hopper
Out-of-band managementNoWorkstation BMC on some OEM SKUs; not standardIsland Mountain
Form factorTwo books and a cableA tower under a deskIntuitionLabs review, Feb 2026Island Mountain
ProcurementUnder every IT small-purchase thresholdUsually under it; some counties' $10k–$15k lines will trip a quote requirementIsland Mountain

Context lengths behind these figures: every PRO 6000 tok/s figure is Hardware Corner's at 4K context (the same card falls to about half at 128K–256K: dense 70B from 32 to 16.6 tok/s, gpt-oss 120B from 210 to 99.8); the Kubesimplify head-to-head and its 32-stream aggregate used a 512-token prompt and 128-token output; Spark figures are as stated under the first table. The seat counts and the per-month task figures are this report's arithmetic from those sources, not measurements.

What the report says about these figures

All Spark and PRO 6000 figures above are published third-party or NVIDIA benchmarks on varying stacks; we have measured nothing ourselves.

The Round 9 token profile is extrapolated from coding-agent studies, not from a measured municipal workload.

The derate (3–5×) is Round 9's, derived from enterprise GPU utilization patterns, not from a municipal workload.

Worked example: one document task in Spark time

From the report section: The work, converted to Spark time

Inputs

A 10-page submission, a three-call agentic workflow, about 30,000 tokens in and 2,400 out. Prefill is 77% of the compute.

  • Input30,000 tokens in
  • Output2,400 out
  • Prefill sharePrefill is 77% of the compute
  • Derate3–5× derate

Steps

  1. Prefill: 30,000 tokens at ~1,500 tok/s ≈ 20 seconds.
  2. Decode: 2,400 tokens at ~200 tok/s aggregate (batched) ≈ 12 seconds of machine time, or ~80 seconds wall-clock for one person waiting on one stream.
  3. Call it 30–35 seconds of machine time per task at full batching.
  4. An 18-hour day, 30 days: ~1.9 million seconds a month, so a saturated, perfectly scheduled Spark clears about 55,000–60,000 of these tasks a month.
  5. Apply Round 9's 3–5× derate for burst, business-hours concentration and p95 sizing: 12,000–20,000 document tasks a month per Spark. Two Sparks: 25,000–40,000.

What it means

Round 9's composite 500,000-resident city generates ~76,000 transactions a month across every function, of which 20,000–40,000 are document-shaped. Two Sparks cover it. A 150,000-resident county is inside one. This is the fact the report exists to state plainly.

Q&A drill

Concede first. Then lead with the fact.

  1. Why not two Sparks?

    Say this on the call

    For one department's documents with nobody but your own staff touching the machine, a workstation running our software is the right answer and we will sell you that; the cabinet is for several departments on one machine, a vendor maintaining it without being able to read it, or a queue that outgrows a desktop.

    Concede first

    If your workload is one department's documents and nobody but your own staff ever touches the machine, a workstation running our software is the right answer and we will sell you that — and we will tell you which workstation. Internal reasoning — do not say verbatim.

    Lead with

    Those four are exactly what Furnace sells, and none of them is the heat.

    The reasoning behind it

    One or two DGX Sparks running a 20–35B open-weight model will do the document work of a mid-size city. Where a Spark stops is not "it cannot run the model." It stops at four things: interactive concurrency past about a dozen simultaneous staff, a hardware guarantee that the vendor cannot read the records, isolation between departments on one box, and anything that has to be run lights-out as a managed appliance.

    Do not say or claim

    • Do not say: “any superlative attached to an unmeasured quantity”claims register

    Report section: The short version

  2. Why not one RTX 6000?

    Say this on the call

    For most county workloads that is the right desktop-class machine, and if that is where you are we will tell you so at the scoping call.

    Concede first

    For every task in this report, the PRO 6000 finishes it three to six times sooner, serves two to three times the seats, and handles the long-document cases — the 100-page permit packet, the 400-page records file — that make a Spark pair crawl, because prefill is the bottleneck in Round 9's profile and prefill is where the bandwidth gap shows most.

    Lead with

    The cabinet is for the point where several departments share one machine, a vendor has to maintain it without being able to read it, or the queue outgrows a desktop — and we will tell you which side of that line you are on at the scoping call. Internal reasoning — do not say verbatim.

    The reasoning behind it

    If a county is going to buy a desktop-class machine, the honest advice is one RTX PRO 6000 workstation, not two Sparks: same money, three to six times the throughput, long documents handled, ECC, and a 70B-class model on the table. The security ceiling is identical on both — software hardening, TPM, secure boot, no enclave — so neither changes the "administer without reading" argument, and both are the same distance from a Furnace on that axis.

    Figures here are third-party estimates; sources are under The numbers.

    Do not say or claim

    • Do not say: “fast”claims register

    Report section: One RTX PRO 6000 in a workstation versus two paired Sparks

  3. Isn't the heat just my electricity bill back?

    Say this on the call

    It is electricity you would have spent on the computing anyway, returned as hot water, and whether that is worth anything depends on your building — we run that with your numbers, not ours.

    Concede first

    Round 8 said it was a door-opener, not a reason to buy. Internal reasoning — do not say verbatim.

    Lead with

    Keep the heat argument for 7kW and up. Internal reasoning — do not say verbatim.

    The reasoning behind it Internal reasoning — do not say verbatim.

    At the 2kW tier the Spark comparison makes it a liability: a buyer who runs the calculator on a 2.3 kW machine against a 480 W pair of Sparks sees that the "heat you can use" is the electricity we asked them to spend. 2kW: marginal. 7kW+: real in a year-round hot-water building.

    Do not say or claim

    • Do not say: “pays for itself”claims register
    • Do not claim: “A measured heat-recovery case study”claims register

    Report section: What the Spark takes away from the Furnace pitch

  4. What does the cabinet do that a workstation can't?

    Say this on the call

    Three things: it lets us administer the machine without being able to read your records, it keeps departments on one machine from seeing each other, and it runs as an appliance without a person at the desk.

    Concede first

    The rule engine is the product, and it is CPU code.

    Lead with

    Those four are exactly what Furnace sells, and none of them is the heat.

    The reasoning behind it

    What does not run on a Spark is the part that depends on the silicon: hardware-enforced "we can administer it and cannot read it," per-department attested isolation, and a management plane we can operate without a person at the desk. Those three are the real reason to buy the box, and they only bind for specific customers: anyone with CJI or multi-agency data on one machine, anyone who wants a vendor to run it, and anyone past about 40,000 document tasks or 15 concurrent seats.

    Do not say or claim

    • Do not say: “CJIS-compliant”claims register
    • Do not say: “CJIS-compatible”claims register

    Report section: What survives, and it is the software

  5. Can you run it without being able to read our records — on a Spark?

    Say this on the call

    No; a Spark has no hardware enclave, so that guarantee cannot be enforced on it, which is the reason the cabinet exists.

    Concede first

    A county-administered Spark with no vendor access is a different question.

    Lead with

    No, if the vendor administers it.

    The reasoning behind it

    There is no NVIDIA confidential-computing mode on GB10 — the TEE with encrypted GPU memory and a signed attestation of what is running exists on H100, H200 and the B200 class, not on the consumer Blackwell SoC. Builders targeting Spark describe their security ceiling as "hardened" (secure boot, TPM 2.0 measurements, software isolation), not "confidential." Spark has no enclave; "administer without reading" cannot be enforced by hardware.

    Do not say or claim

    • Do not say: “CJIS-compliant”claims register
    • Do not say: “CJIS-compatible”claims register

    Report section: What a Spark is, in the numbers that matter

  6. What about the new 64 GB Spark, or the RTX Spark workstation?

    Say this on the call

    Same chip, same memory bandwidth, so everything we have said about the Spark applies to both; the 64 GB holds exactly the model class a county would run.

    Concede first

    The comparator for Furnace's entry tier is no longer a developer box; it is a workstation.

    Lead with

    The 64 GB is $4,999 — more than the Founders Edition, not less, because NVIDIA is pricing the chip, not the memory.

    The reasoning behind it

    There is no second-generation Spark. What ships on 23 October 2026 is the DGX Spark 64 GB: the same GB10 Grace Blackwell superchip, the same GPU, half the unified memory, from $4,999, US only at launch, with the same ConnectX-7 ports so two units pool to 128 GB. So the hardware this report compares against is unchanged by the launch. The RTX Spark makes the "Spark under a desk" scenario a Windows PC purchase.

    Do not say or claim

    • Do not say: “any superlative attached to an unmeasured quantity”claims register

    Report section: Which Spark, and what is actually new

  7. How many of my staff can use it at once?

    Say this on the call

    Roughly ten to fifteen on a Spark pair and thirty to forty on a PRO 6000 workstation at a readable speed; hundreds is a cabinet.

    Concede first

    Yes, for one department. ~10–15 concurrent at readable speed.

    Lead with

    Seat count.

    The reasoning behind it

    And interactive use: 150–300 tok/s aggregate is about 10–15 staff getting a readable 15–20 tok/s each. A 7,000-seat city is not a Spark customer for assistants; a 40-person planning department is.

    On the PRO workstation: Yes, 30–40 concurrent.

    On the cabinet: Yes, hundreds.

    Figures here are third-party estimates; sources are under The numbers.

    Do not say or claim

    • Do not claim: “Any tokens-per-second, tokens-per-kWh, queries-per-day or throughput figure”claims register

    Report section: The work, converted to Spark time

  8. What happens with a 400-page records file?

    Say this on the call

    On a Spark it is a slow one-at-a-time job; on a PRO 6000 it is routine; the difference is how quickly the machine reads the input, which is most of the work.

    Concede first

    For every task in this report, the PRO 6000 finishes it three to six times sooner, serves two to three times the seats, and handles the long-document cases — the 100-page permit packet, the 400-page records file — that make a Spark pair crawl, because prefill is the bottleneck in Round 9's profile and prefill is where the bandwidth gap shows most.

    Lead with

    A 100-page file is ~75,000 input tokens; at 1,500 tok/s prefill that is 50 seconds before the first output token, per call, per pass, and the KV cache for it crowds out concurrency.

    The reasoning behind it

    Where the arithmetic turns: long documents. A permit packet with drawings, a records request against a 400-page file, a benefits case with a year of statements — the Spark does them, one at a time, slowly.

    Figures here are third-party estimates; sources are under The numbers.

    Do not say or claim

    • Do not say: “fast”claims register

    Report section: The work, converted to Spark time

  9. Is the 2 kW cabinet just a Spark in a nicer box?

    Say this on the call

    The small cabinet earns its place on isolation and serviceability, not on raw speed, and if your workload does not need those we will say so.

    Concede first

    The security ceiling is identical on both — software hardening, TPM, secure boot, no enclave — so neither changes the "administer without reading" argument, and both are the same distance from a Furnace on that axis.

    Lead with

    Either the 2kW tier is H200/B200-class (expensive, but the claim holds) or it is the Spark tier under another name. Internal reasoning — do not say verbatim.

    The reasoning behind it Internal reasoning — do not say verbatim.

    Against two Sparks at under $10,000, or one PRO 6000 workstation at $13,000–$16,000 that already does everything a county asks, it has to justify itself on isolation, TEE and serviceability alone, and the TEE argument does not hold for RTX PRO silicon any more than for GB10 — confidential computing is a data-center-GPU feature.

    Do not say or claim

    • Do not claim: “No price, range, or "starting from" figure.”claims register

    Report section: What this means for the product line

  10. Doesn't any machine in our own building keep the records in the building?

    Say this on the call

    Yes — data residency is what any on-premises machine gives you; what the cabinet adds is that we can maintain it without being able to read what is on it.

    Concede first

    A Spark under a desk keeps them in the room. Internal reasoning — do not say verbatim.

    Lead with

    Data residency is table stakes for any on-prem box; it differentiates Furnace from the cloud, not from a Spark. Internal reasoning — do not say verbatim.

    The reasoning behind it Internal reasoning — do not say verbatim.

    "Only Furnace keeps the records in the building." False. A Spark under a desk keeps them in the room. Data residency is table stakes for any on-prem box; it differentiates Furnace from the cloud, not from a Spark.

    Do not say or claim

    • Do not say: “Only Furnace keeps the records in the building.”report

    Report section: What the Spark takes away from the Furnace pitch

  11. Isn't the cabinet cheaper per document than a Spark?

    Say this on the call

    No, and we do not sell it on per-document cost; we sell it on what a desktop machine cannot do.

    Concede first

    Against a Spark it is worse: two Sparks are under $10,000, draw 480 W, and need a power strip. Internal reasoning — do not say verbatim.

    Lead with

    Those four are exactly what Furnace sells, and none of them is the heat.

    The reasoning behind it Internal reasoning — do not say verbatim.

    Round 9 already killed this against the cloud. Furnace's capital line has to be justified on what the Spark cannot do, never on per-task cost.

    Do not say or claim

    • Do not say: “pays for itself”claims register
    • Do not claim: “Any efficiency multiple, including the "90%+ cost saving" in internal vision documents”claims register

    Report section: What the Spark takes away from the Furnace pitch

  12. Can a Spark handle our SNAP benefits volume?

    Say this on the call

    At county volume, yes; what decides it is whether other departments' data shares the machine and who administers it.

    Concede first

    Yes at county volume.

    Lead with

    ~87 applications per 1,000 residents a year → a 300k county is ~2,200/month.

    The reasoning behind it

    Yes at county volume. ~87 applications per 1,000 residents a year → a 300k county is ~2,200/month. Isolation from other departments' data; the data is not CJI, so no TEE requirement.

    Do not say or claim

    • Do not say: “guarantees correct outcomes”claims register

    Report section: Task by task

  13. Can several departments share one Spark?

    Say this on the call

    Only with software separation; if one department must be unable to observe another, that is hardware isolation, and that is the cabinet.

    Concede first

    Yes, for one department.

    Lead with

    Hardware isolation.

    The reasoning behind it

    Software separation only; no MIG, no per-tenant attested key, shared memory.

    On the PRO workstation: No. Same software-only separation.

    On the cabinet: Yes, five-axis separation.

    Do not say or claim

    • Do not say: “any superlative attached to an unmeasured quantity”claims register

    Report section: Task by task

  14. We have a records backlog to digitize. Can a Spark do it?

    Say this on the call

    Current volume yes; a backlog of tens of thousands of pages is weeks on a Spark and days on a PRO 6000 workstation.

    Concede first

    Yes for current volume (4–55 requests per 1,000 residents a year).

    Lead with

    Page volume.

    The reasoning behind it

    No for a backlog: a 20,000-page digitization is weeks of prefill.

    On the PRO workstation: Yes, and a backlog is days, not weeks.

    Do not say or claim

    • Do not say: “fast”claims register

    Report section: Task by task

  15. Can we run a 70B model on a Spark?

    Say this on the call

    It loads but it is not usable at interactive speed; a PRO 6000 runs it comfortably.

    Concede first

    Mixture-of-experts models with a few billion active parameters are the sweet spot: a 120B MoE outruns a dense 70B on this hardware by a factor of eight, because what moves through the 273 GB/s pipe per token is the active parameters, not the total.

    Lead with

    The one number to carry is 273 GB/s.

    The reasoning behind it

    So a Spark can hold a large model and run a small one. A 70B dense model, unusable on Sparks, is an ordinary working model on the PRO 6000.

    On a Spark: No (4–6 tok/s)

    On the PRO workstation: Yes (~32 tok/s)

    Figures here are third-party estimates; sources are under The numbers.

    Do not say or claim

    • Do not say: “the best open weight models”claims register

    Report section: What a Spark is, in the numbers that matter

  16. Can it run without anyone at the desk?

    Say this on the call

    A Spark cannot; a workstation mostly cannot; the cabinet has a management plane built for exactly that.

    Concede first

    The pair is two independent machines, so one can be imaged, patched or tested while the other serves.

    Lead with

    It is a desktop computer with a very good GPU in it.

    The reasoning behind it

    No BMC. Updates and restarts need a person at the desk, or an SSH path that is itself the security hole.

    On the PRO workstation: Mostly no. Some OEM towers have a BMC; it is not an appliance.

    On the cabinet: Yes, management plane.

    Do not say or claim

    • Do not say: “insurable”claims register

    Report section: Task by task

  17. Why would we need a plant room for any of this?

    Say this on the call

    You do not for a workstation; the plant room is for the cabinet, where the heat goes into your hot-water loop instead of the room.

    Concede first

    A Spark is a listed consumer product; it plugs into a wall. Internal reasoning — do not say verbatim.

    Lead with

    The 7kW and 70kW tiers are the real hardware business, and the sales page's customer is the right one: counties and COGs with aggregated volume, several departments, a sheriff's office, and a building with year-round hot water demand. Internal reasoning — do not say verbatim.

    The reasoning behind it Internal reasoning — do not say verbatim.

    Everything Round 7 cataloged as facility work — the thing we said was our uncopyable asset — is also the thing a Spark makes unnecessary. A Spark is a listed consumer product; it plugs into a wall. The "one PO, one responsible party" transaction is only a differentiator when there is facility work to be responsible for.

    Do not say or claim

    • Do not say: “data center / data centre as a description of our product”claims register

    Report section: What the Spark takes away from the Furnace pitch

  18. What does the cabinet cost?

    Say this on the call

    It is quoted per configuration — accelerator count, memory, heat-rejection arrangement and installation all vary by site — and the quotation itemizes every line and marks the estimated ones; we will not give you a number on a call that we would not put in writing.

    Concede first

    Furnace's capital line has to be justified on what the Spark cannot do, never on per-task cost. Internal reasoning — do not say verbatim.

    Do not say or claim

    • Do not claim: “No price, range, or "starting from" figure.”claims register
    • Do not claim: “No preferential terms, discount, favorable rate or "better deal" for a pilot cohort, an early customer, a reference site or a design partner, until someone has decided what that actually is and priced it.”claims register

    Report section: What to say on a call

  19. Have you benchmarked any of this yourselves?

    Say this on the call

    Not yet; every figure we quote is a published third-party or NVIDIA benchmark with its source, and the first pilot replaces them with measurements.

    Concede first

    All Spark and PRO 6000 figures above are published third-party or NVIDIA benchmarks on varying stacks; we have measured nothing ourselves. Internal reasoning — do not say verbatim.

    Lead with

    One week with a Spark and our actual harness on the Round 9 task profile would replace every estimate in this document with a measurement, and it is the cheapest benchmark available to us (A3 in the document plan). Internal reasoning — do not say verbatim.

    The reasoning behind it Internal reasoning — do not say verbatim.

    The derate (3–5×) is Round 9's, derived from enterprise GPU utilization patterns, not from a municipal workload. The pilot fixes this.

    Do not say or claim

    • Do not claim: “Any tokens-per-second, tokens-per-kWh, queries-per-day or throughput figure”claims register

    Report section: Open numbers

  20. Can we start with a pilot without a capital vote?

    Say this on the call

    Yes — a workstation running our software sits under a small-purchase threshold, needs no facility work, and produces the operating data that makes the next conversation concrete.

    Concede first

    A Spark sits under every IT small-purchase threshold in the country. Internal reasoning — do not say verbatim.

    Lead with

    That is a weakness for our nine commitments (there is nothing to commit to) and a strength for a pilot (there is nothing to approve). Internal reasoning — do not say verbatim.

    The reasoning behind it Internal reasoning — do not say verbatim.

    It is also the pilot vehicle: zero facility work, zero UL, zero plant room, under the small-purchase threshold, deliverable in a week. Round 6 said a pilot is an insurance prerequisite and Round 9 said the first pilot's token-consumption data is the most valuable dataset we could own. A Spark pilot produces both for under $10,000 of hardware.

    Do not say or claim

    • Do not claim: “No preferential terms, discount, favorable rate or "better deal" for a pilot cohort, an early customer, a reference site or a design partner, until someone has decided what that actually is and priced it.”claims register
    • Do not say: “insurable”claims register

    Report section: What this means for the product line

What we have not measured

The report's open numbers, verbatim.

Ticking a box here does not change the page's numbers; the report does.

Ticks are saved in this browser only.