7 cards published

Project Hearthmind, research

Research cards

Every technology this project evaluates gets one of these: evidence class stated up front, demonstrated and not-demonstrated kept separate, and a relevance score that is opinion, clearly labelled as opinion.

Research-to-action coverage

Every source gets an honest destination.

Covered sources feed at least one card, essay, or build record. Uncovered sources stay visible so the next lesson or experiment is a deliberate choice rather than an implied claim.

CoveredDGX Spark hardware specificationResearch card · Essay
CoveredDGX Spark systems and pricingResearch card
CoveredDGX Spark review: specs and performanceResearch card · Essay
CoveredNVIDIA DGX Spark power draw under loadResearch card
Needs artifactLocal AI power consumption benchmarksNo published card, essay, BOM, lesson, or story yet.
Needs artifactApple Silicon local LLM cost vs. cloudNo published card, essay, BOM, lesson, or story yet.
Needs artifactSelf-hosted AI inference cost modellingNo published card, essay, BOM, lesson, or story yet.
CoveredZFS RAID-Z and mirror reliability characteristicsResearch card
CoveredHard drive annualized failure ratesResearch card
CoveredOpportunities for district heating in the changing energy landscapeResearch card · Essay
CoveredDGX Spark in-depth review: a new standard for local AIEssay
Needs artifactDGX Spark complete guideNo published card, essay, BOM, lesson, or story yet.
Needs artifactDGX Spark first look: a personal AI supercomputer on your deskNo published card, essay, BOM, lesson, or story yet.
Needs artifactDGX Spark: specs, price, and who should buy itNo published card, essay, BOM, lesson, or story yet.
Needs artifactDGX Spark datasheetNo published card, essay, BOM, lesson, or story yet.
CoveredDGX Spark review: the AI appliance bringing datacentre capabilities to desktopsEssay
CoveredSetup guide for ML training on DGX Spark (GB10, CUDA 13, aarch64)Research card
CoveredDGX Spark ML guideResearch card
CoveredRun vLLM on 1-to-N DGX Spark serversResearch card
Needs artifactDGX Spark specs, FLOPS, benchmarksNo published card, essay, BOM, lesson, or story yet.
Needs artifactDGX Spark research and tests — containers, benchmarks, investigation notesNo published card, essay, BOM, lesson, or story yet.
Needs artifactDGX Spark: 128GB local-AI appliance — specs, honest reviewNo published card, essay, BOM, lesson, or story yet.
Needs artifactOpenZFS documentation FAQ (ECC RAM guidance)No published card, essay, BOM, lesson, or story yet.
CoveredCORE hardware guideHearthvault BOM
CoveredTrueNAS hardware guideResearch card
Needs artifactBuild a ZFS storage server: 3 TrueNAS tiersNo published card, essay, BOM, lesson, or story yet.
CoveredTrueNAS hardware requirements and best server picksResearch card
CoveredBest hardware for TrueNASResearch card
Needs artifactTrueNAS hardware sizing: RAM, HBA, NIC and boot diskNo published card, essay, BOM, lesson, or story yet.
Needs artifactStorage controllers and hardware guideNo published card, essay, BOM, lesson, or story yet.
Needs artifactBuilding a ZFS-based NAS with TrueNAS ScaleNo published card, essay, BOM, lesson, or story yet.
Needs artifactEnergy-efficient home NAS and home lab guideNo published card, essay, BOM, lesson, or story yet.
CoveredDGX Spark Founders Edition product listingEssay
CoveredDCDM1: lessons learned from the world’s most energy-efficient datacentreResearch card
CoveredHigh performance computing datacentre fact sheet (ESIF)Research card
Needs artifactLessons learned from the world’s most energy-efficient datacentreNo published card, essay, BOM, lesson, or story yet.
CoveredLiquid in the rack: liquid cooling your datacentreResearch card
Needs artifactSoaking up energy savings from water projects (ESIF)No published card, essay, BOM, lesson, or story yet.
CoveredUsing thermosyphon hybrid cooling to optimise datacentre energyResearch card
CoveredOptimal environmental and economic performance trade-offs for 5th-gen district heatingResearch card
Needs artifactValuation of novel waste heat sourcesNo published card, essay, BOM, lesson, or story yet.
CoveredData center waste heat recovery techno-economic analysisEssay
CoveredBASOPRA — battery schedule optimizer for residential applicationsResearch card
Needs artifactOptimal energy system scheduling combining mixed-integer programming and deep reinforcement learningNo published card, essay, BOM, lesson, or story yet.
Needs artifactEnergy arbitrage via reinforcement learningNo published card, essay, BOM, lesson, or story yet.
CoveredGridMind — personal energy automation and dashboardResearch card
Needs artifactLessons learned from the world’s most energy-efficient datacentre (fy18osti/71928)No published card, essay, BOM, lesson, or story yet.
Needs artifactAEON DGX Spark — ComfyUI for GB10/Blackwell/sm_121aNo published card, essay, BOM, lesson, or story yet.
CoveredLair — private AI framework (DGX Spark deployment notes)Research card
CoveredTrueNAS Mini R reference platformHearthvault BOM
CoveredAPC UPS UK product rangeHearthvault BOM
CoveredShelly Plug S Gen3 UK price comparisonHearthpower BOM
CoveredAPC Easy UPS BV 1000VA UK listingHearthpower BOM
Vendor spec

2025

NVIDIA DGX Spark as household-scale AI baseline

NVIDIA (hardware specification); independent reviews for real-world power and performance figures

The first commercially available desktop unit with enough unified memory to run serious open-weight models entirely locally, at a household-affordable price and power envelope.

Demonstrated

  • 01A single desktop unit can hold and run open-weight models in the tens-of-billions-of-parameters range locally.
  • 02Independent reviewers measured idle draw around 35–40W and typical inference load around 170W — well under the 240W ceiling.
  • 03List pricing sits in the £3,999–£4,699 band depending on configuration.

Not demonstrated

  • 01Sustained multi-user household workloads over months of real use (no long-run field data yet).
  • 02Direct like-for-like FP4 vs FP16/FP8 throughput comparison against datacentre GPUs for the same task.
  • 03Repairability or long-term maintainability under household (non-datacentre) conditions.
CURRENT: household → rented cloud inference, metered by the token, with data leaving the premises for every query.POSSIBLE: household → owned local inference for most everyday tasks, with cloud used only for capability the Spark genuinely lacks.

How it works — A unified memory architecture lets the CPU and GPU share the same 128GB pool, avoiding the classic consumer-GPU ceiling where VRAM (typically 16–24GB) forces model size or precision compromises.

Why Solystopia cares — This is the first machine in the Hearthmind trajectory: the concrete answer to "can a household afford and power meaningful local AI capability" without appeal to hypothetical future hardware.

Practical constraints

  • 01Capital cost: £3,999–£4,699 up front, no financing model assumed.
  • 02Electricity: continuous idle draw plus load draw, metered against local tariff.
  • 03Knowledge: household needs baseline comfort installing and maintaining a local model runner.
  • 04Maintenance: no long-run reliability data yet from household (rather than datacentre) use.
  • 05Regulation: none currently anticipated for personal-use household compute.

Readiness — TRL 8–9 (commercially available, shipping product) for the hardware; TRL 4–6 for validated household-scale daily-use workflows built on top of it.

What would need to happen next

  • 01Independent long-run (6–12 month) household reliability and maintenance data.
  • 02Published tokens/kWh benchmarks specific to Spark hardware (not extrapolated from Apple Silicon or RTX 4090 figures).
  • 03A published open model-selection guide matched to 128GB unified memory.
  • 04Comparative capital-efficiency data against a self-built alternative (e.g. used enterprise GPUs).
  • 05Documented solar/battery coupling case studies at household scale.

Related work

  • 01Apple Silicon unified-memory local inference (comparable architecture, different ecosystem).
  • 02Consumer GPU (RTX 4090-class) local inference (higher power draw, lower unified memory).

Questions worth investigating

  • 01What is the actual measured tokens/kWh for common open models on Spark hardware, under household (not benchmark-lab) conditions?
  • 02How does capital efficiency (£/petaFLOP) compare to a self-assembled used-hardware alternative?
  • 03What is the realistic multi-year maintenance burden for a household without a systems administrator?
  • 04At what household size or workload does a Spark stop being sufficient, and a neighbourhood-scale node become justified?

Sources, primary first

92/ 100 relevanceEditorial judgement of fit with the Hearthmind subsidiarity framework — not a benchmark score.
Vendor spec

2025

Household-scale redundant storage for models, data, and backups (Hearthvault)

OpenZFS documentation (mirror/RAID-Z mechanics); Backblaze (real-world drive failure-rate data)

A single point of failure in local storage would silently undo the entire sovereignty argument: a dead drive that takes the model library and the household’s own data with it is worse than never owning the hardware.

Demonstrated

  • 01ZFS mirroring and RAID-Z provide well-documented, well-understood protection against single-drive failure at a household-manageable level of complexity.
  • 02Backblaze’s multi-year, large-sample drive data gives a realistic (not vendor-quoted) annualized failure rate to plan around.
  • 03A mirrored NVMe pair plus a separate cold-backup tier is a standard, proven pattern — nothing about it requires enterprise-grade hardware.

Not demonstrated

  • 01No household-specific field data yet on how this setup performs under Hearthmind’s actual read/write patterns (model loading, dataset ingestion, nightly snapshots).
  • 02No measured figure yet for how much local storage a genuinely useful model library plus household document set requires in practice.
  • 03No comparison yet of NVMe mirror vs. simple rsync-based redundancy for a single-operator household, where snapshot rollback matters less than in a multi-user environment.
CURRENT: household data and model weights live on a single unmirrored drive, or partly in a cloud account, with no tested recovery plan.POSSIBLE: household data and model weights survive a single-drive failure automatically, with a separate cold backup surviving a mistake on the live copy.

How it works — Two identical NVMe drives are mirrored so every write lands on both; if one fails, the household loses no data and keeps running on the survivor while it is replaced. A separate NAS holds periodic snapshots so a mistake or ransomware event on the live mirror does not propagate into the only backup.

Why Solystopia cares — Hearthvault is what makes "the household owns its intelligence" true at the data layer, not just the compute layer. Without it, the Spark is a fast processor for data that still, in practice, has no guaranteed local home.

Practical constraints

  • 01Capital cost: roughly £650 for a mirrored NVMe pair plus an 8TB NAS at current retail prices.
  • 02Physical space: a NAS enclosure is a second piece of hardware in the household, not a line in a cloud invoice.
  • 03Maintenance: mirrors need monitoring for silent single-drive failure, or the household is one more failure away from data loss without realising it.
  • 04Knowledge: household needs baseline comfort with a snapshot/backup schedule, even if it is scripted.

Readiness — TRL 9 (mirrored storage and NAS backup are mature, widely deployed patterns) for the mechanism; TRL 3 (early prototype, not yet built) for Hearthvault’s specific household implementation.

What would need to happen next

  • 01Measured figure for actual model-library + household-document storage requirements after 3–6 months of real use.
  • 02A tested, scripted nightly snapshot routine specific to the Hearthmind data layout.
  • 03A documented recovery drill: simulate a drive failure and time how long restoration actually takes.
  • 04A comparison of ZFS mirror vs. simpler rsync-based redundancy once real usage patterns are known.

Related work

  • 01The First Machine (HM-001) — Hearthvault exists specifically to protect the model weights and data that machine produces and consumes.
  • 02Hearthmind Software and Hearthdata Foundry side quests, both of which depend on Hearthvault as their storage substrate.

Questions worth investigating

  • 01Is a ZFS mirror worth its added operational complexity for a single-operator household, versus simpler redundancy?
  • 02How much local storage does a genuinely useful model library plus household document set require, once measured rather than estimated?
  • 03What is the real, tested time-to-recovery after a simulated drive failure — not the theoretical figure?

Sources, primary first

78/ 100 relevanceEditorial judgement of fit with the Hearthmind subsidiarity framework — not a benchmark score.
Vendor spec

2025

A household-scale network fabric for a small AI cluster (Hearthnet)

NVIDIA (DGX Spark networking specification); TrueNAS and homelab documentation (VLAN, HBA, and fabric practice); community multi-Spark clustering guides

Sovereignty at the compute and storage layers is undermined if the machines that hold it cannot talk to each other, or to the household, without routing through someone else’s infrastructure — Hearthnet is the wiring that keeps the boundary real.

Demonstrated

  • 01A single Spark integrates into a household network over standard 10GbE with no special hardware beyond a capable switch.
  • 02Two Sparks can be linked by a direct 200Gbps QSFP cable, and three or more via a switched fabric, to run models too large for one unit alone.
  • 03VLAN segmentation and dedicated firewalling between storage, compute, and household devices is a mature, well-documented homelab pattern, not a novel design.

Not demonstrated

  • 01No household-specific deployment yet of the VLAN/firewall boundary described in the evidence base — Hearthnet exists as a design, not a built network.
  • 02No measured real-world throughput or latency for a two-Spark cluster under Hearthmind’s own workloads, as opposed to vendor or community benchmark figures.
  • 03No tested secure remote-access pattern (e.g. WireGuard) specific to this household’s router and ISP setup.
CURRENT: household devices share a flat, unsegmented network with no dedicated boundary around AI compute or storage traffic.POSSIBLE: AI compute and storage sit on an isolated VLAN, reachable from inside the household and, deliberately, from outside it — never exposed to the open internet by default.

How it works — The household network is split into VLANs — one for Hearthvault and compute, one for general household devices, one for energy telemetry — with a firewall enforcing which traffic can cross between them. The Spark’s QSFP ports, when a second unit is added, bypass the household LAN entirely for machine-to-machine model-serving traffic.

Why Solystopia cares — Hearthnet is what stops the Spark and Hearthvault from being isolated appliances that happen to sit in the same room. It is the layer that lets compute, storage, and (eventually) energy telemetry cohere into one household-owned system rather than three unrelated boxes.

Practical constraints

  • 01Capital cost: a managed switch with VLAN support and, if scaling to a second Spark, a second unit plus QSFP cabling.
  • 02Knowledge: VLAN configuration, firewall rules, and (for remote access) VPN setup require intermediate networking skill.
  • 03Physical: a second Spark for clustering is a substantial additional capital outlay, not a casual upgrade.
  • 04Maintenance: network hardware (switches, routers) is commodity and locally serviceable, unlike the Spark’s SoC.

Readiness — TRL 9 (VLAN segmentation, firewalling, and 10GbE homelab networking are mature, widely deployed patterns) for the mechanism; TRL 2 (design only, not yet built) for Hearthnet’s specific household implementation.

What would need to happen next

  • 01A built and tested VLAN/firewall boundary around Hearthmind’s compute and storage traffic.
  • 02A documented, tested WireGuard (or equivalent) remote-access setup for when the household is away.
  • 03Measured throughput and latency figures for the household’s actual switch and cabling, not vendor specification sheets.
  • 04A decision on whether a second Spark and 200Gbps clustering is ever justified at this household’s scale, versus staying single-node.

Related work

  • 01The First Machine (HM-001) and Hearthvault (HM-002) — Hearthnet is the connective layer between the compute and storage nodes those quests establish.
  • 02Multi-Spark vLLM clustering guides, which validate that the 200Gbps fabric works in practice, even though Hearthmind has not yet built a two-node cluster.

Questions worth investigating

  • 01Is a commercial mesh router acceptable for the household boundary, or does a self-hosted router OS (e.g. OPNsense) better match the sovereignty argument?
  • 02How is remote access handled when the household is away — a permanent WireGuard tunnel, or no remote access at all by default?
  • 03At what point, if any, does a second Spark and clustered fabric become justified for this specific household’s workload?

Sources, primary first

75/ 100 relevanceEditorial judgement of fit with the Hearthmind subsidiarity framework — not a benchmark score.
Peer reviewed

2025

Measuring and coupling household AI compute to real energy metrics (Hearthpower)

NREL and Lawrence Berkeley National Laboratory (data-centre energy and waste-heat metrics); IEA (district heating commentary); open-source residential energy-scheduling projects

The entire "household electricity bill, not a datacentre one" claim is only as strong as the measurement behind it — Hearthpower turns an estimated wattage into a metered fact, and tests whether the household’s own energy system can absorb or even benefit from the load.

Demonstrated

  • 01PUE, ERE, and WUE are rigorously defined, peer-reviewed metrics for assessing energy and water efficiency of computing infrastructure, directly reusable at household scale.
  • 02Open-source household energy schedulers (BASOPRA, GridMind) already optimise battery and PV usage against demand and tariff signals for other loads.
  • 03Smart plugs and platforms like Home Assistant and Prometheus provide the per-device telemetry needed to treat the Spark as one more scheduled load rather than an unmeasured constant draw.

Not demonstrated

  • 01No household has yet applied PUE/ERE-style metrics specifically to an AI compute load, nor published a tokens-per-kWh figure derived from metered (rather than estimated) household data.
  • 02No tested integration yet between an AI workload scheduler and a household’s actual PV/battery system — the pattern is proven for other loads, not for this one.
  • 03No data on how AI inference workloads behave as a "controllable load" in practice, e.g. how well jobs can be deferred to periods of surplus generation without degrading household usefulness.
CURRENT: the Spark’s power draw is estimated from third-party reviewer figures, not measured against this household’s own meter, solar, or battery system.POSSIBLE: every watt the Spark draws is logged against real household generation and tariff data, and heavy workloads are scheduled to prefer surplus renewable generation over grid draw.

How it works — A metered smart plug reports the Spark’s real-time power draw into a time-series store (e.g. Prometheus, via Home Assistant); the same store holds solar generation and battery state of charge. A scheduler can then compare the two and prioritise heavy inference jobs during surplus generation, deferring them when the battery is low or the grid tariff is high.

Why Solystopia cares — Hearthpower is what makes the Spark’s energy footprint a fact rather than an estimate, and what opens the possibility of the household’s AI capability becoming energy-positive rather than purely additive load. It is also the quest that gives every other Hearthmind claim about cost and sustainability a number to stand on.

Practical constraints

  • 01Capital cost: a metered smart plug is inexpensive; a household solar array or battery, if not already present, is a substantial separate investment.
  • 02Attribution: a shared household solar array complicates cleanly attributing "renewable share" to the Spark specifically without a dedicated sub-meter.
  • 03Knowledge: building a scheduler that ties workload timing to energy signals requires basic home-automation and scripting skill.
  • 04Data: a full year of metered data is needed before seasonal patterns (and a genuine annual TCO figure) can be trusted.

Readiness — TRL 9 (PUE/ERE metrics and household energy-scheduling tools are mature and proven) for the underlying techniques; TRL 2 (not yet instrumented) for Hearthpower’s specific application to this household’s Spark.

What would need to happen next

  • 01Installation of a metered smart plug and at least one full month of logged kWh data, replacing the independent-reviewer estimate with a household-measured figure.
  • 02Definition of locally relevant metrics — tokens per kWh, renewable fraction, heat reused — matched to this household’s actual generation mix.
  • 03A first working integration between workload scheduling and existing PV/battery telemetry, even in a minimal form.
  • 04A full twelve-month dataset to support a genuine, seasonally-aware total-cost-of-ownership report.

Related work

  • 01The First Machine (HM-001) — Hearthpower supplies the metered energy data that machine’s TCO report ultimately depends on.
  • 02Hearththermal, which depends on Hearthpower’s telemetry to establish whether waste heat is worth capturing at this household’s scale.

Questions worth investigating

  • 01Does a shared household solar array need a dedicated sub-meter to keep the "renewable fraction" figure honest, or is a proportional estimate acceptable?
  • 02How well, in practice, can AI inference jobs be deferred to periods of surplus generation without degrading day-to-day usefulness?
  • 03What is the real payback period, if any, for adding dedicated battery capacity specifically to smooth the Spark’s load?

Sources, primary first

88/ 100 relevanceEditorial judgement of fit with the Hearthmind subsidiarity framework — not a benchmark score.
Peer reviewed

2025

Managing and reusing the heat a household AI node actually produces (Hearththermal)

NREL and Lawrence Berkeley National Laboratory (data-centre liquid-cooling and heat-reuse case studies); NREL 5th-generation district heating research

A household is not a datacentre with engineered airflow and liquid cooling loops — Hearththermal exists to turn "runs fine on a desk" from an assumption into a measured fact, and to test whether even a small amount of reusable heat is worth capturing.

Demonstrated

  • 01Liquid-cooled data centres can reuse return-loop heat in the 25–40°C range for real building heating and snow-melt applications, at NREL’s facility scale.
  • 02Research-grade district heating studies show this kind of heat reuse can materially offset separate heating energy demand when integrated properly.
  • 03The physics of low-grade heat reuse is well understood and peer-reviewed — the open question for Hearththermal is scale, not feasibility in principle.

Not demonstrated

  • 01No evidence yet, from NREL or otherwise, on whether a single desktop-class device drawing under 250W produces enough usable heat differential to justify any capture equipment at all.
  • 02No household-specific ambient or surface temperature data has been logged for the Spark in its actual placement.
  • 03No cost-benefit case has been built for even a minimal heat-reuse experiment (e.g. ducting exhaust air into a specific room) at this household’s scale.
CURRENT: no temperature monitoring exists around the Spark; any heat it produces is assumed to simply dissipate into the room.POSSIBLE: measured ambient and surface temperatures establish whether the Spark’s heat output is worth redirecting, with a documented placement or ducting recommendation either way.

How it works — Ambient and surface temperature loggers placed near the Spark establish a baseline for how much heat it actually adds to a room under sustained load. If the differential and duration are large enough, a simple passive measure (placement, ducting) or a small supplemental fan could redirect that heat somewhere useful rather than simply venting it.

Why Solystopia cares — Hearththermal is the honesty check on any claim that a household AI node is "efficient" in a fuller sense than electricity cost alone. It also tests, at the smallest possible scale, the same heat-reuse logic that operates at datacentre scale in the peer-reviewed literature this quest draws on.

Practical constraints

  • 01Capital cost: temperature loggers are inexpensive; any active heat-reuse equipment (ducting, fans) is a small but real additional cost.
  • 02Scale: at roughly 170–240W, the Spark’s heat output is orders of magnitude below the datacentre-scale systems this quest’s evidence base is drawn from — findings may not transfer directly.
  • 03Honesty: this quest should be willing to conclude that heat reuse is not worthwhile at this scale, if that is what the data shows.

Readiness — TRL 9 (heat reuse from liquid-cooled data centres is proven at scale) for the underlying physics; TRL 1 (no measurement yet taken) for Hearththermal’s household-scale application.

What would need to happen next

  • 01Baseline ambient and surface temperature logging around the Spark during sustained workloads.
  • 02A documented placement and airflow recommendation based on that baseline data.
  • 03A go/no-go assessment of whether any active heat-reuse experiment is justified at this power scale, made honestly even if the answer is no.

Related work

  • 01Hearthpower (HM-004), whose telemetry establishes the load duration and intensity that determines how much heat is actually available to reuse.
  • 02NREL’s ESIF and 5GDHC research, the datacentre-scale analogue this household-scale experiment is testing the limits of.

Questions worth investigating

  • 01Is any meaningful heat reuse realistic at Spark-level power draw, or is that overreach for a single desktop unit?
  • 02Does room placement alone (rather than any active ducting or fan) already solve any practical heat problem this device creates?
  • 03What is the minimum measured temperature differential that would make even a passive heat-reuse measure worth the effort?

Sources, primary first

82/ 100 relevanceEditorial judgement of fit with the Hearthmind subsidiarity framework — not a benchmark score.
Open-source prototype

2025

A minimal, rebuildable software stack for household-owned AI (Hearthmind Software)

Community GB10/Blackwell setup guides; open-source model-serving, vector-database, and observability projects; self-hosting configuration-as-code practice

Owning the hardware and the data means little if the software layer that turns them into daily capability cannot be understood, audited, or rebuilt by the household itself — Hearthmind Software is the difference between an appliance and a system the household actually controls.

Demonstrated

  • 01Community setup guides (e.g. natolambert/dgx-spark-setup) document working configurations for ML tooling on GB10/CUDA 13/aarch64, despite the platform’s novelty.
  • 02Standard open-source components for a local AI stack — model servers, vector databases, observability tooling — are all available and have documented deployment patterns on Linux and container platforms generally.
  • 03Configuration-as-code practice (Ansible, Terraform, NixOS) is a proven pattern in self-hosting communities for making complex systems rebuildable from bare metal in hours rather than days.

Not demonstrated

  • 01Many existing kernels and libraries, including FlashAttention 3, do not yet fully support Blackwell’s sm_121 compute capability, and require workarounds that may not be stable across updates.
  • 02No Hearthmind-specific infrastructure-as-code has yet been written; the household-facing interface, model library, and full stack integration are still design decisions, not built systems.
  • 03No tested, documented recovery procedure exists yet for rebuilding this specific household’s stack from bare metal.
CURRENT: no local model-serving stack is installed; any AI capability runs on hardware the household does not own and software it cannot inspect.POSSIBLE: a versioned, rebuildable stack the household can audit, modify, and reconstruct from bare metal, running models it has chosen deliberately for a 128GB ceiling.

How it works — A minimal stack layers a container runtime, a model server (such as Ollama or vLLM) serving quantised open-weight models sized to the 128GB unified-memory ceiling, a vector database for retrieval-augmented generation, and an observability layer (Prometheus/Grafana) for monitoring — all defined as versioned configuration so the entire stack can be torn down and rebuilt deterministically.

Why Solystopia cares — This is the layer that makes "the household controls its AI workflows" literally true rather than aspirational. Without it, the Spark is fast hardware running whatever default software ships with it, auditable by no one in the household.

Practical constraints

  • 01Platform maturity: GB10’s Blackwell architecture is new enough that some AI tooling requires nightly builds or manual workarounds rather than stable releases.
  • 02Knowledge: operating CUDA 13, container orchestration, and infrastructure-as-code requires intermediate-to-advanced technical skill.
  • 03Maintenance: keeping pace with a fast-moving toolchain (nightly PyTorch builds, evolving sm_121 support) is an ongoing burden, not a one-time setup cost.

Readiness — TRL 7–8 (the individual open-source components are mature and widely deployed) for the ecosystem; TRL 2 (design and early experimentation) for Hearthmind’s specific integrated stack.

What would need to happen next

  • 01Selection and integration of a minimal, rebuildable stack: Linux base, container runtime, model server, vector database, and observability layer.
  • 02A finalised model library matched to the 128GB unified-memory ceiling, chosen for real household use rather than benchmark scores.
  • 03A decision on whether the household-facing interface should work from any device on Hearthnet or remain a single local web UI.
  • 04A documented, periodically-tested recovery procedure for rebuilding the stack from bare metal.

Related work

  • 01The First Machine (HM-001), which supplies the hardware this stack runs on, and Hearthvault (HM-002), which supplies its storage substrate.
  • 02Hearthdata Foundry (HM-007), which depends on this stack’s vector database and model-serving layer to make household documents queryable.

Questions worth investigating

  • 01Should the household interface be a simple local web UI, or does it need to work from any device on Hearthnet?
  • 02Which model sizes actually deliver a good day-to-day experience inside the 128GB memory ceiling, once tested rather than assumed?
  • 03How much of the stack’s ongoing maintenance burden can realistically be automated away versus requiring continued manual attention?

Sources, primary first

78/ 100 relevanceEditorial judgement of fit with the Hearthmind subsidiarity framework — not a benchmark score.
Vendor spec

2025

Turning household documents into a queryable, provenance-tracked knowledge base (Hearthdata Foundry)

ZFS dataset-management and snapshot documentation; DGX Spark fine-tuning capacity guidance; self-hosting provenance and licensing practice

Generic model capability is not the same as household-specific usefulness — Hearthdata Foundry is the difference between "a chatbot" and a system that actually knows this household’s own manuals, records, and notes, while keeping that material inside the household’s own storage.

Demonstrated

  • 01ZFS snapshot retention schemes (hourly/daily/weekly/monthly) are a proven, low-overhead pattern for versioned dataset and checkpoint management.
  • 02Fine-tuning and adapter-based training on domain-specific data is achievable on single-node or small-cluster hardware like the Spark, and can yield material task-specific improvements.
  • 03Self-hosting and NAS practice provides a mature template for tracking dataset provenance, licensing, and retention when mixing personal, open-source, and third-party data.

Not demonstrated

  • 01No household document set has yet been ingested, embedded, or indexed for retrieval — Hearthdata Foundry exists as a staged plan (prompting → RAG → tool use → memory → evaluation → fine-tuning), not a built pipeline.
  • 02No evaluation methodology has been built yet to actually measure whether fine-tuning on household-specific data improves task performance, as opposed to the general literature suggesting it should.
  • 03No decision has been made on embedding model or vector index technology suited to household (not enterprise) scale.
CURRENT: household-specific knowledge lives in scattered physical or cloud documents, unindexed and unusable by any AI system without manual lookup.POSSIBLE: household documents are ingested, embedded, and retrievable locally, with clear provenance and licensing metadata, entirely inside the household’s own storage.

How it works — Household documents are ingested, embedded with a local embedding model, and indexed for retrieval so the model can answer questions grounded in this household’s actual manuals, records, and notes rather than generic training data. Each dataset carries metadata on source, licence, and allowed use, kept physically inside Hearthvault rather than any external service.

Why Solystopia cares — This is the quest that makes local AI capability specific to the household that owns it, rather than a generic assistant that happens to run locally. It also keeps the most sensitive material — the household’s own documents — inside the sovereignty boundary Hearthvault and Hearthnet establish.

Practical constraints

  • 01Tooling maturity: no embedding model or vector index has yet been selected specifically for household scale, as opposed to enterprise deployments.
  • 02Evaluation: without a documented methodology, claimed improvements from fine-tuning cannot yet be measured, only assumed from the general literature.
  • 03Provenance discipline: keeping a clean separation between public research data, private household data, and licensed third-party corpora requires ongoing metadata hygiene, not a one-time setup step.

Readiness — TRL 6–7 (the underlying techniques — RAG, fine-tuning, dataset versioning — are proven in general AI practice) for the mechanism; TRL 1 (no pipeline built yet) for Hearthdata Foundry’s specific household implementation.

What would need to happen next

  • 01An initial household document inventory with provenance and licensing metadata attached to each source.
  • 02Selection of a local embedding model and a vector index sized appropriately for household (not enterprise) scale.
  • 03An evaluation design capable of actually measuring whether fine-tuning or RAG improves task performance on household-specific queries.
  • 04The first household document set fully ingested and queryable end to end.

Related work

  • 01Hearthvault (HM-002), whose storage substrate this quest’s datasets and indices depend on.
  • 02Hearthmind Software (HM-006), whose model-serving and vector-database layer this quest’s retrieval pipeline runs on top of.

Questions worth investigating

  • 01Is a full vector database overkill at household scale, or does simplicity actually argue for one anyway rather than a simpler search index?
  • 02What evaluation methodology would actually prove that fine-tuning on household-specific data improves real task performance, rather than just plausibly should?
  • 03How should the household draw the line between what it fine-tunes on versus what it only retrieves via RAG?

Sources, primary first

76/ 100 relevanceEditorial judgement of fit with the Hearthmind subsidiarity framework — not a benchmark score.