Built watts-first.
Off-grid means electricity is the scarcest resource in the system, so it is the first constraint rather than the last. Every other decision — which model, which storage, which interface — is downstream of a power budget. Four rules govern the build, in this order.
Power efficiency
Performance per watt beats raw performance. Always.
Data durability
Silicon is replaceable. Weights and corpora are irreplaceable once connectivity is gone.
Commodity repairability
No proprietary interfaces, no license servers, no part that cannot be substituted.
Graceful degradation
Still useful when the primary node is dead. And recoverable by a second, named person.
Two nodes, three copies, one panel.
| Component | Specification | Why this, specifically |
|---|---|---|
| Primary node | Apple Silicon desktop 128 GB unified memory ~5 W idle |
Unified memory lets a 27–31B model live entirely in RAM with room for a second model beside it. The ~5 W idle figure is the reason this part is not a gaming PC: a machine idling at 30 W burns roughly 700 Wh a day doing nothing, a meaningful fraction of the entire battery bank. This is the top of four rungs — the same archive runs on a $580 Raspberry Pi, more slowly. |
| Secondary node | Raspberry Pi 5, 16 GB or low-power x86 mini PC |
Serves the archives and map tiles over local WiFi so reading never queues behind inference, runs a small model when the primary is off, and becomes the surviving system in a total primary failure. |
| Working store | 8 TB external SSD checksumming filesystem |
ZFS or Btrfs so silent bit-rot is detected rather than quietly inherited by every backup made afterward. |
| Cold copy A | 8 TB, powered down rotated quarterly |
Powered off is the only true protection against a controller failure taking the array with it. |
| Cold copy B | 8 TB, separate location | Because fire, flood and theft do not respect a shelf. |
| Power | 600 W – 1 kW solar 3–5 kWh LiFePO₄ pure sine inverter |
LiFePO₄ for cycle life and thermal safety. A 12 V DC path straight to the node where possible, avoiding inverter conversion losses entirely. |
| Audio | Room array mic + headset powered speaker physical PTT control |
The close-talk headset is required, not optional — recognition accuracy at a quiet desk is no evidence of accuracy where the system actually gets used. |
| Resilience stock | Sealed conductive container duplicate everything |
Spare PSUs and cables for every device, a spare drive enclosure — enclosures fail more often than the drives inside them — and two of each audio component. |
Two details in that drawing are the whole design rather than decoration. The battery feeds the node at 12 volts directly, skipping the inverter — an inverter converts DC to AC so the node's supply can convert it back to DC, and both conversions are paid for in watts that were the scarcest thing on the site to begin with.
And the cold copies are drawn detached because that is their function. A backup that stays plugged in shares a controller, a power rail and a filesystem with the thing it is backing up, which means it also shares the failure. The gap in the line is the protection.
What the cut line means
An air gap is usually drawn as an enclosure, which gets it backwards — a wall suggests something is being kept out. Nothing is trying to get in. The point is that the one edge which would normally leave has been removed, so there is no upstream to fail, throttle, reprice or revoke. The local network is real and busy; it simply has nothing above it.
Five tiers, two independent lineages.
Every model in the build is open-weight and permissively licensed — Apache 2.0 or MIT — so nothing here depends on a company's continued permission. Two model families are carried at the top tier so that asking both the same question and treating disagreement as a warning becomes a cheap, everyday practice.
| Tier | Model | Size | Role | License |
|---|---|---|---|---|
| 1 · Reasoning | Qwen3.8-27B Primary | ~17 GB @ 4-bit ~29 GB @ 8-bit | The workhorse. Dense, natively vision-capable, 262K context. Runs 8-bit on the production node, which directly attacks the confident-wrong-answer failure mode. | Apache 2.0 |
| 1 · Cross-check | Gemma 4 31B | ~20 GB @ 4-bit | A completely separate training lineage kept resident alongside. Strong document parsing, multilingual OCR and handwriting recognition. Ask both; disagreement means go to the archive. | Apache 2.0 |
| 2 · Fallback | Gemma 4 12B | ~8 GB | Lives on the backup node. Chosen over its siblings because the smaller Gemma 4 variants include native speech recognition and speech-to-translated-text — a third, independent spoken-Spanish path on the node most likely to still be running. | Apache 2.0 |
| 2 · Minimum | Phi-4-mini | ~2–4 GB | No GPU required. Runs on nearly anything with a few gigabytes of RAM — the floor of the degradation ladder. | MIT |
| 3 · Engineering | Qwen3-Coder-30B-A3B | ~20 GB | The model that keeps other machines running. Reads a charge controller's register map or a radio's command set straight out of a datasheet and writes the glue — microcontroller sketches, shell scripts, cron logic, log parsers. A different job from general reasoning, and general models are measurably worse at it. Sparse mixture-of-experts, so only a fraction of the 30B activates per token, which is how it fits the power budget. | Apache 2.0 |
| 4 · Retrieval | BGE-M3 Critical | 2.27 GB | The component that makes the whole system work, and the one most builds forget. See below. | MIT |
| 5 · Hearing | Whisper large-v3 & large-v3-turbo | ~10 GB | Speech recognition, both models in two runtime formats each — one tuned for Apple Silicon and CPU, one for CUDA — because format conversion later would require a working system. Small and base variants for the backup node. | MIT |
| 5 · Voice | Kokoro-82M primary · Piper fallback | ~330 MB · ~60 MB | 54 voices across eight language and accent groups, 24 kHz, faster than real time on CPU. Piper is the low-power path on the backup node — and, critically, it embeds its own phoneme engine, so the two fail independently rather than sharing a dependency. | Apache 2.0 · GPL-3.0 |
Full-precision master weights (211 GB) are archived alongside every quantized copy. Quantization is one-way — you cannot recover 8-bit from 4-bit any more than you can recover a RAW file from a JPEG — so every level that might ever be wanted is acquired while connectivity still exists.
How a question finds its answer.
The hard problem is not generating an answer; the models do that well. It is finding the four paragraphs inside a terabyte that can confirm or correct it, and doing it when the user's words and the document's words have nothing in common.
The one component nobody remembers
Ask the Ark "my well water tastes metallic, is it safe?" and keyword search will not save you. The article that answers it is titled Iron and manganese in groundwater and never contains the word "metallic."
The embedding model solves this by turning every passage in the archive into a position in a 1,024-dimensional space, arranged so that passages about similar things land near each other regardless of the words they use. Retrieval becomes geometry: convert the question to coordinates, look at what is nearby.
Here is the part that makes it fragile. The same model is used to build the index and to run every future query. Its coordinates are meaningful only relative to itself — a different embedding model produces a different space, and its numbers are not comparable, like overlaying grid references from two different map projections.
Lose it, and every stored vector in an index built over the whole archive becomes an unreadable list of numbers. There is no substitution fix. The only remedy is re-embedding the entire archive with whatever remains, which requires a working system, days of compute, and stable power — which, years from now and offline, may simply not exist.
So 2.27 GB is acquired, hashed, pinned to an exact commit, and copied to cold storage before the 110 GB of Wikipedia. The failure is invisible until the moment it becomes unfixable.
Why BGE-M3 specifically
100+ languages in one shared space — a Spanish question can retrieve an English passage without translating anything, and version two's seven added languages need no re-index. 8,192-token chunks, so a whole repair procedure stays intact instead of being fragmented across pieces. Dense and sparse and multi-vector retrieval from a single model — one component to restore from cold storage instead of two.
Not a small language model
A different class of thing entirely: 568 million parameters, an encoder rather than a generator. Ask it a question and you do not get a wrong answer — you get coordinates. It has no text-generation machinery at all. If the language model is a writer, this is the librarian who has read everything and knows where each thing belongs.
Nothing that needs an update to keep working.
The offline voice and inference ecosystem churns fast — within one year a leading text-to-speech project was archived, another was declared obsolete by its own maintainers, and a third went silent. An offline node cannot chase a moving ecosystem. Take what works today, verify the whole chain with the network physically disconnected, pin exact versions, and freeze.
llama.cpp, one server per model
Two llama-server processes, one per model, started directly rather than through a manager — a manager is one more layer to vendor, pin and repair offline. The 27B reasoning model holds about 14.8 GB on the graphics card; the cross-check model runs entirely in system memory, about 21 GB, and needs no graphics card at all — so the second opinion survives a machine the first one cannot run on. Ollama is carried as the higher-level alternative and is not what runs. Keeping the layer beneath is a principle applied throughout.
One page, served by one file
The surface is a single Python file importing nothing outside the standard library — no web framework to vendor, no compiled dependency to rebuild years from now. Passages on the left, the answer on the right, a map when the question names a place. Reached from any phone, tablet or laptop on the local network, with no client to install. Open WebUI is vendored as the conversational alternative; it is not what the node presents.
Kiwix server
Serves the ZIM corpora as browsable local websites over WiFi, and it is what every citation opens into. Today it runs beside the surface on the primary node. Moving it to the second node, so archive reads never queue behind inference, is planned rather than built.
Keyword and meaning, fused by rank
Keyword search lives in SQLite's own full-text index and needs no model at all, so it still works when nothing else does. Meaning-based search runs a multilingual embedding model over a local vector store. The two are combined by rank rather than by score — the numbers are not comparable — and the page says so when the two disagree. Both retrieval frameworks and both vector stores are vendored offline because that choice cannot be revisited once connectivity is gone; neither framework is on the path a query actually takes.
Every answer opens its source
Each of 39 million passages carries the artifact, the document and the byte offset it came from, so every result resolves to a link that opens the original — a page in a PDF, an article in the archive. It is a table in SQLite rather than a service: nothing to start, nothing to repair. It was emitted while the index was being built, because adding it afterwards means reading a terabyte again. The answer is a summary; the passage is the source, and any figure worth acting on is read there.
PMTiles, served by one static binary
The whole planet's tile pyramid packed into a single file — no database, no thousands of tiny files, and no tile-server daemon: one static binary publishes it and the browser draws it. A conventional tile server was dropped deliberately, because vendoring it meant also vendoring a JavaScript runtime and a native build tree for four architectures — the one component nobody could rebuild from printed instructions. On phones, CoMaps: an Apache-2.0-throughout fork chosen over an alternative whose map data is proprietary. A local gazetteer of 271,848 place names turns a question that names a place into coordinates, with no model on the path.
A four-stage chain
Capture → recognition → inference → synthesis. The models for every stage are on the drive and verified offline. How the four are glued together is not settled, and the project records it as an open question rather than deciding it in copy. Whatever wins has to be rebuildable from printed instructions, which rules out most of the field.
Second runtimes for key models
The embedding and speech models are also carried as ONNX exports, which run without the deep-learning framework underneath them. Protection not against corruption but against dependency rot — the scenario where the framework itself can no longer be installed years from now.
Every package, vendored
The full Python dependency set is downloaded, then installed once with the network forbidden outright to prove it is genuinely self-sufficient, and the resolved versions pinned to a lock file. A dependency conflict discovered offline, years later, is unfixable.
Offline installer images
Installers for the operating systems this build can carry, plus every driver and runtime, so bare hardware can be brought up from cold storage with no network at any point. macOS is the deliberate exception: an installer vendored today may refuse to install on hardware bought years from now, so the project instead records the versions it is known to work on — and the recovery drill targets the Linux node, the one that can always be rebuilt.
Where the terabyte goes.
| Category | Size |
|---|---|
| Map data — planet vector tiles + terrain | 849 GB |
| Geographic, agricultural, practical, supplemental | 367 GB |
| Core corpora — EN + ES Wikipedia, Gutenberg, Stack Overflow, Stack Exchange | 297 GB |
| Quantized models, tiers 1–4 | 287 GB |
| Vector and keyword indexes | 258 GB |
| Full-precision master weights | 211 GB |
| Literature — Gutenberg EN + ES, Standard Ebooks, Wikisource | 86 GB |
| Mathematics, economics, language sets | 26 GB |
| Offline OS and dependency installers | 26 GB |
| Speech models, both runtime formats | 12 GB |
| Total | 2.42 TB |
Measured on the drive, 2026-09-26 — 30% of an 8 TB store, against a specification that projected 1.15 to 1.3 TB. Maps are most of the difference: a whole-planet terrain model was not in the original budget and is 706 GB of the 849 above. The search index is the other large addition: 52 GB when this chart was first drawn, 258 GB after two more indexing passes took it to 39 million passages. The headroom is still deliberate — seven more languages are a planned addition, and migrating drive capacity later would mean re-cloning both cold copies and repeating every verification.
The two line items amateurs omit
Full-precision weights and the search indexes together are 469 GB, 19% of the total — a larger share than when this was first drawn, because the indexes grew fivefold as two more passes extended them across the archive. Leave them out and you have a system that works today and cannot be rebuilt tomorrow — the index because regenerating it takes days of compute, the master weights because you can never re-derive a higher precision from a lower one.
Every file, hashed
Each artifact is verified against its publisher's checksum on arrival and recorded in a manifest with its source, its exact version, and the commit it came from. A repository name does not identify what is on the disk. A hash does.
Quarterly, annually
Cold drives powered on, checksums verified, drives rotated — every quarter. Corpora refreshed and model choices re-evaluated while connectivity still exists — every year. And a full recovery drill run against the printed pages, annually.
Four floors, and the last one has no computer on it.
Graceful degradation is a design constraint, not an aspiration. Each rung below assumes everything above it is gone — and each one is also a configuration you can simply buy, which is what makes the failure plan and the price list the same document.
- FULLPrimary node, everything on. 27B reasoning at 8-bit with a second family beside it for cross-checking, full archive, hybrid retrieval, voice in both languages, planet maps, served to every device on the local network.
- −1Primary node down → the Raspberry Pi. A 12B model with its own built-in speech recognition, the full archives served over WiFi, map tiles, and the low-power voice engine. Slower and less accurate. Still answering.
- −2Both nodes down → any computer at all. Phi-4-mini runs without a GPU on almost anything, and the archives are browsable directly from a cold copy without an AI in the loop at all — a Kiwix reader and a drive is a complete, functioning offline library.
- −3No working computer → paper. The full recovery procedure is printed, in the recovery operator's own reading language, with the exact commands to rebuild the entire system on bare hardware from cold storage. One copy stored with each set of cold drives. Alongside it: printed logarithm, trigonometric and physical-constant tables, and a physical scientific calculator with spare batteries — the analog fallback for the one thing a language model is structurally worst at.
A recovery drill run by the builder tests the drives. A drill run by a second, named person holding the printed pages tests the documentation — which is the actual point. Any step they cannot complete is a defect in the documentation, not in the operator. "The household" is not an acceptable answer to who that person is: diffuse responsibility means nobody ever runs the drill.
The complete pipeline is tested with the network physically disconnected — not merely with WiFi switched off in software. And both languages are tested separately, because a system can pass every English test while Spanish is silently broken. That specific asymmetry is what the test exists to catch.