Ark of Noahledge Offline · Solar · AI Knowledge node · Open source forever — democratizing knowledge where it's needed most
Architecture

Built watts-first.

Off-grid means electricity is the scarcest resource in the system, so it is the first constraint rather than the last. Every other decision — which model, which storage, which interface — is downstream of a power budget. Four rules govern the build, in this order.

Constraint 1

Power efficiency

Performance per watt beats raw performance. Always.

Constraint 2

Data durability

Silicon is replaceable. Weights and corpora are irreplaceable once connectivity is gone.

Constraint 3

Commodity repairability

No proprietary interfaces, no license servers, no part that cannot be substituted.

Constraint 4

Graceful degradation

Still useful when the primary node is dead. And recoverable by a second, named person.

Hardware

Two nodes, three copies, one panel.

ComponentSpecificationWhy this, specifically
Primary node Apple Silicon desktop
128 GB unified memory
~5 W idle
Unified memory lets a 27–31B model live entirely in RAM with room for a second model beside it. The ~5 W idle figure is the reason this part is not a gaming PC: a machine idling at 30 W burns roughly 700 Wh a day doing nothing, a meaningful fraction of the entire battery bank. This is the top of four rungs — the same archive runs on a $580 Raspberry Pi, more slowly.
Secondary node Raspberry Pi 5, 16 GB
or low-power x86 mini PC
Serves the archives and map tiles over local WiFi so reading never queues behind inference, runs a small model when the primary is off, and becomes the surviving system in a total primary failure.
Working store 8 TB external SSD
checksumming filesystem
ZFS or Btrfs so silent bit-rot is detected rather than quietly inherited by every backup made afterward.
Cold copy A 8 TB, powered down
rotated quarterly
Powered off is the only true protection against a controller failure taking the array with it.
Cold copy B 8 TB, separate location Because fire, flood and theft do not respect a shelf.
Power 600 W – 1 kW solar
3–5 kWh LiFePO₄
pure sine inverter
LiFePO₄ for cycle life and thermal safety. A 12 V DC path straight to the node where possible, avoiding inverter conversion losses entirely.
Audio Room array mic + headset
powered speaker
physical PTT control
The close-talk headset is required, not optional — recognition accuracy at a quiet desk is no evidence of accuracy where the system actually gets used.
Resilience stock Sealed conductive container
duplicate everything
Spare PSUs and cables for every device, a spare drive enclosure — enclosures fail more often than the drives inside them — and two of each audio component.
SOLAR PANEL 600 W – 1 kW CHARGE CONTROLLER MPPT LiFePO₄ BATTERY 3–5 kWh DC 12 V 12 V DIRECT — NO INVERTER LOSS AUDIO mic · speaker push-to-talk button PRIMARY NODE 27B + cross-check family ~5 W idle · 90 W thinking SECONDARY NODE Raspberry Pi 5 fallback model + voice USB ARCHIVE + MAP SERVING READS NEVER QUEUE BEHIND INFERENCE USB-C WORKING SSD 8 TB · checksumming filesystem LOCAL WIFI no uplink PHONES & TABLETS read the archive, view maps AIR GAP SATELLITE · VENDOR PAYMENT · INTERNET CLONED QUARTERLY, THEN DISCONNECTED COLD COPY A powered down · rotated quarterly COLD COPY B separate building
Amber is the power path and nothing else. One line leaves the system, toward the infrastructure a connected build would depend on — and it is cut. Nothing else crosses the boundary.

Two details in that drawing are the whole design rather than decoration. The battery feeds the node at 12 volts directly, skipping the inverter — an inverter converts DC to AC so the node's supply can convert it back to DC, and both conversions are paid for in watts that were the scarcest thing on the site to begin with.

And the cold copies are drawn detached because that is their function. A backup that stays plugged in shares a controller, a power rail and a filesystem with the thing it is backing up, which means it also shares the failure. The gap in the line is the protection.

What the cut line means

An air gap is usually drawn as an enclosure, which gets it backwards — a wall suggests something is being kept out. Nothing is trying to get in. The point is that the one edge which would normally leave has been removed, so there is no upstream to fail, throttle, reprice or revoke. The local network is real and busy; it simply has nothing above it.

Models

Five tiers, two independent lineages.

Every model in the build is open-weight and permissively licensed — Apache 2.0 or MIT — so nothing here depends on a company's continued permission. Two model families are carried at the top tier so that asking both the same question and treating disagreement as a warning becomes a cheap, everyday practice.

TierModelSizeRoleLicense
1 · ReasoningQwen3.8-27B Primary~17 GB @ 4-bit
~29 GB @ 8-bit
The workhorse. Dense, natively vision-capable, 262K context. Runs 8-bit on the production node, which directly attacks the confident-wrong-answer failure mode.Apache 2.0
1 · Cross-checkGemma 4 31B~20 GB @ 4-bitA completely separate training lineage kept resident alongside. Strong document parsing, multilingual OCR and handwriting recognition. Ask both; disagreement means go to the archive.Apache 2.0
2 · FallbackGemma 4 12B~8 GBLives on the backup node. Chosen over its siblings because the smaller Gemma 4 variants include native speech recognition and speech-to-translated-text — a third, independent spoken-Spanish path on the node most likely to still be running.Apache 2.0
2 · MinimumPhi-4-mini~2–4 GBNo GPU required. Runs on nearly anything with a few gigabytes of RAM — the floor of the degradation ladder.MIT
3 · EngineeringQwen3-Coder-30B-A3B~20 GBThe model that keeps other machines running. Reads a charge controller's register map or a radio's command set straight out of a datasheet and writes the glue — microcontroller sketches, shell scripts, cron logic, log parsers. A different job from general reasoning, and general models are measurably worse at it. Sparse mixture-of-experts, so only a fraction of the 30B activates per token, which is how it fits the power budget.Apache 2.0
4 · RetrievalBGE-M3 Critical2.27 GBThe component that makes the whole system work, and the one most builds forget. See below.MIT
5 · HearingWhisper large-v3 & large-v3-turbo~10 GBSpeech recognition, both models in two runtime formats each — one tuned for Apple Silicon and CPU, one for CUDA — because format conversion later would require a working system. Small and base variants for the backup node.MIT
5 · VoiceKokoro-82M primary · Piper fallback~330 MB · ~60 MB54 voices across eight language and accent groups, 24 kHz, faster than real time on CPU. Piper is the low-power path on the backup node — and, critically, it embeds its own phoneme engine, so the two fail independently rather than sharing a dependency.Apache 2.0 · GPL-3.0

Full-precision master weights (211 GB) are archived alongside every quantized copy. Quantization is one-way — you cannot recover 8-bit from 4-bit any more than you can recover a RAW file from a JPEG — so every level that might ever be wanted is acquired while connectivity still exists.

The retrieval layer

How a question finds its answer.

The hard problem is not generating an answer; the models do that well. It is finding the four paragraphs inside a terabyte that can confirm or correct it, and doing it when the user's words and the document's words have nothing in common.

QUESTION spoken or typed BGE-M3 text → 1024 numbers + keyword weights HYBRID SEARCH meaning · nearest neighbours exact · part nos, doses 39M passages indexed THE ARCHIVES Wikipedia · medical · repair textbooks · agriculture · forums not indexed: literature, maps indexed once, read forever QWEN / GEMMA answers, then checks against the passages ANSWER + its source AIR GAP No radio. No cable. No upstream. Nothing in this diagram leaves the box.
The retrieval loop — every stage runs on the device

The one component nobody remembers

Ask the Ark "my well water tastes metallic, is it safe?" and keyword search will not save you. The article that answers it is titled Iron and manganese in groundwater and never contains the word "metallic."

The embedding model solves this by turning every passage in the archive into a position in a 1,024-dimensional space, arranged so that passages about similar things land near each other regardless of the words they use. Retrieval becomes geometry: convert the question to coordinates, look at what is nearby.

Here is the part that makes it fragile. The same model is used to build the index and to run every future query. Its coordinates are meaningful only relative to itself — a different embedding model produces a different space, and its numbers are not comparable, like overlaying grid references from two different map projections.

Lose it, and every stored vector in an index built over the whole archive becomes an unreadable list of numbers. There is no substitution fix. The only remedy is re-embedding the entire archive with whatever remains, which requires a working system, days of compute, and stable power — which, years from now and offline, may simply not exist.

So 2.27 GB is acquired, hashed, pinned to an exact commit, and copied to cold storage before the 110 GB of Wikipedia. The failure is invisible until the moment it becomes unfixable.

Why BGE-M3 specifically

100+ languages in one shared space — a Spanish question can retrieve an English passage without translating anything, and version two's seven added languages need no re-index. 8,192-token chunks, so a whole repair procedure stays intact instead of being fragmented across pieces. Dense and sparse and multi-vector retrieval from a single model — one component to restore from cold storage instead of two.

Not a small language model

A different class of thing entirely: 568 million parameters, an encoder rather than a generator. Ask it a question and you do not get a wrong answer — you get coordinates. It has no text-generation machinery at all. If the language model is a writer, this is the librarian who has read everything and knows where each thing belongs.

Software

Nothing that needs an update to keep working.

The offline voice and inference ecosystem churns fast — within one year a leading text-to-speech project was archived, another was declared obsolete by its own maintainers, and a third went silent. An offline node cannot chase a moving ecosystem. Take what works today, verify the whole chain with the network physically disconnected, pin exact versions, and freeze.

Inference

llama.cpp, one server per model

Two llama-server processes, one per model, started directly rather than through a manager — a manager is one more layer to vendor, pin and repair offline. The 27B reasoning model holds about 14.8 GB on the graphics card; the cross-check model runs entirely in system memory, about 21 GB, and needs no graphics card at all — so the second opinion survives a machine the first one cannot run on. Ollama is carried as the higher-level alternative and is not what runs. Keeping the layer beneath is a principle applied throughout.

Interface

One page, served by one file

The surface is a single Python file importing nothing outside the standard library — no web framework to vendor, no compiled dependency to rebuild years from now. Passages on the left, the answer on the right, a map when the question names a place. Reached from any phone, tablet or laptop on the local network, with no client to install. Open WebUI is vendored as the conversational alternative; it is not what the node presents.

Archives

Kiwix server

Serves the ZIM corpora as browsable local websites over WiFi, and it is what every citation opens into. Today it runs beside the surface on the primary node. Moving it to the second node, so archive reads never queue behind inference, is planned rather than built.

Retrieval

Keyword and meaning, fused by rank

Keyword search lives in SQLite's own full-text index and needs no model at all, so it still works when nothing else does. Meaning-based search runs a multilingual embedding model over a local vector store. The two are combined by rank rather than by score — the numbers are not comparable — and the page says so when the two disagree. Both retrieval frameworks and both vector stores are vendored offline because that choice cannot be revisited once connectivity is gone; neither framework is on the path a query actually takes.

Provenance

Every answer opens its source

Each of 39 million passages carries the artifact, the document and the byte offset it came from, so every result resolves to a link that opens the original — a page in a PDF, an article in the archive. It is a table in SQLite rather than a service: nothing to start, nothing to repair. It was emitted while the index was being built, because adding it afterwards means reading a terabyte again. The answer is a summary; the passage is the source, and any figure worth acting on is read there.

Maps

PMTiles, served by one static binary

The whole planet's tile pyramid packed into a single file — no database, no thousands of tiny files, and no tile-server daemon: one static binary publishes it and the browser draws it. A conventional tile server was dropped deliberately, because vendoring it meant also vendoring a JavaScript runtime and a native build tree for four architectures — the one component nobody could rebuild from printed instructions. On phones, CoMaps: an Apache-2.0-throughout fork chosen over an alternative whose map data is proprietary. A local gazetteer of 271,848 place names turns a question that names a place into coordinates, with no model on the path.

Voice

A four-stage chain

Capture → recognition → inference → synthesis. The models for every stage are on the drive and verified offline. How the four are glued together is not settled, and the project records it as an open question rather than deciding it in copy. Whatever wins has to be rebuildable from printed instructions, which rules out most of the field.

Runtime insurance

Second runtimes for key models

The embedding and speech models are also carried as ONNX exports, which run without the deep-learning framework underneath them. Protection not against corruption but against dependency rot — the scenario where the framework itself can no longer be installed years from now.

Dependencies

Every package, vendored

The full Python dependency set is downloaded, then installed once with the network forbidden outright to prove it is genuinely self-sufficient, and the resolved versions pinned to a lock file. A dependency conflict discovered offline, years later, is unfixable.

Operating systems

Offline installer images

Installers for the operating systems this build can carry, plus every driver and runtime, so bare hardware can be brought up from cold storage with no network at any point. macOS is the deliberate exception: an installer vendored today may refuse to install on hardware bought years from now, so the project instead records the versions it is known to work on — and the recovery drill targets the Linux node, the one that can always be rebuilt.

Storage

Where the terabyte goes.

2.42 TB IN USE 5.58 TB UNUSED HEADROOM CORE CORPORA 297 GB MAP DATA 849 GB FULL-PRECISION 211 GB INDEXES 258 GB PRACTICAL 367 GB QUANTIZED 287 GB THE TWO LINE ITEMS AMATEURS OMIT · WEIGHTS + INDEXES · 469 GB
Hue carries one distinction only — the two items most builds leave out — and a diagonal hatch carries it too, so it survives greyscale. The other six are separated by the gaps between them, not by colour; the table below names every one. Segment widths are proportional to bytes measured on the drive rather than to the specification.
CategorySize
Map data — planet vector tiles + terrain849 GB
Geographic, agricultural, practical, supplemental367 GB
Core corpora — EN + ES Wikipedia, Gutenberg, Stack Overflow, Stack Exchange297 GB
Quantized models, tiers 1–4287 GB
Vector and keyword indexes258 GB
Full-precision master weights211 GB
Literature — Gutenberg EN + ES, Standard Ebooks, Wikisource86 GB
Mathematics, economics, language sets26 GB
Offline OS and dependency installers26 GB
Speech models, both runtime formats12 GB
Total2.42 TB

Measured on the drive, 2026-09-26 — 30% of an 8 TB store, against a specification that projected 1.15 to 1.3 TB. Maps are most of the difference: a whole-planet terrain model was not in the original budget and is 706 GB of the 849 above. The search index is the other large addition: 52 GB when this chart was first drawn, 258 GB after two more indexing passes took it to 39 million passages. The headroom is still deliberate — seven more languages are a planned addition, and migrating drive capacity later would mean re-cloning both cold copies and repeating every verification.

The two line items amateurs omit

Full-precision weights and the search indexes together are 469 GB, 19% of the total — a larger share than when this was first drawn, because the indexes grew fivefold as two more passes extended them across the archive. Leave them out and you have a system that works today and cannot be rebuilt tomorrow — the index because regenerating it takes days of compute, the master weights because you can never re-derive a higher precision from a lower one.

Every file, hashed

Each artifact is verified against its publisher's checksum on arrival and recorded in a manifest with its source, its exact version, and the commit it came from. A repository name does not identify what is on the disk. A hash does.

Quarterly, annually

Cold drives powered on, checksums verified, drives rotated — every quarter. Corpora refreshed and model choices re-evaluated while connectivity still exists — every year. And a full recovery drill run against the printed pages, annually.

Failure

Four floors, and the last one has no computer on it.

Graceful degradation is a design constraint, not an aspiration. Each rung below assumes everything above it is gone — and each one is also a configuration you can simply buy, which is what makes the failure plan and the price list the same document.

AS BUILT FULL ~5 / 90 W 27B at 8-bit second family for cross-check voice · planet maps · LAN LOST nothing IF THE PRIMARY NODE FAILS −1 10–15 W Gemma 4 12B on the Pi full archive over WiFi map tiles · low-power voice LOST tier-1 reasoning, cross-check answers in minutes IF BOTH NODES FAIL −2 whatever it draws Phi-4-mini, no GPU needed or Kiwix alone, no AI at all LOST voice, maps, retrieval conversational speed IF NO COMPUTER RUNS −3 0 W printed recovery procedure printed log & trig tables a scientific calculator LOST everything that computes THE ARCHIVE · 1.62 TB ON AT LEAST TWO DRIVES · UNCHANGED AT EVERY STEP
Vertical position is remaining capability. The band beneath is the library, at the same height under all four steps — and every step still connects down to it.
  • FULLPrimary node, everything on. 27B reasoning at 8-bit with a second family beside it for cross-checking, full archive, hybrid retrieval, voice in both languages, planet maps, served to every device on the local network.
  • −1Primary node down → the Raspberry Pi. A 12B model with its own built-in speech recognition, the full archives served over WiFi, map tiles, and the low-power voice engine. Slower and less accurate. Still answering.
  • −2Both nodes down → any computer at all. Phi-4-mini runs without a GPU on almost anything, and the archives are browsable directly from a cold copy without an AI in the loop at all — a Kiwix reader and a drive is a complete, functioning offline library.
  • −3No working computer → paper. The full recovery procedure is printed, in the recovery operator's own reading language, with the exact commands to rebuild the entire system on bare hardware from cold storage. One copy stored with each set of cold drives. Alongside it: printed logarithm, trigonometric and physical-constant tables, and a physical scientific calculator with spare batteries — the analog fallback for the one thing a language model is structurally worst at.
The test that actually matters

A recovery drill run by the builder tests the drives. A drill run by a second, named person holding the printed pages tests the documentation — which is the actual point. Any step they cannot complete is a defect in the documentation, not in the operator. "The household" is not an acceptable answer to who that person is: diffuse responsibility means nobody ever runs the drill.

Verified disconnected

The complete pipeline is tested with the network physically disconnected — not merely with WiFi switched off in software. And both languages are tested separately, because a system can pass every English test while Spanish is silently broken. That specific asymmetry is what the test exists to catch.