Two million years of knowledge. One box. No internet.
The Ark of Noahledge is a single self-contained computer that runs an open-source AI alongside more than one and a half terabytes of humanity's written knowledge — medicine, engineering, mathematics, agriculture, law, literature, the whole planet's maps — and keeps answering questions when the network is gone, the grid is down, and nobody is coming to help.
Pay for the hardware once. It answers for the rest of its life.
A mind that reasons.
A library that remembers.
Every AI you have used is a rented mind at the end of a wire. Cut the wire and it becomes nothing. It also makes things up — fluent, confident, and wrong — because a language model recalls unreliably by construction.
Every offline encyclopedia has the opposite problem. It is trustworthy and completely inert. It cannot answer a question. It can only be searched by someone who already knows what the answer is called.
The Ark carries both, and wires them together. Ask it something and it answers the way a knowledgeable person would, fully and in your language, then checks that answer against the archive: the passages that bear on your question come back beside it, cited, so you can open the source and read it yourself, and whatever the archive does not back is marked as the model's own.
That is the entire design. The model brings the knowledge and the reasoning. The library keeps it honest.
A system that can only tell you how to purify water and set a bone is a survival tool. A system that also carries the literature is a civilization backup.
Why grounding matters
Semantic search alone fails on the things that hurt people when they are wrong — dosages, part numbers, voltages, chemical names. The Ark runs hybrid retrieval: meaning-based search to find the right passage even when it uses none of your words, plus exact keyword matching so M8×1.25 stays M8×1.25.
One question,
all the way through.
A real session on the node, captured from its own screen. Nothing here is a mockup, and nothing left the machine to make it.
-
Screenshot of the Ark surface. Left column: source passages, the first from Hesperian's A Community Guide to Environmental Health. Right column: the answer, giving drops and teaspoons of bleach per litre and per gallon, each amount followed by bracketed citation numbers, under an amber Verify before acting banner.Ask it like you would ask a person.
The surface answers like someone who knows the subject: how much bleach, for how much water, what to do when the water is cloudy. Every figure carries a citation, and the passages it came from sit in the column beside it. Because the topic is a chemical, an amber banner asks you to read the source before you act.
-
Page 116 of Nuclear War Survival Skills, opened from Kiwix in the browser's PDF viewer, under the heading Disinfecting Water: add 1 scant teaspoonful of household bleach to each 10 gallons of clear water, or 2 drops to each quart.Every citation opens the page it came from.
Citation [4] opens Nuclear War Survival Skills in Kiwix, on the same machine, at the paragraph the answer used: one scant teaspoon of 5.25% bleach to each ten gallons of clear water, two drops to a quart. The figure is checked against the original, not against a summary of it.
-
The same page, scrolled. Left: passages from a LibreTexts course for water operators and from Nuclear War Survival Skills. Right: a block headed Answer from Qwen, not from the archive, where the model adds its own guidance on bleach strength and a smell test.Then it adds what it knows, and says so.
After the cited answer, the model offers what it knows beyond the passages: testing by smell, adjusting for stronger bleach. The block is labelled as its own. No passage supports it and nothing on the node can check it, so it is never mistaken for the archive.
-
Further down the same page: Gemma's answer from the same passages, listing the same drop counts with citation numbers, and a green note reading The two model families agree.A second model family reads the same passages.
Gemma, from a different developer than Qwen, answers from the same sources on its own, and the node compares the two cited answers. Here they agree. When they do not, the card says so and sends you to the passages: a signal, not a verdict.
-
The Kiwix library page, 58 books in its English view: ArchWiki, Wikipedia, True Prepper, WikiVet, Libre Pathology, Wiktionary, Appropedia, MDWiki and more, each a card with its logo.Behind the answers, the whole library.
Kiwix serves every book on the node, 67 of them, from Wikipedia to WikEM. Forty-four feed the answers; the other twenty-three are reached here, through Kiwix's own search. All of them open and read page by page. This view is filtered to English.
-
The offline map over Pereira, Colombia: roads over hillshaded hills, and a panel with hillshade and 3D terrain switched on and a clicked coordinate.The planet, offline.
A 137.5 GB basemap of the whole world and 706 GB of terrain, drawn by the node with hillshade and 3D relief. Click a point and it hands you the command that looks up its climate zone. No tile server, no internet.
-
The llama.cpp chat page with the question Explain how to disinfect drinking water with household bleach, and Qwen's numbered answer, with no citations.Or ask the model directly.
The model has its own chat page too, llama.cpp's, reachable from the node. It is quicker, and it is labelled for what it is: no archive and no citations. The same bleach question, answered from training alone.
-
The Ark surface at phone width: the search box, then the cited answer with its citation numbers.From any phone in the room.
Everything above opens in a phone's browser over the node's own WiFi. No app to install, no account, nothing to update.
Captured on the node on 28 September 2026 and not retouched: where a model slipped, the slip is still in the picture.
Passages quoted from the archive: © Hesperian Health Guides; LibreTexts; Nuclear War Survival Skills, public domain; Chemistry Stack Exchange, CC BY-SA 4.0. Map data © OpenStreetMap contributors, ODbL, via Protomaps; terrain from Mapterhorn. Chat page: llama.cpp, MIT.
Not a survival kit.
The whole inheritance.
Most offline knowledge devices stop at water, wounds, fire and navigation. The Ark carries those — and then keeps going, into the material that lets a community not merely survive but teach, trade, build, argue and read.
Wikipedia, complete
Full English and full Spanish editions, with images and diagrams — not the stripped text-only builds. Plus Wiktionary in both languages and bilingual dictionaries.
Field medical reference
The Wikipedia medical subset alongside freely distributable field manuals — Where There Is No Doctor, Where There Is No Dentist — in English and Spanish, plus regional disease references from CDC and WHO material.
iFixit, Restarters & Appropedia
Photographed step-by-step repair manuals for electronics, appliances, vehicles and tools, in English and Spanish — alongside a community repair archive of diagnoses and what usually fails, and appropriate-technology guides for solar, water and sanitation.
Stack Exchange archives
Electronics, chemistry, physics, biology, programming. Question-and-answer form matches the shape of a real problem far better than encyclopedia prose does.
OpenStax & LibreTexts
Actual textbooks — prealgebra through calculus, statistics and physics — with worked examples, graded exercises and answers. Substantial Spanish catalogs taken alongside the English. Wikipedia is a poor mathematics teacher; these are not.
Project Gutenberg
Roughly 75 GB, and the pre-1928 holdings are the quiet treasure: engineering, agriculture, chemistry and metallurgy written in an era when those subjects were taught procedurally, for people working without a supply chain.
Ledgers, contracts, trade
Double-entry bookkeeping — five hundred years old, needs no technology, and any group larger than one household needs it. Contract fundamentals, monetary history, price formation, weights and measures, cooperative and credit-union structures.
Agriculture, climate-indexed
FAO publications organized by climate zone, US Cooperative Extension material across many zones, sustainable-agriculture guides, seed saving, food preservation, veterinary reference — plus Köppen and hardiness maps so an operator can find their zone and navigate the rest.
The entire planet, mapped
Global OpenStreetMap vector tiles as a single self-contained file, servable to any phone on the local network. No tile server on the internet, no region to choose in advance, no relocation that breaks it.
Literature & culture
Project Gutenberg in English and Spanish — poetry, drama, philosophy and the novels. Standard Ebooks and the Biblioteca Cervantes are scheduled and not yet on the drive, which makes Spanish literature the thinnest shelf in the Ark today. Naming that is more useful than implying otherwise.
Handbooks & manuals
Chemistry and materials handbooks, water treatment and sanitation engineering, amateur radio operating and equipment manuals, hydrology, seismic and hazard data, construction norms by climate.
Learning in both directions
Grammar references and language-learning material for Spanish speakers learning English and English speakers learning Spanish — two different corpora, both carried.
You can just talk to it.
The Ark listens and speaks, in English and in Spanish. This is not a convenience feature bolted on at the end. It is the interface that works when your hands are covered in grease, when the person asking does not type well or cannot read, and when the screen is powered down to save the battery.
- 01Press to speak. A physical button, footswitch or key — not a wake word. Always-on listening is a constant power draw, and a false trigger wakes a 27-billion-parameter model. A button costs nothing at rest and works in a noisy workshop where wake-word detection fails.
- 02It hears you. Open-source speech recognition running entirely on the device, accurate across both languages, with a smaller fast model on the backup node for when the main one is off.
- 03It shows you what it heard. The transcription appears for confirmation before the question runs. A misheard question looks exactly like a correct one.
- 04It answers aloud. Natural neural speech, 54 voices across eight language and accent groups, faster than real time — from a text-to-speech model that is under 400 MB and Apache-licensed.
- 05Except when it must not. Any answer containing a dose, voltage, tolerance, pressure, temperature or part number is written to the screen or paper, never delivered by audio alone. Speech recognition errs most on exactly those things, and a spoken answer cannot be re-read. Voice is an input and a convenience layer, not an output channel for anything that can hurt someone.
What this unlocks
A child who cannot yet read can ask why the moon changes shape and be answered. An elder who never learned to type can ask about a medication. A mechanic with both hands inside an engine can ask for a torque figure. A family with no screen on can have a novel read aloud in the evening.
Redundancy, by design
Two independent text-to-speech engines that fail in different ways rather than sharing a weakness, two speech-recognition model sizes, and a third recognition path built directly into the backup node's own language model. Voice survives the loss of any single component.
Honest expectation
A substantive spoken answer takes on the order of ten to twenty seconds end to end. That is a reference tool you talk to. It is not a chat companion, and pretending otherwise is how voice interfaces get abandoned.
It can also keep the
machines running.
The Ark carries a dedicated coding model, and the reason is not software development. It is that an off-grid site is full of small computers nobody has documentation for — and the node that knows how to talk to them is worth more than the node that merely knows about them.
A charge controller, a pump timer, an inverter, a two-way radio, an irrigation valve. Every one of them is a computer, every one speaks some protocol — Modbus registers, AT commands, a serial menu buried on page forty of a manual — and none of them came with the one program you actually need.
That program is always specific to your site. Run the pump when the battery is above 60% and the panels are producing. Nobody sells it. It has to be written where it will run, against the hardware that is actually there.
So the Ark is not only something you ask. It is something that can be wired into the thing it is advising you about — reading a datasheet at one end and producing working control code at the other.
Reads the hardware
Hand it the manual and it will tell you which register reports state of charge, which one reports PV current, and what the undocumented status byte in the reply actually means. This is the work that otherwise requires an engineer who has seen that exact device before.
Writes the glue
Microcontroller sketches, shell scripts, cron logic, config files, the parser for a log nobody documented. Forty lines that make one machine respond to another — the kind of code that is trivial to a specialist and impossible to everyone else.
A different job
General models are measurably worse at this than a model trained for it, which is why the coding tier is carried separately rather than assumed. It costs about 20 GB — nothing against an 8 TB drive — and it is the difference between a reference library and a node that can maintain a site.
And the rule that applies
Generated code that will drive a pump, a relay or a charge controller is read and dry-run before anything is energized. Same doctrine as a dosage: the model proposes, the operator verifies, and where there is consequence the answer goes on paper first.
One cost. Then free, forever.
Every offline-AI product on the market today is a subscription, and a subscription is a license server by another name. The Ark has no account, no key, no telemetry, no per-token fee, and nothing that can be switched off remotely by a company that no longer exists. It also is not one price — it comes in four sizes, and the library is identical in all of them.
| Configuration | On the grid | Off-grid | What it runs |
|---|---|---|---|
| Reader — Raspberry Pi 5 | $580 | $1,080 | Full archive and maps to every phone in range. A small model, slowly. |
| Classroom — Mac Mini 32 GB | $1,740 | $2,840 | The cheapest rung that runs the full 27B model at conversational speed. |
| Community — Mac Mini Pro 64 GB | $3,800 | $5,600 | Named in the spec as fully adequate. Secondary node included. |
| Hub — Mac Studio 128 GB | $5,800 | $8,500 | Two model families resident at once. The only rung that builds indexes. |
| All software and all knowledge | $0 | $0 | Apache 2.0 and MIT throughout; every corpus openly licensed. |
| Every question, forever | $0 | $0 | No network, no metering, nothing to renew. |
Three of the four cost less than ten years of satellite connectivity — before a single AI subscription is bought. The full working is on the What it costs page.
The rule that governs every choice
No proprietary interfaces. No license servers. No component that cannot be substituted with a commodity part. If a piece of this cannot be bought generically or rebuilt from printed instructions, it does not go in.
Why sealed appliances don't count
Several offline knowledge appliances already exist, and they are good. All of them fail the same test: when the device dies, the ability to rebuild the search layer dies with it, even though the archives survive. The Ark keeps the embedding model, the index, the full-precision weights and the conversion toolchain in cold storage — so the whole system can be reconstructed on new hardware years from now, offline, by someone else.
Amortized
One Ark serves a classroom, a clinic, a village hall or a household over the local network at once, and the two-hundredth reader costs nothing. Split across a community and a decade that is between $1.93 and $28 per person — and the drives carry the corpus, so a Reader bought this year becomes a larger node's archive next year without re-downloading anything.
Anywhere the wire doesn't reach.
A patient tutor for every student at once, in two languages, with the textbooks behind it — in buildings that have never had reliable bandwidth.
Field medical reference and drug information at the point of care, where the nearest specialist is a day's travel and the nearest search engine is further.
One node on a solar panel, serving phones and tablets across a hall over local WiFi. Maps, encyclopedia, repair guides, agriculture, and someone to ask.
Homework help that cannot wander onto the open internet, a repair manual for the washing machine, and a novel read aloud at night.
Weeks or months beyond coverage, with the reference load of an entire technical library and none of its weight.
When infrastructure is down and the answers are needed most, the Ark is unaffected — because it was never connected to any of it.
Nothing to block, throttle, filter or revoke. There is no upstream to cut.
And if the worst version of the future arrives: the knowledge to rebuild, in a box that runs on sunlight, with printed instructions for restoring it from cold storage.
What it is not.
A project like this attracts overclaiming. These limits are designed in, documented, and taught to the operator — because a tool you trust incorrectly is more dangerous than one you don't trust at all.
- →It is a competent generalist, not a domain expert. It will state incorrect specifics with complete confidence. That is why every answer shows what the archive says beside what the model says, with the passage one click away, and why the operating doctrine says to verify anything with a safety consequence against the archive or the printed references before acting.
- →It cannot see well enough to diagnose. Image understanding is useful and shallow. It will not reliably identify a plant, a wound or a component at the accuracy a consequential decision requires.
- →It does not learn. Its knowledge is frozen at the last archive refresh. That is a feature — nothing drifts, nothing is silently updated, nothing is retracted from under you — but it is a real boundary.
- →It cannot check what isn't in the archive. Where the library is silent, the model still answers from what it learned, and the screen says plainly that nothing in the archive backs that part. Retrieval quality bounds what can be checked, no matter how fluent the response sounds.
- →Version one serves two languages. Anyone who reads neither English nor Spanish is served only by the model's own memory, with no archive behind it. The search layer was chosen specifically so that adding seven more languages is an addition, not a rebuild.
- →It cannot call for help. Satellite is two-way; no rung of this node is. It will not summon a helicopter, reach a relative or report a collapse. In the first seventy-two hours after a disaster a connection is the more valuable object, and it is not close. This belongs beneath connectivity, not instead of it.
- →Speech degrades in noise. Which describes a workshop, a field, and an emergency. A close-talk headset is required equipment, not an accessory.