the question comes first

Someone asks whether a drug is recalled, where civil legal help exists, what a college costs, or what work an occupation actually contains. The answer matters; an elegant unsourced paragraph is insufficient.

The public-data machine starts with the source agency or registry and compiles a domain into a human page, machine-readable data, a citation path, a stated boundary, and a corrections mechanism. It marks a number as recorded, modeled, derived or interpreted. These are different warrants. A recorded fact can still be incomplete; a modeled score is not an observation; a calculation should expose its inputs.

Nine public-data atlas experiments and a work atlas provide the test bed. The method that builds them is itself recorded. The next research step is to make the family traversable by consequential question and show where the answers disagree or remain unknown.

begin with a question that has consequences

“Is this product recalled?” “Where can I get free legal help?” “What does this college actually cost?” “What work does this occupation involve?” Each question points to public data, but none is answered by dumping a CSV onto a page. The person needs the right record, its date, its jurisdiction or scope, and the consequence of getting it wrong. A machine answering on their behalf needs the same things in a form it can cite.

That suggests a production line. Start with the authoritative record. Preserve the raw source and its update cycle. Normalize entities and identifiers. Define what a question means in that domain. Render an answer for a person and a structured representation for software. Carry the citation and correction route with both. If a source changes, the answer should be traceable back to the version that produced it.

four warrants, four different promises

A recorded figure says a source system stored a value. It does not say the system saw everything. A derived figure comes from a disclosed calculation over inputs. A modeled score is an estimate under assumptions. An interpretation is an argument about what the facts mean. These labels are useful because public sites often flatten them into one confident number.

Take college cost. A published tuition figure is recorded. A net-cost estimate may be derived from aid rules and household inputs. A predicted graduation outcome is modeled. Advice about whether the program is worth it is interpreted. If the interface makes all four look equally observed, it asks the reader to trust the wrong thing. The same pattern appears in health, labor, government services and civic data.

one source can serve two readers

A public atlas should be legible to a tired person on a phone and to an agent looking for a precise field. Those readers need different presentation, but they should not get different facts. The human page can explain the question, the key answer and the caveat in ordinary language. The machine surface can expose stable identifiers, field meanings, update times and source links. Both should resolve to the same underlying record.

The existing atlas experiments show why this is a research program rather than one template repeated across domains. Recall data changes on a different clock from occupation statistics. Service eligibility depends on jurisdiction. A modeled ranking may be useful but must expose its recipe. Each domain needs a compiler that knows what its records can and cannot say.

A good public-data machine eventually becomes a way to ask consequential questions over many domains without pretending they share one ontology. The test is whether a reader can move from an answer to the record and back, notice when the answer is stale, and correct an error without hunting for the person who built the page.

the correction path is part of the answer

Public data has errors, delays and conflicting versions. A person affected by a wrong answer needs more than a footnote saying the source might be imperfect. They need to see the record used, the date it was read, and a route to challenge the interpretation or report a broken link. The system needs to distinguish a source correction from a local parsing bug and from a dispute over what the answer implies.

A machine consumer needs that distinction too. If a field was derived and the input later changes, the machine should know to recompute or retrieve a newer answer. If a modeled score was revised, it should know which model version produced the old one. If an interpretation changed, it should not treat the new wording as if the original public record changed.

That is why the “public-data machine” is more than a publishing system. It is a way to carry the conditions under which an answer deserves to be used. The best version lets someone act quickly while keeping a short path back to the source when the answer matters enough to check.